The open side of a desktop workstation: a large graphics card mounted low in the case, a liquid cooler's radiator and three fans across the top, and coolant tubes running between them.
LLM

Putting a self-hosted model on real production work without betting the codebase on it

Bob OdellSeptember 28, 202616 min read

Here is the short version, for anyone who does not want the details.

I changed the engine without rebuilding the car, and I kept the handbrake within reach.

The engine is the language model that writes the code. The car is everything around it: the machinery that hands the agent a task, keeps it in its own sandbox, stops it doing things it should not, and checks its work. The handbrake is a single switch that sends everything back to the old engine on the very next task.

And the part people ask about first: nothing the local model writes goes out unreviewed. Every change it makes gets a full code review from a stronger cloud model. If that review finds a bug or wants something changed, the work goes back to the local model to fix, and then it gets reviewed again. I do not read every change myself. The review is the gate, and it is the reason I can let a small model do real work.

The rest of this post is the long version.

I run coding agents that pick up small engineering tasks from a board, do the work in an isolated copy of the repository, and open a change for review. For most of this series I have written about the box under my desk that serves a model locally: tuning it, making its inference server event-triggered, the engine bake-off that prefix caching decided, the attention architecture that re-read 162,000 tokens every turn, what actually fits in 64 GB of VRAM, the quantization formats that fail while looking healthy, and the 99,000-token base prompt I had never measured.

This post is about the part none of those covered: how you let that box touch a real codebase without betting the codebase on it.

The live routing settings I quote below were read on 28 September 2026. Everything else is dated where it matters, and anything that has since changed is written in the past tense.

Replace the model behind the harness, not the harness

In June 2026 the economics of the cloud plan my agents ran on changed. The planning estimate for moving all of that work onto the metered Anthropic API came out somewhere between $1,500 and $3,000 a month. A local box had to take over the routine coding lane.

The obvious framing of that problem is "which model do I run locally?" That turned out to be the easy half. The hard half was everything wrapped around the model.

My agents run inside a harness that took months of small corrections to get right. Every task gets its own isolated git worktree, so one agent cannot trample another's work. Hooks fire before risky tool calls and can stop them. A parser reads the agent's output stream and pulls out the structured parts. An audit checks the agent's final message against what it actually did, and blocks the reply if it claims an action that never happened. Holding all of that together are a couple of thousand lines of code and more than fifty tests.

None of that discipline cares which model is generating the tokens. All of it was at risk the moment I considered rewriting the agent loop around a new framework to suit a new model.

So the decision that framed everything, logged on 19 May 2026, was one sentence long: replace the model behind the harness, not the harness.

The mechanism is a translation proxy. Claude Code, which my harness drives, speaks Anthropic's API. Local inference servers speak the OpenAI-compatible API. Claude Code Router, an open-source local gateway, sits between them. The harness points Claude Code at the proxy instead of the cloud endpoint, and the proxy translates each request and sends it to whichever destination the routing rules pick. The harness never learns that anything changed.

  task board
      |
      v
  agent harness (unchanged)
  worktrees, hooks, parser, audit
      |
      v
  translation proxy <-- routing settings
      |               |
      v               v
  local lane       cloud lane
  (small tasks)    (larger tasks)
      |               |
      +-------+-------+
              |
              v
  proposed change
              |
              v
  full review by a stronger
  cloud model (sends it back
  to be fixed if it finds
  problems)
              |
              v
  merged

Two properties fall out of that shape, and they are why I chose it.

I could validate against the real workload. No synthetic benchmark, no rewritten agent loop. The same tasks, the same hooks, the same tests, with only the model varying. When something failed, the failure belonged to the model or the proxy, never to a half-ported harness.

A partial failure stays partial. If the local model is bad at some slice of the work, that slice routes elsewhere. The plan does not collapse, and I never have to choose between "all local" and "all cloud". The proxy was either going to be the durable architecture or a cheap experiment, and I did not need to know which in advance.

Routing by task size, and where size breaks down

With a proxy in place, the next question is what to send where. I route by task size.

Every task on my board gets a size estimate: small, medium, large or extra large. Each size has its own routing row, a setting that names the model for that tier. When the plan was written, small and medium work was meant to run locally, larger work was meant to escalate to a stronger cloud model, and the largest tasks, along with anything that edits a document I treat as authoritative, were not automated at all. Those wait for me to do them in my own interactive session.

Size is a usable proxy for three reasons. It is already there, because tasks get sized when they are written. It tracks the things that actually break a small model: how many files a change touches, how long the agent loop runs, and how much transcript piles up in the context window. And it is legible. When I look at a row, I know what "medium" means.

It also breaks down in predictable places.

A size is a guess made before the work starts. A task written up as small can hide a bug that needs reasoning across four subsystems. The size was right about the diff and wrong about the thinking.

Size does not measure context. The binding constraint on my box is the context window, and a small task can still pile up a large transcript if the agent reads a dozen files to find a one-line fix.

The right cut point moves. As of 28 September 2026, only the small tier routes to the local box. Medium and large work both run on cloud models, and both of those rows have changed in the last few days, the large one as recently as this morning. That is not the plan failing. It is exactly what making each row a setting was for: the matrix is a starting guess, and rows move when the evidence says so.

If you copy this pattern, route by size to start, but do not treat the cut point as a decision. Treat it as a dial.

The local LLM code review gate: a stronger model reads every change

This is the part I would keep if I could only keep one.

Every change the local model writes gets a full code review from a stronger cloud model before it goes anywhere. That reviewer has one power that matters: it can send the work back. If it finds a bug, or wants something done differently, the change goes back to the local model with the findings attached. The local model fixes it, and the fix is reviewed again.

I do not read every change myself. That is the point of the design. The small model does the typing, and the stronger model does the judging.

  queued
    |
    v
  picked up by the local agent
    |
    v
  change written, tests run
    |
    v
  full review by a stronger
  cloud model
    |                 |
    | passes          | bugs or changes
    v                 v
  merged and       back to the local
  deployed         model to fix, then
                   reviewed again

Why be this strict with a model whose tests pass? Because the failure I am guarding against is not "the code does not compile". Tests catch that. It is "the code compiles, the tests pass, and it quietly does the wrong thing", and a smaller model is more prone to exactly that. That is the kind of mistake a second, stronger reader is there to catch.

The gate also changes what "good enough" has to mean for the local model. Without it, the bar would be "never wrong", which no model clears. With it, the bar becomes "reliable tool calling that fails gracefully, and a change a stronger reviewer can judge quickly". A small local model can clear that bar, and it is the bar I actually hold the lane to.

The honest cost is that the local lane is not fully off the cloud. Every change it writes also buys a review from a cloud model, and a change that gets sent back buys another one. I pay that on purpose, because the review is what makes the lane safe to run. And if the reviewer keeps sending the same kind of task back, the right response is to route that kind of task elsewhere, not to loosen the review.

An LLM agent kill switch has to be a setting, not a deploy

The whole local lane sits behind one setting. Turn it off and every size tier falls back to the cloud on the next task. No deploy. No code revert. No restart I would have to think about at the wrong hour.

I want to be blunt about the principle: if your kill switch requires a deploy, you do not have a kill switch. You have a rollback procedure, and rollback procedures are what you run when you are calm. A kill switch is what you reach for when you are not.

I know this one works because I have used it for real. On 2 June 2026, the day the lane went live, every local-routed task started failing and my board raised an alert that tasks were stalling in the queue. I turned the setting off, the queue drained through the cloud, and I worked out what had gone wrong with nothing on fire. That story is the first failure below.

The routing rows follow the same rule. When medium and large work moved to cloud models in the last few days, that was a settings change, not a code change. Each row is a dial, and the master switch is the handbrake.

The failures that taught me the most

Three failures, in the order they happened. None of them was the model being bad at code.

A smoke test that passed while every real request failed

Before turning the lane on, I ran an end-to-end smoke test through the proxy to the local server. It passed: a clean round trip and a real completion back.

Then I turned the lane on and every real task failed with an HTTP 400.

The local server had been started with a 32,000-token context window. A real task prompt, with the instructions and every connected tool's schema, was about 95,000 tokens. The smoke test had used a toy prompt that never came near the limit, so it proved the wiring worked and proved nothing about the workload.

The fix was to restart the server at the model's full 256,000-token window, in a single slot, and then send it a deliberately realistic 96,000-token request. It returned 200, where the same size of request had returned 400 before. The next day, on a different model with a different native window, a long task ran off the end again, and I raised that limit to the new model's native maximum. Context limits are something you revisit every time the model changes.

Test with realistic-size input, not a hello-world. A smoke test that uses a toy prompt is testing your plumbing, not your system. Take the prompt from a real task, send that, and check the status code.

A restart that did not reload anything

Adding an API key to the local server meant adding the same key to the proxy's environment. The proxy runs as a background service under macOS's launchd, with its environment defined in the service file. I edited the file, restarted the service, and every request through the proxy started coming back 401 Unauthorized.

Everything about that error says "wrong key". I checked the key. It was right.

The cause was the restart itself. The quick restart command (launchctl kickstart -k) restarts the process but does not re-read the service definition, so the process came back with its old environment and no key at all. The fix was a full unload and reload of the service (launchctl bootout, then launchctl bootstrap), after which the same key worked first time.

A restart is not a reload. When a config change produces an authentication error, confirm the process actually picked up the new value before you start doubting the credential. The cost of this one was not the fix. It was the time spent suspecting the wrong thing.

Sampling settings that made the model loop

During testing, the model I moved the lane to had a habit that would have been fatal in production. At temperature zero, which is greedy decoding, it looped: it repeated itself indefinitely and never produced a final answer. It is a model that reasons out loud before it answers, and with greedy decoding it could talk itself into a circle.

The tempting fix is to change the sampling defaults everywhere. I did not, because the cloud models on the other rows were behaving well, and changing their sampling to suit a local model would have meant changing a working system to fix a different one.

Instead, on 5 June 2026 I attached a sampling override to the local destination in the proxy, and only to that destination: temperature 0.7 and top-p 1.0, the values the model's own documentation gives for agentic work. Then I checked it on the wire. A request sent straight to the server ran at the server's default temperature. The same request through the proxy ran at 0.7. A forced sequence of six tool calls through the lane completed cleanly, with no looping.

Fix it per lane, not globally. Sampling belongs to a model, not to a system, and a system with several models in it needs several sets of sampling settings, each attached to the lane that needs it.

An airliner cockpit seen from behind the seats, its whole front panel covered in rows of round analog dials, each one reporting a single reading.

Observability: tune the routing matrix from data, not vibes

Everything above depends on one habit. Every run records which model served it and how long it took, and the proxy passes token usage through, so the cost of each cloud run can be worked out afterwards. Without that, the routing matrix is a set of opinions. With it, moving a row is a decision with evidence behind it.

The detail that taught me to care about this is embarrassing, which is why it is worth telling.

When I switched the local lane to a different model in June, I kept the old model's name on the new endpoint. It was a minimal-change cutover: the proxy config, the settings and the code all kept referring to the old name, and the server simply answered to it. It worked. It also meant that for a stretch, my run log said one model had done the work while a completely different model had actually done it. I fixed it by serving the model under both names at once, so nothing had to change on a single flag day, and I wrote the mapping down where the next person reading the log would find it.

Record what actually answered, not what you asked for. A label in a run log is a claim. If you rename, alias or swap anything behind a proxy, the label and the reality can drift apart without a sound, and every decision you make from that log inherits the drift.

The other thing observability buys is the ability to tell a model problem from an infrastructure problem. Two of the three failures above looked like model failures at first glance, because a 400 and a 401 do not announce where they came from, and neither had anything to do with the model. Per-run records are how you find that out in minutes rather than days.

Is a local model a real alternative to cloud API costs?

Here is the honest ledger.

What it bought. The routine lane runs on hardware I own, with no usage caps and no per-token meter on the writing. The marginal cost of a small task on the local lane is electricity plus the cloud review of what it wrote. A change in a cloud provider's pricing or limits no longer stops the small work. And I understand my own stack far better than I did in May, which is most of what this series has been.

What it cost. A machine to run, and the previous six posts are the bill: unstable memory, a driver stack, thermal limits, quantization formats that produce garbage, and an inference engine chosen for prefix caching rather than headline speed. Then the operational surface on top of the box: a proxy, a serving container, a service file with its own reload rules, a second set of sampling settings, a model name that can drift from reality, and the monitoring to catch all of it. Then the review, because every local change is read by a cloud model, and some are read more than once. And capability: as of today, the local lane handles small tasks only.

I am deliberately not quoting a savings figure. The planning estimate was a projection made in May, and the lane has changed shape several times since, most recently this week. Any single number would flatter one version of it.

Who should not do this.

  • Anyone who will not put a review gate in front of the merge. A stronger model reviewing every change, with the power to send it back, is what makes a small model safe. Without it you are betting the codebase.
  • Anyone whose kill switch is a deploy. Build the switch first.
  • Anyone whose work is mostly large. If the routine lane is thin, the box spends its life idle while the cloud does the real work.
  • Anyone who does not want to own a GPU box. It needs looking after, and it will need it at inconvenient times.
  • Anyone whose cloud bill is already small. The economics only work when the thing you are replacing is expensive.

The pattern: running a self-hosted LLM coding agent in production

If you want the reusable architecture without the stories, it is this.

  1. Keep the harness. Put a translation proxy in front of it. Change the model, not the discipline around it.
  1. Route by task size, one setting per row. Start with a guess and move the rows as the data comes in.
  1. Put a stronger model's full review in front of every merge, and give it the power to send work back to be fixed.
  1. Make the kill switch a setting, and use it once on purpose so you know it works before you need it.
  1. Smoke-test with a real-size prompt. Not a hello-world.
  1. Reload, do not just restart, after any change to a service's environment.
  1. Set sampling per lane, attached to the model that needs it.
  1. Log what actually served each run, and never trust a label you have aliased.

The model behind the local lane has changed more than once since June. The car is the same car. The handbrake has been pulled for real, and it worked.

Image credits

Images from Unsplash.

Share

Related Posts