A concrete building under construction, wrapped floor to floor in orange steel scaffolding with zigzag stairways, against a clear blue sky.
LLM

When an autonomous coding agent fails, it is almost never the model

Bob OdellOctober 5, 202614 min read

Here is the short version, for anyone who does not want the details.

I assumed the new hire was underperforming. Then I looked properly and found that the onboarding, the tools and the review process were each broken, and each in a different way.

The new hire is a coding agent running on a model I host myself. It picks up small engineering tasks from a board and works through them on its own. In July I audited 14 of those tasks in a row, and the results looked bad. Only about a third passed the first time. The obvious conclusion was that the model was not good enough.

It was the wrong conclusion. When I traced every failure back to where it started, not one of them was the model reaching the limit of what it could do. Every one started in the machinery around it: a check that asked for the impossible, a check that waved nonsense through, a change that leaked further than it should, or a recovery path that could not recover.

That is the finding this post is about, and it is the one I would most want another builder to hear. If you are judging an agent by how often it succeeds first time, you are mostly measuring your scaffolding.

The rest of this post is the long version.

The audit: a 36% first-pass success rate

This is a dated snapshot. The audit was written on 13 July 2026, and it covered the 14 most recent tasks the local lane had run on its own, across roughly the previous day and a half. Each was a small or medium task. Each was meant to go from the board to merged code with nobody stepping in.

Here are the numbers, without spin.

  • First-pass success: 5 of 14, about 36%. The target was 85%.
  • Done within two rounds: 8 of 14, about 57%. The target was 95%.
  • Median churn: about an hour and five minutes per task. The clean tasks took between 23 and 66 minutes.
  • Five of the 14 needed rescuing, by a cloud model, by an interactive session, or by me.
  • One task froze for 10 hours and 47 minutes.
  • One change caused a brief incident in production, and was reverted.

And the line that made me write this post: zero of the failures were a model-capability ceiling. Every one traced to a harness, gate or routing seam. That was not a one-off. The stress test the day before, which loaded the queue with real tasks overnight, came to the same answer after a forensic sweep of every stall: nine root causes, all in the harness. The one genuine coding mistake the model made that night was caught by the review gate, which is what the review gate is for.

14 tasks, July 2026

OUTCOME
first pass        #####           5
by round two      ###             3
3+ rounds/stuck   ######          6

needed a rescue   #####       5 of 14

WHERE FAILURES STARTED
test gate         dominant seam
change scope      3 tasks named
CI blind spot     1 near miss
recovery path     ~8 hours each
wrong lane        3 of 5 rescues

model ceiling     0

I want to be precise about what "almost never the model" means, because the title makes a strong claim. The agents did make mistakes. One test checked nothing. One operations task was diagnosed as "already fixed" when it was not. But each mistake was either caught by a gate, or slipped through a gate with a hole in it, or happened on a kind of work the lane should never have been given. None of them was a task the model could not do once the machinery around it was right.

Failure mode one: an agent test gate deadlock

The single biggest source of churn was the test gate. In the two weeks before the audit, the same failure pattern had turned up 35 times across 19 tasks, and the count was rising.

Some context on how the gate works. When the agent finishes a change, a separate test-writing agent writes tests for it. Then a reviewer, a stronger cloud model, reads the change and the tests. If the reviewer thinks a test is missing, it says so, and the work goes back for another round. After three failed rounds, the task is held for a human.

The gate failed in two opposite directions at once.

It was too strict. On one task, the reviewer asked for a test that rendered a user-interface component and controlled its timers, the kind of test that needs a simulated browser. But the test runner in that part of the codebase was configured for a plain server environment, with no browser simulation at all. The test the reviewer asked for could not run there.

So the test writer wrote it, the runner failed it, and the reviewer asked again. Three rounds later the task was held. That was the task that froze for nearly eleven hours.

reviewer: "add a test that
renders the component"
          |
          v
test writer writes that test
          |
          v
runner has no browser,
so the test cannot run
          |
          v
round 1 -> round 2 -> round 3
          |
          v
task held for a human

The cruel detail is that the test writer did read the runner's configuration. It just never checked which environment the configuration set up. The reviewer had no idea there was an environment to check. Neither side was wrong on its own terms. Together they made a deadlock that nothing inside the loop could break.

It was also too loose. On another task, the test writer produced a test that only checked a "can this handle the input?" function and never ran the code that did the real work, along with a claim that the rest was "tested elsewhere". It passed. On a third task, the change was judged to have no testable surface at all, so it shipped with zero test coverage. That change was a dialog that was meant to apply to one table and instead applied across the whole board. That was the production incident, and it was reverted.

The root cause is one thing. Two components were judging what counted as testable, and they did not share a definition of what is testable here. One demanded tests that could not run. The other waved through tests that proved nothing. They were the same bug.

There was a third twist. The test writer's own instructions used to recommend a shortcut: read the source file and check that it contains a certain string. The reviewer, rightly, rejected those as tests that only mirror the code. So the harness was telling one agent to do the thing it was telling the other agent to reject.

Failure mode two: scope discipline

Three tasks in the audit did the right thing in the wrong place.

  • A change specified for one table's dialog went into a click handler that every card on the board shared. It changed the whole board. That was the incident above.
  • A task asked to hide a warning. It hid the warning's smaller line of explanatory text and left the warning's header still showing.
  • A task asked to change how often one view refreshed. It changed a global setting instead, one that set the pace of a whole background service and meant four times the load on it, and it left behind a reference to a name that no longer existed.

None of those is a reasoning failure in the usual sense. Each change did what it was meant to do. It just also touched things nobody asked it to touch, and nothing in the loop asked "what else does this touch?" before it merged.

Failure mode three: blind spots in the checks themselves

This one is the most uncomfortable, because every light was green.

A change removed an import, the line that brings a function in from another file, but left a call to that function in place. In most codebases that fails loudly. In this one it passed all six automated checks.

Each check had its own reason not to look.

  • The linter had the rules for undefined and unused names switched off for that part of the codebase.
  • The type checker was set up to read only typed files, and this was a plain JavaScript file.
  • The import checker confirmed that every import that was there pointed somewhere real. It had nothing to say about an import that was missing.
  • The tests never loaded that file at all, because nothing imported it.

Six green checks, and one latent production break. Had it shipped, the call would have failed every two minutes, inside the loop that runs the service's health checks, and that loop would have quietly stopped doing its job.

The model wrote the bug. That part is a model mistake. But a missing import is exactly the kind of mistake a linter exists to catch, and it was a configuration choice, not the model, that turned it off.

Failure mode four: a deadlock in the recovery path

The most expensive failure, per incident, was not in the work at all. It was in the part of the harness meant to recover from failure. Each one cost about eight hours.

When the test gate held a task, the plan was simple: someone pushes a fix, and the harness notices and runs the gate again. In practice, the background job that checks for work skipped every held task, unconditionally, before it looked at anything else. The same job even contained a check that announced it was re-running the test writer when a task's code had moved on. It never took effect, because the skip came first.

A second watcher, the one that notices when a task's code has changed, only looked at tasks in a different state. So a held task waiting for approval was invisible to both. A real fix could land, the automated checks could turn green, and the task would still sit there, frozen, with no way back in.

That is what "stuck in needs approval" meant on my board, for hours at a time.

Prefer harness fixes over prompt fixes

The plan that came out of the audit followed one rule: own the fix in the harness, not in the prompt.

A prompt instruction is a request. A gate is a guarantee. A model will follow a clear instruction most of the time. "Most of the time", multiplied across hundreds of tasks, is a steady stream of exceptions, and each exception costs a round, an hour, or a night. A check in code runs the same way every time.

That does not mean prompts are useless. Here is how the July plan split the work. I am describing the plan as it was written on 13 July; the one piece I cite as landed carries its own date.

Prompts carry facts. The fix for the test deadlock starts by reading the test runner's configuration once, in code, and putting the result into both agents' instructions: this part of the codebase has no browser, so do not write or ask for browser tests. A model cannot follow a rule about an environment it was never told about. That is a prompt fix, and it is the right one, because the problem was missing information.

Gates enforce outcomes. The same fix then adds checks that do not depend on anyone following instructions:

  • One shared verdict on testability. The test writer already decided when a change had nothing to test. The plan makes the reviewer use that same verdict, so the two can no longer disagree. A rule that had lived in the instructions as a hint becomes a check both sides read.
  • Reject impossible tests before they cost a round. A test that uses browser features where there is no browser, or that points at a file path that only existed on one machine, is sent straight back with guidance and does not count toward the three-round limit.
  • Close the loose side. A change that alters real behaviour, like a dialog or a click handler, cannot pass as "nothing to test". It needs a test that exercises the logic, or a check that runs in a real browser.
  • Turn the linter back on. Undefined and unused names, checked on every change in that part of the codebase, and required to pass.
  • Let a fix thaw a held task. Allow exactly one re-run of the test gate when the code has moved on, the new commit came from the implementer, and the checks are green. Anything else stays held.
  • Retry the reviewer when the reviewer breaks. If the review itself fails, for example because the cloud model hit a usage limit, retry later. Do not record it as a request for changes, which spends one of the three rounds on a failure that had nothing to do with the work.
  • Stop sending the wrong work to the lane. Three of the five rescues were operations tasks that needed access to the host machine, which the lane does not have by design. The fix is to route that kind of task elsewhere when it is first sized, not to ask the model to try harder.

One of these pieces is a clean example of both halves working together. The bad advice in the test writer's instructions, the "check the file contains a string" shortcut, was fixed on 21 July 2026 in both places at once. The instructions now forbid it and describe a test that exercises real behaviour instead, and an automated check rejects that kind of test on every change, whoever wrote it. The prompt says what to do. The gate makes sure it happened.

That is the pattern I would keep. Use the prompt to tell the model what is true. Use the gate to make sure the result is right.

A workshop wall of hand tools, each on its own hook or slot: saws, a row of wooden planes, chisels, screwdrivers and pliers.

Why does my AI agent fail? Count the seams, not the model

If you are evaluating coding agents, here is the uncomfortable part. Your first-pass rate is measuring your scaffolding, not the model.

A first-pass rate is the product of everything between the task and the merge: how the task was written, which lane it was sent to, what the agent was told about the codebase, what the checks can see, what the reviewer believes is possible, and whether a stuck task can get unstuck. The model is one term in that product. In my audit it was the term that never set the ceiling.

The same logic applies to any SWE agent evaluation. A score belongs to the model and its scaffolding together. If you only ever change the model, every gap looks like a model gap.

So here is the diagnostic I would give anyone looking at a disappointing number. It takes an afternoon.

  1. Take your last 20 failures. Not a sample of hard ones. The last 20, in order.
  1. For each one, find where it started, not where it ended. "The test failed" is where it ended. "The reviewer asked for a test the runner could not execute" is where it started.
  1. Put each in one bucket: the task was unclear or unachievable; the task went to the wrong lane; the agent was missing a fact about the environment; a check could not see the problem; a check could not be satisfied; the recovery path could not recover; or the model was given everything it needed and still could not do it.
  1. Count the last bucket. That is your capability number. Everything else is scaffolding you can fix, usually in code.

When I did this for the 14 tasks in the audit, the last bucket was empty. Yours may not be. But I would bet it is smaller than your first-pass rate makes it look.

What I would do differently

If I were building the lane again, three things would come first.

  • One definition of "testable", in code, read by everything that judges tests. Two judges with two definitions will deadlock on one side and wave things through on the other.
  • An audit of what every check cannot see. For each part of the codebase, list which checks actually look at its files. A green check that never read the file is not evidence of anything.
  • A recovery path that is tested like a feature. If nothing proves that a held task can be thawed, assume it cannot be.

None of that is about the model. That is the point.

Closing the series

This is the last of eight posts about running coding agents on hardware I own. Read in order, they are one story about the machinery around a model.

It started with an inference server that burned CPU while doing nothing, and an engine bake-off that prefix caching decided. Then came an attention architecture that re-read 162,000 tokens every turn, what actually fits in 64 GB of VRAM, and the quantization formats that fail while looking healthy. After that, a 99,000-token base prompt I had never measured, and how I let the box touch real production code behind a review gate and a kill switch.

Most of those posts were about something around the model rather than the model itself: a server's idle behaviour, a cache, a memory budget, a quantization format, a prompt nobody had measured, a gate. This post is the same lesson at its largest. The routing has changed again since July, as recently as 4 October, and the lane will keep changing. The lesson has held so far: when an autonomous coding agent fails, look at the machinery before you blame the model.

Image credits

Images from Unsplash.

Share

Related Posts