Skip to content

Guides · Applied AI · Updated October 7, 2026 · 6 min read

13 Failure Modes: What Running an Agent Fleet on Real Code Taught Me

I run coding agents in separate worktrees. These are the failures and review catches that changed how I brief, check and integrate their work.

In this guide

·

Short answer

A coding agent fleet needs more than parallel prompts. I give each branch a bounded brief, require a fresh review verdict and check the integrated output. These 13 failure modes come from my runner notes, independent reviews and a documented site integration.

Key takeaways

  • A process finishing does not prove that its branch is ready.
  • Worktrees isolate files; shared providers, databases and acceptance rules still need coordination.
  • Review can introduce a wrong expectation. Its findings need evidence too.
  • I test the output that reaches a user, including files that should never be public.
Separate folders send change slips to a review clipboard; a magnifying glass reveals a torn slip and a check seal marks a reviewed result.

I use agents to implement, review and translate bounded pieces of work. Each gets its own git worktree. I keep the briefs, reports and review results because they explain a failed run better than the last chat message does.

This is a field notebook across several runs, including my public-site integration on 7 October 2026. The 13 items are my classification of documented failures and review catches. They are not an incident rate, and I am not claiming they all happened in one night. Examples below omit client details.

The handoff I want

The unit of work is a branch with evidence. An implementation report says what changed. A separate review checks the diff and the required gates. Integration then checks the combined application. Those are three different results.

StageWhat I need before moving on
BriefOwned files, permitted actions and a result someone can observe
ImplementationA diff, checks actually run and explicit unresolved items
ReviewA fresh verdict with evidence for each blocker
IntegrationCombined build, routes and rendered output checked again

Starting and keeping a run alive

The first six failures sit around the model rather than in the code it writes. A useful brief cannot compensate for a runner that fails to start, drops its foreground process or loses its provider.

1. Context fills before useful work starts

My runner notes record a large skill catalogue and mirrored skill directories being injected into sessions. Much of the initial context described tools rather than the task. Cleaning the catalogue reduced the starting prompt substantially. I now measure input tokens after changing the configuration. A larger context window does not fix irrelevant input.

2. Provider configuration fails before the agent can work

One endpoint required a session header. Without it, requests were rejected. Another configuration failure came from invalid keys in the orchestration settings. I treat a successful small request as a prerequisite for a batch. A rejected request is a runner problem to diagnose before reviewing the model’s work.

3. Simultaneous starts collide in shared storage

The fleet notes record database locks and initialization hangs when several runner instances started together. Separate worktrees did not isolate the runner’s own database. The recorded workaround was staggered starts. I keep that distinction visible: source isolation and process isolation solve different problems.

4. The foreground process exits while children are still working

An agent could finish its answer to wait for background subagents, ending the process the chain depended on. The runner notes include a continuation procedure for this case. My batch brief now requires the foreground task to remain alive until its bounded result exists. A waiting message is not a handoff.

5. A provider becomes unavailable

The notes record endpoints being unavailable for hours. The worktree and brief survived; the model connection did not. I want a recoverable task with a known next step, so replacing a provider does not mean reconstructing the task from chat history. The replacement still owes the same checks.

6. Parallel reviews hit a shared rate limit

The review provider also rate-limited concurrent work. This is separate from an outage: adding more tasks can make the bottleneck worse. I cap concurrent reviews independently of implementations. A queue with visible pending work is easier to operate than repeated starts whose outcomes I cannot distinguish.

Implementation and review catches

The next six cases concern the meaning of a result: who permitted the work, what evidence supports a claim and which behavior a passing test actually covers. Independent review caught these problems, and also made mistakes of its own.

7. The implementer changes the gate it should pass

The fleet notes record an implementer editing the client-data scanner. That changes what a green result means. I keep gate changes outside an ordinary implementation brief and inspect them separately. When a content slug triggers a rule, the first question is whether the content is allowed, not how to silence the rule.

8. The brief assumes permission it does not have

An architecture review caught production reads written as dependencies in a pilot limited to code, tests and local previews. The revised brief made the missing measurement an explicit prerequisite supplied by the lead. The executor could then complete independent work without treating a convenient verification method as authorization.

9. A citation points to a claim instead of its implementation

The same review found a “GET writes nothing” assertion supported by a limitations string. A string describing intended behavior does not prove that behavior. The revision cited the read path and tests, and narrowed the claim to the service it actually covered. Host-side behavior needed its own qualification.

10. The reviewer supplies the wrong expected result

In the next review round, the reviewer corrected several of its own line references. It also acknowledged steering a check toward an API message that the page did not render. I ask reviewers to follow a result through to its consumer. A confident finding can become a false failure if it checks the wrong boundary.

11. Unknown becomes zero in user-facing copy

A UI review found a fallback that described an unavailable event count as zero beside a recorded event. The evidence row still said the count was unavailable. The repair omitted the count sentence when the count was unknown and added a regression test. This was a meaning error that typechecking could not catch.

12. Green model tests leave the rendered behavior unproven

That review also ran a negative control: removing a visibility listener’s cleanup did not fail the existing tests. The renderer was outside the review scope. The report correctly separated inspected code from tested behavior. I want that distinction in every handoff; a passing suite proves the assertions it contains, not every nearby behavior.

The integration failure a branch review can miss

The final failure came from the combined build. Route filtering was correct, but another path through the module graph made draft content downloadable. I need both checks to protect unpublished work.

13. Draft content reaches public assets without a page

During the 7 October site integration, a shared route helper pulled localized MDX into the client module graph. The content imports were lazy and unused, but the build still emitted content chunks. Hiding a draft from routes had not hidden its bytes from the public build.

The integration moved the localized-path index to build-time code that excluded drafts. Its leak gate searched the whole output for distinctive draft sentences and content-named JavaScript. Negative controls checked that deliberately leaked content was detected. I now ask what files a visitor can download, alongside which pages the router serves.

What changed in my briefs

I use the same small set of requirements at each handoff:

  • Own named files and state which dependencies are shared.
  • Preserve the protective gates; report a blocker when they fail.
  • Keep observed behavior, assumptions and untested behavior separate.
  • Write the review verdict into a fresh artifact for this run.
  • Check the combined result after merging, including forbidden output.

My runner checks for a newly written review file and a verdict. That is a useful guard against a stale report. It does not turn every verdict into approval: the report still has to be read, and unresolved findings still belong to the integration decision.

The 7 October integration combined 15 branches. Its report records a locale-map conflict, the draft leak repair and checks of the resulting sites. A later publication packet tracked the approved subset separately. I keep that separation because “merged,” “checked” and “approved for publication” are different states.

I still make the integration decision. The fleet gives me more candidate work and more evidence to inspect. It also gives me more ways to produce a convincing wrong answer. The useful improvement is a shorter path from a failure to its cause, with enough evidence to decide what to run next.

Discuss your project — free.

Send one paragraph: the systems involved, what breaks or is missing today, and what “done” would look like. If the task is already clear, I can quote the implementation directly. If we first need to establish the cause or scope, we agree a separate diagnostic engagement.