At 11:57 on September 16, the shared five-hour usage window on the provider account sat at 18%. By 15:57 it was at 97%, with up to five coding agents overlapping across two parallel tracks. I paused two of them myself and set a timer to resume them. When the coordinating agent offered to speed things up, I wrote back: "i don't ask for faster, sequential is fine."
The work was for a team I work with: moving its agent runs off a separate AWS ECS task per run and onto an ECS service that consumes a queue. The runs belong to pi, a minimal open-source coding-agent harness. The parallel phase landed five PRs in about 19.5 hours, with two conflicts on one PR, repeat approvals and overnight waits. The sequential finish the next morning landed five PRs and the production service in 7 hours 40 minutes, with no conflicts inside the epic. It still had an auth failure, a disconnect, two merges ahead of their gates, and a wrong design premise that shipped and needed a follow-up. Two different packages of work, the second built on the first. Not a race.
Parallelism doesn't make agents slower by itself. It exposes whatever part of delivery was already the bottleneck: a gate, a shared resource, an integration step. In my records, more parallel coding agents were not faster, by one way of counting they were costlier, and they were more fragile. And in a controlled study I ran with coding agents as the workers, on the public pypa/packaging repository (4 change bundles, 2 orders, 3 repetitions), the same parallel workers' output went from 9 to 17 of 24 accepted when only the integration policy changed. Below are the four places I've watched parallel work get lost, and a way to decide which constraint to change before you add the next agent.
What the Pause Actually Cost
Pausing two agents sounds harmless. In my setup then, it wasn't. A paused agent's half-done work sat inside its container, uncommitted. Across that session, four coding agents also died when their auth tokens were revoked, and not all of that belongs to the quota or the pause. In those days a dead session meant copying its container's data into a replacement container by hand. My summary at the time: "semi manual work, far from autonomy level needed for parallel work."
Why pause instead of letting them run? My bet was that with fewer agents running, one would finish and commit before the limit hit, and then I'd switch to another provider and keep going. With five going, none had anything I could commit: every change sat half-done inside a container. A quota or an expired token turned each of those containers into recovery work.
That resume timer is where I started designing automatic recovery into Maestro, the orchestrator I run my agents in. Part of it works today, part is still planned. Provider switching on resume came later, driven by incidents of this kind and not by parallel runs alone. Parallelism just raises the odds of burning the remaining usage too fast.
The standard advice on parallel coding agents is a worktree each, separate files, three to five agents, and Anthropic's Claude Code docs add that returns diminish. True, and silent on half-done agents, the machine, colliding changes and acceptance. Those are the four places I've watched parallel work get lost: at the seams between tasks, in the resources every agent shares, in the start-up cost every extra agent pays, and at the human gates that turn delivery into a queue.
Work Gets Lost at the Seams
Isolation doesn't remove conflicts. It moves them to the merge.
In late September, a feature for that team unifying upload routing and submission validation ran on a feature branch, two tasks at a time, each with its own files. One PR still conflicted the moment its sibling merged. Then I had the whole branch reviewed again, and it turned up five defects between tasks that every PR's own review had passed. One was serious: a routing override that passed validation when you saved it and then got rejected on every submit. It went into a cumulative fix on the branch.
In that second review I look for the seams where PRs meet and guess where merge order mattered: import order, the same function implemented twice with slightly different behavior, the same column added to a model twice. It's a habit, not luck. For bigger work, the coordinating agent collects the changes on a feature branch, each PR gets its own review on the way in, and the whole branch gets one more.
Bugs invisible in the increments aren't rare, and human teams get them too: two tasks start out independent and later both change a shared resource. That branch later spent days in conflict with main too, but that came from a wait for manual QA, not parallel work.
Shopify described the same thing in 2019 as "soft conflicts": "two pull requests that pass CI independently, but fail when merged together." Their merge queue ran CI on a predictive branch, the future combined state, instead of on each PR alone.
Mechanical merging loses work in its own way. In my study, the naive parallel policy cherry-picked each worker's commit and threw out the whole commit on any conflict. It omitted 21 commits across 14 blocks.
The conflicts behind that were file events in a changelog (11), docs (8) and one test file (4), and none in source files. Clean source changes dropped out of the result because a changelog line collided, while the worker commits themselves still existed and were simply left out.
So whose fault is a collision nobody predicted? The planner's. Two parallel tasks touching one contract or one file is a risk whoever planned the task graph should have seen, and then either steered around or run anyway with a plan to resolve it. Not every pair collides: in September, one of that team's tasks removing a tool from an always-granted set and one exposing grantable tools through a new endpoint composed cleanly.
Agents Compete for More Than Session Slots
The resource graph is not the task graph.
In late July, a batch of platform work ran four coding agents and four review agents at once on one machine, several of them running test suites of roughly 13,000 tests. Within 11 minutes, eight sessions stopped reporting and were declared dead. Nothing showed in the agents' output first. The machine went laggy, sessions took longer to dispatch, and my task board and orchestrator slowed to a crawl. The signals were the boring low-level ones: CPU, RAM, network bandwidth, disk.
A cap of about four concurrent sessions, reviewers included, was recommended for that batch because of machine starvation. My reason for wanting a small group was different: I wanted the coordinating agent to pause all of it quickly when the usage window got close. With 16 containers, pausing everything takes a while. Four was never a standing limit.
Then there's the test runner. One unconstrained pytest -n auto from inside one container can try to use the machine's available cores, and every agent session makes that choice on its own. Nothing in pytest-xdist allocates a budget across a fleet. One agent's test run can take every core on the machine, which is why per-container CPU and RAM caps are on my list.
What agents compete for is a longer list than session slots: CPU, RAM, the provider's usage window, network bandwidth, disk I/O, and the coordinating agent's attention. A count of admitted sessions budgets none of them, and I haven't solved it: this month, the harness I built to measure exactly this over-parallelized and starved my own machine.
Cursor's February write-up on running hundreds of agents found that once RAM was constrained, compilation and disk I/O dominated: "Hundreds of agents compiling simultaneously would result in many GB/s reads and writes of build artifacts." Anthropic's C compiler project hit a different wall with 16 agents: "Every agent would hit the same bug, fix that bug, and then overwrite each other's changes." The agents already running only became useful once failures could be isolated, by compiling most files with GCC and a random subset with the new compiler.
Every Extra Agent Pays the Start-Up Cost Again
The optimistic model of parallel coding agents: slice one agent's work into N pieces, spend the same tokens, cut the wall-clock time by N. It forgets the start-up. Every fresh agent loads the project instructions, reads the docs, finds the source files and fetches remote resources. One agent does that once. Five agents do it five times.
Task size makes this a trade-off. A small task is cheap to redo when it fails partway, and an hour of lost work stings. But every piece pays its own start-up cost, so splitting finer hits diminishing returns. I don't have a universal answer for where the two costs cross. Treat task size as a knob, experiment on your own tasks, and watch both costs as you turn it.
Caching softens this, under conditions. On the Claude API, separate agents share cache hits only when they run in the same workspace with caching enabled and their prompts start with an identical prefix: the same tools, system prompt and project instructions. A cache entry appears only after the first response begins, so agents launched together into an empty cache all start cold. Past the shared prefix, each agent reads files in its own order, and that first read is billed as uncached input or a cache write. Its own later turns can then reuse that history from cache.
So the bill has three input categories: uncached input, cache writes and cache reads. As of October 4, 2026, a cache read costs a tenth of base input on most Claude models and a twentieth on Claude Opus 5.5, and a five-minute cache write costs 1.25 times base. That's a ratio for reused input, not a multiplier on an agent's whole bill. Reuse is where most input goes: across my study, about 165 million of 176 million input tokens were cache reads.
My own records: across 117 batches of related tasks with that team (81 run by a single coding agent, 31 with two or three at peak, five with four or more), median lead time from ready to done was about 24-25 hours in every bucket. No general slowdown, and no speedup either.
Cost per finished task depends on how you count. The batch-level median was about $18, $25 and $27 across the three buckets, while the task-level medians were $17, $16 and $25. All of these are estimates from incomplete records, with costs missing most often for single-worker batches. Task size, code surface and model mix confound them, the top bucket holds five batches, and they don't show that start-up cost caused anything.
One idea I want to test: a cache shaped like the task graph, where parallel agents branch from one warmed parent session instead of each starting cold. pi already stores sessions as trees.
Four Ways to Merge the Same Work
When integration is where your work is being lost, change the integration policy before adding workers, and test it. The only controlled comparison I have is my study: four bundles of three related tasks, two orders, three repetitions, so 24 matched blocks per policy. Its limits: one repository, three clusters of overlapping tasks, one model alias. "Accepted" means frozen behavior tests passed and the source compiled, not that it was ready to release.
| Policy | Accepted (of 24) | Model calls per block | On conflict | Caveat |
|---|---|---|---|---|
| Sequential | 16 | 3 | No merge step: three fresh workers in turn | Not one persistent agent |
| Naive parallel | 9 | 3 | Aborts any conflicting commit | Omitted 21 commits in 14 blocks |
| Mechanical queue | 13 | Reuses the parallel calls | Replays parallel commits, unions documentation conflicts | Not a production queue. A filter bug hit 6 artifacts (3 accepted), the corrected replay kept the source results |
| Parallel + integrator | 17 | 4 | A fresh agent reconciles the parallel workers' patches | One extra call per block |
Against sequential, the integrator won two matched blocks and lost one, which settles neither "you need an integrator" nor "it never helps." What the table does show: on the same parallel workers' output, the policy alone moved acceptance from 9 to 17. Don't adopt the study's merger as production code: its documentation filter is exactly the kind of bug a real queue can't have.
How I'd choose: sequential is the baseline when tasks can't be separated reliably. A queue that checks the combined state fits when independent work fails on composition. An integrating agent is worth a trial when reconciliation needs reasoning across requirements, as long as its extra calls and waiting go into the same measurement.
Cursor went the other way at a much larger scale. In its February report, the integrator role "quickly became an obvious bottleneck. There were hundreds of workers and one gate (i.e. 'red tape') that all work must pass through." They removed it and tolerated intermediate commits that were briefly wrong, and they suggest a periodically repaired green branch would be needed before release. That's a research prototype on one branch, and its earlier lock contention was a separate problem.
The queue has knobs too. Gas Town, an open-source multi-agent system, bisects a failing batch to find the culprit. GitHub's merge queue needs the merge_group event to trigger Actions workflows and has its own concurrent-build limit. Coding concurrency, CI capacity and merge batch size are separate settings. Add agents while CI backs up and you've only moved the waiting.
My guess is no dedicated integrator: the coordinating agent decides the merge order.
When the Combined Suite Is Red, Who Decides?
On September 25, in Maestro's own repository, two PRs went in, each with its own checks passed. One changed a contract: a failed turn on an agent whose connection was still live would now report "waiting for feedback" instead of "failed". The other PR's tests still expected "failed", and the combined CI went red at 19:33. I wasn't watching. Whole batches were landing as trains of stacked PRs, and catching this is the coordinating agent's job.
It did. It diagnosed a stale test and handed it to a coding agent, which fixed the test with no runtime change, and 2,213 tests passed. The dedicated integration agent launched at 19:36, after the failure was already recorded, so it wasn't the one that caught it. My part came about seven hours later: the go for the deploy.
Who decides whether the code or the test is wrong? My rule: whoever owns the artifact encoding the old assumption fixes it, here the test's author. The contract's owner reports the collision and does nothing else. The coordinating agent breaks ties.
The contract itself gets settled earlier. I define contracts or sign off on proposed ones during grooming. A task can carry explicit permission to change one, and some tasks are literally "find the right shape for this API".
The direction I want is for the two agents to settle it themselves: both see main go red, find the stale test and agree who pushes the fix. That's not what happened that night. Cognition argued in June 2025 that "agents today are not quite able to engage in this style of long-context proactive discourse with much more reliability than you would get with a single agent." A fair read of 2025 agents, and the wrong bet on where they're heading.
Put the Human at the Plan, Not the Merge
Back to the upload branch. Its PR waited days for manual QA while main kept moving, until a conflict blocked the deploy to the QA environment. That wasn't parallel implementation. It was an acceptance queue, and the conflict grew out of the wait.
In my own projects, my attention goes to grooming: I use agents to talk the problem through, gather context and draft the plan, and only a plan I'm happy with goes to the coordinating agent. For small tasks I trust, the coordinating agent grooms too and merges once the required gates are green.
Here's my strong opinion: if the goal is shipping fast, a human shouldn't sit at the merge gate. Good CI gates cover most of the risk and let agents iterate on their PRs in parallel. A human is serial and scarce, review quality drops after a few PRs, and rubber-stamping follows. Then the whole PR rate falls to one person's throughput.
It comes with conditions. The contract was agreed at grooming, the combined candidate runs the required checks rather than each PR alone, those checks actually block the merge, and someone owns a red result. Where a change's required behavior can't be checked independently yet, or the contract isn't settled, I'd keep a human at acceptance and cap incoming work until the checks exist.
The best evidence for the drift is a 2026 study of one company's 2x mandate, 802 developers. Authored PRs per developer reached 2.09x, per-reviewer load roughly doubled, and PRs with a substantive human review (a human-written comment) fell from about 39% to 21%. The median reviewer's silent approvals roughly doubled, and automated review took over most coverage. What it doesn't show: a silent approval isn't proof nobody looked, steady merge and revert rates aren't a clean bill of health, and the study never removed the human gate. DORA's 2025 report also ties AI adoption to lower stability.
Simon Willison keeps the human in review: "the natural bottleneck on all of this is how fast I can review the results." That team keeps a human on merge too. I disagree, and expect that policy to go away. My own review gate can technically be bypassed today: agents wait for review because their instructions say so.
In my Claude Code quality workflow I put the emphasis on reviewing AI code myself. What changed is where my attention goes, not whether review happens. No human at the merge never means no review.
Three Rules Before You Double
Audit your gates first. Raise the bar on CI tests, security checks, validators, dependency and license scanners, and make feedback fast. If more agents raise the volume of changes, weak gates let it all through. As I put it: "they will happily ship more shit faster", and cleaning that up becomes debt that grows fast. The gates have to be able to fail, which is the point of when AI agents declare victory too early, and the AI slop checklist covers what they should catch.
Measure per accepted task. Tokens and dollars per accepted task, with cache reads, cache writes and uncached input kept separate. Then three clocks: task latency (ready to accepted, which exposes queues), batch elapsed time (first task started to last result accepted) and active human effort (minutes a person spent, not minutes spent waiting for one). Overlapping task durations don't add up to either of the last two. METR found in February 2026 that its "measurements of time-spent on each task are unreliable for the fraction of developers who use multiple AI agents concurrently": that's the human-effort clock, which is why it stays separate.
I use these numbers to analyze past sessions. What makes me stop a session live is a provider nearing its limit with work that won't make it.
Grow in steps, never in bursts. A burst launch can eat RAM, disks and databases faster than you can react. Add one agent, let the metrics settle, then decide.
Where I'd look before adding a worker:
| Symptom | What to measure | What to change before adding a worker |
|---|---|---|
| Agents waiting on a shared resource | CPU, RAM, disk waits, provider throttling per session | Budget tests and builds per session, stage launches |
| The same context read again and again | Uncached input, cache writes, cache reads and output per task | Tune task size, share a stable prompt prefix |
| Rework at integration | Omitted commits, combined-branch failures, whole-branch findings | Change the integration policy, plan overlaps out |
| Ready work waiting for acceptance | Task latency split into waiting and gate time | Automate or restaff the gate, move human attention to grooming |
A protocol I'd suggest: define acceptance before both runs. Change only the worker count, and keep task type and size, model, gates and integration policy the same. Total tokens and dollars across workers, retries, reviews and integration, then divide by accepted tasks. Record batch elapsed time and active human minutes separately.
Then read the result plainly. Zero accepted tasks means the run failed. A mixed result means you repeat comparable work before you add a worker.
Pick one batch of related tasks this week. Record tokens and dollars per accepted task, task latencies, batch elapsed time and active human minutes at today's concurrency. Add one worker, run a comparable batch, and compare before you add the next.
