Task Decomposition Strategies for Parallel Agent Runs
Plan task dependencies before running agents in parallel.

Parallel agent runs fail for a reason that has nothing to do with model quality or infrastructure. They fail because the task was never actually split into independent pieces before anyone hit go. This is a decomposition problem, and the fix lives entirely in the planning that happens before the first line of code gets written.
Why parallel agent runs break down before any agent writes a line
The path to running agents in parallel is usually accidental. Codacy's analysis documents this as the typical origin story: parallelism appears as an improvised workaround for idle time, not as a plan anyone sat down and designed.
That improvisation produces a predictable failure. The paper "Effective Strategies for Asynchronous Software Engineering Agents," which introduces the CAID coordination framework, uses exactly this scenario to illustrate the trap: both tasks complete correctly in isolation, and the combined result doesn't run.
Nothing about this failure is a runtime glitch. The two agents didn't malfunction, and the model didn't misunderstand its prompt. The work wasn't divided into pieces that could be run at the same time without touching each other. Once the task is framed this way, the fix requires doing the division correctly before any agent is handed a task, not resolving conflicts faster after the fact. That shift, from fixing things at runtime to deciding things during preparation, is where the rest of this argument lives.
What makes a unit of work genuinely independent
A task can run safely in parallel with another task only if the two share no mutable state. That is the entire test. It sounds narrow, but it is the single criterion that separates parallelism that costs nothing from parallelism that generates hours of conflict resolution later.
The canonical failure described above comes from mutable shared state. When two agents can each read and write the same file, the same function, or the same interface contract, their work is coupled, even if the two task descriptions sound completely unrelated on paper. A ticket that says "update the billing logic" and a ticket that says "add logging to the payment flow" can look like separate jobs and still collide the moment both agents open the same module.
Shared state hides in more places than the obvious one. And test infrastructure counts too. Agents that each modify shared test helpers or fixtures will interfere with each other even when their actual application code never overlaps.
CAID's research treats this as a modeling requirement, not a hope: a manager process has to build a task plan that is explicitly aware of dependencies, tracking which subtasks sit upstream of others and which can run at the same time without touching each other's output. Codacy's analysis frames the underlying rule simply: work is safe to split when it's independent and touches no shared state, full stop, regardless of whether it happens to also be small or mechanical. Tasks that touch a shared core module, like authentication or billing, carry coordination costs that tend to cancel out whatever time was saved by running agents side by side.
Building a dependency graph before delegating any task
Knowing the independence criterion doesn't automatically tell an engineer how to apply it to a real codebase. The practical move is to build a dependency graph before delegating anything, so that dependencies which would otherwise stay implicit in someone's head get written down and checked.
CAID's method puts a central manager process in charge of this step. Each parallelizable group then gets its own isolated workspace. The graph itself is the artifact that makes concurrency safe to attempt. Without it, "run these agents in parallel" is just a guess.
Reading the graph is straightforward once it exists. Apodex 1.1's training approach for agentic coordination makes the same separation explicit at the model level: the system is trained to decompose a long-horizon task, delegate the pieces that can run in parallel, integrate the results once they come back, and replan from there. Decomposition is treated as its own step, distinct from the execution that follows it.
The graph also isn't something drawn once and left alone. CAID's workflow has the manager merge finished subtasks back into the main line of work and update the delegation plan before handing out the next round of tasks, because the graph that was accurate at the start of a run is often stale by the middle of it. New dependencies appear as work completes, and a static graph drawn at hour zero will miss them.
It's fair to ask whether building this graph before anything starts just adds overhead that eats into the time parallelism was supposed to save. It does add time, but that time is the coordination cost moving to where it's cheapest to pay it. If the graph is skipped, the same cost does not disappear. It reappears later as merge conflicts to debug, failed integration runs to replay, and hours spent figuring out which of two agents broke the other's assumption, at a point in the process where that cost is far more expensive and far harder to trace back to its source.
Using git worktrees to enforce isolation at the file system level
A dependency graph says which tasks can run at the same time. It says nothing about stopping agents from physically stepping on each other's files while they work, and that's a separate problem that needs a separate fix.
This is where git worktrees come in. A worktree gives each agent its own checked-out directory tied to its own branch, so one agent's uncommitted changes are invisible to every other agent running at the same time. Index locks stop contending across sessions, because there's no longer a single shared index for multiple agents to fight over. Conflicts that would otherwise happen silently while agents are mid-task get pushed to merge time instead, where ordinary git tooling can detect them and show them to a human explicitly, rather than letting them hide inside two reports of individual success.
Practitioners have converged on this independently. The worktree is the physical layer the whole strategy depends on to actually hold up.
Worktrees work because they enforce isolation at the level agents actually touch: files and the index. That's the right layer to intervene at. Isolating agents at the network level or the process level doesn't stop two of them from writing incompatible changes to the same branch, because neither of those layers has anything to do with where the actual conflict occurs. For agents running in the cloud rather than on a local machine, the same idea holds at the level of the VM or sandbox: each agent gets its own isolated environment, able to install packages, run services, and modify its own filesystem, and that environment is simply the worktree concept applied one layer up.
Putting the graph and the worktrees together completes the picture: the graph decides which agents are allowed to run concurrently, and the worktree is what makes that decision stick once agents are actually working.
Knowing the upper bound: how task structure sets the ceiling on useful parallelism
More parallel agents is not automatically better. Past a certain point, adding agents makes the result worse, and the point where that happens is set by the structure of the task itself, not by how much compute or patience is available.
Two things set this ceiling, and both come directly from the dependency graph built earlier. One is the intrinsic parallel structure of the task: how many subtasks the graph actually marks as independent of each other. A task with three real points of independence has three genuinely useful slots for parallel agents, not four and not eight. The practical rule is to count the leaf nodes in the dependency graph that share no edges with each other; that count is the ceiling on how many agents are worth running at once for that particular task.
Over-decomposition is a specific, nameable failure mode, not a vague risk. Splitting a task into more pieces than its actual dependency structure supports forces agents into exactly the fine-grained negotiation over shared state that the whole decomposition exercise was built to avoid. The number of agents worth running is a property of the task's structure, a figure to read off the graph.
Writing specifications that are precise enough to keep agents from drifting into each other's scope
A dependency graph can be exactly right, and the task can still fall apart if the instructions handed to each agent leave room for interpretation. Vague task descriptions are a separate source of conflict from a bad graph, and they cause damage even when the graph itself is sound.
The paper "Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork" treats the specification as the actual contract between a human and an agent. A specification has to be precise enough to bound what the agent is allowed to touch, so the agent doesn't quietly decide, on its own initiative, to expand into territory another agent is already working on. The same paper documents what it calls a productivity paradox: individual developers get more done when working with agents, but team-level throughput, review capacity, and overall stability all degrade once the discipline that team-scale engineering requires gets skipped. Vague task delegation is named as a leading cause of that decline.
A specification that actually bounds scope for parallel work needs to name the files or modules the agent may modify as a positive list. It needs to name the interface contracts the agent must leave untouched: function signatures, API shapes, schema versions, anything another agent's work depends on staying stable. And it needs to say, in plain terms, what's out of scope, including anything that might look relevant to the agent but that it should leave alone regardless.
The SDD paper frames this as restoring a kind of accountability that loose, informal delegation erodes: a clear specification is what lets a human later check whether an agent actually stayed inside the boundary it was given. That approach changes the engineer's actual job. Writing the spec well enough that several different but equally acceptable implementations can come back and be fairly compared is the real work, and the quality of that spec decides whether the comparison means anything.
Structured context switching as the human skill that holds parallel runs together
None of the structural work above pays off unless the engineer running these agents has a disciplined rhythm for checking in on them. Agent output can pile up faster than a person can review it, and when that happens, a setup built for speed turns into a backlog.
Codacy's analysis names this shift directly: the actual bottleneck in parallel agent work has moved from how fast an agent can produce code to how much attention and verification capacity a human has left to review it. More output running concurrently does not mean more concurrent scrutiny is available to match it. One person still has one set of eyes, and switching attention between sessions carries a real cost every time it happens.
The discipline that addresses this is structured context switching: treating each parallel effort as a labeled thread with its own branch, its own defined scope, and its own connection to a specific outcome, rather than bouncing between terminals based on whichever one happens to have new output on screen. An engineer working this way moves through a short, deliberate list of active threads, checking the status of each one before moving to the next.
This labeling discipline isn't separate from the decomposition work covered earlier. Simon Willison's own practice, as documented in Codacy's analysis, illustrates the real limit here: even running four agents at once across separate worktrees, he can still only focus on reviewing and landing one significant change at a time. The ceiling on useful parallel work is ultimately set by how much review bandwidth the engineer has, not by how many agent sessions happen to be open.
Integration and merge: where decomposition quality becomes visible
Merge is the moment every decision made earlier in the process gets tested at once. The dependency graph, the worktrees, the specifications, and the engineer's review rhythm all either hold together here or they don't, and there's no step after this one where a bad decomposition can still be papered over.
A clean merge, where independent branches combine without conflicts and every test passes on the first run, is evidence that the graph correctly identified which tasks were truly independent and that the specifications genuinely kept each agent inside its own lane. A messy merge, by contrast, is diagnostic information about exactly where the planning broke down. A conflict between two branches usually traces back to a shared-state dependency the graph missed. A test failure that only appears after integration usually traces back to a specification that didn't name its interface contracts tightly enough. Read this way, the number of conflicts at merge time is a direct measurement of how well the task was decomposed before anyone started, not a random cost of doing parallel work.
Integration is a diagnostic step that reveals exactly where the planning broke down. A team that reviews its merge conflicts for patterns, rather than just resolving them and moving on, is building the same kind of knowledge back into its next dependency graph and its next round of specifications. The quality of a decomposition becomes visible at the moment its separate pieces have to come back together.

