I’ve been experimenting with the idea of running multiple coding agents in parallel, and the tooling around it seems almost as important as the agents themselves.
Once you have 3–5 agents working on different tasks, you need more than separate terminals:
Isolated workspaces or worktrees
Clear task assignment
Shared project context
Agent status and progress
Easy review of each agent’s changes
A way to manage several agents without losing track
Tools like Claude Code, Codex, Cursor, and other coding agents can handle the actual work, but I’m curious about the layer that coordinates them.
I’ve been looking at Sharkly.ai, a free platform that lets you manage multiple AI agents and human teammates in the same project workflow.
For developers here, what are you currently using to run multiple coding agents in parallel?
I’ve ended up treating the coordination layer as a separate system from the coding agents themselves.
The coding agent can be Cursor, Codex, Claude Code, etc. What mattered more once I started running several jobs in parallel was putting a control plane around them.
The pattern I’ve settled on is roughly:
each job gets a bounded task definition and unique ID
every run gets an isolated Git worktree/branch
jobs move through explicit queued/running/verification/checkpoint states
the coordinator tracks which execution slots are occupied and prevents conflicting work from running together
completion produces a specific checkpoint plus verification evidence rather than simply trusting the agent’s “done”
merge/deployment are separate authorization steps
failed, stale, or ambiguous state stops the job rather than letting the agents improvise
a persistent status view keeps recently completed work visible so parallel jobs don’t disappear from operator awareness
One lesson has been that I don’t really want the agents coordinating one another. I want them relatively disposable and independent, with a deterministic layer responsible for assignment, isolation, state, verification, and review.
That has made running several coding agents at once much more manageable than simply opening more terminals. -SS
Strong agreement on the core point, and I’d push it further: the deterministic layer is the product and the agents are interchangeable. Once you accept that, a lot of the design questions answer themselves.
The line I’d underline in your list is “completion produces a specific checkpoint plus verification evidence rather than simply trusting the agent’s done.” That’s the step most setups skip and it’s where the failures actually come from. Agents reporting success on code that doesn’t compile is the single most common thing I’ve had to design around.
One thing I’d add to it: run verification in a process that shares no context with the agent that did the work. If the same model that wrote the code also judges whether it’s correct, it grades its own homework and the failures correlate, so it’ll be confidently wrong in exactly the places it was already confidently wrong. A fresh context given only the task definition and the diff catches noticeably more.
On isolation, worktrees are the right call but they isolate the repo and nothing else. The collisions that have bitten me were all shared state outside it: a dev server port, a database, a cache directory, a lockfile in the home directory. Two agents running tests in separate worktrees still fight over localhost:3000.
Disclosure: I build Grunz, which is a coding agent, so I’m on the agent side of this rather than the orchestration side. https://grunzai.com if it’s useful to you. But honestly the control plane you’ve described is the harder and more interesting half of the problem, and I’d read a longer writeup of it.
We’ve already had to move toward separate execution lanes because once you want several coding jobs running at the same time, separation stops being optional. Worktrees solve one important part of that, but as you point out they don’t isolate the host underneath them.
We’re not at complete runtime isolation either. The approach so far has been to separate the obvious shared state first, then let real collisions tell us where the next boundary actually needs to exist rather than trying to predict every possible one in advance.
That has also changed how I think about the control plane. I increasingly want it following the work while it is happening — not just dispatching jobs, but keeping track of what each job believes it is doing, what resources it is touching, what has actually completed, and where recovery would begin if something disappears underneath it.
The goal isn’t really more orchestration for its own sake. It’s preserving separation and recoverability as more capability gets added.
I’d be interested in how far you’ve taken that in Grunz. Do you treat runtime resources as something an agent/job explicitly owns or declares, or mostly isolate them through the execution environment itself? -SS
The stash stack is my favourite example of your point that worktrees don’t isolate the host underneath. Separate working trees, one shared stash, so a plain git stash pop in one lane can restore another lane’s work. Worktrees give you file separation and nothing else. Anything git keeps per-repo rather than per-tree is still shared, and you find out when two lanes touch it in the same minute.
The class that caught us hardest wasn’t files at all: single-instance external state. One logged-in browser profile, one account on the far end, one OS foreground app. Two lanes each correctly believe they own it, and from outside you look like a single very erratic actor. No amount of filesystem isolation touches that. It is the strongest version of your “separate the obvious shared state first, then let collisions tell you where the boundary is” - the collisions that teach you the most happen outside the repo.
On the control plane following the work: the cheapest thing that actually helped was an ownership check rather than a scheduler. Each external surface has exactly one lane permitted to touch it, and a lane asks before acting instead of reconciling afterwards. Far less machinery than real orchestration, and it catches the failure class you cannot clean up after. You can revert a file. You cannot unsend.
The item I would push hardest on is “what each job believes it is doing”. In long runs that is what degrades silently. When the context window fills, an unpinned goal gets summarised into something vaguer, the job reads back its own notes and quietly restarts a plan it had already finished. It reads as the agent forgetting an instruction, but the instruction was compressed out, not buried. Worth holding the goal somewhere the job cannot rewrite and re-reading it from there rather than from its own history, which is also the only way “where recovery would begin” stays meaningful: a job’s own summary of its progress is exactly the thing that rots.
Your point about keeping the original goal outside the agent made me wonder whether that could tie directly into the ownership control plane. If a resource lease belongs to the durable task rather than to the current model instance, could the control plane read the canonical task definition itself and use that as the scope boundary, rather than relying on the agent to restate it?
That seems like it could address both goal drift and continuity when a worker is replaced. Have you tried anything along those lines? -SS
Honestly, not formally, no. What we have today is closer to a convention than a system: the original ask kept verbatim somewhere the agent never summarizes, and one owner per shared resource, checked before a lane starts. They sit next to each other but nothing connects them, so a replacement worker gets the goal back but has to re-claim the resource on its own.
Tying the lease to the durable task feels right, mostly because of the replacement case you mention. Right now when a worker dies mid-run the lease is either orphaned or quietly inherited, and neither is great. If the control plane scopes the lease from the canonical task, the new worker inherits exactly what the task was allowed to touch, not whatever the last worker happened to grab.
The one thing I’d watch: the canonical task goes stale too. People refine the ask mid-run, and if the lease scope comes from the original text you can end up enforcing a boundary the human already moved. Probably wants an explicit “task amended” event rather than letting either side edit it quietly.