Workflow · 4 min read ·
Multi-agent coding workflows with approvals
Design a multi-agent coding workflow that holds up: fixed steps, one job per agent, checks between steps, and approval gates where a person has to decide.
When do you need a workflow instead of one agent?
When the steps are known and you repeat them. Anthropic’s guide to building effective agents draws a useful line. Workflows are systems where models and tools are orchestrated through predefined code paths. Agents are systems where the model directs its own process and tool use. Anthropic’s advice is to find the simplest solution possible and add complexity only when it is needed.
For coding, that means a single agent is right for exploring an unfamiliar codebase or chasing a strange bug. A workflow is right for the job you do every week: take a ticket, plan the change, implement it, check it, get it reviewed. The steps do not change, so do not make an agent rediscover them each time.
What is the basic shape of a coding workflow?
A chain with gates. Anthropic calls the pattern prompt chaining: it decomposes a task into a sequence of steps, where each model call processes the output of the previous one, and you can add programmatic checks, which it calls gates, on any intermediate step to make sure the process is still on track.
In coding the gates write themselves, because you already have programmatic checks. Tests pass or fail, the type checker is clean or it is not, the linter agrees or objects. Put one of those after the step that edits files, and a failed check stops the line before a reviewer ever sees it.
How do you give each step one job?
Write down four things for every step: its role, what it receives, what it must produce and what it may touch.
Pass artifacts, not transcripts. The planning step produces a plan file. The implementing step receives that file and produces a diff. The review step receives the diff and produces a list of findings. Each agent starts clean with exactly the context it needs, which also keeps token use down.
Limit what each step can touch. A planning step should be read-only. Only the implementing step needs permission to edit files. SwarmPane’s workflow docs make the same point: configure edit permission deliberately for every step, and remember that those settings state your intended access, not a guarantee that every external CLI behaves identically.
Use a different agent for the review than the one that wrote the code. Anthropic’s best practices suggest a separate verification step so that the agent doing the work is not the one grading it.
Where should a person approve?
Before anything that is hard to undo or leaves your machine: merging, deploying, running a migration, deleting data, spending money, sending a message to a customer. In Anthropic’s description of agents, they can pause for human feedback at checkpoints or when they hit blockers. That is the right model: a few meaningful gates, placed before the consequence.
Show the evidence at the gate. An approval step that says “continue?” with nothing attached trains you to click. The gate should carry the diff, the test output and the review findings, so the decision takes a minute and is a real one.
Keep the number of gates low. Anthropic’s own docs admit the problem with asking for everything: after the tenth approval, you are clicking through rather than reviewing. Let deterministic checks handle what a script can judge, such as tests and formatting, and spend people on what only people can judge, such as design and risk.
What happens when a step fails?
Decide before you run it. The usual answer is a bounded loop: if the check fails, send the failing output back to the implementing step, at most twice, and then stop and report. Anthropic notes that it is common to include stopping conditions, such as a maximum number of iterations, to maintain control. Without a cap, a stuck loop spends your budget and gets nowhere.
Make failures visible. A workflow that quietly retries hides the evidence you need. Record each attempt and what changed. Also check for partial changes before you retry any step that edits files, because a failed run can leave edits behind.
When is a chain not enough?
Anthropic describes three more patterns. Parallelization runs several calls at once, either on separate subtasks or as a vote. Its example of a vote is reviewing code for vulnerabilities with several different prompts, each of which flags what it finds. Orchestrator-workers lets one model break a task into subtasks and delegate them, which Anthropic says suits coding changes where the number of files to touch depends on the task. Evaluator-optimizer pairs one call that generates with another that critiques, in a loop.
Start with the chain. Add another pattern when you can name the problem it solves, such as reviews that take too long or tasks that split in unpredictable ways.
What does a complete example look like?
Six steps for a feature ticket.
One, plan. A read-only agent reads the ticket and the relevant code and writes plan.md: the files it will touch, the tests it will add and the risks. Two, approval. You read the plan. This is the cheapest place to catch a wrong direction.
Three, implement. An agent with edit permission, working on a branch, follows the plan and runs the check command after each change. Four, gate. A script runs tests, type checks and lint. If it fails, the workflow loops back to step three, twice at most.
Five, review. A different, read-only agent reads the diff and the test output and writes findings. Six, approval. You read the diff and the findings, and merge or send it back.
Build it in SwarmPane
SwarmPane’s Workflow mode lets you connect up to 16 steps on a canvas. Each result moves to the next agent, approval steps wait for you, and every run is saved, so you can see what happened at each step afterwards.
SwarmPane runs the agent CLIs and accounts you already have. Start with a 7-day trial for $1 and turn your weekly routine into a workflow.