All field notes

Comparison · 4 min read ·

Which AI coding agent is best for your code?

Skip the leaderboards. Run two or three coding agents on the same real task in your repo, score the results and let your own tests decide. A simple, fair bake-off.

Why can’t a leaderboard answer this for you?

Because a leaderboard measures someone else’s code. SWE-bench, a public benchmark built from real GitHub issues, is a good idea. According to its paper, it consists of 2,294 problems drawn from issues and pull requests across 12 popular Python repositories. Your repository has a different language, different tests, different conventions and a different size.

Tools also change faster than any article about them. Gemini CLI’s README says it publishes a new stable release every week. A comparison from last quarter may describe software that has already changed.

So the useful question is narrower: which agent does best on the work I actually do? You can answer that in an afternoon.

How do you run a fair bake-off?

Five steps, and every one exists to keep the test fair.

First, pick three to five real tasks from your own history. Mix the kinds: a bug with a reproduction, a small feature, a refactor, a test-writing task and, if you build a UI, a visual change. Each task needs a check that returns pass or fail, such as your test command, the type checker, the linter or the build. If a task has no check, you are judging by feel, and the bake-off is not worth running.

Second, freeze the starting point. Commit your work and note the hash. Give each agent its own worktree at that hash, for example git worktree add ../bake-a -b bake/a <hash>, so no agent sees another’s changes. The git worktree docs cover the commands.

Third, give every agent the same inputs: the same instructions file, the same prompt pasted word for word, and the same permission setup. If one agent can edit freely and another asks at every step, you are comparing settings, not agents.

Fourth, do not nudge. Send the prompt once. If an agent needs a follow-up to get unstuck, write down that it did. That counts against it.

Fifth, run the checks yourself in each worktree. Do not use the agent’s report. Its summary describes what it believes happened, and your terminal tells you what did.

What does a good bake-off prompt look like?

One that any agent can act on without guessing. Name the place, the symptom, the boundary, the check and the evidence:

“Fix the bug in src/cart/total.ts: a coupon over 100% produces a negative total. Write a failing test first, then fix it. Do not change other files. You are done when npm test passes. Paste the command output.”

That prompt works for every tool because it does not rely on any tool’s habits. It also makes the result easy to score: the check passed or it did not, and the diff either stayed inside one file plus its test or it did not.

What should you score?

Keep the scorecard short, so you will actually fill it in.

Did it pass your checks? Yes or no, from your own run.

How big and how clean is the diff? git diff --stat shows the size, and unrelated files in the list are a mark against it.

Would you merge it as it is, with edits or not at all? This is the score that matters most, because it measures review effort, which is where your time goes.

How many times did you have to step in? Count every follow-up and correction.

What did it use? Claude Code shows plan usage with /usage, and OpenAI’s docs point Codex users to a usage dashboard and /status. Compare like with like, and remember that a subscription limit and a pay-as-you-go bill are different things.

Add the scores up however you like. A sensible default is correctness first, reviewability second and cost third. Decide your weights before you see the results, so the results cannot pick them for you.

What ruins a bake-off?

Different starting commits. Different instruction files. Letting one agent see a hint that the others did not get. Running agents in parallel against one dev server port or one local database, so they interfere with each other. Using only tasks that you already know one tool handles well. And drawing a conclusion from a single run.

That last one matters. The same prompt can produce a different diff on a second run, so for any task where the result will decide something, run it twice and treat a difference between runs as information.

What do you do with the result?

Decide per kind of task, not once for all. You might find that one agent gives you the cleanest refactors and another gives you the most reviewable bug fixes. Write that down in your project’s instructions or your own notes.

Then repeat the bake-off when something big changes: a new model, a new version of a tool, a different repository. An hour of testing beats weeks of arguing about leaderboards. For what each tool’s own documentation says today, see our Claude Code vs Codex vs Gemini CLI comparison.

Let Race run the bake-off

SwarmPane’s Agent Race is this test built into the app. Two or three agents take the same task, each in its own copy of the project. Your tests, typecheck and lint rank them, and nothing merges without your click. Before a race starts, SwarmPane tells you it costs about one run per agent, so the cost is visible up front.

Broadcast is the lighter version. It sends one prompt to several panes at once and tells you how many took it.

SwarmPane runs the agent CLIs and accounts you already have. Start with a 7-day trial for $1 and race your own agents on your own code.