All field notes

Review · 4 min read ·

How to review AI-generated code fast

A ten-minute routine for reviewing AI-written code: shrink the diff, read the tests first, run the checks yourself, and read the risky paths line by line.

Why does AI-written code need its own review routine?

Because you were not there when it was written, and there is no author to ask. An agent can produce a large diff in the time it takes you to make coffee. The risk is not that the code is always bad. It is that the diff is big, the summary is confident, and approving is easier than reading.

The routine below takes about ten minutes for a typical change. It is built from Google’s published code review guidelines, Git’s own tools and Anthropic’s and OpenAI’s advice for their agents.

Step 1: how do you shrink the diff before you read it?

Small changes get better reviews. Google’s guide says small changes are reviewed more quickly and more thoroughly, are less likely to introduce bugs and are simpler to roll back. It defines the right size as one self-contained change that makes a minimal change addressing just one thing, with its tests. Ask the agent for one thing at a time, and the diff stays reviewable.

Then measure what you got. git diff --stat lists every changed file with a count of lines. If files appear that the task never needed, that is your first review comment.

Two flags make a diff easier to read. git diff -w ignores whitespace changes, which hides formatting churn. git diff --color-moved colors moved lines differently from added and removed ones, so a function that only changed place stops looking like new code.

If the diff mixes a good change with an unrelated one, split it. git add -p lets you choose hunks interactively, and git’s docs describe it as a chance to review the difference before adding it.

Step 2: why read the tests first?

Tests tell you what the agent thinks the code should do. Read them before the implementation, and ask three questions. Do they describe the behavior you asked for? Will they fail if the code breaks? Did any existing test get weaker, skipped or deleted?

Google’s guide is blunt about the second question: tests do not test themselves, and a human must ensure they are valid. If you want a deeper treatment of the third, see how to stop agents editing your tests.

Step 3: should you trust “all tests pass”?

Check it. Run the test suite, the type checker, the linter and the build yourself, or let your CI run them, and read the output instead of the agent’s summary.

Anthropic’s best practices page says the same thing from the other side. It tells you to have Claude show evidence rather than assert success: the test output, the command it ran and what it returned, or a screenshot. Evidence is faster to check than a claim, and it still needs you to look at it.

For changes you can see, such as a layout, run the app and look. A passing test suite says nothing about whether the button is where it should be.

Step 4: which parts do you read line by line?

Read every line, and spend your attention where the damage would be worst. Google’s guidance is to look at every line you are assigned to review, and not to skim a human-written class, function or block and assume it is fine. The same applies when the author is an agent.

Decide the risky paths before you start, because that is where the slow, careful reading goes. A good default list: authentication and permissions, anything that handles money, data deletion, database migrations, configuration and secrets, and new dependencies.

For each new dependency, ask whether you need it, whether the package name is exactly the one you meant, and whether it is maintained. A package nobody asked for is a review comment.

Use Google’s lenses for the rest of the diff: design, functionality, complexity, naming, comments and consistency with the code around it. Be especially vigilant about over-engineering, which the guide describes as code made more generic than it needs to be. Solve the problem you have today.

Step 5: how do you get a second opinion?

Ask a different agent to review the diff, with instructions to find problems rather than fix them. Anthropic’s docs suggest a separate verification step so that the agent doing the work is not the one grading it. Codex’s CLI has a dedicated review mode that checks uncommitted changes, a commit or a base branch and reports prioritized findings without modifying your working tree, as its CLI docs describe.

A second opinion adds a pair of eyes. It does not replace yours. Read its findings, and decide which ones are real.

When should you send it back?

Send the diff back, instead of fixing it yourself, when it is bigger than the task, when the design is wrong, when tests got weaker, or when you cannot explain a change.

Be specific: “Revert the changes in billing/ and config/. Keep only what the login fix needs. Run npm test and paste the output.” Fix small things yourself. A one-line naming change is faster to make than to describe.

Review where the agents run

SwarmPane’s Files, Git and Find tools sit beside the terminals. You can read and edit files, review the diff, stage and commit, and search the whole project without leaving the window.

When you want two agents on the same task, Agent Race has each work in its own copy of the project, ranks them by your own checks, and merges nothing without your click.

SwarmPane runs the CLIs and accounts you already have. Start with a 7-day trial for $1 and review your next change next to the agent that wrote it.