Overnight · 4 min read ·
Run AI coding agents overnight, safely
How to let a coding agent work while you sleep without waking up to a mess: isolation, a check as the finish line, caps, no pushes, and a morning review.
Can you leave an AI agent running overnight?
Yes, if you set it up so that being wrong is cheap. Unattended means nobody answers a permission prompt, nobody notices a loop, and a mistake keeps running for hours.
Anthropic’s guide to building agents is direct about the trade. Autonomous agents mean higher costs and the potential for compounding errors, and Anthropic recommends extensive testing in sandboxed environments along with appropriate guardrails. The five rules below are those guardrails, in order.
Which tasks suit an overnight run?
Bounded ones with a clear check. Anthropic’s overview lists the kind of work Claude Code can take off your list: writing tests for untested code, fixing lint errors across a project, resolving merge conflicts, updating dependencies and writing release notes. Add the bug you never got to, as long as it has a failing test.
Avoid tasks that need a design decision, touch production data or need credentials you would not leave lying around. If you would want to ask a question halfway through, it is a daytime task.
Rule 1: how do you isolate the work?
Never point an overnight run at the folder you work in. Give it its own copy. A git worktree is the lightest option: claude --worktree overnight-fix creates one on its own branch, and git worktree add does the same for any agent, as Git’s worktree documentation explains.
Anthropic’s scheduled tasks docs call out the default. A task runs against whatever state your working directory is in, including uncommitted changes, unless you turn on the worktree option so each run gets its own isolated Git worktree.
Before you start, commit or stash your own work. A clean tree means the morning diff contains only what the agent did.
Rule 2: how do you define done?
As a check the agent can run, not a feeling. Anthropic’s best practices say that a check it can run is the difference between a session you watch and one you walk away from. A test suite, a build, a linter or a script that compares output to an expected file all work.
Both Anthropic and OpenAI document this for long-running work. Claude Code can enforce a check with a Stop hook, which runs your check as a script and blocks the turn from ending until it passes, or with a /goal condition. Codex’s Goal mode asks for three things: the outcome, the constraints and the verification, meaning the tests or measurements that prove it is done.
A prompt that works overnight: “Make the checkout tests pass without changing them. Do not touch files outside src/checkout/. Run npm test after each change. Stop when it passes, or after three failed attempts, and write what you tried to NOTES.md.”
Rule 3: what do you cap?
Turns, spend and permissions.
For scripted runs, claude -p accepts --max-turns, which ends the run with an error at the limit, and --max-budget-usd, which stops at a dollar amount using Claude Code’s own estimate. Use --allowedTools and a permission mode such as dontAsk, which lets the run read files and use the tools you listed, and denies everything else. See our post on the risks of bypass modes before you reach for the bigger switch.
Codex’s codex exec starts in a read-only sandbox, and you opt in to edits with --sandbox workspace-write. OpenAI’s docs say to use full access only in a controlled environment, such as an isolated CI runner or container.
Rule 4: how do you make sure nothing leaves the machine?
Nothing should leave without you. Commit to a local branch, and push in the morning after you have read the diff.
Make that a rule the agent cannot break. In Claude Code, a deny rule such as Bash(git push *) blocks matching pushes in every mode, and Anthropic’s permissions docs use that very rule as an example. Keep production credentials out of the environment the run can see. If the agent does not need the network, switch it off, as Codex’s sandbox does by default.
Rule 5: what do you plan for besides the agent?
Your machine and your plan.
A laptop that sleeps stops the run. Anthropic’s docs say a local scheduled task only fires while the app is open and your computer is awake, and a Codex setting in the ChatGPT desktop app, Prevent sleep while running, exists for the same reason. If you do not want your machine involved, cloud options exist: Anthropic’s routines run in the cloud even when your computer is off, though on a fresh clone without your local files.
Usage limits can end a run early. In an interactive Claude Code session, the tool can wait for a limit to reset and continue, and it still stops at permission prompts. Anthropic’s docs say background sessions and -p runs do not get that wait. Size your overnight task to fit the window you have left.
What should you check in the morning?
Read the evidence in this order. First git log and git diff --stat for the size and shape of the change, then the agent’s notes. Then run the checks yourself rather than trusting “all tests pass.” Then read the diff, with extra care in anything touching money, authentication or data.
If something is wrong, you have a branch to delete, not a mess to untangle.
Overnight Autopilot (Beta) in SwarmPane
Overnight Autopilot is a task queue for exactly this. Each task runs alone in a copy of the project, and counts as ready only if your checks pass. Nothing is pushed without your approval, never to your default branch, and the morning report comes with a replay. It is a Beta feature, so check the product page for its current behavior.
SwarmPane runs the agent CLIs and accounts you already have. Start with a 7-day trial for $1 and queue a first small task tonight.