Voice · 4 min read ·
Voice coding: dictate to your AI agents
Four ways to talk to a coding agent instead of typing, where each one sends your audio, and a safe habit: dictate into a draft, read it, then send.
Does voice coding work with AI agents?
For prompts, yes. A coding agent takes plain language, so the part of the job that is prose, such as describing a bug, explaining what a screen should do or leaving review feedback, is faster to say than to type. Speech is a poor fit for the other half: file names, flags and symbols are exactly what recognition mishears.
So think of it as prompting by voice, not programming by voice. You dictate the intent, you check the text, and the agent writes the code. The rest of this post covers the options and one habit that makes them safe.
When does voice beat typing?
When the thought is bigger than the sentence you want to type. Explaining a bug while you look at the screen. Leaving feedback on a diff you just read. Describing how a page should behave. Dumping everything you know about a feature before you ask the agent to plan it. These are natural speech tasks, and you can say a paragraph in the time it takes to type a line.
Typing wins for anything exact: a command, a path, a regular expression, a version number. A good habit is to speak the intent and type the exact bits.
What are your options for dictating to an agent?
Four, as of October 2026.
Built into your agent. Claude Code has /voice. Per Anthropic’s voice dictation docs, you hold a key to record or tap once to start and again to send, and your speech is transcribed live into the prompt. It needs a claude.ai account and a local microphone, and it streams your audio to Anthropic’s servers for transcription rather than processing it locally. The docs say transcription does not consume Claude messages or tokens, and that it is tuned for coding vocabulary. OpenAI’s ChatGPT desktop app has Dictate in the composer, which transcribes into the composer so you can review and edit before you send (prompting guide), and a separate ChatGPT Voice feature for talking through Codex tasks (voice docs).
Your operating system. macOS Dictation lets you speak to enter text anywhere you can type it, and Apple’s support page says you can check Keyboard settings to see whether it is processed on your device. Windows voice typing starts with the Windows key plus H, and Microsoft’s support page says it uses online speech recognition powered by Azure Speech services, so it needs an internet connection.
A speech model you run yourself. OpenAI’s Whisper is a general-purpose speech recognition model that handles multilingual speech. whisper.cpp runs it in plain C and C++ and supports CPU-only inference, so audio never leaves your computer. The price is setup: you wire the recording and the paste yourself.
SwarmPane Voice, described below.
Where should the words land?
In a draft, never in a command. The risk with dictation is not a typo in a sentence. It is a misheard command that runs.
Look at what your tool does after you stop speaking. Claude Code’s default is to insert the transcript and wait for you to press Enter. Its tap mode, per the same docs, submits automatically when the transcript is at least three words, so use hold mode and keep auto-submit off if you want to read first. The ChatGPT app’s Dictate puts text in the composer for review.
In a terminal, an agent’s prompt is a line of input, and a recognizer that presses Return for you will submit whatever it heard. Choose a tool that types the words and stops.
How do you dictate a good prompt?
Speak the frame, not a stream of consciousness: what should change, what must not change, how to check it, and what to show you. “Fix the settings page so the Save button stays on one line on narrow screens. Do not touch other pages. Run the linter and show me the output.”
Say identifiers only when you must, and read them back. Better, paste file names or use your agent’s file mention feature, and speak around them. Be careful with spoken formatting. Apple’s docs say the phrase “new line” is equivalent to pressing the Return key once, which in some prompt boxes sends the message.
Dictate one idea at a time. A rambling two-minute recording produces a rambling prompt. Read the draft before you send it, as you would an email to someone who will act on every word.
Where does your audio go?
Know before you speak about code you cannot share. Claude Code’s voice streams audio to Anthropic. Windows voice typing uses Microsoft’s online recognition. macOS may process on the device, and its settings tell you. A model you run yourself keeps audio on your machine. If your employer restricts where code can be discussed, check the policy before you pick a tool.
SwarmPane Voice
SwarmPane Voice comes with Pro and is macOS only for now. Press ⌘⇧Space and talk: the words land as a draft in your agent’s composer or terminal, and nothing reaches the agent until you send it. In a terminal the words are typed on the input line and never run. It works with whichever agent is in the pane, not one vendor’s.
Cloud voice, the default, transcribes with Whisper large v3 turbo and drops the audio; Pro includes 1,500 minutes a month, about 25 hours. Whisper on your Mac works offline and uses none of those minutes. Starter has no voice.
See SwarmPane Voice for the details, or the voice page to try a simulated demo in your browser. SwarmPane runs the agent CLIs and accounts you already have. Start with a 7-day trial for $1 and choose Pro to try voice, with 300 cloud minutes during the trial.