Voice coding for Claude Code — talk to your agent while it works
6.8K
Talk to Claude Code while it works — Jarvis-style voice coding. It's your voice that's noisy, not your code.
Claude speaks short summaries aloud. An always-on listener turns your speech into messages Claude receives mid-task, without stopping it — no push-to-send, no copy-pasting transcripts. Step away from the keyboard and keep steering your agent.
Speech-to-text and text-to-speech run on the Grok (xAI) Voice API — extremely cheap in practice (a small one-time budget lasts months of daily use).
The backend ships as a hardware-free Docker image
(noisy/noisy-coding):
the dashboard browser tab is the microphone and the speaker. You need
Docker and a browser — no Python, no git, no environment variables.
# terminal: marketplace + plugin in one line
claude plugin marketplace add noisy/noisy-coding && claude plugin install noisy-coding@noisy
# inside Claude Code (new session):
/noisy-coding:setup
The setup command starts the published image and walks you through first contact. Then finish in the browser at http://127.0.0.1:8765: paste your xAI API key (console.x.ai) and click the amber ENABLE TAB AUDIO banner — that one click makes the tab your microphone and speaker. Keep the tab open and just talk.
Prefer staying inside Claude Code? Same thing, four commands:
/plugin marketplace add noisy/noisy-coding →
/plugin install noisy-coding@noisy → /reload-plugins →
/noisy-coding:setup.
Other setups — plain Docker without the plugin, native install with hardware mic/speakers, remote hosts, all configuration knobs — live in docs/INSTALL.md.
All speech logic lives in one listener daemon — the single owner of
the microphone, the playback queue and the speakers. The MCP server is a
thin messenger that forwards speak requests; Claude Code hooks deliver
your transcribed speech back into the session (see
docs/hooks.md).
mic (hardware or browser tab via WS :8766)
-> VAD -> Grok STT -> transcript queue -> HTTP :8765
^ polled by Claude Code hooks
speak (MCP, stdio or HTTP :8767) -> POST /speak -> daemon queue
-> Grok TTS -> speakers (hardware or browser tab)
| Tool | What it does |
|---|---|
speak(text, interrupt?) | Queues text for speech and waits until it has played. Voice/speed/language come from the daemon (dashboard character), not the call. |
announce(text) | Fire-and-forget variant: returns immediately, plays in the background. |
change_voice(voice_id) | Deliberately switches this agent's voice (persists, shows on the dashboard). |
list_voices() | Lists the built-in Grok voices (ara, eve, leo, rex, …). |
MIT © Krzysztof Szumny
Content type
Image
Digest
sha256:2e78da9d4…
Size
108.7 MB
Last updated
18 days ago
docker pull noisy/noisy-coding