Hold a key, speak, and text appears in whatever you are typing into. Select text, speak a change, and it is rewritten in place.
Your audio never leaves your machine.
Open-source software, proprietary model. Every line of this project is MIT
licensed and auditable. The speech recognition it currently depends on is
Apple's on-device SpeechTranscriber, whose weights are Apple's and closed.
That is a real limitation and it is stated here rather than in a footnote:
-
The recognizer runs on your device, so "nothing leaves your machine" holds. It is local, and it is not open.
-
The default recognizer requires macOS 26 or newer, because
SpeechTranscriberis a macOS 26 API. On macOS 13-25 the app installs and runs, and dictation works only once you point it at a whisper model:# a ggml model from https://huggingface.co/ggerganov/whisper.cpp export OUTLOUD_WHISPER_MODEL=~/.outloud/models/ggml-base.en.bin outloud --asr whisper
Without that the recognizer never becomes ready and the hotkey appears to do nothing.
doctorsays so on those versions rather than reporting a clean bill of health. -
whisper.cpp is implemented and is the open-weights path. Parakeet TDT is still a stub. Windows and Linux have no working recognizer yet.
If a fully open stack is what you need today, this is not yet that. It is the
rest of the machine built around one, and the recognizer is a seam designed to
be swapped (crates/asr).
Status: working prototype. Dictation, edit-by-voice, and shell command-line editing are verified end to end on macOS. Windows backends (UI Automation, keyboard hook, layered overlay) are implemented and compile in CI on real Windows runners, but have not been exercised on Windows hardware. Linux is still designed and stubbed. See what works.
The orange step is the only closed component. Everything else is in this repo.
Text is delivered through whichever transport the focused application actually supports, best first, falling back until something works:
Green paths can read the existing text, which is what makes edit-by-voice
possible. Orange paths are insert-only: dictation works, editing does not.
Why each tier exists, and which applications land in which, is in
docs/compat-matrix.md.
Aqua Voice and Wispr Flow are excellent and are both cloud products: your audio leaves your device, transcripts are retained unless you opt out, and there is no offline mode. The open-source alternatives (Handy, VoiceInk, Whispering) are local but stop at dictation. None of them can edit text you have already written.
Three things here are, as far as we can tell, not available anywhere else in open source:
- Edit-by-voice. Select a sentence, say "change hello to goodbye", and it is rewritten in place through the accessibility API, preserving the host application's undo where possible.
- Terminal and shell support. A terminal exposes no writable accessibility
field, so we cooperate with the line editor directly. You can rewrite a
kubectlcommand by voice, and^Xuundoes it through zsh's own undo. - Headless operation. A build with no display libraries linked at all, for SSH sessions and servers.
It is also faster. Measured end to end on an M4 Pro: 131-215ms from key release to text on screen, against Aqua Voice's advertised ~450ms insert latency. The spread is real and depends on the transport: an accessibility write into a native field is the fast end, synthesized keys into a terminal the slow end.
Every number below was measured on this machine, not estimated.
| Path | Result | Latency |
|---|---|---|
| Dictation into a native app | "The rain in Spain falls mainly on the plain." | 189ms |
| Edit-by-voice on a selection | "quick" → "slow", in place | 131ms |
| Shell command line | --namespace prod-web → staging-web, zsh undo intact |
verified |
| Clipboard fallback (unfocused) | text still delivered | 445ms |
Recognition is Apple's on-device SpeechTranscriber (macOS 26+), which needs no
model download and whose weights are Apple's, not ours. Parakeet TDT and
whisper.cpp backends are stubbed, not implemented: the files exist with a
documented integration plan and model URLs, and they return an error if you
select them. Implementing one is the single highest-value contribution
available here, because it is what would make the stack open end to end and
give Windows and Linux a working recognizer.
Not yet built: Linux transports, a Windows shell integration (the unix-socket
bridge needs a named-pipe plus PSReadLine equivalent), streaming partial
injection wired into the daemon, and a settings UI. See
docs/planning/00-roadmap.md.
Most of it has now been run on real Windows hardware (an RTX 5090 box, whisper.cpp with CUDA), and doing that found several defects that CI could not see because CI compiles but cannot exercise GUI, input, or focus. What each row claims below is what was actually observed, not what compiles:
| Piece | Mechanism | State |
|---|---|---|
| Hotkey | WH_KEYBOARD_LL hook, RegisterHotKey conflict probe |
works on hardware: chord drives listening -> transcribing -> idle |
| Read + in-place write | UI Automation TextPattern / ValuePattern |
works on hardware; no undo preservation (needs TSF) |
| Typing fallback | SendInput with KEYEVENTF_UNICODE |
works on hardware (live dictation landed this way) |
| Clipboard fallback | clip.exe / Get-Clipboard + synthetic Ctrl+V |
works on hardware (live dictation landed this way) |
Undo ("scratch that") |
reads the field back, resolves against the ring | works on hardware, after a fix: it used to type the words instead |
| Per-app transport rules | process name -> AX / typing / clipboard | works on hardware; check any app with outloud --route NAME |
| Single-instance guard | Global\ named mutex |
works on hardware: the second daemon refuses |
| Overlay | layered, click-through, topmost, non-activating window | draws and runs; nobody has confirmed how it looks |
| Terminal | ConPTY bracketed paste (owned pseudoconsole) | implemented; foreign console needs a helper process |
| Shell integration | unix-socket bridge | not ported (needs named pipe + PSReadLine module) |
Windows needs a whisper model: put a ggml-*.bin beside the executable, or
point OUTLOUD_WHISPER_MODEL at one. Apple's recognizer is macOS-only, so
--asr whisper is the default everywhere else.
docs/windows-handoff.md has the traps worth knowing before changing any of
this, and the scripts that verify each row.
The trap that will bite first: UIPI. A non-elevated process cannot see keys
typed into, or inject text into, an elevated window. Dictation goes silent
while an admin app has focus and recovers when focus moves. Details in
docs/hotkeys.md and
docs/compat-matrix.md.
The quickest path, which downloads the latest release and needs no toolchain:
curl -fsSL https://raw.githubusercontent.com/blarer/outloud/main/scripts/install.sh | bashApple Silicon only, and it will tell you by name if your machine cannot run
it rather than failing halfway through. It installs to /Applications
(override with OUTLOUD_INSTALL_DIR), quits a running copy first, and
prints the two permissions macOS will not prompt loudly for.
To build from source instead, which is what the rest of this section covers:
Requires macOS 13 or newer. On 26+ the recognizer is built in and needs no
download; on 13-25 it needs a whisper model (see above). Also needs a Rust
toolchain, and Xcode Command Line Tools for swiftc. Windows builds and
installs (scripts/build-windows.sh ships outloud.exe and outloud-spike.exe)
but is untested on hardware; Linux does not work yet.
git clone https://github.com/blarer/outloud
cd outloud
# Builds the daemon, compiles the Swift speech helper, packages the .app, and
# signs it. Use this rather than a bare `cargo build`: the recognizer is a
# Swift child process, not a linked library, so cargo alone does not produce
# it and the daemon comes up unable to transcribe anything.
./scripts/bundle-outloud-macos.sh
# Grant Accessibility against the bundle. macOS attaches the grant to a signed
# bundle rather than to a bare binary, and reading and rewriting text in other
# applications is exactly what that permission governs.
open "x-apple.systempreferences:com.apple.preference.security?Privacy_Accessibility"Then launch it through LaunchServices, so the app is its own responsible process rather than inheriting your terminal's permissions:
open -a "$PWD/dist/OutLoud.app"It has no Dock icon by design. Look for its icon at the right end of your menu
bar. To remove it later, ./scripts/uninstall-macos.sh (add --purge to
delete your settings too, or --dry-run to see the plan first).
These builds are unsigned and un-notarized. Gatekeeper rejects them, so an app copied from another machine will not open by double-clicking. Building locally, as above, avoids the problem entirely because locally built files carry no quarantine flag. See known limitations.
If anything misbehaves, run the doctor before anything else. Almost every failure in this category is environmental rather than a bug, and each check names the exact next action:
./scripts/doctor.shFor a copy-pasteable path from clone to dictating, including which permission
dialogs to approve and the responsible-process trap that makes a granted
permission look denied, see
docs/macos-quickstart.md.
Start the daemon and leave it running:
open -a "$PWD/dist/OutLoud.app"Then, in any application:
- Put your cursor where you want text.
- Hold right-option. The overlay appears.
- Speak.
- Release. Your words appear at the cursor.
To try it without a microphone, feed it synthesized speech. Note the path: this runs the bundled binary, which is the one that ships with the speech helper beside it.
./dist/OutLoud.app/Contents/MacOS/OutLoud --once --say "hello from a local dictation daemon" --no-overlayOutLoud has no Dock icon and no window on purpose: it types into whatever field you are focused on, so it must never steal that focus. Its whole visible presence is one icon at the right of the menu bar, and the glyph is the answer to "is it on?" without a click. A waveform means ready; a filled microphone means the microphone is open right now.
Clicking it gives you the current state, the hotkey it actually bound, the microphone it actually opened, Pause Dictation, a Settings submenu, Run Diagnostics, and Quit OutLoud. When a permission is missing, a row appears that opens the exact System Settings pane rather than telling you to go find it.
Settings are written straight into your config.toml, comments preserved,
and edits you make in an editor show up in the menu within a second. The menu
deliberately offers only the settings that are implemented today; the config
file lists every key the schema knows, and if you set one that nothing reads
yet, the menu says so instead of ignoring you.
This is the part other dictation tools do not do.
- Select some text in any application.
- Hold right-option and speak a command.
- Release. The selection is rewritten in place.
Commands are matched literally, so they are predictable rather than clever:
| Say | Effect |
|---|---|
| "change X to Y" | replaces every X with Y |
| "replace X with Y" | same |
| "make X into Y" | same |
| "swap X for Y" | same |
| "delete X" | removes X |
| "remove X" / "get rid of X" / "scratch X" | same |
| "add X" / "append X" | appends X |
| "all caps" / "uppercase" | THE WHOLE SELECTION |
| "lowercase" | the whole selection |
| "title case" | The Whole Selection |
| "sentence case" | The whole selection |
Matching ignores case, because speech recognition will not reproduce the casing on your screen. If nothing matches, you are told so rather than having the text silently changed.
Anything that is not one of the above ("tighten this up", "make it more formal") is a freeform edit and needs the local language model, which is built but not yet wired into the daemon. Today the daemon says so instead of doing nothing.
A terminal exposes no writable text field to the accessibility API, so this works by cooperating with your shell's line editor directly.
# Install the plugin for your shell. It appends one guarded line to your rc
# file and composes with oh-my-zsh and friends.
cargo run --release -p shell-bridge -- install
# Run the bridge alongside the daemon.
cargo run --release -p shell-bridge -- serveThen, at a prompt:
- Type a command but do not run it.
- Speak an edit (same deterministic commands as above) while the terminal is the focused window. The daemon stages it on the bridge instead of typing it; the menu bar shows what was staged.
- Press Ctrl-X Ctrl-A. The command line rewrites in place.
- Ctrl-X u undoes it, through your shell's own undo.
$ kubectl get pods --namespace prod-web --output wide
say: "change prod-web to staging-web", then ^X^A
$ kubectl get pods --namespace staging-web --output wide
Bash and fish are supported too. The bridge never executes anything: the protocol has no execution verb at all, and a rewritten line always waits for you to press enter.
Two boundaries, both deliberate. Only the deterministic commands (change/replace/delete/add/case) are routed to the bridge: a freeform phrase spoken at a prompt is dictation (a commit message, a grep pattern) and is typed as text, so voice typing at a terminal keeps working. And if no bridge is serving, spoken edits are typed as text too. You can also stage an edit by hand, no microphone involved:
cargo run --release -p shell-bridge -- intent "change prod-web to staging-web"--once run one dictation cycle and exit
--say TEXT synthesize TEXT with `say` instead of using the microphone
--wav FILE feed a WAV file instead of the microphone
--chord CHORD change the hotkey (default: right-option)
--asr apple|mock choose the recognizer
--no-overlay log state changes instead of drawing the panel
--realtime pace file audio like live speech
Hotkeys are written the way you would say them: right-option, fn,
cmd+shift+space. Conflicts with existing system shortcuts are detected and
reported rather than silently failing.
| Crate | Responsibility |
|---|---|
outloud |
The daemon. Wires everything together and owns the state machine |
audio |
Capture, ring buffer, resampling, VAD, speech segmentation |
asr |
Streaming recognizer trait, Apple/Parakeet/whisper backends, model manager |
stream |
Commit horizon, minimal diffs, coalescing, undo ring |
edit-intent |
Spoken command → deterministic text transformation |
llm |
Local model for freeform edits, with guardrails and preview |
text-target |
Picks and drives a transport for any destination |
ax-edit |
macOS accessibility read/rewrite |
shell-bridge |
Unix socket + shell plugins for command-line editing |
hotkey |
Global push-to-talk, tap-to-latch, conflict detection |
overlay |
Non-activating floating panel |
config |
Layered configuration, per-app profiles, vocabulary |
diag |
Environmental checks, timing, redacted bug reports |
spike-cli |
Development harness for the accessibility layer |
A language model is the fallback, not the first resort. Most edit commands are a small closed set: replace, delete, append, recase. A deterministic parser handles them in microseconds with no GPU and no chance of a model rewriting text nobody asked it to touch. Only open-ended instructions escalate to a local model, and those are previewed before they apply.
The overlay must never take focus. Taking focus would destroy the text field
we are about to edit, so this is a correctness requirement rather than polish.
It is a non-activating NSPanel that cannot become key.
Committed text is never retracted. A streaming recognizer revises itself: "recognise speech" can become "wreck a nice beach" three words later. Text is only committed once several consecutive hypotheses agree on it, so the user never watches their document rewrite itself.
Transports are chosen by capability, not by guesswork. Selection is a pure
function of an Env trait, so every branch is unit tested rather than only
reachable on a machine that happens to have that software installed.
Headless is a compile-time gate. Building without the display feature drops
the GUI dependencies entirely, so a display library reaching the default feature
set is a compile error rather than a runtime crash on a server.
Honest list of what will go wrong, so a first run is not a surprise. Detail and
evidence in docs/beta-readiness.md.
| Limitation | What you will see | Workaround |
|---|---|---|
| Unsigned and un-notarized | An app copied or downloaded from another machine silently refuses to open. spctl -a -t exec dist/OutLoud.app says rejected |
Build it locally; local builds carry no quarantine flag |
cargo build alone is not enough |
recognizer failed to load (speech helper not found...) |
Use ./scripts/bundle-outloud-macos.sh, which compiles the Swift helper |
| Only one copy may run | A second launch is refused, naming the pid to quit | Quit the first from the menu bar, or kill N |
| Accessibility grant dies on every rebuild | Toggle reads "on", every call fails. The menu bar glyph turns into a warning triangle within a second | tccutil reset Accessibility dev.outloud.outloud, then re-grant |
| macOS 13-25 has no bundled recognizer | recognizer never becomes ready |
SpeechTranscriber needs macOS 26+; set OUTLOUD_WHISPER_MODEL and run --asr whisper |
| Most config settings are not read yet | Changing them has no effect and no warning | Only hotkey, enabled, microphone.sensitivity, microphone.warm-hold-ms, silence-timeout-ms, and overlay.position are wired today |
| Quiet or distant speech is dropped | Words go missing unless you lean in and enunciate | Raise Microphone Sensitivity in the menu bar (or microphone.sensitivity in config) |
| Bluetooth clips the first word | The headset needs ~200ms to start capturing; the daemon warns you once per device | Set microphone.warm-hold-ms = 2000, or hold the key a beat before speaking |
| Freeform edits are not wired up | "tighten this up" reports that it needs the language model | Use the literal commands listed above |
| Linux cannot type at all | outloud compiles for Linux, but every text delivery tier refuses: there is no X11/Wayland key synthesis. It declines without touching your clipboard |
macOS only for now |
| Windows is unexercised | Built and lint-checked for Windows every commit, but not run there recently | Treat Windows as unverified |
This works. You can dictate into a Discord, FaceTime, Zoom, or Meet call while that app is using the microphone, and neither side loses audio: CoreAudio shares input devices between processes rather than granting one of them exclusive ownership.
Two implementation choices keep it that way, and both are guarded by
crates/audio/tests/shared_device.rs so they cannot be undone by accident:
- We never take hog mode, which is the one call that would seize the device and would knock a call app off the microphone the instant you pressed the hotkey.
- We accept whatever format the device is already running and resample to 16kHz ourselves, instead of demanding a sample rate and forcing a reconfiguration on whoever got there first.
The microphone is also only open while you hold the hotkey, so macOS's orange recording dot means exactly what it appears to mean.
One caveat is specific to Bluetooth: a headset switched into its low-quality call profile is quieter and more compressed for everything using it, so recognition accuracy drops for the same reason the other participants sound worse. Wired and built-in microphones are unaffected.
Because the microphone opens on key-down, whatever a device takes to deliver its first sample lands between your keypress and the first audio anything can hear. The built-in microphone measures 71ms, which is harmless. Bluetooth headsets must negotiate their hands-free profile first, which is typically several hundred milliseconds.
Measured cost, through the real recognizer, saying "the quick brown fox jumps over the lazy dog" with the head of the audio removed:
| Audio lost at the start | What gets transcribed |
|---|---|
| 0ms | "The quick brown fox jumps..." |
| 100ms | "Quick brown fox jumps..." |
| 200ms | "Like brown fox jumps..." |
| 500ms | "fox jumps over the lazy dog." |
The 200ms row is the one to worry about, and it is not the one that lost the most audio: a half-captured word is not dropped, it is misheard. A missing word is obvious. A plausible wrong word is not.
There is no handling for this yet, and no warning. If you dictate on AirPods
and the first word is wrong more often than it should be, this is why. Pausing
briefly between pressing the key and speaking avoids it entirely. Full
measurements and the options considered are in
docs/input-latency.md.
When something goes wrong, ./scripts/doctor.sh classifies it as permission,
configuration, or bug, and only the last belongs in an issue.
| Document | Read it when |
|---|---|
docs/beta-readiness.md |
You want the honest state of the rough edges |
docs/M0-results.md |
You want the measured result and what it cost |
docs/latency.md |
You care where the milliseconds go |
docs/macos-permissions.md |
Anything permission-shaped is behaving strangely |
docs/debugging.md |
Something works in one application and not another |
docs/compat-matrix.md |
You need to know what a destination supports |
docs/shell-integration.md |
You are working on terminal support |
docs/streaming.md |
You are touching partial text commitment |
docs/configuration.md |
You are adding or changing a setting |
docs/signing-runbook.md |
Certificates, or why grants keep dying |
docs/testing.md |
You want to know what is tested and what cannot be |
docs/neural-engine.md |
You are wondering whether the ANE is actually used |
docs/overlay-performance.md |
You are changing anything that draws |
docs/investigations/robustness.md |
You want the ranked failure list, with reproductions |
docs/ux/ |
You are designing user-facing behaviour |
docs/planning/ |
You are picking up work or planning a milestone |
CONTRIBUTING.md |
You are about to write code here |
docs/investigations/release-ships-the-spike.md |
You are about to cut a release, or wonder why the download does not dictate |
docs/investigations/install-link-depends-on-a-branch.md |
You are about to delete a branch, or the install one-liner 404s |
These were all discovered the hard way. They are why doctor exists.
- The system-wide
AXUIElementdoes not work. Asking it forAXFocusedUIElementreturnskAXErrorCannotCompleteeven for a fully trusted process. Resolve the focused application first, then ask it. - Accessibility grants follow the responsible process. A binary run from a shell is judged against your terminal's permission, so the app can appear enabled in System Settings and still be denied. Launch through LaunchServices.
- Ad-hoc signatures invalidate grants on rebuild. TCC pins approval to the
binary's
cdhash. The toggle keeps reading "on" while nothing works. Usetccutil resetduring development, and a Developer ID certificate for real. - Windows hang off
AXWindows, notAXChildren. An application element's children are its menu bar.
cargo test --workspace # 400 tests, no permissions needed
./scripts/doctor.sh # environmental checks
./scripts/test-real-apps.sh # drives TextEdit and Safari, skips cleanly if absent
./scripts/verify-shell-bridge.sh # rewrites a command line in a real zsh
./scripts/bench-latency.sh # criterion benchmarks against live applications
cargo bench -p ax-edit --bench gate # latency regression gateThe fuzz suite found a real panic and a silent over-edit within minutes of being written, which is the entire argument for having it.
MIT for all code. Model weights carry their own licences and are documented
separately in docs/asr-integration.md and
docs/llm.md.