Skip to content
arhamhiPublic

About

Private, on-device voice-to-text for macOS. Hold a key, talk, release: text lands in any app. Open-source Wispr Flow alternative built on Apple SpeechAnalyzer. No subscription, audio never leaves your Mac.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Spokesman app icon

Spokesman

Private, on-device voice-to-text for macOS.
An open-source Wispr Flow alternative. Hold a key, talk, let go, and it types for you in any app.

macOS 26 Swift 6.2 On-device MIT license No subscription

Spokesman dashboard: push-to-talk status, cleanup mode, last dictation and today's word count

Okay so, Spokesman is a dictation app for Mac. You hold Control+Option, talk like you normally would, let go, and the text just shows up in whatever app you're in. Claude, ChatGPT, Codex, Slack, Mail, your editor, anywhere with a text box.

I built it because I wanted something like Wispr Flow, but that's a subscription and your voice goes to the cloud. I wanted the same feel without either of those. So this runs on Apple's on-device speech model (SpeechAnalyzer + DictationTranscriber), it transcribes while you're still talking so there's barely any wait when you let go, and the audio only lives in memory and then gets thrown away. It never touches disk and it never gets uploaded.

Why I use it

Most of what I dictate is long, messy prompts into Codex and Claude. I say "um", I restart sentences, I go "Tuesday, no, Thursday". I don't want that cleaned into some generic corporate paragraph though, I want it to still sound like me, just without the mess. So the cleanup here is pretty careful about what it touches.

What it does

  • Push-to-talk or hands-free. Hold the shortcut to talk, or double-tap it and it keeps recording hands-free. Esc cancels. You can change the shortcut.
  • Fast, on-device transcription. Apple Dictation runs while you speak and stays warmed up, so the wait after you let go is short.
  • Pastes into the app you're in. A small HUD shows you what's happening and it doesn't steal focus from the field you were typing in.
  • Cleanup that keeps your voice. Takes out "um", "uh", "like", repeated words and the self-corrections, and leaves your actual phrasing alone.
  • Your own dictionary. Add names and jargon once and it steers recognition toward your spelling, instead of guessing and patching words after.
  • Learns from your fixes. If you fix a misheard word right after dictating, it picks that up and learns the spelling.
  • History and stats. Search everything you've dictated, compare the cleaned text against what you actually said, and see your words, time saved, streaks and top words. All of it stays on your Mac.
  • Smart Formatting, if you want it. Bring your own OpenAI key and it'll format dictated text into lists and paragraphs. It's off by default, only the text gets sent (never audio), and if the call fails it just falls back to local cleanup.
  • An actual Mac app. Sits in the menu bar, can launch at login, native AppKit. No Electron, no Python, no terminal window you have to keep open.

Dictation HUD states: recording with live waveform, cleaning up, delivered 38 words to Claude

Screenshots

History Dictionary
Searchable dictation history with cleaned text and compare-raw toggle Personal vocabulary with learned spelling corrections

Stats: lifetime words, time saved versus typing, fillers removed, streak and top words

Everything in these screenshots is demo data, not my actual dictations.

Requirements

  • macOS 26 (Tahoe) or later (developed and tested on an M1 Pro)
  • Xcode 26 (only needed to build it)

Install

There's no prebuilt release yet, so for now you build it from source. It takes a couple of minutes:

git clone https://github.com/arhamhi/spokesman.git
cd spokesman
bash scripts/create-signing-identity.sh   # one time; keeps permissions working across rebuilds
bash scripts/build-macos-app.sh --install

This builds, signs, and installs Spokesman.app into ~/Applications and opens it. Say yes to Microphone, Speech Recognition, Accessibility and Input Monitoring when it asks. After that you don't need the terminal again.

If permissions act weird, the full walkthrough is in docs/MACOS-SETUP.md.

Privacy

  • Audio is processed in memory and discarded. It's never saved and never uploaded.
  • History, stats, and your dictionary live on your Mac. You can prune or erase them.
  • No telemetry, no analytics, no crash reporting.
  • Nothing leaves your machine unless you turn on Smart Formatting, which sends only the dictated text to OpenAI (with store: false).

How it works

Control+Option ─► AVAudioEngine (in-memory buffers)
                     │
                     ▼
        SpeechAnalyzer + DictationTranscriber  (on-device, streaming)
                     │   biased by your personal vocabulary
                     ▼
        Local cleanup (fillers, stutters, self-corrections)
                     │   optional: Smart Formatting via your OpenAI key
                     ▼
        Accessibility / CGEvent paste into the focused app

The native app is in macos/ (Swift Package, AppKit, WKWebView dashboard, SMAppService for launch at login). The earlier Windows/Python version is still in src/spokesman/ for reference.

Development

swift test --package-path macos
bash scripts/build-macos-app.sh

Product direction and design notes are in PLAN.md, docs/PRD.md, and DESIGN.md.

Contributing

Issues and PRs are welcome. This started as my own daily driver, so honestly the stuff most likely to get merged is anything that keeps it fast, private and simple. If you're not sure whether something fits, just open an issue and ask.

License

MIT

About

Private, on-device voice-to-text for macOS. Hold a key, talk, release: text lands in any app. Open-source Wispr Flow alternative built on Apple SpeechAnalyzer. No subscription, audio never leaves your Mac.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages