The problem I saw
Every dictation app I tried sent my voice to the cloud. Ten to twenty dollars a month for something my Mac could already do locally. Private thoughts, medical notes, half-formed journal entries — all uploaded to someone else's servers, kept who-knows-where, used who-knows-how.
Apple Silicon is fast enough to run Whisper on-device in real time. There was no good reason any of this audio needed to leave the machine. So I built the thing I wanted to use.
What it does
Hold Right Option, speak, release. The transcribed text gets pasted at your cursor in whatever app is focused — VS Code, Terminal, Chrome, Slack, Notes, Pages, anything you can type into.
All speech recognition runs on the Apple Neural Engine via WhisperKit. There's an optional grammar-cleanup pass through a local LLM (Ollama) that strips filler words and fixes punctuation, also entirely on-device. No accounts. No cloud. No subscription. No telemetry.
Install
One command to clone, build, and launch. The first build pulls WhisperKit dependencies (~2 minutes); after that, builds take about 2 seconds.
git clone https://github.com/Rajvardhman05/openwhisper-app.git
cd openwhisper-app
bash build.sh
open build/OpenWhisper.app
On first launch, grant Microphone and Accessibility permissions when prompted. OpenWhisper lives in the menu bar — look for the microphone icon.
How it works
Hold the hotkey to start recording. Speak naturally. Release to transcribe. Text appears at your cursor wherever you are.
- 100% local and private. Speech recognition runs on-device using the Apple Neural Engine. No audio ever leaves your Mac.
- Works offline. No internet needed after the one-time model download.
- Optional AI cleanup. A local LLM via Ollama can remove filler words and fix punctuation — also fully on-device.
- 29 languages. English, Spanish, French, German, Hindi, Chinese, Japanese, and 22 more, with auto-detection.
- Lightweight. Under 100 MB RAM and less than 1% CPU when idle.
- Works in any app. If you can type into it, you can dictate into it.
vs. cloud dictation
Every other dictation tool sends your voice to someone else's servers. OpenWhisper doesn't.
| OpenWhisper | Cloud services | |
|---|---|---|
| Privacy | 100% local | Voice uploaded to servers |
| Internet | Works offline | Requires connection |
| Cost | Free and open-source | $10–20 / month |
| Latency | Instant on-device | Network round-trip |
| RAM usage | <100 MB idle | 400–800 MB |
| Data collection | None | Voice stored on servers |
Choose your Whisper model
Models are downloaded once from Hugging Face and cached locally. Pick the right balance of speed and accuracy for your use case.
| Model | Size | Speed | Best for |
|---|---|---|---|
tiny |
39 MB | Fastest | Quick notes, short phrases |
base |
140 MB | Fast | General dictation (recommended) |
small |
460 MB | Moderate | Longer passages, multilingual |
small.en |
460 MB | Moderate | English-only, highest accuracy |
Architecture
Built with Swift and SwiftUI. No Xcode project required — it builds entirely from the command line with Swift Package Manager.
- Audio engine: AVAudioEngine, 16 kHz mono resampling.
- Transcription: WhisperKit (CoreML + Apple Neural Engine).
- LLM cleanup: Ollama HTTP API on localhost (optional).
- Text injection: NSPasteboard plus a synthetic
Cmd+Vvia CGEvent. - Global hotkey: Right Option key via NSEvent monitors.
- UI: SwiftUI menu bar popover plus a floating FlowBar status indicator.
- Language: Swift 5.10, ~93% of the codebase.
- Requirements: macOS 14.0+ (Sonoma), Apple Silicon (M1–M4).
Where things stand
Shipped and working today:
- Core transcription engine via WhisperKit
- Hold-to-talk global hotkey
- Auto-paste into any app
- Local LLM grammar cleanup via Ollama
- Multiple model selection (tiny / base / small / small.en)
- 29 languages with auto-detection
- FlowBar animated status indicator
- Open-source on GitHub (MIT license)
In progress:
- Homebrew Cask distribution
On the roadmap:
- Custom hotkey configuration
- Mac App Store release