Public 0.9 beta · MIT · signed and notarized · macOS 13+ on Apple Silicon
Plainsong
Dictation into any app and meeting notes without a bot joining the call, transcribed on your own Mac. Formerly a commercial product, now MIT and free.
Executive summary
Dictation and meeting notes are the two places where people hand raw speech to somebody else's server without thinking about it. A meeting bot joins the call as a visible participant and takes the audio somewhere; a dictation tool streams every sentence you speak to an endpoint you did not choose. Plainsong does both jobs on the machine in front of you. Press a hotkey, speak, and the text appears at the cursor in whatever app has focus. Record a meeting without sending anything into the call.
This started as a commercial app called Nautilus Bot. On September 5, 2026 it became public under the MIT license, with no trial, no tiers, and no paid version held back. The source is on GitHub and the signed beta is a download anyone can take, which means the privacy claims on the product site are checkable rather than promised.


What local-first has to mean to be worth saying
Local-first is a marketing phrase in most products that use it. In Plainsong it is a property with a short list of exceptions, and the exceptions are named.
There are no Plainsong servers. There is no account system, no telemetry, no analytics, and no crash reporting. Transcription runs on your Mac. Dictation audio goes to a temporary file that is removed after processing, and meeting audio is kept only under the retention setting you chose.
- Downloading a speech model reaches Hugging Face once per model. The default, Parakeet TDT 0.6B v3, is a 639 MiB download. Whisper base.en at 142 MiB is the smaller alternative.
- The optional Silero voice-activity model and the speaker-diarization models are pinned to a specific upstream revision and checked against an expected SHA-256 before they are accepted.
- An update check happens only when you ask for one. There is no check on launch.
- Cloud transcription and cloud cleanup are opt-in, use your own provider keys stored in the system keychain, and name the provider at the point you choose it. Usage is billed to you because nothing is proxied through anything of mine.
Three ways to start talking, because one never fits
The activation mode is the part of a dictation tool people either love or uninstall over. Plainsong ships three and lets the hotkey stay where you want it.
| Mode | How it behaves | Where it fits |
|---|---|---|
| Toggle | Press once to start, once to stop. The onboarding default. | Longer passages and anything where holding a key is tiring. |
| Hold to talk | A native key listener starts on press and stops on release, falling back to toggle if the helper is not running. | Short bursts, replies, and command-like phrases. |
| Hands-free | Voice-activity detection starts and stops on its own, with an optional Silero model for better accuracy. | Working with your hands busy, at the cost of the least reliable mode today. |
- A focus-preserving overlay shows live partial text as you speak, but the text actually inserted is always the finished transcription made after you stop. The live preview never changes what gets typed.
- Secure fields are excluded completely. A password box, or any field while the system secure-input flag is set, never receives an insertion, never gets text staged on the clipboard, and never triggers the copy Plainsong uses to read a selection.
- Spoken editing commands undo, delete, or replace the last word, clause, sentence, or paragraph, optionally by count.
- Snippets and a personal dictionary handle the names and terms that every speech model gets wrong the first time.
Picking a model without pretending there is a best one
Local speech models trade size, speed, language coverage, and behavior in silence against each other, and no single choice wins. Rather than hide that behind an automatic pick, Plainsong offers presets that state what each one buys and what it costs.
The Balanced preset uses Parakeet for both dictation and meetings, and its stated advantage is specific: a transducer that emits silence during silence, so stopping mid-sentence to think does not become invented words. Whisper large-v3-turbo covers roughly a hundred languages upstream, and its stated cost is that it is slower per utterance and fills long pauses with text that was never spoken. Both facts matter more than a quality score.
Apple's on-device speech is available with no download at all, using SpeechAnalyzer on macOS 26 or later, which returns per-segment timestamps and can serve meetings, and the older recognizer on earlier systems, which cannot. Apple's server fallback is switched off in both cases.
Meetings without a participant nobody invited
- Recording covers the microphone, plus system audio where the system allows it. Native system-audio capture needs macOS 14.7 or later; below that it requires an already-configured loopback device such as BlackHole, or a microphone-only meeting.
- Transcripts get speaker separation, summaries, action items, and search that runs across meetings rather than inside one.
- Analysis can run against a local model through Ollama, with a catalog including GPT-OSS 20B, DeepSeek R1 Distill 8B, Ministral 3 8B, Llama 3 8B, and smaller options, or through your own cloud keys if you would rather.
- Exports are Markdown, DOCX, plain text, JSON, SRT, and VTT, so a transcript is a file you own rather than a row in somebody's database.
Publishing the limits at the same time as the download
The beta 4 release notes and the product site's ledger page carry the same list of what is not finished, including the parts that are unflattering. Hands-free dictation did not activate in the last two test runs, while toggle and hold-to-talk passed. The longest meeting verified end to end is 45 seconds, for a feature built to run for hours. Nobody has walked an install-to-install update on real hardware yet. The launch-time target was missed, on a candidate run made under extreme host load.
Publishing that alongside a download button is uncomfortable, and it is the correct trade. A beta whose failure modes are listed is usable, because you know which parts to check. A beta that only lists features teaches people to distrust the next release note too.
What did pass is stated with the same precision: ten source gates, 1,897 front-end tests, 1,731 Rust tests, clean JavaScript and Rust dependency audits, Developer ID signing, notarization, stapling, and packaged runs of first launch, the dictation hotkey, microphone meeting capture, and combined microphone plus system audio capture.
Limits
Limits and open questions
- This is a beta. Keep your own backups and read a transcript before relying on it.
- Apple Silicon only. The build is not qualified for Intel Macs, and Windows and Linux are roadmap items rather than dates.
- Hands-free activation failed in the maker's last two test runs. Toggle and hold-to-talk passed.
- Meetings longer than 45 seconds have not been soak-tested, though the feature is built for hours.
- No Homebrew cask yet, and no install-to-install update has been walked on real hardware, so treat the first update as a beta step too.
- The interface is English only, and while the default model's vendor lists 25 European languages, only English has been qualified here.
- Optional cloud providers are a real exit from local-first. The app names the provider at the point of choice, but sending text to a third party is still sending text to a third party.