Skip to content

Repository files navigation

simple-ptt

simple-ptt demo

A fast, minimal push-to-talk app for macOS with live Deepgram transcription and optional LLM cleanup before paste.

simple-ptt is intentionally small: menu bar app, global hotkey, live on-screen transcript, fast paste into the currently focused app. In normal use it aims to stay around 35 MB of RAM. The goal is not to be feature-rich. The goal is to stay fast, understandable, and out of the way.

Quick start

Important

The bundled app is ad-hoc signed but not notarized. macOS will block it on first launch, run this to fix:

xattr -dr com.apple.quarantine /Applications/simple-ptt.app

Expect the usual macOS prompts for: Microphone and Accessibility for the global hotkey and synthetic paste workflow.

Deepgram Costs

Deepgram usage for this kind of developer push-to-talk workflow is usually cheap. As a rough example, 10 hours of speech-to-text should cost less than $5 USD, which covers close to a month of personal usage for this workflow. Actual cost depends on how long you dictate, which Deepgram plan you are on, which model you use, and Deepgram's current pricing. Check the official Deepgram pricing page before treating that number as current.

How it works

Data flow

  • Audio is streamed to Deepgram for transcription.
  • If transformation is enabled, buffered transcript text is sent to your configured LLM provider.

Default workflow

  • Press the record hotkey (F5 by default) to start listening.
  • Speak and watch the live transcript overlay update in real time.
  • Stop recording to paste the buffered text into the focused app.
  • If transformation is configured and enabled, the app can clean up the transcript before pasting.

Additional controls

  • Tap vs hold: short press behaves like toggle; holding past mic.hold_ms turns the same hotkey into hold-to-talk.
  • Editable overlay: You can click into the overlay at any time to manually type, fix, or delete words before pasting.
  • Correction key (LeftMeta, shown as Cmd, by default): hold the configured correction key during dictation or while a buffered annotation is visible, speak a correction request, then release the key to apply that correction to the current annotation.
  • Transform hotkey (F6 by default): transform the current transcript without auto-pasting it. If you press F6 while dictating, you can keep talking — your audio is buffered and will seamlessly append to the transformed text once the LLM finishes.
  • Resume dictation: If you have transformed text (or manually stopped recording), pressing F5 again will seamlessly resume dictating onto the end of your existing text.
  • Escape: abort recording, cancel background work, or discard a ready buffer.
  • Cmd+V while recording: splice the current plain-text clipboard contents into the active transcript.

Microphones and Instant Recording

To make sure your voice is captured the exact millisecond you press the hotkey, simple-ptt keeps your microphone "warm" and ready in the background.

Starting up a microphone in macOS normally takes about a quarter of a second, which would cut off the first word or two of your dictation. Keeping it warm ensures a zero-delay experience at a tiny trade-off of around 1.0% to 1.5% CPU usage when the app is idle. To save that idle CPU at the cost of the startup delay, clear Keep microphone connection open in Settings > Microphone (mic.always_on = false); the microphone stream is then paused except while you record or while the Settings window is open.

Additionally, the app automatically and instantly detects when you switch your default system microphone (like plugging in USB headphones or connecting AirPods) without using any heavy background polling or lagging your system.

Features

LLM text transformation

Simple PTT can optionally send your dictation through an LLM to remove filler words, correct punctuation, and format technical terms before pasting.

Deepgram Keyterms

You can specify custom keyterms in the configuration to boost the transcription accuracy for specific vocabulary like product names, technical jargon, or acronyms.

Configuration

simple-ptt looks for config in this order:

  1. SIMPLE_PTT_CONFIG
  2. $XDG_CONFIG_HOME/simple-ptt/config.toml
  3. ~/.config/simple-ptt/config.toml

If no config file is found, defaults are used where possible and the app opens Settings so you can create one. For normal app launches, ~/.config/simple-ptt/config.toml is the correct default.

Settings groups the options into toolbar panes: General (record and correction shortcuts, overlay font and meter, updates, start on login), Microphone (input device, sample rate, gain in dB with a live meter, silence pad, keep microphone connection open), Deepgram (API key, project ID, language, keyterms, model, endpointing, utterance end), and Transformation (transform shortcut, auto-transform, provider, API key, model, and editors for the dictation and correction prompts, which are sent only to the transformation model). Choosing a transformation provider fills the model list from the model cache (transformation-models.toml in $XDG_CACHE_HOME/simple-ptt, by default ~/.cache/simple-ptt), or fetches the provider's models when none are cached for that provider and API key. Providers other than Ollama and Hugging Face need their API key, in the field or its environment variable, before their models can be fetched. The API key field is shared by every provider, so a key typed there is sent automatically only when it is the key saved for the selected provider; otherwise enter that provider's key and click Fetch models. Fetch models reloads the list, and Check tests the connection. Save writes every pane to the config file. The transformation model and the two prompts are written only when they differ from the built-in defaults, so a config that leaves them out picks up improved defaults in later releases; a key the file already has is kept even when it matches the default. Reset to Default above each prompt editor in the Transformation pane puts the built-in prompt back; saving it without further edits removes that prompt from the config file even if the file had it. If the audio input devices can't be listed, or mic.gain is outside the slider's 0 to 10 dB range, Settings still loads every pane and explains the problem in its status area; the configured device stays selected, and Save writes the gain the slider shows. Settings opens once macOS brings simple-ptt to the front; if you choose Settings… and nothing appears, click the simple-ptt icon in the Dock.

The correction interrupt is configured separately from the record and transform hotkeys via ui.correction_key. This must be a single specific key such as LeftMeta, RightMeta, LeftAlt, or F7, and it must not overlap with the record or transform triggers.

Transformation now has two separate prompts:

  • transformation.system_prompt for normal cleanup or rewrite of dictated text.
  • transformation.correction_system_prompt for correction mode, where the model receives both the current annotation and the spoken correction request.

Minimal config

If you prefer to edit the file by hand, this is enough to get transcription working. See config.example.toml for all available options.

[deepgram]
api_key = "YOUR_DEEPGRAM_API_KEY"

Replace YOUR_DEEPGRAM_API_KEY with your Deepgram API key. A non-empty deepgram.api_key takes precedence over the DEEPGRAM_API_KEY environment variable, which simple-ptt reads only when the key is omitted or empty. Apps opened from Finder or the Dock don't reliably inherit shell environment variables, so keep the key in the config file for normal launches.

Development

For the repo-local development config (./config.toml, gitignored; just run creates it from config.example.toml when it is missing):

just run

Useful helper targets:

just run-config path/to/config.toml
just run-xdg
just bundle-release
just bundle-dmg
just install-app
just start
just list-devices

License

MIT. See LICENSE.

About

A macOS menu bar dictation app with Deepgram transcription and paste-to-focused-app flow

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages