kesha-voice-kit
Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Documentation
Kesha Voice Kit
Give your local tools and LLM agents a voice.Fast speech-to-text, text-to-speech, voice-activity detection, and language detection in one local-first CLI — CoreML on Apple Silicon, ONNX on Linux and Windows.
- Transcribe locally — 25 languages, up to ~19x faster than Whisper on Apple Silicon, ~2.5x on CPU
- Speak back — text-to-speech in 9 languages
- Plug into agents — ship voice workflows as CLI commands, an MCP server, an skill, or a agent
- Small Rust engine — single ~65MB binary, no ffmpeg, no Python, no native Node addons
Quick Start
Runtime: **Bun** >= 1.3.0.
# 1. Install Bun (skip if you have it)
curl -fsSL https://bun.sh/install | bash # macOS/Linux — or: brew install oven-sh/bun/bun
powershell -c "irm bun.sh/install.ps1 | iex" # Windows
# 2. Install Kesha
bun add -g @drakulavich/kesha-voice-kit
kesha --version # confirms `kesha` resolved on PATH
# 3. Download the engine and models — pick one path
kesha init # guided: TTS languages and optional VAD / diarization
kesha install --plan && kesha install # manual: preview the sizes, then download
# 4. Transcribe
kesha audio.ogg # transcript to stdout`kesha install` pulls ~2.5 GB on Linux/Windows and ~0.6 GB on Apple Silicon, whose CoreML engine reads a smaller model set. It is always explicit — nothing downloads behind your back — and reports download progress on stderr. If `bun --version` fails right after step 1, reload your PATH: `exec $SHELL -l`.
Prefer Homebrew or Docker? See Other install methods.
Air-gapped or behind a corporate mirror? See docs/model-mirror.md.
Platform support
All three targets transcribe, detect the spoken language, run VAD, and speak. The macOS-only rows need Apple frameworks — they are not a missing port. Windows is a tested path rather than a published binary nobody ran: CI does a cold `kesha install` on `windows-latest`, transcribes a fixture, and round-trips a synthesis (#216, #667).
| macOS arm64 | Linux x64 | Windows x64 | |
|---|---|---|---|
| Transcribe · audio language ID · VAD | CoreML / ANE | ONNX CPU | ONNX CPU |
| TTS — `en` `ru` `es` `fr` `it` `pt` | ✅ | ✅ | ✅ |
| TTS — `hi` `ja` `zh` and macOS system voices | ✅ | — | — |
| Mic capture and live dictation (`kesha record`) | ✅ | — | — |
| Speaker diarization (`--speakers`) | ✅ | — | — |
| Word-level timestamps (`words` in `--json`) | ✅ | ✅ | ✅ |
| Voice auto-routing from the text's language | ✅ | pass `--lang` | pass `--lang` |
Intel Macs get no published engine binary. Full matrix with maturity labels: docs/product-positioning.md.
Speech-to-text
kesha audio.ogg # transcribe (plain text)
kesha --format transcript audio.ogg # text + language/confidence
kesha --format json audio.ogg # full JSON with lang fields
kesha --json --timestamps audio.ogg # JSON with timestamped segments
kesha --itn audio.ogg # spelled-out numbers -> digits
kesha --toon audio.ogg # compact LLM-friendly TOON
kesha status # show installed backend info
kesha status --disk # + recursive cache disk usage
kesha status --json # machine-readable, for scriptsMultiple files get `head`-style headers; stdout is the transcript, stderr is errors — pipe-friendly:
$ kesha freedom.ogg tahiti.ogg
=== freedom.ogg ===
Свободу попугаям! Свободу!
=== tahiti.ogg ===
Таити, Таити! Не были мы ни в какой Таити! Нас и тут неплохо кормят.- Record from the mic (macOS): `kesha record --out hello.wav` writes microphone audio to a WAV file (`kesha hello.wav` transcribes it). macOS prompts for microphone access on first use — grant it under System Settings → Privacy & Security → Microphone if it was denied. On Linux/Windows or headless boxes, pass any existing audio file straight to `kesha` instead.
- Dictate straight to text (darwin-arm64): `kesha record --live` transcribes the mic as it captures and prints the transcript to stdout — no WAV in between, so it pipes (`kesha record --live | pbcopy`). To end after trailing silence, explicitly install VAD then opt in: `kesha install --vad && kesha record --live --auto-stop`. The defaults are 1,000 ms of silence after 250 ms of speech; tune them with `--auto-stop-silence-ms`, `--auto-stop-min-speech-ms`, and `--auto-stop-threshold`. Progress goes to stderr. Linux and Windows do not capture the microphone; pass an existing audio file to `kesha` to transcribe it. An interruption is recoverable: Ctrl-C (or SIGTERM) stops the session, still prints what you dictated, and exits 130/143, and the audio is spilled to a recovery WAV under `~/.cache/kesha/recordings/` — named on stderr when the session starts, deleted once the transcript has actually been delivered, kept if anything — a signal, a crash, a closed terminal, a dead pipe — got in the way first (#962).
- Long / silence-heavy audio: install VAD (`kesha install --vad`); Kesha auto-uses it past 120 s. Without VAD, long audio falls back to fixed ASR chunks. See docs/vad.md.
- Speaker diarization (darwin-arm64): `kesha install --diarize` (which installs VAD too), then `kesha --json --speakers meeting.m4a` stamps each segment with a `speaker` id. `--speakers` engages VAD windowing itself at any duration, so it cannot be combined with `--no-vad`. Linux/Windows return a clear "darwin-arm64 only" error (#199).
- Word-level timestamps (every platform): `kesha --json --timestamps audio.ogg` adds a `words` array to each segment — `{ "word": "email", "start": 0.72, "end": 1.12 }` — on the same file-relative clock as the segment, so a word always lies inside the segment carrying it. Read them off the decoder's own frame grid, so: times are quantised to 0.08 s, consecutive spans may overlap (each `end` is a per-word duration prediction, not the next word's `start`), `end >= start` rather than strictly greater, and punctuation stays attached to its word. The key is simply absent where a segment has none — any segment `--itn` rewrote, for one — so check the `transcribe.words` capability rather than expecting an empty array (#720).
- Text-language detection: JSON and TOON results include `textLanguage` with a language code, confidence, and its `source`. On macOS Kesha uses Apple `NLLanguageRecognizer`; elsewhere it uses the bundled `tinyld` fallback, whose confidence scale is different. This is separate from `audioLanguage`, which identifies the spoken audio when available.
- Written-form numbers: `--itn` rewrites what the model spells out — `"two hundred thirty two"` → `"232"`, `"five dollars and fifty cents"` → `"$5.50"`. Opt-in, every platform, timestamps untouched. English-only in practice; Russian and the rest pass through unchanged. Spoken punctuation names stay words (`"dot"`, `"comma"`, `"the period of growth"`) because Kesha transcribes speech rather than dictation — so `"example dot com"` keeps its words too (#822). A sentence `"and"` survives the number that follows it (`"cats and three dogs"` → `"cats and 3 dogs"`), while an `"and"` the number owns still joins it (`"three hundred and five"` → `"305"`) (#1000) — and no longer splits the number around it (`"two hundred and thirty two"` → `"232"`, not `"230 2"`) (#1006). A hyphenated number reads the same as the spaced form (`"twenty-five apples"` → `"25 apples"`), while a hyphen between ordinary words is left alone (`"well-known"`, `"state-of-the-art"`, `"twenty-something"`) (#1004).
Text-to-speech
Kesha speaks back in 9 languages. Kokoro runs natively through FluidAudio CoreML/ANE on Apple Silicon and through ONNX on Linux and Windows; Russian uses Vosk-TTS, while `macos-*` system voices need no model download. On macOS Kesha picks the voice from the text's own language; on Linux and Windows, state the language with `--lang ` (or the voice with `--voice `) — otherwise the engine default speaks.
kesha install --tts # English voices; sizes differ per platform — preview: kesha install --plan
kesha install --tts en ru # + Russian (+~890 MB, Vosk)
kesha say "Hello, world" > hello.wav
kesha say "Привет, мир" > privet.wav # auto-routes by language (macOS)
kesha say --lang ru "Привет, мир" > privet.wav # explicit — the Linux/Windows path
kesha say --voice ru-vosk-m02 "Голос в текст." > ru.wavOutput formats (`--format`, or inferred from the `--out` extension):
kesha say "Hello" --out hi.wav # WAV (default, uncompressed)
kesha say "Hello" --format ogg-opus --out hi.ogg # OGG/Opus — messenger voice notes
kesha say "Hello" --format flac --out hi.flac # FLAC — lossless, plays in every browser incl. Safari/iOS`kesha say --list-voices` lists what's installed. Voices, the full catalogue, macOS system voices, SSML, speaking rate (`--rate`, ``), Russian word stress, and Russian/English abbreviation handling are all in **docs/tts.md**.
Languages
Speech-to-text spans 25 languages and text-to-speech 9 — full tables with codes, flags, and per-platform availability in **docs/languages.md**. Audio language detection identifies 107 languages.
Performance
> Up to ~19x faster than Whisper on Apple Silicon (M2), ~2.5x faster on CPU
Compared against Whisper `large-v3-turbo`, all engines auto-detecting language:
Full per-file breakdown (Russian + English): BENCHMARK.md. The CPU figure is the ONNX engine on an M2's CPU cores; no x86 numbers are published yet.
Other install methods
All of these install the Bun CLI wrapper; engine + models still download explicitly via `kesha install`. (Nix is the exception — it currently builds only the engine from source; see below.)
- Homebrew — `brew install drakulavich/tap/kesha-voice-kit` · docs/homebrew.md
- Linux packages (`.deb`/`.rpm`, x64) — published on CLI releases, see docs/linux-packages.md
- Docker (GHCR image) — docs/docker.md
- Nix (`aarch64-darwin` / `x86_64-linux`) — builds the engine from source (`nix build github:drakulavich/kesha-voice-kit#kesha-engine`). The full `kesha` CLI via `nix run` / `nix profile install` is not yet available — it needs a maintainer with Nix to populate a build hash (#946). · docs/nix-install.md
- Shell completions + manpage — `kesha completions bash|zsh|fish` and `kesha manpage` print the packaged files to install wherever your shell expects them.
Integrations
- MCP server — `kesha mcp` exposes transcribe/synthesize/list tools to any MCP client (Claude, Cursor, Codex, Gemini). Setup: docs/mcp.md.
- OpenClaw — give your LLM agent ears. Install & config: docs/openclaw.md.
- Hermes Agent — local STT/TTS through Hermes command providers. Setup: docs/hermes.md.
- Raycast (macOS) — offline microphone dictation from the launcher: *Dictate to Clipboard* records with a live signal meter, auto-stops on silence, transcribes locally, and copies the text. Install from the Raycast Store · source: `raycast/`.
- Programmatic API — `@drakulavich/kesha-voice-kit/core` for use inside a Bun program. See docs/api.md.
More
- Architecture — runtime data flow, the models that ship, the CLI ↔ Rust engine boundary, model pinning, and where tests live.
- Use cases — copy-paste recipes (transcribe a meeting, speak from OpenClaw, run offline, move the cache).
- Product positioning — supported workflows, non-goals, maturity labels, platform matrix.
- Changelog — every release, with the behaviour changes spelled out.
- Diagnostics: `kesha doctor`, `kesha support-bundle` (redacted `.tar.gz` for issues), and `kesha logs` produce local, content-free diagnostics — see docs/diagnostic-logs.md. Every failure prints a stable `error [CODE]: …` line and a documented process exit code.
- Scripting & CI: `--json` (or `--toon`) for machine-readable output, `--include-errors` (with either) to get per-file failures on stdout alongside the results, `--quiet`/`-q` to silence progress, and `--no-color` (or `NO_COLOR=1`) for plain logs. Colors switch off automatically when `CI=true`.
- Privacy / Local Stats: Stats are off by default and fully local. Opt in with `kesha stats enable` to record content-free operational metrics in a local SQLite database — never networked, never storing audio, transcripts, text, or paths. Full commands & lifecycle: docs/local-stats.md.
Contributing
See CONTRIBUTING.md, the Roadmap (Now / Next / Later), and the Decision log (why platform/model choices were made — and reversed). Dev setup: `just dev-setup` (Bun, Rust, nextest, platform libs).
License
Made with 💛🩵 and 🥤 energy under MIT License
Frequently asked questions
What is kesha-voice-kit?
kesha-voice-kit is Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
How do I install kesha-voice-kit?
Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.
Is kesha-voice-kit open source?
Yes — it is hosted on GitHub at https://github.com/drakulavich/kesha-voice-kit and has 73 stars.
Related MCP tools
YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
Open-source cross-agent memory layer for coding agents via MCP. Compatible with Claude Code, Codex, Cursor, Windsurf, Gemini CLI, Antigravity, OpenClaw, Hermes Agent, Oh-my-Pi, Pi, Copilot, Kiro, OpenCode, and Trae.
Markdown knowledge graph — LSP for your editor, CLI + MCP memory for your AI agents
Natural voice conversations with Claude Code
MCP server that enables AI assistants to interact with Google Gemini CLI, leveraging Gemini's massive token window for large file analysis and codebase understanding
Run your own MCP server? See who uses it and what to fix.
Measure it with TrackMCP