AverVOX 0.5.7 is live. Free edition installs from PyPI; integrations are published for Hermes and OpenClaw.See integrations

Private, system-wide speech for Linux

Dictate anywhere. Talk to any LLM.

AverVOX is a local speech layer for Linux. Speak into any focused application, read selected text aloud, hold hands-free conversations with OpenAI-compatible models, and give agents the same private STT/TTS engine.

Linux Mint & UbuntuLocal STT/TTSNo telemetryNo subscription
AVERVOX // LOCAL SPEECH LAYER
Any focused Linux application

Turn rough meeting notes into a clear follow-up email, preserving the action items and deadlines.

Ctrl+Alt+Space · listening locally
Ollama · LM Studio · any compatible endpoint
Summarize the trade-offs and give me a recommendation.
The fastest path is to keep speech local and point AverVOX at the model endpoint you already use…
Streaming reply sentence by sentence
One speech engine · multiple agents
OpenClaw / Hermeshost application
AverVOX daemonlocal STT + TTS
$ avrvx --install-integration openclaw
speech provider verified
Works with your stack
OllamaLM StudiovLLMLocalAIOpenAI-compatible APIsHermesOpenClaw

Private by architecture

Your microphone talks to your machine—not a speech vendor.

Speech recognition, synthesis, voice activity detection, and desktop text insertion run locally. Conversation text goes only to the LLM endpoint you choose.

Microphonelocal audio
Faster Whisperlocal transcription
Your LLMlocal, LAN, or remote
Piper / Kokorolocal speech

AverVOX collects no usage analytics, crash reports, or product telemetry. A remote LLM receives conversation text only when you configure one.

Measured through real integrations

Fast enough for natural turn-taking—even on older hardware.

The warm daemon keeps speech models loaded, avoiding process startup and model-load cost on every agent response.

Hermes · warm0.32sFirst audio
OpenClaw · warm0.36sFirst audio
Hermes · whole reply0.77sComplete playback file
OpenClaw · whole reply0.91sComplete playback file

Reference benchmark, not a universal guarantee. Measured on a 2017 laptop with Piper en_US-lessac-high, warm daemon enabled, and each host’s native provider interface. Your hardware, voice, and text length will change the result.

Agent-ready

Replace cloud speech APIs with one local installation.

Install host-native packages for Hermes and OpenClaw, or call the bridge CLI directly from your own tool.

PyPITTS + STTStreaming

Hermes Agent

A Python package that plugs AverVOX into Hermes through its native speech-provider interface.

npm + ClawHubTTSStreaming

OpenClaw

A native plugin published for normal OpenClaw discovery and installation, with automatic daemon fallback.

CLI + Unix socketJSONPCM streaming

Your application

Synthesize, transcribe, discover capabilities, cancel work, and stream PCM without embedding speech models.

Free core, focused upgrade

Start open source. Upgrade when voice becomes part of your day.

AverVOX OSS contains the complete core workflow. Pro adds higher-quality voice, hands-free activation, persistence, LAN sharing, and management tools.

CapabilityOSSPro
System-wide dictation, selected-text TTS, and Converse
Faster Whisper, Piper, bridge CLI, warm daemon, streaming, barge-in
Kokoro voices and playback-speed control
Wake word, persistent sessions, and per-profile personas
LAN voice server, dashboard, transcripts, and dictation logs
PriceFree · MIT$39 one time

Buy with confidence

Test the core workflow before purchasing Pro.

The OSS edition is the compatibility check: install it, verify your desktop and audio path, connect your endpoint, and make sure voice fits your workflow.

Install OSS

Use the guided installer or the published Python package.

Test all three modes

Dictation, selected-text speech, and Converse use the same underlying desktop and audio integrations as Pro.

Confirm your session

X11 and XWayland are supported. Pure Wayland remains best-effort because global hotkeys and text insertion vary by compositor.

Questions before you install

A clear path from first hotkey to daily workflow.

Does AverVOX send my microphone audio to the cloud?
No. Speech recognition, synthesis, and voice activity detection run locally. In Converse mode, the transcription is sent only to the LLM endpoint you configure; that endpoint may be local, on your LAN, or remote.
Which model servers work?
Any service exposing a compatible OpenAI-style chat endpoint can work, including Ollama, LM Studio, vLLM, LocalAI, and remote inference providers. You can keep multiple profiles and switch between them.
Is Pro a subscription?
No. Pro is a one-time $39 purchase. You may use the purchased version indefinitely, and the license includes every release in the version 1 line, up to but not including version 2.0.
Can I evaluate compatibility before buying?
Yes. Install AverVOX OSS and test the core workflows first. Pro has no trial or post-delivery refund, so the free edition is the recommended pre-purchase evaluation path.
Does AverVOX work on pure Wayland?
Pure Wayland is best-effort. X11 and XWayland are the supported paths because global hotkeys, clipboard access, and synthetic text insertion differ among Wayland compositors.
Can another application use AverVOX without launching the desktop UI?
Yes. The bridge CLI supports file-based synthesis, transcription, capability discovery, and streaming PCM. The optional warm daemon exposes the same work over a private Unix socket.

Start with the workflow

Put a private speech layer between you and every app.

Install the MIT-licensed core today. Move to Pro when you want Kokoro, wake-word activation, persistent memory, LAN sharing, and local history.