AverVOX 0.5.7 is live. Free edition installs from PyPI; integrations are published for Hermes and OpenClaw.See integrations

Offline dictation for Linux

Speak into any Linux app.

Press a global hotkey, speak naturally, and let AverVOX insert the transcription wherever your cursor already is—from a browser and email client to VS Code and a terminal.

Faster Whisper on-deviceX11 / XWaylandNo browser extensionPiper TTS included

What it gives you

Dictation that follows your focus—not a special editor.

AverVOX keeps the interaction at the desktop layer, so the same hotkey works across the applications you already use.

System-wide text insertion

The transcript goes to the active window through supported desktop input and clipboard backends. There is no AverVOX-only document silo.

Voice activity detection

Press once, speak at your own pace, and let local VAD detect pauses and turns. Configure timing for your voice and environment.

Read selected text aloud

Highlight text in a compatible app and use a second global hotkey to hear it through the local Piper engine—or Kokoro in Pro.

Setup

From install to first sentence in three steps.

The repository installer handles the Linux system packages, isolated Python environment, voice model, desktop entry, and autostart setup.

Install the OSS edition

Clone the public repository and run its guided installer. The Python package is also published as both avrvx and the avervox alias.

Launch the tray app

Run avrvx. A desktop notification confirms that the global hotkeys are active; normal operation lives in the tray rather than a permanent main window.

Focus an app and speak

Place the cursor where text should appear, press Ctrl+Alt+Space, speak, and press again when needed. Tune pause timing in Settings if phrases are split too early or too late.

Built for real Linux desktops

Choose the insertion path that matches your session.

AverVOX supports xdotool, ydotool, and clipboard-based workflows. Test the exact desktop session and target apps you rely on before purchasing Pro.

X11

The most predictable path for global hotkeys, selection capture, and text insertion.

Supported

XWayland

Supported for the documented Mint and Ubuntu workflows, subject to the behavior of the target app.

Supported

Pure Wayland

May work with ydotool and wl-paste, but behavior varies by compositor and security policy.

Best-effort

Common questions

Linux dictation, without pretending every desktop is identical.

Which applications can receive dictation?
Most applications that accept normal keyboard or clipboard text can work. AverVOX is designed around the focused window rather than app-specific plugins, but unusual toolkits or protected input fields may behave differently.
Is speech recognition really offline?
Yes. Faster Whisper runs on the local machine. No raw audio or transcript is sent to AverVOX. Network traffic is used only when you intentionally use Converse with a configured LLM endpoint.
Does it need a GPU?
No GPU is required for the default CPU-optimized STT workflow. A modern multi-core CPU and at least 8 GB of RAM are recommended for comfortable latency.
What about Wayland?
X11 and XWayland are supported. Pure Wayland is best-effort because global hotkeys, selection access, and synthetic text insertion are compositor-specific. Test OSS before buying Pro.
Can I change the shortcuts?
Yes. Hotkeys, STT model, timing, insertion backend, and other behavior are configurable through Settings and the YAML configuration file.

Start with the workflow

Give your hands a break without changing applications.

Start with the free MIT-licensed edition. Pro adds Kokoro voices, a wake word, persistent sessions, LAN sharing, transcripts, and a daily dictation log.