Composable architecture
Use the desktop tray, command line, bridge daemon, host-native integration, or a combination. Speech remains a reusable service instead of one chat app feature.
Build a private Linux voice workflow
AverVOX is the reusable speech layer around the model, application, or agent you already prefer. Keep voice local, swap LLM endpoints freely, and use the same hotkeys throughout Linux.
What it gives you
AverVOX separates microphone and speaker workflows from the model service, giving you a stable interface while the AI stack changes.
Use the desktop tray, command line, bridge daemon, host-native integration, or a combination. Speech remains a reusable service instead of one chat app feature.
Raw microphone audio, STT, TTS, voice activity detection, and text insertion stay on the machine. You choose where conversation text is sent.
Run the LLM on the same laptop, a more capable workstation, a private server, or a provider—without replacing the Linux speech workflow.
Setup
Start with a supported Linux desktop, then connect the speech layer to the model endpoint that matches your privacy and performance goals.
Use Ollama, LM Studio, vLLM, LocalAI, another OpenAI-compatible server, or a remote inference provider.
AverVOX supplies Faster Whisper, Piper, VAD, hotkeys, desktop insertion, streaming TTS, and the conversation state machine.
Use the bridge CLI and daemon for scripts, or install the Hermes and OpenClaw packages when an agent should reuse the same speech engine.
Use the right machine for each job
AverVOX can run responsive local speech on the Linux desktop while a different LAN machine handles a larger model.
Hotkeys, microphone, Faster Whisper, text insertion, and TTS stay close to the user.
Only conversation text and streaming reply text cross to the endpoint you configured.
A larger workstation or server can run the model without becoming your everyday desktop.
Common questions
Start with the workflow
Use the free core as the stable speech layer. Upgrade only when the Pro automation, persistence, and voice-quality features earn a place in your daily workflow.