| Audio capture and voice activity detection | Choose libraries, devices, chunking, silence rules | Configured VAD and recording flow |
| System-wide activation | Implement global hotkeys and desktop-session behavior | Three configurable global modes |
| Text insertion and selection | Handle X11/Wayland tools and app differences | Supported insertion and selection backends |
| Streaming model replies into TTS | Split text, strip markup, manage buffering and gaps | Sentence-level streaming and cleanup |
| Conversation state | Build listening, processing, speaking, re-arm, exit logic | Conversation loop, HUD, silence and goodbye handling |
| Interruption and echo control | Implement cancellation, microphone muting, re-arm timing | Barge-in and documented headphone safeguards |
| Multiple endpoints and models | Create configuration and switching UI | Named profiles and tray switching |
| Use from other agents | Write another adapter for each host | Bridge CLI, daemon, Hermes and OpenClaw packages |
| Installation and desktop lifecycle | Package dependencies, autostart, tray, logs, updates | Guided OSS installer and packaged Pro distribution |
| Ongoing ownership | Full flexibility and full maintenance burden | Public core plus a one-time Pro option |