computer (0.3.0)
Installation
pip install --index-url computerAbout this package
Low-latency speech-to-speech pipeline
"Computer!" - a low latency tool capable voice assistant
If you've ever watched Star Trek: The Next Generation and felt that you wanted your own Computer! - here you are. This started out as a fork of the huggingface/speech-to-speech project where every change and addition that has been made has had that goal in mind.
Notable functionality:
1. Wake word detection. The system only responds initially after hearing a configured wake word (e.g. "computer"). Uses openWakeWord with a custom-trained computer.onnx wake word model — training is complete, nothing pending there. If you want a different wake word, train your own ONNX model at openwakeword.com/train. During the conversation that follows no wake word is needed until after a configurable cooldown period.
2. Audio chimes. Plays a WAV chime on wake word detection (signals the user can speak) and when a server-side tool/search returns (signals the answer is incoming).
3. Barge-in (interrupt). Speak over the assistant mid-response to interrupt TTS output. The system cancels in-flight generation, clears pending audio, plays a wake chime, and processes your speech as new input without requiring the wake word.
4. Dedicated WebSocket client(s). A standalone computer-client CLI tool handles wake word detection, microphone input, and speaker output. One server can respond to multiple clients.
5. Can use multimodal LLM backend for STT. If the backing LLM is multimodal, like Gemma4 12B, no dedicated Speech-To-Text (STT) component is needed. This saves VRAM and helps allowing the full stack to run on regular consumer GPUs.
6. Speaker ID. With a [crew] user service configured, each finalized utterance is identified against enrolled crew voices (CAM++ embedder) and the transcript is prefixed with [name], so the LLM knows who is speaking. See Speaker ID.
Caveat: For obvious reasons this project does not include the specific chimes, the trained wake word model (computer.onnx), and voice cloning data you might want to use.
Example configurations
Four stacked configurations using different amounts of GPU VRAM. Every option runs VAD (Silero VAD v5, ~6 MB), so it is not listed separately. The LLM (Gemma4 12B) runs as a separate llama-server process; all STT and TTS models are shared across pipeline units.
Option 1: faster-whisper + Qwen3-TTS
| Component | VRAM |
|---|---|
| faster-whisper tiny.en | ~150 MB |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| Pipeline total | ~3.4 GB |
computer \
--stt faster-whisper --faster_whisper_stt_model_name tiny.en \
--tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
--mode realtime
Option 2: Parakeet TDT + Qwen3-TTS
| Component | VRAM |
|---|---|
| Parakeet TDT 0.6B | ~2 GB |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| Pipeline total | ~5.2 GB |
computer \
--stt parakeet-tdt \
--tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
--mode realtime
Option 3: Parakeet int8 ONNX + Qwen3-TTS
Parakeet TDT 0.6B v3 as int8 ONNX via onnx-asr (ONNX Runtime). Roughly a fifth of the PyTorch handler's VRAM with a hard-capped CUDA arena, sized to coexist with a local llama-server on the same GPU. Measured on the RTX A2000 12 GB reference server (llama-server ~7.9 GB resident): load +361 MB static, a 10 s turn peaks +533 MB, a 30 s turn +1.35 GB — and the arena never exceeds --parakeet_onnx_gpu_mem_limit_bytes (default 1.5 GB, kSameAsRequested), so nothing can grow unbounded. Trade-offs: final-turn-only transcription (no live partials) and int8 RTFx ~6-10, i.e. a 10 s utterance transcribes in ~1-1.7 s.
| Component | VRAM |
|---|---|
| parakeet-tdt-0.6b-v3 int8 ONNX | ~0.4 GB (arena capped at 1.5 GB) |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| Pipeline total | ~3.6 GB |
uv pip install -e '.[parakeet-onnx]' # onnx-asr[hub] + onnxruntime-gpu (Linux x86_64)
computer \
--stt parakeet-onnx --parakeet_onnx_device cuda \
--tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
--mode realtime
Option 4: Native LLM audio (no STT) + Qwen3-TTS FP16
Drops the STT model entirely. Audio is transcribed by the multimodal LLM itself. Frees enough VRAM to run Qwen3-TTS at full FP16 precision.
| Component | VRAM |
|---|---|
| Qwen3-TTS FP16 | ~3.6 GB |
| Pipeline total | ~3.6 GB |
computer \
--stt native-llm --llm_backend chat-completions \
--tts qwen3 --qwen3_tts_backend ggml \
--mode realtime
Caveat: with --stt native-llm, the text-based echo filter is not active. The assistant's own speech is transcribed by the LLM as empty text (there is no separate STT producing text), so the transcript-similarity echo filter never runs. Echo rejection then relies entirely on the VAD's RMS echo gate (compare mic energy against the speaker's output energy), which only protects while the speaker is actively playing. If you hear the assistant answering itself, enable an STT backend (e.g. faster-whisper) so the text echo filter engages as well.
Reference server configuration
The file server.sh in the repository root is the author's daily-driver configuration:
- LLM: Gemma4 12B QAT served by llama-server on another machine (
--llm_backend responses-api,--responses_api_base_url http://<llm-host>:11434). - STT: faster-whisper tiny.en (
--stt faster-whisper,--faster_whisper_stt_model_name tiny.en). - TTS: Qwen3-TTS with voice cloning to a Star Trek TNG computer voice (
--tts qwen3,--qwen3_tts_quant Q8_0,--qwen3_tts_ref_audio computer.wav). - System prompt: Star Trek TNG computer persona with tool-use instructions (
--init_chat_prompt). - MCP tools: Home Assistant and web search via two MCP servers (
--mcp-config mcp.json). - Two concurrent clients: (
--num_pipelines 2). - Wake word: Wake word model handled on the client side; server plays chimes on wake (
--wake_word_wake_chime here.wav) and after tool results (--wake_word_search_chime ready.wav).
Adjust --responses_api_base_url to point at your own llama-server, swap in your own --qwen3_tts_ref_audio and --qwen3_tts_instruct for voice cloning, and replace --init_chat_prompt with your desired persona.
Client configuration
A companion computer-client CLI handles microphone, speaker, and optional local wake word:
computer-client \
--host <server-host> --port 8766 \
--wake-word-model computer.onnx \
--wake-inactivity-timeout 10
Connect multiple clients at the same time: the server pool (--num_pipelines) determines how many can talk simultaneously. Extra clients are rejected when the pool is full.
Model Setup
The pipeline uses several models (STT, VAD, TTS). These are never auto-downloaded — all network access is blocked at startup. There are two ways to get them:
One-time download (recommended):
uv run python scripts/download_models.py
Downloads all default models (NLTK data, Silero VAD, Parakeet TDT, Qwen3-TTS) into local caches. After this completes, the pipeline runs fully offline.
Qwen3-TTS GGML Quantizations and CUDA Wheel Rebuild
server.sh uses Qwen/Qwen3-TTS-12Hz-1.7B-Base via faster-qwen3-tts ggml (qwentts.cpp):
| Quant | Talker | Tokenizer | TTS VRAM (weights + 896 MB KV) | nvidia-smi total (LLM 7658 + TTS + 546 uvicorn) |
|---|---|---|---|---|
Q8_0 (default) |
1765 MB | 278 MB | ~3.1 GB | 11439 / 12282 MiB (467 free) |
Q4_K_M |
1012 MB | 244 MB | ~2.24 GB | 10597 / 12282 MiB (1309 free) — saves ~842 MiB |
BF16/F32 exist but are larger; Q4_0 has no upstream Serveurperso/Qwen3-TTS-GGUF GGUF and is rejected (qwentts_cpp/models.py: _normalize_quant allows F32/BF16/Q8_0/Q4_K_M + aliases Q4/Q8).
The PyPI qwentts-cpp-python 0.3.1 wheel pins qwentts.cpp 7df559a (2026-07-17) which lacks the CUDA getrows kernel for GGML_TYPE_Q6_K (ggml/src/ggml-cuda/getrows.cu: case Q6_K) used by Q4_K_M. Running Q4_K_M on that wheel aborts: ggml_cuda_get_rows_switch_src0_type: unsupported src0 type: q6_K (see OOM.md and 12:23:47 crash after switching). Rebuild against main / a8a7716+ fixes it — the new GGML ships case Q6_K.
Rebuild and install a CUDA wheel (headless: RTX A2000 8.6, CUDA 13.3):
with-build-server ./scripts/build_qwentts_wheel.sh
# or pin: QWENTTS_REF=main ./scripts/build_qwentts_wheel.sh
# verify: python3 -c "from qwentts_cpp import QwenLibrary; print(QwenLibrary().version())"
# then:
XDG_RUNTIME_DIR=/run/user/$(id -u) systemctl --user restart computer-server
The script clones ServeurpersoCom/qwentts.cpp and andimarafioti/qwentts-cpp-python to /tmp (override with QWENTTS_CPP_SOURCE, QWENTTS_PY_SOURCE), runs scripts/build_native.py --backend cuda --clean with CMAKE_CUDA_ARCHITECTURES=86-real — the deployed A2000's arch only (override CUDA_ARCHITECTURES for other GPUs), builds with QWENTTS_CPP_WHEEL_BUILD_TAG=1cu128, copies the wheel to wheels/ (the pinned location), and installs it into .venv (venv pip if present, else uv pip; system pip otherwise). Requires cmake ninja-build patchelf + CUDA toolkit (nvcc). No fork is needed unless you must patch ggml-cuda itself — the upstream main already contains the Q6_K fix; pinning a newer QWENTTS_REF is sufficient.
The script also applies an idempotent ABI-v4 patch to the binding: upstream qwentts-cpp-python 0.3.1 speaks QT_ABI_VERSION 2, but the rebuilt native lib (a8a7716+) is ABI v4 — its qt_init_params gained max_batch/codec_chunk_sec and qt_tts_params dropped codec_chunk_sec/codec_left_context_sec. Without the patch, qt_init_default_params() writes codec_chunk_sec 4 bytes past the ctypes struct and corrupts the heap, causing random SIGSEGVs while the TTS model loads. Building a clean QWENTTS_PY_SOURCE checkout without this patch reproduces the crashes.
The rebuilt wheel is pinned. [tool.uv.sources] in pyproject.toml points qwentts-cpp-python at wheels/qwentts_cpp_python-0.3.1-1cu128-py3-none-linux_x86_64.whl, so uv sync / uv pip install always install the rebuilt wheel and can never regress to the broken PyPI 0.3.1. The wheels/ directory is gitignored (build artifact): after building, copy the wheel there — uv lock / uv sync then fail if the file is missing from a fresh checkout, which is the intended guard.
Empirical VRAM (empty / typical / worst + per-concurrent-user)
scripts/measure_vram.py polls nvidia-smi --query-gpu=memory.used,free,total at 50 ms and records static (post-load idle) vs peak (max used during qt_synthesize) — transient = peak - static, min free = headroom. Run on the headless RTX A2000 (12282 MiB) with llama-server 7658 MiB + uvicorn 546 MiB already resident (--num_pipelines 2).
Scenarios: empty (no synth), typical (~100 chars / ~7s audio), worst (~700 chars / ~1536 tokens max, DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS in src/computer/TTS/qwen3_tts_handler.py:45). concurrent=1 is the single-user peak; concurrent=2 runs 2 parallel Qwen3TTSHandler.process calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = peak_2 - peak_1.
Regenerate on the server (requires nvidia-smi + cached Qwen/Qwen3-TTS GGUFs):
# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md # paste into table below
Last measured 2026-08-27 headless RTX A2000 12GB (12282 MiB, driver 580.173, qwentts.cpp a8a7716 (2026-08-07), torch 2.11+cu130) via scripts/measure_vram.py:1 (50 ms nvidia-smi poll, LLM 8256 baseline on this host). Both Q4_K_M and Q8_0 CustomVoice measured after CUDA wheel rebuild with q6_K fix (ggml/src/ggml-cuda/getrows.cu: case Q6_K):
| Quant | Scenario | Concurrent | Static used |
Peak used |
Transient | Min free |
Total | Notes |
|---|---|---|---|---|---|---|---|---|
Q8_0 |
empty | 1 | 11641 | 11641 | 0 | 265 | 12282 | LLM 8256 + TTS 1765+896+overhead |
Q8_0 |
typical | 1 | 11637 | 11655 | 18 | 251 | 12282 | 102 chars / ~5.4s audio |
Q8_0 |
typical | 2 | 11651 | 11663 | 12 | 243 | 12282 | 2nd concurrent +8 MB |
Q8_0 |
worst | 1 | 11659 | 11685 | 26 | 221 | 12282 | 625 chars / ~42s audio |
Q8_0 |
worst | 2 | 11641 | 11693 | 52 | 213 | 12282 | 2nd concurrent +8 MB (isolated run) |
Q4_K_M |
empty | 1 | 10799 | 10799 | 0 | 1107 | 12282 | LLM 8256 + TTS 1012+896+overhead |
Q4_K_M |
typical | 1 | 10799 | 10813 | 14 | 1093 | 12282 | 102 chars / ~5.4s audio |
Q4_K_M |
typical | 2 | 10809 | 10821 | 12 | 1085 | 12282 | 2nd concurrent +8 MB |
Q4_K_M |
worst | 1 | 10817 | 10843 | 26 | 1063 | 12282 | 625 chars / ~44s audio |
Q4_K_M |
worst | 2 | 10823 | 10851 | 28 | 1055 | 12282 | 2nd concurrent +8 MB |
Q4_K_M saves ~842 MiB vs Q8_0 (peak worst 10851 vs 11693). Raw JSON in artifacts/vram/2026-08-27_*.json. Static is shared weights — extra concurrent pipelines do not add VRAM when idle, only when their qt_synthesize overlap (marginal +8 MB for 2nd pipeline on both quants). Server was stopped during measurement (baseline 8256 = LLM alone; prior 11439/10597 totals in earlier logs included 546 uvicorn + server overhead).
Empirical VRAM (empty / typical / worst + per-concurrent-user)
scripts/measure_vram.py polls nvidia-smi --query-gpu=memory.used,free,total at 50 ms and records static (post-load idle) vs peak (max used during qt_synthesize) — transient = peak - static, min free = headroom. Run on the headless RTX A2000 (12282 MiB) with llama-server 7658 MiB + uvicorn 546 MiB already resident (--num_pipelines 2).
Scenarios: empty (no synth), typical (~100 chars / ~7s audio), worst (~700 chars / ~1536 tokens max, DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS in src/computer/TTS/qwen3_tts_handler.py:45). concurrent=1 is the single-user peak; concurrent=2 runs 2 parallel Qwen3TTSHandler.process calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = peak_2 - peak_1.
Regenerate on the server (requires nvidia-smi + cached Qwen/Qwen3-TTS GGUFs):
# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md # paste into table below
Last measured 2026-08-27 headless RTX A2000 12GB (12282 MiB, driver 580.173, qwentts.cpp a8a7716 (2026-08-07), torch 2.11+cu130) via scripts/measure_vram.py:1 (50 ms nvidia-smi poll, LLM 8256 baseline on this host). Both Q4_K_M and Q8_0 CustomVoice measured after CUDA wheel rebuild with q6_K fix (ggml/src/ggml-cuda/getrows.cu: case Q6_K):
| Quant | Scenario | Concurrent | Static used |
Peak used |
Transient | Min free |
Total | Notes |
|---|---|---|---|---|---|---|---|---|
Q8_0 |
empty | 1 | 11641 | 11641 | 0 | 265 | 12282 | LLM 8256 + TTS 1765+896+overhead |
Q8_0 |
typical | 1 | 11637 | 11655 | 18 | 251 | 12282 | 102 chars / ~5.4s audio |
Q8_0 |
typical | 2 | 11651 | 11663 | 12 | 243 | 12282 | 2nd concurrent +8 MB |
Q8_0 |
worst | 1 | 11659 | 11685 | 26 | 221 | 12282 | 625 chars / ~42s audio |
Q8_0 |
worst | 2 | 11641 | 11693 | 52 | 213 | 12282 | 2nd concurrent +8 MB (isolated run) |
Q4_K_M |
empty | 1 | 10799 | 10799 | 0 | 1107 | 12282 | LLM 8256 + TTS 1012+896+overhead |
Q4_K_M |
typical | 1 | 10799 | 10813 | 14 | 1093 | 12282 | 102 chars / ~5.4s audio |
Q4_K_M |
typical | 2 | 10809 | 10821 | 12 | 1085 | 12282 | 2nd concurrent +8 MB |
Q4_K_M |
worst | 1 | 10817 | 10843 | 26 | 1063 | 12282 | 625 chars / ~44s audio |
Q4_K_M |
worst | 2 | 10823 | 10851 | 28 | 1055 | 12282 | 2nd concurrent +8 MB |
Q4_K_M saves ~842 MiB vs Q8_0 (peak worst 10851 vs 11693). Raw JSON in artifacts/vram/2026-08-27_*.json. Static is shared weights — extra concurrent pipelines do not add VRAM when idle, only when their qt_synthesize overlap (marginal +8 MB for 2nd pipeline on both quants). Server was stopped during measurement (baseline 8256 = LLM alone; prior 11439/10597 totals in earlier logs included 546 uvicorn + server overhead).
Realtime API
Realtime mode streams audio over a WebSocket using the OpenAI Realtime protocol, with live transcription and low-latency turn-taking. The server exposes /v1/realtime, and any OpenAI Realtime-compatible client can connect. A dedicated computer-client CLI is included.
See src/computer/api/openai_realtime/README.md for the supported event set, architecture, and design details.
Speaker ID (crew user system)
The pipeline can tag each finalized utterance with a speaker id (or guest),
sourced from the separate crew user service. See
docs/specs/2026-09-02-crew-user-system-design.md.
Enable it by passing --users_crew_url <crew-base-url> (plus the other
--users_* flags; see server.sh.example). Leave --users_crew_url empty to
disable speaker ID entirely.
When crew serves per-user prompts, the realtime assistant starts each session
with the device owner's crew-composed prompt instead of --init_chat_prompt.
Whenever a known speaker is identified — including mid-session — the session
prompt is swapped to that user's prompt. Per-user prompts are cached with a TTL
controlled by --users_session_prompt_ttl_s (default 300 s); unknown speakers
and crew outages fall back to the configured default --init_chat_prompt.
- The CAM++ embedder runs on the GPU when CUDA is available (falls back to CPU): ~62 MiB VRAM total cost (27 MiB weights + CUDA context) and ~60 ms per 2 s segment warm, versus 68–878 ms on CPU — the difference between resolving the speaker inside the join budget and dropping the tag.
- Identified speakers reach the LLM as a
[name]transcript prefix; the voice system prompt teaches the model to treat it as the speaker's name, not part of the utterance. - Requires the
paraformerextra on the server:uv pip install -e '.[paraformer]'. - Add users and enroll reference clips via the crew CLI (
crew create-user,crew enroll clip.wav --user <id>). - Tune
--users_match_thresholdwithscripts/smoke_speaker_id.pyon your real voices (start ~0.80; set it just below the within-person cosine). - Enrollment runs through
POST /internal/users/enrollon this server;--users_enroll_tokenprotects it. - Clients identify themselves:
computer-clientsends a stable--device-id(generated once, persisted at~/.config/computer/device_id) anddevice_typein the WebSocket query string; the server keeps the connected-device registry behindGET /internal/devices. - Guided enrollment and verification run server-side through
POST /internal/devices/{device_id}/capture(modesenroll/verify): the server diverts the device's mic into a VAD-gated capture (the assistant pipeline and the client's wake-word detector pause for its duration) and returns the speaker vector plus abest_usermatch. Crew's admin UI drives the flow, so someone standing at the device can enroll or verify without uploading a file.
License
Citations (from the upstream Speech-to-Speech project)
If you use this pipeline, please also cite the component models you run. The defaults are:
Gemma4
@misc{gemma4-2026,
author = {Google DeepMind},
title = {Gemma 4},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/google/gemma-4-12B-it}}
}
Silero VAD
@misc{SileroVAD,
author = {Silero Team},
title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},
year = {2021},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/snakers4/silero-vad}},
email = {hello@silero.ai}
}
Parakeet TDT
@misc{parakeet-tdt,
author = {NVIDIA},
title = {Parakeet TDT 0.6B v3},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3}}
}
Qwen3-TTS
@misc{qwen3-tts,
author = {Qwen Team},
title = {Qwen3-TTS},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice}}
}
Citations for optional backends such as Kokoro, Pocket TTS, ChatTTS, Whisper variants, Paraformer, and MMS live in the respective component READMEs.