computer (0.3.0)

Published 2026-09-07 22:54:45 +02:00 by troed

Installation

pip install --index-url  computer

About this package

Low-latency speech-to-speech pipeline

Computer logo

Computer! demo video

"Computer!" - a low latency tool capable voice assistant

If you've ever watched Star Trek: The Next Generation and felt that you wanted your own Computer! - here you are. This started out as a fork of the huggingface/speech-to-speech project where every change and addition that has been made has had that goal in mind.

Notable functionality:

1. Wake word detection. The system only responds initially after hearing a configured wake word (e.g. "computer"). Uses openWakeWord with a custom-trained computer.onnx wake word model — training is complete, nothing pending there. If you want a different wake word, train your own ONNX model at openwakeword.com/train. During the conversation that follows no wake word is needed until after a configurable cooldown period.

2. Audio chimes. Plays a WAV chime on wake word detection (signals the user can speak) and when a server-side tool/search returns (signals the answer is incoming).

3. Barge-in (interrupt). Speak over the assistant mid-response to interrupt TTS output. The system cancels in-flight generation, clears pending audio, plays a wake chime, and processes your speech as new input without requiring the wake word.

4. Dedicated WebSocket client(s). A standalone computer-client CLI tool handles wake word detection, microphone input, and speaker output. One server can respond to multiple clients.

5. Can use multimodal LLM backend for STT. If the backing LLM is multimodal, like Gemma4 12B, no dedicated Speech-To-Text (STT) component is needed. This saves VRAM and helps allowing the full stack to run on regular consumer GPUs.

6. Speaker ID. With a [crew] user service configured, each finalized utterance is identified against enrolled crew voices (CAM++ embedder) and the transcript is prefixed with [name], so the LLM knows who is speaking. See Speaker ID.

Caveat: For obvious reasons this project does not include the specific chimes, the trained wake word model (computer.onnx), and voice cloning data you might want to use.

Example configurations

Four stacked configurations using different amounts of GPU VRAM. Every option runs VAD (Silero VAD v5, ~6 MB), so it is not listed separately. The LLM (Gemma4 12B) runs as a separate llama-server process; all STT and TTS models are shared across pipeline units.

Option 1: faster-whisper + Qwen3-TTS

Component VRAM
faster-whisper tiny.en ~150 MB
Qwen3-TTS Q8_0 ~3.2 GB
Pipeline total ~3.4 GB
computer \
    --stt faster-whisper --faster_whisper_stt_model_name tiny.en \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime

Option 2: Parakeet TDT + Qwen3-TTS

Component VRAM
Parakeet TDT 0.6B ~2 GB
Qwen3-TTS Q8_0 ~3.2 GB
Pipeline total ~5.2 GB
computer \
    --stt parakeet-tdt \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime

Option 3: Parakeet int8 ONNX + Qwen3-TTS

Parakeet TDT 0.6B v3 as int8 ONNX via onnx-asr (ONNX Runtime). Roughly a fifth of the PyTorch handler's VRAM with a hard-capped CUDA arena, sized to coexist with a local llama-server on the same GPU. Measured on the RTX A2000 12 GB reference server (llama-server ~7.9 GB resident): load +361 MB static, a 10 s turn peaks +533 MB, a 30 s turn +1.35 GB — and the arena never exceeds --parakeet_onnx_gpu_mem_limit_bytes (default 1.5 GB, kSameAsRequested), so nothing can grow unbounded. Trade-offs: final-turn-only transcription (no live partials) and int8 RTFx ~6-10, i.e. a 10 s utterance transcribes in ~1-1.7 s.

Component VRAM
parakeet-tdt-0.6b-v3 int8 ONNX ~0.4 GB (arena capped at 1.5 GB)
Qwen3-TTS Q8_0 ~3.2 GB
Pipeline total ~3.6 GB
uv pip install -e '.[parakeet-onnx]'   # onnx-asr[hub] + onnxruntime-gpu (Linux x86_64)

computer \
    --stt parakeet-onnx --parakeet_onnx_device cuda \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime

Option 4: Native LLM audio (no STT) + Qwen3-TTS FP16

Drops the STT model entirely. Audio is transcribed by the multimodal LLM itself. Frees enough VRAM to run Qwen3-TTS at full FP16 precision.

Component VRAM
Qwen3-TTS FP16 ~3.6 GB
Pipeline total ~3.6 GB
computer \
    --stt native-llm --llm_backend chat-completions \
    --tts qwen3 --qwen3_tts_backend ggml \
    --mode realtime

Caveat: with --stt native-llm, the text-based echo filter is not active. The assistant's own speech is transcribed by the LLM as empty text (there is no separate STT producing text), so the transcript-similarity echo filter never runs. Echo rejection then relies entirely on the VAD's RMS echo gate (compare mic energy against the speaker's output energy), which only protects while the speaker is actively playing. If you hear the assistant answering itself, enable an STT backend (e.g. faster-whisper) so the text echo filter engages as well.

Reference server configuration

The file server.sh in the repository root is the author's daily-driver configuration:

  • LLM: Gemma4 12B QAT served by llama-server on another machine (--llm_backend responses-api, --responses_api_base_url http://<llm-host>:11434).
  • STT: faster-whisper tiny.en (--stt faster-whisper, --faster_whisper_stt_model_name tiny.en).
  • TTS: Qwen3-TTS with voice cloning to a Star Trek TNG computer voice (--tts qwen3, --qwen3_tts_quant Q8_0, --qwen3_tts_ref_audio computer.wav).
  • System prompt: Star Trek TNG computer persona with tool-use instructions (--init_chat_prompt).
  • MCP tools: Home Assistant and web search via two MCP servers (--mcp-config mcp.json).
  • Two concurrent clients: (--num_pipelines 2).
  • Wake word: Wake word model handled on the client side; server plays chimes on wake (--wake_word_wake_chime here.wav) and after tool results (--wake_word_search_chime ready.wav).

Adjust --responses_api_base_url to point at your own llama-server, swap in your own --qwen3_tts_ref_audio and --qwen3_tts_instruct for voice cloning, and replace --init_chat_prompt with your desired persona.

Client configuration

A companion computer-client CLI handles microphone, speaker, and optional local wake word:

computer-client \
    --host <server-host> --port 8766 \
    --wake-word-model computer.onnx \
    --wake-inactivity-timeout 10

Connect multiple clients at the same time: the server pool (--num_pipelines) determines how many can talk simultaneously. Extra clients are rejected when the pool is full.

Model Setup

The pipeline uses several models (STT, VAD, TTS). These are never auto-downloaded — all network access is blocked at startup. There are two ways to get them:

One-time download (recommended):

uv run python scripts/download_models.py

Downloads all default models (NLTK data, Silero VAD, Parakeet TDT, Qwen3-TTS) into local caches. After this completes, the pipeline runs fully offline.

Qwen3-TTS GGML Quantizations and CUDA Wheel Rebuild

server.sh uses Qwen/Qwen3-TTS-12Hz-1.7B-Base via faster-qwen3-tts ggml (qwentts.cpp):

Quant Talker Tokenizer TTS VRAM (weights + 896 MB KV) nvidia-smi total (LLM 7658 + TTS + 546 uvicorn)
Q8_0 (default) 1765 MB 278 MB ~3.1 GB 11439 / 12282 MiB (467 free)
Q4_K_M 1012 MB 244 MB ~2.24 GB 10597 / 12282 MiB (1309 free) — saves ~842 MiB

BF16/F32 exist but are larger; Q4_0 has no upstream Serveurperso/Qwen3-TTS-GGUF GGUF and is rejected (qwentts_cpp/models.py: _normalize_quant allows F32/BF16/Q8_0/Q4_K_M + aliases Q4/Q8).

The PyPI qwentts-cpp-python 0.3.1 wheel pins qwentts.cpp 7df559a (2026-07-17) which lacks the CUDA getrows kernel for GGML_TYPE_Q6_K (ggml/src/ggml-cuda/getrows.cu: case Q6_K) used by Q4_K_M. Running Q4_K_M on that wheel aborts: ggml_cuda_get_rows_switch_src0_type: unsupported src0 type: q6_K (see OOM.md and 12:23:47 crash after switching). Rebuild against main / a8a7716+ fixes it — the new GGML ships case Q6_K.

Rebuild and install a CUDA wheel (headless: RTX A2000 8.6, CUDA 13.3):

with-build-server ./scripts/build_qwentts_wheel.sh
# or pin: QWENTTS_REF=main ./scripts/build_qwentts_wheel.sh
# verify:  python3 -c "from qwentts_cpp import QwenLibrary; print(QwenLibrary().version())"
# then:
XDG_RUNTIME_DIR=/run/user/$(id -u) systemctl --user restart computer-server

The script clones ServeurpersoCom/qwentts.cpp and andimarafioti/qwentts-cpp-python to /tmp (override with QWENTTS_CPP_SOURCE, QWENTTS_PY_SOURCE), runs scripts/build_native.py --backend cuda --clean with CMAKE_CUDA_ARCHITECTURES=86-real — the deployed A2000's arch only (override CUDA_ARCHITECTURES for other GPUs), builds with QWENTTS_CPP_WHEEL_BUILD_TAG=1cu128, copies the wheel to wheels/ (the pinned location), and installs it into .venv (venv pip if present, else uv pip; system pip otherwise). Requires cmake ninja-build patchelf + CUDA toolkit (nvcc). No fork is needed unless you must patch ggml-cuda itself — the upstream main already contains the Q6_K fix; pinning a newer QWENTTS_REF is sufficient.

The script also applies an idempotent ABI-v4 patch to the binding: upstream qwentts-cpp-python 0.3.1 speaks QT_ABI_VERSION 2, but the rebuilt native lib (a8a7716+) is ABI v4 — its qt_init_params gained max_batch/codec_chunk_sec and qt_tts_params dropped codec_chunk_sec/codec_left_context_sec. Without the patch, qt_init_default_params() writes codec_chunk_sec 4 bytes past the ctypes struct and corrupts the heap, causing random SIGSEGVs while the TTS model loads. Building a clean QWENTTS_PY_SOURCE checkout without this patch reproduces the crashes.

The rebuilt wheel is pinned. [tool.uv.sources] in pyproject.toml points qwentts-cpp-python at wheels/qwentts_cpp_python-0.3.1-1cu128-py3-none-linux_x86_64.whl, so uv sync / uv pip install always install the rebuilt wheel and can never regress to the broken PyPI 0.3.1. The wheels/ directory is gitignored (build artifact): after building, copy the wheel there — uv lock / uv sync then fail if the file is missing from a fresh checkout, which is the intended guard.

Empirical VRAM (empty / typical / worst + per-concurrent-user)

scripts/measure_vram.py polls nvidia-smi --query-gpu=memory.used,free,total at 50 ms and records static (post-load idle) vs peak (max used during qt_synthesize) — transient = peak - static, min free = headroom. Run on the headless RTX A2000 (12282 MiB) with llama-server 7658 MiB + uvicorn 546 MiB already resident (--num_pipelines 2).

Scenarios: empty (no synth), typical (~100 chars / ~7s audio), worst (~700 chars / ~1536 tokens max, DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS in src/computer/TTS/qwen3_tts_handler.py:45). concurrent=1 is the single-user peak; concurrent=2 runs 2 parallel Qwen3TTSHandler.process calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = peak_2 - peak_1.

Regenerate on the server (requires nvidia-smi + cached Qwen/Qwen3-TTS GGUFs):

# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md   # paste into table below

Last measured 2026-08-27 headless RTX A2000 12GB (12282 MiB, driver 580.173, qwentts.cpp a8a7716 (2026-08-07), torch 2.11+cu130) via scripts/measure_vram.py:1 (50 ms nvidia-smi poll, LLM 8256 baseline on this host). Both Q4_K_M and Q8_0 CustomVoice measured after CUDA wheel rebuild with q6_K fix (ggml/src/ggml-cuda/getrows.cu: case Q6_K):

Quant Scenario Concurrent Static used Peak used Transient Min free Total Notes
Q8_0 empty 1 11641 11641 0 265 12282 LLM 8256 + TTS 1765+896+overhead
Q8_0 typical 1 11637 11655 18 251 12282 102 chars / ~5.4s audio
Q8_0 typical 2 11651 11663 12 243 12282 2nd concurrent +8 MB
Q8_0 worst 1 11659 11685 26 221 12282 625 chars / ~42s audio
Q8_0 worst 2 11641 11693 52 213 12282 2nd concurrent +8 MB (isolated run)
Q4_K_M empty 1 10799 10799 0 1107 12282 LLM 8256 + TTS 1012+896+overhead
Q4_K_M typical 1 10799 10813 14 1093 12282 102 chars / ~5.4s audio
Q4_K_M typical 2 10809 10821 12 1085 12282 2nd concurrent +8 MB
Q4_K_M worst 1 10817 10843 26 1063 12282 625 chars / ~44s audio
Q4_K_M worst 2 10823 10851 28 1055 12282 2nd concurrent +8 MB

Q4_K_M saves ~842 MiB vs Q8_0 (peak worst 10851 vs 11693). Raw JSON in artifacts/vram/2026-08-27_*.json. Static is shared weights — extra concurrent pipelines do not add VRAM when idle, only when their qt_synthesize overlap (marginal +8 MB for 2nd pipeline on both quants). Server was stopped during measurement (baseline 8256 = LLM alone; prior 11439/10597 totals in earlier logs included 546 uvicorn + server overhead).

Empirical VRAM (empty / typical / worst + per-concurrent-user)

scripts/measure_vram.py polls nvidia-smi --query-gpu=memory.used,free,total at 50 ms and records static (post-load idle) vs peak (max used during qt_synthesize) — transient = peak - static, min free = headroom. Run on the headless RTX A2000 (12282 MiB) with llama-server 7658 MiB + uvicorn 546 MiB already resident (--num_pipelines 2).

Scenarios: empty (no synth), typical (~100 chars / ~7s audio), worst (~700 chars / ~1536 tokens max, DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS in src/computer/TTS/qwen3_tts_handler.py:45). concurrent=1 is the single-user peak; concurrent=2 runs 2 parallel Qwen3TTSHandler.process calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = peak_2 - peak_1.

Regenerate on the server (requires nvidia-smi + cached Qwen/Qwen3-TTS GGUFs):

# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md   # paste into table below

Last measured 2026-08-27 headless RTX A2000 12GB (12282 MiB, driver 580.173, qwentts.cpp a8a7716 (2026-08-07), torch 2.11+cu130) via scripts/measure_vram.py:1 (50 ms nvidia-smi poll, LLM 8256 baseline on this host). Both Q4_K_M and Q8_0 CustomVoice measured after CUDA wheel rebuild with q6_K fix (ggml/src/ggml-cuda/getrows.cu: case Q6_K):

Quant Scenario Concurrent Static used Peak used Transient Min free Total Notes
Q8_0 empty 1 11641 11641 0 265 12282 LLM 8256 + TTS 1765+896+overhead
Q8_0 typical 1 11637 11655 18 251 12282 102 chars / ~5.4s audio
Q8_0 typical 2 11651 11663 12 243 12282 2nd concurrent +8 MB
Q8_0 worst 1 11659 11685 26 221 12282 625 chars / ~42s audio
Q8_0 worst 2 11641 11693 52 213 12282 2nd concurrent +8 MB (isolated run)
Q4_K_M empty 1 10799 10799 0 1107 12282 LLM 8256 + TTS 1012+896+overhead
Q4_K_M typical 1 10799 10813 14 1093 12282 102 chars / ~5.4s audio
Q4_K_M typical 2 10809 10821 12 1085 12282 2nd concurrent +8 MB
Q4_K_M worst 1 10817 10843 26 1063 12282 625 chars / ~44s audio
Q4_K_M worst 2 10823 10851 28 1055 12282 2nd concurrent +8 MB

Q4_K_M saves ~842 MiB vs Q8_0 (peak worst 10851 vs 11693). Raw JSON in artifacts/vram/2026-08-27_*.json. Static is shared weights — extra concurrent pipelines do not add VRAM when idle, only when their qt_synthesize overlap (marginal +8 MB for 2nd pipeline on both quants). Server was stopped during measurement (baseline 8256 = LLM alone; prior 11439/10597 totals in earlier logs included 546 uvicorn + server overhead).

Realtime API

Realtime mode streams audio over a WebSocket using the OpenAI Realtime protocol, with live transcription and low-latency turn-taking. The server exposes /v1/realtime, and any OpenAI Realtime-compatible client can connect. A dedicated computer-client CLI is included.

See src/computer/api/openai_realtime/README.md for the supported event set, architecture, and design details.

Speaker ID (crew user system)

The pipeline can tag each finalized utterance with a speaker id (or guest), sourced from the separate crew user service. See docs/specs/2026-09-02-crew-user-system-design.md.

Enable it by passing --users_crew_url <crew-base-url> (plus the other --users_* flags; see server.sh.example). Leave --users_crew_url empty to disable speaker ID entirely.

When crew serves per-user prompts, the realtime assistant starts each session with the device owner's crew-composed prompt instead of --init_chat_prompt. Whenever a known speaker is identified — including mid-session — the session prompt is swapped to that user's prompt. Per-user prompts are cached with a TTL controlled by --users_session_prompt_ttl_s (default 300 s); unknown speakers and crew outages fall back to the configured default --init_chat_prompt.

  • The CAM++ embedder runs on the GPU when CUDA is available (falls back to CPU): ~62 MiB VRAM total cost (27 MiB weights + CUDA context) and ~60 ms per 2 s segment warm, versus 68–878 ms on CPU — the difference between resolving the speaker inside the join budget and dropping the tag.
  • Identified speakers reach the LLM as a [name] transcript prefix; the voice system prompt teaches the model to treat it as the speaker's name, not part of the utterance.
  • Requires the paraformer extra on the server: uv pip install -e '.[paraformer]'.
  • Add users and enroll reference clips via the crew CLI (crew create-user, crew enroll clip.wav --user <id>).
  • Tune --users_match_threshold with scripts/smoke_speaker_id.py on your real voices (start ~0.80; set it just below the within-person cosine).
  • Enrollment runs through POST /internal/users/enroll on this server; --users_enroll_token protects it.
  • Clients identify themselves: computer-client sends a stable --device-id (generated once, persisted at ~/.config/computer/device_id) and device_type in the WebSocket query string; the server keeps the connected-device registry behind GET /internal/devices.
  • Guided enrollment and verification run server-side through POST /internal/devices/{device_id}/capture (modes enroll/verify): the server diverts the device's mic into a VAD-gated capture (the assistant pipeline and the client's wake-word detector pause for its duration) and returns the speaker vector plus a best_user match. Crew's admin UI drives the flow, so someone standing at the device can enroll or verify without uploading a file.

License

Apache 2.0

Citations (from the upstream Speech-to-Speech project)

If you use this pipeline, please also cite the component models you run. The defaults are:

Gemma4

@misc{gemma4-2026,
  author = {Google DeepMind},
  title = {Gemma 4},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/google/gemma-4-12B-it}}
}

Silero VAD

@misc{SileroVAD,
  author = {Silero Team},
  title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},
  year = {2021},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/snakers4/silero-vad}},
  email = {hello@silero.ai}
}

Parakeet TDT

@misc{parakeet-tdt,
  author = {NVIDIA},
  title = {Parakeet TDT 0.6B v3},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3}}
}

Qwen3-TTS

@misc{qwen3-tts,
  author = {Qwen Team},
  title = {Qwen3-TTS},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice}}
}

Citations for optional backends such as Kokoro, Pocket TTS, ChatTTS, Whisper variants, Paraformer, and MMS live in the respective component READMEs.

Requirements

Requires Python: >=3.10
Details
PyPI
2026-09-07 22:54:45 +02:00
18
troed
762 KiB
Assets (2)
Versions (1) View all
0.3.0 2026-09-07