Metadata-Version: 2.4
Name: computer
Version: 0.3.0
Summary: Low-latency speech-to-speech pipeline
Author: troed
License-Expression: Apache-2.0
Project-URL: Homepage, https://git.sync.wtf/starfleet/computer
Project-URL: Repository, https://git.sync.wtf/starfleet/computer
Project-URL: Issues, https://git.sync.wtf/starfleet/computer/issues
Keywords: speech-to-speech,voice-agents,speech-recognition,text-to-speech,openai
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: mcp>=2.0
Requires-Dist: fastapi>=0.115.0
Requires-Dist: httpx>=0.28.0
Requires-Dist: nltk==3.9.4
Requires-Dist: numpy<2.4.4,>=1.26.0; platform_system == "Darwin"
Requires-Dist: numpy>=1.26.0; platform_system != "Darwin"
Requires-Dist: openai==2.28.0
Requires-Dist: openwakeword<0.5,>=0.4.0
Requires-Dist: pillow>=10.0.0
Requires-Dist: pydantic>=2.0
Requires-Dist: rich>=13.0
Requires-Dist: scipy>=1.10.0
Requires-Dist: sounddevice==0.5.3; platform_system == "Darwin"
Requires-Dist: sounddevice>=0.5.0; platform_system != "Darwin"
Requires-Dist: soundfile>=0.13.0; platform_system == "Darwin"
Requires-Dist: torch==2.11.0; platform_system == "Darwin"
Requires-Dist: torch>=2.4.0; platform_system != "Darwin"
Requires-Dist: torchaudio==2.11.0; platform_system == "Darwin"
Requires-Dist: torchaudio>=2.4.0; platform_system != "Darwin"
Requires-Dist: transformers==5.6.2; platform_system == "Darwin"
Requires-Dist: transformers>=4.57.0; platform_system != "Darwin"
Requires-Dist: uvicorn>=0.30.0
Requires-Dist: websockets>=12.0
Requires-Dist: nano-parakeet>=0.2.0; platform_system != "Darwin"
Requires-Dist: faster-qwen3-tts[ggml]>=0.3.2; platform_system != "Darwin" and platform_system != "Windows"
Requires-Dist: faster-qwen3-tts>=0.3.2; platform_system == "Windows"
Requires-Dist: qwentts-cpp-python==0.3.1; sys_platform != "darwin" and sys_platform != "win32"
Requires-Dist: lingua-language-detector>=2.0.2
Requires-Dist: miniaudio==1.61; platform_system == "Darwin"
Requires-Dist: mlx==0.31.1; platform_system == "Darwin"
Requires-Dist: mlx-audio==0.4.2; platform_system == "Darwin"
Requires-Dist: mlx-lm==0.31.1; platform_system == "Darwin"
Requires-Dist: mlx-metal==0.31.1; platform_system == "Darwin"
Requires-Dist: misaki>=0.9.4; platform_system == "Darwin"
Requires-Dist: espeakng-loader>=0.2.4; platform_system == "Darwin"
Requires-Dist: spacy>=3.8.4; platform_system == "Darwin"
Requires-Dist: phonemizer-fork>=3.3.2; platform_system == "Darwin"
Provides-Extra: chattts
Requires-Dist: ChatTTS>=0.1.1; extra == "chattts"
Provides-Extra: facebook-mms
Requires-Dist: transformers>=4.57.0; extra == "facebook-mms"
Provides-Extra: faster-whisper
Requires-Dist: faster-whisper>=1.0.3; extra == "faster-whisper"
Provides-Extra: kokoro
Requires-Dist: kokoro>=0.9.2; platform_system != "Darwin" and extra == "kokoro"
Provides-Extra: language-detection
Requires-Dist: lingua-language-detector>=2.0.2; extra == "language-detection"
Provides-Extra: mlx-lm
Requires-Dist: mlx-lm==0.31.1; platform_system == "Darwin" and extra == "mlx-lm"
Requires-Dist: mlx-vlm==0.4.1; platform_system == "Darwin" and extra == "mlx-lm"
Provides-Extra: parakeet-onnx
Requires-Dist: onnx-asr[hub]>=0.5.0; extra == "parakeet-onnx"
Requires-Dist: onnxruntime-gpu[cuda,cudnn]>=1.27; (platform_system == "Linux" and platform_machine == "x86_64" and python_version >= "3.11") and extra == "parakeet-onnx"
Requires-Dist: onnxruntime>=1.20; (platform_system != "Linux" or platform_machine != "x86_64" or python_version < "3.11") and extra == "parakeet-onnx"
Provides-Extra: paraformer
Requires-Dist: funasr>=1.1.6; extra == "paraformer"
Requires-Dist: modelscope>=1.17.1; extra == "paraformer"
Requires-Dist: onnxruntime<1.24; python_version < "3.11" and extra == "paraformer"
Provides-Extra: pocket
Requires-Dist: pocket-tts>=0.1.0; extra == "pocket"
Provides-Extra: omnivoice
Requires-Dist: omnivoice>=0.2.1; extra == "omnivoice"
Provides-Extra: webrtc
Requires-Dist: aiortc>=1.9.0; extra == "webrtc"
Provides-Extra: websocket
Requires-Dist: websockets>=12.0; extra == "websocket"
Provides-Extra: whisper-mlx
Requires-Dist: lightning-whisper-mlx>=0.0.10; platform_system == "Darwin" and extra == "whisper-mlx"
Provides-Extra: client
Requires-Dist: sounddevice>=0.5.0; extra == "client"
Requires-Dist: websockets>=12.0; extra == "client"
Dynamic: license-file

<p align="center">
  <img src="logo.png" alt="Computer logo" width="220">
</p>

<p align="center">
  <a href="https://video.troed.se/w/vwTvn6qsVtQGsSW8TJMBDj">
    <img src="https://video.troed.se/static/thumbnails/vwTvn6qsVtQGsSW8TJMBDj.jpg" alt="Computer! demo video" width="560">
  </a>
</p>

# *"Computer!"* - a low latency tool capable voice assistant

If you've ever watched Star Trek: The Next Generation and felt that you wanted your own *Computer!* - here you are. This started out as a fork of the [huggingface/speech-to-speech](https://github.com/huggingface/speech-to-speech) project where every change and addition that has been made has had that goal in mind.

Notable functionality:

**1. Wake word detection.** The system only responds initially after hearing a configured wake word (e.g. "computer"). Uses [openWakeWord](https://github.com/dscripka/openWakeWord) with a custom-trained `computer.onnx` wake word model — training is complete, nothing pending there. If you want a different wake word, train your own ONNX model at [openwakeword.com/train](https://openwakeword.com/train). During the conversation that follows no wake word is needed until after a configurable cooldown period.

**2. Audio chimes.** Plays a WAV chime on wake word detection (signals the user can speak) and when a server-side tool/search returns (signals the answer is incoming).

**3. Barge-in (interrupt).** Speak over the assistant mid-response to interrupt TTS output. The system cancels in-flight generation, clears pending audio, plays a wake chime, and processes your speech as new input without requiring the wake word.

**4. Dedicated WebSocket client(s).** A standalone `computer-client` CLI tool handles wake word detection, microphone input, and speaker output. One server can respond to multiple clients.

**5. Can use multimodal LLM backend for STT.** If the backing LLM is multimodal, like Gemma4 12B, no dedicated Speech-To-Text (STT) component is needed. This saves VRAM and helps allowing the full stack to run on regular consumer GPUs.

**6. Speaker ID.** With a [crew] user service configured, each finalized utterance is identified against enrolled crew voices (CAM++ embedder) and the transcript is prefixed with `[name]`, so the LLM knows who is speaking. See [Speaker ID](#speaker-id-crew-user-system).

Caveat: For obvious reasons this project does not include the specific chimes, the trained wake word model (`computer.onnx`), and voice cloning data you might want to use.

## Example configurations

Four stacked configurations using different amounts of GPU VRAM. Every option runs VAD (Silero VAD v5, ~6 MB), so it is not listed separately. The LLM (Gemma4 12B) runs as a separate llama-server process; all STT and TTS models are shared across pipeline units.

### Option 1: faster-whisper + Qwen3-TTS

| Component | VRAM |
|---|---|
| faster-whisper tiny.en | ~150 MB |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| **Pipeline total** | **~3.4 GB** |

```bash
computer \
    --stt faster-whisper --faster_whisper_stt_model_name tiny.en \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime
```

### Option 2: Parakeet TDT + Qwen3-TTS

| Component | VRAM |
|---|---|
| Parakeet TDT 0.6B | ~2 GB |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| **Pipeline total** | **~5.2 GB** |

```bash
computer \
    --stt parakeet-tdt \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime
```

### Option 3: Parakeet int8 ONNX + Qwen3-TTS

Parakeet TDT 0.6B v3 as int8 ONNX via [onnx-asr](https://github.com/istupakov/onnx-asr) (ONNX Runtime). Roughly a fifth of the PyTorch handler's VRAM with a hard-capped CUDA arena, sized to coexist with a local llama-server on the same GPU. Measured on the RTX A2000 12 GB reference server (llama-server ~7.9 GB resident): load +361 MB static, a 10 s turn peaks +533 MB, a 30 s turn +1.35 GB — and the arena never exceeds `--parakeet_onnx_gpu_mem_limit_bytes` (default 1.5 GB, `kSameAsRequested`), so nothing can grow unbounded. Trade-offs: final-turn-only transcription (no live partials) and int8 RTFx ~6-10, i.e. a 10 s utterance transcribes in ~1-1.7 s.

| Component | VRAM |
|---|---|
| parakeet-tdt-0.6b-v3 int8 ONNX | ~0.4 GB (arena capped at 1.5 GB) |
| Qwen3-TTS Q8_0 | ~3.2 GB |
| **Pipeline total** | **~3.6 GB** |

```bash
uv pip install -e '.[parakeet-onnx]'   # onnx-asr[hub] + onnxruntime-gpu (Linux x86_64)

computer \
    --stt parakeet-onnx --parakeet_onnx_device cuda \
    --tts qwen3 --qwen3_tts_backend ggml --qwen3_tts_quant Q8_0 \
    --mode realtime
```

### Option 4: Native LLM audio (no STT) + Qwen3-TTS FP16

Drops the STT model entirely. Audio is transcribed by the multimodal LLM itself. Frees enough VRAM to run Qwen3-TTS at full FP16 precision.

| Component | VRAM |
|---|---|
| Qwen3-TTS FP16 | ~3.6 GB |
| **Pipeline total** | **~3.6 GB** |

```bash
computer \
    --stt native-llm --llm_backend chat-completions \
    --tts qwen3 --qwen3_tts_backend ggml \
    --mode realtime
```

Caveat: with `--stt native-llm`, the text-based echo filter is **not** active. The assistant's own speech is transcribed by the LLM as empty text (there is no separate STT producing text), so the transcript-similarity echo filter never runs. Echo rejection then relies entirely on the VAD's RMS echo gate (compare mic energy against the speaker's output energy), which only protects while the speaker is actively playing. If you hear the assistant answering itself, enable an STT backend (e.g. faster-whisper) so the text echo filter engages as well.

### Reference server configuration

The file `server.sh` in the repository root is the author's daily-driver configuration:

- **LLM**: Gemma4 12B QAT served by llama-server on another machine (`--llm_backend responses-api`, `--responses_api_base_url http://<llm-host>:11434`).
- **STT**: faster-whisper tiny.en (`--stt faster-whisper`, `--faster_whisper_stt_model_name tiny.en`).
- **TTS**: Qwen3-TTS with voice cloning to a Star Trek TNG computer voice (`--tts qwen3`, `--qwen3_tts_quant Q8_0`, `--qwen3_tts_ref_audio computer.wav`).
- **System prompt**: Star Trek TNG computer persona with tool-use instructions (`--init_chat_prompt`).
- **MCP tools**: Home Assistant and web search via two MCP servers (`--mcp-config mcp.json`).
- **Two concurrent clients**: (`--num_pipelines 2`).
- **Wake word**: Wake word model handled on the client side; server plays chimes on wake (`--wake_word_wake_chime here.wav`) and after tool results (`--wake_word_search_chime ready.wav`).

Adjust `--responses_api_base_url` to point at your own llama-server, swap in your own `--qwen3_tts_ref_audio` and `--qwen3_tts_instruct` for voice cloning, and replace `--init_chat_prompt` with your desired persona.

### Client configuration

A companion `computer-client` CLI handles microphone, speaker, and optional local wake word:

```bash
computer-client \
    --host <server-host> --port 8766 \
    --wake-word-model computer.onnx \
    --wake-inactivity-timeout 10
```

Connect multiple clients at the same time: the server pool (`--num_pipelines`) determines how many can talk simultaneously. Extra clients are rejected when the pool is full.

### Model Setup

The pipeline uses several models (STT, VAD, TTS). These are **never auto-downloaded** — all network access is blocked at startup. There are two ways to get them:

**One-time download (recommended):**

```bash
uv run python scripts/download_models.py
```

Downloads all default models (NLTK data, Silero VAD, Parakeet TDT, Qwen3-TTS) into local caches. After this completes, the pipeline runs fully offline.

### Qwen3-TTS GGML Quantizations and CUDA Wheel Rebuild

`server.sh` uses `Qwen/Qwen3-TTS-12Hz-1.7B-Base` via `faster-qwen3-tts` `ggml` (`qwentts.cpp`):

| Quant | Talker | Tokenizer | TTS VRAM (weights + 896 MB KV) | nvidia-smi total (LLM 7658 + TTS + 546 uvicorn) |
|-------|--------|-----------|--------------------------------|------------------------------------------------|
| `Q8_0` (default) | 1765 MB | 278 MB | ~3.1 GB | 11439 / 12282 MiB (467 free) |
| `Q4_K_M` | 1012 MB | 244 MB | ~2.24 GB | 10597 / 12282 MiB (1309 free) — saves ~842 MiB |

`BF16/F32` exist but are larger; `Q4_0` has no upstream `Serveurperso/Qwen3-TTS-GGUF` GGUF and is rejected (`qwentts_cpp/models.py: _normalize_quant` allows `F32/BF16/Q8_0/Q4_K_M` + aliases `Q4/Q8`).

The PyPI `qwentts-cpp-python 0.3.1` wheel pins `qwentts.cpp 7df559a (2026-07-17)` which lacks the CUDA `getrows` kernel for `GGML_TYPE_Q6_K` (`ggml/src/ggml-cuda/getrows.cu: case Q6_K`) used by `Q4_K_M`. Running `Q4_K_M` on that wheel aborts: `ggml_cuda_get_rows_switch_src0_type: unsupported src0 type: q6_K` (see `OOM.md` and `12:23:47` crash after switching). Rebuild against `main` / `a8a7716+` fixes it — the new GGML ships `case Q6_K`.

Rebuild and install a CUDA wheel (headless: RTX A2000 8.6, CUDA 13.3):

```bash
with-build-server ./scripts/build_qwentts_wheel.sh
# or pin: QWENTTS_REF=main ./scripts/build_qwentts_wheel.sh
# verify:  python3 -c "from qwentts_cpp import QwenLibrary; print(QwenLibrary().version())"
# then:
XDG_RUNTIME_DIR=/run/user/$(id -u) systemctl --user restart computer-server
```

The script clones `ServeurpersoCom/qwentts.cpp` and `andimarafioti/qwentts-cpp-python` to `/tmp` (override with `QWENTTS_CPP_SOURCE`, `QWENTTS_PY_SOURCE`), runs `scripts/build_native.py --backend cuda --clean` with `CMAKE_CUDA_ARCHITECTURES=86-real` — the deployed A2000's arch only (override `CUDA_ARCHITECTURES` for other GPUs), builds with `QWENTTS_CPP_WHEEL_BUILD_TAG=1cu128`, copies the wheel to `wheels/` (the pinned location), and installs it into `.venv` (venv `pip` if present, else `uv pip`; system `pip` otherwise). Requires `cmake ninja-build patchelf` + CUDA toolkit (`nvcc`). No fork is needed unless you must patch `ggml-cuda` itself — the upstream `main` already contains the `Q6_K` fix; pinning a newer `QWENTTS_REF` is sufficient.

The script also applies an idempotent ABI-v4 patch to the binding: upstream `qwentts-cpp-python` 0.3.1 speaks `QT_ABI_VERSION 2`, but the rebuilt native lib (a8a7716+) is ABI v4 — its `qt_init_params` gained `max_batch`/`codec_chunk_sec` and `qt_tts_params` dropped `codec_chunk_sec`/`codec_left_context_sec`. Without the patch, `qt_init_default_params()` writes `codec_chunk_sec` 4 bytes past the ctypes struct and corrupts the heap, causing random SIGSEGVs while the TTS model loads. Building a clean `QWENTTS_PY_SOURCE` checkout without this patch reproduces the crashes.

**The rebuilt wheel is pinned.** `[tool.uv.sources]` in `pyproject.toml` points `qwentts-cpp-python` at `wheels/qwentts_cpp_python-0.3.1-1cu128-py3-none-linux_x86_64.whl`, so `uv sync` / `uv pip install` always install the rebuilt wheel and can never regress to the broken PyPI 0.3.1. The `wheels/` directory is gitignored (build artifact): after building, copy the wheel there — `uv lock` / `uv sync` then fail if the file is missing from a fresh checkout, which is the intended guard.

### Empirical VRAM (empty / typical / worst + per-concurrent-user)

`scripts/measure_vram.py` polls `nvidia-smi --query-gpu=memory.used,free,total` at 50 ms and records **static** (post-load idle) vs **peak** (max `used` during `qt_synthesize`) — `transient = peak - static`, `min free` = headroom. Run on the headless RTX A2000 (12282 MiB) with `llama-server` 7658 MiB + `uvicorn` 546 MiB already resident (`--num_pipelines 2`).

Scenarios: `empty` (no synth), `typical` (~100 chars / ~7s audio), `worst` (~700 chars / ~1536 tokens max, `DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS` in `src/computer/TTS/qwen3_tts_handler.py:45`). `concurrent=1` is the single-user peak; `concurrent=2` runs 2 parallel `Qwen3TTSHandler.process` calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = `peak_2 - peak_1`.

Regenerate on the server (requires `nvidia-smi` + cached `Qwen/Qwen3-TTS` GGUFs):

```bash
# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md   # paste into table below
```

Last measured `2026-08-27` headless `RTX A2000 12GB` (12282 MiB, driver `580.173`, `qwentts.cpp a8a7716 (2026-08-07)`, `torch 2.11+cu130`) via `scripts/measure_vram.py:1` (50 ms `nvidia-smi` poll, `LLM 8256` baseline on this host). Both `Q4_K_M` and `Q8_0` `CustomVoice` measured after CUDA wheel rebuild with `q6_K` fix (`ggml/src/ggml-cuda/getrows.cu: case Q6_K`):

| Quant | Scenario | Concurrent | Static `used` | Peak `used` | Transient | Min `free` | Total | Notes |
|-------|----------|------------|---------------|-------------|-----------|------------|-------|-------|
| `Q8_0` | empty | 1 | 11641 | 11641 | 0 | 265 | 12282 | LLM 8256 + TTS 1765+896+overhead |
| `Q8_0` | typical | 1 | 11637 | 11655 | 18 | 251 | 12282 | 102 chars / ~5.4s audio |
| `Q8_0` | typical | 2 | 11651 | 11663 | 12 | 243 | 12282 | 2nd concurrent +8 MB |
| `Q8_0` | worst | 1 | 11659 | 11685 | 26 | 221 | 12282 | 625 chars / ~42s audio |
| `Q8_0` | worst | 2 | 11641 | 11693 | 52 | 213 | 12282 | 2nd concurrent +8 MB (isolated run) |
| `Q4_K_M` | empty | 1 | 10799 | 10799 | 0 | 1107 | 12282 | LLM 8256 + TTS 1012+896+overhead |
| `Q4_K_M` | typical | 1 | 10799 | 10813 | 14 | 1093 | 12282 | 102 chars / ~5.4s audio |
| `Q4_K_M` | typical | 2 | 10809 | 10821 | 12 | 1085 | 12282 | 2nd concurrent +8 MB |
| `Q4_K_M` | worst | 1 | 10817 | 10843 | 26 | 1063 | 12282 | 625 chars / ~44s audio |
| `Q4_K_M` | worst | 2 | 10823 | 10851 | 28 | 1055 | 12282 | 2nd concurrent +8 MB |

`Q4_K_M` saves ~842 MiB vs `Q8_0` (peak worst `10851` vs `11693`). Raw JSON in `artifacts/vram/2026-08-27_*.json`. Static is shared weights — extra concurrent pipelines do **not** add VRAM when idle, only when their `qt_synthesize` overlap (marginal `+8 MB` for 2nd pipeline on both quants). Server was stopped during measurement (`baseline 8256` = LLM alone; prior `11439`/`10597` totals in earlier logs included `546 uvicorn` + server overhead).

### Empirical VRAM (empty / typical / worst + per-concurrent-user)

`scripts/measure_vram.py` polls `nvidia-smi --query-gpu=memory.used,free,total` at 50 ms and records **static** (post-load idle) vs **peak** (max `used` during `qt_synthesize`) — `transient = peak - static`, `min free` = headroom. Run on the headless RTX A2000 (12282 MiB) with `llama-server` 7658 MiB + `uvicorn` 546 MiB already resident (`--num_pipelines 2`).

Scenarios: `empty` (no synth), `typical` (~100 chars / ~7s audio), `worst` (~700 chars / ~1536 tokens max, `DEFAULT_QWEN3_TTS_MAX_NEW_TOKENS` in `src/computer/TTS/qwen3_tts_handler.py:45`). `concurrent=1` is the single-user peak; `concurrent=2` runs 2 parallel `Qwen3TTSHandler.process` calls sharing the single loaded weights (static shared, scratch is per concurrent synth) — marginal = `peak_2 - peak_1`.

Regenerate on the server (requires `nvidia-smi` + cached `Qwen/Qwen3-TTS` GGUFs):

```bash
# headless, after ./deploy.sh
uv run python scripts/measure_vram.py --quant q4_k_m,q8_0 --scenario empty,typical,worst --concurrent 1,2
cat artifacts/vram/README_table.md   # paste into table below
```

Last measured `2026-08-27` headless `RTX A2000 12GB` (12282 MiB, driver `580.173`, `qwentts.cpp a8a7716 (2026-08-07)`, `torch 2.11+cu130`) via `scripts/measure_vram.py:1` (50 ms `nvidia-smi` poll, `LLM 8256` baseline on this host). Both `Q4_K_M` and `Q8_0` `CustomVoice` measured after CUDA wheel rebuild with `q6_K` fix (`ggml/src/ggml-cuda/getrows.cu: case Q6_K`):

| Quant | Scenario | Concurrent | Static `used` | Peak `used` | Transient | Min `free` | Total | Notes |
|-------|----------|------------|---------------|-------------|-----------|------------|-------|-------|
| `Q8_0` | empty | 1 | 11641 | 11641 | 0 | 265 | 12282 | LLM 8256 + TTS 1765+896+overhead |
| `Q8_0` | typical | 1 | 11637 | 11655 | 18 | 251 | 12282 | 102 chars / ~5.4s audio |
| `Q8_0` | typical | 2 | 11651 | 11663 | 12 | 243 | 12282 | 2nd concurrent +8 MB |
| `Q8_0` | worst | 1 | 11659 | 11685 | 26 | 221 | 12282 | 625 chars / ~42s audio |
| `Q8_0` | worst | 2 | 11641 | 11693 | 52 | 213 | 12282 | 2nd concurrent +8 MB (isolated run) |
| `Q4_K_M` | empty | 1 | 10799 | 10799 | 0 | 1107 | 12282 | LLM 8256 + TTS 1012+896+overhead |
| `Q4_K_M` | typical | 1 | 10799 | 10813 | 14 | 1093 | 12282 | 102 chars / ~5.4s audio |
| `Q4_K_M` | typical | 2 | 10809 | 10821 | 12 | 1085 | 12282 | 2nd concurrent +8 MB |
| `Q4_K_M` | worst | 1 | 10817 | 10843 | 26 | 1063 | 12282 | 625 chars / ~44s audio |
| `Q4_K_M` | worst | 2 | 10823 | 10851 | 28 | 1055 | 12282 | 2nd concurrent +8 MB |

`Q4_K_M` saves ~842 MiB vs `Q8_0` (peak worst `10851` vs `11693`). Raw JSON in `artifacts/vram/2026-08-27_*.json`. Static is shared weights — extra concurrent pipelines do **not** add VRAM when idle, only when their `qt_synthesize` overlap (marginal `+8 MB` for 2nd pipeline on both quants). Server was stopped during measurement (`baseline 8256` = LLM alone; prior `11439`/`10597` totals in earlier logs included `546 uvicorn` + server overhead).

## Realtime API

Realtime mode streams audio over a WebSocket using the OpenAI Realtime protocol, with live transcription and low-latency turn-taking. The server exposes `/v1/realtime`, and any OpenAI Realtime-compatible client can connect. A dedicated `computer-client` CLI is included.

See [src/computer/api/openai_realtime/README.md](./src/computer/api/openai_realtime/README.md) for the supported event set, architecture, and design details.

## Speaker ID (crew user system)

The pipeline can tag each finalized utterance with a speaker id (or `guest`),
sourced from the separate `crew` user service. See
`docs/specs/2026-09-02-crew-user-system-design.md`.

Enable it by passing `--users_crew_url <crew-base-url>` (plus the other
`--users_*` flags; see `server.sh.example`). Leave `--users_crew_url` empty to
disable speaker ID entirely.

When crew serves per-user prompts, the realtime assistant starts each session
with the device owner's crew-composed prompt instead of `--init_chat_prompt`.
Whenever a known speaker is identified — including mid-session — the session
prompt is swapped to that user's prompt. Per-user prompts are cached with a TTL
controlled by `--users_session_prompt_ttl_s` (default 300 s); unknown speakers
and crew outages fall back to the configured default `--init_chat_prompt`.

- The CAM++ embedder runs on the GPU when CUDA is available (falls back to CPU): **~62 MiB VRAM** total cost (27 MiB weights + CUDA context) and ~60 ms per 2 s segment warm, versus 68–878 ms on CPU — the difference between resolving the speaker inside the join budget and dropping the tag.
- Identified speakers reach the LLM as a `[name]` transcript prefix; the voice system prompt teaches the model to treat it as the speaker's name, not part of the utterance.
- Requires the `paraformer` extra on the server: `uv pip install -e '.[paraformer]'`.
- Add users and enroll reference clips via the crew CLI (`crew create-user`,
  `crew enroll clip.wav --user <id>`).
- Tune `--users_match_threshold` with `scripts/smoke_speaker_id.py` on your real
  voices (start ~0.80; set it just below the within-person cosine).
- Enrollment runs through `POST /internal/users/enroll` on this server;
  `--users_enroll_token` protects it.
- Clients identify themselves: `computer-client` sends a stable `--device-id`
  (generated once, persisted at `~/.config/computer/device_id`) and
  `device_type` in the WebSocket query string; the server keeps the
  connected-device registry behind `GET /internal/devices`.
- Guided enrollment and verification run server-side through
  `POST /internal/devices/{device_id}/capture` (modes `enroll`/`verify`):
  the server diverts the device's mic into a VAD-gated capture (the assistant
  pipeline and the client's wake-word detector pause for its duration) and
  returns the speaker vector plus a `best_user` match. Crew's admin UI drives
  the flow, so someone standing at the device can enroll or verify without
  uploading a file.

## License

[![Apache 2.0](https://img.shields.io/badge/license-Apache%202.0-blue)](./LICENSE)

## Citations (from the upstream Speech-to-Speech project)

If you use this pipeline, please also cite the component models you run. The defaults are:

### Gemma4

```bibtex
@misc{gemma4-2026,
  author = {Google DeepMind},
  title = {Gemma 4},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/google/gemma-4-12B-it}}
}
```

### Silero VAD

```bibtex
@misc{SileroVAD,
  author = {Silero Team},
  title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},
  year = {2021},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/snakers4/silero-vad}},
  email = {hello@silero.ai}
}
```

### Parakeet TDT

```bibtex
@misc{parakeet-tdt,
  author = {NVIDIA},
  title = {Parakeet TDT 0.6B v3},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3}}
}
```

### Qwen3-TTS

```bibtex
@misc{qwen3-tts,
  author = {Qwen Team},
  title = {Qwen3-TTS},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice}}
}
```

Citations for optional backends such as Kokoro, Pocket TTS, ChatTTS, Whisper variants, Paraformer, and MMS live in the respective [component READMEs](./src/computer).
