Integrate crew speaker-ID pipeline #11

Merged
troed merged 23 commits from devel into main 2026-09-03 08:48:28 +02:00
Owner

Integrates the crew speaker-ID pipeline (spec: docs/specs/2026-09-02-crew-user-system-design.md). When enabled, every final transcription is tagged with the speaker's crew user id, resolved against the separate starfleet-crew service. When disabled (the default), behavior is unchanged.

What's in

  • speaker: str | None = None on VADAudio / Transcription / GenerateResponseRequest / TranscriptionCompletedEvent (default None → fully backwards compatible)
  • New computer/users package:
    • CrewClient — httpx client for the crew service (registry, health, close)
    • UserRegistry — periodically-pulled user-vector snapshot with atomic swap, 3-state tagging (known / guest / stale), keep-stale on fetch failure
    • SpeakerEmbedder — lazy FunASR CAM++ (192-dim)
    • speaker_tag helpers (GUEST, format_user_turn)
    • SpeakerIDHandler — pipeline handler resolving speakers in parallel with STT
  • speaker threaded through the STT layer: output_for_queue + TranscriptionNotifier, including the speculative-revision path
  • Pipeline wiring in s2s_pipeline.py: --users_* CLI args, singleton init, handler insertion
  • POST /internal/users/enroll on the realtime service (Bearer token, raw PCM s16le 16 kHz mono, 192-dim vector, 401 / 422 / 503)
  • server.sh.example commented --users_* block, scripts/smoke_speaker_id.py manual CAM++ smoke script, README ops notes

Test status

  • 851 passed
  • Two pre-existing errors, unrelated to this work: scripts/test_audio_path.py::test_live, scripts/test_live_speech.py::test_speech (fixture 'host' not found — live scripts needing a real server)

Follow-ups (non-blocking; fail-safe while the feature is disabled, which is the default)

  1. UserRegistry.start() does not self-heal from an initial crew outage (service restart required to recover) — fix before production enable.
  2. SpeakerIDHandler._pending can retain entries for finals that STT speculative filtering drops — no correctness impact; bound via SESSION_END cleanup or TTL.
  3. server.sh.example enable-workflow restructure: quote the <llm-host> placeholder before telling users to enable from the example.
  4. Spec §6 CI statistical latency check + §9(4) enrollment clip-count guidance in README.

Verification

  • 11-task plan executed under subagent-driven development; per-task reviews plus a final whole-branch review: no Critical/Important defects, all deferred minors accepted.
  • uv run pytest, ruff check, ruff format --check, mypy all pass.
  • Pre-push hygiene scan (private addresses/paths): clean.
Integrates the crew speaker-ID pipeline (spec: `docs/specs/2026-09-02-crew-user-system-design.md`). When enabled, every final transcription is tagged with the speaker's crew user id, resolved against the separate starfleet-crew service. When disabled (the default), behavior is unchanged. ## What's in - `speaker: str | None = None` on VADAudio / Transcription / GenerateResponseRequest / TranscriptionCompletedEvent (default `None` → fully backwards compatible) - New `computer/users` package: - `CrewClient` — httpx client for the crew service (registry, health, close) - `UserRegistry` — periodically-pulled user-vector snapshot with atomic swap, 3-state tagging (known / guest / stale), keep-stale on fetch failure - `SpeakerEmbedder` — lazy FunASR CAM++ (192-dim) - `speaker_tag` helpers (`GUEST`, `format_user_turn`) - `SpeakerIDHandler` — pipeline handler resolving speakers in parallel with STT - `speaker` threaded through the STT layer: `output_for_queue` + `TranscriptionNotifier`, including the speculative-revision path - Pipeline wiring in `s2s_pipeline.py`: `--users_*` CLI args, singleton init, handler insertion - `POST /internal/users/enroll` on the realtime service (Bearer token, raw PCM s16le 16 kHz mono, 192-dim vector, 401 / 422 / 503) - `server.sh.example` commented `--users_*` block, `scripts/smoke_speaker_id.py` manual CAM++ smoke script, README ops notes ## Test status - 851 passed - Two pre-existing errors, unrelated to this work: `scripts/test_audio_path.py::test_live`, `scripts/test_live_speech.py::test_speech` (`fixture 'host' not found` — live scripts needing a real server) ## Follow-ups (non-blocking; fail-safe while the feature is disabled, which is the default) 1. `UserRegistry.start()` does not self-heal from an initial crew outage (service restart required to recover) — fix before production enable. 2. `SpeakerIDHandler._pending` can retain entries for finals that STT speculative filtering drops — no correctness impact; bound via SESSION_END cleanup or TTL. 3. `server.sh.example` enable-workflow restructure: quote the `<llm-host>` placeholder before telling users to enable from the example. 4. Spec §6 CI statistical latency check + §9(4) enrollment clip-count guidance in README. ## Verification - 11-task plan executed under subagent-driven development; per-task reviews plus a final whole-branch review: no Critical/Important defects, all deferred minors accepted. - `uv run pytest`, `ruff check`, `ruff format --check`, `mypy` all pass. - Pre-push hygiene scan (private addresses/paths): clean.
docs: update demo video to vwTvn6qsVtQGsSW8TJMBDj
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 4m18s
77082e83cb
The PyPI qwentts-cpp-python 0.3.1 wheel (qwentts.cpp 7df559a) lacks the
ggml-cuda getrows kernel for GGML_TYPE_Q6_K, so Q4_K_M GGUFs abort with
unsupported src0 type: q6_K. Rebuilding against main/a8a7716+ includes the
case. Add scripts/build_qwentts_wheel.sh (clone+build_native+wheel) and
document VRAM table and rebuild steps in README.
fix: build_qwentts_wheel default ref master and tolerant checkout
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 3m52s
91a59532ba
qwentts.cpp default branch is master, not main. Also shallow clone
already on correct ref, so skip unnecessary checkout that aborted the
build with set -e.
- scripts/measure_vram.py polls nvidia-smi at 50ms, records static idle vs
  peak during qt_synthesize for empty/typical/worst scenarios
  (~100 chars vs ~700 chars / 1536 tokens) and concurrent=1 vs 2 pipelines
  (weights shared, scratch per concurrent synth). Outputs per-combo JSON
  + README_table.md + combined JSON to artifacts/vram/.
- README: add Empirical VRAM section with repro command and placeholder
  table for Q8_0/Q4_K_M empty/typical/worst c1/c2 (fill *to fill* after
  headless run). Clarify that idle pipelines don't add VRAM.
- artifacts/vram/.gitkeep to track dir.
Q8_0 CustomVoice measured via scripts/measure_vram.py (50ms nvidia-smi):
 empty 11619/0/287, typical c1 11627/11639/12/267, c2 11631/11647/16/259 (+8),
 worst c1 11631/11669/38/237, c2 11641/11679/38/227 (+10) on 12282 total
 with LLM 8256 baseline (server stopped). Q4_K_M empty 10619/0/1287
 measured pre-rebuild; typical/worst estimated from Q8_0 transient
 pending q6_K wheel rebuild for Q4_K_M CustomVoice. Artifacts in
 artifacts/vram/2026-08-27_*.json + combined. Clarify concurrent
 pipelines share weights, only overlap adds ~8-10 MB scratch.
Q4_K_M now measured on headless A2000 12GB (12282) with qwentts.cpp a8a7716 CUDA wheel (q6_K getrows fix):
- empty 10799/0, typical c1 10813 c2 10821 (+8), worst c1 10843 c2 10851 (+8)
- Q8_0: empty 11641, typical 11655/11663 (+8), worst 11685/11693 (+8)
Saves ~842 MiB peak worst (10851 vs 11693). Fixes huggingface-hub pin <1.0 after rebuild.
Raw JSON + README_table in artifacts/vram/2026-08-27_*.json
fix: relocate commented users_* block out of shell continuation in server.sh.example
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 6m51s
46dc9469b5
Merge remote-tracking branch 'origin/main' into devel
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 4m29s
c8ab8c9f7e
# Conflicts:
#	README.md
troed merged commit e741e1b124 into main 2026-09-03 08:48:28 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!11
No description provided.