parakeet-onnx STT + LLM/session reliability fixes #13

Merged
troed merged 11 commits from devel into main 2026-09-06 20:38:32 +02:00
Owner

Summary

Brings devel to main: the new low-VRAM STT backend plus the reliability fixes found during live deployment.

New: --stt parakeet-onnx (int8 ONNX Runtime)

  • Runs parakeet-tdt-0.6b-v3 via onnx-asr (parakeet-onnx extra) with a hard-capped CUDA arena (gpu_mem_limit 1.5 GiB default, kSameAsRequested), sized to coexist with a local llama-server: load +361 MB, 10 s turn +533 MB, 30 s worst case +1.35 GB (measured on RTX A2000 12 GB).
  • Final-turn-only transcription; same 25 European languages (shares the lingua detector).
  • New CLI flags: --parakeet_onnx_model_name, --parakeet_onnx_device, --parakeet_onnx_quantization, --parakeet_onnx_gpu_mem_limit_bytes, --parakeet_onnx_language; benchmark support in scripts/benchmark_stt.py; README + STT README document the VRAM profile.

LLM reliability

  • Runaway generations bounded: max_output_tokens (default 512, 0 disables) sent on every chat-completions request; compaction calls exempt.
  • First-output watchdog (first_output_timeout_s, default 15 s, 0 disables): a stream that yields nothing in time is cancelled with the spoken apology instead of silently consuming for minutes; audio text that exceeds 2000 chars without a sentence boundary is force-flushed so TTS is never silent.
  • Real-shape warmup: warmup now sends the actual init_chat_prompt with max_tokens=1 (llama prefix-cache holds it) — first question after a restart drops from ~7.5 s to ~1 s; warmup failure no longer aborts startup.

Turn identity / session hygiene

  • VAD turn ids are monotonic across sessions — a reconnect no longer collides with persisted STT completed-turn state (which wedged the pipeline as "stale").
  • Dropped FINAL-mode STT inputs always log at INFO; on_session_end logs cleared key counts; parakeet-onnx handler no longer shadows the base on_session_end.

Language / echo correctness

  • Reply-language instruction moved to the system prompt and embedded in the last user message, re-embedded on every tool round — Swedish tool output no longer derails the reply language or tool calling; exactly one "Working." status chunk per turn regardless of tool count.
  • Transcription notifier no longer records user transcripts into the echo filter (repeated questions were discarded as echoes of themselves).
  • server.sh.example documents the TTS-language guard flags (--enable_lang_prompt --tts_supported_languages ... --tts_default_language).

Test plan

  • Full pytest suite green (919 passed), ruff + mypy clean
  • Deployed to the reference RTX A2000 server; verified live: warmup 0.7-0.8 s cold / ~0.1 s cached, fast first question, turn-id collision gone, single Working. per multi-tool turn, Swedish question → tool calls + English reply, repeated question no longer echo-discarded
## Summary Brings `devel` to `main`: the new low-VRAM STT backend plus the reliability fixes found during live deployment. ### New: `--stt parakeet-onnx` (int8 ONNX Runtime) - Runs parakeet-tdt-0.6b-v3 via onnx-asr (`parakeet-onnx` extra) with a hard-capped CUDA arena (`gpu_mem_limit` 1.5 GiB default, `kSameAsRequested`), sized to coexist with a local llama-server: load +361 MB, 10 s turn +533 MB, 30 s worst case +1.35 GB (measured on RTX A2000 12 GB). - Final-turn-only transcription; same 25 European languages (shares the lingua detector). - New CLI flags: `--parakeet_onnx_model_name`, `--parakeet_onnx_device`, `--parakeet_onnx_quantization`, `--parakeet_onnx_gpu_mem_limit_bytes`, `--parakeet_onnx_language`; benchmark support in `scripts/benchmark_stt.py`; README + STT README document the VRAM profile. ### LLM reliability - **Runaway generations bounded**: `max_output_tokens` (default 512, 0 disables) sent on every chat-completions request; compaction calls exempt. - **First-output watchdog** (`first_output_timeout_s`, default 15 s, 0 disables): a stream that yields nothing in time is cancelled with the spoken apology instead of silently consuming for minutes; audio text that exceeds 2000 chars without a sentence boundary is force-flushed so TTS is never silent. - **Real-shape warmup**: warmup now sends the actual `init_chat_prompt` with `max_tokens=1` (llama prefix-cache holds it) — first question after a restart drops from ~7.5 s to ~1 s; warmup failure no longer aborts startup. ### Turn identity / session hygiene - VAD turn ids are monotonic across sessions — a reconnect no longer collides with persisted STT completed-turn state (which wedged the pipeline as "stale"). - Dropped FINAL-mode STT inputs always log at INFO; `on_session_end` logs cleared key counts; parakeet-onnx handler no longer shadows the base `on_session_end`. ### Language / echo correctness - Reply-language instruction moved to the system prompt *and* embedded in the last user message, re-embedded on every tool round — Swedish tool output no longer derails the reply language or tool calling; exactly one "Working." status chunk per turn regardless of tool count. - Transcription notifier no longer records user transcripts into the echo filter (repeated questions were discarded as echoes of themselves). - `server.sh.example` documents the TTS-language guard flags (`--enable_lang_prompt --tts_supported_languages ... --tts_default_language`). ## Test plan - [x] Full pytest suite green (919 passed), ruff + mypy clean - [x] Deployed to the reference RTX A2000 server; verified live: warmup 0.7-0.8 s cold / ~0.1 s cached, fast first question, turn-id collision gone, single Working. per multi-tool turn, Swedish question → tool calls + English reply, repeated question no longer echo-discarded
New --stt parakeet-onnx choice runs parakeet-tdt-0.6b-v3 via onnx-asr
with a bounded CUDA arena (kSameAsRequested + gpu_mem_limit, default
1.5 GiB), sized to coexist with a local LLM server on a shared GPU.
Final-turn-only transcription; language detection reuses the parakeet
lingua detector. Installed via the new 'parakeet-onnx' extra.
rename_args() adds gen_kwargs to every parsed argument dataclass, so
STT handler setup signatures must tolerate it (as parakeet-tdt does).
istupakov/parakeet-tdt-0.6b-v3-onnx is the repo that hosts the int8
ONNX weights (and the one cached on the server); the non-onnx repo
401s on config.json.
The built-in registry entry resolves to the istupakov int8 ONNX files
that are already cached on the server; passing the raw HF repo name
made onnx-asr fall back to config.json fetching from a different repo.
Three production failures addressed:
- gemma runaway generation (47k tokens/170s) silently consumed by the
  sentence-batching consumer: cap output tokens (max_output_tokens,
  default 512, 0 disables), raise FirstOutputTimeout when a stream
  yields nothing within first_output_timeout_s (default 15s) and emit
  the spoken apology, and force-flush audio text that exceeds 2000
  chars without a sentence boundary.
- toy warmup never warmed the real prompt (first question paid 7.5s
  prefill): warmup now sends the actual init_chat_prompt with
  max_tokens=1 so llama prefix-cache holds it; warmup failure no
  longer aborts setup.
- reconnect after a session wedged the pipeline: STT turn ids restart
  at turn_1 while BaseSTTHandler persists completed (turn, revision)
  keys, so the next session's first turn was dropped as stale.
  VAD turn ids are now monotonic across sessions, dropped FINAL inputs
  always log at INFO, on_session_end logs cleared key counts, and the
  parakeet-onnx handler no longer shadows the base on_session_end.
The per-turn language constraint (46219e2) only activates with
--enable_lang_prompt + --tts_supported_languages + --tts_default_language;
the example never showed them, so deployments could silently ship
without it (as the live server.sh had).
- The 'reply in <lang>' instruction was appended as a trailing user-role
  message after the actual question; gemma then treated it as the final
  instruction and answered conversationally instead of calling tools.
  It now rides in the system prompt via the instructions text.
- TranscriptionNotifier recorded every user transcript into the echo
  filter history, so repeating a question verbatim was discarded as an
  echo of itself (turn_3 live: repeated weather question dropped). The
  echo history stays assistant-output only (lm_output_processor).
gemma-4-12b ignored the system-prompt language stamp once Swedish tool
output dominated the context (probes A-D); embedding '(Reply in X.)'
in the same user message as the question survives it and does not
derail tool calls. Multiple server-side tools in one turn now emit a
single 'Working.' status chunk instead of one per tool.
_generate re-serializes the chat for each server-side tool iteration,
but _serialize consumed the note after the first round — the round
that actually answers (after Swedish tool output) lost it. Keep the
note on the handler; process() resets it per turn.
docs: document parakeet-onnx VRAM profile and usage
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 2m50s
eaf6ced819
troed merged commit a0e3f384a1 into main 2026-09-06 20:38:32 +02:00
troed deleted branch devel 2026-09-06 20:38:32 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!13
No description provided.