feat: per-turn response language constrained to TTS-supported languages #5

Merged
troed merged 5 commits from devel into main 2026-08-26 15:07:01 +02:00
Owner

Per-turn response language constrained to TTS-supported languages

The pipeline already plumbed language_code from STT through LLM output to TTSInput, but the LLM could reply in any language and Qwen3-TTS discarded the per-turn code, always synthesizing with its setup-time language.

What this adds

  • resolve_effective_language() (src/computer/LLM/utils.py): maps a detected language onto the TTS-supported set, falling back to a configurable default.
  • Variant 1 — discrete STT (e.g. parakeet): when both knobs are set, detected input languages outside the supported set are remapped to the default before the --enable_lang_prompt instruction is built, so the model never writes text TTS cannot speak; the resolved code is stamped on the turn for TTS.
  • Variant 2 — unified STT+LLM (Gemma, audio in): an instruction listing supported languages is injected next to the input_audio part, and the reply's language is detected from generated text via lingua (≥20 chars) and stamped onto output chunks.
  • Qwen3-TTS threading: process() now resolves each turn's language_code through QWEN3_LANGUAGE_ALIASES (base-prefix fallback, configured default otherwise) and threads it into _process_voice_clone/_process_custom_voice/_process_voice_design; existing mixed-language coalescing guard unchanged.
  • Config knobs: --tts_supported_languages / --tts_default_language added to shared LM args (verified parseable by HfArgumentParser); documented in src/computer/LLM/README.md.

Also updates the main README: wake word model training is complete (computer.onnx), noted as a non-distributed asset like chimes/voice data.

Testing

  • New tests: tests/test_llm_utils.py, tests/test_language_model_base_arguments.py, plus additions to qwen3-tts backend and chat-completions backend tests (RED→GREEN per unit).
  • Full suite: 804 passed. Pre-existing failures unrelated to this work were verified on clean HEAD via git stash and left alone: conformance test_server_starts_and_delivers_pushed_event, 3 files failing ruff format --check.
## Per-turn response language constrained to TTS-supported languages The pipeline already plumbed `language_code` from STT through LLM output to `TTSInput`, but the LLM could reply in any language and Qwen3-TTS discarded the per-turn code, always synthesizing with its setup-time language. ### What this adds - **`resolve_effective_language()`** (`src/computer/LLM/utils.py`): maps a detected language onto the TTS-supported set, falling back to a configurable default. - **Variant 1 — discrete STT (e.g. parakeet):** when both knobs are set, detected input languages outside the supported set are remapped to the default before the `--enable_lang_prompt` instruction is built, so the model never writes text TTS cannot speak; the resolved code is stamped on the turn for TTS. - **Variant 2 — unified STT+LLM (Gemma, audio in):** an instruction listing supported languages is injected next to the `input_audio` part, and the reply's language is detected from generated text via lingua (≥20 chars) and stamped onto output chunks. - **Qwen3-TTS threading:** `process()` now resolves each turn's `language_code` through `QWEN3_LANGUAGE_ALIASES` (base-prefix fallback, configured default otherwise) and threads it into `_process_voice_clone/_process_custom_voice/_process_voice_design`; existing mixed-language coalescing guard unchanged. - **Config knobs:** `--tts_supported_languages` / `--tts_default_language` added to shared LM args (verified parseable by HfArgumentParser); documented in `src/computer/LLM/README.md`. Also updates the main README: wake word model training is complete (`computer.onnx`), noted as a non-distributed asset like chimes/voice data. ### Testing - New tests: `tests/test_llm_utils.py`, `tests/test_language_model_base_arguments.py`, plus additions to qwen3-tts backend and chat-completions backend tests (RED→GREEN per unit). - Full suite: 804 passed. Pre-existing failures unrelated to this work were verified on clean HEAD via `git stash` and left alone: conformance `test_server_starts_and_delivers_pushed_event`, 3 files failing `ruff format --check`.
Follow-up to the public-repo security audit: tracked files carried real internal network details (LAN IP, hostnames, username, home paths, SSH port). This removes them; real values stay in gitignored local files and the deploy environment, mirroring the policy already used in starfleet/roster.

Changes:
- `AGENTS.md`: added a **Public repo hygiene** section (no internal hostnames/IPs/usernames/home paths/credentials in tracked files; pre-push scan command). Sanitized the deploy/fj sections that documented the live server's address, user, repo path, and SSH commands.
- `deploy.sh`: `SERVER_HOST`, `SERVER_USER`, `SERVER_REPO_DIR` are now **required env vars** (were hardcoded LAN defaults); setup comments use placeholders.
- `client.sh.example` / `server.sh.example` / `mcp.json.example`: placeholder hosts (`<server-host>`, `<llm-host>`, `<home-assistant-host>`).
- `README.md` + `src/computer/client/__init__.py`: example IPs replaced with placeholders.

Deploy note: run deploys with `SERVER_HOST=... SERVER_USER=... SERVER_REPO_DIR=... ./deploy.sh` (export them in your shell or a gitignored env file).

Verification: `bash -n deploy.sh` OK; edited Python module compiles; hygiene scan (`git grep -nIEi ...`) clean over tracked tree.
Co-authored-by: Troed Sångberg <github@troed.se>
Reviewed-on: #4
docs: note wake word model training is complete
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 3m33s
8a67ed25fd
troed merged commit 46219e2f66 into main 2026-08-26 15:07:01 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!5
No description provided.