chore: keeping documentation up to date #12

Merged
troed merged 56 commits from devel into main 2026-09-04 12:20:58 +02:00
Owner

Docs: Speaker ID section gains the device-id / connected-device registry and guided-capture (enroll/verify) bullets; CONFORMANCE.md gains the optional server.capture_control extension.

Includes eeede3f: fixes 9 pre-existing mypy errors and two unformatted files from the capture-mode work so CI passes on the merge.

Docs: Speaker ID section gains the device-id / connected-device registry and guided-capture (enroll/verify) bullets; CONFORMANCE.md gains the optional server.capture_control extension. Includes eeede3f: fixes 9 pre-existing mypy errors and two unformatted files from the capture-mode work so CI passes on the merge.
docs: update demo video to vwTvn6qsVtQGsSW8TJMBDj
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 4m18s
77082e83cb
The PyPI qwentts-cpp-python 0.3.1 wheel (qwentts.cpp 7df559a) lacks the
ggml-cuda getrows kernel for GGML_TYPE_Q6_K, so Q4_K_M GGUFs abort with
unsupported src0 type: q6_K. Rebuilding against main/a8a7716+ includes the
case. Add scripts/build_qwentts_wheel.sh (clone+build_native+wheel) and
document VRAM table and rebuild steps in README.
fix: build_qwentts_wheel default ref master and tolerant checkout
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 3m52s
91a59532ba
qwentts.cpp default branch is master, not main. Also shallow clone
already on correct ref, so skip unnecessary checkout that aborted the
build with set -e.
- scripts/measure_vram.py polls nvidia-smi at 50ms, records static idle vs
  peak during qt_synthesize for empty/typical/worst scenarios
  (~100 chars vs ~700 chars / 1536 tokens) and concurrent=1 vs 2 pipelines
  (weights shared, scratch per concurrent synth). Outputs per-combo JSON
  + README_table.md + combined JSON to artifacts/vram/.
- README: add Empirical VRAM section with repro command and placeholder
  table for Q8_0/Q4_K_M empty/typical/worst c1/c2 (fill *to fill* after
  headless run). Clarify that idle pipelines don't add VRAM.
- artifacts/vram/.gitkeep to track dir.
Q8_0 CustomVoice measured via scripts/measure_vram.py (50ms nvidia-smi):
 empty 11619/0/287, typical c1 11627/11639/12/267, c2 11631/11647/16/259 (+8),
 worst c1 11631/11669/38/237, c2 11641/11679/38/227 (+10) on 12282 total
 with LLM 8256 baseline (server stopped). Q4_K_M empty 10619/0/1287
 measured pre-rebuild; typical/worst estimated from Q8_0 transient
 pending q6_K wheel rebuild for Q4_K_M CustomVoice. Artifacts in
 artifacts/vram/2026-08-27_*.json + combined. Clarify concurrent
 pipelines share weights, only overlap adds ~8-10 MB scratch.
Q4_K_M now measured on headless A2000 12GB (12282) with qwentts.cpp a8a7716 CUDA wheel (q6_K getrows fix):
- empty 10799/0, typical c1 10813 c2 10821 (+8), worst c1 10843 c2 10851 (+8)
- Q8_0: empty 11641, typical 11655/11663 (+8), worst 11685/11693 (+8)
Saves ~842 MiB peak worst (10851 vs 11693). Fixes huggingface-hub pin <1.0 after rebuild.
Raw JSON + README_table in artifacts/vram/2026-08-27_*.json
fix: relocate commented users_* block out of shell continuation in server.sh.example
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 6m51s
46dc9469b5
Merge remote-tracking branch 'origin/main' into devel
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 4m29s
c8ab8c9f7e
# Conflicts:
#	README.md
PyPI qwentts-cpp-python 0.3.1 (qwentts.cpp 7df559a) is missing the CUDA getrows kernel for K-quants, so any Q4_K_M GGUF TTS attempt SIGABRTs and takes down the whole server (silent ESP32 after the 2026-09-03 deploy).

- scripts/build_qwentts_wheel.sh: default arch 86-real (deployed GPU), uv fallback when python3 lacks the build module, copy the wheel into gitignored wheels/ and install into .venv
- pyproject.toml: qwentts-cpp-python==0.3.1 as a direct dep + [tool.uv.sources] path pin to wheels/ (uv source overrides are not applied to transitive deps)
- .gitignore: wheels/
- README.md: build + pinning docs
resolve_speaker() previously waited on a deadline set at VAD-final spawn,
but it is called after STT finishes (60-250 ms later), so the budget was
already spent and tags were always dropped. The join now gets the full
join_budget_ms from the point the resolution is requested.

SpeakerEmbedder now uses CUDA when available (warm ~60 ms / 62 MiB VRAM
vs 68-878 ms on CPU) so identification finishes inside the join budget.
The voice system prompt now explains the [name] transcript prefix so the
LLM knows who is speaking. README documents the VRAM cost.
In sync mode (--stt native-llm), SpeakerIDHandler sets VADAudio.speaker
inline and registers no pending entry, but NativeLLMSTTHandler built the
Transcription without speaker=, dropping the tag before
TranscriptionNotifier. The resolver then correctly returned None (nothing
pending in sync mode), so the LLM never saw the [name] prefix.

Now Transcription carries speaker=vad_audio.speaker. Regression test added.
The rebuilt wheel pairs the native lib (ServeurpersoCom/qwentts.cpp
a8a7716+, ABI v4) with the stale upstream binding (andimarafioti/
qwentts-cpp-python 0.3.1, ABI v2). qt_init_default_params() writes
codec_chunk_sec 4 bytes past the undersized QtInitParams ctypes struct,
corrupting the heap and causing random SIGSEGVs during TTS model load.
qt_tts_params also carries two fields the new lib dropped.

build_qwentts_wheel.sh now applies an idempotent patch after the binding
checkout: bump QT_ABI_VERSION to 4, add max_batch/codec_chunk_sec to
QtInitParams, and remove the stale codec_chunk_sec/codec_left_context_sec
from QtTTSParams and its two write sites. A clean QWENTTS_PY_SOURCE
checkout reproduces the crash without this patch; the rebuilt wheel and
both venvs (server + local) are updated and verified (836 tests pass,
server stays up with max_batch=1 loading cleanly).
This reverts commit 75326416dc.
This reverts commit b32ff03f81.
chore: keeping documentation up to date
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 10s
0ea279cb3a
Merge remote-tracking branch 'origin/main' into devel
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 10s
8c191b30fe
# Conflicts:
#	README.md
#	src/computer/api/openai_realtime/server.py
#	src/computer/api/openai_realtime/websocket_router.py
#	src/computer/s2s_pipeline.py
#	src/computer/users/speaker_id_handler.py
#	tests/users/test_speaker_id_handler.py
fix: resolve qwentts-cpp-python from PyPI in CI when the local wheel is absent
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 5m34s
8cc8fad3ce
The [tool.uv.sources] pin points qwentts-cpp-python at a gitignored
local CUDA wheel (wheels/, built by scripts/build_qwentts_wheel.sh).
CI checkouts don't have that wheel, so uv sync failed with 'No such
file or directory'. Drop the pin in CI when wheels/ is empty and let
uv resolve the package from PyPI, restoring the pre-pin resolution.
troed merged commit df0de6a4bf into main 2026-09-04 12:20:58 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!12
No description provided.