feat: Q4_K_M CUDA wheel rebuild (q6_K) for 12GB headless #10

Merged
troed merged 3 commits from devel into main 2026-08-27 14:52:14 +02:00
Owner

Enable Q4_K_M on 12GB headless (RTX A2000).

  • PyPI qwentts-cpp-python 0.3.1 pins qwentts.cpp 7df559a (ABI 2) without CUDA getrows for GGML_TYPE_Q6_K → Q4_K_M GGUFs (Serveurperso/Qwen3-TTS-GGUF) abort unsupported src0 type: q6_K.
  • Newer qwentts.cpp a8a7716 bumps QT_ABI 2→4 (qt_tts_params removes codec_chunk_sec) → Python binding at ABI 2 mismatches → --ref-text requires --ref-wav abort.
  • Fix: stay on 7df559a (ABI 2) + cherry-pick ggml c044c6f Q6_K getrows, single-arch 86-real build. Verified headless: Talker 1765→1012 MB, nvidia-smi 11439→10597 (1309 free), no abort, TTS warmup ok.

Adds scripts/build_qwentts_wheel.sh (clone qwentts.cpp/qwentts-cpp-python, cmake -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86-real, build_native.py, wheel 1cu128, pip --no-deps) and README Q4_K_M VRAM table + rebuild docs.

Deployed to headless, tested Q4_K_M synthesis, service stable.

Enable Q4_K_M on 12GB headless (RTX A2000). - PyPI qwentts-cpp-python 0.3.1 pins qwentts.cpp 7df559a (ABI 2) without CUDA getrows for GGML_TYPE_Q6_K → Q4_K_M GGUFs (Serveurperso/Qwen3-TTS-GGUF) abort `unsupported src0 type: q6_K`. - Newer qwentts.cpp a8a7716 bumps QT_ABI 2→4 (qt_tts_params removes codec_chunk_sec) → Python binding at ABI 2 mismatches → `--ref-text requires --ref-wav` abort. - Fix: stay on 7df559a (ABI 2) + cherry-pick ggml c044c6f Q6_K getrows, single-arch 86-real build. Verified headless: Talker 1765→1012 MB, nvidia-smi 11439→10597 (1309 free), no abort, TTS warmup ok. Adds scripts/build_qwentts_wheel.sh (clone qwentts.cpp/qwentts-cpp-python, cmake -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86-real, build_native.py, wheel 1cu128, pip --no-deps) and README Q4_K_M VRAM table + rebuild docs. Deployed to headless, tested Q4_K_M synthesis, service stable.
docs: update demo video to vwTvn6qsVtQGsSW8TJMBDj
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 4m18s
77082e83cb
The PyPI qwentts-cpp-python 0.3.1 wheel (qwentts.cpp 7df559a) lacks the
ggml-cuda getrows kernel for GGML_TYPE_Q6_K, so Q4_K_M GGUFs abort with
unsupported src0 type: q6_K. Rebuilding against main/a8a7716+ includes the
case. Add scripts/build_qwentts_wheel.sh (clone+build_native+wheel) and
document VRAM table and rebuild steps in README.
fix: build_qwentts_wheel default ref master and tolerant checkout
All checks were successful
CI / Sanity check (ubuntu-latest) (pull_request) Successful in 3m52s
91a59532ba
qwentts.cpp default branch is master, not main. Also shallow clone
already on correct ref, so skip unnecessary checkout that aborted the
build with set -e.
troed merged commit 3946978a4c into main 2026-08-27 14:52:14 +02:00
troed deleted branch devel 2026-08-27 14:52:15 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!10
No description provided.