feat: add OmniVoice TTS backend with quantization and FlashInfer options #15

Merged
troed merged 4 commits from feat/omnivoice-tts into devel 2026-09-07 22:20:28 +02:00
Owner

Summary

Adds OmniVoice (k2-fsa) as a TTS backend: 600+ languages, zero-shot voice cloning and voice design, 0.6B fp16 weights + 805MB audio tokenizer — far under the Qwen Q4_K_M VRAM budget.

  • Ported handler + argument class (adapted from upstream speech-to-speech PR #527, with fork's simple EndOfResponse semantics), wired into s2s_pipeline mirroring pocket_tts (backend arg --tts omnivoice, computer[omnivoice] pip extra).
  • bitsandbytes quantization on load: --omnivoice_quant none|int8|nf4|fp4 (CUDA only); NF4 measured fastest unaccelerated (RTF ~0.19, static ~1.8 GiB).
  • Optional FlashInfer acceleration (--omnivoice_flashinfer, --omnivoice_cuda_graph): RTF ~0.05 (~4x). Requires flashinfer-python + git-source omnivoice (see TTS README). Quant+flashinfer combo is guarded (FlashInfer weight packing requires unquantized Linear).
  • Docs: TTS README section with usage, measured VRAM/RTF table, and license notes (Apache-2.0 code; weights CC-BY-NC 4.0 + Boson Higgs Community License — non-commercial).
  • Tests: 18 OmniVoice handler tests (TDD), full suite 930 passed / 16 skipped, ruff clean.
## Summary Adds OmniVoice (k2-fsa) as a TTS backend: 600+ languages, zero-shot voice cloning and voice design, 0.6B fp16 weights + 805MB audio tokenizer — far under the Qwen Q4_K_M VRAM budget. - Ported handler + argument class (adapted from upstream speech-to-speech PR #527, with fork's simple EndOfResponse semantics), wired into s2s_pipeline mirroring pocket_tts (backend arg `--tts omnivoice`, `computer[omnivoice]` pip extra). - bitsandbytes quantization on load: `--omnivoice_quant none|int8|nf4|fp4` (CUDA only); NF4 measured fastest unaccelerated (RTF ~0.19, static ~1.8 GiB). - Optional FlashInfer acceleration (`--omnivoice_flashinfer`, `--omnivoice_cuda_graph`): RTF ~0.05 (~4x). Requires flashinfer-python + git-source omnivoice (see TTS README). Quant+flashinfer combo is guarded (FlashInfer weight packing requires unquantized Linear). - Docs: TTS README section with usage, measured VRAM/RTF table, and license notes (Apache-2.0 code; weights CC-BY-NC 4.0 + Boson Higgs Community License — non-commercial). - Tests: 18 OmniVoice handler tests (TDD), full suite 930 passed / 16 skipped, ruff clean.
Port k2-fsa OmniVoice (600+ languages, voice cloning, voice design)
as --tts omnivoice, thread its kwargs through s2s_pipeline mirroring
pocket TTS, add the omnivoice pip extra and a section in the TTS
README. Weights are CC-BY-NC 4.0 (non-commercial).
feat: guard quant+flashinfer combo, document measured VRAM/RTF
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 2m22s
fdb092939e
troed merged commit 1a376270a2 into devel 2026-09-07 22:20:28 +02:00
troed deleted branch feat/omnivoice-tts 2026-09-07 22:20:28 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!15
No description provided.