feat: add OmniVoice TTS backend with quantization and FlashInfer options #15
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/omnivoice-tts"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Adds OmniVoice (k2-fsa) as a TTS backend: 600+ languages, zero-shot voice cloning and voice design, 0.6B fp16 weights + 805MB audio tokenizer — far under the Qwen Q4_K_M VRAM budget.
--tts omnivoice,computer[omnivoice]pip extra).--omnivoice_quant none|int8|nf4|fp4(CUDA only); NF4 measured fastest unaccelerated (RTF ~0.19, static ~1.8 GiB).--omnivoice_flashinfer,--omnivoice_cuda_graph): RTF ~0.05 (~4x). Requires flashinfer-python + git-source omnivoice (see TTS README). Quant+flashinfer combo is guarded (FlashInfer weight packing requires unquantized Linear).