parakeet-onnx STT + LLM/session reliability fixes #13
Loading…
Reference in a new issue
No description provided.
Delete branch "devel"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Brings
develtomain: the new low-VRAM STT backend plus the reliability fixes found during live deployment.New:
--stt parakeet-onnx(int8 ONNX Runtime)parakeet-onnxextra) with a hard-capped CUDA arena (gpu_mem_limit1.5 GiB default,kSameAsRequested), sized to coexist with a local llama-server: load +361 MB, 10 s turn +533 MB, 30 s worst case +1.35 GB (measured on RTX A2000 12 GB).--parakeet_onnx_model_name,--parakeet_onnx_device,--parakeet_onnx_quantization,--parakeet_onnx_gpu_mem_limit_bytes,--parakeet_onnx_language; benchmark support inscripts/benchmark_stt.py; README + STT README document the VRAM profile.LLM reliability
max_output_tokens(default 512, 0 disables) sent on every chat-completions request; compaction calls exempt.first_output_timeout_s, default 15 s, 0 disables): a stream that yields nothing in time is cancelled with the spoken apology instead of silently consuming for minutes; audio text that exceeds 2000 chars without a sentence boundary is force-flushed so TTS is never silent.init_chat_promptwithmax_tokens=1(llama prefix-cache holds it) — first question after a restart drops from ~7.5 s to ~1 s; warmup failure no longer aborts startup.Turn identity / session hygiene
on_session_endlogs cleared key counts; parakeet-onnx handler no longer shadows the baseon_session_end.Language / echo correctness
server.sh.exampledocuments the TTS-language guard flags (--enable_lang_prompt --tts_supported_languages ... --tts_default_language).Test plan