Realtime follow-up robustness: playback ack, grace windows and VAD re-arm #18

Merged
troed merged 7 commits from devel into main 2026-09-08 18:30:18 +02:00
Owner

Summary

Fixes the follow-up-after-reply failures reported on the live voice assistant:

  • Per-response playback-ack accounting (client): the desktop client acked playback instantly (lifetime-cumulative played bytes vs per-connection received bytes), so the server's wake-word-free grace window opened while the answer was still playing and expired mid-playback. The client now tracks bytes per response (same contract as the esp32 firmware player), acking at true playback end.
  • Zoom into stale grace windows (client): a leftover grace deadline from a previous reply could expire during the next reply's playback and re-gate the mic. New responses now cancel any stale window; a wedged response still arms one.
  • VAD re-arm on playback end (server): the shared Silero model's recurrent state degraded across the long near-silence + playback stream, underscoring the next utterance into discarded fragments (dump + offline analysis). VAD state is now reset when response playback ends.
  • VAD tuning (runtime config): --min_speech_ms 256 so short utterances survive.

Test plan

  • pytest: 983 passed, 2 skipped; ruff check+format clean.
  • TDD regression tests for each fix (per-response ack, stale-grace cancel, VAD re-arm live on server).
  • Verified on the live server across repeated wake → question → follow-up rounds; user-confirmed working ("this seems to work perfectly").

Release notes

Session code proportional to instrumentation probes was added and fully removed afterwards; COMPUTER_INPUT_DUMP env diagnostic included only in diagnostic commits that cancel out within this PR.

## Summary Fixes the follow-up-after-reply failures reported on the live voice assistant: - **Per-response playback-ack accounting (client)**: the desktop client acked playback instantly (lifetime-cumulative played bytes vs per-connection received bytes), so the server's wake-word-free grace window opened while the answer was still playing and expired mid-playback. The client now tracks bytes per response (same contract as the esp32 firmware player), acking at true playback end. - **Zoom into stale grace windows (client)**: a leftover grace deadline from a previous reply could expire during the next reply's playback and re-gate the mic. New responses now cancel any stale window; a wedged response still arms one. - **VAD re-arm on playback end (server)**: the shared Silero model's recurrent state degraded across the long near-silence + playback stream, underscoring the next utterance into discarded fragments (dump + offline analysis). VAD state is now reset when response playback ends. - **VAD tuning (runtime config)**: `--min_speech_ms 256` so short utterances survive. ## Test plan - `pytest`: 983 passed, 2 skipped; ruff check+format clean. - TDD regression tests for each fix (per-response ack, stale-grace cancel, VAD re-arm live on server). - Verified on the live server across repeated wake → question → follow-up rounds; user-confirmed working ("this seems to work perfectly"). ## Release notes Session code proportional to instrumentation probes was added and fully removed afterwards; `COMPUTER_INPUT_DUMP` env diagnostic included only in diagnostic commits that cancel out within this PR.
The ack compared a lifetime-cumulative played counter against per-connection
received bytes, so after any chime or reconnect it fired ~0s after
response_done — the wake-word-free grace window then expired while a long
answer was still playing and follow-up replies needed the wake word again.

Now the client snapshots the played baseline when a response's first audio
delta arrives and counts the response's own received bytes, mirroring the
esp32 firmware contract (exactly one output_audio.played per response).
Playback-era chunks consumed by the barge-in branch leave the shared Silero
model's recurrent state biased so a follow-up utterance right after playback
scores far below threshold and is discarded (the 'no response after long
answers' bug). Proven offline with server-side input dumps: the same audio
confirms with fresh state (p90 0.976, 1.4s of >=0.6 speech) but collapses
(~0.2-0.45) when fed sequentially through a response cycle. Reset the
iterator on the response_playing set->clear transition, mirroring the
existing barge-in reset. Also adds an env-gated COMPUTER_INPUT_DUMP probe
used to capture the server-side input audio for this diagnosis.
Reply 1's grace deadline (ack+10s) kept counting into reply 2's playback and
expired mid-playback (17:32 turn: Grace period expired at 210944/276480
bytes), after which the client re-gated the mic and the follow-up never
reached the server. Zero grace_deadline at each response's first audio
delta, and arm a fresh window when the ack-never-received give-up path
fires so a wedged response still ends with a usable listening window.
Regression test drives the exact sequence with a 0.3s window and asserts
no mid-playback expiry.
chore: remove temporary VAD/ack instrumentation probes
Some checks failed
CI / Sanity check (ubuntu-latest) (pull_request) Failing after 2m37s
b92ff80f3e
troed merged commit 73a9a2a559 into main 2026-09-08 18:30:18 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
starfleet/computer!18
No description provided.