ESP32-class MCU client for the Computer! voice assistant
  • C 68.5%
  • Python 20.8%
  • C++ 8.1%
  • Shell 2%
  • CMake 0.6%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-10-10 18:50:51 +02:00
components fix(ota): harden manifest verification and schedule handling 2026-09-11 21:04:35 +02:00
docs docs: ESP32 OTA implementation plan (A/B partitions, signed manifests, wss) 2026-09-11 18:13:06 +02:00
host fix(tools): resolve flash_models.sh partition CSV from the build config 2026-10-10 16:27:27 +00:00
main fix(ota): harden manifest verification and schedule handling 2026-09-11 21:04:35 +02:00
tools fix(tools): resolve flash_models.sh partition CSV from the build config 2026-10-10 16:27:27 +00:00
.gitignore feat(release): carry Wi-Fi SSID/password in release build (Kconfig-only provisioning) 2026-09-14 15:25:14 +02:00
AGENTS.md docs: note the partition-layout drift trap; drop stale sg dialout guidance 2026-10-10 16:28:31 +00:00
CMakeLists.txt feat: IDF project scaffold for ESP32-S3-AUDIO-Board 2026-08-22 14:19:29 +02:00
LICENSE chore(license): replace Apache-2.0 (Hugging Face boilerplate) with CC0 1.0 2026-08-26 21:15:01 +02:00
partitions_ota16m.csv feat(partition): A/B OTA slots with otadata and bootloader rollback 2026-09-11 18:20:14 +02:00
partitions_sr16m.csv feat(mww): wire microwakeword as sole wake source; drop wn9 models 2026-08-24 01:43:26 +02:00
POWER.md docs: add POWER.md (power architecture, low-power idle, drain measurement) 2026-09-10 19:57:59 +02:00
README.md docs: A/B OTA migration, release tooling, and trust material 2026-09-11 20:32:30 +02:00
sdkconfig.defaults fix(tools): isolate release build config and assert release identity 2026-09-11 21:04:58 +02:00

communicator-esp32

Wake-word voice communicator client for the Computer! realtime assistant (OpenAI-Realtime-compatible WebSocket at ws://{host}:{port}/v1/realtime). Third client in the starfleet family alongside the PC/server package (computer) and the Sailfish OS client (communicator-sailfish). Say the wake word, hear the chime, talk hands-free — server-side VAD decides turn ends, barge-in interrupts playback, and the device drops its connection after idle so it never holds a server pipeline slot.

Hardware

Waveshare ESP32-S3-AUDIO-Board: ESP32-S3 (dual-core @ 240 MHz), 16 MB flash, 8 MB octal PSRAM, Wi-Fi 2.4 GHz. Onboard ES8311 DAC (speaker) + ES7210 ADC (mic) behind an I2C control bus, WS2812 x7 RGB ring, and three user keys. Flashing is USB (baseline) plus over-the-air for app firmware: A/B OTA slots with bootloader rollback, Ed25519-signed release manifests, and TLS pinned to a local CA (design: starfleet/crew:docs/specs/2026-09-10-ota-design.md).

Firmware: ESP-IDF v5.5.0 (exact; installed at ~/esp/esp-idf, toolchains in ~/.espressif). Target esp32s3, partition table partitions_ota16m.csv (16 MB layout, A/B OTA: otadata + ota_0/ota_1 3 MB slots + mww 5000 KB). Managed components resolve via the IDF component registry espressif/esp-sr, espressif/esp_codec_dev, espressif/esp_websocket_client, espressif/led_strip, …; the board layer itself is a custom BSP under components/bsp).

Board pins (canonical source: components/bsp/include/bsp.h):

Signal GPIO Notes
I2C SDA 11 vendor wiki pinout confirmed by bus scan
I2C SCL 10 (the plan's SDA=10/SCL=11 was transposed)
I2S MCLK 12
I2S BCLK 13
I2S WS 14
Mic DIN 15 ES7210 ASDOUT (capture)
Spk DOUT 16 ES8311 DSDIN (playback)
LED ring data 38 WS2812 x7

Keys K1/K2/K3 are not native ESP32 GPIOs — they hang off a TCA9555 I2C expander at address 0x20 (EXIO9/10/11), read active-low:

  • K1 — volume down (−10)
  • K2 — volume up (+10)
  • K3 — mic mute toggle (blocks upstream audio)

The speaker amp enable is also behind the TCA9555 (EXIO8, active-high).

Build & flash

. $HOME/esp/esp-idf/export.sh
idf.py set-target esp32s3      # once; sdkconfig.defaults pins it anyway
idf.py menuconfig              # provision Wi-Fi/server + MicroWakeWord (table below)
idf.py build
idf.py -p /dev/ttyACM0 flash monitor
# model partition (once per model update; survives ordinary idf.py flash):
sg dialout -c "tools/flash_models.sh"

dialout caveat: if your user is not in the dialout group (and you don't want to re-login after adding it), wrap every port-touching command in sg:

sg dialout -c "idf.py -p /dev/ttyACM0 flash"
sg dialout -c "idf.py -p /dev/ttyACM0 monitor"
sg dialout -c "tools/flash_models.sh"

Build-only calls need no wrapper. idf.py monitor refuses non-TTY stdin in some capture setups; plain pyserial works for scripted serial logging. The mww model partition is never flashed by idf.py — always use tools/flash_models.sh (it also embeds the selftest reference wav when ../computer/test_wakeword.wav is present). Never use idf.py flash to update the model partition — it can replace the partition table and orphan the payload.

One-time migration to A/B OTA (from the old single-app table)

The first move to partitions_ota16m.csv requires a USB reflash per device, then model content:

sg dialout -c "idf.py -p /dev/ttyACM0 erase-flash"
sg dialout -c "idf.py -p /dev/ttyACM0 flash"
sg dialout -c "tools/flash_models.sh"

After this, app firmware updates happen over the air from crew; mww model updates still go through tools/flash_models.sh. USB-flashing an app after OTA (e.g. a recovery) writes ota_0 and resets the OTA slot selection:

sg dialout -c "idf.py -p /dev/ttyACM0 app-flash"

Releases are built and published with tools/release_firmware.sh (build machine, keys in tools/ota-keys/); devices download and verify them.

Provisioning (Kconfig)

All configuration lives in main/Kconfig.projbuild → Communicator Configuration and components/mww/Kconfig → MicroWakeWord. Reflash to change network settings (no NVS persistence).

Option Default Meaning
WIFI_SSID (empty) Wi-Fi SSID
WIFI_PASSWORD (empty) Wi-Fi password
SERVER_HOST (empty) computer server host
SERVER_PORT 8766 computer server port
CONNECT_TIMEOUT_SEC 5 WebSocket connect timeout
IDLE_DISCONNECT_SEC 15 Disconnect after N s idle past last server event
DEFAULT_VOLUME 60 Speaker volume 0–100 (runtime-adjustable)
BSP_SELFTEST n Boot-time tone + mic RMS selftest
IDLE_LOW_POWER y Low-power idle: amp + ring off, 2 Hz key poll while IDLE (restores on wake)
UI_COLOR_CALIBRATE n Ring R→G→B channel-map probe
UI_STATE_WALK n Cycles all six ring states
AFE_REF_DELAY_MS 80 AEC reference delay compensation
MWW_THRESHOLD 0 Wake probability threshold (permille, 0 = defer to computer.json cutoff)
MWW_REFRACTORY_SEC 2 Minimum seconds between wake detections
MWW_INPUT_GAIN 30000 Digital input gain into wake path (permille, 30×) — stages raw mic peaks of a few hundred LSBs into the trained domain
MWW_PMAX_FLOOR 10 Anti-transient floor — fire requires best window P(wake) ≥ this (permille, 0 = off)
MWW_MIN_ABOVE 3 Spread evidence — ≥ this many windows in average must individually reach threshold (0 = off)
MWW_ARM_DELAY_MS 250 Burst-age arming — suppress fires until gated burst has lasted this long (0 = off)
MWW_DEBUG_METRICS n Log sliding-window P(wake) / peak every ~0.5 s (calibration aid)
MWW_FA_CAPTURE n Hex-dump 6 s gated audio ring on every wake fire to serial (FAD lines, for harvest — adds ~192 KB PSRAM)
MWW_SELFTEST n Boot-time deterministic selftest — stream partitioned test wav through the full wake chain

Contributor rules — including secrets never get committed — live in AGENTS.md.

Wake word — "computer" via microWakeWord

The device wakes hands-free on "computer" (same word as the computer server) via a custom microWakeWord streaming classifier running entirely on-device — no WakeNet / wn9_hiesp ("Hi ESP") remains.

  • Model: tools/wakeword-model/computer.tflite + computer.json — int8 mixednet (~59 KB), 40-dim mel front-end (30 ms window / 10 ms stride), stride 3, sliding window 5. Tensor arena 34304 B (PSRAM). Shipped v2 is fine-tuned on ear-verified device captures: 7 positives (4 train / 1 val / 1 test / 1 held-out reference) mixed with 6500 synthetic piper-TTS positives. The held-out real reference peaks at P(wake) = 1.00 with operating probability_cutoff = 0.50 (recall 1.0, 1.875 FAH on testing_ambient). The synthetic-only predecessor was deaf on real speech (live peaks 0.016–0.04, only its silence floor ever fired). See tools/wakeword-model/REPRODUCIBILITY.md for the full config, seed (20260823), cutoff sweep, and arena derivation.
  • Engine: components/mww/ (MIT, ours; mel preprocessor derived from Apache-2.0 espressif/esp-tflite-micro examples/micro_speech) on the standard TFLM runtime (espressif/esp-tflite-micro + esp-nn). Model + JSON load at boot from the mww flash partition (5000 KB, data/spiffs); build embeds expected size + SHA256 prefix so stale partitions are caught loudly. Flash with tools/flash_models.sh (not idf.py flash):
    . $HOME/esp/esp-idf/export.sh
    sg dialout -c "tools/flash_models.sh"              # model + reference wav
    sg dialout -c "tools/flash_models.sh --wav FILE"   # override test wav
    sg dialout -c "tools/flash_models.sh --image-only" # build/verify only
    
    Payload layout is documented in tools/flash_models.sh header (32 B MWW1 header → json → tflite → optional 16 kHz mono test wav).
  • Training pipeline: tools/wakeword-training/ synthesizes positives with piper-sample-generator (LibriTTS-R, 5k + 1500 phonetic variants) and trains with OHF-Voice/micro-wake-word @ 4665173 via the wrapper train_computer.py (config_computer.json = synthetic baseline, config_computer_real.json = v2 real-capture fine-tune). Real positives are harvested on-device with CONFIG_MWW_FA_CAPTURE=y + CONFIG_MWW_DEBUG_METRICS=y (each wake fire snapshots [trigger−2 s … trigger+4 s] to a PSRAM ring and hex-dumps it over serial — extract_capture_log.py → ear-verify → slice_captures.py), then weighted positives_real: 10.0 vs synthetic 2.0 with explicit real_split lists. Full workflow, capture rig, and lessons are in docs/wake-word-pipeline.md and tools/wakeword-training/README.md.
  • Runtime: continuous streaming inference — one mww_features_step + TFLM Invoke per 480-sample stride (30 ms) while IDLE; during LISTENING/SPEAKING the ring is drained without inference (stays sample-synced, zero CPU). Trigger fires app_state_wake_detected() only when IDLE && !mic-muted, via a sliding-window average crossing MWW_THRESHOLD (or the JSON probability_cutoff when 0) plus spread-evidence conditions (MWW_MIN_ABOVE, MWW_PMAX_FLOOR, MWW_ARM_DELAY_MS) and a refractory gate. Audio reaches the model post-AEC, post MWW_INPUT_GAIN (default 30000) — raw mic peaks of a few hundred LSBs are staged into the trained domain; the offline trainer is amplitude-invariant but the device PCAN/NR front-end is not.
  • Selftest / calibration: CONFIG_MWW_SELFTEST=y streams the partitioned test wav through the full chain (bypassing mic) logging P(wake) per invoke plus a summary — deterministic regression of quantization/PCAN/streaming; CONFIG_MWW_DEBUG_METRICS=y exposes the live sliding average / peak and adaptive noise state for threshold tuning.

Host conformance gate

cd host/tests && python3 -m pytest -q

Must be green before any on-device network work. The suite unit-tests the wire-format helpers, runs the protocol layer against the shared conformance harness from the sibling checkout: ../computer/tests/conformance/ fake_realtime_server.py, and host-builds the portable wake-word core (host/mwwtest → test_mww_features.py / test_mww_trigger.py — features vs the Python reference oracle, trigger state-machine). No SDK needed — pure host pytest.

Architecture (one paragraph)

afe_task runs two FreeRTOS tasks (feed/fetch) around the ESP-SR audio frontend in AEC-only mode (wakenet_init=false, afe_config_init("MR", …) with empty model list; AFE_REF_DELAY_MS aligns the playback reference). The enhanced mono fetch path is injected via mww_inject_pcm() into mww_task (prio 6, 8 KiB), which computes 40-dim mel features in C and invokes the int8 streaming classifier (computer.tflite) every 30 ms (every 3 slices) while IDLE; a sliding-window trigger fires the wake. A detection wakes the app; ws_client opens the WebSocket to computer, sends session.update immediately on connect to arm server-side VAD, then shuttles events. player buffers outbound audio in a 256 KB PSRAM jitter ring (~8.2 s cap) feeding 20 ms slices and does ack accounting (exactly one output_audio.played per response). app_state is the machine IDLE → CONNECTING → LISTENING ⇄ SPEAKING → DISCONNECTING; ui_ring maps states onto the WS2812 x7 ring; buttons polls the TCA9555 keys.

Behavior notes

  • Low-power idle (CONFIG_IDLE_LOW_POWER, default on): while IDLE the speaker amp is disabled, the LED ring is cleared, and the key poll drops to 2 Hz; all restore on wake. Wake-word detection (mic + AEC + mww inference) stays live at all times, and Wi-Fi only ever connects from the first wake onward. A power_hb task persists a virtual run clock to NVS every 60 s so each boot logs the previous run's duration + reset reason (power: prev run ... (reset BROWNOUT ...)) ‒ with the brownout detector enabled, a battery discharge self-reports how long the run lasted. Full power architecture, log semantics, and the no-power-meter drain-measurement method: POWER.md.
  • Server-side VAD: the server commits utterances on silence; the device just streams mic audio while LISTENING.
  • Barge-in: speaking over playback makes the server cancel and start a new response; the player drops buffered audio and any chime tail.
  • Follow-up window: after each response the mic stays live until the idle watchdog fires — the device disconnects after 15 s without server events (CONFIG_IDLE_DISCONNECT_SEC), so follow-ups must land within ~15 s.
  • Idle disconnect: after 15 s without server events the device disconnects itself, releasing the server pipeline slot (a second client can connect while this one sits idle).
  • session.update on connect arms the server's VAD before any audio flows.
  • Wake chime is local: the canonical family cue (here.wav) is embedded in the firmware and injected into the playback ring when CONNECTING starts.
  • Wire contract is owned by computer: see ../computer/src/computer/api/openai_realtime/CONFORMANCE.md (canonical — this repo's docs/conformance.md and host tests track it).

Status

Ships the custom "computer" wake-word model (v2, 2026-08-26) — a microWakeWord-format int8 mixednet distilled from real device captures. Previous stock wn9_hiesp ("Hi ESP") is gone: esp-sr runs AEC-only (wakenet_init=false, empty srmodels.bin) and wake originates solely in components/mww (mww_task). The v2 model (tools/wakeword-model/ computer.tflite 59392 B + computer.json cutoff 0.50) was fine-tuned on 7 ear-verified harvests; the held-out reference peaks at P(wake) = 1.00 and wakes live on the first natural utterance. Promotion path is copy-computer.tflite+computer.json into tools/wakeword-model/ → tools/flash_models.sh → reflash app if the embedded SHA prefix changed. See docs/wake-word-pipeline.md for the harvest → slice → train → validate → ship workflow and tools/wakeword-model/REPRODUCIBILITY.md for the bit-exact trainer record.