- C 68.5%
- Python 20.8%
- C++ 8.1%
- Shell 2%
- CMake 0.6%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| components | ||
| docs | ||
| host | ||
| main | ||
| tools | ||
| .gitignore | ||
| AGENTS.md | ||
| CMakeLists.txt | ||
| LICENSE | ||
| partitions_ota16m.csv | ||
| partitions_sr16m.csv | ||
| POWER.md | ||
| README.md | ||
| sdkconfig.defaults | ||
communicator-esp32
Wake-word voice communicator client for the Computer! realtime
assistant (OpenAI-Realtime-compatible WebSocket at ws://{host}:{port}/v1/realtime).
Third client in the starfleet family alongside the PC/server package
(computer) and the Sailfish OS client (communicator-sailfish). Say the wake
word, hear the chime, talk hands-free — server-side VAD decides turn ends,
barge-in interrupts playback, and the device drops its connection after idle so
it never holds a server pipeline slot.
Hardware
Waveshare ESP32-S3-AUDIO-Board: ESP32-S3 (dual-core @ 240 MHz), 16 MB flash,
8 MB octal PSRAM, Wi-Fi 2.4 GHz. Onboard ES8311 DAC (speaker) + ES7210 ADC
(mic) behind an I2C control bus, WS2812 x7 RGB ring, and three user keys.
Flashing is USB (baseline) plus over-the-air for app firmware: A/B OTA slots
with bootloader rollback, Ed25519-signed release manifests, and TLS pinned to
a local CA (design: starfleet/crew:docs/specs/2026-09-10-ota-design.md).
Firmware: ESP-IDF v5.5.0 (exact; installed at ~/esp/esp-idf, toolchains
in ~/.espressif). Target esp32s3, partition table partitions_ota16m.csv
(16 MB layout, A/B OTA: otadata + ota_0/ota_1 3 MB slots + mww
5000 KB). Managed components resolve via the IDF component registry
espressif/esp-sr, espressif/esp_codec_dev, espressif/esp_websocket_client,
espressif/led_strip, …; the board layer itself is a custom BSP under
components/bsp).
Board pins (canonical source: components/bsp/include/bsp.h):
| Signal | GPIO | Notes |
|---|---|---|
| I2C SDA | 11 | vendor wiki pinout confirmed by bus scan |
| I2C SCL | 10 | (the plan's SDA=10/SCL=11 was transposed) |
| I2S MCLK | 12 | |
| I2S BCLK | 13 | |
| I2S WS | 14 | |
| Mic DIN | 15 | ES7210 ASDOUT (capture) |
| Spk DOUT | 16 | ES8311 DSDIN (playback) |
| LED ring data | 38 | WS2812 x7 |
Keys K1/K2/K3 are not native ESP32 GPIOs — they hang off a TCA9555 I2C
expander at address 0x20 (EXIO9/10/11), read active-low:
- K1 — volume down (−10)
- K2 — volume up (+10)
- K3 — mic mute toggle (blocks upstream audio)
The speaker amp enable is also behind the TCA9555 (EXIO8, active-high).
Build & flash
. $HOME/esp/esp-idf/export.sh
idf.py set-target esp32s3 # once; sdkconfig.defaults pins it anyway
idf.py menuconfig # provision Wi-Fi/server + MicroWakeWord (table below)
idf.py build
idf.py -p /dev/ttyACM0 flash monitor
# model partition (once per model update; survives ordinary idf.py flash):
sg dialout -c "tools/flash_models.sh"
dialout caveat: if your user is not in the dialout group (and you don't
want to re-login after adding it), wrap every port-touching command in sg:
sg dialout -c "idf.py -p /dev/ttyACM0 flash"
sg dialout -c "idf.py -p /dev/ttyACM0 monitor"
sg dialout -c "tools/flash_models.sh"
Build-only calls need no wrapper. idf.py monitor refuses non-TTY stdin in
some capture setups; plain pyserial works for scripted serial logging.
The mww model partition is never flashed by idf.py — always use
tools/flash_models.sh (it also embeds the selftest reference wav when
../computer/test_wakeword.wav is present). Never use idf.py flash to
update the model partition — it can replace the partition table and orphan
the payload.
One-time migration to A/B OTA (from the old single-app table)
The first move to partitions_ota16m.csv requires a USB reflash per device,
then model content:
sg dialout -c "idf.py -p /dev/ttyACM0 erase-flash"
sg dialout -c "idf.py -p /dev/ttyACM0 flash"
sg dialout -c "tools/flash_models.sh"
After this, app firmware updates happen over the air from crew; mww model
updates still go through tools/flash_models.sh. USB-flashing an app after
OTA (e.g. a recovery) writes ota_0 and resets the OTA slot selection:
sg dialout -c "idf.py -p /dev/ttyACM0 app-flash"
Releases are built and published with tools/release_firmware.sh (build
machine, keys in tools/ota-keys/); devices download and verify them.
Provisioning (Kconfig)
All configuration lives in main/Kconfig.projbuild → Communicator
Configuration and components/mww/Kconfig → MicroWakeWord. Reflash to
change network settings (no NVS persistence).
| Option | Default | Meaning |
|---|---|---|
WIFI_SSID |
(empty) | Wi-Fi SSID |
WIFI_PASSWORD |
(empty) | Wi-Fi password |
SERVER_HOST |
(empty) | computer server host |
SERVER_PORT |
8766 |
computer server port |
CONNECT_TIMEOUT_SEC |
5 |
WebSocket connect timeout |
IDLE_DISCONNECT_SEC |
15 |
Disconnect after N s idle past last server event |
DEFAULT_VOLUME |
60 |
Speaker volume 0–100 (runtime-adjustable) |
BSP_SELFTEST |
n |
Boot-time tone + mic RMS selftest |
IDLE_LOW_POWER |
y |
Low-power idle: amp + ring off, 2 Hz key poll while IDLE (restores on wake) |
UI_COLOR_CALIBRATE |
n |
Ring R→G→B channel-map probe |
UI_STATE_WALK |
n |
Cycles all six ring states |
AFE_REF_DELAY_MS |
80 |
AEC reference delay compensation |
MWW_THRESHOLD |
0 |
Wake probability threshold (permille, 0 = defer to computer.json cutoff) |
MWW_REFRACTORY_SEC |
2 |
Minimum seconds between wake detections |
MWW_INPUT_GAIN |
30000 |
Digital input gain into wake path (permille, 30×) — stages raw mic peaks of a few hundred LSBs into the trained domain |
MWW_PMAX_FLOOR |
10 |
Anti-transient floor — fire requires best window P(wake) ≥ this (permille, 0 = off) |
MWW_MIN_ABOVE |
3 |
Spread evidence — ≥ this many windows in average must individually reach threshold (0 = off) |
MWW_ARM_DELAY_MS |
250 |
Burst-age arming — suppress fires until gated burst has lasted this long (0 = off) |
MWW_DEBUG_METRICS |
n |
Log sliding-window P(wake) / peak every ~0.5 s (calibration aid) |
MWW_FA_CAPTURE |
n |
Hex-dump 6 s gated audio ring on every wake fire to serial (FAD lines, for harvest — adds ~192 KB PSRAM) |
MWW_SELFTEST |
n |
Boot-time deterministic selftest — stream partitioned test wav through the full wake chain |
Contributor rules — including secrets never get committed — live in AGENTS.md.
Wake word — "computer" via microWakeWord
The device wakes hands-free on "computer" (same word as the computer
server) via a custom microWakeWord streaming classifier running entirely
on-device — no WakeNet / wn9_hiesp ("Hi ESP") remains.
- Model:
tools/wakeword-model/computer.tflite+computer.json— int8mixednet(~59 KB), 40-dim mel front-end (30 ms window / 10 ms stride), stride 3, sliding window 5. Tensor arena 34304 B (PSRAM). Shipped v2 is fine-tuned on ear-verified device captures: 7 positives (4 train / 1 val / 1 test / 1 held-out reference) mixed with 6500 synthetic piper-TTS positives. The held-out real reference peaks atP(wake) = 1.00with operatingprobability_cutoff = 0.50(recall 1.0, 1.875 FAH ontesting_ambient). The synthetic-only predecessor was deaf on real speech (live peaks 0.016–0.04, only its silence floor ever fired). Seetools/wakeword-model/REPRODUCIBILITY.mdfor the full config, seed (20260823), cutoff sweep, and arena derivation. - Engine:
components/mww/(MIT, ours; mel preprocessor derived from Apache-2.0espressif/esp-tflite-microexamples/micro_speech) on the standard TFLM runtime (espressif/esp-tflite-micro+esp-nn). Model + JSON load at boot from themwwflash partition (5000 KB,data/spiffs); build embeds expected size + SHA256 prefix so stale partitions are caught loudly. Flash withtools/flash_models.sh(notidf.py flash):
Payload layout is documented in. $HOME/esp/esp-idf/export.sh sg dialout -c "tools/flash_models.sh" # model + reference wav sg dialout -c "tools/flash_models.sh --wav FILE" # override test wav sg dialout -c "tools/flash_models.sh --image-only" # build/verify onlytools/flash_models.shheader (32 BMWW1header → json → tflite → optional 16 kHz mono test wav). - Training pipeline:
tools/wakeword-training/synthesizes positives withpiper-sample-generator(LibriTTS-R, 5k + 1500 phonetic variants) and trains withOHF-Voice/micro-wake-word @ 4665173via the wrappertrain_computer.py(config_computer.json= synthetic baseline,config_computer_real.json= v2 real-capture fine-tune). Real positives are harvested on-device withCONFIG_MWW_FA_CAPTURE=y+CONFIG_MWW_DEBUG_METRICS=y(each wake fire snapshots[trigger−2 s … trigger+4 s]to a PSRAM ring and hex-dumps it over serial —extract_capture_log.py→ ear-verify →slice_captures.py), then weightedpositives_real: 10.0vs synthetic2.0with explicitreal_splitlists. Full workflow, capture rig, and lessons are indocs/wake-word-pipeline.mdandtools/wakeword-training/README.md. - Runtime: continuous streaming inference — one
mww_features_step+ TFLMInvokeper 480-sample stride (30 ms) whileIDLE; duringLISTENING/SPEAKINGthe ring is drained without inference (stays sample-synced, zero CPU). Trigger firesapp_state_wake_detected()only whenIDLE && !mic-muted, via a sliding-window average crossingMWW_THRESHOLD(or the JSONprobability_cutoffwhen0) plus spread-evidence conditions (MWW_MIN_ABOVE,MWW_PMAX_FLOOR,MWW_ARM_DELAY_MS) and a refractory gate. Audio reaches the model post-AEC, postMWW_INPUT_GAIN(default 30000) — raw mic peaks of a few hundred LSBs are staged into the trained domain; the offline trainer is amplitude-invariant but the device PCAN/NR front-end is not. - Selftest / calibration:
CONFIG_MWW_SELFTEST=ystreams the partitioned test wav through the full chain (bypassing mic) loggingP(wake)per invoke plus a summary — deterministic regression of quantization/PCAN/streaming;CONFIG_MWW_DEBUG_METRICS=yexposes the live sliding average / peak and adaptive noise state for threshold tuning.
Host conformance gate
cd host/tests && python3 -m pytest -q
Must be green before any on-device network work. The suite unit-tests the
wire-format helpers, runs the protocol layer against the shared conformance
harness from the sibling checkout: ../computer/tests/conformance/ fake_realtime_server.py, and host-builds the portable wake-word core
(host/mwwtest → test_mww_features.py / test_mww_trigger.py —
features vs the Python reference oracle, trigger state-machine). No SDK
needed — pure host pytest.
Architecture (one paragraph)
afe_task runs two FreeRTOS tasks (feed/fetch) around the ESP-SR audio
frontend in AEC-only mode (wakenet_init=false,
afe_config_init("MR", …) with empty model list; AFE_REF_DELAY_MS
aligns the playback reference). The enhanced mono fetch path is injected via
mww_inject_pcm() into mww_task (prio 6, 8 KiB), which computes 40-dim mel
features in C and invokes the int8 streaming classifier (computer.tflite)
every 30 ms (every 3 slices) while IDLE; a sliding-window trigger fires the
wake. A detection wakes the app; ws_client opens the WebSocket to
computer, sends session.update immediately on connect to arm server-side
VAD, then shuttles events. player buffers outbound audio in a 256 KB PSRAM
jitter ring (~8.2 s cap) feeding 20 ms slices and does ack accounting (exactly
one output_audio.played per response). app_state is the machine
IDLE → CONNECTING → LISTENING ⇄ SPEAKING → DISCONNECTING; ui_ring maps
states onto the WS2812 x7 ring; buttons polls the TCA9555 keys.
Behavior notes
- Low-power idle (
CONFIG_IDLE_LOW_POWER, default on): while IDLE the speaker amp is disabled, the LED ring is cleared, and the key poll drops to 2 Hz; all restore on wake. Wake-word detection (mic + AEC + mww inference) stays live at all times, and Wi-Fi only ever connects from the first wake onward. Apower_hbtask persists a virtual run clock to NVS every 60 s so each boot logs the previous run's duration + reset reason (power: prev run ... (reset BROWNOUT ...)) ‒ with the brownout detector enabled, a battery discharge self-reports how long the run lasted. Full power architecture, log semantics, and the no-power-meter drain-measurement method: POWER.md. - Server-side VAD: the server commits utterances on silence; the device just streams mic audio while LISTENING.
- Barge-in: speaking over playback makes the server cancel and start a new response; the player drops buffered audio and any chime tail.
- Follow-up window: after each response the mic stays live until the idle
watchdog fires — the device disconnects after 15 s without server events
(
CONFIG_IDLE_DISCONNECT_SEC), so follow-ups must land within ~15 s. - Idle disconnect: after 15 s without server events the device disconnects itself, releasing the server pipeline slot (a second client can connect while this one sits idle).
- session.update on connect arms the server's VAD before any audio flows.
- Wake chime is local: the canonical family cue (
here.wav) is embedded in the firmware and injected into the playback ring when CONNECTING starts. - Wire contract is owned by
computer: see../computer/src/computer/api/openai_realtime/CONFORMANCE.md(canonical — this repo'sdocs/conformance.mdand host tests track it).
Status
Ships the custom "computer" wake-word model (v2, 2026-08-26) — a
microWakeWord-format int8 mixednet distilled from real device captures.
Previous stock wn9_hiesp ("Hi ESP") is gone: esp-sr runs AEC-only
(wakenet_init=false, empty srmodels.bin) and wake originates solely in
components/mww (mww_task). The v2 model (tools/wakeword-model/
computer.tflite 59392 B + computer.json cutoff 0.50) was fine-tuned on
7 ear-verified harvests; the held-out reference peaks at P(wake) = 1.00
and wakes live on the first natural utterance. Promotion path is
copy-computer.tflite+computer.json into tools/wakeword-model/ →
tools/flash_models.sh → reflash app if the embedded SHA prefix changed.
See docs/wake-word-pipeline.md for the harvest → slice → train →
validate → ship workflow and tools/wakeword-model/REPRODUCIBILITY.md for
the bit-exact trainer record.