Tatlock turns take 10-25s, which is a long silence after a request. Speak
a canned "Let me check on that for you, sir" immediately (synthesized once
and cached), then re-assert the thinking state so the device keeps its
effort-face spinner up until the real reply arrives. On the device, guard
the playback-done handler so the filler audio finishing doesn't drop the
spinner back to idle.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LKPbR6DY2JygHbyLjxm7Uu
Adds a way to capture the live LVGL screen off the device and rebuild it
as a PNG on the host, so UI changes can be verified remotely without a
camera. A watcher task polls the USB-serial-JTAG RX for a trigger byte and
streams the current screen as raw RGB565 straight to the USB FIFO (framed
by ###SHOT_BEGIN/END### with a CRC); firmware/tools/device_shot.py and the
device-screenshot skill drive it from the host.
Writing straight to the USB FIFO bypasses the primary UART console, which
at 115200 baud would take ~37s per frame. The whole capability is behind
DESKLOCK_DEVMODE (off by default, enable at deploy time with
-DDESKLOCK_DEVMODE=ON) so production spends no internal RAM on the watcher
and nothing extra runs on the render path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LKPbR6DY2JygHbyLjxm7Uu
Bump raw_buf_almost_empty_thrd 512 -> 1024 right after display start so
the bridge demands a DMA refill with more slack still in the FIFO,
letting it ride out a PSRAM-bus latency spike instead of draining to the
blue underrun colour. Symptom mitigation for the bus-arbitration
starvation; the register isn't otherwise exposed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LKPbR6DY2JygHbyLjxm7Uu
Start the display via bsp_display_start_with_config with a 16 KB LVGL
task stack. The 8 KB default overflows once LV_INV_BUF_SIZE is raised to
128 (the partial-flush path's stack frame scales with it) — the guard
fires as a "Stack protection fault" boot loop otherwise.
Add FACE_LOADTEST (gated off): a load-emulation harness that cycles
isolated load types — a render-only face ladder, light/heavy render plus
a network stream, render plus playback, and all three stacked. Each
phase logs a marker and shows an on-screen label. It pinned the blue
flicker to a DSI PSRAM-bandwidth underrun (the driver logs "underrun
happens"), which is serial-verifiable without eyes on the screen. Kept
for future bus-contention debugging.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LKPbR6DY2JygHbyLjxm7Uu
Root cause (investigation): the BSP default tear-avoid mode is
TRIPLE_PARTIAL. On ESP-IDF 5.5 the MIPI-DSI driver has no
on_frame_buf_complete callback (added in IDF 6.0 -> the boot warning
"buffer-switch...may not function on MIPI DSI"), so the adapter falls
back to on_refresh_done, which fires every refresh (~60Hz) as a fake
vsync. In PARTIAL mode that release path is NOT submit-gated, so when
one LVGL frame takes >1 refresh to render it over-releases the buffer
that is still being scanned out -> LVGL draws into the live front
buffer -> tearing/flicker.
Why it only started with the wake word, and only under load: at idle
(sparse rain) a frame renders in <16.6ms so exactly one submit per
refresh -> harmless. The always-on WakeNet added constant CPU/PSRAM
load that pushed the heavy-rain frames (listening=16, thinking=40
streams) past one refresh -> triggered the over-release. User
correctly identified it as a resource starve exposing the latent bug.
Fix: switch to TRIPLE_FULL, which IS submit-gated even without the
callback (one release per actual submit). Same 3 framebuffers, zero
memory cost, app-side one-liner via bsp_display_start_with_config so
the managed component is untouched. Trade-off: full-screen redraw per
frame; if the rain gets choppy under load, cheaper rain rendering
(canvas / half-rate tick) is the follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Say "Computer" -> chime + listening -> speak -> AFE VAD detects you
stopped -> auto-sends. No taps. Touch still works as a manual override.
- esp-sr 2.4.6 added; wn9_computer_tts model packed into a new "model"
flash partition (MODEL_IN_FLASH). App moved to 8M, model 4M.
- audio.c: replaced the on-demand capture_task with an AFE pipeline —
feed_task is the SOLE mic reader (-> afe->feed); detect_task fetches,
watches wakeup_state for the wake word and vad_state for end of
speech, and forwards AFE-cleaned audio upstream during an utterance.
One mic reader ever.
- short rising chime acknowledges the wake audibly.
Fixes from adversarial review before trusting it:
1. utterance framing (blocking WS sends) moved OFF the AFE fetch thread
onto an app_task event queue (EV_TOUCH/EV_WAKE/EV_SPEECH_END) — a
1.5s send could stall fetch and drop the first ~1.5s of speech.
2. app_task is now the single serializer of start/end -> no TOCTOU
double-start (was: two utterance_start on a tap during wake).
3. VAD accounting resets on every streaming (re)start (wake OR tap),
not just wake -> a tapped utterance can no longer end instantly on
stale silence.
4. chime/reply set s_playing (+DMA tail hold) and detect_task skips the
mic while s_playing -> our own audio no longer streams into STT or
false-triggers the wake at a playback boundary (no AEC yet).
5. NULL-checked AFE create + feed buffer; tasks only start if AFE is up.
Verified on hardware: model loads, AFE inits with the Computer word,
boots and connects clean, no crash/wedge.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause (verified against our exact IDF tree, not the community guess):
the "258" in "sdio_write_task: Failed to send data: 258" is NOT a timeout
(that is 263). 258 = 0x102 = ESP_ERR_INVALID_ARG. On the ESP32-P4, block-
mode CMD53 writes require the SOURCE buffer to be 64-byte (cache-line)
aligned; the IDF sdmmc driver rejects a misaligned source with INVALID_ARG
BEFORE any bus activity. esp_hosts write loop then declares "Unrecoverable
host sdio state" and reboots the whole P4. The audio TX payload is not
64-aligned, so streaming mic audio wedged on the very FIRST frame (which is
exactly what we saw: listening -> instant Failed to send -> reboot).
This also explains why buffer/queue/clock/retry tuning all did nothing: the
write never reached the bus. And why our symptom was instant, not after
~100 writes (the community block-mode-desync theory) — it is the first
misaligned buffer, every time.
Fix: vendored esp_hosted 2.12.11 as an editable local component (overrides
the registry copy) and bounce a misaligned TX payload through one aligned
DMA scratch buffer in hosted_sdio_write_block (port_esp_hosted_host_sdio.c).
TX is serialized by the bus lock so a single static bounce buffer is safe;
freed in hosted_sdio_deinit. Host-only change — no C6 reflash.
VERIFIED ON HARDWARE (autonomous self-test): 40s of continuous mic-audio
upstream streaming — the traffic that previously wedged on the first frame
— ran clean, zero timeouts, zero reboots. A guarded SDIO_TX_SELFTEST harness
is kept (compiled out) for future SDIO stress testing.
Credit: root cause + patch designed via multi-agent investigation; the
precise 258=INVALID_ARG decode (correcting the upstream community timeout
assumption) came from checking our actual esp_err.h.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TOUCH CRASH (the blue screen): app_on_touch ran a blocking WebSocket
send (up to 5s) directly in the LVGL touch callback, stalling the
MIPI-DSI flush into a garbage/blue frame + task-watchdog reboot on
every tap. Now touch_cb only gives a semaphore; a dedicated app_task
does the blocking sends, audio, and face changes off the render
thread. WS send timeout cut 5s to 1.5s as belt-and-braces.
RAIN = REACHABILITY (user request): rain now falls only while the
gateway WebSocket is live (gated on gw_connected in rain_tick). It
drains gracefully on disconnect, resumes on reconnect, a genuine
glanceable reachable signal. Idle density bumped 2 to 4 so
connected-idle reads distinctly from disconnected-black.
CALM THE WEDGE CHURN (user request): transient wifi/WS drops no longer
slam to the x_x error face or a CONNECTING banner. Boot goes straight
to the calm idle face (dry until connected). Only a sustained 30s+
outage escalates to x_x (clock_cb); the ~15s wedge-recovery just shows
a brief rain pause.
Two fixes from adversarial concurrency review before flashing:
- persistent single capture task (was xTaskCreate per utterance; a
rapid re-tap or WS-stop-vs-app-start race could put two readers on
one mic/I2S handle and corrupt the codec)
- reset s_talking on disconnect (app_on_disconnect) so the first tap
after a mid-utterance drop starts fresh, not the stop branch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Night-watch caught the instability in the act at 7min uptime:
E H_SDIO_DRV: sdio_write_task: Failed to send data: 258 (timeout)
E H_SDIO_DRV: Unrecoverable host sdio state -> SW_CPU_RESET
i.e. esp-hosted-mcu#167. Known mitigation: 1-bit SDIO bus (#148),
still ~80x voice bandwidth. Also: recovery reboots no longer ring
the gong (esp_reset_reason gate) - a butler doesn't bong at 3am.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause of the dead data path (assoc/scan/RPC fine, zero data
frames): esp-hosted 1.4.x is formally incompatible with IDF 5.5
(esp-hosted-mcu#47) and the factory C6 slave firmware was ancient
(couldn't even answer a version RPC). Waveshare's examples pin 1.4.* —
do not follow them.
Fix:
- host: espressif/esp_hosted ^2.12 (+ esp_wifi_remote 1.6)
- slave: 2.12.11 network_adapter.bin embedded in the app (c6_ota.c
streams it to the C6 over the SDIO RPC channel at boot when the
reported version is < 2.x; ~10s, no wires, idempotent)
- RAM diet for hosted 2.x's footprint (MEMPOOL_PREFER_SPIRAM,
reduced WIFI_RMT buffers) — without it internal SRAM famine
boot-loops in xTaskCreateStaticPinnedToCore before app_main
- conservative 20MHz SDIO clock for first verified data path
Verified on hardware: DHCP lease (even that healed), 8/8 pings to
router and tower-of-joy at 1-5ms, WebSocket to the gateway connected,
device registered at /devices as desklock-p4 with fw version.
Also: wifi_diag.c L1 SoftAP diagnostic mode (DESKLOCK-DIAG) with
review fixes, boot ping ladder in net.c, static-IP fallback,
face_status() line, docs for the whole failure taxonomy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The device now runs the complete face contract on glass: seven states
with glyph expressions (custom 140px face font), matrix rain as pooled
label streams (22px DejaVu+Noto katakana font), rage kaomoji frames
(64px), orbit arc + elapsed counter wait cues, blink/talk/thought
animations, idle clock (SNTP, Europe/Amsterdam), and the power ladder
(active/ambient 35%/dormant 5% with rain parked).
Voice path: ES7210 mic capture task streams 16k PCM over
esp_websocket_client to the gateway; reply PCM buffers to PSRAM and
plays via the audio task; touch-to-talk (tap to speak, tap to send).
Protocol mapping per docs: thinking->effort, audio->speaking,
error event->rage (auto-composes after 2 loops), WS loss->error face.
Bring-up fixes: 8MB factory partition (fonts overflowed 1.5M),
bsp_display_lock(0) is try-lock in this adapter (use UINT32_MAX),
gong synth yields to feed IDLE0, wifi scan diagnostics (found SSID
case mismatch), esp_websocket_client pinned ~1.3 (1.4 needs IDF>5.5).
Verified on hardware: boots, connects to Wi-Fi and holds an open
WebSocket to the gateway. Sauron (iMac TTY endpoint + ops console)
planned in architecture.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
User verdict: audible and butler-non-intrusive on the bare 2W speaker.
Recorded as the house sound signature in architecture.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Additive gong synthesis: seven inharmonic partials with per-partial
decay (fundamental 165Hz rings 2.6s, highs die in 0.2s), a detuned
pair for shimmer, 45ms distant attack, and three feedback-comb
reflections (95/210/370ms, net gain .93) for a dense stone-space tail.
Normalized to -12dBFS at codec volume 58: quiet but voluminous.
Synthesized into PSRAM at boot (~4s), plays 5s, blocking is fine for
bring-up — audio moves to its own task in phase 2.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
C5->E5 half-second chime with fade-out at boot, right after the face
renders. Codec opens in slave mode, esp_codec_dev write path works.
Phase 1 bring-up is now complete: display, touch, audio, 200MHz PSRAM.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- firmware boots to a rendered LVGL face in 1.57s (eyes + smile + tag)
- sdkconfig: CONFIG_IDF_EXPERIMENTAL_FEATURES=y unlocks SPIRAM 200MHz
(silently degraded to 20MHz before -> MIPI-DSI underrun -> LVGL lock
starvation -> task watchdog); mirrors official 08_lvgl_demo_v9 config
- architecture.md: power management section (user prime concern) —
active/ambient/dormant/night ladder, levers, hard edges, <=300ms wake
- AGENTS/CLAUDE: replace stale bring-up warnings with verified commands
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DeskLock gives the Tatlock butler a face and voice on a Waveshare
ESP32-P4-WIFI6-Touch-LCD-3.4C round display in the living room.
- firmware/: ESP-IDF project targeting esp32p4 with the Waveshare XC BSP
- gateway/: FastAPI voice bridge (faster-whisper STT, Tatlock chat, Piper TTS)
- docs/architecture.md: component design and device<->gateway WS protocol
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>