jpmschweitzerandClaude Fable 5 cb5826e02b
Test, Build and Push / test-gateway (push) Successful in 12s
Test, Build and Push / release (push) Skipped
Test, Build and Push / build-gateway (push) Skipped
Fork fix: the SDIO wedge is FIXED (esp-hosted-mcu #167)
Root cause (verified against our exact IDF tree, not the community guess):
the "258" in "sdio_write_task: Failed to send data: 258" is NOT a timeout
(that is 263). 258 = 0x102 = ESP_ERR_INVALID_ARG. On the ESP32-P4, block-
mode CMD53 writes require the SOURCE buffer to be 64-byte (cache-line)
aligned; the IDF sdmmc driver rejects a misaligned source with INVALID_ARG
BEFORE any bus activity. esp_hosts write loop then declares "Unrecoverable
host sdio state" and reboots the whole P4. The audio TX payload is not
64-aligned, so streaming mic audio wedged on the very FIRST frame (which is
exactly what we saw: listening -> instant Failed to send -> reboot).

This also explains why buffer/queue/clock/retry tuning all did nothing: the
write never reached the bus. And why our symptom was instant, not after
~100 writes (the community block-mode-desync theory) — it is the first
misaligned buffer, every time.

Fix: vendored esp_hosted 2.12.11 as an editable local component (overrides
the registry copy) and bounce a misaligned TX payload through one aligned
DMA scratch buffer in hosted_sdio_write_block (port_esp_hosted_host_sdio.c).
TX is serialized by the bus lock so a single static bounce buffer is safe;
freed in hosted_sdio_deinit. Host-only change — no C6 reflash.

VERIFIED ON HARDWARE (autonomous self-test): 40s of continuous mic-audio
upstream streaming — the traffic that previously wedged on the first frame
— ran clean, zero timeouts, zero reboots. A guarded SDIO_TX_SELFTEST harness
is kept (compiled out) for future SDIO stress testing.

Credit: root cause + patch designed via multi-agent investigation; the
precise 258=INVALID_ARG decode (correcting the upstream community timeout
assumption) came from checking our actual esp_err.h.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 09:50:27 +02:00

DeskLock

A living-room visual/audio endpoint for Tatlock, the homelab butler. DeskLock gives Tatlock a face and a voice on a Waveshare round touch display: you talk to it, it listens, thinks, and answers — the first line of contact with the butler backend running on tower-of-joy.

Hardware

Waveshare ESP32-P4-WIFI6-Touch-LCD-3.4C

Component Details
SoC ESP32-P4NRW32 — dual-core RISC-V @ 400 MHz + LP core
Memory 32 MB PSRAM (in-package), 32 MB NOR flash
Display 3.4" round IPS, 800×800, MIPI-DSI 2-lane, capacitive touch
Radio ESP32-C6-MINI-1 (Wi-Fi 6 + BLE 5) over SDIO via ESP-Hosted
Audio in Dual onboard microphones + ES7210 echo-cancellation ADC
Audio out ES8311 codec, PH2.0 2-pin speaker connector (8Ω 2W recommended)
Flashing USB-C (hold BOOT during reset for download mode)

The device is currently connected over USB-C directly to tower-of-joy, so build/flash happens on this server.

Architecture

┌──────────────────────┐  WebSocket: PCM audio + JSON events
│  DeskLock device     │◄───────────────────────────────────┐
│  (ESP32-P4)          │                                    │
│  • LVGL face         │     ┌──────────────────────────────┴───────────┐
│  • touch / wake word │     │  DeskLock Gateway  (container, :8600)    │
│  • mic capture + AEC │     │  thin orchestrator — no ML dependencies  │
│  • TTS playback      │     └───────┬──────────────────┬───────────────┘
└──────────────────────┘             │                  │  OpenAI-format HTTP
                                     │                  ▼
                          HTTP (LAN) │       ┌─────────────────────────────┐
                                     ▼       │  Speaches  (container, GPU) │
                     ┌────────────────────┐  │  • STT: faster-whisper      │
                     │  Tatlock (butler)  │  │  • TTS: Kokoro / Piper      │
                     │  tatlock.schweitz. │  │  also usable by Open WebUI, │
                     │  internal :8000    │  │  Home Assistant, …          │
                     └────────────────────┘  └─────────────────────────────┘

Tatlock stays a text-only brain. The gateway orchestrates speech-to-text, chat, and text-to-speech; the Speaches container owns the actual STT/TTS models on the GPU, shared homelab-wide. The device firmware stays thin: audio transport, wake word, and face rendering only. Everything runs on the LAN — no cloud in the voice path.

See docs/architecture.md for the full design.

Repository layout

  • firmware/ — ESP-IDF (C, LVGL 9) application for the ESP32-P4
  • gateway/ — Python FastAPI voice gateway, deployed as a container on tower-of-joy
  • sim/face/ — browser simulator of the face (design source of truth; serve with python3 -m http.server and open index.html, or use ?state=…&nochrome=1 for screenshots)
  • docs/ — architecture and design notes

Roadmap

  1. Bring-up — display, touch, audio, 200 MHz PSRAM, cathedral gong
  2. Face + voice loop (in hardware test) — full LVGL face (7 states, rain, orbit, power ladder), Wi-Fi, WebSocket, touch-to-talk
  3. Voice (touch-to-talk) — tap to talk → gateway → Tatlock → spoken reply
  4. Wake word — esp-sr WakeNet on-device, echo cancellation, barge-in
  5. Polish — Tatlock-initiated notifications, presence, OTA updates

References

S
Description
No description provided
Readme
10 MiB
2026-07-19 12:32:13 +02:00
Languages
C 87.2%
C++ 8.5%
M4 1.6%
CMake 1%
Python 0.8%
Other 0.9%