Files
desklock/README.md
T
jpmschweitzerandClaude Fable 5 ed0e8221f1
Test, Build and Push / test-gateway (push) Successful in 10s
Test, Build and Push / release (push) Skipped
Test, Build and Push / build-gateway (push) Skipped
Phase 2: full LVGL face + Wi-Fi + gateway WebSocket + touch-to-talk
The device now runs the complete face contract on glass: seven states
with glyph expressions (custom 140px face font), matrix rain as pooled
label streams (22px DejaVu+Noto katakana font), rage kaomoji frames
(64px), orbit arc + elapsed counter wait cues, blink/talk/thought
animations, idle clock (SNTP, Europe/Amsterdam), and the power ladder
(active/ambient 35%/dormant 5% with rain parked).

Voice path: ES7210 mic capture task streams 16k PCM over
esp_websocket_client to the gateway; reply PCM buffers to PSRAM and
plays via the audio task; touch-to-talk (tap to speak, tap to send).
Protocol mapping per docs: thinking->effort, audio->speaking,
error event->rage (auto-composes after 2 loops), WS loss->error face.

Bring-up fixes: 8MB factory partition (fonts overflowed 1.5M),
bsp_display_lock(0) is try-lock in this adapter (use UINT32_MAX),
gong synth yields to feed IDLE0, wifi scan diagnostics (found SSID
case mismatch), esp_websocket_client pinned ~1.3 (1.4 needs IDF>5.5).

Verified on hardware: boots, connects to Wi-Fi and holds an open
WebSocket to the gateway. Sauron (iMac TTY endpoint + ops console)
planned in architecture.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 21:42:51 +02:00

4.6 KiB
Raw Blame History

DeskLock

A living-room visual/audio endpoint for Tatlock, the homelab butler. DeskLock gives Tatlock a face and a voice on a Waveshare round touch display: you talk to it, it listens, thinks, and answers — the first line of contact with the butler backend running on tower-of-joy.

Hardware

Waveshare ESP32-P4-WIFI6-Touch-LCD-3.4C

Component Details
SoC ESP32-P4NRW32 — dual-core RISC-V @ 400 MHz + LP core
Memory 32 MB PSRAM (in-package), 32 MB NOR flash
Display 3.4" round IPS, 800×800, MIPI-DSI 2-lane, capacitive touch
Radio ESP32-C6-MINI-1 (Wi-Fi 6 + BLE 5) over SDIO via ESP-Hosted
Audio in Dual onboard microphones + ES7210 echo-cancellation ADC
Audio out ES8311 codec, PH2.0 2-pin speaker connector (8Ω 2W recommended)
Flashing USB-C (hold BOOT during reset for download mode)

The device is currently connected over USB-C directly to tower-of-joy, so build/flash happens on this server.

Architecture

┌──────────────────────┐  WebSocket: PCM audio + JSON events
│  DeskLock device     │◄───────────────────────────────────┐
│  (ESP32-P4)          │                                    │
│  • LVGL face         │     ┌──────────────────────────────┴───────────┐
│  • touch / wake word │     │  DeskLock Gateway  (container, :8600)    │
│  • mic capture + AEC │     │  thin orchestrator — no ML dependencies  │
│  • TTS playback      │     └───────┬──────────────────┬───────────────┘
└──────────────────────┘             │                  │  OpenAI-format HTTP
                                     │                  ▼
                          HTTP (LAN) │       ┌─────────────────────────────┐
                                     ▼       │  Speaches  (container, GPU) │
                     ┌────────────────────┐  │  • STT: faster-whisper      │
                     │  Tatlock (butler)  │  │  • TTS: Kokoro / Piper      │
                     │  tatlock.schweitz. │  │  also usable by Open WebUI, │
                     │  internal :8000    │  │  Home Assistant, …          │
                     └────────────────────┘  └─────────────────────────────┘

Tatlock stays a text-only brain. The gateway orchestrates speech-to-text, chat, and text-to-speech; the Speaches container owns the actual STT/TTS models on the GPU, shared homelab-wide. The device firmware stays thin: audio transport, wake word, and face rendering only. Everything runs on the LAN — no cloud in the voice path.

See docs/architecture.md for the full design.

Repository layout

  • firmware/ — ESP-IDF (C, LVGL 9) application for the ESP32-P4
  • gateway/ — Python FastAPI voice gateway, deployed as a container on tower-of-joy
  • sim/face/ — browser simulator of the face (design source of truth; serve with python3 -m http.server and open index.html, or use ?state=…&nochrome=1 for screenshots)
  • docs/ — architecture and design notes

Roadmap

  1. Bring-up — display, touch, audio, 200 MHz PSRAM, cathedral gong
  2. Face + voice loop (in hardware test) — full LVGL face (7 states, rain, orbit, power ladder), Wi-Fi, WebSocket, touch-to-talk
  3. Voice (touch-to-talk) — tap to talk → gateway → Tatlock → spoken reply
  4. Wake word — esp-sr WakeNet on-device, echo cancellation, barge-in
  5. Polish — Tatlock-initiated notifications, presence, OTA updates

References