Every figure in the latency budget was stale, in both directions. TTS was listed at ~1.9 s per sentence but measures ~0.24 s warm for 4.5 s of audio; the full Tatlock flow was listed at 11-25 s but measures ~10-13 s for simple turns. Both sets of numbers predate the current model. The VRAM section now carries real figures and the reason they matter: on 2026-08-07 Tatlock ran against a 9.3 GB model, leaving 7 MiB free, and every transcription failed with CUDA out of memory while the Speaches container still reported healthy. The budget is the constraint, not slack. Also replaces the retired tatlock.schweitz.internal hostname in the topology diagram with the docker container name. Co-Authored-By: Claude <noreply@anthropic.com>
DeskLock
A living-room visual/audio endpoint for Tatlock, the homelab butler. DeskLock gives Tatlock a face and a voice on a Waveshare round touch display: you talk to it, it listens, thinks, and answers — the first line of contact with the butler backend running on tower-of-joy.
Hardware
Waveshare ESP32-P4-WIFI6-Touch-LCD-3.4C
| Component | Details |
|---|---|
| SoC | ESP32-P4NRW32 — dual-core RISC-V @ 400 MHz + LP core |
| Memory | 32 MB PSRAM (in-package), 32 MB NOR flash |
| Display | 3.4" round IPS, 800×800, MIPI-DSI 2-lane, capacitive touch |
| Radio | ESP32-C6-MINI-1 (Wi-Fi 6 + BLE 5) over SDIO via ESP-Hosted |
| Audio in | Dual onboard microphones + ES7210 echo-cancellation ADC |
| Audio out | ES8311 codec, PH2.0 2-pin speaker connector (8Ω 2W recommended) |
| Flashing | USB-C (hold BOOT during reset for download mode) |
The device is currently connected over USB-C directly to tower-of-joy, so build/flash happens on this server.
Architecture
┌──────────────────────┐ WebSocket: PCM audio + JSON events
│ DeskLock device │◄───────────────────────────────────┐
│ (ESP32-P4) │ │
│ • LVGL face │ ┌──────────────────────────────┴───────────┐
│ • touch / wake word │ │ DeskLock Gateway (container, :8600) │
│ • mic capture + AEC │ │ thin orchestrator — no ML dependencies │
│ • TTS playback │ └───────┬──────────────────┬───────────────┘
└──────────────────────┘ │ │ OpenAI-format HTTP
│ ▼
HTTP (LAN) │ ┌─────────────────────────────┐
▼ │ Speaches (container, GPU) │
┌────────────────────┐ │ • STT: faster-whisper │
│ Tatlock (butler) │ │ • TTS: Kokoro / Piper │
│ http://tatlock │ │ also usable by Open WebUI, │
│ :8000 │ │ Home Assistant, … │
└────────────────────┘ └─────────────────────────────┘
Tatlock stays a text-only brain. The gateway orchestrates speech-to-text, chat, and text-to-speech; the Speaches container owns the actual STT/TTS models on the GPU, shared homelab-wide. The device firmware stays thin: audio transport, wake word, and face rendering only. Everything runs on the LAN — no cloud in the voice path.
See docs/architecture.md for the full design.
Repository layout
firmware/— ESP-IDF (C, LVGL 9) application for the ESP32-P4gateway/— Python FastAPI voice gateway, deployed as a container on tower-of-joysim/face/— browser simulator of the face (design source of truth; serve withpython3 -m http.serverand openindex.html, or use?state=…&nochrome=1for screenshots)docs/— architecture and design notes
Roadmap
Bring-up✅ — display, touch, audio, 200 MHz PSRAM, cathedral gong- Face + voice loop (in hardware test) — full LVGL face (7 states, rain, orbit, power ladder), Wi-Fi, WebSocket, touch-to-talk
- Voice (touch-to-talk) — tap to talk → gateway → Tatlock → spoken reply
- Wake word — esp-sr WakeNet on-device, echo cancellation, barge-in
- Polish — Tatlock-initiated notifications, presence, OTA updates