The homelab is retiring the *.schweitz.internal domain; in-network
machine-to-machine traffic uses docker container names on the
docker-dataplane network. The deployed tatlock-ui stack already overrides
DESKLOCK_TATLOCK_BASE_URL (Tatlock runs on the host), so only the
fallback default changes.
Docs follow: AGENTS.md M2M guidance now points at container names with
*.schweitz.net reserved for browsers, architecture.md drops the retired
domain (the registry name now matches what CI actually pushes since
c477019), and the README diagram loses the stale hostname.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
20 KiB
DeskLock Architecture
Goal
An always-on, glanceable butler face in the living room. You speak to it; it relays your words to Tatlock and speaks the reply back, with a face that reflects what it's doing (idle, listening, thinking, speaking). It is deliberately a thin endpoint: all intelligence lives in Tatlock, all heavy audio processing lives server-side on tower-of-joy. Everything is local — no audio, transcript, or reply ever leaves the LAN.
System overview
┌──────────────────────┐ WebSocket: PCM audio + JSON events
│ DeskLock device │◄───────────────────────────────────┐
│ (ESP32-P4) │ │
│ • LVGL face │ ┌──────────────────────────────┴───────────┐
│ • touch / wake word │ │ DeskLock Gateway (container, :8600) │
│ • mic capture + AEC │ │ thin orchestrator — no ML dependencies │
│ • TTS playback │ └───────┬──────────────────┬───────────────┘
└──────────────────────┘ │ │ OpenAI-format HTTP
│ ▼
HTTP (LAN) │ ┌─────────────────────────────┐
▼ │ Speaches (container, GPU) │
┌────────────────────┐ │ • STT: faster-whisper │
│ Tatlock (butler) │ │ • TTS: Kokoro / Piper │
│ tatlock.schweitz. │ │ also usable by Open WebUI, │
│ internal :8000 │ │ Home Assistant, … │
└────────────────────┘ └─────────────────────────────┘
Tatlock stays a text-only brain. The gateway orchestrates Tatlock's ears and mouth; the speech layer (Speaches) owns the actual STT/TTS models on the GPU.
Components
1. Firmware (firmware/) — ESP32-P4
Responsibilities:
- Face rendering (LVGL 9 on the 800×800 round MIPI-DSI panel via the
waveshare/esp32_p4_wifi6_touch_lcd_xcBSP). Six expression states — see Face design for the visual contract. - Audio capture: dual mics through the ES7210 (hardware echo cancellation reference from the playback path), 16 kHz 16-bit mono PCM.
- Audio playback: ES8311 codec → speaker. Plays PCM streamed from the gateway.
- Transport: a single WebSocket to the gateway carrying binary PCM frames plus JSON
control events (
state,transcript,reply_text, errors). Device reconnects with backoff; face shows a disconnected state when the gateway is unreachable.
On-device speech processing — what runs on the P4 and what deliberately doesn't. The P4 (dual RISC-V @ 400 MHz, 32 MB PSRAM) has a hard ceiling; the split is:
| On-device (planned) | Why |
|---|---|
| Wake word — esp-sr WakeNet (phase 2) | Must be local: always-listening audio should never leave the device until the wake word fires |
| Voice-activity detection (end-of-utterance) | Removes tap-to-stop; cheap on-device |
| Echo cancellation — ES7210 hardware | Enables barge-in while TTS is playing |
| esp-sr MultiNet fixed commands (optional, later) | ~200-phrase closed vocabulary recognized entirely on-device — instant "lights off"-style commands with zero round trip |
Full open-vocabulary STT on-device is out of scope permanently: even whisper-tiny needs hundreds of MB and orders of magnitude more compute than the P4 offers. Anything open-ended goes to the speech layer.
Non-responsibilities: no STT, no TTS, no conversation state. If it's not rendering, recording, or playing, it doesn't belong in firmware.
2. Gateway (gateway/) — container on tower-of-joy, port 8600
A FastAPI service bridging device audio to Tatlock text. It owns orchestration, not models — the container stays a slim pure-Python image with no CUDA/ML dependencies:
- Accepts the device WebSocket (
/ws/voice). - Buffers inbound PCM until end-of-utterance (client-signalled in phase 1; VAD later).
- STT: POST to Speaches
/v1/audio/transcriptions. - Chat: POST the transcript to Tatlock
/v1/chat/completions(http://tatlock:8000, OpenAI-compatible, streaming), maintaining the conversation history so follow-ups have context. - TTS: as Tatlock's token stream completes each sentence, POST it to Speaches
/v1/audio/speechand forward the PCM immediately — see Latency budget.
stt.py / tts.py are pluggable backends selected by config
(DESKLOCK_STT_BACKEND / DESKLOCK_TTS_BACKEND):
speaches(default) — OpenAI-format HTTP to the shared speech container.embedded— in-process faster-whisper / Piper. Kept as a fallback so the gateway can run standalone (dev on a laptop, speech container down), at the cost of a fat image.
The gateway is stateless apart from in-flight conversations; it can restart freely.
3. Speech layer — Speaches (container, GPU)
Speaches (successor to faster-whisper-server)
is a self-hosted, OpenAI-API-compatible speech server: STT via faster-whisper, TTS via
Kokoro/Piper, dynamic model load/offload with a TTL, and a /v1/realtime WebSocket API
we may adopt later for streaming transcription.
- Deployed 2026-07-14:
ghcr.io/speaches-ai/speaches:latest-cudaon host port 8601, withSystran/faster-whisper-small(STT) andspeaches-ai/Kokoro-82M-v1.0-ONNX(TTS — Kokoro is natively 24 kHz, but the gateway requests the 16 kHz device contract directly via Speaches'sample_rateextension, verified live; default voicebm_george, en-GB male). LAN-only like the Tatlock internal route — do not expose through NPM without auth. Register inCONTAINERS.md. - Measured (live round trip through the gateway code, warm): STT ~0.3 s for a ~3 s utterance; TTS ~1.9 s for a ~3 s sentence. Cold start after model TTL offload adds ~5–10 s to the first request.
- Why a shared layer instead of models inside the gateway: one GPU-resident model
instance serves the whole homelab. Open WebUI is currently configured with
AUDIO_STT_ENGINE=openai/AUDIO_TTS_ENGINE=openai(OpenAI cloud) — pointing its audio base URL at Speaches makes it fully local with a config change. Home Assistant can share it too. Meanwhile the gateway image needs no CUDA and rebuilds in seconds. - VRAM budget: RTX 2080 Ti, 11 GB, shared with Ollama (~3.6 GB in use as of
2026-07). whisper
smallat int8 is <1 GB; Kokoro is a few hundred MB. Speaches' model TTL offload keeps idle pressure near zero. If VRAM contention ever bites, faster-whispersmallon CPU is an acceptable fallback (int8, a few seconds per utterance).
4. Tatlock — existing backend (/mnt/media/Projects/tatlock)
Untouched by this project. DeskLock consumes its OpenAI-compatible API over the internal LAN (port 8000, bypassing the Authentik-protected public route). If device auth is needed later, the gateway holds the credential — never the firmware.
Face design
Aesthetic: pure black screen; a face drawn from ASCII/terminal glyphs in green
phosphor (#adffc8 face, dimmer greens for secondary info); Matrix-style digital rain
whose density encodes activity — barely-there drips when idle, a downpour while
Tatlock works. No bitmaps, no skeuomorphism: glyphs only.
Source of truth: sim/face/index.html — a self-contained browser simulator of the
800×800 round panel. Design changes land there first, get approved visually, then get
ported to LVGL. The STATES table in the sim defines the contract:
| State | Eyes | Mouth | Rain | Extra cues |
|---|---|---|---|---|
idle |
- - |
\_/ |
2 slow streams | clock (HH:MM), breathing bob, blinks |
listening |
O O |
o |
16 streams | blinks |
pensive |
· · |
~ |
7 streams | cycling ... thought dots |
effort |
> < |
~ |
40 fast streams | orbit arc on bezel + [ Ns ] elapsed counter, face jitter |
speaking |
^ ^ |
cycles o O - O = o |
14 streams | mouth animates ~150 ms/frame |
rage |
— | — | 34 fast streams | 3-frame kaomoji loop through the eyes slot: (°□°) ┬─┬ → (╯°□°)╯︵ ┻━┻ → ┬─┬ ノ( º_º ノ) — flips the table, then composes itself and puts it back |
error |
x x |
- |
none (rain dies) | face dims to 45% |
Sound signature: the cathedral gong (play_boot_gong in firmware — 220 Hz
inharmonic partial stack, feedback-comb reflections, ~-7 dBFS) is the approved house
sound: "audible and butler-non-intrusive." Plays at boot; planned as the
wake-from-dormant sound. Tune by adjective: tail = taus, cathedral size = echo
delays, depth = fundamental, presence = GONG_PEAK/GONG_VOLUME.
Wait cues are a hard requirement (user-stated): Tatlock turns take 10–25 s, so
effort must always show alive-and-working signals — the orbiting bezel arc, the
elapsed-seconds counter, and max rain. Never a bare static face during a wait, and no
fake progress bars — only honest cues.
Protocol → face mapping: gateway state: thinking → effort; transcription and
other short local waits → pensive; listening/speaking map 1:1; an in-flight
request failure (STT/Tatlock/TTS error) → rage for a few loops, then idle;
WebSocket disconnected → error (quiet, persistent); otherwise idle.
LVGL port notes (for phase 2):
- Drive everything from fixed-step
lv_timers (~30 fps rain tick) — the sim deliberately usessetInterval, notrequestAnimationFrame, to mirror this. - Rain:
lv_canvas(or a pooled label grid) with per-frame fade; orbit arc =lv_arc. - Fonts: generate a large monospace glyph font including the katakana subset used in
GLYPHSvialv_font_conv; the built-inunsciifonts are too small for 800 px. Therageframes additionally need╯ ︵ ┻ ━ ┬ ─ ノ ° □ ºin the subset. - The sim's text glow (
text-shadow) is browser flair — the device renders flat glyphs.
Power management (prime concern)
User requirement: the device idles on a wall 95%+ of its life — low power when nothing is happening is a first-class design goal, not a phase-5 nicety. The face state machine is therefore built around a power ladder from day one:
| Power state | Backlight | Rendering | CPU | Entered when |
|---|---|---|---|---|
active |
100% | full animation | full clock | conversation in progress (listening→speaking) |
ambient |
~35% | idle face, sparse rain | DFS enabled | idle, but activity in the last few minutes |
dormant |
off (or ≤5%) | no redraws — render loop parked | min clock via DFS | no voice/touch for N min (default 10) |
night |
off, panel sleep | none | min clock | schedule or "goodnight" command |
Levers, in order of impact:
- Backlight — this is an IPS LCD: black pixels still burn backlight (unlike OLED),
so brightness is the dominant lever.
bsp_display_brightness_set()drives it. - Render idleness — rain off and animations parked means LVGL stops producing frames, which is what lets DFS actually reach its floor.
- DFS / power management (
CONFIG_PM_ENABLE) — automatic frequency scaling when tasks are quiet. Note the MIPI-DSI constraint below. - Radio — the C6 runs Wi-Fi modem power-save; the gateway WebSocket widens its
ping interval when the device reports
dormant.
Wake triggers (any → ambient/active): wake word (phase 5), touch (from phase 2),
local VAD "someone is speaking" pre-warm, a gateway-initiated event (butler wants to
say something), scheduled morning end of night.
Hard edges — what limits how low we can go:
- The hands-free promise sets the power floor. Wake word requires mics + the AFE pipeline running continuously; deep sleep is permanently off the table while the device promises to answer its name. The floor is "CPU lightly loaded at min clock, radios in power-save, backlight off."
- DSI needs clocks while the panel is active — the deepest CPU savings only unlock
in
dormant/nightwhen the panel stops being refreshed (panel sleep / blank). - Wake latency budget: ≤ ~300 ms from trigger to visible face (backlight ramp + first render). Anything slower reads as "it's off," which kills the butler illusion.
- Touch stays powered in all states except possibly
night— its idle draw is negligible and tap-to-wake must always work. - No invented numbers: actual draw gets measured with a USB power meter at each
phase; working target is
dormant≤ ⅓ ofactive. (Always-on device: every watt saved ≈ 9 kWh/year.)
Implementation order: backlight dimming + dormant timeout + touch wake land in
phase 2 with the face state machine (timeout-driven); voice-linked triggers upgrade
it in phase 5.
Latency budget & streaming
Measured/known numbers that shape the design (Tatlock figures per tatlock CLAUDE.md, GPU-resident benchmarks of 2026-07-14, gemma4:e2b at ~100 tok/s):
| Stage | Cost |
|---|---|
STT (Speaches whisper small) |
~0.3 s warm (measured) |
| TTS (Speaches Kokoro) | ~1.9 s per ~3 s sentence, warm (measured) |
| Tatlock Steward analysis | ~6 s warm |
| Tatlock, full local flow | 11–25 s end-to-end (librarian-routed ~20–25 s) |
| Tatlock cold start (>2 h idle) | +~8 s (OLLAMA_KEEP_ALIVE=2h) |
(Older "~35 s Steward / ~2 min flow" figures were from a CPU-only driver-mismatch era — do not plan against them.)
Speech is not the bottleneck — Tatlock is, by one to two orders of magnitude. Constraints this imposes:
- The gateway must consume Tatlock's streaming response and synthesize
sentence-by-sentence, forwarding audio as each sentence is ready. The device starts
speaking after the first sentence instead of waiting for the full reply — with
streaming, first audio should land roughly at Steward-time + first-sentence-time,
well under the 11–25 s full-flow figure. The WS protocol already supports this: one
audio_start… PCM …audio_endenvelope with chunks arriving as they're synthesized — the device just plays a continuous stream. - The
thinkingface state is a first-class feature, not decoration — it's what makes a 10–25 s Tatlock turn feel intentional instead of broken. Consider progress cues (e.g. surface Tatlock's reasoning summaries on-screen) later. - A fast lane may eventually be needed: MultiNet on-device commands for instant home-automation phrases, and/or a low-latency intent path in Tatlock itself. Out of scope for now, but don't design it out.
WebSocket protocol (device ↔ gateway)
Binary frames: raw 16 kHz s16le mono PCM (mic upstream, TTS downstream). Text frames: JSON control messages.
device → gateway: {"type": "utterance_start"}
device → gateway: <binary PCM frames>
device → gateway: {"type": "utterance_end"}
gateway → device: {"type": "state", "value": "thinking"}
gateway → device: {"type": "transcript", "text": "..."}
gateway → device: {"type": "reply_text", "text": "..."}
gateway → device: {"type": "audio_start", "sample_rate": 16000}
gateway → device: <binary PCM frames> (may arrive sentence-by-sentence; play as a stream)
gateway → device: {"type": "audio_end"}
gateway → device: {"type": "command", "action": "volume_up"} (LLM-bypass; see below)
command (gateway → device) is an alternative to the reply path: when the
gateway recognizes a simple device command in the transcript (volume/mute), it
sends a command instead of calling Tatlock — no reply_text/audio — then returns
to idle. Actions: volume_up, volume_down, mute, unmute, and volume_set
with an extra "level" field (0–11, the on-device volume scale). Matched by the
gateway's commands.py; applied on the device in gw_client.c → face.c.
Planned additions (documented before implemented, here first):
reply_delta(gateway → device): incremental reply text for on-screen streaming while audio is synthesized.- An interrupt event (device → gateway) for barge-in during playback (phase 4).
Keep this protocol documented here and mirrored in firmware/ and gateway/ constants —
it is the one contract between the two halves of the repo.
Key decisions & rationale
- ESP-IDF native (not Arduino/ESPHome): the P4 + MIPI-DSI + esp-sr stack is only first-class in ESP-IDF; Waveshare recommends it, and the BSP targets it.
- Server-side STT/TTS, device does wake word + VAD + AEC only: server whisper is dramatically better than anything embeddable, the GPU is already there, and the P4 physically can't run open-vocabulary STT. Wake word must be on-device (privacy: no audio leaves the device until it fires).
- STT/TTS as a shared Speaches service (not embedded in the gateway): one model
instance for the whole homelab (DeskLock, Open WebUI, potentially HA), slim gateway
image, models upgradable independently.
embeddedbackend retained as a dev/fallback mode.- Rejected — Wyoming protocol containers (
wyoming-faster-whisper/wyoming-piper): native to Home Assistant's ecosystem, but Tatlock and Open WebUI already speak OpenAI format, so Speaches fits the lab better. Revisit only if HA Assist becomes a first-class consumer. - Rejected — cloud STT/TTS: violates the local-first premise; also adds WAN latency and per-minute cost.
- Rejected — Wyoming protocol containers (
- Separate gateway (not extending Tatlock): keeps Tatlock's API text-only and clean; audio concerns (codecs, VAD, streaming, sentence segmentation) stay at the edge. The gateway is also where a future second endpoint (kitchen, office) would connect.
- Monorepo: the WS protocol couples firmware and gateway; versioning them together avoids contract drift.
Additional endpoints — sauron (planned)
sauron is an old iMac (Linux, text-only console) that will run the same UX as a second butler endpoint plus an ops console. Because the gateway protocol is endpoint-agnostic and each WS connection gets its own conversation, extra endpoints are architecturally free.
- Client: a terminal UI (
clients/sauron/, Python + curses/textual) — the face design is already ASCII, so a TTY renders it natively: glyph face states, character-cell matrix rain, green-on-black. Audio via ALSA (arecord/aplay-level, 16 kHz mono PCM), same WebSocket protocol, same state machine. - Screensaver model (user-confirmed): the face+voice layer is the idle mode — full-screen butler when nobody's working. The workspace mode is an SSH ops console: live stats and remote control of tower-of-joy and forge. Any keypress drops from face to console; idle timeout (and wake word later) raises the face again. Voice stays available in both modes.
- Not started — planned after the device reaches phase 4/5. No browser/kiosk stack needed unless we later want the glow.
CI & deployment
Gitea Actions (.gitea/workflows/build.yml), following the tatlock/tatlock-ui pattern:
- Every push to
main: lint + tests for the gateway (Python 3.12). - Version tags (
v0.1.0, …): tests, then buildgateway/intogit.schweitz.net/jpmschweitzer/desklock-gateway:{latest,tag}, push to the Gitea registry, create a release, and trigger Watchtower to roll the running container. - Required repo/org secrets:
REGISTRY_USER,REGISTRY_PASSWORD,WATCHTOWER_HTTP_API_TOKEN(same trio tatlock uses). - The gateway is a service in the
tatlock-uiPortainer stack (system-admin-toj/containers/stacks/tatlock-ui.yml, registered inCONTAINERS.md): it sharesdocker-dataplanewith Speaches (service-name URLhttp://speaches:8000) and reaches the host-run Tatlock via LAN IP.
Firmware is not containerized: it's flashed over USB (idf.py flash), with OTA planned
for phase 5.
Versioning: 0.x while interfaces are still moving — roughly one minor bump per
roadmap phase (0.1.x gateway server-side, 0.2.x first device firmware, 0.3.x
hands-free). v1.0.0 is reserved for the wall milestone: the device mounted in the
living room, talking to Tatlock end to end.