Files
desklock/docs/architecture.md
T
jpmschweitzerandClaude Fable 5 ed0e8221f1
Test, Build and Push / test-gateway (push) Successful in 10s
Test, Build and Push / release (push) Skipped
Test, Build and Push / build-gateway (push) Skipped
Phase 2: full LVGL face + Wi-Fi + gateway WebSocket + touch-to-talk
The device now runs the complete face contract on glass: seven states
with glyph expressions (custom 140px face font), matrix rain as pooled
label streams (22px DejaVu+Noto katakana font), rage kaomoji frames
(64px), orbit arc + elapsed counter wait cues, blink/talk/thought
animations, idle clock (SNTP, Europe/Amsterdam), and the power ladder
(active/ambient 35%/dormant 5% with rain parked).

Voice path: ES7210 mic capture task streams 16k PCM over
esp_websocket_client to the gateway; reply PCM buffers to PSRAM and
plays via the audio task; touch-to-talk (tap to speak, tap to send).
Protocol mapping per docs: thinking->effort, audio->speaking,
error event->rage (auto-composes after 2 loops), WS loss->error face.

Bring-up fixes: 8MB factory partition (fonts overflowed 1.5M),
bsp_display_lock(0) is try-lock in this adapter (use UINT32_MAX),
gong synth yields to feed IDLE0, wifi scan diagnostics (found SSID
case mismatch), esp_websocket_client pinned ~1.3 (1.4 needs IDF>5.5).

Verified on hardware: boots, connects to Wi-Fi and holds an open
WebSocket to the gateway. Sauron (iMac TTY endpoint + ops console)
planned in architecture.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 21:42:51 +02:00

346 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DeskLock Architecture
## Goal
An always-on, glanceable butler face in the living room. You speak to it; it relays your
words to Tatlock and speaks the reply back, with a face that reflects what it's doing
(idle, listening, thinking, speaking). It is deliberately a *thin* endpoint: all
intelligence lives in Tatlock, all heavy audio processing lives server-side on
tower-of-joy. Everything is local — no audio, transcript, or reply ever leaves the LAN.
## System overview
```
┌──────────────────────┐ WebSocket: PCM audio + JSON events
│ DeskLock device │◄───────────────────────────────────┐
│ (ESP32-P4) │ │
│ • LVGL face │ ┌──────────────────────────────┴───────────┐
│ • touch / wake word │ │ DeskLock Gateway (container, :8600) │
│ • mic capture + AEC │ │ thin orchestrator — no ML dependencies │
│ • TTS playback │ └───────┬──────────────────┬───────────────┘
└──────────────────────┘ │ │ OpenAI-format HTTP
│ ▼
HTTP (LAN) │ ┌─────────────────────────────┐
▼ │ Speaches (container, GPU) │
┌────────────────────┐ │ • STT: faster-whisper │
│ Tatlock (butler) │ │ • TTS: Kokoro / Piper │
│ tatlock.schweitz. │ │ also usable by Open WebUI, │
│ internal :8000 │ │ Home Assistant, … │
└────────────────────┘ └─────────────────────────────┘
```
Tatlock stays a text-only brain. The **gateway** orchestrates Tatlock's ears and mouth;
the **speech layer** (Speaches) owns the actual STT/TTS models on the GPU.
## Components
### 1. Firmware (`firmware/`) — ESP32-P4
Responsibilities:
- **Face rendering** (LVGL 9 on the 800×800 round MIPI-DSI panel via the
`waveshare/esp32_p4_wifi6_touch_lcd_xc` BSP). Six expression states — see
[Face design](#face-design) for the visual contract.
- **Audio capture**: dual mics through the ES7210 (hardware echo cancellation reference
from the playback path), 16 kHz 16-bit mono PCM.
- **Audio playback**: ES8311 codec → speaker. Plays PCM streamed from the gateway.
- **Transport**: a single WebSocket to the gateway carrying binary PCM frames plus JSON
control events (`state`, `transcript`, `reply_text`, errors). Device reconnects with
backoff; face shows a disconnected state when the gateway is unreachable.
**On-device speech processing — what runs on the P4 and what deliberately doesn't.**
The P4 (dual RISC-V @ 400 MHz, 32 MB PSRAM) has a hard ceiling; the split is:
| On-device (planned) | Why |
|---------------------|-----|
| Wake word — esp-sr WakeNet (phase 2) | Must be local: always-listening audio should never leave the device until the wake word fires |
| Voice-activity detection (end-of-utterance) | Removes tap-to-stop; cheap on-device |
| Echo cancellation — ES7210 hardware | Enables barge-in while TTS is playing |
| esp-sr MultiNet fixed commands (optional, later) | ~200-phrase closed vocabulary recognized entirely on-device — instant "lights off"-style commands with zero round trip |
Full open-vocabulary STT on-device is **out of scope permanently**: even whisper-tiny
needs hundreds of MB and orders of magnitude more compute than the P4 offers. Anything
open-ended goes to the speech layer.
Non-responsibilities: no STT, no TTS, no conversation state. If it's not rendering,
recording, or playing, it doesn't belong in firmware.
### 2. Gateway (`gateway/`) — container on tower-of-joy, port 8600
A FastAPI service bridging device audio to Tatlock text. It owns *orchestration*, not
models — the container stays a slim pure-Python image with no CUDA/ML dependencies:
1. Accepts the device WebSocket (`/ws/voice`).
2. Buffers inbound PCM until end-of-utterance (client-signalled in phase 1; VAD later).
3. **STT**: POST to Speaches `/v1/audio/transcriptions`.
4. **Chat**: POST the transcript to Tatlock `/v1/chat/completions`
(`http://tatlock.schweitz.internal:8000`, OpenAI-compatible, **streaming**),
maintaining the conversation history so follow-ups have context.
5. **TTS**: as Tatlock's token stream completes each sentence, POST it to Speaches
`/v1/audio/speech` and forward the PCM immediately — see
[Latency budget](#latency-budget--streaming).
`stt.py` / `tts.py` are pluggable backends selected by config
(`DESKLOCK_STT_BACKEND` / `DESKLOCK_TTS_BACKEND`):
- `speaches` (default) — OpenAI-format HTTP to the shared speech container.
- `embedded` — in-process faster-whisper / Piper. Kept as a fallback so the gateway can
run standalone (dev on a laptop, speech container down), at the cost of a fat image.
The gateway is stateless apart from in-flight conversations; it can restart freely.
### 3. Speech layer — Speaches (container, GPU)
[Speaches](https://github.com/speaches-ai/speaches) (successor to faster-whisper-server)
is a self-hosted, OpenAI-API-compatible speech server: STT via faster-whisper, TTS via
Kokoro/Piper, dynamic model load/offload with a TTL, and a `/v1/realtime` WebSocket API
we may adopt later for streaming transcription.
- **Deployed 2026-07-14**: `ghcr.io/speaches-ai/speaches:latest-cuda` on host port
**8601**, with `Systran/faster-whisper-small` (STT) and
`speaches-ai/Kokoro-82M-v1.0-ONNX` (TTS — Kokoro is natively 24 kHz, but the
gateway requests the 16 kHz device contract directly via Speaches' `sample_rate`
extension, verified live; default voice `bm_george`, en-GB male). LAN-only like the
Tatlock internal route — do not expose through NPM without auth. Register in
`CONTAINERS.md`.
- **Measured** (live round trip through the gateway code, warm): STT ~0.3 s for a
~3 s utterance; TTS ~1.9 s for a ~3 s sentence. Cold start after model TTL offload
adds ~510 s to the first request.
- **Why a shared layer instead of models inside the gateway**: one GPU-resident model
instance serves the whole homelab. Open WebUI is currently configured with
`AUDIO_STT_ENGINE=openai` / `AUDIO_TTS_ENGINE=openai` (OpenAI *cloud*) — pointing its
audio base URL at Speaches makes it fully local with a config change. Home Assistant
can share it too. Meanwhile the gateway image needs no CUDA and rebuilds in seconds.
- **VRAM budget**: RTX 2080 Ti, 11 GB, shared with Ollama (~3.6 GB in use as of
2026-07). whisper `small` at int8 is <1 GB; Kokoro is a few hundred MB. Speaches'
model TTL offload keeps idle pressure near zero. If VRAM contention ever bites,
faster-whisper `small` on CPU is an acceptable fallback (int8, a few seconds per
utterance).
### 4. Tatlock — existing backend (`/mnt/media/Projects/tatlock`)
Untouched by this project. DeskLock consumes its OpenAI-compatible API over the internal
LAN (port 8000, bypassing the Authentik-protected public route). If device auth is needed
later, the gateway holds the credential — never the firmware.
## Face design
**Aesthetic**: pure black screen; a face drawn from ASCII/terminal glyphs in green
phosphor (`#adffc8` face, dimmer greens for secondary info); Matrix-style digital rain
whose **density encodes activity** — barely-there drips when idle, a downpour while
Tatlock works. No bitmaps, no skeuomorphism: glyphs only.
**Source of truth**: `sim/face/index.html` — a self-contained browser simulator of the
800×800 round panel. Design changes land there first, get approved visually, then get
ported to LVGL. The `STATES` table in the sim defines the contract:
| State | Eyes | Mouth | Rain | Extra cues |
|-------|------|-------|------|------------|
| `idle` | `- -` | `\_/` | 2 slow streams | clock (HH:MM), breathing bob, blinks |
| `listening` | `O O` | `o` | 16 streams | blinks |
| `pensive` | `· ·` | `~` | 7 streams | cycling `...` thought dots |
| `effort` | `> <` | `~` | 40 fast streams | **orbit arc on bezel + `[ Ns ]` elapsed counter**, face jitter |
| `speaking` | `^ ^` | cycles `o O - O = o` | 14 streams | mouth animates ~150 ms/frame |
| `rage` | — | — | 34 fast streams | 3-frame kaomoji loop through the eyes slot: `(°□°) ┬─┬``(╯°□°)╯︵ ┻━┻``┬─┬ ( º_º )` — flips the table, then composes itself and puts it back |
| `error` | `x x` | `-` | none (rain dies) | face dims to 45% |
**Sound signature**: the cathedral gong (`play_boot_gong` in firmware — 220 Hz
inharmonic partial stack, feedback-comb reflections, ~-7 dBFS) is the approved house
sound: "audible and butler-non-intrusive." Plays at boot; planned as the
wake-from-`dormant` sound. Tune by adjective: tail = `tau`s, cathedral size = echo
delays, depth = fundamental, presence = `GONG_PEAK`/`GONG_VOLUME`.
**Wait cues are a hard requirement** (user-stated): Tatlock turns take 1025 s, so
`effort` must always show *alive-and-working* signals — the orbiting bezel arc, the
elapsed-seconds counter, and max rain. Never a bare static face during a wait, and no
fake progress bars — only honest cues.
**Protocol → face mapping**: gateway `state: thinking``effort`; transcription and
other short local waits → `pensive`; `listening`/`speaking` map 1:1; an in-flight
request failure (STT/Tatlock/TTS error) → `rage` for a few loops, then `idle`;
WebSocket disconnected → `error` (quiet, persistent); otherwise `idle`.
**LVGL port notes** (for phase 2):
- Drive everything from fixed-step `lv_timer`s (~30 fps rain tick) — the sim
deliberately uses `setInterval`, not `requestAnimationFrame`, to mirror this.
- Rain: `lv_canvas` (or a pooled label grid) with per-frame fade; orbit arc = `lv_arc`.
- Fonts: generate a large monospace glyph font including the katakana subset used in
`GLYPHS` via `lv_font_conv`; the built-in `unscii` fonts are too small for 800 px.
The `rage` frames additionally need `╯ ︵ ┻ ━ ┬ ─ ノ ° □ º` in the subset.
- The sim's text glow (`text-shadow`) is browser flair — the device renders flat glyphs.
## Power management (prime concern)
User requirement: the device idles on a wall 95%+ of its life — low power when nothing
is happening is a first-class design goal, not a phase-5 nicety. The face state machine
is therefore built around a **power ladder** from day one:
| Power state | Backlight | Rendering | CPU | Entered when |
|-------------|-----------|-----------|-----|--------------|
| `active` | 100% | full animation | full clock | conversation in progress (listening→speaking) |
| `ambient` | ~35% | idle face, sparse rain | DFS enabled | idle, but activity in the last few minutes |
| `dormant` | off (or ≤5%) | **no redraws** — render loop parked | min clock via DFS | no voice/touch for N min (default 10) |
| `night` | off, panel sleep | none | min clock | schedule or "goodnight" command |
Levers, in order of impact:
1. **Backlight** — this is an IPS LCD: black pixels still burn backlight (unlike OLED),
so brightness is the dominant lever. `bsp_display_brightness_set()` drives it.
2. **Render idleness** — rain off and animations parked means LVGL stops producing
frames, which is what lets DFS actually reach its floor.
3. **DFS / power management** (`CONFIG_PM_ENABLE`) — automatic frequency scaling when
tasks are quiet. Note the MIPI-DSI constraint below.
4. **Radio** — the C6 runs Wi-Fi modem power-save; the gateway WebSocket widens its
ping interval when the device reports `dormant`.
Wake triggers (any → `ambient`/`active`): wake word (phase 5), touch (from phase 2),
local VAD "someone is speaking" pre-warm, a gateway-initiated event (butler wants to
say something), scheduled morning end of `night`.
**Hard edges — what limits how low we can go:**
- **The hands-free promise sets the power floor.** Wake word requires mics + the AFE
pipeline running continuously; deep sleep is permanently off the table while the
device promises to answer its name. The floor is "CPU lightly loaded at min clock,
radios in power-save, backlight off."
- **DSI needs clocks while the panel is active** — the deepest CPU savings only unlock
in `dormant`/`night` when the panel stops being refreshed (panel sleep / blank).
- **Wake latency budget: ≤ ~300 ms** from trigger to visible face (backlight ramp +
first render). Anything slower reads as "it's off," which kills the butler illusion.
- **Touch stays powered** in all states except possibly `night` — its idle draw is
negligible and tap-to-wake must always work.
- **No invented numbers**: actual draw gets measured with a USB power meter at each
phase; working target is `dormant` ≤ ⅓ of `active`. (Always-on device: every watt
saved ≈ 9 kWh/year.)
Implementation order: backlight dimming + `dormant` timeout + touch wake land in
**phase 2** with the face state machine (timeout-driven); voice-linked triggers upgrade
it in phase 5.
## Latency budget & streaming
Measured/known numbers that shape the design (Tatlock figures per tatlock CLAUDE.md,
GPU-resident benchmarks of 2026-07-14, gemma4:e2b at ~100 tok/s):
| Stage | Cost |
|-------|------|
| STT (Speaches whisper `small`) | ~0.3 s warm (measured) |
| TTS (Speaches Kokoro) | ~1.9 s per ~3 s sentence, warm (measured) |
| Tatlock Steward analysis | ~6 s warm |
| **Tatlock, full local flow** | **1125 s end-to-end** (librarian-routed ~2025 s) |
| Tatlock cold start (>2 h idle) | +~8 s (`OLLAMA_KEEP_ALIVE=2h`) |
(Older "~35 s Steward / ~2 min flow" figures were from a CPU-only driver-mismatch era —
do not plan against them.)
Speech is not the bottleneck — **Tatlock is**, by one to two orders of magnitude.
Constraints this imposes:
1. **The gateway must consume Tatlock's streaming response and synthesize
sentence-by-sentence**, forwarding audio as each sentence is ready. The device starts
speaking after the first sentence instead of waiting for the full reply — with
streaming, first audio should land roughly at Steward-time + first-sentence-time,
well under the 1125 s full-flow figure. The WS protocol already supports this: one
`audio_start` … PCM … `audio_end` envelope with chunks arriving as they're
synthesized — the device just plays a continuous stream.
2. **The `thinking` face state is a first-class feature**, not decoration — it's what
makes a 1025 s Tatlock turn feel intentional instead of broken. Consider progress
cues (e.g. surface Tatlock's reasoning summaries on-screen) later.
3. A **fast lane** may eventually be needed: MultiNet on-device commands for instant
home-automation phrases, and/or a low-latency intent path in Tatlock itself. Out of
scope for now, but don't design it out.
## WebSocket protocol (device ↔ gateway)
Binary frames: raw 16 kHz s16le mono PCM (mic upstream, TTS downstream).
Text frames: JSON control messages.
```
device → gateway: {"type": "utterance_start"}
device → gateway: <binary PCM frames>
device → gateway: {"type": "utterance_end"}
gateway → device: {"type": "state", "value": "thinking"}
gateway → device: {"type": "transcript", "text": "..."}
gateway → device: {"type": "reply_text", "text": "..."}
gateway → device: {"type": "audio_start", "sample_rate": 16000}
gateway → device: <binary PCM frames> (may arrive sentence-by-sentence; play as a stream)
gateway → device: {"type": "audio_end"}
```
Planned additions (documented before implemented, here first):
- `reply_delta` (gateway → device): incremental reply text for on-screen streaming while
audio is synthesized.
- An interrupt event (device → gateway) for barge-in during playback (phase 4).
Keep this protocol documented here and mirrored in `firmware/` and `gateway/` constants —
it is the one contract between the two halves of the repo.
## Key decisions & rationale
- **ESP-IDF native (not Arduino/ESPHome)**: the P4 + MIPI-DSI + esp-sr stack is only
first-class in ESP-IDF; Waveshare recommends it, and the BSP targets it.
- **Server-side STT/TTS, device does wake word + VAD + AEC only**: server whisper is
dramatically better than anything embeddable, the GPU is already there, and the P4
physically can't run open-vocabulary STT. Wake word must be on-device (privacy: no
audio leaves the device until it fires).
- **STT/TTS as a shared Speaches service (not embedded in the gateway)**: one model
instance for the whole homelab (DeskLock, Open WebUI, potentially HA), slim gateway
image, models upgradable independently. `embedded` backend retained as a dev/fallback
mode.
- *Rejected — Wyoming protocol containers* (`wyoming-faster-whisper`/`wyoming-piper`):
native to Home Assistant's ecosystem, but Tatlock and Open WebUI already speak
OpenAI format, so Speaches fits the lab better. Revisit only if HA Assist becomes a
first-class consumer.
- *Rejected — cloud STT/TTS*: violates the local-first premise; also adds WAN latency
and per-minute cost.
- **Separate gateway (not extending Tatlock)**: keeps Tatlock's API text-only and clean;
audio concerns (codecs, VAD, streaming, sentence segmentation) stay at the edge. The
gateway is also where a future second endpoint (kitchen, office) would connect.
- **Monorepo**: the WS protocol couples firmware and gateway; versioning them together
avoids contract drift.
## Additional endpoints — sauron (planned)
**sauron** is an old iMac (Linux, text-only console) that will run the same UX as a
second butler endpoint plus an ops console. Because the gateway protocol is
endpoint-agnostic and each WS connection gets its own conversation, extra endpoints
are architecturally free.
- **Client**: a terminal UI (`clients/sauron/`, Python + curses/textual) — the face
design is already ASCII, so a TTY renders it natively: glyph face states,
character-cell matrix rain, green-on-black. Audio via ALSA (arecord/aplay-level,
16 kHz mono PCM), same WebSocket protocol, same state machine.
- **Screensaver model** (user-confirmed): the face+voice layer is the *idle mode*
full-screen butler when nobody's working. The *workspace mode* is an SSH ops console:
live stats and remote control of tower-of-joy and forge. Any keypress drops from face
to console; idle timeout (and wake word later) raises the face again. Voice stays
available in both modes.
- **Not started** — planned after the device reaches phase 4/5. No browser/kiosk stack
needed unless we later want the glow.
## CI & deployment
Gitea Actions (`.gitea/workflows/build.yml`), following the tatlock/tatlock-ui pattern:
- **Every push to `main`**: lint + tests for the gateway (Python 3.12).
- **Version tags (`v0.1.0`, …)**: tests, then build `gateway/` into
`git.schweitz.internal/jpmschweitzer/desklock-gateway:{latest,tag}`, push to the
Gitea registry, create a release, and trigger Watchtower to roll the running
container.
- Required repo/org secrets: `REGISTRY_USER`, `REGISTRY_PASSWORD`,
`WATCHTOWER_HTTP_API_TOKEN` (same trio tatlock uses).
- The gateway is a service in the **`tatlock-ui` Portainer stack**
(`system-admin-toj/containers/stacks/tatlock-ui.yml`, registered in
`CONTAINERS.md`): it shares `docker-dataplane` with Speaches (service-name URL
`http://speaches:8000`) and reaches the host-run Tatlock via LAN IP.
Firmware is not containerized: it's flashed over USB (`idf.py flash`), with OTA planned
for phase 5.
**Versioning**: `0.x` while interfaces are still moving — roughly one minor bump per
roadmap phase (`0.1.x` gateway server-side, `0.2.x` first device firmware, `0.3.x`
hands-free). **`v1.0.0` is reserved for the wall milestone**: the device mounted in the
living room, talking to Tatlock end to end.