Files
desklock/docs/architecture.md
T
jpmschweitzerandClaude Fable 5 79d43ec199
Test, Build and Push / test-gateway (push) Successful in 10s
Test, Build and Push / release (push) Successful in 2s
Test, Build and Push / build-gateway (push) Failing after 8s
Deployment home is the tatlock-ui Portainer stack, not a standalone file
The gateway service now lives in system-admin-toj/containers/stacks/
tatlock-ui.yml (shared docker-dataplane network: speaches by service
name, host-run Tatlock by LAN IP). Drop deploy/desklock-gateway.yml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:08:50 +02:00

268 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DeskLock Architecture
## Goal
An always-on, glanceable butler face in the living room. You speak to it; it relays your
words to Tatlock and speaks the reply back, with a face that reflects what it's doing
(idle, listening, thinking, speaking). It is deliberately a *thin* endpoint: all
intelligence lives in Tatlock, all heavy audio processing lives server-side on
tower-of-joy. Everything is local — no audio, transcript, or reply ever leaves the LAN.
## System overview
```
┌──────────────────────┐ WebSocket: PCM audio + JSON events
│ DeskLock device │◄───────────────────────────────────┐
│ (ESP32-P4) │ │
│ • LVGL face │ ┌──────────────────────────────┴───────────┐
│ • touch / wake word │ │ DeskLock Gateway (container, :8600) │
│ • mic capture + AEC │ │ thin orchestrator — no ML dependencies │
│ • TTS playback │ └───────┬──────────────────┬───────────────┘
└──────────────────────┘ │ │ OpenAI-format HTTP
│ ▼
HTTP (LAN) │ ┌─────────────────────────────┐
▼ │ Speaches (container, GPU) │
┌────────────────────┐ │ • STT: faster-whisper │
│ Tatlock (butler) │ │ • TTS: Kokoro / Piper │
│ tatlock.schweitz. │ │ also usable by Open WebUI, │
│ internal :8000 │ │ Home Assistant, … │
└────────────────────┘ └─────────────────────────────┘
```
Tatlock stays a text-only brain. The **gateway** orchestrates Tatlock's ears and mouth;
the **speech layer** (Speaches) owns the actual STT/TTS models on the GPU.
## Components
### 1. Firmware (`firmware/`) — ESP32-P4
Responsibilities:
- **Face rendering** (LVGL 9 on the 800×800 round MIPI-DSI panel via the
`waveshare/esp32_p4_wifi6_touch_lcd_xc` BSP). Six expression states — see
[Face design](#face-design) for the visual contract.
- **Audio capture**: dual mics through the ES7210 (hardware echo cancellation reference
from the playback path), 16 kHz 16-bit mono PCM.
- **Audio playback**: ES8311 codec → speaker. Plays PCM streamed from the gateway.
- **Transport**: a single WebSocket to the gateway carrying binary PCM frames plus JSON
control events (`state`, `transcript`, `reply_text`, errors). Device reconnects with
backoff; face shows a disconnected state when the gateway is unreachable.
**On-device speech processing — what runs on the P4 and what deliberately doesn't.**
The P4 (dual RISC-V @ 400 MHz, 32 MB PSRAM) has a hard ceiling; the split is:
| On-device (planned) | Why |
|---------------------|-----|
| Wake word — esp-sr WakeNet (phase 2) | Must be local: always-listening audio should never leave the device until the wake word fires |
| Voice-activity detection (end-of-utterance) | Removes tap-to-stop; cheap on-device |
| Echo cancellation — ES7210 hardware | Enables barge-in while TTS is playing |
| esp-sr MultiNet fixed commands (optional, later) | ~200-phrase closed vocabulary recognized entirely on-device — instant "lights off"-style commands with zero round trip |
Full open-vocabulary STT on-device is **out of scope permanently**: even whisper-tiny
needs hundreds of MB and orders of magnitude more compute than the P4 offers. Anything
open-ended goes to the speech layer.
Non-responsibilities: no STT, no TTS, no conversation state. If it's not rendering,
recording, or playing, it doesn't belong in firmware.
### 2. Gateway (`gateway/`) — container on tower-of-joy, port 8600
A FastAPI service bridging device audio to Tatlock text. It owns *orchestration*, not
models — the container stays a slim pure-Python image with no CUDA/ML dependencies:
1. Accepts the device WebSocket (`/ws/voice`).
2. Buffers inbound PCM until end-of-utterance (client-signalled in phase 1; VAD later).
3. **STT**: POST to Speaches `/v1/audio/transcriptions`.
4. **Chat**: POST the transcript to Tatlock `/v1/chat/completions`
(`http://tatlock.schweitz.internal:8000`, OpenAI-compatible, **streaming**),
maintaining the conversation history so follow-ups have context.
5. **TTS**: as Tatlock's token stream completes each sentence, POST it to Speaches
`/v1/audio/speech` and forward the PCM immediately — see
[Latency budget](#latency-budget--streaming).
`stt.py` / `tts.py` are pluggable backends selected by config
(`DESKLOCK_STT_BACKEND` / `DESKLOCK_TTS_BACKEND`):
- `speaches` (default) — OpenAI-format HTTP to the shared speech container.
- `embedded` — in-process faster-whisper / Piper. Kept as a fallback so the gateway can
run standalone (dev on a laptop, speech container down), at the cost of a fat image.
The gateway is stateless apart from in-flight conversations; it can restart freely.
### 3. Speech layer — Speaches (container, GPU)
[Speaches](https://github.com/speaches-ai/speaches) (successor to faster-whisper-server)
is a self-hosted, OpenAI-API-compatible speech server: STT via faster-whisper, TTS via
Kokoro/Piper, dynamic model load/offload with a TTL, and a `/v1/realtime` WebSocket API
we may adopt later for streaming transcription.
- **Deployed 2026-07-14**: `ghcr.io/speaches-ai/speaches:latest-cuda` on host port
**8601**, with `Systran/faster-whisper-small` (STT) and
`speaches-ai/Kokoro-82M-v1.0-ONNX` (TTS — Kokoro is natively 24 kHz, but the
gateway requests the 16 kHz device contract directly via Speaches' `sample_rate`
extension, verified live; default voice `bm_george`, en-GB male). LAN-only like the
Tatlock internal route — do not expose through NPM without auth. Register in
`CONTAINERS.md`.
- **Measured** (live round trip through the gateway code, warm): STT ~0.3 s for a
~3 s utterance; TTS ~1.9 s for a ~3 s sentence. Cold start after model TTL offload
adds ~510 s to the first request.
- **Why a shared layer instead of models inside the gateway**: one GPU-resident model
instance serves the whole homelab. Open WebUI is currently configured with
`AUDIO_STT_ENGINE=openai` / `AUDIO_TTS_ENGINE=openai` (OpenAI *cloud*) — pointing its
audio base URL at Speaches makes it fully local with a config change. Home Assistant
can share it too. Meanwhile the gateway image needs no CUDA and rebuilds in seconds.
- **VRAM budget**: RTX 2080 Ti, 11 GB, shared with Ollama (~3.6 GB in use as of
2026-07). whisper `small` at int8 is <1 GB; Kokoro is a few hundred MB. Speaches'
model TTL offload keeps idle pressure near zero. If VRAM contention ever bites,
faster-whisper `small` on CPU is an acceptable fallback (int8, a few seconds per
utterance).
### 4. Tatlock — existing backend (`/mnt/media/Projects/tatlock`)
Untouched by this project. DeskLock consumes its OpenAI-compatible API over the internal
LAN (port 8000, bypassing the Authentik-protected public route). If device auth is needed
later, the gateway holds the credential — never the firmware.
## Face design
**Aesthetic**: pure black screen; a face drawn from ASCII/terminal glyphs in green
phosphor (`#adffc8` face, dimmer greens for secondary info); Matrix-style digital rain
whose **density encodes activity** — barely-there drips when idle, a downpour while
Tatlock works. No bitmaps, no skeuomorphism: glyphs only.
**Source of truth**: `sim/face/index.html` — a self-contained browser simulator of the
800×800 round panel. Design changes land there first, get approved visually, then get
ported to LVGL. The `STATES` table in the sim defines the contract:
| State | Eyes | Mouth | Rain | Extra cues |
|-------|------|-------|------|------------|
| `idle` | `- -` | `\_/` | 2 slow streams | clock (HH:MM), breathing bob, blinks |
| `listening` | `O O` | `o` | 16 streams | blinks |
| `pensive` | `· ·` | `~` | 7 streams | cycling `...` thought dots |
| `effort` | `> <` | `~` | 40 fast streams | **orbit arc on bezel + `[ Ns ]` elapsed counter**, face jitter |
| `speaking` | `^ ^` | cycles `o O - O = o` | 14 streams | mouth animates ~150 ms/frame |
| `rage` | — | — | 34 fast streams | 3-frame kaomoji loop through the eyes slot: `(°□°) ┬─┬``(╯°□°)╯︵ ┻━┻``┬─┬ ( º_º )` — flips the table, then composes itself and puts it back |
| `error` | `x x` | `-` | none (rain dies) | face dims to 45% |
**Wait cues are a hard requirement** (user-stated): Tatlock turns take 1025 s, so
`effort` must always show *alive-and-working* signals — the orbiting bezel arc, the
elapsed-seconds counter, and max rain. Never a bare static face during a wait, and no
fake progress bars — only honest cues.
**Protocol → face mapping**: gateway `state: thinking``effort`; transcription and
other short local waits → `pensive`; `listening`/`speaking` map 1:1; an in-flight
request failure (STT/Tatlock/TTS error) → `rage` for a few loops, then `idle`;
WebSocket disconnected → `error` (quiet, persistent); otherwise `idle`.
**LVGL port notes** (for phase 2):
- Drive everything from fixed-step `lv_timer`s (~30 fps rain tick) — the sim
deliberately uses `setInterval`, not `requestAnimationFrame`, to mirror this.
- Rain: `lv_canvas` (or a pooled label grid) with per-frame fade; orbit arc = `lv_arc`.
- Fonts: generate a large monospace glyph font including the katakana subset used in
`GLYPHS` via `lv_font_conv`; the built-in `unscii` fonts are too small for 800 px.
The `rage` frames additionally need `╯ ︵ ┻ ━ ┬ ─ ノ ° □ º` in the subset.
- The sim's text glow (`text-shadow`) is browser flair — the device renders flat glyphs.
## Latency budget & streaming
Measured/known numbers that shape the design (Tatlock figures per tatlock CLAUDE.md,
GPU-resident benchmarks of 2026-07-14, gemma4:e2b at ~100 tok/s):
| Stage | Cost |
|-------|------|
| STT (Speaches whisper `small`) | ~0.3 s warm (measured) |
| TTS (Speaches Kokoro) | ~1.9 s per ~3 s sentence, warm (measured) |
| Tatlock Steward analysis | ~6 s warm |
| **Tatlock, full local flow** | **1125 s end-to-end** (librarian-routed ~2025 s) |
| Tatlock cold start (>2 h idle) | +~8 s (`OLLAMA_KEEP_ALIVE=2h`) |
(Older "~35 s Steward / ~2 min flow" figures were from a CPU-only driver-mismatch era —
do not plan against them.)
Speech is not the bottleneck — **Tatlock is**, by one to two orders of magnitude.
Constraints this imposes:
1. **The gateway must consume Tatlock's streaming response and synthesize
sentence-by-sentence**, forwarding audio as each sentence is ready. The device starts
speaking after the first sentence instead of waiting for the full reply — with
streaming, first audio should land roughly at Steward-time + first-sentence-time,
well under the 1125 s full-flow figure. The WS protocol already supports this: one
`audio_start` … PCM … `audio_end` envelope with chunks arriving as they're
synthesized — the device just plays a continuous stream.
2. **The `thinking` face state is a first-class feature**, not decoration — it's what
makes a 1025 s Tatlock turn feel intentional instead of broken. Consider progress
cues (e.g. surface Tatlock's reasoning summaries on-screen) later.
3. A **fast lane** may eventually be needed: MultiNet on-device commands for instant
home-automation phrases, and/or a low-latency intent path in Tatlock itself. Out of
scope for now, but don't design it out.
## WebSocket protocol (device ↔ gateway)
Binary frames: raw 16 kHz s16le mono PCM (mic upstream, TTS downstream).
Text frames: JSON control messages.
```
device → gateway: {"type": "utterance_start"}
device → gateway: <binary PCM frames>
device → gateway: {"type": "utterance_end"}
gateway → device: {"type": "state", "value": "thinking"}
gateway → device: {"type": "transcript", "text": "..."}
gateway → device: {"type": "reply_text", "text": "..."}
gateway → device: {"type": "audio_start", "sample_rate": 16000}
gateway → device: <binary PCM frames> (may arrive sentence-by-sentence; play as a stream)
gateway → device: {"type": "audio_end"}
```
Planned additions (documented before implemented, here first):
- `reply_delta` (gateway → device): incremental reply text for on-screen streaming while
audio is synthesized.
- An interrupt event (device → gateway) for barge-in during playback (phase 4).
Keep this protocol documented here and mirrored in `firmware/` and `gateway/` constants —
it is the one contract between the two halves of the repo.
## Key decisions & rationale
- **ESP-IDF native (not Arduino/ESPHome)**: the P4 + MIPI-DSI + esp-sr stack is only
first-class in ESP-IDF; Waveshare recommends it, and the BSP targets it.
- **Server-side STT/TTS, device does wake word + VAD + AEC only**: server whisper is
dramatically better than anything embeddable, the GPU is already there, and the P4
physically can't run open-vocabulary STT. Wake word must be on-device (privacy: no
audio leaves the device until it fires).
- **STT/TTS as a shared Speaches service (not embedded in the gateway)**: one model
instance for the whole homelab (DeskLock, Open WebUI, potentially HA), slim gateway
image, models upgradable independently. `embedded` backend retained as a dev/fallback
mode.
- *Rejected — Wyoming protocol containers* (`wyoming-faster-whisper`/`wyoming-piper`):
native to Home Assistant's ecosystem, but Tatlock and Open WebUI already speak
OpenAI format, so Speaches fits the lab better. Revisit only if HA Assist becomes a
first-class consumer.
- *Rejected — cloud STT/TTS*: violates the local-first premise; also adds WAN latency
and per-minute cost.
- **Separate gateway (not extending Tatlock)**: keeps Tatlock's API text-only and clean;
audio concerns (codecs, VAD, streaming, sentence segmentation) stay at the edge. The
gateway is also where a future second endpoint (kitchen, office) would connect.
- **Monorepo**: the WS protocol couples firmware and gateway; versioning them together
avoids contract drift.
## CI & deployment
Gitea Actions (`.gitea/workflows/build.yml`), following the tatlock/tatlock-ui pattern:
- **Every push to `main`**: lint + tests for the gateway (Python 3.12).
- **Version tags (`v0.1.0`, …)**: tests, then build `gateway/` into
`git.schweitz.internal/jpmschweitzer/desklock-gateway:{latest,tag}`, push to the
Gitea registry, create a release, and trigger Watchtower to roll the running
container.
- Required repo/org secrets: `REGISTRY_USER`, `REGISTRY_PASSWORD`,
`WATCHTOWER_TOKEN` (same trio tatlock uses).
- The gateway is a service in the **`tatlock-ui` Portainer stack**
(`system-admin-toj/containers/stacks/tatlock-ui.yml`, registered in
`CONTAINERS.md`): it shares `docker-dataplane` with Speaches (service-name URL
`http://speaches:8000`) and reaches the host-run Tatlock via LAN IP.
Firmware is not containerized: it's flashed over USB (`idf.py flash`), with OTA planned
for phase 5.