Files
desklock/docs/architecture.md
T
jpmschweitzerandClaude Fable 5 22f00fda0f
Test, Build and Push / test-gateway (push) Successful in 33s
Test, Build and Push / release (push) Skipped
Test, Build and Push / build-gateway (push) Skipped
Drop audioop: request 16 kHz directly via Speaches sample_rate extension
Settles the Python version question: floor >=3.11, no ceiling.
Container moves to python:3.13-slim, CI tests on 3.13. The embedded
Piper fallback resamples with numpy (already present via [speech]).
Verified with a live round trip at 16 kHz.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:13:58 +02:00

15 KiB
Raw Blame History

DeskLock Architecture

Goal

An always-on, glanceable butler face in the living room. You speak to it; it relays your words to Tatlock and speaks the reply back, with a face that reflects what it's doing (idle, listening, thinking, speaking). It is deliberately a thin endpoint: all intelligence lives in Tatlock, all heavy audio processing lives server-side on tower-of-joy. Everything is local — no audio, transcript, or reply ever leaves the LAN.

System overview

┌──────────────────────┐  WebSocket: PCM audio + JSON events
│  DeskLock device     │◄───────────────────────────────────┐
│  (ESP32-P4)          │                                    │
│  • LVGL face         │     ┌──────────────────────────────┴───────────┐
│  • touch / wake word │     │  DeskLock Gateway  (container, :8600)    │
│  • mic capture + AEC │     │  thin orchestrator — no ML dependencies  │
│  • TTS playback      │     └───────┬──────────────────┬───────────────┘
└──────────────────────┘             │                  │  OpenAI-format HTTP
                                     │                  ▼
                          HTTP (LAN) │       ┌─────────────────────────────┐
                                     ▼       │  Speaches  (container, GPU) │
                     ┌────────────────────┐  │  • STT: faster-whisper      │
                     │  Tatlock (butler)  │  │  • TTS: Kokoro / Piper      │
                     │  tatlock.schweitz. │  │  also usable by Open WebUI, │
                     │  internal :8000    │  │  Home Assistant, …          │
                     └────────────────────┘  └─────────────────────────────┘

Tatlock stays a text-only brain. The gateway orchestrates Tatlock's ears and mouth; the speech layer (Speaches) owns the actual STT/TTS models on the GPU.

Components

1. Firmware (firmware/) — ESP32-P4

Responsibilities:

  • Face rendering (LVGL 9 on the 800×800 round MIPI-DSI panel via the waveshare/esp32_p4_wifi6_touch_lcd_xc BSP). Six expression states — see Face design for the visual contract.
  • Audio capture: dual mics through the ES7210 (hardware echo cancellation reference from the playback path), 16 kHz 16-bit mono PCM.
  • Audio playback: ES8311 codec → speaker. Plays PCM streamed from the gateway.
  • Transport: a single WebSocket to the gateway carrying binary PCM frames plus JSON control events (state, transcript, reply_text, errors). Device reconnects with backoff; face shows a disconnected state when the gateway is unreachable.

On-device speech processing — what runs on the P4 and what deliberately doesn't. The P4 (dual RISC-V @ 400 MHz, 32 MB PSRAM) has a hard ceiling; the split is:

On-device (planned) Why
Wake word — esp-sr WakeNet (phase 2) Must be local: always-listening audio should never leave the device until the wake word fires
Voice-activity detection (end-of-utterance) Removes tap-to-stop; cheap on-device
Echo cancellation — ES7210 hardware Enables barge-in while TTS is playing
esp-sr MultiNet fixed commands (optional, later) ~200-phrase closed vocabulary recognized entirely on-device — instant "lights off"-style commands with zero round trip

Full open-vocabulary STT on-device is out of scope permanently: even whisper-tiny needs hundreds of MB and orders of magnitude more compute than the P4 offers. Anything open-ended goes to the speech layer.

Non-responsibilities: no STT, no TTS, no conversation state. If it's not rendering, recording, or playing, it doesn't belong in firmware.

2. Gateway (gateway/) — container on tower-of-joy, port 8600

A FastAPI service bridging device audio to Tatlock text. It owns orchestration, not models — the container stays a slim pure-Python image with no CUDA/ML dependencies:

  1. Accepts the device WebSocket (/ws/voice).
  2. Buffers inbound PCM until end-of-utterance (client-signalled in phase 1; VAD later).
  3. STT: POST to Speaches /v1/audio/transcriptions.
  4. Chat: POST the transcript to Tatlock /v1/chat/completions (http://tatlock.schweitz.internal:8000, OpenAI-compatible, streaming), maintaining the conversation history so follow-ups have context.
  5. TTS: as Tatlock's token stream completes each sentence, POST it to Speaches /v1/audio/speech and forward the PCM immediately — see Latency budget.

stt.py / tts.py are pluggable backends selected by config (DESKLOCK_STT_BACKEND / DESKLOCK_TTS_BACKEND):

  • speaches (default) — OpenAI-format HTTP to the shared speech container.
  • embedded — in-process faster-whisper / Piper. Kept as a fallback so the gateway can run standalone (dev on a laptop, speech container down), at the cost of a fat image.

The gateway is stateless apart from in-flight conversations; it can restart freely.

3. Speech layer — Speaches (container, GPU)

Speaches (successor to faster-whisper-server) is a self-hosted, OpenAI-API-compatible speech server: STT via faster-whisper, TTS via Kokoro/Piper, dynamic model load/offload with a TTL, and a /v1/realtime WebSocket API we may adopt later for streaming transcription.

  • Deployed 2026-07-14: ghcr.io/speaches-ai/speaches:latest-cuda on host port 8601, with Systran/faster-whisper-small (STT) and speaches-ai/Kokoro-82M-v1.0-ONNX (TTS — Kokoro is natively 24 kHz, but the gateway requests the 16 kHz device contract directly via Speaches' sample_rate extension, verified live; default voice bm_george, en-GB male). LAN-only like the Tatlock internal route — do not expose through NPM without auth. Register in CONTAINERS.md.
  • Measured (live round trip through the gateway code, warm): STT ~0.3 s for a ~3 s utterance; TTS ~1.9 s for a ~3 s sentence. Cold start after model TTL offload adds ~510 s to the first request.
  • Why a shared layer instead of models inside the gateway: one GPU-resident model instance serves the whole homelab. Open WebUI is currently configured with AUDIO_STT_ENGINE=openai / AUDIO_TTS_ENGINE=openai (OpenAI cloud) — pointing its audio base URL at Speaches makes it fully local with a config change. Home Assistant can share it too. Meanwhile the gateway image needs no CUDA and rebuilds in seconds.
  • VRAM budget: RTX 2080 Ti, 11 GB, shared with Ollama (~3.6 GB in use as of 2026-07). whisper small at int8 is <1 GB; Kokoro is a few hundred MB. Speaches' model TTL offload keeps idle pressure near zero. If VRAM contention ever bites, faster-whisper small on CPU is an acceptable fallback (int8, a few seconds per utterance).

4. Tatlock — existing backend (/mnt/media/Projects/tatlock)

Untouched by this project. DeskLock consumes its OpenAI-compatible API over the internal LAN (port 8000, bypassing the Authentik-protected public route). If device auth is needed later, the gateway holds the credential — never the firmware.

Face design

Aesthetic: pure black screen; a face drawn from ASCII/terminal glyphs in green phosphor (#adffc8 face, dimmer greens for secondary info); Matrix-style digital rain whose density encodes activity — barely-there drips when idle, a downpour while Tatlock works. No bitmaps, no skeuomorphism: glyphs only.

Source of truth: sim/face/index.html — a self-contained browser simulator of the 800×800 round panel. Design changes land there first, get approved visually, then get ported to LVGL. The STATES table in the sim defines the contract:

State Eyes Mouth Rain Extra cues
idle - - \_/ 2 slow streams clock (HH:MM), breathing bob, blinks
listening O O o 16 streams blinks
pensive · · ~ 7 streams cycling ... thought dots
effort > < ~ 40 fast streams orbit arc on bezel + [ Ns ] elapsed counter, face jitter
speaking ^ ^ cycles o O - O = o 14 streams mouth animates ~150 ms/frame
rage 34 fast streams 3-frame kaomoji loop through the eyes slot: (°□°) ┬─┬(╯°□°)╯︵ ┻━┻┬─┬ ( º_º ) — flips the table, then composes itself and puts it back
error x x - none (rain dies) face dims to 45%

Wait cues are a hard requirement (user-stated): Tatlock turns take 1025 s, so effort must always show alive-and-working signals — the orbiting bezel arc, the elapsed-seconds counter, and max rain. Never a bare static face during a wait, and no fake progress bars — only honest cues.

Protocol → face mapping: gateway state: thinkingeffort; transcription and other short local waits → pensive; listening/speaking map 1:1; an in-flight request failure (STT/Tatlock/TTS error) → rage for a few loops, then idle; WebSocket disconnected → error (quiet, persistent); otherwise idle.

LVGL port notes (for phase 2):

  • Drive everything from fixed-step lv_timers (~30 fps rain tick) — the sim deliberately uses setInterval, not requestAnimationFrame, to mirror this.
  • Rain: lv_canvas (or a pooled label grid) with per-frame fade; orbit arc = lv_arc.
  • Fonts: generate a large monospace glyph font including the katakana subset used in GLYPHS via lv_font_conv; the built-in unscii fonts are too small for 800 px. The rage frames additionally need ╯ ︵ ┻ ━ ┬ ─ ° □ º in the subset.
  • The sim's text glow (text-shadow) is browser flair — the device renders flat glyphs.

Latency budget & streaming

Measured/known numbers that shape the design (Tatlock figures per tatlock CLAUDE.md, GPU-resident benchmarks of 2026-07-14, gemma4:e2b at ~100 tok/s):

Stage Cost
STT (Speaches whisper small) ~0.3 s warm (measured)
TTS (Speaches Kokoro) ~1.9 s per ~3 s sentence, warm (measured)
Tatlock Steward analysis ~6 s warm
Tatlock, full local flow 1125 s end-to-end (librarian-routed ~2025 s)
Tatlock cold start (>2 h idle) +~8 s (OLLAMA_KEEP_ALIVE=2h)

(Older "~35 s Steward / ~2 min flow" figures were from a CPU-only driver-mismatch era — do not plan against them.)

Speech is not the bottleneck — Tatlock is, by one to two orders of magnitude. Constraints this imposes:

  1. The gateway must consume Tatlock's streaming response and synthesize sentence-by-sentence, forwarding audio as each sentence is ready. The device starts speaking after the first sentence instead of waiting for the full reply — with streaming, first audio should land roughly at Steward-time + first-sentence-time, well under the 1125 s full-flow figure. The WS protocol already supports this: one audio_start … PCM … audio_end envelope with chunks arriving as they're synthesized — the device just plays a continuous stream.
  2. The thinking face state is a first-class feature, not decoration — it's what makes a 1025 s Tatlock turn feel intentional instead of broken. Consider progress cues (e.g. surface Tatlock's reasoning summaries on-screen) later.
  3. A fast lane may eventually be needed: MultiNet on-device commands for instant home-automation phrases, and/or a low-latency intent path in Tatlock itself. Out of scope for now, but don't design it out.

WebSocket protocol (device ↔ gateway)

Binary frames: raw 16 kHz s16le mono PCM (mic upstream, TTS downstream). Text frames: JSON control messages.

device → gateway:  {"type": "utterance_start"}
device → gateway:  <binary PCM frames>
device → gateway:  {"type": "utterance_end"}
gateway → device:  {"type": "state", "value": "thinking"}
gateway → device:  {"type": "transcript", "text": "..."}
gateway → device:  {"type": "reply_text", "text": "..."}
gateway → device:  {"type": "audio_start", "sample_rate": 16000}
gateway → device:  <binary PCM frames>   (may arrive sentence-by-sentence; play as a stream)
gateway → device:  {"type": "audio_end"}

Planned additions (documented before implemented, here first):

  • reply_delta (gateway → device): incremental reply text for on-screen streaming while audio is synthesized.
  • An interrupt event (device → gateway) for barge-in during playback (phase 4).

Keep this protocol documented here and mirrored in firmware/ and gateway/ constants — it is the one contract between the two halves of the repo.

Key decisions & rationale

  • ESP-IDF native (not Arduino/ESPHome): the P4 + MIPI-DSI + esp-sr stack is only first-class in ESP-IDF; Waveshare recommends it, and the BSP targets it.
  • Server-side STT/TTS, device does wake word + VAD + AEC only: server whisper is dramatically better than anything embeddable, the GPU is already there, and the P4 physically can't run open-vocabulary STT. Wake word must be on-device (privacy: no audio leaves the device until it fires).
  • STT/TTS as a shared Speaches service (not embedded in the gateway): one model instance for the whole homelab (DeskLock, Open WebUI, potentially HA), slim gateway image, models upgradable independently. embedded backend retained as a dev/fallback mode.
    • Rejected — Wyoming protocol containers (wyoming-faster-whisper/wyoming-piper): native to Home Assistant's ecosystem, but Tatlock and Open WebUI already speak OpenAI format, so Speaches fits the lab better. Revisit only if HA Assist becomes a first-class consumer.
    • Rejected — cloud STT/TTS: violates the local-first premise; also adds WAN latency and per-minute cost.
  • Separate gateway (not extending Tatlock): keeps Tatlock's API text-only and clean; audio concerns (codecs, VAD, streaming, sentence segmentation) stay at the edge. The gateway is also where a future second endpoint (kitchen, office) would connect.
  • Monorepo: the WS protocol couples firmware and gateway; versioning them together avoids contract drift.

CI & deployment

Gitea Actions (.gitea/workflows/build.yml), following the tatlock/tatlock-ui pattern:

  • Every push to main: lint + tests for the gateway (Python 3.12).
  • Version tags (v0.1.0, …): tests, then build gateway/ into git.schweitz.internal/jpmschweitzer/desklock-gateway:{latest,tag}, push to the Gitea registry, create a release, and trigger Watchtower to roll the running container.
  • Required repo/org secrets: REGISTRY_USER, REGISTRY_PASSWORD, WATCHTOWER_TOKEN (same trio tatlock uses).
  • The runtime stack definition lives in deploy/desklock-gateway.yml; copy it into system-admin-toj/containers/stacks/ to deploy, and register port 8600 in CONTAINERS.md.

Firmware is not containerized: it's flashed over USB (idf.py flash), with OTA planned for phase 5.