Every figure in the latency budget was stale, in both directions. TTS was listed at ~1.9 s per sentence but measures ~0.24 s warm for 4.5 s of audio; the full Tatlock flow was listed at 11-25 s but measures ~10-13 s for simple turns. Both sets of numbers predate the current model. The VRAM section now carries real figures and the reason they matter: on 2026-08-07 Tatlock ran against a 9.3 GB model, leaving 7 MiB free, and every transcription failed with CUDA out of memory while the Speaches container still reported healthy. The budget is the constraint, not slack. Also replaces the retired tatlock.schweitz.internal hostname in the topology diagram with the docker container name. Co-Authored-By: Claude <noreply@anthropic.com>
21 KiB
DeskLock Architecture
Goal
An always-on, glanceable butler face in the living room. You speak to it; it relays your words to Tatlock and speaks the reply back, with a face that reflects what it's doing (idle, listening, thinking, speaking). It is deliberately a thin endpoint: all intelligence lives in Tatlock, all heavy audio processing lives server-side on tower-of-joy. Everything is local — no audio, transcript, or reply ever leaves the LAN.
System overview
┌──────────────────────┐ WebSocket: PCM audio + JSON events
│ DeskLock device │◄───────────────────────────────────┐
│ (ESP32-P4) │ │
│ • LVGL face │ ┌──────────────────────────────┴───────────┐
│ • touch / wake word │ │ DeskLock Gateway (container, :8600) │
│ • mic capture + AEC │ │ thin orchestrator — no ML dependencies │
│ • TTS playback │ └───────┬──────────────────┬───────────────┘
└──────────────────────┘ │ │ OpenAI-format HTTP
│ ▼
HTTP (LAN) │ ┌─────────────────────────────┐
▼ │ Speaches (container, GPU) │
┌────────────────────┐ │ • STT: faster-whisper │
│ Tatlock (butler) │ │ • TTS: Kokoro / Piper │
│ container name: │ │ also usable by Open WebUI, │
│ tatlock:8000 │ │ Home Assistant, … │
└────────────────────┘ └─────────────────────────────┘
Tatlock stays a text-only brain. The gateway orchestrates Tatlock's ears and mouth; the speech layer (Speaches) owns the actual STT/TTS models on the GPU.
Components
1. Firmware (firmware/) — ESP32-P4
Responsibilities:
- Face rendering (LVGL 9 on the 800×800 round MIPI-DSI panel via the
waveshare/esp32_p4_wifi6_touch_lcd_xcBSP). Six expression states — see Face design for the visual contract. - Audio capture: dual mics through the ES7210 (hardware echo cancellation reference from the playback path), 16 kHz 16-bit mono PCM.
- Audio playback: ES8311 codec → speaker. Plays PCM streamed from the gateway.
- Transport: a single WebSocket to the gateway carrying binary PCM frames plus JSON
control events (
state,transcript,reply_text, errors). Device reconnects with backoff; face shows a disconnected state when the gateway is unreachable.
On-device speech processing — what runs on the P4 and what deliberately doesn't. The P4 (dual RISC-V @ 400 MHz, 32 MB PSRAM) has a hard ceiling; the split is:
| On-device (planned) | Why |
|---|---|
| Wake word — esp-sr WakeNet (phase 2) | Must be local: always-listening audio should never leave the device until the wake word fires |
| Voice-activity detection (end-of-utterance) | Removes tap-to-stop; cheap on-device |
| Echo cancellation — ES7210 hardware | Enables barge-in while TTS is playing |
| esp-sr MultiNet fixed commands (optional, later) | ~200-phrase closed vocabulary recognized entirely on-device — instant "lights off"-style commands with zero round trip |
Full open-vocabulary STT on-device is out of scope permanently: even whisper-tiny needs hundreds of MB and orders of magnitude more compute than the P4 offers. Anything open-ended goes to the speech layer.
Non-responsibilities: no STT, no TTS, no conversation state. If it's not rendering, recording, or playing, it doesn't belong in firmware.
2. Gateway (gateway/) — container on tower-of-joy, port 8600
A FastAPI service bridging device audio to Tatlock text. It owns orchestration, not models — the container stays a slim pure-Python image with no CUDA/ML dependencies:
- Accepts the device WebSocket (
/ws/voice). - Buffers inbound PCM until end-of-utterance (client-signalled in phase 1; VAD later).
- STT: POST to Speaches
/v1/audio/transcriptions. - Chat: POST the transcript to Tatlock
/v1/chat/completions(http://tatlock:8000, OpenAI-compatible, streaming), maintaining the conversation history so follow-ups have context. - TTS: as Tatlock's token stream completes each sentence, POST it to Speaches
/v1/audio/speechand forward the PCM immediately — see Latency budget.
stt.py / tts.py are pluggable backends selected by config
(DESKLOCK_STT_BACKEND / DESKLOCK_TTS_BACKEND):
speaches(default) — OpenAI-format HTTP to the shared speech container.embedded— in-process faster-whisper / Piper. Kept as a fallback so the gateway can run standalone (dev on a laptop, speech container down), at the cost of a fat image.
The gateway is stateless apart from in-flight conversations; it can restart freely.
3. Speech layer — Speaches (container, GPU)
Speaches (successor to faster-whisper-server)
is a self-hosted, OpenAI-API-compatible speech server: STT via faster-whisper, TTS via
Kokoro/Piper, dynamic model load/offload with a TTL, and a /v1/realtime WebSocket API
we may adopt later for streaming transcription.
- Deployed 2026-07-14:
ghcr.io/speaches-ai/speaches:latest-cudaon host port 8601, withSystran/faster-whisper-small(STT) andspeaches-ai/Kokoro-82M-v1.0-ONNX(TTS — Kokoro is natively 24 kHz, but the gateway requests the 16 kHz device contract directly via Speaches'sample_rateextension, verified live; default voicebm_george, en-GB male). LAN-only like the Tatlock internal route — do not expose through NPM without auth. Register inCONTAINERS.md. - Measured (live round trip, warm, 2026-08-07): STT ~0.30 s for a ~4.8 s utterance; TTS ~0.24 s for a ~4.5 s sentence (real-time factor ~0.05). The first call after an idle gap costs ~1.2 s; a full cold start after model TTL offload adds ~4 s.
- Why a shared layer instead of models inside the gateway: one GPU-resident model
instance serves the whole homelab. Open WebUI is currently configured with
AUDIO_STT_ENGINE=openai/AUDIO_TTS_ENGINE=openai(OpenAI cloud) — pointing its audio base URL at Speaches makes it fully local with a config change. Home Assistant can share it too. Meanwhile the gateway image needs no CUDA and rebuilds in seconds. - VRAM budget: RTX 2080 Ti, 11,264 MiB, shared with Ollama. As of 2026-08-07 the
steady state is ~4.9 GB used / ~5.9 GB free with everything resident:
gemma4:e2b1.9 GB andnomic-embed-text0.3 GB (both pinned), whispersmallint8 <1 GB, Kokoro a few hundred MB. Speaches' model TTL offload keeps idle pressure near zero. This budget is not slack — it is the constraint. On 2026-08-07 Tatlock was deployed againstmistral-nemo:latest(9.3 GB, 2 h keep-alive), which left 7 MiB free and made every transcription fail withCUDA failed with error out of memorywhile the Speaches container still reported healthy. Keep Tatlock's model at or below ~4 GB resident, and checknvidia-smifree VRAM before changing it. If contention ever bites anyway, faster-whispersmallon CPU is an acceptable fallback (int8, a few seconds per utterance).
4. Tatlock — existing backend (/mnt/media/Projects/tatlock)
Untouched by this project. DeskLock consumes its OpenAI-compatible API over the internal LAN (port 8000, bypassing the Authentik-protected public route). If device auth is needed later, the gateway holds the credential — never the firmware.
Face design
Aesthetic: pure black screen; a face drawn from ASCII/terminal glyphs in green
phosphor (#adffc8 face, dimmer greens for secondary info); Matrix-style digital rain
whose density encodes activity — barely-there drips when idle, a downpour while
Tatlock works. No bitmaps, no skeuomorphism: glyphs only.
Source of truth: sim/face/index.html — a self-contained browser simulator of the
800×800 round panel. Design changes land there first, get approved visually, then get
ported to LVGL. The STATES table in the sim defines the contract:
| State | Eyes | Mouth | Rain | Extra cues |
|---|---|---|---|---|
idle |
- - |
\_/ |
2 slow streams | clock (HH:MM), breathing bob, blinks |
listening |
O O |
o |
16 streams | blinks |
pensive |
· · |
~ |
7 streams | cycling ... thought dots |
effort |
> < |
~ |
40 fast streams | orbit arc on bezel + [ Ns ] elapsed counter, face jitter |
speaking |
^ ^ |
cycles o O - O = o |
14 streams | mouth animates ~150 ms/frame |
rage |
— | — | 34 fast streams | 3-frame kaomoji loop through the eyes slot: (°□°) ┬─┬ → (╯°□°)╯︵ ┻━┻ → ┬─┬ ノ( º_º ノ) — flips the table, then composes itself and puts it back |
error |
x x |
- |
none (rain dies) | face dims to 45% |
Sound signature: the cathedral gong (play_boot_gong in firmware — 220 Hz
inharmonic partial stack, feedback-comb reflections, ~-7 dBFS) is the approved house
sound: "audible and butler-non-intrusive." Plays at boot; planned as the
wake-from-dormant sound. Tune by adjective: tail = taus, cathedral size = echo
delays, depth = fundamental, presence = GONG_PEAK/GONG_VOLUME.
Wait cues are a hard requirement (user-stated): Tatlock turns take 10–25 s, so
effort must always show alive-and-working signals — the orbiting bezel arc, the
elapsed-seconds counter, and max rain. Never a bare static face during a wait, and no
fake progress bars — only honest cues.
Protocol → face mapping: gateway state: thinking → effort; transcription and
other short local waits → pensive; listening/speaking map 1:1; an in-flight
request failure (STT/Tatlock/TTS error) → rage for a few loops, then idle;
WebSocket disconnected → error (quiet, persistent); otherwise idle.
LVGL port notes (for phase 2):
- Drive everything from fixed-step
lv_timers (~30 fps rain tick) — the sim deliberately usessetInterval, notrequestAnimationFrame, to mirror this. - Rain:
lv_canvas(or a pooled label grid) with per-frame fade; orbit arc =lv_arc. - Fonts: generate a large monospace glyph font including the katakana subset used in
GLYPHSvialv_font_conv; the built-inunsciifonts are too small for 800 px. Therageframes additionally need╯ ︵ ┻ ━ ┬ ─ ノ ° □ ºin the subset. - The sim's text glow (
text-shadow) is browser flair — the device renders flat glyphs.
Power management (prime concern)
User requirement: the device idles on a wall 95%+ of its life — low power when nothing is happening is a first-class design goal, not a phase-5 nicety. The face state machine is therefore built around a power ladder from day one:
| Power state | Backlight | Rendering | CPU | Entered when |
|---|---|---|---|---|
active |
100% | full animation | full clock | conversation in progress (listening→speaking) |
ambient |
~35% | idle face, sparse rain | DFS enabled | idle, but activity in the last few minutes |
dormant |
off (or ≤5%) | no redraws — render loop parked | min clock via DFS | no voice/touch for N min (default 10) |
night |
off, panel sleep | none | min clock | schedule or "goodnight" command |
Levers, in order of impact:
- Backlight — this is an IPS LCD: black pixels still burn backlight (unlike OLED),
so brightness is the dominant lever.
bsp_display_brightness_set()drives it. - Render idleness — rain off and animations parked means LVGL stops producing frames, which is what lets DFS actually reach its floor.
- DFS / power management (
CONFIG_PM_ENABLE) — automatic frequency scaling when tasks are quiet. Note the MIPI-DSI constraint below. - Radio — the C6 runs Wi-Fi modem power-save; the gateway WebSocket widens its
ping interval when the device reports
dormant.
Wake triggers (any → ambient/active): wake word (phase 5), touch (from phase 2),
local VAD "someone is speaking" pre-warm, a gateway-initiated event (butler wants to
say something), scheduled morning end of night.
Hard edges — what limits how low we can go:
- The hands-free promise sets the power floor. Wake word requires mics + the AFE pipeline running continuously; deep sleep is permanently off the table while the device promises to answer its name. The floor is "CPU lightly loaded at min clock, radios in power-save, backlight off."
- DSI needs clocks while the panel is active — the deepest CPU savings only unlock
in
dormant/nightwhen the panel stops being refreshed (panel sleep / blank). - Wake latency budget: ≤ ~300 ms from trigger to visible face (backlight ramp + first render). Anything slower reads as "it's off," which kills the butler illusion.
- Touch stays powered in all states except possibly
night— its idle draw is negligible and tap-to-wake must always work. - No invented numbers: actual draw gets measured with a USB power meter at each
phase; working target is
dormant≤ ⅓ ofactive. (Always-on device: every watt saved ≈ 9 kWh/year.)
Implementation order: backlight dimming + dormant timeout + touch wake land in
phase 2 with the face state machine (timeout-driven); voice-linked triggers upgrade
it in phase 5.
Latency budget & streaming
Measured 2026-08-07 against the deployed stack (gemma4:e2b at ~95 tok/s, GPU-resident):
| Stage | Cost |
|---|---|
STT (Speaches whisper small) |
~0.30 s warm, for ~4.8 s of audio |
| TTS (Speaches Kokoro) | ~0.24 s warm, for ~4.5 s of audio (RTF ~0.05) |
| Tatlock, full local flow | ~10–13 s end-to-end for simple turns |
| Tatlock cold model load | +~36 s — avoided while the model is pinned |
A Tatlock turn costs 3 sequential Ollama calls (Steward routing → tool orchestration → butler-tone synthesis) and ~710 generated tokens even for "what is 61 plus 12?". Most of that is the model's own reasoning: gemma4 thinks by default, and the effort is spent three times per turn.
(Older figures — "~35 s Steward / ~2 min flow" from the CPU-only era, and "11–25 s full flow" from 2026-07-14 — are superseded. Do not plan against them.)
Speech is not the bottleneck — Tatlock is, by one to two orders of magnitude. Constraints this imposes:
- The gateway must consume Tatlock's streaming response and synthesize
sentence-by-sentence, forwarding audio as each sentence is ready. The device starts
speaking after the first sentence instead of waiting for the full reply — with
streaming, first audio should land roughly at Steward-time + first-sentence-time,
well under the ~10–13 s full-flow figure. The WS protocol already supports this: one
audio_start… PCM …audio_endenvelope with chunks arriving as they're synthesized — the device just plays a continuous stream. - The
thinkingface state is a first-class feature, not decoration — it's what makes a ~10 s Tatlock turn feel intentional instead of broken. Consider progress cues (e.g. surface Tatlock's reasoning summaries on-screen) later. - A fast lane may eventually be needed: MultiNet on-device commands for instant home-automation phrases, and/or a low-latency intent path in Tatlock itself. Out of scope for now, but don't design it out.
WebSocket protocol (device ↔ gateway)
Binary frames: raw 16 kHz s16le mono PCM (mic upstream, TTS downstream). Text frames: JSON control messages.
device → gateway: {"type": "utterance_start"}
device → gateway: <binary PCM frames>
device → gateway: {"type": "utterance_end"}
gateway → device: {"type": "state", "value": "thinking"}
gateway → device: {"type": "transcript", "text": "..."}
gateway → device: {"type": "reply_text", "text": "..."}
gateway → device: {"type": "audio_start", "sample_rate": 16000}
gateway → device: <binary PCM frames> (may arrive sentence-by-sentence; play as a stream)
gateway → device: {"type": "audio_end"}
gateway → device: {"type": "command", "action": "volume_up"} (LLM-bypass; see below)
command (gateway → device) is an alternative to the reply path: when the
gateway recognizes a simple device command in the transcript (volume/mute), it
sends a command instead of calling Tatlock — no reply_text/audio — then returns
to idle. Actions: volume_up, volume_down, mute, unmute, and volume_set
with an extra "level" field (0–11, the on-device volume scale). Matched by the
gateway's commands.py; applied on the device in gw_client.c → face.c.
Planned additions (documented before implemented, here first):
reply_delta(gateway → device): incremental reply text for on-screen streaming while audio is synthesized.- An interrupt event (device → gateway) for barge-in during playback (phase 4).
Keep this protocol documented here and mirrored in firmware/ and gateway/ constants —
it is the one contract between the two halves of the repo.
Key decisions & rationale
- ESP-IDF native (not Arduino/ESPHome): the P4 + MIPI-DSI + esp-sr stack is only first-class in ESP-IDF; Waveshare recommends it, and the BSP targets it.
- Server-side STT/TTS, device does wake word + VAD + AEC only: server whisper is dramatically better than anything embeddable, the GPU is already there, and the P4 physically can't run open-vocabulary STT. Wake word must be on-device (privacy: no audio leaves the device until it fires).
- STT/TTS as a shared Speaches service (not embedded in the gateway): one model
instance for the whole homelab (DeskLock, Open WebUI, potentially HA), slim gateway
image, models upgradable independently.
embeddedbackend retained as a dev/fallback mode.- Rejected — Wyoming protocol containers (
wyoming-faster-whisper/wyoming-piper): native to Home Assistant's ecosystem, but Tatlock and Open WebUI already speak OpenAI format, so Speaches fits the lab better. Revisit only if HA Assist becomes a first-class consumer. - Rejected — cloud STT/TTS: violates the local-first premise; also adds WAN latency and per-minute cost.
- Rejected — Wyoming protocol containers (
- Separate gateway (not extending Tatlock): keeps Tatlock's API text-only and clean; audio concerns (codecs, VAD, streaming, sentence segmentation) stay at the edge. The gateway is also where a future second endpoint (kitchen, office) would connect.
- Monorepo: the WS protocol couples firmware and gateway; versioning them together avoids contract drift.
Additional endpoints — sauron (planned)
sauron is an old iMac (Linux, text-only console) that will run the same UX as a second butler endpoint plus an ops console. Because the gateway protocol is endpoint-agnostic and each WS connection gets its own conversation, extra endpoints are architecturally free.
- Client: a terminal UI (
clients/sauron/, Python + curses/textual) — the face design is already ASCII, so a TTY renders it natively: glyph face states, character-cell matrix rain, green-on-black. Audio via ALSA (arecord/aplay-level, 16 kHz mono PCM), same WebSocket protocol, same state machine. - Screensaver model (user-confirmed): the face+voice layer is the idle mode — full-screen butler when nobody's working. The workspace mode is an SSH ops console: live stats and remote control of tower-of-joy and forge. Any keypress drops from face to console; idle timeout (and wake word later) raises the face again. Voice stays available in both modes.
- Not started — planned after the device reaches phase 4/5. No browser/kiosk stack needed unless we later want the glow.
CI & deployment
Gitea Actions (.gitea/workflows/build.yml), following the tatlock/tatlock-ui pattern:
- Every push to
main: lint + tests for the gateway (Python 3.12). - Version tags (
v0.1.0, …): tests, then buildgateway/intogit.schweitz.net/jpmschweitzer/desklock-gateway:{latest,tag}, push to the Gitea registry, create a release, and trigger Watchtower to roll the running container. - Required repo/org secrets:
REGISTRY_USER,REGISTRY_PASSWORD,WATCHTOWER_HTTP_API_TOKEN(same trio tatlock uses). - The gateway is a service in the
tatlock-uiPortainer stack (system-admin-toj/containers/stacks/tatlock-ui.yml, registered inCONTAINERS.md): it sharesdocker-dataplanewith Speaches (service-name URLhttp://speaches:8000) and reaches the host-run Tatlock via LAN IP.
Firmware is not containerized: it's flashed over USB (idf.py flash), with OTA planned
for phase 5.
Versioning: 0.x while interfaces are still moving — roughly one minor bump per
roadmap phase (0.1.x gateway server-side, 0.2.x first device firmware, 0.3.x
hands-free). v1.0.0 is reserved for the wall milestone: the device mounted in the
living room, talking to Tatlock end to end.