feat(bench): serving benchmark — the baseline instrument for the migration

A scripted 8-turn conversation through the real pipeline (tool call,
memory recall, librarian route, small talk), run against any base URL so
the same instrument measures Ollama today and llama-server after the
cutover. Non-stream mode records wall time and usage; --stream records
time-to-first-token; a VRAM poller and an optional interleaved second
tenant capture the contention the serving plan calls out.

Planted facts carry a BENCH-MARKER prefix so anything the biographer
learns from a bench run stays recognizable in the memory stores.

First smoke against production: 9.7s mean per simple turn, and TTFT
equals wall time — the client sees nothing until the whole
steward-orchestrate-synthesize pipeline has finished. Time-to-first-word
for desklock is currently the full pipeline latency, which makes
streamed synthesis a first-class goal of the serving work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-11 21:07:57 +02:00
co-authored by Claude Fable 5
parent 014afc960c
commit f578ebded8
3 changed files with 366 additions and 1 deletions
+7
View File
@@ -7,6 +7,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- `make bench-serving` — serving-layer benchmark replaying a scripted
conversation through the pipeline; measures per-turn wall time, TTFT
(`--stream`), token usage and VRAM peak, with an optional interleaved
second tenant. Baseline instrument for the serving migration.
### Changed
- `make setup` now ends with a `pytest --collect-only` pass so a broken environment