feat(bench): serving benchmark — the baseline instrument for the migration
A scripted 8-turn conversation through the real pipeline (tool call, memory recall, librarian route, small talk), run against any base URL so the same instrument measures Ollama today and llama-server after the cutover. Non-stream mode records wall time and usage; --stream records time-to-first-token; a VRAM poller and an optional interleaved second tenant capture the contention the serving plan calls out. Planted facts carry a BENCH-MARKER prefix so anything the biographer learns from a bench run stays recognizable in the memory stores. First smoke against production: 9.7s mean per simple turn, and TTFT equals wall time — the client sees nothing until the whole steward-orchestrate-synthesize pipeline has finished. Time-to-first-word for desklock is currently the full pipeline latency, which makes streamed synthesis a first-class goal of the serving work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -7,6 +7,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
|
||||
- `make bench-serving` — serving-layer benchmark replaying a scripted
|
||||
conversation through the pipeline; measures per-turn wall time, TTFT
|
||||
(`--stream`), token usage and VRAM peak, with an optional interleaved
|
||||
second tenant. Baseline instrument for the serving migration.
|
||||
|
||||
### Changed
|
||||
|
||||
- `make setup` now ends with a `pytest --collect-only` pass so a broken environment
|
||||
|
||||
Reference in New Issue
Block a user