docs: correct stale tooling and model references
Three migrations left their documentation behind: wakeup.sh was replaced by the Makefile during the project structure consolidation, but AGENTS.md and the e2e README still tell you to run it. The log path moved to build/logs/server.log at the same time. The local model moved to gemma4:e2b, but the e2e prerequisites and the benchmark recommendation still name mistral-nemo. The benchmark figures in CLAUDE.md predate the current model. Measured 2026-08-07: ~95 tok/s, full flow ~10-13s for simple turns, cold model load ~36s rather than ~8s. A turn costs three sequential Ollama calls and ~710 generated tokens regardless of how trivial the question is. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -268,7 +268,7 @@ async def run_benchmarks(iterations: int = 10, verbose: bool = False):
|
||||
print(f" Max: {overall_max:.3f}s (target: ≤5.0s)")
|
||||
print(f" Avg: {overall_avg:.3f}s (target: ≤1.67s)")
|
||||
print(f"\n Recommendations:")
|
||||
print(f" - Switch to a faster model (current: mistral-nemo)")
|
||||
print(f" - Switch to a faster model (current: gemma4:e2b)")
|
||||
print(f" - Reduce system prompt complexity")
|
||||
print(f" - Limit tool calls (currently limited to 3)")
|
||||
print(f" - Consider caching household registry responses")
|
||||
|
||||
Reference in New Issue
Block a user