docs: correct stale tooling and model references

Three migrations left their documentation behind:

wakeup.sh was replaced by the Makefile during the project structure
consolidation, but AGENTS.md and the e2e README still tell you to run it.
The log path moved to build/logs/server.log at the same time.

The local model moved to gemma4:e2b, but the e2e prerequisites and the
benchmark recommendation still name mistral-nemo.

The benchmark figures in CLAUDE.md predate the current model. Measured
2026-08-07: ~95 tok/s, full flow ~10-13s for simple turns, cold model load
~36s rather than ~8s. A turn costs three sequential Ollama calls and ~710
generated tokens regardless of how trivial the question is.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-08-07 15:07:10 +02:00
co-authored by Claude
parent 2cf3252a19
commit 99569e786e
4 changed files with 9 additions and 9 deletions
+5 -5
View File
@@ -4,8 +4,8 @@ These tests make real HTTP requests to the running Tatlock API server to verify
## Prerequisites
1. **Server must be running** on `http://localhost:8777` (use `./wakeup.sh`)
2. **Ollama must be running** with `mistral-nemo:latest` model
1. **Server must be running** on `http://localhost:8777` (use `make run`)
2. **Ollama must be running** with the `gemma4:e2b` model
3. **Redis must be running** (for benchmarking)
4. **Qdrant must be running** on `http://localhost:6333` (for memory tests)
@@ -15,9 +15,9 @@ These tests make real HTTP requests to the running Tatlock API server to verify
```bash
# Terminal 1: Start the server (auto-reload enabled)
./wakeup.sh
make run
# Logs are written to logs/server.log - tail them in another terminal:
# Logs are written to build/logs/server.log - tail them in another terminal:
tail -f logs/server.log
```
@@ -132,7 +132,7 @@ memory = await qdrant.find_memory_by_key("memories_llm_tester", "favorite_color"
Make sure the server is running:
```bash
./wakeup.sh
make run
curl http://localhost:8777/health # Should return 200
```