3.0 KiB
CLAUDE.md
Claude Code-specific notes for this project. For general development instructions, architecture, coding standards, and deployment — see AGENTS.md.
Setup & Commands
make setup # Create venv and install all dependencies
make test # Unit tests (no external services)
make test-integration # Integration tests (needs Claude/Ollama)
make test-contracts # Wire-level contract tests against live service boundaries
make run # Start dev server on port 8777
make lint # Ruff linter + formatter check
make typecheck # Mypy
make clean # Remove caches and build artifacts
Dependencies are in pyproject.toml ([project.dependencies] and [project.optional-dependencies.dev]).
Critical Gotchas
ASGITransport does NOT trigger FastAPI lifespan events. The session-scoped _initialize_app fixture in tests/conftest.py calls initialize_application() explicitly via asyncio.run(). Without this, the Ollama/Claude health checks never run: _ollama_available stays None (treated as available, so requests go to Ollama) and _claude_available stays None (treated as unavailable, so the Claude fallback never engages).
AsyncIO scope mismatch. asyncio_default_fixture_loop_scope = function is set in pyproject.toml. Session-scoped async fixtures cause ScopeMismatch errors. The fix is to use a sync fixture with asyncio.run() for session-scoped initialization.
The butler persona prompt suppresses local-model tool calling. With TATLOCK_SYSTEM_PROMPT attached, gemma4 reasons about calling the calculator, then answers from memory with wrong arithmetic (a different wrong product each run). orchestrate_tool_calls() therefore uses the terse TATLOCK_ORCHESTRATION_PROMPT; the persona is applied in synthesize_from_results(). Do not reattach the persona prompt to a tool-phase agent. tool_choice: "required" via extra_body does NOT force Ollama to call tools — it is advisory at best.
Claude Sonnet 5+ rejects sampling parameters. temperature/top_p/top_k return a 400. Use get_sampling_settings() from the model selector instead of passing ModelSettings(temperature=...) directly to agents that can run on the Claude fallback. The contract test suite pins this (make test-contracts).
Integration test timeouts. Set to 120s to match OLLAMA_TIMEOUT config (300s for the pure-Ollama fallback test, which cannot be rescued by Claude). The full local Steward → orchestrate → synthesize flow takes ~2 minutes on gemma4. Steward analysis alone needs ~35s warm — STEWARD_TIMEOUT defaults to 60s.
get_benchmark_store does not exist. The benchmarking module (src/core/benchmarks.py) was never implemented. scripts/benchmark_analysis.py also references it and is broken. Do not add mocks for it in tests.
Steward tests need household registry. Use register_household_members() (sync) in fixtures, not initialize_application() (async). The steward extracts capabilities from the registry.