Files
tatlock/CLAUDE.md
T

3.0 KiB

CLAUDE.md

Claude Code-specific notes for this project. For general development instructions, architecture, coding standards, and deployment — see AGENTS.md.

Setup & Commands

make setup          # Create venv and install all dependencies
make test           # Unit tests (no external services)
make test-integration  # Integration tests (needs Claude/Ollama)
make test-contracts # Wire-level contract tests against live service boundaries
make run            # Start dev server on port 8777
make lint           # Ruff linter + formatter check
make typecheck      # Mypy
make clean          # Remove caches and build artifacts

Dependencies are in pyproject.toml ([project.dependencies] and [project.optional-dependencies.dev]).

Critical Gotchas

ASGITransport does NOT trigger FastAPI lifespan events. The session-scoped _initialize_app fixture in tests/conftest.py calls initialize_application() explicitly via asyncio.run(). Without this, the Ollama/Claude health checks never run: _ollama_available stays None (treated as available, so requests go to Ollama) and _claude_available stays None (treated as unavailable, so the Claude fallback never engages).

AsyncIO scope mismatch. asyncio_default_fixture_loop_scope = function is set in pyproject.toml. Session-scoped async fixtures cause ScopeMismatch errors. The fix is to use a sync fixture with asyncio.run() for session-scoped initialization.

The butler persona prompt suppresses local-model tool calling. With TATLOCK_SYSTEM_PROMPT attached, gemma4 reasons about calling the calculator, then answers from memory with wrong arithmetic (a different wrong product each run). orchestrate_tool_calls() therefore uses the terse TATLOCK_ORCHESTRATION_PROMPT; the persona is applied in synthesize_from_results(). Do not reattach the persona prompt to a tool-phase agent. tool_choice: "required" via extra_body does NOT force Ollama to call tools — it is advisory at best.

Claude Sonnet 5+ rejects sampling parameters. temperature/top_p/top_k return a 400. Use get_sampling_settings() from the model selector instead of passing ModelSettings(temperature=...) directly to agents that can run on the Claude fallback. The contract test suite pins this (make test-contracts).

Integration test timeouts. Set to 120s to match OLLAMA_TIMEOUT config (300s for the pure-Ollama fallback test, which cannot be rescued by Claude). The full local Steward → orchestrate → synthesize flow takes ~2 minutes on gemma4. Steward analysis alone needs ~35s warm — STEWARD_TIMEOUT defaults to 60s.

get_benchmark_store does not exist. The benchmarking module (src/core/benchmarks.py) was never implemented. scripts/benchmark_analysis.py also references it and is broken. Do not add mocks for it in tests.

Steward tests need household registry. Use register_household_members() (sync) in fixtures, not initialize_application() (async). The steward extracts capabilities from the registry.