Files
tatlock/tests/e2e
jpmschweitzerandClaude Opus 4.5 583c407edd
Build and Push / build (release) Successful in 53s
fix: Redis bool storage, tool tracking matching, e2e fixture scope
- Convert booleans to strings for Redis hset (Redis doesn't accept bool)
- Extract capability from delegate_to_X tool names for tracking
- Use loop_scope="module" for pytest-asyncio module-scoped fixtures
- Add note about using venv for tests in AGENTS.md

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-16 14:53:51 +01:00
..

End-to-End API Tests

These tests make real HTTP requests to the running Tatlock API server to verify the complete stack works correctly.

Prerequisites

  1. Server must be running on http://localhost:8777 (use ./wakeup.sh)
  2. Ollama must be running with mistral-nemo:latest model
  3. Redis must be running (for benchmarking)
  4. Qdrant must be running on http://localhost:6333 (for memory tests)

Running the Tests

Start the server first:

# Terminal 1: Start the server (auto-reload enabled)
./wakeup.sh

# Logs are written to logs/server.log - tail them in another terminal:
tail -f logs/server.log

Run the E2E tests:

# Run all E2E tests
pytest tests/e2e/ -v -m e2e

# Run orchestration tests specifically
pytest tests/e2e/test_orchestration_e2e.py -v

# Run API endpoint tests
pytest tests/e2e/test_api_endpoints.py -v

Run specific test categories:

# Memory system tests
pytest tests/e2e/test_orchestration_e2e.py::TestMemoryStorage -v
pytest tests/e2e/test_orchestration_e2e.py::TestMemoryRecall -v

# Steward delegation tests
pytest tests/e2e/test_orchestration_e2e.py::TestStewardDelegation -v

# Direct delegation bypass tests (new feature)
pytest tests/e2e/test_orchestration_e2e.py::TestDirectDelegationBypass -v

# User isolation tests
pytest tests/e2e/test_orchestration_e2e.py::TestUserContextIsolation -v

# Orchestration scenario tests
pytest tests/e2e/test_orchestration_e2e.py::TestScenario1WeatherWithMemory -v
pytest tests/e2e/test_orchestration_e2e.py::TestScenario4SimpleExpertDelegation -v
pytest tests/e2e/test_orchestration_e2e.py::TestScenario6WikiCreation -v

# Generate evaluation report
pytest tests/e2e/test_orchestration_e2e.py::TestEvaluationReport -v -s

Test Organization

test_api_endpoints.py - Core API Tests

  • Chat Completions endpoint (/v1/chat/completions)
  • Responses API endpoint (/v1/responses)
  • Streaming responses
  • Error handling
  • OpenAI format compliance

test_orchestration_e2e.py - Orchestration Scenario Tests

Based on ORCHESTRATION_SCENARIOS.md:

Class Scenario What it Tests
TestMemoryStorage Memory storage Store -> Qdrant verification
TestMemoryRecall Memory recall Store -> Recall flow
TestStewardDelegation Steward routing Capability recommendations
TestDirectDelegation Direct bypass Pure memory/librarian requests
TestScenario1WeatherWithMemory Weather check Multi-step with memory lookup
TestScenario4SimpleExpertDelegation Calculator/datetime Simple tool use
TestScenario6WikiCreation Wiki operations Librarian delegation
TestScenario8MultiExpertCoordination Complex requests Multiple capabilities
TestUserContextIsolation User isolation llm_tester vs production
TestDataVerification Data presence Qdrant structure verification
TestIntegrationHealth System health API/Qdrant reachability
TestEvaluationReport Diagnostic Generates behavior reports

User Isolation

Tests use the llm_tester user (development environment default) to isolate test data from production:

  • Test memories: memories_llm_tester (Qdrant collection)
  • Production memories: memories_jpmschweitzer (never modified by tests)

Handling LLM Non-Determinism

LLM outputs are non-deterministic. Tests handle this by:

  1. Flexible assertions - Check for behavior patterns, not exact text
  2. assert_llm_behavior() - Helper for pattern matching with confidence levels
  3. Soft failures (pytest.xfail) - Some tests may fail due to LLM variance without failing the suite
  4. Evaluation reports - Generate diagnostic reports for human review

Example:

result = assert_llm_behavior(
    message_text,
    expected_patterns=[r"(remember|noted|stored)", r"purple"],
    min_matches=1,
)
if not result.passed:
    pytest.xfail(f"LLM response unclear: {result.evidence}")

Data Verification

Tests verify data presence in Qdrant:

# QdrantVerifier helper
qdrant = QdrantVerifier()
points = await qdrant.scroll_points("memories_llm_tester")
memory = await qdrant.find_memory_by_key("memories_llm_tester", "favorite_color")

Troubleshooting

Tests fail with connection error

Make sure the server is running:

./wakeup.sh
curl http://localhost:8777/health  # Should return 200

Tests timeout

  • Check Ollama is running: curl http://localhost:11434/api/tags
  • Increase timeout if needed (default: 120s for LLM calls)

Memory tests fail

  • Check Qdrant is running: curl http://localhost:6333/collections
  • Verify memories_llm_tester collection exists

Inconsistent results

  • LLM responses vary - this is expected
  • Check the evaluation report for detailed diagnostics:
    pytest tests/e2e/test_orchestration_e2e.py::TestEvaluationReport -v -s
    

Tests pollute production data

  • This shouldn't happen - tests use llm_tester user
  • If it does, check ENVIRONMENT is set to development in .env

Adding New Tests

  1. Use existing fixtures (client, qdrant, clean_test_memories)
  2. Use assert_llm_behavior() for flexible LLM output checking
  3. Add @pytest.mark.e2e decorator
  4. Consider adding soft failures for non-deterministic checks
  5. Add test keys to clean_test_memories fixture if storing new memories

Example:

@pytest.mark.e2e
@pytest.mark.asyncio
class TestNewScenario:
    async def test_something(
        self,
        client: httpx.AsyncClient,
        qdrant: QdrantVerifier,
        clean_test_memories,
    ):
        response = await client.post("/v1/responses", json={...})
        # Use assert_llm_behavior for flexible checking
        result = assert_llm_behavior(response_text, expected_patterns=[...])