feat(ai): add comprehensive quality test suite for core-ai agent

Add automated test suite for regression detection and performance tracking
of the core-ai agent behavior across code changes.

Changes:
- Add test_ai_flow_quality.py with 5 core test scenarios
  - Simple knowledge queries (no tools)
  - Web search integration
  - Mathematical calculations
  - Date/time operations
  - Multi-tool reasoning tasks
- Add QUALITY_TESTS.md documentation
  - Usage guide and test descriptions
  - Baseline establishment workflow
  - Model benchmarking procedures
  - Troubleshooting guide
- Add performance baseline tests
- Add regression detection tests
- Generate text and JSON reports with git tagging
- Update .gitignore to exclude generated test reports
- Update CHANGELOG.md with test suite details

Baseline Results:
- 4/5 tests passing (80% success rate)
- Average response time: 2-8s per query
- Agent: OllamaNativeAgent with PydanticAI
- Model: mistral-nemo:latest

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2025-12-02 18:29:47 +01:00
co-authored by Claude
parent cd81442a8b
commit 5368496f6f
4 changed files with 985 additions and 0 deletions
+1
View File
@@ -91,6 +91,7 @@ cache/
coverage/
.coverage
test-results/
services/core-ai/tests/reports/
# Documentation builds
docs/_build/
+15
View File
@@ -7,6 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **AI Quality Test Suite**
- Comprehensive test suite for core-ai agent performance (`test_ai_flow_quality.py`)
- Test documentation and usage guide (`QUALITY_TESTS.md`)
- 5 core test scenarios: simple knowledge, web search, calculation, date operations, multi-tool reasoning
- Performance baseline tests and regression detection
- Automated report generation (text and JSON formats)
- Git-tagged reports for version comparison and rollback
- Baseline established: 4/5 tests passing (80% success rate)
### Changed
- **Gitignore Updates**
- Added `services/core-ai/tests/reports/` to prevent committing test output files
- Maintains proper separation of code vs generated artifacts
### In Progress
- **System Monitoring:** Post-migration stability monitoring and performance optimization
+364
View File
@@ -0,0 +1,364 @@
# Core-AI Quality Test Suite
Comprehensive test suite for benchmarking AI agent performance and detecting regressions across code changes.
## Purpose
This test suite validates:
- **Tool calling decision-making** - Does the agent choose the right tools?
- **Response quality** - Are responses accurate and complete?
- **Performance** - Are responses delivered within acceptable timeframes?
- **Regression detection** - Has quality degraded since the last version?
## Current Implementation
- **Agent**: OllamaNativeAgent (Ollama native API with tool calling)
- **Model**: mistral-nemo:latest
- **Framework**: PydanticAI
- **Tools**: web_search, calculate, date/time operations
## Quick Start
### Run All Tests
```bash
# From services/core-ai directory
python tests/test_ai_flow_quality.py
```
This will:
1. Run all test scenarios
2. Generate a comprehensive report
3. Save reports to `tests/reports/` with timestamp and git tag
4. Output results to console
### Run with pytest
```bash
# Run all tests
pytest tests/test_ai_flow_quality.py -v
# Run specific scenario
pytest tests/test_ai_flow_quality.py::test_scenario2_web_search -v
# Run regression tests only
pytest tests/test_ai_flow_quality.py -k regression -v
```
## Test Scenarios
### Scenario 1: Simple Knowledge Query
- **Query**: "What is Docker?"
- **Expected**: Direct answer without tools
- **Performance Target**: < 10s
### Scenario 2: Web Search
- **Query**: "What are the latest Kubernetes security best practices?"
- **Expected**: Uses web_search tool, synthesizes results
- **Performance Target**: < 30s
### Scenario 3: Mathematical Calculation
- **Query**: "Calculate 2847 * 1923 + 5612 - 999"
- **Expected**: Uses calculate tool for precision
- **Performance Target**: < 15s
- **Expected Result**: 5,479,394
### Scenario 4: Date/Time Operations
- **Query**: "What is the current date and what will it be in 30 days?"
- **Expected**: Uses get_current_date and add_days_to_date tools
- **Performance Target**: < 15s
### Scenario 5: Multi-Tool Complex Query
- **Query**: "Get current time in NYC and Tokyo, calculate difference"
- **Expected**: Multiple tool calls (get_current_time × 2, synthesis)
- **Performance Target**: < 30s
## Report Format
Reports are saved in two formats:
### 1. Text Report (`quality-report-YYYYMMDD-HHMMSS.txt`)
```
================================================================================
CORE-AI QUALITY REPORT
Generated: 2025-12-01T10:30:45
Git Tag: v2.1.0
Git Commit: a3b2c1d
Git Branch: main
Implementation:
Agent: OllamaNativeAgent
Model: mistral-nemo:latest
Framework: PydanticAI
API: Ollama native (/api/chat)
================================================================================
Total Tests: 5
Passed: 5 (100.0%)
Failed: 0
Performance:
Average response time: 8.45s
Fastest response: 3.21s
Slowest response: 15.67s
Test Details:
--------------------------------------------------------------------------------
1. Scenario 1: Simple Knowledge: ✓ PASS
Query: What is Docker?...
Response time: 3.21s
Tools available: 6
Response length: 245 chars
...
================================================================================
REVERT INSTRUCTIONS:
If quality has degraded, revert to: v2.1.0
git checkout v2.1.0
================================================================================
```
### 2. JSON Report (`quality-report-YYYYMMDD-HHMMSS.json`)
Machine-readable format for programmatic analysis and trend tracking:
```json
{
"timestamp": "2025-12-01T10:30:45",
"git_info": {
"tag": "v2.1.0",
"commit": "a3b2c1d",
"branch": "main"
},
"summary": {
"total": 5,
"passed": 5,
"failed": 0
},
"results": [...]
}
```
## Workflow: Before Making Changes
### 1. Establish Baseline
Before making any code changes, run the test suite to establish a quality baseline:
```bash
cd /home/jpmschweitzer/Projects/portainer-core/services/core-ai
python tests/test_ai_flow_quality.py
```
**Save the report location** - you'll compare against this later.
### 2. Make Your Changes
Edit agent code, tools, prompts, etc.
### 3. Run Tests Again
```bash
python tests/test_ai_flow_quality.py
```
### 4. Compare Reports
Compare the new report against the baseline:
```bash
# List recent reports
ls -lh tests/reports/
# Compare two reports
diff tests/reports/quality-report-20251201-103045.txt \
tests/reports/quality-report-20251201-115522.txt
```
**Key metrics to watch**:
- Pass rate (should stay 100%)
- Average response time (should not significantly increase)
- Individual test failures (investigate immediately)
### 5. Revert if Quality Degrades
If tests fail or performance degrades significantly:
```bash
# Check the git tag from the failing report
cat tests/reports/quality-report-20251201-115522.txt | grep "Git Tag"
# Revert to that tag
git checkout v2.1.0
```
## Benchmarking Models
To compare different models:
### 1. Run baseline with current model
```bash
python tests/test_ai_flow_quality.py
# Save this as baseline
```
### 2. Change model in config
Edit `services/core-ai/src/config.py` or environment variable:
```python
# Change from mistral-nemo:latest to gemma2:9b
AGENT_MODEL = "gemma2:9b"
```
Restart core-ai:
```bash
cd /home/jpmschweitzer/Projects/portainer-core/stacks
docker restart core-ai
```
### 3. Run tests with new model
```bash
python tests/test_ai_flow_quality.py
```
### 4. Compare results
```bash
# Check both JSON reports for performance comparison
cat tests/reports/quality-report-BASELINE.json | jq '.summary'
cat tests/reports/quality-report-NEW_MODEL.json | jq '.summary'
```
Look for:
- **Pass rate changes** - Did the new model fail any tests?
- **Response time changes** - Is it faster or slower?
- **Response quality** - Are answers as good?
## Troubleshooting
### Tests Fail: "Connection refused"
**Problem**: core-ai service not running
**Solution**:
```bash
cd /home/jpmschweitzer/Projects/portainer-core/stacks
docker restart core-ai
docker logs core-ai # Check for startup errors
```
### Tests Timeout
**Problem**: Model too slow or stuck
**Solution**:
1. Check Ollama GPU usage: `nvidia-smi`
2. Check model is loaded: `docker exec ollama ollama list`
3. Increase timeout in test file if needed
### Web Search Tests Fail
**Problem**: SearXNG not available
**Solution**:
```bash
docker restart searxng
curl "http://localhost:8087/search?q=test&format=json"
```
### Calculation Tests Fail
**Problem**: Agent not using calculate tool
**Solution**: Check tool registration:
```bash
curl http://localhost:8086/v1/tools | jq '.tools[] | .name'
```
## Adding New Test Scenarios
### 1. Add test function
```python
@pytest.mark.asyncio
async def test_my_new_scenario():
"""
Scenario: My New Feature
Expected: Describe expected behavior
Performance target: < Xs
"""
async with AIFlowTester() as tester:
result = await tester.chat("My test query")
tester.assert_response_quality(
result,
expected_keywords=["keyword1", "keyword2"],
min_length=50,
max_time=15.0
)
# Custom assertions
assert "expected result" in result["response"]
print(f"✓ My scenario: {result['total_time']:.2f}s")
```
### 2. Add to scenario list
In `run_full_quality_check()`:
```python
test_scenarios = [
# ... existing scenarios ...
{
"name": "Scenario X: My New Feature",
"query": "My test query",
"test": test_my_new_scenario
},
]
```
### 3. Run to verify
```bash
pytest tests/test_ai_flow_quality.py::test_my_new_scenario -v
```
## Best Practices
### ✅ DO:
- Run tests before committing major changes
- Compare reports to detect regressions
- Save baseline reports for each release
- Document expected behavior in test docstrings
- Use meaningful git tags for easy reversion
### ❌ DON'T:
- Skip tests when making agent changes
- Ignore performance degradation warnings
- Delete old reports (keep for trend analysis)
- Change test expectations to make tests pass
- Commit without running tests first
## Report Retention
Keep reports organized:
```bash
# Keep last 30 days of reports
find tests/reports/ -name "*.txt" -mtime +30 -delete
find tests/reports/ -name "*.json" -mtime +30 -delete
# Archive reports by month
mkdir -p tests/reports/archive/2025-12/
mv tests/reports/quality-report-202512*.* tests/reports/archive/2025-12/
```
---
**Remember**: These tests protect quality. If they fail, investigate before proceeding!
@@ -0,0 +1,605 @@
"""
AI Flow Quality Test Suite
Purpose: Benchmark core-ai agent behavior for regression detection
and performance tracking over time.
This test suite validates:
- Tool calling decision-making (OllamaNativeAgent)
- Multi-step reasoning capability
- Response quality and accuracy
- Performance characteristics
- Regression detection across code changes
Current Implementation:
- Agent: OllamaNativeAgent (Ollama native API with tool calling)
- Model: mistral-nemo:latest (primary reasoning model)
- Tools: Local tools (web_search, calculate, date/time operations)
- Framework: PydanticAI
Usage:
pytest tests/test_ai_flow_quality.py -v
pytest tests/test_ai_flow_quality.py::test_simple_knowledge -v
python tests/test_ai_flow_quality.py # Run and generate report
Reports saved to: tests/reports/
"""
import asyncio
import time
import json
import httpx
import subprocess
from pathlib import Path
from typing import Dict, List, Any, Optional
from datetime import datetime
import pytest
# Test Configuration
BASE_URL = "http://localhost:8086"
TIMEOUT = 60.0
REPORTS_DIR = Path(__file__).parent / "reports"
def get_git_info() -> Dict[str, str]:
"""Get current git tag and commit hash"""
try:
# Get current tag
tag = subprocess.check_output(
["git", "describe", "--tags", "--exact-match"],
stderr=subprocess.DEVNULL,
text=True
).strip()
except subprocess.CalledProcessError:
# No exact tag, get latest tag + commit
try:
tag = subprocess.check_output(
["git", "describe", "--tags", "--always"],
text=True
).strip()
except subprocess.CalledProcessError:
tag = "unknown"
try:
commit = subprocess.check_output(
["git", "rev-parse", "--short", "HEAD"],
text=True
).strip()
except subprocess.CalledProcessError:
commit = "unknown"
try:
branch = subprocess.check_output(
["git", "rev-parse", "--abbrev-ref", "HEAD"],
text=True
).strip()
except subprocess.CalledProcessError:
branch = "unknown"
return {"tag": tag, "commit": commit, "branch": branch}
class AIFlowTester:
"""Test harness for AI flow quality checks"""
def __init__(self):
self.results = []
self.client = None
self.git_info = get_git_info()
async def __aenter__(self):
self.client = httpx.AsyncClient(timeout=TIMEOUT)
return self
async def __aexit__(self, exc_type, exc_val, exc_tb):
if self.client:
await self.client.aclose()
async def chat(
self,
message: str,
enable_tools: bool = True,
stream: bool = False
) -> Dict[str, Any]:
"""
Send a chat request and measure performance
Returns:
{
"response": str,
"time_to_first_token": float,
"total_time": float,
"tools_count": int,
"success": bool,
"error": str | None
}
"""
start_time = time.time()
time_to_first_token = None
response_text = ""
try:
payload = {
"messages": [{"role": "user", "content": message}],
"stream": stream,
"enable_tools": enable_tools
}
response = await self.client.post(
f"{BASE_URL}/v1/chat/completions",
json=payload
)
response.raise_for_status()
if not stream:
time_to_first_token = time.time() - start_time
data = response.json()
response_text = data["choices"][0]["message"]["content"]
tools_count = data.get("tools_count", 0)
tools_enabled = data.get("tools_enabled", False)
else:
# Handle streaming response
# TODO: Implement streaming support
raise NotImplementedError("Streaming not yet implemented")
total_time = time.time() - start_time
return {
"response": response_text,
"time_to_first_token": time_to_first_token,
"total_time": total_time,
"tools_count": tools_count if tools_enabled else 0,
"success": True,
"error": None
}
except Exception as e:
total_time = time.time() - start_time
return {
"response": "",
"time_to_first_token": None,
"total_time": total_time,
"tools_count": 0,
"success": False,
"error": str(e)
}
def assert_response_quality(
self,
result: Dict[str, Any],
expected_keywords: List[str] = None,
min_length: int = 20,
max_time: float = 30.0
):
"""
Validate response quality
Args:
result: Test result from chat()
expected_keywords: Keywords that should appear in response
min_length: Minimum response length
max_time: Maximum acceptable response time
"""
assert result["success"], f"Request failed: {result['error']}"
assert len(result["response"]) >= min_length, \
f"Response too short: {len(result['response'])} < {min_length}"
assert result["total_time"] <= max_time, \
f"Response too slow: {result['total_time']:.2f}s > {max_time}s"
if expected_keywords:
response_lower = result["response"].lower()
for keyword in expected_keywords:
assert keyword.lower() in response_lower, \
f"Missing keyword '{keyword}' in response"
def generate_report(self) -> str:
"""Generate a quality report from test results"""
if not self.results:
return "No test results to report"
report = []
report.append("=" * 80)
report.append("CORE-AI QUALITY REPORT")
report.append(f"Generated: {datetime.now().isoformat()}")
report.append(f"Git Tag: {self.git_info['tag']}")
report.append(f"Git Commit: {self.git_info['commit']}")
report.append(f"Git Branch: {self.git_info['branch']}")
report.append("")
report.append("Implementation:")
report.append(" Agent: OllamaNativeAgent")
report.append(" Model: mistral-nemo:latest")
report.append(" Framework: PydanticAI")
report.append(" API: Ollama native (/api/chat)")
report.append("=" * 80)
report.append("")
# Summary statistics
total_tests = len(self.results)
passed_tests = sum(1 for r in self.results if r.get("passed", False))
failed_tests = total_tests - passed_tests
report.append(f"Total Tests: {total_tests}")
report.append(f"Passed: {passed_tests} ({100*passed_tests/total_tests:.1f}%)")
report.append(f"Failed: {failed_tests}")
report.append("")
# Performance metrics
response_times = [r["result"]["total_time"] for r in self.results
if r["result"]["success"]]
if response_times:
avg_time = sum(response_times) / len(response_times)
min_time = min(response_times)
max_time = max(response_times)
report.append("Performance:")
report.append(f" Average response time: {avg_time:.2f}s")
report.append(f" Fastest response: {min_time:.2f}s")
report.append(f" Slowest response: {max_time:.2f}s")
report.append("")
# Individual test results
report.append("Test Details:")
report.append("-" * 80)
for i, test in enumerate(self.results, 1):
status = "✓ PASS" if test.get("passed", False) else "✗ FAIL"
report.append(f"\n{i}. {test['name']}: {status}")
report.append(f" Query: {test['query'][:60]}...")
result = test["result"]
if result["success"]:
report.append(f" Response time: {result['total_time']:.2f}s")
report.append(f" Tools available: {result['tools_count']}")
report.append(f" Response length: {len(result['response'])} chars")
if test.get("error"):
report.append(f" ⚠ Assertion failed: {test['error']}")
else:
report.append(f" ✗ Error: {result['error']}")
report.append("")
report.append("=" * 80)
report.append("REVERT INSTRUCTIONS:")
report.append(f"If quality has degraded, revert to: {self.git_info['tag']}")
report.append(f" git checkout {self.git_info['tag']}")
report.append("=" * 80)
return "\n".join(report)
# ============================================================================
# TEST SCENARIOS
# ============================================================================
@pytest.mark.asyncio
async def test_scenario1_simple_knowledge():
"""
Scenario 1: Simple Knowledge Query (No Tools Needed)
Expected: Direct answer without requiring tools
Performance target: < 10s
"""
async with AIFlowTester() as tester:
result = await tester.chat("What is Docker?")
tester.assert_response_quality(
result,
expected_keywords=["container", "platform"],
min_length=50,
max_time=10.0
)
assert result["success"], "Request failed"
print(f"✓ Simple knowledge query: {result['total_time']:.2f}s")
@pytest.mark.asyncio
async def test_scenario2_web_search():
"""
Scenario 2: Web Search Required
Expected: Uses web_search tool via SearXNG, synthesizes results
Performance target: < 30s
"""
async with AIFlowTester() as tester:
result = await tester.chat(
"What are the latest Kubernetes security best practices?"
)
tester.assert_response_quality(
result,
expected_keywords=["kubernetes", "security"],
min_length=100,
max_time=30.0
)
assert result["success"], "Request failed"
print(f"✓ Web search query: {result['total_time']:.2f}s")
print(f" Tools available: {result['tools_count']}")
@pytest.mark.asyncio
async def test_scenario3_calculation():
"""
Scenario 3: Mathematical Calculation
Expected: Uses calculate tool for accurate results
Performance target: < 15s
"""
async with AIFlowTester() as tester:
result = await tester.chat(
"Calculate 2847 * 1923 + 5612 - 999"
)
tester.assert_response_quality(
result,
min_length=20,
max_time=15.0
)
assert result["success"], "Request failed"
# Verify correct answer: 5,479,394
response_clean = result["response"].replace(",", "").replace(" ", "")
assert "5479394" in response_clean, \
"Calculation result not found in response"
print(f"✓ Calculation query: {result['total_time']:.2f}s")
@pytest.mark.asyncio
async def test_scenario4_date_operations():
"""
Scenario 4: Date/Time Operations
Expected: Uses date/time tools for accurate results
Performance target: < 15s
"""
async with AIFlowTester() as tester:
result = await tester.chat(
"What is the current date and what will the date be 30 days from now?"
)
tester.assert_response_quality(
result,
expected_keywords=["date", "2025"],
min_length=50,
max_time=15.0
)
assert result["success"], "Request failed"
print(f"✓ Date operation query: {result['total_time']:.2f}s")
@pytest.mark.asyncio
async def test_scenario5_multi_tool_reasoning():
"""
Scenario 5: Multi-Tool Complex Query
Expected: Multiple tool calls, synthesizes results
Performance target: < 30s
"""
async with AIFlowTester() as tester:
result = await tester.chat(
"Get the current time in New York and Tokyo, then calculate the time difference in hours"
)
tester.assert_response_quality(
result,
expected_keywords=["time", "hour"],
min_length=80,
max_time=30.0
)
assert result["success"], "Request failed"
print(f"✓ Multi-tool query: {result['total_time']:.2f}s")
@pytest.mark.asyncio
async def test_tools_disabled():
"""
Test: Agent with Tools Disabled
Expected: Works without tool access, generates knowledge-based response
"""
async with AIFlowTester() as tester:
result = await tester.chat(
"What is Python?",
enable_tools=False
)
tester.assert_response_quality(
result,
expected_keywords=["python", "programming"],
min_length=30,
max_time=10.0
)
assert result["tools_count"] == 0, "Tools should be disabled"
print(f"✓ No-tools query: {result['total_time']:.2f}s")
# ============================================================================
# PERFORMANCE BENCHMARKS
# ============================================================================
@pytest.mark.asyncio
async def test_performance_baseline():
"""
Performance Baseline Test
Establishes baseline metrics for comparison across versions
"""
async with AIFlowTester() as tester:
queries = [
"What is containerization?",
"Explain Kubernetes in one sentence",
"What does API stand for?",
]
times = []
for query in queries:
result = await tester.chat(query)
assert result["success"], f"Query failed: {query}"
times.append(result["total_time"])
avg_time = sum(times) / len(times)
print(f"\n✓ Performance baseline:")
print(f" Average response time: {avg_time:.2f}s")
print(f" Range: {min(times):.2f}s - {max(times):.2f}s")
# Assert reasonable performance
assert avg_time < 10.0, f"Average response too slow: {avg_time:.2f}s"
# ============================================================================
# REGRESSION TESTS
# ============================================================================
@pytest.mark.asyncio
async def test_regression_tool_availability():
"""
Regression: Verify all expected tools are available
"""
async with AIFlowTester() as tester:
response = await tester.client.get(f"{BASE_URL}/v1/tools")
response.raise_for_status()
data = response.json()
tools = {tool["name"] for tool in data["tools"]}
# Expected tools from local.py
expected_tools = {
"get_current_time",
"get_current_date",
"calculate_date_difference",
"add_days_to_date",
"calculate",
"web_search"
}
for tool in expected_tools:
assert tool in tools, f"Missing tool: {tool}"
print(f"✓ Tool availability: {len(tools)} tools registered")
@pytest.mark.asyncio
async def test_regression_health_check():
"""
Regression: Health check endpoint works
"""
async with AIFlowTester() as tester:
response = await tester.client.get(f"{BASE_URL}/health")
response.raise_for_status()
data = response.json()
assert data["status"] == "ok", "Service not healthy"
assert data["service"] == "core-ai"
assert data["agents"]["ollama-native"] is True, \
"OllamaNativeAgent not available"
print(f"✓ Health check passed")
print(f" Default agent: {data['default_agent']}")
print(f" Tools count: {data['tools_count']}")
# ============================================================================
# MAIN RUNNER (for standalone execution)
# ============================================================================
async def run_full_quality_check():
"""Run all tests and generate comprehensive report"""
tester = AIFlowTester()
test_scenarios = [
{
"name": "Scenario 1: Simple Knowledge",
"query": "What is Docker?",
"test": test_scenario1_simple_knowledge
},
{
"name": "Scenario 2: Web Search",
"query": "What are the latest Python 3.13 features?",
"test": test_scenario2_web_search
},
{
"name": "Scenario 3: Calculation",
"query": "Calculate 2847 * 1923 + 5612 - 999",
"test": test_scenario3_calculation
},
{
"name": "Scenario 4: Date Operations",
"query": "What is the current date and what will it be in 30 days?",
"test": test_scenario4_date_operations
},
{
"name": "Scenario 5: Multi-tool Reasoning",
"query": "Get current time in NYC and Tokyo, calculate difference",
"test": test_scenario5_multi_tool_reasoning
},
]
print("\n" + "=" * 80)
print("CORE-AI QUALITY CHECK")
print(f"Started: {datetime.now().isoformat()}")
print(f"Git Tag: {tester.git_info['tag']}")
print(f"Git Commit: {tester.git_info['commit']}")
print("=" * 80 + "\n")
async with tester:
for scenario in test_scenarios:
print(f"\nRunning: {scenario['name']}")
print(f"Query: {scenario['query']}")
print("-" * 80)
try:
await scenario["test"]()
tester.results.append({
"name": scenario["name"],
"query": scenario["query"],
"result": {"success": True, "total_time": 0, "tools_count": 0, "response": ""},
"passed": True
})
except Exception as e:
tester.results.append({
"name": scenario["name"],
"query": scenario["query"],
"result": {"success": False, "error": str(e), "total_time": 0, "tools_count": 0, "response": ""},
"passed": False,
"error": str(e)
})
print(f"✗ FAILED: {e}")
# Generate and print report
report = tester.generate_report()
print("\n" + report)
# Ensure reports directory exists
REPORTS_DIR.mkdir(parents=True, exist_ok=True)
# Save report to file
timestamp = datetime.now().strftime('%Y%m%d-%H%M%S')
report_file = REPORTS_DIR / f"quality-report-{timestamp}.txt"
with open(report_file, "w") as f:
f.write(report)
print(f"\nReport saved to: {report_file}")
# Also save as JSON for programmatic analysis
json_file = REPORTS_DIR / f"quality-report-{timestamp}.json"
json_data = {
"timestamp": datetime.now().isoformat(),
"git_info": tester.git_info,
"summary": {
"total": len(tester.results),
"passed": sum(1 for r in tester.results if r.get("passed", False)),
"failed": sum(1 for r in tester.results if not r.get("passed", False))
},
"results": tester.results
}
with open(json_file, "w") as f:
json.dump(json_data, f, indent=2)
print(f"JSON report saved to: {json_file}")
if __name__ == "__main__":
asyncio.run(run_full_quality_check())