# Core Code API OpenAPI-compatible functions for Open WebUI, providing web scraping and data processing capabilities. ## Features ### Web Scraper - Intelligent content extraction using Trafilatura - BeautifulSoup fallback for complex pages - Configurable content length limits - Optional link extraction - Perfect for feeding webpage content to LLMs ## Architecture ``` src/ ├── config.py # Global application settings ├── logging_config.py # Logging configuration ├── base_schema.py # Base Pydantic models ├── main.py # FastAPI application entry point └── web_scraper/ # Web scraper module ├── __init__.py ├── config.py # Module-specific settings ├── schemas.py # Pydantic request/response models ├── service.py # Business logic ├── router.py # API routes └── exceptions.py # Custom exceptions ``` ## Development ### Requirements - Python 3.12+ - Docker (for containerized deployment) ### Local Development ```bash # Install dependencies pip install -r requirements.txt # Run locally uvicorn src.main:app --reload --host 0.0.0.0 --port 8083 ``` ### Adding New Dependencies **Important**: Dependencies use major version pinning (`~=`) for automatic patch updates while preventing breaking changes. 1. Add package to `requirements.txt` with major version constraint: ``` package-name~=1.2.0 # Allows 1.2.x, blocks 1.3.0 ``` 2. Restart the container to install: ```bash docker restart core-api ``` The container automatically runs `pip install -r requirements.txt` on every boot, so new dependencies are installed immediately on restart. **Version Pinning Best Practices**: - Use `~=` (compatible release) for most packages: `fastapi~=0.115.0` - Use `>=X,=0.3.17,<0.4.0` - Allows automatic security patches without breaking changes - Documented in PEP 440 ### Docker Build ```bash # Build image docker build -t core-code:latest . # Run container docker run -p 8083:8083 core-code:latest ``` ## Deployment ### Portainer Stack 1. Navigate to Portainer UI 2. Go to **Stacks** → **Add Stack** 3. Name: `core-code` 4. Upload `stacks/core-code.yml` or paste contents 5. Deploy ### Environment Variables See `.env.example` for all available configuration options. ## API Documentation Once deployed, access documentation at: - **Swagger UI**: http://192.168.86.149:8083/docs - **ReDoc**: http://192.168.86.149:8083/redoc - **OpenAPI Spec**: http://192.168.86.149:8083/openapi.json ## Integration with Open WebUI ### Method 1: Functions (OpenAPI Import) 1. In Open WebUI, navigate to Functions 2. Import from OpenAPI spec: `http://192.168.86.149:8083/openapi.json` 3. Use functions directly in chat ### Method 2: Pipelines 1. Create a pipeline that calls Core Code API endpoints 2. Use as data source for LLM workflows ### Method 3: Direct API Calls ```python import httpx async with httpx.AsyncClient() as client: response = await client.post( "http://192.168.86.149:8083/web-scraper/scrape", json={ "url": "https://example.com", "extract_main_content": True } ) data = response.json() ``` ## API Endpoints ### Web Scraper **POST /web-scraper/scrape** Scrape and extract content from a website. Request: ```json { "url": "https://example.com/article", "extract_main_content": true, "include_links": false, "max_length": 10000 } ``` Response: ```json { "url": "https://example.com/article", "title": "Article Title", "content": "Extracted article content...", "extracted_at": "2025-11-12T19:30:00Z", "content_length": 5432, "links": null } ``` ## Logging Logs are written to: - **Console**: stdout (captured by Docker) - **File**: `/app/logs/app.log` (persisted via volume mount) Log format: ``` 2025-11-12 19:30:00 | INFO | src.web_scraper.service:scrape_url:45 | Starting scrape for URL: https://example.com ``` ## Health Checks - **Endpoint**: `GET /health` - **Docker**: Automatic health checks configured - **Response**: `{"status": "healthy"}` ## Security - Runs as non-root user (uid 1000) - No authentication required (internal network only) - CORS configured for same-network access - Rate limiting: Not implemented (internal use only) ## Future Modules The architecture supports adding new modules: - Data transformation functions - API integrations - File processing - Database queries Each module follows the same structure: ``` src/ └── module_name/ ├── config.py ├── schemas.py ├── service.py ├── router.py └── exceptions.py ``` ## Troubleshooting ### Container won't start ```bash docker logs core-code ``` ### API not responding ```bash curl http://192.168.86.149:8083/health ``` ### Check OpenAPI spec ```bash curl http://192.168.86.149:8083/openapi.json | jq ``` ## License Internal use only.