Files
odysseus/specs/model-providers/ollama.md
T
RaresKeYandStressTestor 7026cf40b5 docs: bootstrap specs ground truth (#5794)
* docs(specs): restore bootstrap after dev rewrite

* docs(specs): remove runtime inventory snapshot

* docs(specs): reconcile current dev truth

* docs(specs): document scheduled task actions as an owner-attribution source

Owner Attribution covered cookie, bearer-token and internal-loopback
requests. Scheduled task actions are a fourth source and behave
differently: _execute_action passes owner=task.owner off the stored
ScheduledTask row, so no request and no resolved principal are in
flight, and route-level require_user() never runs.

Webhook triggers are the sharp case. They are unauthenticated by
design with the token as the only credential and execute under the
stored task.owner.

Paths cite routes/task/task_routes.py, the canonical location after
the task subpackage move (#6081); routes/task_routes.py on current dev
is the backward-compat shim.

* docs(specs): add chained tasks to the trigger list, refresh dev stamp

Review feedback from RaresKeY on the previous commit.

"Every trigger path" was too broad: success-chained tasks are another
path into _execute_action. Added them with their own citation, and
noted that chaining additionally requires the target task to share
task.owner and rejects cycles, which is stricter than the trigger-side
checks. Softened the lead-in to "these trigger paths".

Line 56 still pointed at routes/task_routes.py for webhook credential
validation. That path is the backward-compat shim on current dev after
the task subpackage move (#6081); repointed to the canonical
routes/task/task_routes.py.

Stamp moved to dev@2a6b09b. Inspection backing that bump was scoped:
every file path cited in this spec was mechanically checked to resolve
on 2a6b09b, and every file:line in the Owner Attribution additions was
read against it. Behavioral claims elsewhere in the file were not
re-audited.

* docs(specs): correct SECURE_COOKIES description to match current behavior

Third of the stale details RaresKeY enumerated. The cookie section
described SECURE_COOKIES as purely opt-in, which stopped being true.

_secure_cookie() (routes/auth_routes.py:89) treats an explicit true or
false as authoritative and derives the Secure attribute from the
request otherwise, including when the variable is unset and when
docker-compose injects it present-but-empty. Either the connection
scheme or the first X-Forwarded-Proto hop being https is enough.

* docs(specs): refresh current dev truth

---------

Co-authored-by: StressTestor <212606152+StressTestor@users.noreply.github.com>
2026-08-25 14:18:44 +02:00

2.5 KiB

Ollama Provider Shape

Last updated: dev@e71f8ce | 2026-08-25

Scope

Canonical provider ID ollama; native Ollama chat/generate plus OpenAI compatibility; reader src/model_capability_readers/ollama.py; discovery and runtime code in routes/model_routes.py and src/llm_core.py.

Catalog And Detail Shapes

Use two native steps:

  1. GET /api/tags returns models[] identity (name/model, digest, details.family|families, format, parameter size, quantization). Tags do not claim capabilities.
  2. POST /api/show for a selected model returns explicit capabilities[], details, and model_info. Map completion/chat, embedding, vision, tools, and thinking/reasoning tokens. Map context from exact context_length or native <architecture>.context_length fields.

The reader does not parse model names or architecture names. It does parse a two-column serialized parameters value and can take num_ctx from it before falling back to exact or suffix *.context_length keys in structured mappings. The parameters text is used only for that keyed limit lookup, not capability inference.

Request And Response Shape

Native chat uses /api/chat, messages, optional OpenAI-shaped tool definitions, format, options, and model-dependent think. Responses use message.content, message.thinking, and message.tool_calls. Generate uses top-level response and thinking. OpenAI compatibility is a separate dialect and can change control names independently.

Manual Ollama endpoints registered against the OpenAI-compatible /v1 surface default to text/prompted tools unless the operator explicitly enables supports_tools; model naming alone does not opt that dialect into native function schemas.

Thinking control is model-specific: most documented reasoning families accept a native bool, while GPT-OSS accepts low/medium/high and cannot be fully disabled. A reported Ollama 0.20.6 Qwen3.5 OpenAI-compat path requires reasoning_effort: none rather than think: false (#5503); keep it versioned and low-confidence until corroborated.

Fallback And Safety

Current reader detection identifies port 11434 as Ollama, in addition to an explicit endpoint kind or an exact/label-bounded ollama.com hostname. This is a normalization hint, not endpoint trust or capability evidence. Names that contain vision, embed, or qwen are not capability evidence (#3743, #4487).

Current Gaps

  • List discovery needs an orchestrated /api/show detail step per model.
  • Runtime OpenAI-compat thinking suppression still contains name heuristics.