Files
odysseus/specs/model-providers/vllm.md
RaresKeYandStressTestor 7026cf40b5 docs: bootstrap specs ground truth (#5794)
* docs(specs): restore bootstrap after dev rewrite

* docs(specs): remove runtime inventory snapshot

* docs(specs): reconcile current dev truth

* docs(specs): document scheduled task actions as an owner-attribution source

Owner Attribution covered cookie, bearer-token and internal-loopback
requests. Scheduled task actions are a fourth source and behave
differently: _execute_action passes owner=task.owner off the stored
ScheduledTask row, so no request and no resolved principal are in
flight, and route-level require_user() never runs.

Webhook triggers are the sharp case. They are unauthenticated by
design with the token as the only credential and execute under the
stored task.owner.

Paths cite routes/task/task_routes.py, the canonical location after
the task subpackage move (#6081); routes/task_routes.py on current dev
is the backward-compat shim.

* docs(specs): add chained tasks to the trigger list, refresh dev stamp

Review feedback from RaresKeY on the previous commit.

"Every trigger path" was too broad: success-chained tasks are another
path into _execute_action. Added them with their own citation, and
noted that chaining additionally requires the target task to share
task.owner and rejects cycles, which is stricter than the trigger-side
checks. Softened the lead-in to "these trigger paths".

Line 56 still pointed at routes/task_routes.py for webhook credential
validation. That path is the backward-compat shim on current dev after
the task subpackage move (#6081); repointed to the canonical
routes/task/task_routes.py.

Stamp moved to dev@2a6b09b. Inspection backing that bump was scoped:
every file path cited in this spec was mechanically checked to resolve
on 2a6b09b, and every file:line in the Owner Attribution additions was
read against it. Behavioral claims elsewhere in the file were not
re-audited.

* docs(specs): correct SECURE_COOKIES description to match current behavior

Third of the stale details RaresKeY enumerated. The cookie section
described SECURE_COOKIES as purely opt-in, which stopped being true.

_secure_cookie() (routes/auth_routes.py:89) treats an explicit true or
false as authoritative and derives the Secure attribute from the
request otherwise, including when the variable is unset and when
docker-compose injects it present-but-empty. Either the connection
scheme or the first X-Forwarded-Proto hop being https is enough.

* docs(specs): refresh current dev truth

---------

Co-authored-by: StressTestor <212606152+StressTestor@users.noreply.github.com>
2026-08-25 14:18:44 +02:00

1.9 KiB

vLLM Provider Shape

Last updated: dev@e57f60b | 2026-07-20

Scope

Canonical placeholder provider ID vllm; OpenAI Chat and Responses serving; generic identity-only inventory normalization. There is no dedicated vLLM reader or model-card detector on current dev.

Catalog Shape

Current GET /v1/models returns object: list, data[] model cards with id, object, owned_by: vllm, root, parent, max_model_len, and permission[]. The generic reader retains only identity/raw data and does not inspect owned_by, root, parent, max_model_len, or permission. The card does not prove chat template, tools, reasoning parser, vision assets, embeddings, transcription, or rerank.

LoRA cards can use a different id, root path, and parent. Keep each served ID endpoint scoped and do not merge it globally with the base checkpoint.

Runtime Capability

vLLM's supported API surface is broad, but actual behavior depends on the loaded model task, chat template, multimodal assets, tool-call parser, reasoning parser, structured-output configuration, and launch flags. Current Odysseus reasoning regressions cover structured reasoning, legacy reasoning_content, and compatible fields (#602). These response channels are transport evidence, not a claim that every vLLM model reasons.

Fallback And Safety

Current reader detection identifies port 8000 as vLLM, or accepts an explicit endpoint kind, then dispatches to the generic identity-only reader. It does not infer vLLM from the model-card payload. Do not consume /server_info environment/config dumps for normal discovery because they can be large and operationally sensitive.

Current Gaps

  • A small safe native capability endpoint is not part of the canonical probe.
  • Deployment parser/template flags are not persisted with endpoint capability.
  • No dedicated reader maps vLLM model-card fields today.