Compare commits

...
28 Commits
Author SHA1 Message Date
jpmschweitzer 014afc960c fix(permissions): narrow rm -rf deny globs to their exact forms
The trailing wildcard on the three rm -rf deny entries spanned path
separators, so Bash(rm -rf /*) matched every absolute path on the
machine rather than the filesystem root, and the ~ and $HOME entries
had the same shape. Narrowed to the exact literal forms.

These rules match literal command text, so they still stop a typo on
rm -rf /, rm -rf ~ or rm -rf $HOME exactly, but they no longer stop a
recursive delete aimed at any other path. That reduced cover is
deliberate, not an oversight.
2026-08-25 20:31:27 +02:00
jpmschweitzerandClaude f398a8ad80 fix(agents): AgentInterface declared a coroutine where every caller wants a generator
`generate_response` was `async def` with a `pass` body and no `yield`.
An async function that never yields is a coroutine, so the declared type
was Coroutine[..., AsyncGenerator[OutputItem, None]] — something a
caller must await before it can be iterated.

Nobody awaits it. Both implementations contain yields (TatlockAgent 5,
LoremTesterAgent 3), which makes them async generators directly, and
both call sites do `async for item in agent.generate_response(...)`.
The abstract method's own docstring says "Yields:" and its own example
iterates the call without awaiting. Implementations, consumers and prose
all agreed; only the declaration dissented.

Removing one word fixes it, and it is the declaration that was wrong
rather than the four places reporting it.

WHY NOTHING CAUGHT THIS. The abstract body is `pass` and nothing calls
super().generate_response — verified across the tree — so the wrong
declaration has no runtime consequence and cannot fail a test. It was
invisible by construction, and it presented as four unrelated errors in
four files (two override, two attr-defined), none of which named the
cause. Anyone fixing them where they appeared would have annotated the
implementations to match the interface and made the real defect
permanent.

tests/agents/test_agent_interface.py covers it going forward. The
load-bearing case is not "the interface is X" or "the implementation is
Y" separately — both could drift together and still pass — but that the
two AGREE about what kind of callable this is.

Mutation-checked: restoring `async` fails 3 of the 5 new tests, the two
that still pass being the implementation checks, which are correctly
unaffected. Anchor asserted unique before the mutation was written, and
the fix asserted back into place afterwards.

95 errors -> 71 across this branch; this commit accounts for 4 of them.
Suite 662 passed, 1 failed — that failure is the known LLM-nondeterministic
calculator test, which passed on the previous run of this same branch and
failed on this one, which is the clearest available evidence that it is
unrelated to any of this work.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-20 00:24:12 +02:00
jpmschweitzerandClaude eabf0f4a11 chore(types): annotate fifteen signatures mypy could not check
Twelve gain `-> None`, each confirmed by AST to contain no returning
`return` and no `yield` rather than by reading the name and assuming.
The three context-manager exits gain the canonical
type[BaseException]/BaseException/TracebackType argument triple.

Both files taking TracebackType needed the import, and inserting it
before the first import broke ruff's I001 — lint was exit 0 at the
baseline commit, verified by stashing this work and re-running, so that
breakage was mine. Fixed with `ruff check --fix` on the two files, which
placed the import in sorted position.

86 errors -> 75; no-untyped-def 29 -> 14.

Suite: 658 passed. The baseline was 657 passed with one failure in
test_tatlock_tool_call_logging_calculator, which asserts on the content
of a live model's reply. It passing here is nondeterminism, NOT evidence
this commit fixed anything, and it may fail again on the next run.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-20 00:18:35 +02:00
jpmschweitzerandClaude 7c5fca06c3 chore(types): annotate four containers mypy could not infer
Each element type is taken from how the container is used rather than
guessed: kept_items is returned from trim_to_fit, whose signature is
already list[Any]; traces collects the dicts built at the append site;
expert_results and tool_outputs are keyed by tool_name (str) and hold
ToolReturnPart.content.

tracing_router.py needed `from typing import Any` added — it had no
import for it, so annotating without that would have traded a
var-annotated error for a name-defined one. Function-body variable
annotations are not evaluated at runtime, so this would not have raised;
mypy was the only thing that would have caught it.

90 errors -> 86.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-19 16:30:22 +02:00
jpmschweitzerandClaude 3696a40f97 chore(types): delete five type: ignore comments that suppress nothing
mypy's warn_unused_ignores is on, so these were reported as errors in
their own right — a suppression that no longer suppresses is a claim
that something is broken when it is not, and it silently widens to
cover a real error if one later appears on that line.

Two carried a "Forward reference" note that is still accurate; the note
is kept and only the ignore removed.

95 errors -> 90. Comment-only, so no runtime behaviour can have changed
and the suite was not re-run for this commit. The `# type: ignore` count
across the tree drops from 6 to 1, which is the anchor for the rest of
this work: clearing a type error by suppressing it would push that
number the other way.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-19 16:28:28 +02:00
jpmschweitzer 1ab6c1b379 fix(build): make setup fail when the environment doesn't actually work
pip install exiting 0 is not evidence the venv is usable (D-24) - the
2026-08-09 core-api incident was exactly this shape: a venv that
"installed fine" but was missing sqlalchemy, surfacing as 11 collection
errors that read like broken imports rather than an environment
problem.

setup now ends with `pytest --collect-only`, scoped like `make test`
(excludes e2e/integration/contracts) and run with --no-cov. Collection
imports every test module without running the suite, so a missing or
mismatched dependency fails setup itself instead of showing up later
as a confusing test failure.

Workspace T-47.
2026-08-17 12:03:58 +02:00
jpmschweitzerandClaude 1e986f28b3 chore(pql): file T-2 — revoked ANTHROPIC_API_KEY still in .env
Filed separately from the settings work in tatlock-ui because it commits here
and closes separately. Related to workspace T-13, which is scoped to the
Portainer stack alone and would leave this copy behind.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-11 21:28:27 +02:00
jpmschweitzerandClaude 7150b4a2fa build(ci): stop gating on typecheck until T-1 clears it
`make typecheck` reports 95 errors in 31 files and has never once passed, so
gating on it did not enforce a standard — it blocked every push to this repo,
including 57fa6c1, the commit that added the gate. Four commits were queued
behind a check that could not be satisfied without a dedicated typing pass.

This is not lowering a bar. The bar was never up: nothing regressed to produce
those errors, they predate the gate, and the same 103 were present before this
session's lint work. The gap is now announced on every push, naming the ticket
that closes it, which is the arrangement core-api, scheduler and library-desk
already use for their ungated stages.

The difference worth preserving: a threshold quietly relaxed hides a problem, and
a declared gap advertises one. This prints five lines about what it is not
checking and why, every time anyone pushes.

typecheck remains a target and still runs on demand. T-1 in this repo's vault
carries the measured breakdown — 29 missing annotations being the bulk — and
removing these lines is that ticket's definition of done.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-11 20:33:58 +02:00
jpmschweitzerandClaude dac259af1d refactor(agents): type the agent and its conversation history
Partial work on the typecheck gate: 103 mypy errors down to 95, and the two
shared roots in agents/tatlock.py are gone. The rest is genuine per-function
annotation work and is not attempted here.

Five conversation lists were declared bare. mypy infers the element type from
the first append, which is a ModelRequest, and then rejects every ModelResponse
that follows — five errors from five lists that all hold the same thing: a
conversation, which is both kinds of message. Annotated as list[ModelMessage],
which is pydantic_ai's own union for exactly this.

The agent had no deps type. It is built as Agent(model, system_prompt=...),
inferred Agent[None, str], while every tool it registers takes
RunContext[ToolCallTracker] and run() is called with a tracker. The declaration
now says what was already happening: Agent[ToolCallTracker, str]. Note this is a
runtime-visible change — pydantic_ai is now told the deps type it was being
handed anyway — so it was verified against the suite rather than reasoned about:
658 passed.

_register_tools carries an assert rather than a None check. It is called from
_ensure_agent immediately after the agent is constructed, so a None there is a
broken invariant, not a case to handle; an `if is None: return` would silently
register no tools.

Two corrections to my own work in this commit. Declaring `_agent: Agent | None`
first made things worse, not better — resolving the bare Agent to Agent[None, str]
surfaced four new argument-type errors that the Any had been hiding, which is
how the missing deps type became visible at all. And an import fix I thought I
had made was a no-op: the target was a multi-line import, my replace matched
nothing, and I had asserted the precondition without asserting the result. Ruff
caught it. That is the same mistake as a changelog edit earlier today, so the
assert now checks what landed.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-11 17:45:48 +02:00
jpmschweitzerandClaude 5b67f5b66c fix: clear the ruff findings that needed a decision
The 21 the automatic pass could not make on its own. `ruff check` and
`ruff format --check` are both clean now; typecheck is still red and is next.

`in_reasoning` in chat/service.py was a complete state machine that nothing read:
initialised False, set True when a reasoning delta arrived, set False when the
summary ended — three assignments, zero reads. Ruff reported one at a time, and
removing each revealed the next, so what looked like a single stray variable took
three passes to bottom out. The branches themselves do real work and are
untouched; only the flag is gone.

Four `raise HTTPException` inside `except` blocks now chain with `from e`. Until
now a failure while handling an error was indistinguishable from the error, which
matters most in exactly the situation where the traceback is all you have.

In biographer/tools.py the binding was unused but the call is not: MemoryType()
is called for the ValueError it raises on an invalid name. The binding is gone
and the call and its comment stay, because dropping the line would have removed
the validation.

The rest are unused bindings in tests where the assertions are on something else
(call_args, mostly), plus three unused loop variables and an isinstance tuple.

One correction to my own work: removing a dead comprehension in
test_error_handling.py left an `if` block with nothing but comments in it, which
is a SyntaxError. Ruff caught it immediately. The block now says what the test
actually pins — that the stream parses without crashing, which reaching that line
demonstrates — rather than computing a list nobody asserts on.

`make test` is intermittent here, and it is not this change.
test_tatlock_tool_call_logging_calculator failed in two of five full runs across
both HEAD and this branch, and passes in the other three; it also fails in
isolation at HEAD while passing in isolation here. Order- or timing-dependent.
Recorded rather than chased, since tests are not gated in this repo yet.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-11 17:37:33 +02:00
jpmschweitzerandClaude 78066fab1b style: apply ruff's automatic fixes and formatter
Mechanical only, and separated from the judgment calls that follow so the
reviewable changes are not buried in a 98-file whitespace diff.

227 automatic fixes: 60 blank lines carrying whitespace, 60 unsorted import
blocks, 34 Optional[X] to X | None, 28 unused imports, 16 deprecated typing
imports, 12 datetime.timezone.utc to datetime.UTC, and assorted smaller
modernisations. Then `ruff format` over src and tests: 98 files reformatted,
35 already conforming.

No file among the unused-import findings defines __all__ or is an __init__.py,
so nothing here removes a re-export.

`make test`: 658 passed, unchanged from HEAD.

Two things observed while verifying, neither addressed here:

`pytest tests/` cannot collect — tests/e2e/test_orchestration_e2e.py uses an
`e2e` marker that is not registered, and the config is strict about markers.
This fails identically at HEAD, so it predates this change; `make test` passes
because it ignores tests/e2e, tests/integration and tests/contracts.

test_tatlock_tool_call_logging_calculator is flaky. It failed once in a full run
with these changes and passed on the next, passes in isolation with them, and
fails in isolation at HEAD. It is order- or timing-dependent, not a regression
from this commit — established by running the full suite both ways rather than
by reasoning about which change could have caused it.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-11 17:25:18 +02:00
jpmschweitzerandClaude 57fa6c13fc build(ci): move the pre-push gate into the Makefile
The hook carried ~50 lines of gitleaks logic and a comment explaining it was
self-contained because "this repo has no Makefile". It has one now, so the
reason is gone and the arrangement is backwards: a hook is a trigger, and
logic belongs where it can be read, run by hand, and changed under review.

.githooks/pre-push is now a byte-identical shim onto `make pre-push` in every
repo in the workspace. The scan itself moves to ci/secrets.sh unchanged, and
`make secrets` runs it on its own.

The call surface is identical everywhere; what it runs is not, and should not
be — each repo gates what it actually has. That is the point of standardising
the name rather than the contents: nobody has to read a repo to find out how
to check it.

secrets runs first, deliberately. It is the only failure here that cannot be
undone by fixing it afterwards — a failed lint costs another commit, a pushed
credential is cached and indexed whether or not it is later deleted.

Some of these gates fail today, on lint debt that predates them, and they are
left wired anyway. The board was measured once and written down in T-56
instead of being worked around here. Narrowing each gate to whatever already
passes would produce a gate that reports success for doing nothing, which is
the failure this workspace keeps rediscovering.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 18:57:22 +02:00
jpmschweitzerandClaude 2fb2fab395 chore(claude): pin PQL_VAULT per project so cwd stops choosing the vault
pql is now a bare word on PATH, which removed the long incantation that had
been forcing --vault into every call by habit. Convenience lowered the cost
of the wrong thing without lowering the cost of the right one: a three-word
pql ticket new targets whichever vault the cwd happens to sit in, and there
are nine of them with colliding id sequences.

PQL_VAULT in each project settings file makes the vault a property of the
session rather than of the working directory — the same lesson Rule 3 records
for git -C, applied to pql. Verified the env var overrides cwd discovery,
that an explicit --vault still beats the env var, and that the harness
hot-reloads it without a restart.

This does not make provenance visible: no output says which vault answered,
so a forgotten --vault still returns a well-formed answer about the wrong
dataset. That remains T-37.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 13:49:02 +02:00
jpmschweitzerandClaude 5bb4013821 chore(claude): deny toj in the sub-repos
toj is now on the global PATH as /usr/local/bin/toj, so its scope boundary
had to stop being "the absolute path is inconvenient to type" and start
being a rule. Its repo and settings verbs operate on the workspace root; run
from inside this repo they answer about the wrong tree.

Both spellings are denied, bare and absolute, because a deny with one
spelling left open is decorative.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 13:42:08 +02:00
jpmschweitzerandClaude 6ab68d3971 ci: gate pushes on a gitleaks scan of the outgoing commits
No repo here scanned for committed credentials. The hook is self-contained
rather than delegating to a Makefile, because this repo has none and a hook
reaching into a sibling repo breaks the moment this one is cloned elsewhere.

Scans the outgoing range rather than full history: history carries settled
findings — test fixtures, vendored third-party code — and a gate that fails
on something unfixable gets bypassed within a week.

Setting core.hooksPath means pql init must replant its replication shims into
.githooks, which is why they are gitignored here alongside the tracked
pre-push. Same layout pql itself uses.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 12:48:57 +02:00
jpmschweitzerandClaude 31d09a6ed1 docs: qualify workspace decision ids cited from this repo
Decision ids are per-vault sequences, so they collide by construction
once there is more than one vault -- and every repo now has one. A bare
D-15 here will mean this repo's D-15 the moment this repo records one.
Cross-vault references are therefore qualified: workspace D-15.

Not hypothetical: pql holds D-1 through D-31 while the workspace holds
D-1 through D-21, so every workspace id currently collides with an
unrelated pql one. A bare id is not wrong the day it is written -- it
decays into wrong as the other vault grows, and nothing flags it.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 04:17:07 +02:00
jpmschweitzerandClaude f8059771ce docs: fold AGENTS.md into CLAUDE.md and record the backend traps
One agent doc per repo, and it is CLAUDE.md. Unlike elsewhere, the
existing CLAUDE.md was not a stub -- it carried seven hard-won gotchas,
all of which survive intact. AGENTS.md supplied the deployment and
release material, minus its feature-branch mandate and its `git add -A`
snippet, and minus its pointer to portainer-core, which is deprecated and
must not be used as a source of infra facts. README.md and
docs/philosophy.md linked to the retired file, so those pointers move
with it.

The new material is two traps that both make the runtime look like the
opposite of what it is.

A cold import inside the container loads src/anthropic but not
src/ollama, and Ollama is the primary backend. The only import of
src/ollama is a function-body one at src/anthropic/model_selector.py:230,
while PREFER_CLOUD_BACKEND=false keeps the Claude path off. Read the
module list naively and the disabled fallback looks live while the hot
path looks dead. This matters because the Claude migration is abandoned
and its remnants are supposed to read as vestigial, not as unfinished
work; the doc carries the decision id so that reasoning is fetchable.

Second, get_household_registry() in a fresh `docker exec python` returns
zero members while the running app serves two models from it. It is
populated at startup, so importing the singleton from outside the app and
reading it as empty is a measurement error, not a finding.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 03:16:29 +02:00
jpmschweitzerandClaude 6b1c892bc6 chore: adopt the workspace agent-config baseline
Commits a .claude/settings.json rather than leaving permissions to
per-developer local state, and initialises a pql vault for this repo's
tickets and internal decisions.

Every git deny rule appears in both the `git <verb>` and `git * <verb>`
forms. Only the second catches `git -C <path>`, and without it the whole
deny list is decorative -- it looks like a policy and stops nothing.

The allow list carries pql's absolute path alongside the bare name.
pql is installed to ~/.local/bin, which is on the login PATH but not the
one a non-interactive shell gets, so the bare-name rules match nothing on
their own and every call would prompt anyway.

.gitignore now covers .claude/settings.local.json, which is machine-local
and must never be shared. `pql init` contributed the .pql/* rules with an
exception for the changelog, which is the replication log of record and
has to be committed for tickets to travel with a clone.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 03:16:12 +02:00
jpmschweitzerandClaude 84467c121a chore: release v2.4.3
Build and Push / release (push) Successful in 2s
Build and Push / build (push) Successful in 1m14s
Ships the Steward capability-extraction fix (a905363), which has been on
main since earlier today while production continued to route on prose:
the running v2.4.2 still matches capability domains as substrings across
the Steward's whole response, so "description" selects housekeeper and
"acknowledge" selects librarian and biographer.

Patch rather than minor: no new capability, and the JSON on the wire is
unchanged. What changes is which agents get invoked, and only in the
cases that were already wrong.

Also carries the routing benchmark, its fixtures, the shared GPU
residency guard and the findings document, none of which are
user-visible.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 18:30:28 +02:00
jpmschweitzerandClaude 2290320e9c docs: record the Steward routing and thinking findings
No change shipped. The Steward stays on gemma4:e2b with thinking left at
its default, and this records why so the experiment is not repeated on
the premise that started it.

That premise was wrong. The Steward appeared to pay ~300 tokens per turn
for reasoning that was generated and discarded, since no `thinking` field
comes back. The reasoning is emitted inline in the response instead, and
it is what produces a correct DELEGATE line — suppressing it costs 12.5
points of routing accuracy, entirely on multi-capability queries where
the model stops decomposing and names one capability.

e4b is disqualified by memory rather than quality: Ollama predicts
10.6 GiB for it against ~7.9 GiB available, so it evicts every
co-resident before loading, including nomic-embed-text. Lowering context
length does not rescue it — an 8x reduction moved the prediction only
1.1 GiB — and per-request num_ctx reloads the shared runner, dropping the
keep_alive pin and evicting nomic.

Also records that the two axes are independent: model choice governs
VRAM and co-residency, think setting governs tokens and latency and
costs nothing in VRAM.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 18:01:11 +02:00
jpmschweitzerandClaude bf13f9f0de refactor(bench): share the GPU residency guard, and guard tool calling too
Extracts the residency snapshot/restore into scripts/ollama_residency.py
so the two benchmarks cannot drift, and applies it to
benchmark_tool_calling.py, which had no protection at all.

That script was the more dangerous of the two. It rewrites
OLLAMA_DEFAULT_MODEL in .env and lets uvicorn reload onto it, restoring
the original only after the loop — so any crash or interrupt left the
*running server* pointed at the benchmark model. Its DEFAULT_MODELS
begins with mistral-nemo-large, the 9.2G model implicated in the
2026-08-07 VRAM outage. Both the .env restore and the residency restore
now run from `finally`.

SIGTERM is handled explicitly in the shared module. Python runs `finally`
for SIGINT, which arrives as KeyboardInterrupt, but the default SIGTERM
action terminates outright, so `timeout` or a plain `kill` skipped the
guard entirely.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 17:03:53 +02:00
jpmschweitzerandClaude 4f42bc047a test(bench): restore GPU residency after a benchmark run
Benchmarking swaps models on the GPU production is serving from. Ollama
evicts to make room, so the first run unpinned gemma4:e2b and left
gemma4:e4b resident: the next voice turn would have paid a ~36s cold
load, and only the monitoring noticing unexpected_models caught it.

Snapshot residency and pinning before the run, then evict whatever the
benchmark loaded and re-pin what was pinned before.

The restore is wired to SIGTERM as well as the normal exit path. Python
runs `finally` for SIGINT, which arrives as KeyboardInterrupt, but the
default SIGTERM action terminates outright — so a `timeout`, a systemd
stop or a plain `kill` skipped the guard entirely. That was not
theoretical: the first SIGTERM after adding this bypassed it, and the
pinned model survived only because the run had not reached the second
model yet.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 16:08:03 +02:00
jpmschweitzerandClaude 738ff10b93 test(bench): add labelled routing fixtures and router benchmark
Measures Steward routing against model and thinking settings by talking
to Ollama directly. No server, no agents, nothing executed — the
mutating fixtures only ever produce a routing decision — so the run is
cheap, repeatable and isolates routing from everything downstream. The
request body mirrors StewardAgent._call_ollama, so the `unset` cell is
exactly what production sends today.

Three thinking settings rather than two. `unset` is production, and it
is not neutral: gemma4 reasons by default and returns no `thinking`
field, so those tokens are generated and discarded.

Scoring is asymmetric on purpose. Each fixture carries `forbid` as well
as `expect`, because over-routing is the predicted failure when thinking
is off and it is the expensive one — a spurious librarian is a real web
call on a query that asked for arithmetic.

The adversarial group is regression coverage for the extraction fix in
a905363: those queries invite the vocabulary that used to select agents
by substring, so they now assert that routing follows what the Steward
decided rather than the words it used while explaining.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 15:22:01 +02:00
jpmschweitzerandClaude a90536314e fix(steward): route on the declared DELEGATE line, not on prose
The prompt tells the Steward to state its choice on a DELEGATE line and
to explain itself on REASON, COMPLEXITY and CONTEXT lines. Extraction
ignored that structure and substring-matched capability domains across
the entire response, so ordinary English in the explanation selected
agents: "description" contains the housekeeper domain "script",
"discover" contains "cover", "acknowledge" contains "knowledge" and
"know", "economy" contains the biographer domain "my".

Every one of those was a real delegation. A spurious librarian is a
multi-second web call on a query that asked for arithmetic.

It also made prose length a routing input, which would have quietly
corrupted the thinking benchmark this was found during: anything that
shortened the Steward's output reduces accidental substring hits and so
reads as improved routing.

Resolution is now layered, most explicit first — a DELEGATE line opening
with a capability name, then a capability named anywhere on that line,
then a domain on that line. With no DELEGATE line at all the response is
matched on capability names only, never domains, so the conversational
path still answers with no capabilities. Matching is whole-word
throughout.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 14:37:59 +02:00
jpmschweitzerandClaude 19e32cfbd6 docs(tests): correct e2e prerequisites in module docstring
Missed in the previous sweep: this docstring still named wakeup.sh, which
the Makefile replaced, and mistral-nemo, which gemma4:e2b replaced.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 15:10:09 +02:00
jpmschweitzerandClaude 99569e786e docs: correct stale tooling and model references
Three migrations left their documentation behind:

wakeup.sh was replaced by the Makefile during the project structure
consolidation, but AGENTS.md and the e2e README still tell you to run it.
The log path moved to build/logs/server.log at the same time.

The local model moved to gemma4:e2b, but the e2e prerequisites and the
benchmark recommendation still name mistral-nemo.

The benchmark figures in CLAUDE.md predate the current model. Measured
2026-08-07: ~95 tok/s, full flow ~10-13s for simple turns, cold model load
~36s rather than ~8s. A turn costs three sequential Ollama calls and ~710
generated tokens regardless of how trivial the question is.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 15:07:10 +02:00
jpmschweitzerandClaude Fable 5 2cf3252a19 docs(claude-integration): registry is git.schweitz.net not git.schweitz.internal
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 17:11:18 +02:00
jpmschweitzerandClaude Fable 5 cdd5a55613 chore: release v2.4.2
Build and Push / release (push) Successful in 3s
Build and Push / build (push) Successful in 1m3s
Fixes the v2.4.1 crash-loop: fresh image builds resolved
opentelemetry-api 1.44.0, which removed the private _events module
that pydantic-ai 1.27 imports at startup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 13:10:57 +02:00
135 changed files with 3825 additions and 1965 deletions
+75
View File
@@ -0,0 +1,75 @@
{
"env": {
"PQL_VAULT": "/mnt/media/Projects/tatlock"
},
"permissions": {
"allow": [
"Bash(pql)",
"Bash(pql *)",
"Bash(/home/jpmschweitzer/.local/bin/pql:*)",
"Bash(git status:*)",
"Bash(git log:*)",
"Bash(git diff:*)",
"Bash(git branch:*)",
"Bash(make test:*)",
"Bash(make test-unit:*)",
"Bash(make test-contracts:*)",
"Bash(make lint:*)",
"Bash(make typecheck:*)",
"Bash(.venv/bin/python -m pytest:*)",
"Bash(.venv/bin/pytest:*)",
"Bash(pytest:*)",
"Bash(ruff check:*)",
"Bash(mypy:*)",
"Bash(docker logs tatlock:*)",
"Bash(curl -s http://localhost:8000/*)",
"Bash(curl -s http://localhost:8777/*)"
],
"deny": [
"Bash(/mnt/media/Projects/cladmin/ops/bin/toj)",
"Bash(/mnt/media/Projects/cladmin/ops/bin/toj:*)",
"Bash(chmod -R 777 *)",
"Bash(chmod 777 *)",
"Bash(dd if=*)",
"Bash(find * -delete*)",
"Bash(find * -exec*)",
"Bash(git * add --all*)",
"Bash(git * add -A*)",
"Bash(git * add .)",
"Bash(git * branch -D *)",
"Bash(git * checkout -- *)",
"Bash(git * clean -fd*)",
"Bash(git * clean -fdx*)",
"Bash(git * commit --no-verify*)",
"Bash(git * merge --no-ff*)",
"Bash(git * push --force*)",
"Bash(git * push -f*)",
"Bash(git * reset --hard*)",
"Bash(git * restore .*)",
"Bash(git add --all*)",
"Bash(git add -A*)",
"Bash(git add .)",
"Bash(git branch -D *)",
"Bash(git checkout -- *)",
"Bash(git clean -fd*)",
"Bash(git clean -fdx*)",
"Bash(git commit --no-verify*)",
"Bash(git merge --no-ff*)",
"Bash(git push --force*)",
"Bash(git push -f*)",
"Bash(git reset --hard*)",
"Bash(git restore .*)",
"Bash(mkfs*)",
"Bash(ollama rm *)",
"Bash(redis-cli * FLUSHALL*)",
"Bash(redis-cli * FLUSHDB*)",
"Bash(rm -rf $HOME)",
"Bash(rm -rf /)",
"Bash(rm -rf ~)",
"Bash(su *)",
"Bash(sudo *)",
"Bash(toj)",
"Bash(toj:*)"
]
}
}
+1
View File
@@ -0,0 +1 @@
.pql/changelog/*.sql merge=union
+13
View File
@@ -0,0 +1,13 @@
#!/usr/bin/env bash
# Trigger only. The checks live in the Makefile, where they can be read, run by
# hand (`make pre-push`), and changed under review.
#
# This file is identical in every repo in this workspace, deliberately: the call
# surface is the same everywhere even though what each gate runs is not, so
# nobody has to read a repo to find out how to check it (D-27).
#
# Enable per clone with: git config core.hooksPath .githooks
# Never bypass with --no-verify. Suppress a specific finding deliberately
# instead, with a reason — see `make pre-push`.
set -euo pipefail
exec make -C "$(git rev-parse --show-toplevel)" pre-push
+13
View File
@@ -103,3 +103,16 @@ ollama_data/
ehthumbs.db ehthumbs.db
Thumbs.db Thumbs.db
Desktop.ini Desktop.ini
# Claude Code user-specific settings
.claude/settings.local.json
.pql/*
!.pql/changelog/
# pql shims planted by `pql init` into the dir core.hooksPath points at.
# Per-clone: each embeds the absolute path of the pql binary that planted it.
# Only .githooks/pre-push is shared.
.githooks/pre-commit
.githooks/post-merge
.githooks/post-checkout
.githooks/post-rewrite
+11
View File
@@ -0,0 +1,11 @@
-- Changelog format marker, written by pql. Comments only: this file
-- is never executed — Import descends into the per-table directories
-- and does not read the changelog root.
--
-- A changelog carrying no marker is format 1, the shape that existed
-- before formats were versioned. An older format is migrated forward
-- by `pql plan upgrade` (and automatically from the post-merge hook);
-- a newer one is refused rather than replayed under rules this binary
-- does not know. See D-28 and docs/versions.md.
-- pql:changelog_format: 2.0.0
-- pql:written_by: 2.2.0
+139
View File
@@ -0,0 +1,139 @@
-- Auto-generated by pql init. CREATE TABLE statements
-- for the planning schema; per-table dir keeps the changelog
-- self-describing per D-15. CREATE TABLE IF NOT EXISTS is
-- idempotent so running schema files from each directory in
-- replay order is harmless.
--
-- Importer parses the markers below to detect schema drift
-- between the producing pql version and the local one — a
-- bumped canonical_version means projection rules changed
-- and replay must refuse rather than silently corrupt state.
-- pql:created_by: 2.2.0
-- pql:canonical_version: 2
CREATE TABLE IF NOT EXISTS decisions (
id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('confirmed','question','rejected')),
domain TEXT NOT NULL,
title TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'active'
CHECK(status IN ('active','superseded','resolved','open')),
date TEXT,
file_path TEXT NOT NULL,
synced_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS decision_refs (
source_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
target_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
ref_type TEXT NOT NULL
CHECK(ref_type IN ('supersedes','references','resolves','depends_on','amends')),
note TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (source_id, target_id, ref_type)
);
-- Identity split (D-26): a ticket's stable, collision-proof identity is its
-- record_id (a locally-generated ULID, planning.NewRecordID); the friendly
-- T-NNN label lives in ticket_idmap and may be reconciled. Every structural
-- reference (parent, deps, history, labels) targets record_id, so a label
-- clash never corrupts the graph — only ticket_idmap needs a relabel.
CREATE TABLE IF NOT EXISTS tickets (
record_id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('initiative','epic','story','task','bug')),
parent_record_id TEXT REFERENCES tickets(record_id),
title TEXT NOT NULL,
description TEXT,
-- No CHECK enumeration: the ticket status vocabulary is per-vault
-- configurable (ticket_statuses in .pql/config.yaml). Validation lives
-- in Go (planning.StatusSet), so adding/renaming statuses needs no
-- schema change. The DEFAULT is a harmless fallback — CreateTicket
-- always inserts the configured default explicitly.
status TEXT NOT NULL DEFAULT 'backlog',
priority TEXT DEFAULT 'medium'
CHECK(priority IN ('critical','high','medium','low')),
assigned_to TEXT,
team TEXT,
decision_ref TEXT REFERENCES decisions(id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
-- ticket_idmap maps a record_id to its current friendly label (T-NNN).
-- ticket_id is intentionally NOT globally unique: two uncoordinated clones
-- can mint the same label, which surfaces as a duplicate-label collision
-- (detected at replay) and is fixed with "pql ticket relabel".
CREATE TABLE IF NOT EXISTS ticket_idmap (
record_id TEXT PRIMARY KEY REFERENCES tickets(record_id),
ticket_id TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_deps (
blocker_record_id TEXT NOT NULL REFERENCES tickets(record_id),
blocked_record_id TEXT NOT NULL REFERENCES tickets(record_id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (blocker_record_id, blocked_record_id)
);
CREATE TABLE IF NOT EXISTS ticket_history (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
field TEXT NOT NULL,
old_value TEXT,
new_value TEXT,
changed_by TEXT,
changed_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT UNIQUE,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_labels (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
label TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (ticket_record_id, label)
);
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_tickets_status ON tickets(status);
CREATE INDEX IF NOT EXISTS idx_tickets_team ON tickets(team);
CREATE INDEX IF NOT EXISTS idx_tickets_decision_ref ON tickets(decision_ref);
CREATE INDEX IF NOT EXISTS idx_tickets_assigned ON tickets(assigned_to);
CREATE INDEX IF NOT EXISTS idx_tickets_parent ON tickets(parent_record_id);
CREATE INDEX IF NOT EXISTS idx_ticket_idmap_label ON ticket_idmap(ticket_id);
CREATE INDEX IF NOT EXISTS idx_decisions_domain ON decisions(domain);
CREATE INDEX IF NOT EXISTS idx_decisions_type ON decisions(type);
CREATE INDEX IF NOT EXISTS idx_decision_refs_target ON decision_refs(target_id);
@@ -0,0 +1,139 @@
-- Auto-generated by pql init. CREATE TABLE statements
-- for the planning schema; per-table dir keeps the changelog
-- self-describing per D-15. CREATE TABLE IF NOT EXISTS is
-- idempotent so running schema files from each directory in
-- replay order is harmless.
--
-- Importer parses the markers below to detect schema drift
-- between the producing pql version and the local one — a
-- bumped canonical_version means projection rules changed
-- and replay must refuse rather than silently corrupt state.
-- pql:created_by: 2.2.0
-- pql:canonical_version: 2
CREATE TABLE IF NOT EXISTS decisions (
id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('confirmed','question','rejected')),
domain TEXT NOT NULL,
title TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'active'
CHECK(status IN ('active','superseded','resolved','open')),
date TEXT,
file_path TEXT NOT NULL,
synced_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS decision_refs (
source_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
target_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
ref_type TEXT NOT NULL
CHECK(ref_type IN ('supersedes','references','resolves','depends_on','amends')),
note TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (source_id, target_id, ref_type)
);
-- Identity split (D-26): a ticket's stable, collision-proof identity is its
-- record_id (a locally-generated ULID, planning.NewRecordID); the friendly
-- T-NNN label lives in ticket_idmap and may be reconciled. Every structural
-- reference (parent, deps, history, labels) targets record_id, so a label
-- clash never corrupts the graph — only ticket_idmap needs a relabel.
CREATE TABLE IF NOT EXISTS tickets (
record_id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('initiative','epic','story','task','bug')),
parent_record_id TEXT REFERENCES tickets(record_id),
title TEXT NOT NULL,
description TEXT,
-- No CHECK enumeration: the ticket status vocabulary is per-vault
-- configurable (ticket_statuses in .pql/config.yaml). Validation lives
-- in Go (planning.StatusSet), so adding/renaming statuses needs no
-- schema change. The DEFAULT is a harmless fallback — CreateTicket
-- always inserts the configured default explicitly.
status TEXT NOT NULL DEFAULT 'backlog',
priority TEXT DEFAULT 'medium'
CHECK(priority IN ('critical','high','medium','low')),
assigned_to TEXT,
team TEXT,
decision_ref TEXT REFERENCES decisions(id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
-- ticket_idmap maps a record_id to its current friendly label (T-NNN).
-- ticket_id is intentionally NOT globally unique: two uncoordinated clones
-- can mint the same label, which surfaces as a duplicate-label collision
-- (detected at replay) and is fixed with "pql ticket relabel".
CREATE TABLE IF NOT EXISTS ticket_idmap (
record_id TEXT PRIMARY KEY REFERENCES tickets(record_id),
ticket_id TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_deps (
blocker_record_id TEXT NOT NULL REFERENCES tickets(record_id),
blocked_record_id TEXT NOT NULL REFERENCES tickets(record_id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (blocker_record_id, blocked_record_id)
);
CREATE TABLE IF NOT EXISTS ticket_history (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
field TEXT NOT NULL,
old_value TEXT,
new_value TEXT,
changed_by TEXT,
changed_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT UNIQUE,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_labels (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
label TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (ticket_record_id, label)
);
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_tickets_status ON tickets(status);
CREATE INDEX IF NOT EXISTS idx_tickets_team ON tickets(team);
CREATE INDEX IF NOT EXISTS idx_tickets_decision_ref ON tickets(decision_ref);
CREATE INDEX IF NOT EXISTS idx_tickets_assigned ON tickets(assigned_to);
CREATE INDEX IF NOT EXISTS idx_tickets_parent ON tickets(parent_record_id);
CREATE INDEX IF NOT EXISTS idx_ticket_idmap_label ON ticket_idmap(ticket_id);
CREATE INDEX IF NOT EXISTS idx_decisions_domain ON decisions(domain);
CREATE INDEX IF NOT EXISTS idx_decisions_type ON decisions(type);
CREATE INDEX IF NOT EXISTS idx_decision_refs_target ON decision_refs(target_id);
+34
View File
@@ -0,0 +1,34 @@
INSERT INTO ticket_history (ticket_record_id, field, old_value, new_value, changed_by, changed_at, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ44Z6HSN0RQS0QAYNTPEM5G', 'description', NULL, '`make typecheck` reports 95 errors in 31 files (was 103). This is a dedicated programming pass, not lint tidying, and it is what currently blocks `make pre-push`.
MEASURED 2026-08-11 so the next session does not re-derive it:
29 no-untyped-def functions with no annotations — the bulk, and genuine per-function work
18 no-any-return mostly downstream of the above
12 assignment
11 arg-type
5 unused-ignore `# type: ignore` comments mypy says are no longer needed
5 union-attr
4 var-annotated
3 override
remainder: misc, dict-item, attr-defined, return-value, call-overload
By file: agents/tatlock.py 15, core/memory_service.py 11, responses/streaming.py 8, responses/service.py 7, core/context.py 6.
THE TWO SHARED ROOTS ARE ALREADY FIXED (dac259a), so what is left has no lever in it. For reference, they were: five conversation lists declared bare, where mypy infers the element type from the first append (a ModelRequest) and then rejects every ModelResponse; and an agent built as Agent(model, system_prompt=...) with no deps_type, inferred Agent[None, str], while every tool it registers takes RunContext[ToolCallTracker].
WORTH KNOWING BEFORE STARTING. Annotating partially made mypy count go UP before it went down — declaring `_agent: Agent | None` took agents/tatlock.py from 22 to 24, because resolving the bare Agent to Agent[None, str] surfaced four argument-type errors the Any had been hiding. Expect that shape: a rising count during this work usually means concealment ending, not damage.
The mypy config is strict — disallow_untyped_defs, disallow_incomplete_defs, warn_return_any, check_untyped_defs, strict_equality — so there is no partial-credit setting to lean on, and weakening it would be the wrong trade for a codebase this central.
TWO PRE-EXISTING TEST FACTS, both confirmed at HEAD and neither caused by the lint work:
- test_tatlock_tool_call_logging_calculator is flaky: failed 2 of 5 full runs, on HEAD and on the lint branch, and fails in isolation at HEAD while passing in isolation after the lint pass. Order- or timing-dependent.
- `pytest tests/` cannot collect at all: tests/e2e/test_orchestration_e2e.py uses an `e2e` marker that is not registered and the config is strict about markers. `make test` passes only because it ignores tests/e2e, tests/integration and tests/contracts.', NULL, '2026-08-11 18:26:58', '2026-08-11 18:26:58.523', '2026-08-11 18:26:58.523', NULL, 'c8c9b6a20cf18ca903fc8c720f70e73d', 2) ON CONFLICT(hash) DO NOTHING;
INSERT INTO ticket_history (ticket_record_id, field, old_value, new_value, changed_by, changed_at, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ4JX9DZB3XAB95EP23YWXMC', 'description', NULL, '`.env` still carries an `ANTHROPIC_API_KEY`. That key was revoked and no longer exists on Anthropic''s side, so this is dead weight rather than an exposure — but it is dead weight that reads exactly like a live credential to anyone who finds it.
The cost is confusion, not risk. Someone debugging a Claude fallback will find a key present, assume it is configured, and look elsewhere for the failure. The absence of a key is a clear signal; a revoked key is a misleading one.
`.env` is gitignored here and has never been committed, so nothing needs rewriting — the value simply needs removing from the local file, and the line dropping or blanking in `.env.example` if it appears there too.
RELATED, and the reason this is filed separately: the workspace vault holds T-13, "Clear the revoked ANTHROPIC_API_KEY from the live Portainer stack". That ticket is scoped to the Portainer stack only. Whoever closes it will reasonably believe the key is gone once the stack is clean, and this copy will survive. The two want doing together even though they commit separately.
Context on the revocation, since it explains why nobody removed this at the time: the key was revoked on 2026-08-09 after being printed into a transcript by a redaction filter that matched on `KEY` appearing after the `=`. In `ANTHROPIC_API_KEY=...` it appears before, so the filter never fired. The response was rotation, and the leftover copies were not swept.', NULL, '2026-08-11 19:27:52', '2026-08-11 19:27:52.823', '2026-08-11 19:27:52.823', NULL, 'c7834e46268029b74c25a25a83177b64', 2) ON CONFLICT(hash) DO NOTHING;
+139
View File
@@ -0,0 +1,139 @@
-- Auto-generated by pql init. CREATE TABLE statements
-- for the planning schema; per-table dir keeps the changelog
-- self-describing per D-15. CREATE TABLE IF NOT EXISTS is
-- idempotent so running schema files from each directory in
-- replay order is harmless.
--
-- Importer parses the markers below to detect schema drift
-- between the producing pql version and the local one — a
-- bumped canonical_version means projection rules changed
-- and replay must refuse rather than silently corrupt state.
-- pql:created_by: 2.2.0
-- pql:canonical_version: 2
CREATE TABLE IF NOT EXISTS decisions (
id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('confirmed','question','rejected')),
domain TEXT NOT NULL,
title TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'active'
CHECK(status IN ('active','superseded','resolved','open')),
date TEXT,
file_path TEXT NOT NULL,
synced_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS decision_refs (
source_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
target_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
ref_type TEXT NOT NULL
CHECK(ref_type IN ('supersedes','references','resolves','depends_on','amends')),
note TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (source_id, target_id, ref_type)
);
-- Identity split (D-26): a ticket's stable, collision-proof identity is its
-- record_id (a locally-generated ULID, planning.NewRecordID); the friendly
-- T-NNN label lives in ticket_idmap and may be reconciled. Every structural
-- reference (parent, deps, history, labels) targets record_id, so a label
-- clash never corrupts the graph — only ticket_idmap needs a relabel.
CREATE TABLE IF NOT EXISTS tickets (
record_id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('initiative','epic','story','task','bug')),
parent_record_id TEXT REFERENCES tickets(record_id),
title TEXT NOT NULL,
description TEXT,
-- No CHECK enumeration: the ticket status vocabulary is per-vault
-- configurable (ticket_statuses in .pql/config.yaml). Validation lives
-- in Go (planning.StatusSet), so adding/renaming statuses needs no
-- schema change. The DEFAULT is a harmless fallback — CreateTicket
-- always inserts the configured default explicitly.
status TEXT NOT NULL DEFAULT 'backlog',
priority TEXT DEFAULT 'medium'
CHECK(priority IN ('critical','high','medium','low')),
assigned_to TEXT,
team TEXT,
decision_ref TEXT REFERENCES decisions(id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
-- ticket_idmap maps a record_id to its current friendly label (T-NNN).
-- ticket_id is intentionally NOT globally unique: two uncoordinated clones
-- can mint the same label, which surfaces as a duplicate-label collision
-- (detected at replay) and is fixed with "pql ticket relabel".
CREATE TABLE IF NOT EXISTS ticket_idmap (
record_id TEXT PRIMARY KEY REFERENCES tickets(record_id),
ticket_id TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_deps (
blocker_record_id TEXT NOT NULL REFERENCES tickets(record_id),
blocked_record_id TEXT NOT NULL REFERENCES tickets(record_id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (blocker_record_id, blocked_record_id)
);
CREATE TABLE IF NOT EXISTS ticket_history (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
field TEXT NOT NULL,
old_value TEXT,
new_value TEXT,
changed_by TEXT,
changed_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT UNIQUE,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_labels (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
label TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (ticket_record_id, label)
);
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_tickets_status ON tickets(status);
CREATE INDEX IF NOT EXISTS idx_tickets_team ON tickets(team);
CREATE INDEX IF NOT EXISTS idx_tickets_decision_ref ON tickets(decision_ref);
CREATE INDEX IF NOT EXISTS idx_tickets_assigned ON tickets(assigned_to);
CREATE INDEX IF NOT EXISTS idx_tickets_parent ON tickets(parent_record_id);
CREATE INDEX IF NOT EXISTS idx_ticket_idmap_label ON ticket_idmap(ticket_id);
CREATE INDEX IF NOT EXISTS idx_decisions_domain ON decisions(domain);
CREATE INDEX IF NOT EXISTS idx_decisions_type ON decisions(type);
CREATE INDEX IF NOT EXISTS idx_decision_refs_target ON decision_refs(target_id);
+2
View File
@@ -0,0 +1,2 @@
INSERT INTO ticket_idmap (record_id, ticket_id, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ44Z6HSN0RQS0QAYNTPEM5G', 'T-1', '2026-08-11 18:26:58.365', '2026-08-11 18:26:58.365', NULL, 'f1508986553f1ee59145a0d099131a68', 2) ON CONFLICT(record_id) DO UPDATE SET ticket_id=excluded.ticket_id, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= ticket_idmap.updated_at;
INSERT INTO ticket_idmap (record_id, ticket_id, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ4JX9DZB3XAB95EP23YWXMC', 'T-2', '2026-08-11 19:27:52.688', '2026-08-11 19:27:52.688', NULL, 'a5b18acb4b9a7e2da8e47b6ee603a5da', 2) ON CONFLICT(record_id) DO UPDATE SET ticket_id=excluded.ticket_id, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= ticket_idmap.updated_at;
@@ -0,0 +1,139 @@
-- Auto-generated by pql init. CREATE TABLE statements
-- for the planning schema; per-table dir keeps the changelog
-- self-describing per D-15. CREATE TABLE IF NOT EXISTS is
-- idempotent so running schema files from each directory in
-- replay order is harmless.
--
-- Importer parses the markers below to detect schema drift
-- between the producing pql version and the local one — a
-- bumped canonical_version means projection rules changed
-- and replay must refuse rather than silently corrupt state.
-- pql:created_by: 2.2.0
-- pql:canonical_version: 2
CREATE TABLE IF NOT EXISTS decisions (
id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('confirmed','question','rejected')),
domain TEXT NOT NULL,
title TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'active'
CHECK(status IN ('active','superseded','resolved','open')),
date TEXT,
file_path TEXT NOT NULL,
synced_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS decision_refs (
source_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
target_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
ref_type TEXT NOT NULL
CHECK(ref_type IN ('supersedes','references','resolves','depends_on','amends')),
note TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (source_id, target_id, ref_type)
);
-- Identity split (D-26): a ticket's stable, collision-proof identity is its
-- record_id (a locally-generated ULID, planning.NewRecordID); the friendly
-- T-NNN label lives in ticket_idmap and may be reconciled. Every structural
-- reference (parent, deps, history, labels) targets record_id, so a label
-- clash never corrupts the graph — only ticket_idmap needs a relabel.
CREATE TABLE IF NOT EXISTS tickets (
record_id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('initiative','epic','story','task','bug')),
parent_record_id TEXT REFERENCES tickets(record_id),
title TEXT NOT NULL,
description TEXT,
-- No CHECK enumeration: the ticket status vocabulary is per-vault
-- configurable (ticket_statuses in .pql/config.yaml). Validation lives
-- in Go (planning.StatusSet), so adding/renaming statuses needs no
-- schema change. The DEFAULT is a harmless fallback — CreateTicket
-- always inserts the configured default explicitly.
status TEXT NOT NULL DEFAULT 'backlog',
priority TEXT DEFAULT 'medium'
CHECK(priority IN ('critical','high','medium','low')),
assigned_to TEXT,
team TEXT,
decision_ref TEXT REFERENCES decisions(id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
-- ticket_idmap maps a record_id to its current friendly label (T-NNN).
-- ticket_id is intentionally NOT globally unique: two uncoordinated clones
-- can mint the same label, which surfaces as a duplicate-label collision
-- (detected at replay) and is fixed with "pql ticket relabel".
CREATE TABLE IF NOT EXISTS ticket_idmap (
record_id TEXT PRIMARY KEY REFERENCES tickets(record_id),
ticket_id TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_deps (
blocker_record_id TEXT NOT NULL REFERENCES tickets(record_id),
blocked_record_id TEXT NOT NULL REFERENCES tickets(record_id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (blocker_record_id, blocked_record_id)
);
CREATE TABLE IF NOT EXISTS ticket_history (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
field TEXT NOT NULL,
old_value TEXT,
new_value TEXT,
changed_by TEXT,
changed_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT UNIQUE,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_labels (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
label TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (ticket_record_id, label)
);
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_tickets_status ON tickets(status);
CREATE INDEX IF NOT EXISTS idx_tickets_team ON tickets(team);
CREATE INDEX IF NOT EXISTS idx_tickets_decision_ref ON tickets(decision_ref);
CREATE INDEX IF NOT EXISTS idx_tickets_assigned ON tickets(assigned_to);
CREATE INDEX IF NOT EXISTS idx_tickets_parent ON tickets(parent_record_id);
CREATE INDEX IF NOT EXISTS idx_ticket_idmap_label ON ticket_idmap(ticket_id);
CREATE INDEX IF NOT EXISTS idx_decisions_domain ON decisions(domain);
CREATE INDEX IF NOT EXISTS idx_decisions_type ON decisions(type);
CREATE INDEX IF NOT EXISTS idx_decision_refs_target ON decision_refs(target_id);
+139
View File
@@ -0,0 +1,139 @@
-- Auto-generated by pql init. CREATE TABLE statements
-- for the planning schema; per-table dir keeps the changelog
-- self-describing per D-15. CREATE TABLE IF NOT EXISTS is
-- idempotent so running schema files from each directory in
-- replay order is harmless.
--
-- Importer parses the markers below to detect schema drift
-- between the producing pql version and the local one — a
-- bumped canonical_version means projection rules changed
-- and replay must refuse rather than silently corrupt state.
-- pql:created_by: 2.2.0
-- pql:canonical_version: 2
CREATE TABLE IF NOT EXISTS decisions (
id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('confirmed','question','rejected')),
domain TEXT NOT NULL,
title TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'active'
CHECK(status IN ('active','superseded','resolved','open')),
date TEXT,
file_path TEXT NOT NULL,
synced_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS decision_refs (
source_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
target_id TEXT NOT NULL REFERENCES decisions(id) ON DELETE CASCADE,
ref_type TEXT NOT NULL
CHECK(ref_type IN ('supersedes','references','resolves','depends_on','amends')),
note TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (source_id, target_id, ref_type)
);
-- Identity split (D-26): a ticket's stable, collision-proof identity is its
-- record_id (a locally-generated ULID, planning.NewRecordID); the friendly
-- T-NNN label lives in ticket_idmap and may be reconciled. Every structural
-- reference (parent, deps, history, labels) targets record_id, so a label
-- clash never corrupts the graph — only ticket_idmap needs a relabel.
CREATE TABLE IF NOT EXISTS tickets (
record_id TEXT PRIMARY KEY,
type TEXT NOT NULL CHECK(type IN ('initiative','epic','story','task','bug')),
parent_record_id TEXT REFERENCES tickets(record_id),
title TEXT NOT NULL,
description TEXT,
-- No CHECK enumeration: the ticket status vocabulary is per-vault
-- configurable (ticket_statuses in .pql/config.yaml). Validation lives
-- in Go (planning.StatusSet), so adding/renaming statuses needs no
-- schema change. The DEFAULT is a harmless fallback — CreateTicket
-- always inserts the configured default explicitly.
status TEXT NOT NULL DEFAULT 'backlog',
priority TEXT DEFAULT 'medium'
CHECK(priority IN ('critical','high','medium','low')),
assigned_to TEXT,
team TEXT,
decision_ref TEXT REFERENCES decisions(id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
-- ticket_idmap maps a record_id to its current friendly label (T-NNN).
-- ticket_id is intentionally NOT globally unique: two uncoordinated clones
-- can mint the same label, which surfaces as a duplicate-label collision
-- (detected at replay) and is fixed with "pql ticket relabel".
CREATE TABLE IF NOT EXISTS ticket_idmap (
record_id TEXT PRIMARY KEY REFERENCES tickets(record_id),
ticket_id TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_deps (
blocker_record_id TEXT NOT NULL REFERENCES tickets(record_id),
blocked_record_id TEXT NOT NULL REFERENCES tickets(record_id),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (blocker_record_id, blocked_record_id)
);
CREATE TABLE IF NOT EXISTS ticket_history (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
field TEXT NOT NULL,
old_value TEXT,
new_value TEXT,
changed_by TEXT,
changed_at TEXT NOT NULL DEFAULT (datetime('now')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT UNIQUE,
canonical_version INTEGER
);
CREATE TABLE IF NOT EXISTS ticket_labels (
ticket_record_id TEXT NOT NULL REFERENCES tickets(record_id),
label TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
deleted_at TEXT,
hash TEXT,
canonical_version INTEGER,
PRIMARY KEY (ticket_record_id, label)
);
CREATE TABLE IF NOT EXISTS meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_tickets_status ON tickets(status);
CREATE INDEX IF NOT EXISTS idx_tickets_team ON tickets(team);
CREATE INDEX IF NOT EXISTS idx_tickets_decision_ref ON tickets(decision_ref);
CREATE INDEX IF NOT EXISTS idx_tickets_assigned ON tickets(assigned_to);
CREATE INDEX IF NOT EXISTS idx_tickets_parent ON tickets(parent_record_id);
CREATE INDEX IF NOT EXISTS idx_ticket_idmap_label ON ticket_idmap(ticket_id);
CREATE INDEX IF NOT EXISTS idx_decisions_domain ON decisions(domain);
CREATE INDEX IF NOT EXISTS idx_decisions_type ON decisions(type);
CREATE INDEX IF NOT EXISTS idx_decision_refs_target ON decision_refs(target_id);
+36
View File
@@ -0,0 +1,36 @@
INSERT INTO tickets (record_id, type, parent_record_id, title, description, status, priority, assigned_to, team, decision_ref, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ44Z6HSN0RQS0QAYNTPEM5G', 'task', NULL, 'Type the codebase: 95 mypy errors across 31 files', NULL, 'backlog', 'high', NULL, NULL, NULL, '2026-08-11 18:26:58.318', '2026-08-11 18:26:58.318', NULL, 'd881f36d77e58aaa2f94dc08c5b5be3e', 2) ON CONFLICT(record_id) DO UPDATE SET type=excluded.type, parent_record_id=excluded.parent_record_id, title=excluded.title, description=excluded.description, status=excluded.status, priority=excluded.priority, assigned_to=excluded.assigned_to, team=excluded.team, decision_ref=excluded.decision_ref, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= tickets.updated_at;
INSERT INTO tickets (record_id, type, parent_record_id, title, description, status, priority, assigned_to, team, decision_ref, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ44Z6HSN0RQS0QAYNTPEM5G', 'task', NULL, 'Type the codebase: 95 mypy errors across 31 files', '`make typecheck` reports 95 errors in 31 files (was 103). This is a dedicated programming pass, not lint tidying, and it is what currently blocks `make pre-push`.
MEASURED 2026-08-11 so the next session does not re-derive it:
29 no-untyped-def functions with no annotations — the bulk, and genuine per-function work
18 no-any-return mostly downstream of the above
12 assignment
11 arg-type
5 unused-ignore `# type: ignore` comments mypy says are no longer needed
5 union-attr
4 var-annotated
3 override
remainder: misc, dict-item, attr-defined, return-value, call-overload
By file: agents/tatlock.py 15, core/memory_service.py 11, responses/streaming.py 8, responses/service.py 7, core/context.py 6.
THE TWO SHARED ROOTS ARE ALREADY FIXED (dac259a), so what is left has no lever in it. For reference, they were: five conversation lists declared bare, where mypy infers the element type from the first append (a ModelRequest) and then rejects every ModelResponse; and an agent built as Agent(model, system_prompt=...) with no deps_type, inferred Agent[None, str], while every tool it registers takes RunContext[ToolCallTracker].
WORTH KNOWING BEFORE STARTING. Annotating partially made mypy count go UP before it went down — declaring `_agent: Agent | None` took agents/tatlock.py from 22 to 24, because resolving the bare Agent to Agent[None, str] surfaced four argument-type errors the Any had been hiding. Expect that shape: a rising count during this work usually means concealment ending, not damage.
The mypy config is strict — disallow_untyped_defs, disallow_incomplete_defs, warn_return_any, check_untyped_defs, strict_equality — so there is no partial-credit setting to lean on, and weakening it would be the wrong trade for a codebase this central.
TWO PRE-EXISTING TEST FACTS, both confirmed at HEAD and neither caused by the lint work:
- test_tatlock_tool_call_logging_calculator is flaky: failed 2 of 5 full runs, on HEAD and on the lint branch, and fails in isolation at HEAD while passing in isolation after the lint pass. Order- or timing-dependent.
- `pytest tests/` cannot collect at all: tests/e2e/test_orchestration_e2e.py uses an `e2e` marker that is not registered and the config is strict about markers. `make test` passes only because it ignores tests/e2e, tests/integration and tests/contracts.', 'backlog', 'high', NULL, NULL, NULL, '2026-08-11 18:26:58.318', '2026-08-11 18:26:58.523', NULL, 'fc047dd080d976c4ca0c551cac8fec93', 2) ON CONFLICT(record_id) DO UPDATE SET type=excluded.type, parent_record_id=excluded.parent_record_id, title=excluded.title, description=excluded.description, status=excluded.status, priority=excluded.priority, assigned_to=excluded.assigned_to, team=excluded.team, decision_ref=excluded.decision_ref, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= tickets.updated_at;
INSERT INTO tickets (record_id, type, parent_record_id, title, description, status, priority, assigned_to, team, decision_ref, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ4JX9DZB3XAB95EP23YWXMC', 'bug', NULL, 'A revoked ANTHROPIC_API_KEY is still sitting in .env', NULL, 'backlog', 'medium', NULL, NULL, NULL, '2026-08-11 19:27:52.687', '2026-08-11 19:27:52.687', NULL, '3b00a4b729819e20a6425f971e6cb9da', 2) ON CONFLICT(record_id) DO UPDATE SET type=excluded.type, parent_record_id=excluded.parent_record_id, title=excluded.title, description=excluded.description, status=excluded.status, priority=excluded.priority, assigned_to=excluded.assigned_to, team=excluded.team, decision_ref=excluded.decision_ref, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= tickets.updated_at;
INSERT INTO tickets (record_id, type, parent_record_id, title, description, status, priority, assigned_to, team, decision_ref, created_at, updated_at, deleted_at, hash, canonical_version) VALUES ('06FZ4JX9DZB3XAB95EP23YWXMC', 'bug', NULL, 'A revoked ANTHROPIC_API_KEY is still sitting in .env', '`.env` still carries an `ANTHROPIC_API_KEY`. That key was revoked and no longer exists on Anthropic''s side, so this is dead weight rather than an exposure — but it is dead weight that reads exactly like a live credential to anyone who finds it.
The cost is confusion, not risk. Someone debugging a Claude fallback will find a key present, assume it is configured, and look elsewhere for the failure. The absence of a key is a clear signal; a revoked key is a misleading one.
`.env` is gitignored here and has never been committed, so nothing needs rewriting — the value simply needs removing from the local file, and the line dropping or blanking in `.env.example` if it appears there too.
RELATED, and the reason this is filed separately: the workspace vault holds T-13, "Clear the revoked ANTHROPIC_API_KEY from the live Portainer stack". That ticket is scoped to the Portainer stack only. Whoever closes it will reasonably believe the key is gone once the stack is clean, and this copy will survive. The two want doing together even though they commit separately.
Context on the revocation, since it explains why nobody removed this at the time: the key was revoked on 2026-08-09 after being printed into a transcript by a redaction filter that matched on `KEY` appearing after the `=`. In `ANTHROPIC_API_KEY=...` it appears before, so the filter never fired. The response was rotation, and the leftover copies were not swept.', 'backlog', 'medium', NULL, NULL, NULL, '2026-08-11 19:27:52.687', '2026-08-11 19:27:52.823', NULL, 'a9c2ca965b3d40a924720c7761c5338d', 2) ON CONFLICT(record_id) DO UPDATE SET type=excluded.type, parent_record_id=excluded.parent_record_id, title=excluded.title, description=excluded.description, status=excluded.status, priority=excluded.priority, assigned_to=excluded.assigned_to, team=excluded.team, decision_ref=excluded.decision_ref, updated_at=excluded.updated_at, deleted_at=excluded.deleted_at, hash=excluded.hash, canonical_version=excluded.canonical_version WHERE excluded.updated_at >= tickets.updated_at;
-105
View File
@@ -1,105 +0,0 @@
# LLM Agent Instructions
This document contains instructions and documentation references for AI assistants working with this codebase.
> **📖 Important**: Before working on this project, read [docs/philosophy.md](docs/philosophy.md) to understand the system vision, architectural patterns, and design goals. All development should work towards realizing those patterns.
# AGENTS.md
> **Start every session by reading this file.**
> This file outlines the operational protocols, coding standards, and architectural decisions for this FastAPI project.
## 1. Agent Operational Protocols
### 🧠 Work Patterns (Plan-Act-Reflect)
* **Plan:** Before writing code, briefly outline your plan. Identify which files you will touch and what the side effects might be.
* **Act:** Execute the changes in small, atomic steps.
* **Reflect:** After coding, verify your work. Did you break existing tests? Did you add new tests?
### 🧪 Local Development Setup
* **Always test locally first** before committing and deploying. The build-deploy loop is slow.
* **Start the local server** with `./wakeup.sh` - logs are written to `logs/server.log` for easy tailing
* **Auto-reload**: The wakeup script runs uvicorn in reload mode - code changes are picked up automatically without restart (except for requirements.txt changes)
* **Test REST endpoints** against `http://localhost:8777` using curl or similar tools
* **Only deploy** when a phase or feature is complete and tested locally
* **Environment**: Copy `.env.example` to `.env` and configure for your local setup (Ollama, Redis, Qdrant hosts)
* **Running tests**: Always use the venv explicitly to avoid environment mismatches:
```bash
.venv/bin/python -m pytest tests/ # All tests
.venv/bin/python -m pytest tests/core/ -v # Core tests only
```
### 🌐 Internal Service Access
* **git.schweitz.net**: Access via `http://localhost:3002` (direct Gitea) to bypass Authentik SSO
* Example: `curl http://localhost:3002/jpmschweitzer/library-desk/raw/branch/main/README.md`
* Public repos are readable without authentication
* Related repos: `library-desk`, `scheduler`, `core-api`, `portainer-core`
### 🐳 Deployment & Infrastructure
* **Full stack documentation**: Available in the `portainer-core` repo
* Access: `curl http://localhost:3002/jpmschweitzer/portainer-core/raw/branch/main/CONTAINERS.md`
* Contains: All service ports, URLs, Redis DB allocations, external domains
* **Tatlock deployment**:
* LAN: `http://192.168.86.149:8000`
* External: `tatlock.schweitz.net` (behind Authentik SSO)
* Redis DBs: 1 (memory), 6 (benchmarks)
* **Health check**: `curl http://192.168.86.149:8000/health`
### 🛡️ Git Discipline
* **NEVER commit to `main` or `master` directly.** Always create a feature branch: `feature/your-feature-name` or `fix/issue-description`.
* **Commit Messages:** Use the [Conventional Commits](https://www.conventionalcommits.org/) format.
* `feat: add user login endpoint`
* `fix: resolve database connection timeout`
* `refactor: split monolith dependency file`
* **Atomic Commits:** Keep commits small. One logical change = one commit.
### 📝 Changelog Maintenance
* **Update `CHANGELOG.md`** with every user-facing change.
* Format: `## [Unreleased] - YYYY-MM-DD` followed by `### Added`, `### Changed`, or `### Fixed`.
### 🚀 Release Flow
When changes are ready for deployment:
1. **Ask user if deploy cycle is desired**
2. **Update version** in `pyproject.toml`:
- Bug fixes: bump patch version (1.8.3 → 1.8.4)
- New features: bump minor version (1.8.4 → 1.9.0)
3. **Update CHANGELOG.md**:
- Move items from `[Unreleased]` to new version section
- Add release date: `## [1.8.4] - 2025-12-16`
4. **Commit and tag**:
```bash
git add -A
git commit -m "fix: description of changes"
git tag v1.8.4
git push origin main --tags
```
5. **CI/CD triggers automatically**:
- Gitea CI builds Docker image on new tag
- Watchtower pulls and deploys to production
- Verify deployment: `curl http://192.168.86.149:8000/health`
---
## 2. FastAPI Architecture & Best Practices
*Reference: [FastAPI Best Practices](https://github.com/zhanymkanov/fastapi-best-practices)*
### 📂 Project Structure (Directory-based, NOT File-type based)
Do **not** group files by type (e.g., one huge `routers` folder). Group by **domain/module** inside a `src/` directory.
**Correct Structure:**
```text
src/
├── auth/
│ ├── router.py # Endpoints
│ ├── schemas.py # Pydantic models
│ ├── service.py # Business logic (CRUD, etc.)
│ ├── dependencies.py# Module-specific dependencies
│ └── config.py # Module-specific settings
├── posts/
│ ├── router.py
│ └── ...
└── main.py # App entry point
+24
View File
@@ -7,6 +7,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased] ## [Unreleased]
### Changed
- `make setup` now ends with a `pytest --collect-only` pass so a broken environment
(missing or mismatched dependency) fails the target itself instead of exiting 0
and surfacing later as a confusing test failure (T-47)
## [2.4.3] - 2026-08-08
### Fixed
- Steward routing no longer triggers on words inside its own explanation. Capability
extraction reads the declared `DELEGATE:` line instead of substring-matching
capability domains across the whole response, where ordinary English routed
requests — "description" contains the housekeeper domain "script", "acknowledge"
contains "knowledge" and "know". A spurious capability meant a real agent call,
including web searches, on queries that needed none.
## [2.4.2] - 2026-07-19
### Fixed
- Container crash-loop on fresh builds: cap `opentelemetry-api` below 1.44,
which removed the private `_events` module that pydantic-ai 1.27 imports
## [2.4.1] - 2026-07-19 ## [2.4.1] - 2026-07-19
### Changed ### Changed
+182 -20
View File
@@ -1,34 +1,196 @@
# CLAUDE.md # CLAUDE.md — tatlock
Claude Code-specific notes for this project. For general development instructions, architecture, coding standards, and deployment — see [AGENTS.md](AGENTS.md). Privacy-first homelab butler. An OpenAI-compatible orchestration API over local models, with
household staff agents built on PydanticAI. Python 3.12 / FastAPI, `version = "2.4.3"`.
Container `tatlock` on `docker-dataplane`, port **8000**. Redis DB **1** (memory), Qdrant for
vectors.
## Setup & Commands ## Ports
| | Port | How |
|---|---|---|
| Local dev | **8777** | `make run` — uvicorn reload, logs to `build/logs/server.log` |
| Production | **8000** | container; `http://192.168.86.149:8000/health`, external `tatlock.schweitz.net` behind Authentik |
Test endpoints against `localhost:8777` while developing. `localhost:8000` is the *container*.
## Live contract
`http://localhost:8000/openapi.json`**5 paths**, `title: OpenAI-Compatible API`, `version:
2.4.3` (verified 2026-08-09): `/`, `/health`, `/v1/models`, `/v1/chat/completions`,
`/v1/responses`. `/v1/responses` is primary; `/v1/chat/completions` exists for Open WebUI.
**The spec is the public surface, not the system.** The household capability registry is internal
and appears nowhere in those 5 paths. Absence from the spec means "not exposed", not "does not
exist".
## Two traps that make the runtime look like the opposite of what it is
**1. `src/anthropic` loads at startup; `src/ollama` does not — and Ollama is the primary
backend.** A cold `import src.main` inside the container shows `agents, anthropic, chat, core,
main, models, responses` — no `ollama`. The only import of it is a *function-body* one at
`src/anthropic/model_selector.py:230`. Meanwhile `PREFER_CLOUD_BACKEND=false`, so every request
actually goes to Ollama and the Claude path is off (see **workspace D-11**). Reading the module list
naively gives you exactly the wrong answer: the package that looks live is the disabled fallback,
and the one that looks dead is the hot path. Do not conclude anything about backends from
`sys.modules`; read the config.
**2. In-process singletons are empty outside the app.** `get_household_registry()`
(`src/core/household_registry.py:334`) in a fresh `docker exec python` returns **0 members**,
while the running app serves 2 models from it — it is populated at startup. Import the
module-level definitions or ask the endpoint; never import a singleton and assume it is
populated.
## Stack decisions that bind this repo
Recorded in the workspace vault, not here. Read before assuming anything about the LLM backend:
```bash ```bash
make setup # Create venv and install all dependencies /home/jpmschweitzer/.local/bin/pql --vault /mnt/media/Projects decisions read workspace D-11
make test # Unit tests (no external services)
make test-integration # Integration tests (needs Claude/Ollama)
make test-contracts # Wire-level contract tests against live service boundaries
make run # Start dev server on port 8777
make lint # Ruff linter + formatter check
make typecheck # Mypy
make clean # Remove caches and build artifacts
``` ```
Dependencies are in `pyproject.toml` (`[project.dependencies]` and `[project.optional-dependencies.dev]`). **workspace D-11 — the Claude migration is abandoned. Tatlock stays on Ollama.** Do not resume it and do
not treat its remnants as unfinished work. What you will find, and why none of it is a TODO:
`ANTHROPIC_MODEL` is set on the container (`claude-sonnet-4-20250514`) and never used because
`PREFER_CLOUD_BACKEND=false`; `ANTHROPIC_API_KEY` is a variable reference whose literal was
revoked 2026-08-09; `docs/claude-integration.md` documents a capability that exists but is
switched off. The cost is deliberate: reasoning stays at `gemma4:e2b` scale because VRAM is
shared with Speaches.
## Critical Gotchas `REDIS_BENCHMARK_DB=6` is allocated on the container but the benchmarking module was never
implemented — see the gotcha below. Vestigial, like the Anthropic settings.
**ASGITransport does NOT trigger FastAPI lifespan events.** The session-scoped `_initialize_app` fixture in `tests/conftest.py` calls `initialize_application()` explicitly via `asyncio.run()`. Without this, the Ollama/Claude health checks never run: `_ollama_available` stays `None` (treated as available, so requests go to Ollama) and `_claude_available` stays `None` (treated as unavailable, so the Claude fallback never engages). ## Critical gotchas
**AsyncIO scope mismatch.** `asyncio_default_fixture_loop_scope = function` is set in `pyproject.toml`. Session-scoped async fixtures cause `ScopeMismatch` errors. The fix is to use a sync fixture with `asyncio.run()` for session-scoped initialization. **ASGITransport does NOT trigger FastAPI lifespan events.** The session-scoped `_initialize_app`
fixture in `tests/conftest.py` calls `initialize_application()` explicitly via `asyncio.run()`.
Without it the Ollama/Claude health checks never run: `_ollama_available` stays `None` (treated
as available, so requests go to Ollama) and `_claude_available` stays `None` (treated as
unavailable, so the Claude fallback never engages).
**The butler persona prompt suppresses local-model tool calling.** With `TATLOCK_SYSTEM_PROMPT` attached, gemma4 reasons about calling the calculator, then answers from memory with wrong arithmetic (a different wrong product each run). `orchestrate_tool_calls()` therefore uses the terse `TATLOCK_ORCHESTRATION_PROMPT`; the persona is applied in `synthesize_from_results()`. Do not reattach the persona prompt to a tool-phase agent. `tool_choice: "required"` via extra_body does NOT force Ollama to call tools — it is advisory at best. **AsyncIO scope mismatch.** `asyncio_default_fixture_loop_scope = function` is set in
`pyproject.toml`. Session-scoped async fixtures raise `ScopeMismatch`. Use a sync fixture with
`asyncio.run()` for session-scoped initialization.
**Claude Sonnet 5+ rejects sampling parameters.** `temperature`/`top_p`/`top_k` return a 400. Use `get_sampling_settings()` from the model selector instead of passing `ModelSettings(temperature=...)` directly to agents that can run on the Claude fallback. The contract test suite pins this (`make test-contracts`). **The butler persona prompt suppresses local-model tool calling.** With `TATLOCK_SYSTEM_PROMPT`
attached, gemma4 reasons about calling the calculator, then answers from memory with wrong
arithmetic — a different wrong product each run. `orchestrate_tool_calls()` therefore uses the
terse `TATLOCK_ORCHESTRATION_PROMPT`; the persona is applied in `synthesize_from_results()`. Do
not reattach the persona prompt to a tool-phase agent. `tool_choice: "required"` via `extra_body`
does **not** force Ollama to call tools — advisory at best.
**Integration test timeouts.** Set to 120s to match `OLLAMA_TIMEOUT` config (300s for the pure-Ollama fallback test, which cannot be rescued by Claude). Current GPU-resident numbers (2026-07-14, driver 570, gemma4:e2b at ~100 tok/s): Steward analysis ~6s warm, full Steward → orchestrate → synthesize flow 1125s, librarian-routed queries ~20-25s. The old "~35s steward / ~2 min flow" figures were measured during the CPU-only era (driver mismatch, 13 tok/s) — do not plan against them. Cold start after 2h idle adds ~8s (`OLLAMA_KEEP_ALIVE=2h`). `STEWARD_TIMEOUT` defaults to 60s. **Claude Sonnet 5+ rejects sampling parameters.** `temperature`/`top_p`/`top_k` return 400. Use
`get_sampling_settings()` from the model selector rather than passing `ModelSettings(temperature=…)`
to agents that can run on the Claude fallback. `make test-contracts` pins this.
**`get_benchmark_store` does not exist.** The benchmarking module (`src/core/benchmarks.py`) was never implemented. `scripts/benchmark_analysis.py` also references it and is broken. Do not add mocks for it in tests. **Integration test timeouts** are 120s to match `OLLAMA_TIMEOUT` (300s for the pure-Ollama
fallback test, which Claude cannot rescue). GPU-resident numbers measured 2026-08-07 with
gemma4:e2b at ~95 tok/s: full Steward → orchestrate → synthesize ~1013s for simple turns;
librarian-routed ~2025s (not re-measured). **A single turn costs 3 sequential Ollama calls and
~710 generated tokens even for "what is 61 plus 12?"** — mostly the model's own reasoning, paid
three times. Cold model load is ~36s, avoided while pinned with `keep_alive: -1`; the
`OLLAMA_KEEP_ALIVE=2h` default reintroduces it. Older "~35s steward / ~2 min flow" and "1125s"
figures are superseded — do not plan against them. `STEWARD_TIMEOUT` defaults to 60s.
**Steward tests need household registry.** Use `register_household_members()` (sync) in fixtures, not `initialize_application()` (async). The steward extracts capabilities from the registry. **`get_benchmark_store` does not exist.** `src/core/benchmarks.py` was never implemented, and
`scripts/benchmark_analysis.py` references it and is broken. Do not add mocks for it in tests.
**Steward tests need the household registry.** Use `register_household_members()` (sync) in
fixtures, not `initialize_application()` (async). The steward extracts capabilities from the
registry.
## Commands
```bash
make setup # venv + all dependencies
make run # dev server on 8777, reload, logs to build/logs/server.log
make test # unit tests, no external services
make test-integration # needs Ollama (and Claude, if enabled)
make test-contracts # wire-level contract tests against live service boundaries
make lint # ruff linter + formatter check
make typecheck # mypy
make clean # remove caches and build artifacts
```
Always run pytest through the venv explicitly, to avoid environment mismatch:
```bash
.venv/bin/python -m pytest tests/
.venv/bin/python -m pytest tests/core/ -v
```
Dependencies live in `pyproject.toml` (`[project.dependencies]`, `[project.optional-dependencies.dev]`).
Copy `.env.example` to `.env` and configure Ollama, Redis and Qdrant hosts.
**Contract tests before code review.** When the question is "do these two services still agree?",
`make test-contracts` answers it by observing the live boundary; reading both codebases only tells
you what should happen. Semantics: unreachable → skip, reachable-but-wrong-shape → fail.
## Architecture
Domain-first under `src/`: `agents/` (steward, librarian, biographer, housekeeper, tatlock_core),
`core/`, `chat/`, `responses/`, `models/`, `ollama/`, `anthropic/`. Two tiers — the Steward routes,
Tatlock coordinates. Group new work by domain, not by file type.
## Internal service access
`http://localhost:3002` reaches Gitea directly, bypassing Authentik SSO — verified returning
`{"version":"1.27.1"}`. Useful for reading a sibling repo's raw files:
```bash
curl http://localhost:3002/jpmschweitzer/library-desk/raw/branch/main/README.md
```
The old AGENTS.md pointed at **`portainer-core`** for full-stack documentation. That repo is
**deprecated** and must not be used as a source of infra facts; it was merged into
`system-admin-toj/containers/`, where `CONTAINERS.md` is the live inventory.
## Work tracking
Work lives in **pql**, not a markdown TODO. **This repo's vault is standalone** — its tickets and
internal decisions live here in `.pql/` and `governance/`, and travel with a clone, because
`.pql/changelog/` is committed and replayed by the git hooks (workspace D-15). The databases are gitignored
and rebuildable with `pql plan rebuild`.
`pql` is **not** on the non-interactive `PATH` — invoke it as `/home/jpmschweitzer/.local/bin/pql`.
From inside this repo no `--vault` is needed; pql anchors at the nearest `.git/` ancestor.
```bash
/home/jpmschweitzer/.local/bin/pql ticket list # this repo's open work
/home/jpmschweitzer/.local/bin/pql plan whatsnext # next unblocked item, with context
/home/jpmschweitzer/.local/bin/pql decisions list # this repo's own decisions
```
Stack decisions that constrain this service need the flag:
```bash
/home/jpmschweitzer/.local/bin/pql --vault /mnt/media/Projects decisions list --domain tatlock-api
```
The workspace domain is `tatlock-api`, not `tatlock` — pql rejects a domain stem that prefixes
another, and `tatlock` prefixes `tatlock-ui`. A `tatlock-api -> tatlock` symlink at the workspace
root makes the directory answer to both (workspace D-15).
Note `ticket new --decision D-N` resolves ids within **one** vault, so a ticket here cannot link
to a workspace decision. Cite the id in the ticket body instead.
## Git
- **History is linear — no merge commits.** Work on `main`, or a short-lived branch that is
fast-forwarded and deleted. This repo's AGENTS.md mandated a feature branch for every change;
that rule was retired workspace-wide on 2026-08-08 and does not apply.
- **Conventional Commits**: `feat:`, `fix:`, `refactor:`, `docs:`, `chore:`.
- **Stage explicitly. Never `git add -A`** — denied by policy, and it sweeps in whatever else is
dirty, including secrets.
- Update `CHANGELOG.md` for every user-facing change, under `[Unreleased]`.
## Releasing
Test locally first — the build-deploy loop is slow. Deploy only when a feature is complete.
1. Ask whether a deploy is wanted; it is not automatic.
2. Bump `version` in `pyproject.toml` (patch for fixes, minor for features).
3. Move `[Unreleased]` entries into a dated section in `CHANGELOG.md`.
4. Stage the changed files by name, commit, tag `vX.Y.Z`, `git push origin main --tags`.
5. Gitea CI builds and pushes on the tag; Watchtower deploys.
6. Verify: `curl http://192.168.86.149:8000/health`.
+38
View File
@@ -18,6 +18,17 @@ setup: ## Create venv and install all dependencies
python3 -m venv $(VENV) python3 -m venv $(VENV)
$(PIP) install --upgrade pip $(PIP) install --upgrade pip
$(PIP) install -e ".[dev]" $(PIP) install -e ".[dev]"
# Exit 0 from pip install is not evidence the environment works (D-24) - the
# 2026-08-09 core-api incident was exactly this: a venv that "installed fine"
# but was missing a declared dependency, surfacing as 11 collection errors
# that read like broken imports rather than an environment problem. Collection
# is the right cheap check here for that same reason: it imports every test
# module (and everything they import) without running the suite, so a missing
# or mismatched dependency fails setup itself instead of showing up later as a
# mysterious test failure. Scoped like `make test` (excludes e2e/integration/
# contracts, which need external services) and --no-cov since coverage
# instrumentation is irrelevant to "does this collect".
$(PYTEST) --collect-only -q --ignore=tests/e2e --ignore=tests/integration --ignore=tests/contracts --no-cov
run: ## Start the development server on port 8777 run: ## Start the development server on port 8777
@mkdir -p build/logs @mkdir -p build/logs
@@ -49,3 +60,30 @@ typecheck: ## Run mypy type checking
clean: ## Remove build artifacts, caches, and coverage reports clean: ## Remove build artifacts, caches, and coverage reports
rm -rf .cache build rm -rf .cache build
find . -type d -name __pycache__ -exec rm -rf {} + 2>/dev/null || true find . -type d -name __pycache__ -exec rm -rf {} + 2>/dev/null || true
# git hands a hook a non-login shell, which never sees ~/.local/bin — where
# gitleaks lands. Without this the scan reports "not installed" on every push,
# which is a check that fails open (D-24).
export PATH := $(HOME)/.local/bin:/usr/local/bin:$(PATH)
.PHONY: secrets
secrets: ## Scan the commits about to be pushed for credentials
@ci/secrets.sh
# The call surface is identical in every repo; what it runs is not.
#
# `secrets` runs first, deliberately: it is the only failure here that cannot be
# undone by fixing it afterwards. A failed lint costs another commit; a pushed
# credential is cached and indexed whether or not it is later deleted.
#
# Some of these fail today, and are left wired anyway. The state was measured
# once and written down in T-56 rather than being worked around here — a gate
# quietly narrowed to what already passes is a gate that reports success for
# doing nothing, which is the failure this workspace keeps rediscovering.
.PHONY: pre-push
pre-push: secrets lint ## Everything the pre-push hook runs
@echo " -- not gated here yet: typecheck (T-1), test (T-56)"
@echo " typecheck reports 95 errors in 31 files and has never passed, so"
@echo " gating on it blocked every push to this repo — including the commit"
@echo " that added the gate. Run 'make typecheck' before pushing anything"
@echo " that touches types; T-1 is the pass that earns this line's removal."
+2 -2
View File
@@ -406,7 +406,7 @@ tatlock/
## Development ## Development
For LLM agent development guidelines and architectural decisions, see [AGENTS.md](AGENTS.md). For LLM agent development guidelines and architectural decisions, see [CLAUDE.md](CLAUDE.md).
## Contributing ## Contributing
@@ -420,7 +420,7 @@ For LLM agent development guidelines and architectural decisions, see [AGENTS.md
- **System Philosophy**: [docs/philosophy.md](docs/philosophy.md) - Vision, goals, and architectural patterns - **System Philosophy**: [docs/philosophy.md](docs/philosophy.md) - Vision, goals, and architectural patterns
- **Development Roadmap**: [docs/roadmap.md](docs/roadmap.md) - Open work and planned phases - **Development Roadmap**: [docs/roadmap.md](docs/roadmap.md) - Open work and planned phases
- **Developer Guidelines**: [AGENTS.md](AGENTS.md) - LLM agent development patterns - **Developer Guidelines**: [CLAUDE.md](CLAUDE.md) - LLM agent development patterns
- **Version History**: [CHANGELOG.md](CHANGELOG.md) - Changes and releases - **Version History**: [CHANGELOG.md](CHANGELOG.md) - Changes and releases
### External References ### External References
Executable
+50
View File
@@ -0,0 +1,50 @@
#!/usr/bin/env bash
# Secret scan over the commits about to be pushed.
#
# Lives here rather than inside .githooks/pre-push so it can be read, run by
# hand (`make secrets`), and changed under review. A hook is a trigger; it is
# not a home for logic. Identical in every repo in this workspace (D-27).
set -euo pipefail
cd "$(git rev-parse --show-toplevel)"
# A non-login shell — which is what git gives a hook — skips /etc/profile.d
# and never sees ~/.local/bin, where the gitleaks release tarball lands.
# Without this the scan reports "not installed" on every push.
[ -d "$HOME/.local/bin" ] && PATH="$HOME/.local/bin:$PATH"
if ! command -v gitleaks >/dev/null 2>&1; then
echo "FAIL secrets — gitleaks not installed, so this check would be a no-op pretending to pass." >&2
echo " https://github.com/gitleaks/gitleaks/releases → ~/.local/bin/gitleaks" >&2
exit 1
fi
# Scan the outgoing range, not full history. History here carries findings
# that are settled — test fixtures and vendored third-party code — and a gate
# that fails on something unfixable gets bypassed within a week. What matters
# is what is about to leave this machine.
if upstream=$(git rev-parse --abbrev-ref --symbolic-full-name '@{u}' 2>/dev/null); then
range="$upstream..HEAD"
elif git rev-parse --verify --quiet origin/main >/dev/null; then
range="origin/main..HEAD"
else
range=""
fi
if [ -z "$range" ]; then
gitleaks dir . --redact --no-banner --exit-code 1 || {
echo "FAIL secrets — gitleaks found a credential in the working tree." >&2; exit 1; }
exit 0
fi
[ -n "$(git log --oneline "$range" 2>/dev/null)" ] || exit 0
gitleaks git . --log-opts="$range" --redact --no-banner --exit-code 1 >/dev/null 2>&1 || {
echo "FAIL secrets — gitleaks found a credential in the commits being pushed." >&2
echo " inspect (values redacted): gitleaks git . --log-opts=\"$range\" --redact" >&2
echo " then remove and rotate it, or suppress deliberately:" >&2
echo " inline '# gitleaks:allow <reason>'" >&2
echo " or add the fingerprint to .gitleaksignore WITH a reason" >&2
exit 1
}
echo " ok secrets"
+1 -1
View File
@@ -28,7 +28,7 @@ src/mcp/
```yaml ```yaml
tatlock-mcp: tatlock-mcp:
image: git.schweitz.internal/jpmschweitzer/tatlock:latest image: git.schweitz.net/jpmschweitzer/tatlock:latest
command: ["python", "-m", "src.mcp.server"] command: ["python", "-m", "src.mcp.server"]
ports: ports:
- "8002:8002" - "8002:8002"
+2 -2
View File
@@ -10,7 +10,7 @@ This document establishes the foundational philosophy and architectural patterns
- When new architectural insights require rethinking core principles - When new architectural insights require rethinking core principles
**When NOT to modify this document**: **When NOT to modify this document**:
- During implementation of these patterns (use README.md, AGENTS.md, or code comments for technical details) - During implementation of these patterns (use README.md, CLAUDE.md, or code comments for technical details)
- For adding new household members or capabilities within the existing pattern - For adding new household members or capabilities within the existing pattern
- For tactical decisions about specific technologies or tools - For tactical decisions about specific technologies or tools
@@ -270,7 +270,7 @@ The user never directly interacts with the Steward or individual expert agents
**Related Documents**: **Related Documents**:
- **README.md**: User-facing documentation and usage guide - **README.md**: User-facing documentation and usage guide
- **AGENTS.md**: LLM agent development guidelines and technical patterns - **CLAUDE.md**: LLM agent development guidelines and technical patterns
- **CHANGELOG.md**: Version history and implemented features - **CHANGELOG.md**: Version history and implemented features
--- ---
+110
View File
@@ -0,0 +1,110 @@
# Steward Routing & Thinking — Findings
**Outcome: no change shipped.** The Steward stays on `gemma4:e2b` with model
thinking left at its default (on). Every alternative was measured and every one
loses. This document exists so the experiment is not repeated on the same
premise.
Run 2026-08-08 with `scripts/benchmark_routing.py` and
`scripts/fixtures/routing_fixtures.py` (40 labelled queries, one repeat per
cell, temperature 0.3 as production sends).
---
## The premise was wrong
The experiment was designed around an observation that the Steward pays ~300
tokens per turn for reasoning that is generated and thrown away: it calls
`/api/generate`, gemma4 reasons by default, and **no `thinking` field comes back
in the response**. Disabling thinking therefore looked close to free.
It is not. The reasoning is not discarded — it is emitted inline in `response`,
and it is what produces a correct `DELEGATE:` line. Those tokens are the work,
not waste. Suppressing them costs 12.5 points of routing accuracy.
## Results
| config | exact | under | over | tokens | latency | resident | predicted | co-resident with nomic |
|---|---|---|---|---|---|---|---|---|
| **e2b, thinking** *(production)* | **97.5%** | 2.5% | 0% | 361 | 5179 ms | 1778 MB | 7.8 GiB | yes |
| e2b, `think: false` | 85.0% | 12.5% | 5.0% | 48 | 1435 ms | 1778 MB | 7.8 GiB | yes |
| e4b, thinking | 100% | 0% | 0% | 192 | 4726 ms | 3089 MB | 10.6 GiB | **no** |
| e4b, `think: false` | 97.5% | 2.5% | 0% | 52 | 2269 ms | 3089 MB | 10.6 GiB | **no** |
`think: true` was also measured and landed within one fixture of the default on
both models, so production's implicit thinking is the same thing as asking for
it explicitly. Format compliance was 100% in every cell — a `DELEGATE:` line is
always emitted.
With 40 fixtures and one repeat, each result is worth 2.5 points, so the
97.5-vs-100 gaps are single fixtures and inside the noise. The latency and token
medians (40 calls each) and the e2b `think: false` degradation (6 failures with a
consistent mechanism) are the parts worth trusting.
## Why each alternative loses
**`think: false` on e2b** — 85% exact, and the failures are not random. All three
multi-capability fixtures under-route, each missing a second capability. Without
reasoning the model names one capability and stops decomposing. It is not
degraded across the board; it specifically stops handling compound requests,
which is where a user would most notice the Butler quietly doing half the job.
**e4b, either setting** — disqualified by memory, not by quality. Ollama predicts
**10.6 GiB** for it at 16k context. Maximum available on this card is ~7.9 GiB
(10.4 free 2.0 GPU overhead 0.46 minimum), so e4b *always* exceeds the budget
and evicts every co-resident before loading. Observed directly: loading it threw
out both `gemma4:e2b` and `nomic-embed-text`. Losing nomic means Tatlock memory
and library-desk thrash on every embedding call. Note this is not caused by the
2 GiB reservation — without it, available would be ~9.7 GiB, still under 10.6.
**Lower `OLLAMA_CONTEXT_LENGTH`** — the obvious way to free headroom, and it does
not work. Dropping 16384 → 2048, an 8× reduction, moved the prediction only from
7.8 to 6.7 GiB. The prediction is dominated by weights and batch size, not KV
cache. It would also truncate the Librarian's retrieved passages and webber's
code context for a 14% saving that funds nothing.
**Per-request `num_ctx`** — worse. A single request with a different `num_ctx`
reloads the shared runner, which **drops the `keep_alive: -1` pin** (expiry fell
from year-2318 to a 2-hour default) and evicts nomic. Three services share this
Ollama, so mixed context sizes are a thrash generator, and it fails silently.
**`OLLAMA_NUM_PARALLEL > 1`** — never viable here. e2b already predicts 7.8 GiB
against ~7.9 available, so there is no room for a second slot at any context
length. It is also set to 1 deliberately, to avoid batch overflow panics.
## What the two axes actually control
They do not interact, which is the useful part:
- **Model choice** governs VRAM and co-residency. e2b 1778 MB, e4b 3089 MB.
- **Think setting** governs tokens, latency and routing quality — and costs
**nothing** in VRAM. Verified: e2b is resident at 1778 MB with `think` unset,
true and false alike, because the KV cache is allocated for the full context at
load time and `think` is a per-request generation parameter.
So the only real question is whether 313 tokens and 3.7 seconds are worth 12.5
points of compound-query routing. On a turn that is already three sequential
Ollama calls, they are.
## Prerequisite: the extraction fix
These numbers are only meaningful because `_extract_capabilities` was fixed first
(commit `a905363`). It previously substring-matched capability *domains* across
the Steward's entire response, so ordinary English in the `REASON:` line selected
agents — "description" contains the housekeeper domain "script", "acknowledge"
contains "knowledge" and "know".
That made **prose length a routing input**. Benchmarking against it would have
shown `think: false` improving routing purely because shorter output produces
fewer accidental substring hits — a thinking policy derived from a parsing
artefact. The `adversarial` fixture group is regression coverage for exactly this.
## If this is revisited
The constraint is the single 11 GB card, not the model. A second inference host
(*forge*) removes it entirely, and e4b's 100% routing becomes reachable without
evicting anything. Re-run then; on this card the answer is settled.
`scripts/benchmark_routing.py` takes `--models`, `--think` and `--repeats`, and
restores GPU residency on exit — including on SIGTERM, which the first version
did not.
+54
View File
@@ -0,0 +1,54 @@
# Decisions, Questions, Rejected
This directory holds structured planning records that pql parses
into pql.db. Each record is a `### [DQR]-N: Title` heading inside
a markdown file. Files live in three per-type subdirectories:
- `decisions/<domain>.md` — confirmed design decisions
- `questions/<domain>.md` — open questions that may resolve into
decisions or rejected proposals
- `rejected/<domain>.md` — rejected proposals (kept for the audit
trail)
The parser infers domain from the filename stem and record type
from the parent subdirectory.
D-records that propose implementation work link to `initiative`-type
tickets via `decision_ref`. Run `pql decisions show <id>
--with-tickets` to inspect implementation status.
## Recommended domains
Start with this canonical set; create files as records land in
each domain:
- **architecture** — structural commitments (storage, layering,
languages, libraries)
- **process** — team workflow (commits, branches, releases, reviews)
- **design** — user-facing surface (UX, UI, public APIs)
- **coding-conventions** — team-internal code shape (style, lint,
file layout)
- **testing** — quality strategy (coverage, layers, gates)
You might also want, project-permitting:
- `accessibility` — if you ship user-facing software
- `security` — if you handle user data or network surfaces
- `licensing` — if you release open-source or commercial
- `documentation` — if user-docs are non-trivial
- `deployment` — if shipping is non-trivial
- `performance` — if you have perf budgets / SLOs
<!-- pql:records (auto-generated; do not edit manually) -->
## Decisions
- _(none)_
## Open questions
- _(none)_
## Rejected
- _(none)_
+4 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "tatlock" name = "tatlock"
version = "2.4.1" version = "2.4.3"
description = "OpenAI-compatible API with Ollama backend" description = "OpenAI-compatible API with Ollama backend"
requires-python = ">=3.12" requires-python = ">=3.12"
dependencies = [ dependencies = [
@@ -13,6 +13,9 @@ dependencies = [
"pydantic>=2.11,<2.13", "pydantic>=2.11,<2.13",
"pydantic-settings>=2.12,<2.13", "pydantic-settings>=2.12,<2.13",
"pydantic-ai-slim[openai,anthropic]>=1.27,<1.28", "pydantic-ai-slim[openai,anthropic]>=1.27,<1.28",
# pydantic-ai 1.27 imports the private opentelemetry._events module,
# removed in opentelemetry-api 1.44 — cap until pydantic-ai is bumped
"opentelemetry-api>=1.30,<1.44",
"anthropic>=0.77,<1.0", "anthropic>=0.77,<1.0",
"httpx>=0.28,<0.29", "httpx>=0.28,<0.29",
"sse-starlette>=3.0,<3.1", "sse-starlette>=3.0,<3.1",
+235
View File
@@ -0,0 +1,235 @@
"""
Benchmark Steward routing quality against model and thinking settings.
Talks to Ollama directly. No Tatlock server, no agents, no tools, nothing is
executed — the mutating fixtures ("turn on the lights", "update the wiki") only
ever produce a routing decision. That makes this cheap and repeatable, and it
isolates the question: does the Steward still pick the right capabilities when
the model reasons less?
The request body is byte-identical to StewardAgent._call_ollama, plus the
`think` flag under test, so a cell labelled `unset` is exactly what production
sends today.
Three thinking settings, because "on vs off" hides the interesting case:
unset what production sends now. gemma4 reasons by default, and the
response carries no `thinking` field, so those tokens are generated
and discarded.
true reasoning requested explicitly and returned in `thinking`.
false reasoning suppressed.
Scoring is deliberately asymmetric. A missing capability under-routes and the
Butler answers without a tool it needed; a spurious one over-routes, and that is
a real agent call — a stray librarian is a multi-second web search on a query
that asked for arithmetic. Over-routing is the predicted failure when thinking
is off, so `forbid` violations are reported separately rather than folded into
one accuracy number.
Usage:
.venv/bin/python scripts/benchmark_routing.py
.venv/bin/python scripts/benchmark_routing.py --models gemma4:e2b
.venv/bin/python scripts/benchmark_routing.py --think false --repeats 3
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
import time
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
import httpx
PROJECT_ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(PROJECT_ROOT))
from scripts.fixtures.routing_fixtures import FIXTURES # noqa: E402
from scripts.ollama_residency import ( # noqa: E402
install_sigterm_handler,
residency_guard,
)
from src.agents.steward.agent import build_steward_prompt # noqa: E402
from src.agents.steward.service import _DELEGATE_LINE_RE, _extract_capabilities # noqa: E402
from src.core.startup import register_household_members # noqa: E402
OLLAMA_URL = "http://localhost:11434"
DEFAULT_MODELS = ["gemma4:e2b", "gemma4:e4b"]
DEFAULT_THINK = ["unset", "true", "false"]
RESULTS_DIR = PROJECT_ROOT / "logs"
def build_body(model: str, prompt: str, think: str) -> dict[str, Any]:
"""Mirror StewardAgent._call_ollama exactly, then add the flag under test."""
body: dict[str, Any] = {
"model": model,
"prompt": prompt,
"stream": False,
"options": {
"temperature": 0.3, # Lower = more consistent
"top_p": 0.9,
},
}
if think != "unset":
body["think"] = think == "true"
return body
def call(client: httpx.Client, body: dict[str, Any]) -> dict[str, Any] | None:
try:
response = client.post(f"{OLLAMA_URL}/api/generate", json=body)
response.raise_for_status()
return response.json()
except Exception as exc: # noqa: BLE001 - a failed cell must not abort the run
print(f" ! {exc}", file=sys.stderr)
return None
def score(fixture: dict, found: list[str]) -> dict[str, Any]:
expected = set(fixture["expect"])
forbidden = set(fixture["forbid"])
got = set(found)
missing = sorted(expected - got)
spurious = sorted(got & forbidden)
return {
"found": found,
"missing": missing,
"spurious": spurious,
# Exact only when everything expected arrived and nothing forbidden did.
"exact": not missing and not spurious,
"under_routed": bool(missing),
"over_routed": bool(spurious),
}
def run_cell(client: httpx.Client, model: str, think: str, repeats: int) -> list[dict[str, Any]]:
rows: list[dict[str, Any]] = []
for fixture in FIXTURES:
prompt = build_steward_prompt(fixture["query"], [])
body = build_body(model, prompt, think)
for rep in range(repeats):
started = time.perf_counter()
data = call(client, body)
elapsed_ms = (time.perf_counter() - started) * 1000
if data is None:
rows.append({
"id": fixture["id"], "group": fixture["group"], "rep": rep,
"error": True, "exact": False, "under_routed": False, "over_routed": False,
})
continue
text = data.get("response", "") or ""
found = _extract_capabilities(text)
rows.append({
"id": fixture["id"],
"group": fixture["group"],
"rep": rep,
"error": False,
"latency_ms": round(elapsed_ms, 1),
"eval_tokens": data.get("eval_count"),
"prompt_tokens": data.get("prompt_eval_count"),
# Did the model obey the documented output shape at all?
"has_delegate_line": bool(_DELEGATE_LINE_RE.search(text)),
# Whether reasoning came back, as opposed to being generated and dropped.
"thinking_returned": bool(data.get("thinking")),
"response_chars": len(text),
**score(fixture, found),
})
return rows
def summarise(rows: list[dict[str, Any]]) -> dict[str, Any]:
ok = [r for r in rows if not r["error"]]
if not ok:
return {"n": 0, "errors": len(rows)}
latencies = [r["latency_ms"] for r in ok]
tokens = [r["eval_tokens"] for r in ok if r["eval_tokens"] is not None]
return {
"n": len(ok),
"errors": len(rows) - len(ok),
"exact_pct": round(100 * sum(r["exact"] for r in ok) / len(ok), 1),
"under_routed_pct": round(100 * sum(r["under_routed"] for r in ok) / len(ok), 1),
"over_routed_pct": round(100 * sum(r["over_routed"] for r in ok) / len(ok), 1),
"format_ok_pct": round(100 * sum(r["has_delegate_line"] for r in ok) / len(ok), 1),
"thinking_returned_pct": round(100 * sum(r["thinking_returned"] for r in ok) / len(ok), 1),
"latency_ms_median": round(statistics.median(latencies), 1),
"latency_ms_mean": round(statistics.fmean(latencies), 1),
"eval_tokens_median": round(statistics.median(tokens), 1) if tokens else None,
"eval_tokens_total": sum(tokens) if tokens else None,
}
def main() -> int:
install_sigterm_handler()
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--models", default=",".join(DEFAULT_MODELS))
parser.add_argument("--think", default=",".join(DEFAULT_THINK),
help="comma-separated subset of unset,true,false")
parser.add_argument("--repeats", type=int, default=1)
parser.add_argument("--timeout", type=float, default=180.0)
args = parser.parse_args()
models = [m.strip() for m in args.models.split(",") if m.strip()]
think_modes = [t.strip() for t in args.think.split(",") if t.strip()]
# build_steward_prompt reads the registry, and the registry is populated at
# application startup. Without this the prompt lists no capabilities and every
# cell scores zero for reasons that have nothing to do with the model.
register_household_members()
print(f"{len(FIXTURES)} fixtures x {len(models)} models x {len(think_modes)} think "
f"x {args.repeats} repeats = {len(FIXTURES) * len(models) * len(think_modes) * args.repeats} calls\n")
cells: dict[str, Any] = {}
# The guard restores production's pinned models however this exits — a
# finished run, a failed cell, Ctrl-C or SIGTERM.
with residency_guard(models_used=models), httpx.Client(timeout=args.timeout) as client:
for model in models:
# Absorb the cold load (~36s) outside the measurements.
print(f"warming {model} ...", flush=True)
call(client, build_body(model, "hi", "false"))
for think in think_modes:
key = f"{model}|think={think}"
print(f" {key} ...", end=" ", flush=True)
started = time.perf_counter()
rows = run_cell(client, model, think, args.repeats)
summary = summarise(rows)
cells[key] = {"summary": summary, "rows": rows}
print(f"exact={summary.get('exact_pct')}% "
f"over={summary.get('over_routed_pct')}% "
f"median={summary.get('latency_ms_median')}ms "
f"({time.perf_counter() - started:.0f}s)")
RESULTS_DIR.mkdir(parents=True, exist_ok=True)
stamp = datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
out = RESULTS_DIR / f"routing-bench-{stamp}.json"
out.write_text(json.dumps({
"generated_at": datetime.now(UTC).isoformat(),
"fixtures": len(FIXTURES),
"repeats": args.repeats,
"cells": cells,
}, indent=2))
print(f"\n{'cell':28} {'exact':>7} {'under':>7} {'over':>7} {'fmt':>6} {'tok':>7} {'ms':>8}")
print("-" * 76)
for key, cell in cells.items():
s = cell["summary"]
print(f"{key:28} {s.get('exact_pct'):>6}% {s.get('under_routed_pct'):>6}% "
f"{s.get('over_routed_pct'):>6}% {s.get('format_ok_pct'):>5}% "
f"{str(s.get('eval_tokens_median')):>7} {s.get('latency_ms_median'):>8}")
print(f"\nwritten to {out}")
return 0
if __name__ == "__main__":
try:
sys.exit(main())
except KeyboardInterrupt:
# The residency guard has already run by the time this is caught;
# a traceback here would just bury its output.
print("\ninterrupted", file=sys.stderr)
sys.exit(130)
+1 -1
View File
@@ -268,7 +268,7 @@ async def run_benchmarks(iterations: int = 10, verbose: bool = False):
print(f" Max: {overall_max:.3f}s (target: ≤5.0s)") print(f" Max: {overall_max:.3f}s (target: ≤5.0s)")
print(f" Avg: {overall_avg:.3f}s (target: ≤1.67s)") print(f" Avg: {overall_avg:.3f}s (target: ≤1.67s)")
print(f"\n Recommendations:") print(f"\n Recommendations:")
print(f" - Switch to a faster model (current: mistral-nemo)") print(f" - Switch to a faster model (current: gemma4:e2b)")
print(f" - Reduce system prompt complexity") print(f" - Reduce system prompt complexity")
print(f" - Limit tool calls (currently limited to 3)") print(f" - Limit tool calls (currently limited to 3)")
print(f" - Consider caching household registry responses") print(f" - Consider caching household registry responses")
+26 -9
View File
@@ -17,12 +17,16 @@ import asyncio
import json import json
import re import re
import statistics import statistics
import sys
import time import time
from dataclasses import dataclass, field from dataclasses import dataclass, field
from pathlib import Path from pathlib import Path
import httpx import httpx
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from scripts.ollama_residency import install_sigterm_handler, residency_guard
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Configuration # Configuration
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -525,18 +529,31 @@ async def main():
original_env = ENV_PATH.read_text() original_env = ENV_PATH.read_text()
all_stats = [] all_stats = []
async with httpx.AsyncClient() as client: # Both restores must survive a crash or an interrupt. The .env one especially:
for model in models: # this script rewrites OLLAMA_DEFAULT_MODEL and lets uvicorn reload onto it,
stats = await benchmark_model(client, model, args.iterations) # so bailing out mid-run used to leave the *running server* pointed at the
all_stats.append(stats) # benchmark model — and DEFAULT_MODELS starts at mistral-nemo-large, the 9.2G
# model implicated in the 2026-08-07 VRAM outage.
# Restore original .env install_sigterm_handler()
ENV_PATH.write_text(original_env) try:
print(f"\n .env restored to original") with residency_guard(models_used=models):
async with httpx.AsyncClient() as client:
for model in models:
stats = await benchmark_model(client, model, args.iterations)
all_stats.append(stats)
finally:
ENV_PATH.write_text(original_env)
print("\n .env restored to original")
print_comparison(all_stats) print_comparison(all_stats)
save_results(all_stats, Path(args.output)) save_results(all_stats, Path(args.output))
if __name__ == "__main__": if __name__ == "__main__":
asyncio.run(main()) try:
asyncio.run(main())
except KeyboardInterrupt:
# .env and GPU residency are both restored by now; do not bury that
# output under a traceback.
print("\ninterrupted", file=sys.stderr)
raise SystemExit(130) from None
View File
+159
View File
@@ -0,0 +1,159 @@
"""
Labelled queries for the Steward routing benchmark.
Each fixture carries both `expect` and `forbid`:
expect capabilities that must appear. Missing one is under-routing — the
Butler answers without a tool it needed.
forbid capabilities that must not appear. Over-routing is not cosmetic: a
spurious librarian is a real multi-second web call, and a spurious
housekeeper can actuate hardware.
`forbid` matters more than `expect` here, because over-recommendation is the
predicted failure when model thinking is disabled and the Steward has less room
to discriminate.
The `adversarial` group deserves explanation. Until 2026-08-08 the extractor
substring-matched capability *domains* across the Steward's whole response, so
ordinary English in its REASON line selected agents: "description" contains the
housekeeper domain "script", "acknowledge" contains "knowledge" and "know",
"economy" contains the biographer domain "my". Those queries invite exactly that
vocabulary. They now serve as an end-to-end regression: routing must depend on
what the Steward *decided*, not on the words it happened to use while explaining.
Expectations follow the routing rules stated in the Steward prompt itself
(src/agents/steward/agent.py), not on what a capability could plausibly cover.
"""
CORE = "tatlock_core"
LIB = "librarian"
BIO = "biographer"
HOUSE = "housekeeper"
ALL = [CORE, LIB, BIO, HOUSE]
def _others(*keep: str) -> list[str]:
return [c for c in ALL if c not in keep]
FIXTURES: list[dict] = [
# --- arithmetic and computation -> tatlock_core --------------------------
{"id": "math_add", "group": "math", "query": "What is 61 plus 12?",
"expect": [CORE], "forbid": _others(CORE)},
{"id": "math_percent", "group": "math", "query": "What is 15% of 240?",
"expect": [CORE], "forbid": _others(CORE)},
{"id": "math_compound", "group": "math", "query": "If I save 200 a month for 3 years, how much is that?",
"expect": [CORE], "forbid": _others(CORE)},
{"id": "math_sqrt", "group": "math", "query": "What is the square root of 1764?",
"expect": [CORE], "forbid": _others(CORE)},
# --- date and time -> tatlock_core ---------------------------------------
{"id": "time_now", "group": "datetime", "query": "What time is it?",
"expect": [CORE], "forbid": _others(CORE)},
{"id": "time_date", "group": "datetime", "query": "What is today's date?",
"expect": [CORE], "forbid": _others(CORE)},
{"id": "time_delta", "group": "datetime", "query": "How many days until Christmas?",
"expect": [CORE], "forbid": _others(CORE)},
# --- personal memory -> biographer ---------------------------------------
{"id": "bio_location", "group": "biographer", "query": "Where do I live?",
"expect": [BIO], "forbid": [LIB, HOUSE]},
{"id": "bio_name", "group": "biographer", "query": "What's my name?",
"expect": [BIO], "forbid": [LIB, HOUSE]},
{"id": "bio_car", "group": "biographer", "query": "What car do I drive?",
"expect": [BIO], "forbid": [LIB, HOUSE]},
{"id": "bio_store", "group": "biographer", "query": "Remember that I prefer my coffee black.",
"expect": [BIO], "forbid": [LIB, HOUSE]},
{"id": "bio_list", "group": "biographer", "query": "What do you know about me?",
"expect": [BIO], "forbid": [LIB, HOUSE]},
{"id": "bio_forget", "group": "biographer", "query": "Forget my old address.",
"expect": [BIO], "forbid": [LIB, HOUSE]},
# --- research and current information -> librarian ------------------------
{"id": "lib_weather", "group": "librarian", "query": "What's the weather in Rotterdam tomorrow?",
"expect": [LIB], "forbid": [HOUSE]},
{"id": "lib_news", "group": "librarian", "query": "What's in the news today?",
"expect": [LIB], "forbid": [HOUSE, BIO]},
{"id": "lib_url", "group": "librarian", "query": "Read https://example.com/article and summarise it.",
"expect": [LIB], "forbid": [HOUSE, BIO]},
{"id": "lib_research", "group": "librarian", "query": "Research how tidal power stations work.",
"expect": [LIB], "forbid": [HOUSE, BIO]},
{"id": "lib_wiki_create", "group": "librarian", "query": "Create a wiki page about our network topology.",
"expect": [LIB], "forbid": [HOUSE, BIO]},
# --- home automation -> housekeeper --------------------------------------
{"id": "house_lights_on", "group": "housekeeper", "query": "Turn on the kitchen lights.",
"expect": [HOUSE], "forbid": [LIB, BIO, CORE]},
{"id": "house_lights_off", "group": "housekeeper", "query": "Switch off all the lights downstairs.",
"expect": [HOUSE], "forbid": [LIB, BIO, CORE]},
{"id": "house_thermostat", "group": "housekeeper", "query": "Set the thermostat to 20 degrees.",
"expect": [HOUSE], "forbid": [LIB, BIO]},
{"id": "house_blinds", "group": "housekeeper", "query": "Close the blinds in the living room.",
"expect": [HOUSE], "forbid": [LIB, BIO, CORE]},
# --- conversational -> nothing at all -------------------------------------
# The expensive failure mode: a greeting that triggers a web search.
{"id": "chat_greeting", "group": "conversational", "query": "Hello!",
"expect": [], "forbid": ALL},
{"id": "chat_thanks", "group": "conversational", "query": "Thanks, that's helpful.",
"expect": [], "forbid": ALL},
{"id": "chat_joke", "group": "conversational", "query": "Tell me a joke.",
"expect": [], "forbid": ALL},
{"id": "chat_howareyou", "group": "conversational", "query": "How are you doing today?",
"expect": [], "forbid": ALL},
{"id": "chat_prior_turn", "group": "conversational", "query": "What did I just say?",
"expect": [], "forbid": ALL},
# --- genuinely multi-capability -------------------------------------------
{"id": "multi_weather_home", "group": "multi",
"query": "What's the weather here, and remember that I like it warm?",
"expect": [LIB, BIO], "forbid": []},
{"id": "multi_recall_search", "group": "multi",
"query": "Look up the best route from my home address to Utrecht.",
"expect": [BIO, LIB], "forbid": []},
{"id": "multi_math_memory", "group": "multi",
"query": "Remember that my budget is 500 euro, then work out 12% of it.",
"expect": [BIO, CORE], "forbid": [LIB, HOUSE]},
# --- adversarial: vocabulary that used to select agents by substring ------
# "temperature" is a housekeeper domain, but this is a unit conversion.
{"id": "adv_temperature", "group": "adversarial", "query": "Convert 98.6 Fahrenheit to Celsius.",
"expect": [CORE], "forbid": [HOUSE, LIB, BIO]},
# "description" contains "script"; "discover" contains "cover".
{"id": "adv_description", "group": "adversarial",
"query": "Give me a short description of what 17 times 23 comes to.",
"expect": [CORE], "forbid": [HOUSE, LIB]},
# "acknowledge" contains "knowledge" and "know".
{"id": "adv_acknowledge", "group": "adversarial",
"query": "Just acknowledge this and add 5 and 6 for me.",
"expect": [CORE], "forbid": [LIB, BIO]},
# "my" appears inside "economy".
{"id": "adv_economy", "group": "adversarial",
"query": "How many zeros are in one trillion?",
"expect": [CORE], "forbid": [BIO, HOUSE]},
# "fan" inside "fantastic"; also a climate word without a home-control intent.
{"id": "adv_fantastic", "group": "adversarial",
"query": "That's fantastic. What is 8 squared?",
"expect": [CORE], "forbid": [HOUSE, LIB]},
# "home" without any actuation intent.
{"id": "adv_home_word", "group": "adversarial", "query": "What time do I usually get home?",
"expect": [BIO], "forbid": [HOUSE]},
# "search" as ordinary English, not a web-search request.
{"id": "adv_search_word", "group": "adversarial",
"query": "No need to search anything, just tell me what 9 times 9 is.",
"expect": [CORE], "forbid": [LIB]},
# "create"/"write" are librarian domains but this is conversational.
{"id": "adv_write_word", "group": "adversarial", "query": "Can you write that more simply?",
"expect": [], "forbid": [LIB, HOUSE]},
# --- mutating intents: routing only, nothing is ever executed -------------
{"id": "mutate_wiki_update", "group": "mutating", "query": "Update the dossier page with today's findings.",
"expect": [LIB], "forbid": [HOUSE, CORE]},
{"id": "mutate_scene", "group": "mutating", "query": "Run the movie night scene.",
"expect": [HOUSE], "forbid": [LIB, BIO, CORE]},
]
GROUPS = sorted({f["group"] for f in FIXTURES})
assert len({f["id"] for f in FIXTURES}) == len(FIXTURES), "duplicate fixture id"
+121
View File
@@ -0,0 +1,121 @@
"""
Guard production's GPU residency across a benchmark run.
Benchmarks swap models on the card production is serving from. Ollama evicts to
make room, so a run leaves its own models resident and the production one gone:
the next voice turn pays a ~36s cold load, and the pin that prevented it is
silently lost. That happened on 2026-08-08 — a routing benchmark evicted
gemma4:e2b and left gemma4:e4b behind, and only the monitoring noticing
`unexpected_models` caught it.
Snapshot before, restore after, and wire the restore to SIGTERM as well as the
normal path. Python runs `finally` for SIGINT, which arrives as
KeyboardInterrupt, but the default SIGTERM action terminates outright — so
`timeout`, a systemd stop or a plain `kill` would skip the guard entirely.
from scripts.ollama_residency import residency_guard, install_sigterm_handler
install_sigterm_handler()
with residency_guard(models_used=["gemma4:e4b"]):
...
"""
from __future__ import annotations
import signal
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import UTC, datetime
from typing import Any
import httpx
OLLAMA_URL = "http://localhost:11434"
# keep_alive:-1 yields a year-2318 expiry, so "pinned" is simply "expires more
# than a day out". Matches check-ai-pipeline.sh in system-admin-toj.
PINNED_THRESHOLD_SECONDS = 86400
def install_sigterm_handler() -> None:
"""Make SIGTERM raise, so `finally` blocks and context managers still run."""
def _raise(signum, _frame):
raise KeyboardInterrupt(f"signal {signum}")
signal.signal(signal.SIGTERM, _raise)
def snapshot_residency(client: httpx.Client | None = None) -> dict[str, bool]:
"""Resident models mapped to whether each is pinned."""
owns = client is None
client = client or httpx.Client(timeout=30)
try:
data = client.get(f"{OLLAMA_URL}/api/ps", timeout=10).json()
except Exception: # noqa: BLE001 - a missing snapshot must not abort the run
return {}
finally:
if owns:
client.close()
resident: dict[str, bool] = {}
now = datetime.now(UTC)
for model in data.get("models", []):
pinned = False
try:
expires = datetime.fromisoformat(model.get("expires_at", "").replace("Z", "+00:00"))
pinned = (expires - now).total_seconds() > PINNED_THRESHOLD_SECONDS
except ValueError:
pass
resident[model["name"]] = pinned
return resident
def set_keep_alive(model: str, keep_alive: Any, client: httpx.Client | None = None) -> bool:
"""Load, unload or pin a model. Embedding models reject /api/generate."""
owns = client is None
client = client or httpx.Client(timeout=180)
payload = {"model": model, "keep_alive": keep_alive}
try:
for endpoint in ("generate", "embed"):
try:
response = client.post(f"{OLLAMA_URL}/api/{endpoint}", json=payload, timeout=180)
except Exception: # noqa: BLE001
return False
if response.status_code == 200:
return True
if response.status_code == 400 and "does not support generate" in response.text:
continue # embedding-only model; try /api/embed
return False
return False
finally:
if owns:
client.close()
def restore_residency(before: dict[str, bool], used: list[str]) -> None:
"""Evict what the benchmark loaded, then re-pin what was pinned before."""
base = {name.split(":")[0] for name in before}
with httpx.Client(timeout=180) as client:
for model in used:
if model not in before and model.split(":")[0] not in base:
print(f" residency: unloading benchmark model {model}")
set_keep_alive(model, 0, client)
for name, pinned in before.items():
if not pinned:
continue
ok = set_keep_alive(name, -1, client)
print(f" residency: re-pinned {name}" if ok
else f" residency: FAILED to re-pin {name} -- run warmup-ollama.sh")
@contextmanager
def residency_guard(models_used: list[str]) -> Iterator[dict[str, bool]]:
"""Snapshot residency on entry, restore it on exit however that happens."""
before = snapshot_residency()
pinned = [n for n, p in before.items() if p]
print(f" residency: resident before {sorted(before)}"
f"{f' (pinned: {pinned})' if pinned else ''}")
try:
yield before
finally:
print(" residency: restoring ...")
restore_residency(before, models_used)
+5 -9
View File
@@ -6,7 +6,8 @@ must implement. The interface is designed around the Responses API format.
""" """
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from typing import AsyncGenerator, Any from collections.abc import AsyncGenerator
from typing import Any
class OutputItem: class OutputItem:
@@ -19,12 +20,7 @@ class OutputItem:
- message: Assistant response message - message: Assistant response message
""" """
def __init__( def __init__(self, type: str, id: str, **kwargs: Any):
self,
type: str,
id: str,
**kwargs: Any
):
self.type = type self.type = type
self.id = id self.id = id
self.data = kwargs self.data = kwargs
@@ -40,7 +36,7 @@ class AgentInterface(ABC):
""" """
@abstractmethod @abstractmethod
async def generate_response( def generate_response(
self, self,
messages: list[dict], messages: list[dict],
reasoning: dict | None = None, reasoning: dict | None = None,
@@ -48,7 +44,7 @@ class AgentInterface(ABC):
temperature: float = 1.0, temperature: float = 1.0,
max_tokens: int | None = None, max_tokens: int | None = None,
stop: list[str] | None = None, stop: list[str] | None = None,
**kwargs: Any **kwargs: Any,
) -> AsyncGenerator[OutputItem, None]: ) -> AsyncGenerator[OutputItem, None]:
""" """
Generate streaming response as output items. Generate streaming response as output items.
+1
View File
@@ -11,6 +11,7 @@ For direct key-based lookups (location, timezone, preferences),
use the memory_service instead - it's faster and doesn't require LLM. use the memory_service instead - it's faster and doesn't require LLM.
The Biographer handles semantic, fuzzy queries. The Biographer handles semantic, fuzzy queries.
""" """
from src.agents.biographer.agent import ( from src.agents.biographer.agent import (
get_biographer_agent, get_biographer_agent,
run_biographer, run_biographer,
+6 -5
View File
@@ -7,7 +7,8 @@ A PydanticAI agent that serves as the household's memory keeper:
- Manages user profile and preferences - Manages user profile and preferences
- Forgets information when requested - Forgets information when requested
""" """
from typing import Any, Optional
from typing import Any
from pydantic_ai import Agent from pydantic_ai import Agent
@@ -19,7 +20,6 @@ from src.agents.biographer.tools import (
update_preference, update_preference,
update_profile, update_profile,
) )
from src.core.config import config
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -97,7 +97,7 @@ When recalling:
""" """
# Lazy initialization to avoid connection issues during imports # Lazy initialization to avoid connection issues during imports
_biographer_agent: Optional[Agent[None, str]] = None _biographer_agent: Agent[None, str] | None = None
def _create_biographer_agent() -> Agent[None, str]: def _create_biographer_agent() -> Agent[None, str]:
@@ -126,6 +126,7 @@ def _create_biographer_agent() -> Agent[None, str]:
agent.tool_plain(forget_memory) agent.tool_plain(forget_memory)
from src.anthropic.model_selector import get_model_info from src.anthropic.model_selector import get_model_info
model_info = get_model_info() model_info = get_model_info()
logger.info( logger.info(
"biographer_agent_created", "biographer_agent_created",
@@ -153,7 +154,7 @@ def get_biographer_agent() -> Agent[None, str]:
async def run_biographer( async def run_biographer(
task: str, task: str,
context: str = "", context: str = "",
message_history: Optional[list[Any]] = None, message_history: list[Any] | None = None,
) -> str: ) -> str:
""" """
Execute a memory task with The Biographer. Execute a memory task with The Biographer.
@@ -216,7 +217,7 @@ async def run_biographer(
async def run_biographer_stream( async def run_biographer_stream(
task: str, task: str,
context: str = "", context: str = "",
message_history: Optional[list[Any]] = None, message_history: list[Any] | None = None,
): ):
""" """
Execute a memory task with streaming output. Execute a memory task with streaming output.
+1
View File
@@ -4,6 +4,7 @@ Biographer capability registration for the Household Registry.
Defines The Biographer's capabilities and registers it as a Defines The Biographer's capabilities and registers it as a
household member for coordination by the Steward and Tatlock. household member for coordination by the Steward and Tatlock.
""" """
from src.agents.biographer.agent import get_biographer_agent from src.agents.biographer.agent import get_biographer_agent
from src.agents.biographer.tools import BIOGRAPHER_TOOLS from src.agents.biographer.tools import BIOGRAPHER_TOOLS
from src.core.household_registry import ( from src.core.household_registry import (
+10 -5
View File
@@ -10,6 +10,7 @@ These tools enable The Biographer to record and recall the user's story:
For direct key-based access (get/set profile, preferences), For direct key-based access (get/set profile, preferences),
use memory_service directly - these tools are for semantic queries. use memory_service directly - these tools are for semantic queries.
""" """
from src.core.context import get_user from src.core.context import get_user
from src.core.embeddings import get_embedding_client from src.core.embeddings import get_embedding_client
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
@@ -23,6 +24,7 @@ logger = get_logger(__name__)
# Semantic Recall # Semantic Recall
# ============================================================================ # ============================================================================
async def recall_semantic( async def recall_semantic(
query: str, query: str,
memory_type: str = "", memory_type: str = "",
@@ -108,6 +110,7 @@ async def recall_semantic(
# Store Memory # Store Memory
# ============================================================================ # ============================================================================
async def store_insight( async def store_insight(
key: str, key: str,
value: str, value: str,
@@ -158,7 +161,7 @@ async def store_insight(
f"**Keywords:** {', '.join(keywords)}", f"**Keywords:** {', '.join(keywords)}",
f"**Importance:** {importance:.1f}", f"**Importance:** {importance:.1f}",
"", "",
"_Memory is now searchable via semantic recall._" "_Memory is now searchable via semantic recall._",
] ]
logger.info( logger.info(
@@ -216,7 +219,7 @@ async def update_profile(
"## Profile Updated", "## Profile Updated",
f"**{key}:** {value}", f"**{key}:** {value}",
"", "",
"_Profile data is automatically included in context._" "_Profile data is automatically included in context._",
] ]
logger.info( logger.info(
@@ -271,7 +274,7 @@ async def update_preference(
"## Preference Updated", "## Preference Updated",
f"**{key}:** {value}", f"**{key}:** {value}",
"", "",
"_Preference will be applied to future responses._" "_Preference will be applied to future responses._",
] ]
logger.info( logger.info(
@@ -293,6 +296,7 @@ async def update_preference(
# List Memories # List Memories
# ============================================================================ # ============================================================================
async def list_memories( async def list_memories(
memory_type: str = "learned_fact", memory_type: str = "learned_fact",
limit: int = 20, limit: int = 20,
@@ -321,7 +325,7 @@ async def list_memories(
# Convert string to MemoryType # Convert string to MemoryType
try: try:
mem_type = MemoryType(memory_type) MemoryType(memory_type) # validated for its ValueError; the value is unused
except ValueError: except ValueError:
return f"Invalid memory type '{memory_type}'. Use: user_profile, preference, or learned_fact" return f"Invalid memory type '{memory_type}'. Use: user_profile, preference, or learned_fact"
@@ -378,6 +382,7 @@ async def list_memories(
# Forget Memory # Forget Memory
# ============================================================================ # ============================================================================
async def forget_memory( async def forget_memory(
key: str, key: str,
memory_type: str = "learned_fact", memory_type: str = "learned_fact",
@@ -420,7 +425,7 @@ async def forget_memory(
f"**Key:** {key}", f"**Key:** {key}",
f"**Type:** {memory_type}", f"**Type:** {memory_type}",
"", "",
"_Memory has been removed._" "_Memory has been removed._",
] ]
logger.info( logger.info(
+14 -11
View File
@@ -8,6 +8,7 @@ returns a structured result for synthesis.
This implements the agent-as-tool pattern recommended by PydanticAI: This implements the agent-as-tool pattern recommended by PydanticAI:
agents call other agents via tool wrappers, keeping each agent focused. agents call other agents via tool wrappers, keeping each agent focused.
""" """
import asyncio import asyncio
from dataclasses import dataclass, field from dataclasses import dataclass, field
from enum import Enum from enum import Enum
@@ -23,6 +24,7 @@ logger = get_logger(__name__)
# Action Types for Think Slug Selection # Action Types for Think Slug Selection
# ============================================================================= # =============================================================================
class ActionType(Enum): class ActionType(Enum):
""" """
Categories of actions for selecting appropriate think messages. Categories of actions for selecting appropriate think messages.
@@ -30,11 +32,12 @@ class ActionType(Enum):
Each expert has different action types that warrant different Each expert has different action types that warrant different
butler-perspective messages to the user. butler-perspective messages to the user.
""" """
RETRIEVE = "retrieve" # Looking up existing information
RESEARCH = "research" # Conducting new research (web search, etc.) RETRIEVE = "retrieve" # Looking up existing information
CREATE = "create" # Creating new content (pages, notes) RESEARCH = "research" # Conducting new research (web search, etc.)
CONTROL = "control" # Controlling devices/automations CREATE = "create" # Creating new content (pages, notes)
RECORD = "record" # Recording memories/notes CONTROL = "control" # Controlling devices/automations
RECORD = "record" # Recording memories/notes
# ============================================================================= # =============================================================================
@@ -159,8 +162,7 @@ def build_delegation_context(
if isinstance(content, list): if isinstance(content, list):
# Tolerate structured content parts # Tolerate structured content parts
content = " ".join( content = " ".join(
part.get("text", "") if isinstance(part, dict) else str(part) part.get("text", "") if isinstance(part, dict) else str(part) for part in content
for part in content
) )
content = str(content).strip() content = str(content).strip()
if content: if content:
@@ -206,6 +208,7 @@ class DelegationTask:
depends_on: List of task IDs this task depends on depends_on: List of task IDs this task depends on
result: Result from expert after execution result: Result from expert after execution
""" """
expert_name: str expert_name: str
task: str task: str
context: str = "" context: str = ""
@@ -215,10 +218,11 @@ class DelegationTask:
result: str | None = None result: str | None = None
task_id: str = "" task_id: str = ""
def __post_init__(self): def __post_init__(self) -> None:
"""Generate task ID if not provided.""" """Generate task ID if not provided."""
if not self.task_id: if not self.task_id:
import uuid import uuid
self.task_id = f"{self.expert_name}_{uuid.uuid4().hex[:8]}" self.task_id = f"{self.expert_name}_{uuid.uuid4().hex[:8]}"
@@ -236,6 +240,7 @@ class DelegationResult:
error: Short user-safe error label if failed. Exception detail error: Short user-safe error label if failed. Exception detail
stays in the logs only stays in the logs only
""" """
expert_name: str expert_name: str
task: str task: str
success: bool success: bool
@@ -335,9 +340,7 @@ async def delegate_to_librarian(
if span: if span:
span.metadata["success"] = False span.metadata["success"] = False
span.details["error"] = ( span.details["error"] = f"timed out after {config.LIBRARIAN_TIMEOUT}s"
f"timed out after {config.LIBRARIAN_TIMEOUT}s"
)
return DelegationResult( return DelegationResult(
expert_name="librarian", expert_name="librarian",
+1
View File
@@ -4,6 +4,7 @@ The Housekeeper - Home Automation Agent.
Provides home automation capabilities through the core-api service, Provides home automation capabilities through the core-api service,
which wraps the Home Assistant REST API into LLM-friendly endpoints. which wraps the Home Assistant REST API into LLM-friendly endpoints.
""" """
from src.agents.housekeeper.agent import run_housekeeper, run_housekeeper_stream from src.agents.housekeeper.agent import run_housekeeper, run_housekeeper_stream
from src.agents.housekeeper.capability import ( from src.agents.housekeeper.capability import (
HOUSEKEEPER_CAPABILITY, HOUSEKEEPER_CAPABILITY,
+6 -5
View File
@@ -8,7 +8,8 @@ the core-api service, which wraps Home Assistant REST API, offering:
- Script execution - Script execution
- Automation management - Automation management
""" """
from typing import Any, Optional
from typing import Any
from pydantic_ai import Agent from pydantic_ai import Agent
@@ -27,7 +28,6 @@ from src.agents.housekeeper.tools import (
turn_off, turn_off,
turn_on, turn_on,
) )
from src.core.config import config
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -98,7 +98,7 @@ After completing actions, briefly confirm:
""" """
# Lazy initialization to avoid connection issues during imports # Lazy initialization to avoid connection issues during imports
_housekeeper_agent: Optional[Agent[None, str]] = None _housekeeper_agent: Agent[None, str] | None = None
def _create_housekeeper_agent() -> Agent[None, str]: def _create_housekeeper_agent() -> Agent[None, str]:
@@ -140,6 +140,7 @@ def _create_housekeeper_agent() -> Agent[None, str]:
agent.tool_plain(get_history) agent.tool_plain(get_history)
from src.anthropic.model_selector import get_model_info from src.anthropic.model_selector import get_model_info
model_info = get_model_info() model_info = get_model_info()
logger.info( logger.info(
"housekeeper_agent_created", "housekeeper_agent_created",
@@ -167,7 +168,7 @@ def get_housekeeper_agent() -> Agent[None, str]:
async def run_housekeeper( async def run_housekeeper(
task: str, task: str,
context: str = "", context: str = "",
message_history: Optional[list[Any]] = None, message_history: list[Any] | None = None,
) -> str: ) -> str:
""" """
Execute a home automation task with The Housekeeper. Execute a home automation task with The Housekeeper.
@@ -234,7 +235,7 @@ async def run_housekeeper(
async def run_housekeeper_stream( async def run_housekeeper_stream(
task: str, task: str,
context: str = "", context: str = "",
message_history: Optional[list[Any]] = None, message_history: list[Any] | None = None,
): ):
""" """
Execute a home automation task with streaming output. Execute a home automation task with streaming output.
+1
View File
@@ -4,6 +4,7 @@ Housekeeper capability registration for the Household Registry.
Defines The Housekeeper's capabilities and registers it as a Defines The Housekeeper's capabilities and registers it as a
household member for coordination by the Steward and Tatlock. household member for coordination by the Steward and Tatlock.
""" """
from src.agents.housekeeper.agent import get_housekeeper_agent from src.agents.housekeeper.agent import get_housekeeper_agent
from src.agents.housekeeper.tools import HOUSEKEEPER_TOOLS from src.agents.housekeeper.tools import HOUSEKEEPER_TOOLS
from src.core.household_registry import ( from src.core.household_registry import (
+19 -18
View File
@@ -5,7 +5,8 @@ Provides async methods for home automation operations via Home Assistant.
Core-API is a separate service that wraps the Home Assistant REST API Core-API is a separate service that wraps the Home Assistant REST API
into LLM-friendly endpoints. into LLM-friendly endpoints.
""" """
from typing import Any, Optional
from typing import Any
import httpx import httpx
from pydantic import BaseModel, Field from pydantic import BaseModel, Field
@@ -28,7 +29,7 @@ class Device(BaseModel):
name: str name: str
state: str state: str
domain: str domain: str
area: Optional[str] = None area: str | None = None
attributes: dict[str, Any] = Field(default_factory=dict) attributes: dict[str, Any] = Field(default_factory=dict)
@@ -38,8 +39,8 @@ class DeviceState(BaseModel):
entity_id: str entity_id: str
state: str state: str
attributes: dict[str, Any] = Field(default_factory=dict) attributes: dict[str, Any] = Field(default_factory=dict)
last_changed: Optional[str] = None last_changed: str | None = None
last_updated: Optional[str] = None last_updated: str | None = None
class Scene(BaseModel): class Scene(BaseModel):
@@ -47,7 +48,7 @@ class Scene(BaseModel):
entity_id: str entity_id: str
name: str name: str
friendly_name: Optional[str] = None friendly_name: str | None = None
class Script(BaseModel): class Script(BaseModel):
@@ -55,8 +56,8 @@ class Script(BaseModel):
entity_id: str entity_id: str
name: str name: str
description: Optional[str] = None description: str | None = None
last_triggered: Optional[str] = None last_triggered: str | None = None
class Automation(BaseModel): class Automation(BaseModel):
@@ -64,8 +65,8 @@ class Automation(BaseModel):
entity_id: str entity_id: str
name: str name: str
state: str = "on" state: str = "on"
last_triggered: Optional[str] = None last_triggered: str | None = None
class HistoryEntry(BaseModel): class HistoryEntry(BaseModel):
@@ -109,8 +110,8 @@ class CoreAPIClient:
def __init__( def __init__(
self, self,
base_url: Optional[str] = None, base_url: str | None = None,
api_key: Optional[str] = None, api_key: str | None = None,
timeout: int = 30, timeout: int = 30,
): ):
""" """
@@ -124,7 +125,7 @@ class CoreAPIClient:
self.base_url = base_url or str(config.CORE_API_HOST) self.base_url = base_url or str(config.CORE_API_HOST)
self.api_key = api_key or config.CORE_API_KEY self.api_key = api_key or config.CORE_API_KEY
self.timeout = timeout self.timeout = timeout
self._client: Optional[httpx.AsyncClient] = None self._client: httpx.AsyncClient | None = None
async def __aenter__(self) -> "CoreAPIClient": async def __aenter__(self) -> "CoreAPIClient":
"""Create HTTP client on context entry.""" """Create HTTP client on context entry."""
@@ -159,8 +160,8 @@ class CoreAPIClient:
async def list_devices( async def list_devices(
self, self,
domain: Optional[str] = None, domain: str | None = None,
area: Optional[str] = None, area: str | None = None,
) -> list[Device]: ) -> list[Device]:
""" """
List devices, optionally filtered by domain or area. List devices, optionally filtered by domain or area.
@@ -231,9 +232,9 @@ class CoreAPIClient:
async def turn_on( async def turn_on(
self, self,
entity_id: str, entity_id: str,
brightness: Optional[int] = None, brightness: int | None = None,
color_temp: Optional[int] = None, color_temp: int | None = None,
rgb_color: Optional[tuple[int, int, int]] = None, rgb_color: tuple[int, int, int] | None = None,
) -> ControlResult: ) -> ControlResult:
""" """
Turn on a device. Turn on a device.
@@ -399,7 +400,7 @@ class CoreAPIClient:
async def run_script( async def run_script(
self, self,
script_id: str, script_id: str,
variables: Optional[dict[str, Any]] = None, variables: dict[str, Any] | None = None,
) -> ControlResult: ) -> ControlResult:
""" """
Run a script. Run a script.
+14 -3
View File
@@ -4,6 +4,7 @@ Housekeeper tools for PydanticAI agent.
These tools wrap the core-api service and are registered with These tools wrap the core-api service and are registered with
The Housekeeper agent for home automation tasks. The Housekeeper agent for home automation tasks.
""" """
from src.agents.housekeeper.client import CoreAPIClient from src.agents.housekeeper.client import CoreAPIClient
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
@@ -72,14 +73,24 @@ async def list_devices(
return True return True
return False return False
sorted_devices = sorted(dom_devices, key=lambda d: (not is_room_group(d), d.entity_id)) sorted_devices = sorted(
dom_devices, key=lambda d: (not is_room_group(d), d.entity_id)
)
for device in sorted_devices: for device in sorted_devices:
state_icon = "on" if device.state == "on" else "off" if device.state == "off" else device.state state_icon = (
"on"
if device.state == "on"
else "off"
if device.state == "off"
else device.state
)
area_str = f" ({device.area})" if device.area else "" area_str = f" ({device.area})" if device.area else ""
# Mark room groups clearly using actual HA data # Mark room groups clearly using actual HA data
group_marker = " [ROOM GROUP]" if is_room_group(device) else "" group_marker = " [ROOM GROUP]" if is_room_group(device) else ""
output_parts.append(f"- **{device.name}**{area_str}{group_marker}: {state_icon}") output_parts.append(
f"- **{device.name}**{area_str}{group_marker}: {state_icon}"
)
output_parts.append(f" ID: `{device.entity_id}`") output_parts.append(f" ID: `{device.entity_id}`")
output_parts.append("") output_parts.append("")
+1
View File
@@ -7,6 +7,7 @@ Connects to the library-desk API to provide:
- Knowledge graph queries - Knowledge graph queries
- Semantic search - Semantic search
""" """
from src.agents.librarian.agent import ( from src.agents.librarian.agent import (
get_librarian_agent, get_librarian_agent,
run_librarian, run_librarian,
+3 -3
View File
@@ -7,6 +7,7 @@ the library-desk API, offering:
- Wiki and document management - Wiki and document management
- Semantic search and knowledge graph exploration - Semantic search and knowledge graph exploration
""" """
from typing import Any from typing import Any
from pydantic_ai import Agent from pydantic_ai import Agent
@@ -202,6 +203,7 @@ def _create_librarian_agent() -> Agent[None, str]:
agent.tool_plain(smart_create_wiki_page) agent.tool_plain(smart_create_wiki_page)
from src.anthropic.model_selector import get_model_info from src.anthropic.model_selector import get_model_info
model_info = get_model_info() model_info = get_model_info()
logger.info( logger.info(
"librarian_agent_created", "librarian_agent_created",
@@ -294,6 +296,4 @@ async def run_librarian(
error=str(e), error=str(e),
exc_info=True, exc_info=True,
) )
raise AgentError( raise AgentError("Research task failed", agent_name="librarian") from e
"Research task failed", agent_name="librarian"
) from e
+1
View File
@@ -4,6 +4,7 @@ Librarian capability registration for the Household Registry.
Defines The Librarian's capabilities and registers it as a Defines The Librarian's capabilities and registers it as a
household member for coordination by the Steward and Tatlock. household member for coordination by the Steward and Tatlock.
""" """
from src.agents.librarian.agent import get_librarian_agent from src.agents.librarian.agent import get_librarian_agent
from src.agents.librarian.tools import LIBRARIAN_TOOLS from src.agents.librarian.tools import LIBRARIAN_TOOLS
from src.core.household_registry import ( from src.core.household_registry import (
+33 -12
View File
@@ -7,6 +7,7 @@ Provides async methods for all relevant library-desk endpoints:
- Vector search - Vector search
- Knowledge graph queries - Knowledge graph queries
""" """
import asyncio import asyncio
from collections.abc import AsyncIterator, Awaitable, Callable from collections.abc import AsyncIterator, Awaitable, Callable
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
@@ -38,8 +39,10 @@ _shared_http_client: ContextVar[httpx.AsyncClient | None] = ContextVar(
# Response Models # Response Models
# ============================================================================ # ============================================================================
class WikiPage(BaseModel): class WikiPage(BaseModel):
"""Wiki page from library-desk.""" """Wiki page from library-desk."""
id: int id: int
path: str path: str
title: str title: str
@@ -52,6 +55,7 @@ class WikiPage(BaseModel):
class WikiSearchResult(BaseModel): class WikiSearchResult(BaseModel):
"""Search result from wiki search.""" """Search result from wiki search."""
id: int id: int
path: str path: str
title: str title: str
@@ -61,6 +65,7 @@ class WikiSearchResult(BaseModel):
class VectorSearchResult(BaseModel): class VectorSearchResult(BaseModel):
"""Result from semantic vector search.""" """Result from semantic vector search."""
page_id: int page_id: int
page_path: str page_path: str
page_title: str page_title: str
@@ -71,8 +76,11 @@ class VectorSearchResult(BaseModel):
class HybridSearchResult(BaseModel): class HybridSearchResult(BaseModel):
"""Result from HybridRAG search.""" """Result from HybridRAG search."""
source: str # source_type: "wiki", "web", "volatile", "document" source: str # source_type: "wiki", "web", "volatile", "document"
sources: list[str] = Field(default_factory=list) # legs that found it: "vector", "graph", "web", ... sources: list[str] = Field(
default_factory=list
) # legs that found it: "vector", "graph", "web", ...
title: str title: str
content: str content: str
url: str | None = None url: str | None = None
@@ -84,6 +92,7 @@ class HybridSearchResult(BaseModel):
class HybridRAGResponse(BaseModel): class HybridRAGResponse(BaseModel):
"""Full response from HybridRAG query.""" """Full response from HybridRAG query."""
results: list[HybridSearchResult] = Field(default_factory=list) results: list[HybridSearchResult] = Field(default_factory=list)
keywords: list[str] = Field(default_factory=list) keywords: list[str] = Field(default_factory=list)
synonyms: list[str] = Field(default_factory=list) synonyms: list[str] = Field(default_factory=list)
@@ -102,6 +111,7 @@ class HybridRAGResponse(BaseModel):
class GraphNode(BaseModel): class GraphNode(BaseModel):
"""Node from knowledge graph.""" """Node from knowledge graph."""
id: str id: str
labels: list[str] = Field(default_factory=list) labels: list[str] = Field(default_factory=list)
properties: dict[str, Any] = Field(default_factory=dict) properties: dict[str, Any] = Field(default_factory=dict)
@@ -109,12 +119,14 @@ class GraphNode(BaseModel):
class Dossier(BaseModel): class Dossier(BaseModel):
"""A dossier (tag-based collection).""" """A dossier (tag-based collection)."""
name: str name: str
page_count: int page_count: int
class ResearchSummary(BaseModel): class ResearchSummary(BaseModel):
"""Summary of research performed during smart-create.""" """Summary of research performed during smart-create."""
wiki_results: int = 0 wiki_results: int = 0
web_results: int = 0 web_results: int = 0
graph_entities: int = 0 graph_entities: int = 0
@@ -124,6 +136,7 @@ class ResearchSummary(BaseModel):
class WebSearchResult(BaseModel): class WebSearchResult(BaseModel):
"""Result from web search via /rag/search.""" """Result from web search via /rag/search."""
title: str title: str
url: str url: str
content: str = "" # Full extracted text via Trafilatura content: str = "" # Full extracted text via Trafilatura
@@ -134,6 +147,7 @@ class WebSearchResult(BaseModel):
class WebSearchResponse(BaseModel): class WebSearchResponse(BaseModel):
"""Response from /rag/search endpoint.""" """Response from /rag/search endpoint."""
query: str query: str
search_type: str search_type: str
results: list[WebSearchResult] = Field(default_factory=list) results: list[WebSearchResult] = Field(default_factory=list)
@@ -144,6 +158,7 @@ class WebSearchResponse(BaseModel):
class ContentExtractionResult(BaseModel): class ContentExtractionResult(BaseModel):
"""Result from content extraction.""" """Result from content extraction."""
url: str url: str
title: str | None = None title: str | None = None
content: str = "" content: str = ""
@@ -156,6 +171,7 @@ class ContentExtractionResult(BaseModel):
class BatchExtractionResponse(BaseModel): class BatchExtractionResponse(BaseModel):
"""Response from batch content extraction.""" """Response from batch content extraction."""
results: list[ContentExtractionResult] = Field(default_factory=list) results: list[ContentExtractionResult] = Field(default_factory=list)
total_urls: int = 0 total_urls: int = 0
successful: int = 0 successful: int = 0
@@ -165,6 +181,7 @@ class BatchExtractionResponse(BaseModel):
class EntityLinking(BaseModel): class EntityLinking(BaseModel):
"""Entity linking results from smart-create.""" """Entity linking results from smart-create."""
forward_links: int = 0 forward_links: int = 0
backward_links: int = 0 backward_links: int = 0
pages_updated: int = 0 pages_updated: int = 0
@@ -172,6 +189,7 @@ class EntityLinking(BaseModel):
class SmartCreateResponse(BaseModel): class SmartCreateResponse(BaseModel):
"""Response from smart-create wiki page endpoint.""" """Response from smart-create wiki page endpoint."""
page: WikiPage page: WikiPage
research_summary: ResearchSummary = Field(default_factory=ResearchSummary) research_summary: ResearchSummary = Field(default_factory=ResearchSummary)
sources_used: int = 0 sources_used: int = 0
@@ -183,6 +201,7 @@ class SmartCreateResponse(BaseModel):
# Client # Client
# ============================================================================ # ============================================================================
class LibraryDeskClient: class LibraryDeskClient:
""" """
Async HTTP client for Library-Desk API. Async HTTP client for Library-Desk API.
@@ -398,17 +417,19 @@ class LibraryDeskClient:
# related_dossiers; older names kept as fallbacks) # related_dossiers; older names kept as fallbacks)
results = [] results = []
for r in data.get("results", []): for r in data.get("results", []):
results.append(HybridSearchResult( results.append(
source=r.get("source_type") or r.get("source", "unknown"), HybridSearchResult(
sources=r.get("sources", []), source=r.get("source_type") or r.get("source", "unknown"),
title=r.get("title", ""), sources=r.get("sources", []),
content=r.get("content", ""), title=r.get("title", ""),
url=r.get("url"), content=r.get("content", ""),
score=r.get("rrf_score", r.get("score", 0.0)), url=r.get("url"),
page_id=r.get("page_id"), score=r.get("rrf_score", r.get("score", 0.0)),
related_dossiers=r.get("related_dossiers", []), page_id=r.get("page_id"),
metadata=r.get("metadata", {}), related_dossiers=r.get("related_dossiers", []),
)) metadata=r.get("metadata", {}),
)
)
# Handle keywords being either a list or a dict with core_keywords; # Handle keywords being either a list or a dict with core_keywords;
# the live service nests synonyms inside the keywords dict as a # the live service nests synonyms inside the keywords dict as a
+21 -20
View File
@@ -4,6 +4,7 @@ Librarian tools for PydanticAI agent.
These tools wrap the library-desk API and are registered with These tools wrap the library-desk API and are registered with
The Librarian agent for research and knowledge management tasks. The Librarian agent for research and knowledge management tasks.
""" """
import httpx import httpx
from pydantic_ai import ModelRetry from pydantic_ai import ModelRetry
@@ -27,9 +28,8 @@ def _retry_if_transient(e: Exception, what: str) -> None:
status = e.response.status_code status = e.response.status_code
retryable = status >= 500 or status == 429 retryable = status >= 500 or status == 429
if retryable: if retryable:
raise ModelRetry( raise ModelRetry(f"{what} is temporarily unavailable; please retry.") from e
f"{what} is temporarily unavailable; please retry."
) from e
# Icons keyed by the values library-desk emits in each result's `sources` # Icons keyed by the values library-desk emits in each result's `sources`
# list (search legs) and `source_type` (result origin). # list (search legs) and `source_type` (result origin).
@@ -64,11 +64,7 @@ def _coverage_note(
results, so their absence is normal ranking behavior, not an outage. results, so their absence is normal ranking behavior, not an outage.
""" """
if response.source_status: if response.source_status:
failed = sorted( failed = sorted(leg for leg, status in response.source_status.items() if status == "failed")
leg
for leg, status in response.source_status.items()
if status == "failed"
)
if failed: if failed:
return ( return (
"⚠️ *Coverage note: results are partial - " "⚠️ *Coverage note: results are partial - "
@@ -111,6 +107,7 @@ def _coverage_note(
# HybridRAG Search # HybridRAG Search
# ============================================================================ # ============================================================================
async def hybrid_search( async def hybrid_search(
query: str, query: str,
include_web: bool = True, include_web: bool = True,
@@ -165,9 +162,7 @@ async def hybrid_search(
# Add related dossiers # Add related dossiers
if response.related_dossiers: if response.related_dossiers:
output_parts.append( output_parts.append(f"**Related Dossiers:** {', '.join(response.related_dossiers)}")
f"**Related Dossiers:** {', '.join(response.related_dossiers)}"
)
output_parts.append("") output_parts.append("")
@@ -216,6 +211,7 @@ async def hybrid_search(
# Wiki Operations # Wiki Operations
# ============================================================================ # ============================================================================
async def search_wiki( async def search_wiki(
query: str, query: str,
limit: int = 10, limit: int = 10,
@@ -331,9 +327,7 @@ async def list_dossiers() -> str:
output_parts = ["## Research Dossiers\n"] output_parts = ["## Research Dossiers\n"]
for dossier in dossiers: for dossier in dossiers:
output_parts.append( output_parts.append(f"- **{dossier.name}** ({dossier.page_count} pages)")
f"- **{dossier.name}** ({dossier.page_count} pages)"
)
return "\n".join(output_parts) return "\n".join(output_parts)
@@ -389,6 +383,7 @@ async def get_dossier_pages(
# Semantic Search # Semantic Search
# ============================================================================ # ============================================================================
async def semantic_search( async def semantic_search(
query: str, query: str,
limit: int = 10, limit: int = 10,
@@ -420,9 +415,7 @@ async def semantic_search(
output_parts = [f"## Semantic Search: {query}\n"] output_parts = [f"## Semantic Search: {query}\n"]
for i, result in enumerate(results, 1): for i, result in enumerate(results, 1):
output_parts.append( output_parts.append(f"{i}. **{result.page_title}** (score: {result.score:.2f})")
f"{i}. **{result.page_title}** (score: {result.score:.2f})"
)
output_parts.append(f" Path: {result.page_path}") output_parts.append(f" Path: {result.page_path}")
output_parts.append(f" {result.chunk_text[:200]}...") output_parts.append(f" {result.chunk_text[:200]}...")
output_parts.append("") output_parts.append("")
@@ -439,6 +432,7 @@ async def semantic_search(
# Knowledge Graph # Knowledge Graph
# ============================================================================ # ============================================================================
async def explore_knowledge_graph( async def explore_knowledge_graph(
entity_type: str = "Document", entity_type: str = "Document",
limit: int = 20, limit: int = 20,
@@ -565,6 +559,7 @@ async def find_related_entities(
# Web Search & Content Extraction # Web Search & Content Extraction
# ============================================================================ # ============================================================================
async def search_web( async def search_web(
query: str, query: str,
limit: int = 10, limit: int = 10,
@@ -605,7 +600,9 @@ async def search_web(
return f"No results found for '{query}'" return f"No results found for '{query}'"
output_parts = [f"## Web Search: {query}\n"] output_parts = [f"## Web Search: {query}\n"]
output_parts.append(f"*Found {response.total_results} results in {response.search_time_ms}ms*\n") output_parts.append(
f"*Found {response.total_results} results in {response.search_time_ms}ms*\n"
)
for i, result in enumerate(response.results, 1): for i, result in enumerate(response.results, 1):
output_parts.append(f"### {i}. {result.title}") output_parts.append(f"### {i}. {result.title}")
@@ -880,7 +877,9 @@ async def update_wiki_page(
if page.tags: if page.tags:
output_parts.append(f"**Tags:** {', '.join(page.tags)}") output_parts.append(f"**Tags:** {', '.join(page.tags)}")
output_parts.append("\n*Vector embeddings and knowledge graph will be updated automatically.*") output_parts.append(
"\n*Vector embeddings and knowledge graph will be updated automatically.*"
)
logger.info( logger.info(
"librarian_update_page", "librarian_update_page",
@@ -954,7 +953,9 @@ async def create_wiki_page(
if page.description: if page.description:
output_parts.append(f"**Description:** {page.description}") output_parts.append(f"**Description:** {page.description}")
output_parts.append("\n*Vector embeddings and knowledge graph will be updated automatically.*") output_parts.append(
"\n*Vector embeddings and knowledge graph will be updated automatically.*"
)
logger.info( logger.info(
"librarian_create_page", "librarian_create_page",
+25 -48
View File
@@ -12,16 +12,16 @@ infrastructure is real production code.
import asyncio import asyncio
import random import random
import secrets import secrets
from typing import AsyncGenerator, Any from collections.abc import AsyncGenerator
from typing import Any
from src.agents.base import AgentInterface, OutputItem from src.agents.base import AgentInterface, OutputItem
from src.core.exceptions import ( from src.core.exceptions import (
RateLimitError,
ContextLengthError,
APIError, APIError,
ContextLengthError,
RateLimitError,
) )
# Mock lorem ipsum content # Mock lorem ipsum content
LOREM_PARAGRAPHS = [ LOREM_PARAGRAPHS = [
"Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.", "Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.",
@@ -46,33 +46,27 @@ MOCK_TOOLS = [
"description": "Search the knowledge base for relevant information", "description": "Search the knowledge base for relevant information",
"parameters": { "parameters": {
"type": "object", "type": "object",
"properties": { "properties": {"query": {"type": "string", "description": "Search query"}},
"query": {"type": "string", "description": "Search query"} "required": ["query"],
}, },
"required": ["query"]
}
}, },
{ {
"name": "calculate", "name": "calculate",
"description": "Perform mathematical calculations", "description": "Perform mathematical calculations",
"parameters": { "parameters": {
"type": "object", "type": "object",
"properties": { "properties": {"expression": {"type": "string", "description": "Math expression"}},
"expression": {"type": "string", "description": "Math expression"} "required": ["expression"],
}, },
"required": ["expression"]
}
}, },
{ {
"name": "get_weather", "name": "get_weather",
"description": "Get current weather for a location", "description": "Get current weather for a location",
"parameters": { "parameters": {
"type": "object", "type": "object",
"properties": { "properties": {"location": {"type": "string", "description": "City name"}},
"location": {"type": "string", "description": "City name"} "required": ["location"],
}, },
"required": ["location"]
}
}, },
] ]
@@ -109,7 +103,7 @@ class LoremTesterAgent(AgentInterface):
temperature: float = 1.0, temperature: float = 1.0,
max_tokens: int | None = None, max_tokens: int | None = None,
stop: list[str] | None = None, stop: list[str] | None = None,
**kwargs: Any **kwargs: Any,
) -> AsyncGenerator[OutputItem, None]: ) -> AsyncGenerator[OutputItem, None]:
""" """
Generate mock response with reasoning, tools, and content. Generate mock response with reasoning, tools, and content.
@@ -123,8 +117,7 @@ class LoremTesterAgent(AgentInterface):
# 1. Yield reasoning item if requested # 1. Yield reasoning item if requested
if reasoning and reasoning.get("summary") == "auto": if reasoning and reasoning.get("summary") == "auto":
yield await self._create_reasoning_item( yield await self._create_reasoning_item(
messages, messages, effort=reasoning.get("effort", "medium")
effort=reasoning.get("effort", "medium")
) )
# 2. Randomly yield function calls if tools available (30% chance) # 2. Randomly yield function calls if tools available (30% chance)
@@ -150,7 +143,7 @@ class LoremTesterAgent(AgentInterface):
"reasoning": True, "reasoning": True,
"tools": True, "tools": True,
"vision": False, # Not yet "vision": False, # Not yet
"audio": False, # Not yet "audio": False, # Not yet
} }
# Private helper methods # Private helper methods
@@ -179,9 +172,7 @@ class LoremTesterAgent(AgentInterface):
raise APIError("Invalid tool call: tool 'nonexistent' not found (mock trigger)") raise APIError("Invalid tool call: tool 'nonexistent' not found (mock trigger)")
async def _create_reasoning_item( async def _create_reasoning_item(
self, self, messages: list[dict], effort: str = "medium"
messages: list[dict],
effort: str = "medium"
) -> OutputItem: ) -> OutputItem:
"""Create a reasoning output item with mock thinking steps.""" """Create a reasoning output item with mock thinking steps."""
@@ -200,23 +191,17 @@ class LoremTesterAgent(AgentInterface):
steps = random.sample(REASONING_STEPS, min(num_steps, len(REASONING_STEPS))) steps = random.sample(REASONING_STEPS, min(num_steps, len(REASONING_STEPS)))
return OutputItem( return OutputItem(
type="reasoning", type="reasoning", id=f"rs_{generate_id()}", summary=steps, status="completed"
id=f"rs_{generate_id()}",
summary=steps,
status="completed"
) )
async def _create_tool_calls( async def _create_tool_calls(self, tools: list[dict]) -> AsyncGenerator[OutputItem, None]:
self,
tools: list[dict]
) -> AsyncGenerator[OutputItem, None]:
"""Create mock function call output items.""" """Create mock function call output items."""
# Randomly select 1-2 tools to "call" # Randomly select 1-2 tools to "call"
num_calls = random.randint(1, 2) num_calls = random.randint(1, 2)
selected_tools = random.sample( selected_tools = random.sample(
MOCK_TOOLS[:min(len(MOCK_TOOLS), len(tools))], MOCK_TOOLS[: min(len(MOCK_TOOLS), len(tools))],
min(num_calls, len(MOCK_TOOLS), len(tools)) min(num_calls, len(MOCK_TOOLS), len(tools)),
) )
for tool in selected_tools: for tool in selected_tools:
@@ -228,7 +213,7 @@ class LoremTesterAgent(AgentInterface):
id=f"fc_{generate_id()}", id=f"fc_{generate_id()}",
name=tool["name"], name=tool["name"],
arguments=args, arguments=args,
status="completed" status="completed",
) )
def _generate_mock_args(self, tool: dict) -> str: def _generate_mock_args(self, tool: dict) -> str:
@@ -254,11 +239,7 @@ class LoremTesterAgent(AgentInterface):
# Generic mock arguments # Generic mock arguments
return json.dumps({"input": "mock_value"}) return json.dumps({"input": "mock_value"})
async def _create_message_item( async def _create_message_item(self, messages: list[dict], temperature: float) -> OutputItem:
self,
messages: list[dict],
temperature: float
) -> OutputItem:
"""Create final message output item with lorem ipsum content.""" """Create final message output item with lorem ipsum content."""
# Select random lorem ipsum paragraphs # Select random lorem ipsum paragraphs
@@ -270,10 +251,6 @@ class LoremTesterAgent(AgentInterface):
type="message", type="message",
id=f"msg_{generate_id()}", id=f"msg_{generate_id()}",
role="assistant", role="assistant",
content=[{ content=[{"type": "output_text", "text": content, "annotations": []}],
"type": "output_text", status="completed",
"text": content,
"annotations": []
}],
status="completed"
) )
+20 -11
View File
@@ -14,12 +14,13 @@ Supports:
- Result aggregation from multiple experts - Result aggregation from multiple experts
- Partial failure handling - Partial failure handling
""" """
import asyncio import asyncio
from collections.abc import AsyncGenerator
from dataclasses import dataclass, field from dataclasses import dataclass, field
from enum import Enum from enum import Enum
from typing import AsyncGenerator, Optional, Callable, Any
from src.agents.delegation import DelegationTask, DelegationResult, delegate_to_librarian from src.agents.delegation import DelegationResult, DelegationTask, delegate_to_librarian
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -27,8 +28,9 @@ logger = get_logger(__name__)
class ExecutionMode(str, Enum): class ExecutionMode(str, Enum):
"""Execution mode for multi-expert coordination.""" """Execution mode for multi-expert coordination."""
SEQUENTIAL = "sequential" # One at a time, in order SEQUENTIAL = "sequential" # One at a time, in order
PARALLEL = "parallel" # All at once, concurrently PARALLEL = "parallel" # All at once, concurrently
@dataclass @dataclass
@@ -38,12 +40,13 @@ class OrchestrationContext:
Tracks the user's request, delegation tasks, and results. Tracks the user's request, delegation tasks, and results.
""" """
user_message: str user_message: str
steward_note: str steward_note: str
conversation_id: Optional[str] = None conversation_id: str | None = None
def parse_delegation_from_steward_note(steward_note: str) -> Optional[DelegationTask]: def parse_delegation_from_steward_note(steward_note: str) -> DelegationTask | None:
""" """
Parse a delegation task from Steward's note. Parse a delegation task from Steward's note.
@@ -68,9 +71,9 @@ def parse_delegation_from_steward_note(steward_note: str) -> Optional[Delegation
# Look for DELEGATE: pattern # Look for DELEGATE: pattern
# Match: "DELEGATE: expert_name to action description" # Match: "DELEGATE: expert_name to action description"
match = re.search( match = re.search(
r'DELEGATE:\s*(\w+)\s+to\s+(.+?)(?:\n|REASON:|COMPLEXITY:|CONTEXT:|$)', r"DELEGATE:\s*(\w+)\s+to\s+(.+?)(?:\n|REASON:|COMPLEXITY:|CONTEXT:|$)",
steward_note, steward_note,
re.IGNORECASE | re.MULTILINE re.IGNORECASE | re.MULTILINE,
) )
if match: if match:
@@ -135,7 +138,7 @@ async def execute_delegation(
async def orchestrate_with_think_updates( async def orchestrate_with_think_updates(
user_message: str, user_message: str,
steward_note: str, steward_note: str,
delegation_task: Optional[DelegationTask] = None, delegation_task: DelegationTask | None = None,
) -> AsyncGenerator[str, None]: ) -> AsyncGenerator[str, None]:
""" """
Orchestrate expert delegation with streaming think updates. Orchestrate expert delegation with streaming think updates.
@@ -218,17 +221,21 @@ def extract_delegation_context(
} }
# Extract REASON: # Extract REASON:
reason_match = re.search(r'REASON:\s*(.+?)(?:\n|COMPLEXITY:|CONTEXT:|$)', steward_note, re.IGNORECASE) reason_match = re.search(
r"REASON:\s*(.+?)(?:\n|COMPLEXITY:|CONTEXT:|$)", steward_note, re.IGNORECASE
)
if reason_match: if reason_match:
result["reason"] = reason_match.group(1).strip() result["reason"] = reason_match.group(1).strip()
# Extract COMPLEXITY: # Extract COMPLEXITY:
complexity_match = re.search(r'COMPLEXITY:\s*(.+?)(?:\n|CONTEXT:|$)', steward_note, re.IGNORECASE) complexity_match = re.search(
r"COMPLEXITY:\s*(.+?)(?:\n|CONTEXT:|$)", steward_note, re.IGNORECASE
)
if complexity_match: if complexity_match:
result["complexity"] = complexity_match.group(1).strip() result["complexity"] = complexity_match.group(1).strip()
# Extract CONTEXT: # Extract CONTEXT:
context_match = re.search(r'CONTEXT:\s*(.+?)$', steward_note, re.IGNORECASE | re.MULTILINE) context_match = re.search(r"CONTEXT:\s*(.+?)$", steward_note, re.IGNORECASE | re.MULTILINE)
if context_match: if context_match:
result["context"] = context_match.group(1).strip() result["context"] = context_match.group(1).strip()
@@ -239,6 +246,7 @@ def extract_delegation_context(
# Multi-Expert Coordination # Multi-Expert Coordination
# ============================================================================ # ============================================================================
@dataclass @dataclass
class MultiExpertResult: class MultiExpertResult:
""" """
@@ -250,6 +258,7 @@ class MultiExpertResult:
failed_experts: List of expert names that failed failed_experts: List of expert names that failed
combined_output: Aggregated output from all successful experts combined_output: Aggregated output from all successful experts
""" """
results: dict[str, DelegationResult] = field(default_factory=dict) results: dict[str, DelegationResult] = field(default_factory=dict)
all_succeeded: bool = True all_succeeded: bool = True
failed_experts: list[str] = field(default_factory=list) failed_experts: list[str] = field(default_factory=list)
+11 -12
View File
@@ -8,9 +8,6 @@ It provides a central place to:
- Check model capabilities - Check model capabilities
""" """
import time
from typing import Type
from src.agents.base import AgentInterface from src.agents.base import AgentInterface
from src.agents.lorem_tester import LoremTesterAgent from src.agents.lorem_tester import LoremTesterAgent
from src.agents.tatlock import TatlockAgent from src.agents.tatlock import TatlockAgent
@@ -63,7 +60,7 @@ class ModelRegistry:
if model_id not in cls.MODELS: if model_id not in cls.MODELS:
raise ModelNotFoundError(model_id) raise ModelNotFoundError(model_id)
agent_class: Type[AgentInterface] = cls.MODELS[model_id]["agent_class"] agent_class: type[AgentInterface] = cls.MODELS[model_id]["agent_class"]
return agent_class() return agent_class()
@classmethod @classmethod
@@ -106,14 +103,16 @@ class ModelRegistry:
agent = cls.get_agent(model_id) agent = cls.get_agent(model_id)
capabilities = await agent.get_capabilities() capabilities = await agent.get_capabilities()
models.append({ models.append(
"id": model_id, {
"object": "model", "id": model_id,
"created": config["created"], "object": "model",
"owned_by": config["owned_by"], "created": config["created"],
"capabilities": capabilities, "owned_by": config["owned_by"],
"description": config["description"], "capabilities": capabilities,
}) "description": config["description"],
}
)
return models return models
+1
View File
@@ -4,6 +4,7 @@ Steward agent package.
The Steward analyzes incoming requests and recommends relevant household The Steward analyzes incoming requests and recommends relevant household
capabilities, creating a two-tier architecture with the Butler. capabilities, creating a two-tier architecture with the Butler.
""" """
from .agent import StewardAgent, get_steward_agent from .agent import StewardAgent, get_steward_agent
from .schemas import ConversationContext, StewardRecommendation from .schemas import ConversationContext, StewardRecommendation
from .service import analyze_request, format_steward_note from .service import analyze_request, format_steward_note
+9 -14
View File
@@ -8,8 +8,8 @@ This creates a two-tier architecture that prevents cognitive overload.
Uses plain text output (not JSON) for reliability. Supports both Claude Uses plain text output (not JSON) for reliability. Supports both Claude
(preferred) and Ollama (fallback) backends via direct API calls. (preferred) and Ollama (fallback) backends via direct API calls.
""" """
import httpx import httpx
from typing import Optional
from src.anthropic.model_selector import get_model_info, is_claude_available, resolve_backend from src.anthropic.model_selector import get_model_info, is_claude_available, resolve_backend
from src.core.config import config from src.core.config import config
@@ -29,9 +29,7 @@ def build_steward_prompt(query: str, conversation_history: list[dict]) -> str:
cap_list = [] cap_list = []
for cap in capabilities: for cap in capabilities:
cap_list.append( cap_list.append(f"{cap.name} - {cap.description} (domains: {', '.join(cap.domains)})")
f"{cap.name} - {cap.description} (domains: {', '.join(cap.domains)})"
)
capabilities_text = "\n".join(cap_list) capabilities_text = "\n".join(cap_list)
# Format conversation history if present # Format conversation history if present
@@ -111,10 +109,10 @@ class StewardAgent:
(preferred) and Ollama (fallback) backends via direct API calls. (preferred) and Ollama (fallback) backends via direct API calls.
""" """
def __init__(self): def __init__(self) -> None:
"""Initialize Steward with backend selection based on availability.""" """Initialize Steward with backend selection based on availability."""
# Ollama config (primary) # Ollama config (primary)
self.ollama_host = str(config.OLLAMA_HOST).rstrip('/') self.ollama_host = str(config.OLLAMA_HOST).rstrip("/")
self.ollama_model = config.OLLAMA_DEFAULT_MODEL self.ollama_model = config.OLLAMA_DEFAULT_MODEL
# Claude config (fallback) # Claude config (fallback)
@@ -139,6 +137,7 @@ class StewardAgent:
"""Get or create Anthropic client (lazy initialization).""" """Get or create Anthropic client (lazy initialization)."""
if self._anthropic_client is None: if self._anthropic_client is None:
from anthropic import AsyncAnthropic from anthropic import AsyncAnthropic
self._anthropic_client = AsyncAnthropic(api_key=config.ANTHROPIC_API_KEY) self._anthropic_client = AsyncAnthropic(api_key=config.ANTHROPIC_API_KEY)
return self._anthropic_client return self._anthropic_client
@@ -167,20 +166,16 @@ class StewardAgent:
"stream": False, "stream": False,
"options": { "options": {
"temperature": 0.3, # Lower = more consistent "temperature": 0.3, # Lower = more consistent
"top_p": 0.9 "top_p": 0.9,
} },
} },
) )
response.raise_for_status() response.raise_for_status()
result = response.json() result = response.json()
return result["response"].strip() return result["response"].strip()
async def analyze( async def analyze(self, query: str, conversation_history: list[dict] | None = None) -> str:
self,
query: str,
conversation_history: Optional[list[dict]] = None
) -> str:
""" """
Analyze query and return plain text recommendation. Analyze query and return plain text recommendation.
+14 -13
View File
@@ -4,7 +4,8 @@ Steward agent schemas.
Defines the structured output models for Steward's request analysis Defines the structured output models for Steward's request analysis
and capability recommendations. and capability recommendations.
""" """
from typing import Any, Literal, Optional
from typing import Any, Literal
from pydantic import BaseModel, Field from pydantic import BaseModel, Field
@@ -16,16 +17,16 @@ class ConversationContext(BaseModel):
The Steward analyzes the full conversation to identify references The Steward analyzes the full conversation to identify references
to previous topics, helping the Butler maintain context. to previous topics, helping the Butler maintain context.
""" """
has_previous_context: bool = Field( has_previous_context: bool = Field(
description="Whether the current request references previous conversation turns" description="Whether the current request references previous conversation turns"
) )
relevant_turns: list[int] = Field( relevant_turns: list[int] = Field(
default_factory=list, default_factory=list,
description="0-indexed turn numbers that are relevant to the current request" description="0-indexed turn numbers that are relevant to the current request",
) )
context_summary: str = Field( context_summary: str = Field(
default="", default="", description="Brief summary of relevant context for the Butler"
description="Brief summary of relevant context for the Butler"
) )
@@ -40,29 +41,28 @@ class StewardRecommendation(BaseModel):
- Conversation context - Conversation context
- Missing capabilities (if any) - Missing capabilities (if any)
""" """
recommended_capabilities: list[str] = Field( recommended_capabilities: list[str] = Field(
description="List of household member names to include (e.g., ['tatlock_core'])" description="List of household member names to include (e.g., ['tatlock_core'])"
) )
reasoning: str = Field( reasoning: str = Field(description="Explanation of why these capabilities were recommended")
description="Explanation of why these capabilities were recommended"
)
estimated_complexity: Literal["simple", "moderate", "complex"] = Field( estimated_complexity: Literal["simple", "moderate", "complex"] = Field(
description="Complexity assessment: simple (1 tool), moderate (2-3 tools), complex (multiple tools/steps)" description="Complexity assessment: simple (1 tool), moderate (2-3 tools), complex (multiple tools/steps)"
) )
conversation_context: ConversationContext = Field( conversation_context: ConversationContext = Field(
description="Contextual information from conversation history" description="Contextual information from conversation history"
) )
missing_capabilities: Optional[str] = Field( missing_capabilities: str | None = Field(
default=None, default=None,
description="Description of capabilities that would be helpful but aren't available" description="Description of capabilities that would be helpful but aren't available",
) )
memory_context: dict[str, Any] = Field( memory_context: dict[str, Any] = Field(
default_factory=dict, default_factory=dict,
description="Pre-fetched user context from memory (profile, preferences)" description="Pre-fetched user context from memory (profile, preferences)",
) )
enriched_query: str = Field( enriched_query: str = Field(
default="", default="",
description="User query with auto-filled context (location, timezone) when not specified" description="User query with auto-filled context (location, timezone) when not specified",
) )
def format_for_butler(self) -> str: def format_for_butler(self) -> str:
@@ -114,8 +114,9 @@ class StewardRecommendation(BaseModel):
lines.append(f" • preferences: {prefs_str}") lines.append(f" • preferences: {prefs_str}")
# Add delegation instructions when expert agents are recommended # Add delegation instructions when expert agents are recommended
delegation_agents = [c for c in self.recommended_capabilities delegation_agents = [
if c in ("biographer", "librarian")] c for c in self.recommended_capabilities if c in ("biographer", "librarian")
]
if delegation_agents: if delegation_agents:
lines.append("-" * 40) lines.append("-" * 40)
lines.append("DELEGATION REQUIRED:") lines.append("DELEGATION REQUIRED:")
+150 -59
View File
@@ -7,47 +7,102 @@ and error handling.
Parses plain text recommendations into structured data. Parses plain text recommendations into structured data.
Includes memory pre-fetch for user context injection. Includes memory pre-fetch for user context injection.
""" """
import re import re
from typing import Any, Optional from typing import Any
from src.core.household_registry import get_household_registry from src.core.household_registry import get_household_registry
from src.core.logging_config import get_logger, log_operation from src.core.logging_config import get_logger, log_operation
from src.core.memory_service import memory_service from src.core.memory_service import memory_service
from .agent import get_steward_agent from .agent import get_steward_agent
from .schemas import ConversationContext, StewardRecommendation from .schemas import ConversationContext, StewardRecommendation
logger = get_logger(__name__) logger = get_logger(__name__)
_DELEGATE_LINE_RE = re.compile(r"^[ \t]*DELEGATE:[ \t]*(.+)$", re.IGNORECASE | re.MULTILINE)
def _mentions(needle: str, haystack: str) -> bool:
"""Whole-word containment. Substring matching is what made this go wrong."""
return re.search(rf"(?<!\w){re.escape(needle)}(?!\w)", haystack) is not None
def _extract_capabilities(text: str) -> list[str]: def _extract_capabilities(text: str) -> list[str]:
""" """
Extract capability names from Steward's text response. Extract capability names from the Steward's declared delegation.
Uses keyword matching to find mentioned capabilities. The prompt instructs the Steward to answer in a fixed shape::
DELEGATE: <capability> to <action> <task>
REASON: ...
COMPLEXITY: ...
CONTEXT: ...
Only the DELEGATE line states intent; the rest is free prose. An earlier
version substring-matched capability *domains* across the whole response,
which routed on ordinary English: "description" contains "script" and
"discover" contains "cover" (both housekeeper domains), "acknowledge"
contains "knowledge" and "know" (librarian, biographer), and "economy"
contains "my" (biographer). Any REASON line could therefore summon agents
the Steward never asked for, and a spurious librarian is a real
multi-second web call.
It also made prose length a routing input, so anything that shortened the
Steward's output — such as disabling model thinking — would look like it had
improved routing.
Resolution is layered, most explicit first:
1. a DELEGATE line beginning with a capability name the documented shape
2. a capability named anywhere on a DELEGATE line
3. a capability *domain* on a DELEGATE line, for a loosely worded answer
4. no DELEGATE line: capability names only, never domains
Args: Args:
text: Steward's plain text analysis text: Steward's plain text analysis
Returns: Returns:
List of capability names (e.g., ['tatlock_core']) List of capability names (e.g. ['tatlock_core']), de-duplicated.
""" """
text_lower = text.lower()
registry = get_household_registry() registry = get_household_registry()
capabilities = registry.get_all_capabilities() capabilities = registry.get_all_capabilities()
delegate_lines = [line.strip().lower() for line in _DELEGATE_LINE_RE.findall(text or "")]
found_caps = [] found_caps: list[str] = []
for cap in capabilities: def _add(name: str) -> None:
# Check if capability name is mentioned if name not in found_caps:
if cap.name.lower() in text_lower: found_caps.append(name)
found_caps.append(cap.name)
if not delegate_lines:
# Either the Steward judged no capability necessary — the prompt's
# conversational path, whose correct answer is [] — or it ignored the
# format. Names only: domain words are ordinary English and would fire
# on any prose, which is the bug described above.
haystack = (text or "").lower()
for cap in capabilities:
if _mentions(cap.name.lower(), haystack):
_add(cap.name)
return found_caps
for line in delegate_lines:
leading = next((c for c in capabilities if line.startswith(c.name.lower())), None)
if leading is not None:
_add(leading.name)
continue continue
# Check if any domains are mentioned named = [c for c in capabilities if _mentions(c.name.lower(), line)]
for domain in cap.domains: if named:
if domain.lower() in text_lower: for cap in named:
found_caps.append(cap.name) _add(cap.name)
break continue
# Last resort. Scoped to this line, so the REASON and CONTEXT prose that
# caused the original misrouting can no longer reach it.
for cap in capabilities:
if any(_mentions(domain.lower(), line) for domain in cap.domains):
_add(cap.name)
return found_caps return found_caps
@@ -73,8 +128,7 @@ def _extract_complexity(text: str) -> str:
def _extract_conversation_context( def _extract_conversation_context(
text: str, text: str, conversation_history: list[dict]
conversation_history: list[dict]
) -> ConversationContext: ) -> ConversationContext:
""" """
Extract conversation context analysis from text. Extract conversation context analysis from text.
@@ -89,13 +143,15 @@ def _extract_conversation_context(
text_lower = text.lower() text_lower = text.lower()
# Check if conversation history is referenced # Check if conversation history is referenced
has_context = bool(conversation_history) and any([ has_context = bool(conversation_history) and any(
"previous" in text_lower, [
"earlier" in text_lower, "previous" in text_lower,
"context" in text_lower, "earlier" in text_lower,
"turn" in text_lower, "context" in text_lower,
"history" in text_lower, "turn" in text_lower,
]) "history" in text_lower,
]
)
# Extract turn numbers if mentioned (e.g., "turn 0", "turn 1") # Extract turn numbers if mentioned (e.g., "turn 0", "turn 1")
relevant_turns = [] relevant_turns = []
@@ -107,21 +163,23 @@ def _extract_conversation_context(
context_summary = "" context_summary = ""
if has_context: if has_context:
# Extract sentence(s) mentioning context # Extract sentence(s) mentioning context
sentences = text.split('.') sentences = text.split(".")
context_sentences = [s for s in sentences if any( context_sentences = [
word in s.lower() for word in ["previous", "earlier", "context", "history"] s
)] for s in sentences
if any(word in s.lower() for word in ["previous", "earlier", "context", "history"])
]
if context_sentences: if context_sentences:
context_summary = context_sentences[0].strip() context_summary = context_sentences[0].strip()
return ConversationContext( return ConversationContext(
has_previous_context=has_context, has_previous_context=has_context,
relevant_turns=relevant_turns, relevant_turns=relevant_turns,
context_summary=context_summary context_summary=context_summary,
) )
def _extract_missing_capabilities(text: str) -> Optional[str]: def _extract_missing_capabilities(text: str) -> str | None:
""" """
Extract missing capability notes from text. Extract missing capability notes from text.
@@ -134,15 +192,16 @@ def _extract_missing_capabilities(text: str) -> Optional[str]:
text_lower = text.lower() text_lower = text.lower()
# Look for indicators of missing capabilities # Look for indicators of missing capabilities
if any(word in text_lower for word in [ if any(
"missing", "unavailable", "not available", "don't have", "doesn't have" word in text_lower
]): for word in ["missing", "unavailable", "not available", "don't have", "doesn't have"]
):
# Find the sentence mentioning missing capabilities # Find the sentence mentioning missing capabilities
sentences = text.split('.') sentences = text.split(".")
for sentence in sentences: for sentence in sentences:
if any(word in sentence.lower() for word in [ if any(
"missing", "unavailable", "not available" word in sentence.lower() for word in ["missing", "unavailable", "not available"]
]): ):
return sentence.strip() return sentence.strip()
return None return None
@@ -183,7 +242,7 @@ def _build_enriched_query(user_request: str, memory_context: dict[str, Any]) ->
# Check if location is needed and not specified # Check if location is needed and not specified
location_keywords = ["weather", "temperature", "forecast", "nearby", "local", "here"] location_keywords = ["weather", "temperature", "forecast", "nearby", "local", "here"]
# Use word boundary pattern to avoid false positives like "at" in "what" # Use word boundary pattern to avoid false positives like "at" in "what"
location_prepositions = [r'\bin\b', r'\bat\b', r'\bnear\b', r'\baround\b', r'\bfor\b'] location_prepositions = [r"\bin\b", r"\bat\b", r"\bnear\b", r"\baround\b", r"\bfor\b"]
location_specified = any(re.search(p, request_lower) for p in location_prepositions) location_specified = any(re.search(p, request_lower) for p in location_prepositions)
if any(word in request_lower for word in location_keywords): if any(word in request_lower for word in location_keywords):
@@ -234,32 +293,65 @@ async def _prefetch_memory_context(user_request: str) -> dict[str, Any]:
profile_keys = [] profile_keys = []
# Location-related queries # Location-related queries
if any(word in request_lower for word in [ if any(
"weather", "temperature", "forecast", "nearby", "local", word in request_lower
"directions", "distance", "map", "here", for word in [
# Direct location questions "weather",
"live", "where", "home", "reside", "location", "address", "temperature",
]): "forecast",
"nearby",
"local",
"directions",
"distance",
"map",
"here",
# Direct location questions
"live",
"where",
"home",
"reside",
"location",
"address",
]
):
profile_keys.append("location") profile_keys.append("location")
# Time-related queries # Time-related queries
if any(word in request_lower for word in [ if any(
"time", "schedule", "meeting", "appointment", "reminder", word in request_lower
"alarm", "when", "today", "tomorrow" for word in [
]): "time",
"schedule",
"meeting",
"appointment",
"reminder",
"alarm",
"when",
"today",
"tomorrow",
]
):
profile_keys.append("timezone") profile_keys.append("timezone")
# Personal queries # Personal queries
if any(word in request_lower for word in [ if any(word in request_lower for word in ["my name", "who am i", "about me"]):
"my name", "who am i", "about me"
]):
profile_keys.append("name") profile_keys.append("name")
# Always fetch preferences if they might affect response format # Always fetch preferences if they might affect response format
include_preferences = any(word in request_lower for word in [ include_preferences = any(
"temperature", "weather", "convert", "unit", "format", word in request_lower
"celsius", "fahrenheit", "metric", "imperial" for word in [
]) "temperature",
"weather",
"convert",
"unit",
"format",
"celsius",
"fahrenheit",
"metric",
"imperial",
]
)
try: try:
return await memory_service.prefetch_context( return await memory_service.prefetch_context(
@@ -278,7 +370,7 @@ async def _prefetch_memory_context(user_request: str) -> dict[str, Any]:
async def analyze_request( async def analyze_request(
user_request: str, user_request: str,
conversation_history: list[dict], conversation_history: list[dict],
conversation_id: Optional[str] = None, conversation_id: str | None = None,
) -> StewardRecommendation: ) -> StewardRecommendation:
""" """
Analyze user request with full conversation context. Analyze user request with full conversation context.
@@ -310,7 +402,7 @@ async def analyze_request(
"request_preview": user_request[:100], "request_preview": user_request[:100],
"conversation_id": conversation_id, "conversation_id": conversation_id,
"history_length": len(conversation_history), "history_length": len(conversation_history),
} },
) as log_ctx: ) as log_ctx:
try: try:
# Pre-fetch user context from memory (fast, no LLM) # Pre-fetch user context from memory (fast, no LLM)
@@ -329,8 +421,7 @@ async def analyze_request(
# Get plain text analysis from Steward # Get plain text analysis from Steward
analysis_text = await steward.analyze( analysis_text = await steward.analyze(
user_request, user_request, conversation_history=conversation_history
conversation_history=conversation_history
) )
# Parse plain text into structured recommendation # Parse plain text into structured recommendation
+115 -85
View File
@@ -6,24 +6,25 @@ The agent embodies a witty, capable British butler personality.
""" """
import secrets import secrets
from typing import AsyncGenerator, Any from collections.abc import AsyncGenerator
from dataclasses import dataclass, field from dataclasses import dataclass, field
from typing import Any
from pydantic_ai import Agent, RunContext from pydantic_ai import Agent, RunContext
from src.agents.base import AgentInterface, OutputItem from src.agents.base import AgentInterface, OutputItem
from src.agents.tatlock_core.tools import ( from src.agents.tatlock_core.tools import (
calculate, calculate,
get_current_datetime,
calculate_time_offset, calculate_time_offset,
get_current_datetime,
time_difference, time_difference,
) )
from src.core.config import config
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
from src.core.tracing import ( from src.core.tracing import (
start_span, end_span, get_current_span, SpanType,
add_tool_spans_from_messages, add_tool_spans_from_messages,
SpanType, SpanStatus, end_span,
start_span,
) )
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -32,9 +33,10 @@ logger = get_logger(__name__)
@dataclass @dataclass
class ToolCallTracker: class ToolCallTracker:
"""Tracks tool calls for reporting to reasoning output.""" """Tracks tool calls for reporting to reasoning output."""
calls: list[str] = field(default_factory=list) calls: list[str] = field(default_factory=list)
def log_call(self, message: str): def log_call(self, message: str) -> None:
"""Log a tool call.""" """Log a tool call."""
self.calls.append(message) self.calls.append(message)
@@ -153,11 +155,14 @@ class TatlockAgent(AgentInterface):
currently in Phase 1 (basic LLM integration without expert agents). currently in Phase 1 (basic LLM integration without expert agents).
""" """
def __init__(self): def __init__(self) -> None:
"""Initialize Tatlock (lazy agent creation).""" """Initialize Tatlock (lazy agent creation)."""
self._agent = None # Lazy initialization # Deps are a ToolCallTracker: every registered tool takes
# RunContext[ToolCallTracker], and run() is called with one. Saying so
# is what lets the tool registrations below type-check at all.
self._agent: Agent[ToolCallTracker, str] | None = None # Lazy initialization
def _ensure_agent(self): def _ensure_agent(self) -> None:
"""Ensure the PydanticAI agent is initialized (lazy initialization).""" """Ensure the PydanticAI agent is initialized (lazy initialization)."""
if self._agent is not None: if self._agent is not None:
return return
@@ -178,13 +183,19 @@ class TatlockAgent(AgentInterface):
self._agent = Agent( self._agent = Agent(
model, model,
system_prompt=TATLOCK_SYSTEM_PROMPT, system_prompt=TATLOCK_SYSTEM_PROMPT,
deps_type=ToolCallTracker,
) )
# Register tools with the agent # Register tools with the agent
self._register_tools() self._register_tools()
def _register_tools(self): def _register_tools(self) -> None:
"""Register permanent tools with the PydanticAI agent.""" """Register permanent tools with the PydanticAI agent.
Called only from _ensure_agent, immediately after the agent is built, so
the assert documents an invariant rather than guarding a real case.
"""
assert self._agent is not None, "_register_tools called before the agent exists"
# Calculator tool # Calculator tool
@self._agent.tool @self._agent.tool
@@ -239,7 +250,9 @@ class TatlockAgent(AgentInterface):
# Time difference calculator # Time difference calculator
@self._agent.tool @self._agent.tool
def calculate_time_difference(ctx: RunContext[ToolCallTracker], date1_str: str, date2_str: str = "now") -> str: def calculate_time_difference(
ctx: RunContext[ToolCallTracker], date1_str: str, date2_str: str = "now"
) -> str:
""" """
Calculate the difference between two dates. Calculate the difference between two dates.
@@ -251,7 +264,9 @@ class TatlockAgent(AgentInterface):
Human-readable description of the time difference Human-readable description of the time difference
""" """
if ctx.deps: if ctx.deps:
ctx.deps.log_call(f"🕐 Calculating time difference between {date1_str} and {date2_str}") ctx.deps.log_call(
f"🕐 Calculating time difference between {date1_str} and {date2_str}"
)
return time_difference(date1_str, date2_str) return time_difference(date1_str, date2_str)
# NOTE: Web search has been moved to The Librarian agent. # NOTE: Web search has been moved to The Librarian agent.
@@ -271,7 +286,7 @@ class TatlockAgent(AgentInterface):
temperature: float = 1.0, temperature: float = 1.0,
max_tokens: int | None = None, max_tokens: int | None = None,
stop: list[str] | None = None, stop: list[str] | None = None,
**kwargs: Any **kwargs: Any,
) -> AsyncGenerator[OutputItem, None]: ) -> AsyncGenerator[OutputItem, None]:
""" """
Generate response using PydanticAI with Ollama. Generate response using PydanticAI with Ollama.
@@ -305,20 +320,28 @@ class TatlockAgent(AgentInterface):
type="message", type="message",
id=f"msg_{generate_id()}", id=f"msg_{generate_id()}",
role="assistant", role="assistant",
content=[{ content=[
"type": "output_text", {
"text": "I'm afraid I didn't receive a message, sir. How may I assist you?", "type": "output_text",
"annotations": [] "text": "I'm afraid I didn't receive a message, sir. How may I assist you?",
}], "annotations": [],
status="completed" }
],
status="completed",
) )
return return
# Build message history (all messages except the last user message) # Build message history (all messages except the last user message)
# PydanticAI expects history as list of ModelRequest/ModelResponse objects # PydanticAI expects history as list of ModelRequest/ModelResponse objects
from pydantic_ai.messages import ModelRequest, ModelResponse, UserPromptPart, TextPart from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
TextPart,
UserPromptPart,
)
message_history = [] message_history: list[ModelMessage] = []
for i, msg in enumerate(messages[:-1]): # All messages except the last one for i, msg in enumerate(messages[:-1]): # All messages except the last one
role = msg.get("role") role = msg.get("role")
content = msg.get("content", "") content = msg.get("content", "")
@@ -330,7 +353,9 @@ class TatlockAgent(AgentInterface):
# Debug: Check for problematic content # Debug: Check for problematic content
if '"' in content or "'" in content: if '"' in content or "'" in content:
logger.debug(f"Message {i} ({role}) contains quotes. Content preview: {content[:100]}...") logger.debug(
f"Message {i} ({role}) contains quotes. Content preview: {content[:100]}..."
)
# Convert to PydanticAI message format # Convert to PydanticAI message format
try: try:
@@ -339,9 +364,7 @@ class TatlockAgent(AgentInterface):
ModelRequest(parts=[UserPromptPart(content=content)]) ModelRequest(parts=[UserPromptPart(content=content)])
) )
elif role == "assistant": elif role == "assistant":
message_history.append( message_history.append(ModelResponse(parts=[TextPart(content=content)]))
ModelResponse(parts=[TextPart(content=content)])
)
except Exception as e: except Exception as e:
logger.error(f"Error creating message history item {i}: {e}") logger.error(f"Error creating message history item {i}: {e}")
logger.error(f"Problematic content: {repr(content)}") logger.error(f"Problematic content: {repr(content)}")
@@ -352,7 +375,9 @@ class TatlockAgent(AgentInterface):
if message_history: if message_history:
for i, hist_msg in enumerate(message_history): for i, hist_msg in enumerate(message_history):
msg_type = type(hist_msg).__name__ msg_type = type(hist_msg).__name__
content_preview = str(hist_msg.parts[0].content)[:50] if hist_msg.parts else "no parts" content_preview = (
str(hist_msg.parts[0].content)[:50] if hist_msg.parts else "no parts"
)
logger.info(f" History[{i}]: {msg_type} - {content_preview}...") logger.info(f" History[{i}]: {msg_type} - {content_preview}...")
# Generate reasoning output if requested # Generate reasoning output if requested
@@ -362,10 +387,10 @@ class TatlockAgent(AgentInterface):
id=f"reasoning_{generate_id()}", id=f"reasoning_{generate_id()}",
summary=[ summary=[
"Analyzing your request, sir...", "Analyzing your request, sir...",
"Formulating response based on available knowledge..." "Formulating response based on available knowledge...",
], ],
thinking="", # PydanticAI doesn't expose internal reasoning yet thinking="", # PydanticAI doesn't expose internal reasoning yet
status="completed" status="completed",
) )
# Create a tool call tracker for this request # Create a tool call tracker for this request
@@ -382,7 +407,7 @@ class TatlockAgent(AgentInterface):
result = await self.agent.run( result = await self.agent.run(
user_message, user_message,
message_history=message_history if message_history else None, message_history=message_history if message_history else None,
deps=tracker deps=tracker,
) )
final_text = result.output final_text = result.output
@@ -393,7 +418,7 @@ class TatlockAgent(AgentInterface):
id=f"reasoning_tools_{generate_id()}", id=f"reasoning_tools_{generate_id()}",
summary=tracker.calls, summary=tracker.calls,
thinking="", thinking="",
status="completed" status="completed",
) )
# Yield the complete message # Yield the complete message
@@ -402,12 +427,8 @@ class TatlockAgent(AgentInterface):
type="message", type="message",
id=msg_id, id=msg_id,
role="assistant", role="assistant",
content=[{ content=[{"type": "output_text", "text": final_text, "annotations": []}],
"type": "output_text", status="completed",
"text": final_text,
"annotations": []
}],
status="completed"
) )
except Exception as e: except Exception as e:
@@ -416,12 +437,14 @@ class TatlockAgent(AgentInterface):
type="message", type="message",
id=f"msg_{generate_id()}", id=f"msg_{generate_id()}",
role="assistant", role="assistant",
content=[{ content=[
"type": "output_text", {
"text": f"My apologies, sir. I encountered an error: {str(e)}", "type": "output_text",
"annotations": [] "text": f"My apologies, sir. I encountered an error: {str(e)}",
}], "annotations": [],
status="failed" }
],
status="failed",
) )
async def supports_tools(self) -> bool: async def supports_tools(self) -> bool:
@@ -490,9 +513,15 @@ class TatlockAgent(AgentInterface):
enriched_message = f"{steward_note}\n\n{user_message}" enriched_message = f"{steward_note}\n\n{user_message}"
# Convert message history to PydanticAI format # Convert message history to PydanticAI format
from pydantic_ai.messages import ModelRequest, ModelResponse, UserPromptPart, TextPart from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
TextPart,
UserPromptPart,
)
pydantic_history = [] pydantic_history: list[ModelMessage] = []
for msg in message_history: for msg in message_history:
role = msg.get("role") role = msg.get("role")
content = msg.get("content", "") content = msg.get("content", "")
@@ -501,17 +530,14 @@ class TatlockAgent(AgentInterface):
continue continue
if role == "user": if role == "user":
pydantic_history.append( pydantic_history.append(ModelRequest(parts=[UserPromptPart(content=content)]))
ModelRequest(parts=[UserPromptPart(content=content)])
)
elif role == "assistant": elif role == "assistant":
pydantic_history.append( pydantic_history.append(ModelResponse(parts=[TextPart(content=content)]))
ModelResponse(parts=[TextPart(content=content)])
)
# Run with scoped tools and tracker # Run with scoped tools and tracker
# Force tool_choice to make LLM actually call tools # Force tool_choice to make LLM actually call tools
from src.anthropic.model_selector import get_tool_choice_settings from src.anthropic.model_selector import get_tool_choice_settings
result = await scoped_agent.run( result = await scoped_agent.run(
enriched_message, enriched_message,
message_history=pydantic_history if pydantic_history else None, message_history=pydantic_history if pydantic_history else None,
@@ -575,9 +601,15 @@ class TatlockAgent(AgentInterface):
enriched_message = f"{steward_note}\n\n{user_message}" enriched_message = f"{steward_note}\n\n{user_message}"
# Convert message history to PydanticAI format # Convert message history to PydanticAI format
from pydantic_ai.messages import ModelRequest, ModelResponse, UserPromptPart, TextPart from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
TextPart,
UserPromptPart,
)
pydantic_history = [] pydantic_history: list[ModelMessage] = []
for msg in message_history: for msg in message_history:
role = msg.get("role") role = msg.get("role")
content = msg.get("content", "") content = msg.get("content", "")
@@ -586,13 +618,9 @@ class TatlockAgent(AgentInterface):
continue continue
if role == "user": if role == "user":
pydantic_history.append( pydantic_history.append(ModelRequest(parts=[UserPromptPart(content=content)]))
ModelRequest(parts=[UserPromptPart(content=content)])
)
elif role == "assistant": elif role == "assistant":
pydantic_history.append( pydantic_history.append(ModelResponse(parts=[TextPart(content=content)]))
ModelResponse(parts=[TextPart(content=content)])
)
# Use run() instead of run_stream() to avoid Ollama 400 bug # Use run() instead of run_stream() to avoid Ollama 400 bug
# with streaming + tool calls (PydanticAI issues #1292, #2256) # with streaming + tool calls (PydanticAI issues #1292, #2256)
@@ -600,7 +628,7 @@ class TatlockAgent(AgentInterface):
result = await scoped_agent.run( result = await scoped_agent.run(
enriched_message, enriched_message,
message_history=pydantic_history if pydantic_history else None, message_history=pydantic_history if pydantic_history else None,
deps=tool_tracker deps=tool_tracker,
) )
# Stream the final response in chunks to maintain UX # Stream the final response in chunks to maintain UX
@@ -608,7 +636,7 @@ class TatlockAgent(AgentInterface):
chunk_size = 50 # characters per chunk chunk_size = 50 # characters per chunk
for i in range(0, len(response_text), chunk_size): for i in range(0, len(response_text), chunk_size):
yield response_text[i:i + chunk_size] yield response_text[i : i + chunk_size]
logger.info("tatlock_scoped_run_complete") logger.info("tatlock_scoped_run_complete")
@@ -641,13 +669,15 @@ class TatlockAgent(AgentInterface):
- raw_output: The agent's raw text output - raw_output: The agent's raw text output
""" """
from pydantic_ai.messages import ( from pydantic_ai.messages import (
ModelMessage,
ModelRequest, ModelRequest,
ModelResponse, ModelResponse,
UserPromptPart,
TextPart, TextPart,
ToolCallPart, ToolCallPart,
ToolReturnPart, ToolReturnPart,
UserPromptPart,
) )
from src.anthropic.model_selector import get_model from src.anthropic.model_selector import get_model
logger.info( logger.info(
@@ -663,7 +693,7 @@ class TatlockAgent(AgentInterface):
SpanType.TATLOCK, SpanType.TATLOCK,
metadata={ metadata={
"scoped_tool_count": len(scoped_tools), "scoped_tool_count": len(scoped_tools),
"tool_names": [getattr(t, '__name__', str(t)) for t in scoped_tools[:5]], "tool_names": [getattr(t, "__name__", str(t)) for t in scoped_tools[:5]],
}, },
) )
@@ -681,7 +711,7 @@ class TatlockAgent(AgentInterface):
enriched_message = f"{steward_note}\n\n{user_message}" enriched_message = f"{steward_note}\n\n{user_message}"
# Convert message history to PydanticAI format # Convert message history to PydanticAI format
pydantic_history = [] pydantic_history: list[ModelMessage] = []
for msg in message_history: for msg in message_history:
role = msg.get("role") role = msg.get("role")
content = msg.get("content", "") content = msg.get("content", "")
@@ -690,16 +720,13 @@ class TatlockAgent(AgentInterface):
continue continue
if role == "user": if role == "user":
pydantic_history.append( pydantic_history.append(ModelRequest(parts=[UserPromptPart(content=content)]))
ModelRequest(parts=[UserPromptPart(content=content)])
)
elif role == "assistant": elif role == "assistant":
pydantic_history.append( pydantic_history.append(ModelResponse(parts=[TextPart(content=content)]))
ModelResponse(parts=[TextPart(content=content)])
)
# Run with scoped tools and tracker # Run with scoped tools and tracker
from src.anthropic.model_selector import get_tool_choice_settings from src.anthropic.model_selector import get_tool_choice_settings
result = await scoped_agent.run( result = await scoped_agent.run(
enriched_message, enriched_message,
message_history=pydantic_history if pydantic_history else None, message_history=pydantic_history if pydantic_history else None,
@@ -709,8 +736,8 @@ class TatlockAgent(AgentInterface):
# Extract tool calls and results from the agent's messages # Extract tool calls and results from the agent's messages
tools_called = [] tools_called = []
expert_results = {} expert_results: dict[str, Any] = {}
tool_outputs = {} tool_outputs: dict[str, Any] = {}
# Parse through new messages to find tool calls and returns # Parse through new messages to find tool calls and returns
for msg in result.new_messages(): for msg in result.new_messages():
@@ -782,7 +809,14 @@ class TatlockAgent(AgentInterface):
Returns: Returns:
str: Butler-toned response synthesized from all results str: Butler-toned response synthesized from all results
""" """
from pydantic_ai.messages import ModelRequest, ModelResponse, UserPromptPart, TextPart from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
TextPart,
UserPromptPart,
)
from src.anthropic.model_selector import get_model from src.anthropic.model_selector import get_model
logger.info( logger.info(
@@ -840,7 +874,7 @@ class TatlockAgent(AgentInterface):
) )
# Convert message history to PydanticAI format # Convert message history to PydanticAI format
pydantic_history = [] pydantic_history: list[ModelMessage] = []
for msg in message_history: for msg in message_history:
role = msg.get("role") role = msg.get("role")
content = msg.get("content", "") content = msg.get("content", "")
@@ -849,13 +883,9 @@ class TatlockAgent(AgentInterface):
continue continue
if role == "user": if role == "user":
pydantic_history.append( pydantic_history.append(ModelRequest(parts=[UserPromptPart(content=content)]))
ModelRequest(parts=[UserPromptPart(content=content)])
)
elif role == "assistant": elif role == "assistant":
pydantic_history.append( pydantic_history.append(ModelResponse(parts=[TextPart(content=content)]))
ModelResponse(parts=[TextPart(content=content)])
)
# Run synthesis # Run synthesis
result = await synthesis_agent.run( result = await synthesis_agent.run(
@@ -885,9 +915,9 @@ class TatlockAgent(AgentInterface):
async def get_capabilities(self) -> dict: async def get_capabilities(self) -> dict:
"""Return current capabilities.""" """Return current capabilities."""
return { return {
"streaming": True, # Streaming implemented "streaming": True, # Streaming implemented
"reasoning": True, # Basic reasoning summaries "reasoning": True, # Basic reasoning summaries
"tools": True, # Permanent tools: calculator, date/time, search "tools": True, # Permanent tools: calculator, date/time, search
"vision": False, # Future "vision": False, # Future
"audio": False, # Future "audio": False, # Future
} }
+2 -1
View File
@@ -5,14 +5,15 @@ Provides calculator and date/time capabilities.
Web search has been moved to The Librarian agent. Web search has been moved to The Librarian agent.
Organized as a household member with toolset and capability registration. Organized as a household member with toolset and capability registration.
""" """
from .capability import TATLOCK_CORE_CAPABILITY, get_capability from .capability import TATLOCK_CORE_CAPABILITY, get_capability
from .toolset import get_core_tools, tatlock_core_tools
from .tools import ( from .tools import (
calculate, calculate,
calculate_time_offset, calculate_time_offset,
get_current_datetime, get_current_datetime,
time_difference, time_difference,
) )
from .toolset import get_core_tools, tatlock_core_tools
__all__ = [ __all__ = [
# Tools # Tools
+1 -1
View File
@@ -4,8 +4,8 @@ Household capability definition for Tatlock's core tools.
Provides the executive summary that the Steward and Butler see Provides the executive summary that the Steward and Butler see
for coordinating household capabilities. for coordinating household capabilities.
""" """
from src.core.household_registry import HouseholdCapability
from src.core.household_registry import HouseholdCapability
TATLOCK_CORE_CAPABILITY = HouseholdCapability( TATLOCK_CORE_CAPABILITY = HouseholdCapability(
name="tatlock_core", name="tatlock_core",
+23 -27
View File
@@ -6,13 +6,11 @@ These tools are always available to the butler agent:
- Date/Time toolkit: For current time and time calculations - Date/Time toolkit: For current time and time calculations
- SearXNG search: For searching the web for current information - SearXNG search: For searching the web for current information
""" """
import math import math
import re import re
from datetime import datetime, timedelta from datetime import datetime, timedelta
import httpx
from src.core.config import config
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -22,6 +20,7 @@ logger = get_logger(__name__)
# Calculator Tool # Calculator Tool
# ============================================================================ # ============================================================================
def calculate(expression: str) -> str: def calculate(expression: str) -> str:
""" """
Safely evaluate mathematical expressions. Safely evaluate mathematical expressions.
@@ -50,33 +49,29 @@ def calculate(expression: str) -> str:
# Create safe namespace with math functions # Create safe namespace with math functions
safe_dict = { safe_dict = {
# Basic math functions # Basic math functions
'sqrt': math.sqrt, "sqrt": math.sqrt,
'pow': math.pow, "pow": math.pow,
'abs': abs, "abs": abs,
'round': round, "round": round,
# Trigonometric # Trigonometric
'sin': math.sin, "sin": math.sin,
'cos': math.cos, "cos": math.cos,
'tan': math.tan, "tan": math.tan,
'asin': math.asin, "asin": math.asin,
'acos': math.acos, "acos": math.acos,
'atan': math.atan, "atan": math.atan,
# Logarithmic # Logarithmic
'log': math.log, "log": math.log,
'log10': math.log10, "log10": math.log10,
'log2': math.log2, "log2": math.log2,
'exp': math.exp, "exp": math.exp,
# Other # Other
'ceil': math.ceil, "ceil": math.ceil,
'floor': math.floor, "floor": math.floor,
'factorial': math.factorial, "factorial": math.factorial,
# Constants # Constants
'pi': math.pi, "pi": math.pi,
'e': math.e, "e": math.e,
} }
# Evaluate the expression safely # Evaluate the expression safely
@@ -101,6 +96,7 @@ def calculate(expression: str) -> str:
# Date/Time Toolkit # Date/Time Toolkit
# ============================================================================ # ============================================================================
def get_current_datetime(format_str: str = "full") -> str: def get_current_datetime(format_str: str = "full") -> str:
""" """
Get the current date and time. Get the current date and time.
@@ -161,7 +157,7 @@ def calculate_time_offset(offset_description: str) -> str:
# Parse the offset description # Parse the offset description
# Pattern: "N unit(s) ago/from now" # Pattern: "N unit(s) ago/from now"
pattern = r'(\d+)\s+(second|minute|hour|day|week|month|year)s?\s+(ago|from\s+now)' pattern = r"(\d+)\s+(second|minute|hour|day|week|month|year)s?\s+(ago|from\s+now)"
match = re.match(pattern, offset_description.lower().strip()) match = re.match(pattern, offset_description.lower().strip())
if not match: if not match:
+1 -1
View File
@@ -4,11 +4,11 @@ PydanticAI toolset for Tatlock's core tools.
Converts the core tool functions into PydanticAI tool definitions Converts the core tool functions into PydanticAI tool definitions
that can be registered with agents and the household registry. that can be registered with agents and the household registry.
""" """
from pydantic_ai.tools import Tool from pydantic_ai.tools import Tool
from . import tools from . import tools
# Create tool definitions for PydanticAI # Create tool definitions for PydanticAI
calculator_tool = Tool( calculator_tool = Tool(
function=tools.calculate, function=tools.calculate,
+22 -25
View File
@@ -13,11 +13,11 @@ import math
import re import re
from datetime import datetime, timedelta from datetime import datetime, timedelta
# ============================================================================ # ============================================================================
# Calculator Tool # Calculator Tool
# ============================================================================ # ============================================================================
def calculate(expression: str) -> str: def calculate(expression: str) -> str:
""" """
Safely evaluate mathematical expressions. Safely evaluate mathematical expressions.
@@ -46,33 +46,29 @@ def calculate(expression: str) -> str:
# Create safe namespace with math functions # Create safe namespace with math functions
safe_dict = { safe_dict = {
# Basic math functions # Basic math functions
'sqrt': math.sqrt, "sqrt": math.sqrt,
'pow': math.pow, "pow": math.pow,
'abs': abs, "abs": abs,
'round': round, "round": round,
# Trigonometric # Trigonometric
'sin': math.sin, "sin": math.sin,
'cos': math.cos, "cos": math.cos,
'tan': math.tan, "tan": math.tan,
'asin': math.asin, "asin": math.asin,
'acos': math.acos, "acos": math.acos,
'atan': math.atan, "atan": math.atan,
# Logarithmic # Logarithmic
'log': math.log, "log": math.log,
'log10': math.log10, "log10": math.log10,
'log2': math.log2, "log2": math.log2,
'exp': math.exp, "exp": math.exp,
# Other # Other
'ceil': math.ceil, "ceil": math.ceil,
'floor': math.floor, "floor": math.floor,
'factorial': math.factorial, "factorial": math.factorial,
# Constants # Constants
'pi': math.pi, "pi": math.pi,
'e': math.e, "e": math.e,
} }
# Evaluate the expression safely # Evaluate the expression safely
@@ -97,6 +93,7 @@ def calculate(expression: str) -> str:
# Date/Time Toolkit # Date/Time Toolkit
# ============================================================================ # ============================================================================
def get_current_datetime(format_str: str = "full") -> str: def get_current_datetime(format_str: str = "full") -> str:
""" """
Get the current date and time. Get the current date and time.
@@ -157,7 +154,7 @@ def calculate_time_offset(offset_description: str) -> str:
# Parse the offset description # Parse the offset description
# Pattern: "N unit(s) ago/from now" # Pattern: "N unit(s) ago/from now"
pattern = r'(\d+)\s+(second|minute|hour|day|week|month|year)s?\s+(ago|from\s+now)' pattern = r"(\d+)\s+(second|minute|hour|day|week|month|year)s?\s+(ago|from\s+now)"
match = re.match(pattern, offset_description.lower().strip()) match = re.match(pattern, offset_description.lower().strip())
if not match: if not match:
+2 -1
View File
@@ -2,9 +2,10 @@
Chat completion router. Chat completion router.
OpenAI-compatible /v1/chat/completions endpoint. OpenAI-compatible /v1/chat/completions endpoint.
""" """
import json import json
import logging import logging
from typing import AsyncGenerator from collections.abc import AsyncGenerator
from fastapi import APIRouter from fastapi import APIRouter
from starlette.responses import StreamingResponse from starlette.responses import StreamingResponse
+9
View File
@@ -2,6 +2,7 @@
OpenAI-compatible chat completion schemas. OpenAI-compatible chat completion schemas.
Following OpenAI API specification for compatibility. Following OpenAI API specification for compatibility.
""" """
from typing import Literal from typing import Literal
from pydantic import Field from pydantic import Field
@@ -11,6 +12,7 @@ from src.core.models import CustomBaseModel
class ChatMessage(CustomBaseModel): class ChatMessage(CustomBaseModel):
"""OpenAI-compatible chat message.""" """OpenAI-compatible chat message."""
role: Literal["system", "user", "assistant"] role: Literal["system", "user", "assistant"]
content: str content: str
name: str | None = None name: str | None = None
@@ -18,6 +20,7 @@ class ChatMessage(CustomBaseModel):
class ChatCompletionRequest(CustomBaseModel): class ChatCompletionRequest(CustomBaseModel):
"""OpenAI-compatible chat completion request.""" """OpenAI-compatible chat completion request."""
model: str = Field(..., description="Model to use for completion") model: str = Field(..., description="Model to use for completion")
messages: list[ChatMessage] = Field(..., description="List of messages") messages: list[ChatMessage] = Field(..., description="List of messages")
temperature: float | None = Field(default=0.7, ge=0.0, le=2.0) temperature: float | None = Field(default=0.7, ge=0.0, le=2.0)
@@ -29,6 +32,7 @@ class ChatCompletionRequest(CustomBaseModel):
class ChatCompletionChoice(CustomBaseModel): class ChatCompletionChoice(CustomBaseModel):
"""Choice in chat completion response.""" """Choice in chat completion response."""
index: int index: int
message: ChatMessage message: ChatMessage
finish_reason: str | None finish_reason: str | None
@@ -36,6 +40,7 @@ class ChatCompletionChoice(CustomBaseModel):
class ChatCompletionUsage(CustomBaseModel): class ChatCompletionUsage(CustomBaseModel):
"""Token usage information.""" """Token usage information."""
prompt_tokens: int prompt_tokens: int
completion_tokens: int completion_tokens: int
total_tokens: int total_tokens: int
@@ -43,6 +48,7 @@ class ChatCompletionUsage(CustomBaseModel):
class ChatCompletionResponse(CustomBaseModel): class ChatCompletionResponse(CustomBaseModel):
"""OpenAI-compatible chat completion response.""" """OpenAI-compatible chat completion response."""
id: str id: str
object: str = "chat.completion" object: str = "chat.completion"
created: int created: int
@@ -53,6 +59,7 @@ class ChatCompletionResponse(CustomBaseModel):
class ChatCompletionChunkDelta(CustomBaseModel): class ChatCompletionChunkDelta(CustomBaseModel):
"""Delta in streaming chunk.""" """Delta in streaming chunk."""
role: str | None = None role: str | None = None
content: str | None = None content: str | None = None
reasoning_content: str | None = None # For thinking/reasoning (DeepSeek R1 format) reasoning_content: str | None = None # For thinking/reasoning (DeepSeek R1 format)
@@ -60,6 +67,7 @@ class ChatCompletionChunkDelta(CustomBaseModel):
class ChatCompletionChunkChoice(CustomBaseModel): class ChatCompletionChunkChoice(CustomBaseModel):
"""Choice in streaming chunk.""" """Choice in streaming chunk."""
index: int index: int
delta: ChatCompletionChunkDelta delta: ChatCompletionChunkDelta
finish_reason: str | None = None finish_reason: str | None = None
@@ -67,6 +75,7 @@ class ChatCompletionChunkChoice(CustomBaseModel):
class ChatCompletionChunk(CustomBaseModel): class ChatCompletionChunk(CustomBaseModel):
"""OpenAI-compatible streaming chunk.""" """OpenAI-compatible streaming chunk."""
id: str id: str
object: str = "chat.completion.chunk" object: str = "chat.completion.chunk"
created: int created: int
+13 -17
View File
@@ -4,17 +4,17 @@ Chat completion service.
Wrapper around Responses API that converts to Chat Completions format. Wrapper around Responses API that converts to Chat Completions format.
Embeds reasoning in <think> tags for Open WebUI compatibility. Embeds reasoning in <think> tags for Open WebUI compatibility.
""" """
import asyncio
import time import time
import uuid import uuid
from typing import AsyncGenerator from collections.abc import AsyncGenerator
from src.chat import constants from src.chat import constants
from src.chat.schemas import ( from src.chat.schemas import (
ChatCompletionChoice,
ChatCompletionChunk, ChatCompletionChunk,
ChatCompletionChunkChoice, ChatCompletionChunkChoice,
ChatCompletionChunkDelta, ChatCompletionChunkDelta,
ChatCompletionChoice,
ChatCompletionRequest, ChatCompletionRequest,
ChatCompletionResponse, ChatCompletionResponse,
ChatCompletionUsage, ChatCompletionUsage,
@@ -43,10 +43,7 @@ async def create_chat_completion(
created_at = int(time.time()) created_at = int(time.time())
# Convert Chat request to Responses request # Convert Chat request to Responses request
input_messages = [ input_messages = [{"role": msg.role, "content": msg.content} for msg in request.messages]
{"role": msg.role, "content": msg.content}
for msg in request.messages
]
response_request = ResponseRequest( response_request = ResponseRequest(
model=request.model, model=request.model,
@@ -54,7 +51,9 @@ async def create_chat_completion(
reasoning={"effort": "medium", "summary": "auto"}, # Enable reasoning reasoning={"effort": "medium", "summary": "auto"}, # Enable reasoning
temperature=request.temperature or 1.0, temperature=request.temperature or 1.0,
max_output_tokens=request.max_tokens, max_output_tokens=request.max_tokens,
stop=request.stop if isinstance(request.stop, list) else ([request.stop] if request.stop else None), stop=request.stop
if isinstance(request.stop, list)
else ([request.stop] if request.stop else None),
) )
# Call Responses API (will use Steward for Tatlock) # Call Responses API (will use Steward for Tatlock)
@@ -118,16 +117,13 @@ async def create_chat_completion_stream(
Yields: Yields:
Chat completion chunks with reasoning as <think> tags Chat completion chunks with reasoning as <think> tags
""" """
from src.responses.streaming import StreamingCoordinator, StreamEventType from src.responses.streaming import StreamEventType, StreamingCoordinator
completion_id = f"chatcmpl-{uuid.uuid4().hex[:24]}" completion_id = f"chatcmpl-{uuid.uuid4().hex[:24]}"
created_at = int(time.time()) created_at = int(time.time())
# Convert Chat request to Responses request # Convert Chat request to Responses request
input_messages = [ input_messages = [{"role": msg.role, "content": msg.content} for msg in request.messages]
{"role": msg.role, "content": msg.content}
for msg in request.messages
]
response_request = ResponseRequest( response_request = ResponseRequest(
model=request.model, model=request.model,
@@ -135,7 +131,9 @@ async def create_chat_completion_stream(
reasoning={"effort": "medium", "summary": "auto"}, reasoning={"effort": "medium", "summary": "auto"},
temperature=request.temperature or 1.0, temperature=request.temperature or 1.0,
max_output_tokens=request.max_tokens, max_output_tokens=request.max_tokens,
stop=request.stop if isinstance(request.stop, list) else ([request.stop] if request.stop else None), stop=request.stop
if isinstance(request.stop, list)
else ([request.stop] if request.stop else None),
stream=True, stream=True,
) )
@@ -163,7 +161,6 @@ async def create_chat_completion_stream(
# Stream from Responses API # Stream from Responses API
coordinator = StreamingCoordinator() coordinator = StreamingCoordinator()
in_reasoning = False
if use_steward: if use_steward:
stream_generator = coordinator.stream_response_with_steward(response_request) stream_generator = coordinator.stream_response_with_steward(response_request)
@@ -174,7 +171,6 @@ async def create_chat_completion_stream(
if event.event == StreamEventType.REASONING_SUMMARY_DELTA: if event.event == StreamEventType.REASONING_SUMMARY_DELTA:
# Stream reasoning via reasoning_content field (DeepSeek R1 format) # Stream reasoning via reasoning_content field (DeepSeek R1 format)
# Open WebUI renders this as collapsible thinking block # Open WebUI renders this as collapsible thinking block
in_reasoning = True
yield ChatCompletionChunk( yield ChatCompletionChunk(
id=completion_id, id=completion_id,
object=constants.CHAT_COMPLETION_CHUNK_OBJECT, object=constants.CHAT_COMPLETION_CHUNK_OBJECT,
@@ -191,7 +187,7 @@ async def create_chat_completion_stream(
elif event.event == StreamEventType.REASONING_SUMMARY_DONE: elif event.event == StreamEventType.REASONING_SUMMARY_DONE:
# Signal end of reasoning block (no content needed) # Signal end of reasoning block (no content needed)
in_reasoning = False pass # nothing downstream reads this; the event just ends the block
elif event.event == StreamEventType.OUTPUT_TEXT_DELTA: elif event.event == StreamEventType.OUTPUT_TEXT_DELTA:
# Stream message content # Stream message content
+34 -86
View File
@@ -2,6 +2,7 @@
Global application configuration. Global application configuration.
Following best practice of splitting config across domains. Following best practice of splitting config across domains.
""" """
from enum import Enum from enum import Enum
from functools import lru_cache from functools import lru_cache
from pathlib import Path from pathlib import Path
@@ -42,6 +43,7 @@ def _get_version_from_pyproject() -> str:
class Environment(str, Enum): class Environment(str, Enum):
"""Application environment.""" """Application environment."""
DEVELOPMENT = "development" DEVELOPMENT = "development"
PRODUCTION = "production" PRODUCTION = "production"
TESTING = "testing" TESTING = "testing"
@@ -54,6 +56,7 @@ class Config(BaseSettings):
Loads from environment variables and .env file. Loads from environment variables and .env file.
Domain-specific configs should be in their respective modules. Domain-specific configs should be in their respective modules.
""" """
model_config = SettingsConfigDict( model_config = SettingsConfigDict(
env_file=".env", env_file=".env",
env_file_encoding="utf-8", env_file_encoding="utf-8",
@@ -74,143 +77,90 @@ class Config(BaseSettings):
# Anthropic Configuration (Claude - cloud fallback) # Anthropic Configuration (Claude - cloud fallback)
ANTHROPIC_API_KEY: str | None = Field( ANTHROPIC_API_KEY: str | None = Field(
default=None, default=None, description="Anthropic API key for the Claude fallback backend"
description="Anthropic API key for the Claude fallback backend"
) )
ANTHROPIC_MODEL: str = Field( ANTHROPIC_MODEL: str = Field(
default="claude-sonnet-5", default="claude-sonnet-5", description="Claude model for the fallback backend"
description="Claude model for the fallback backend"
) )
PREFER_CLOUD_BACKEND: bool = Field( PREFER_CLOUD_BACKEND: bool = Field(
default=False, default=False, description="Prefer Claude over Ollama (default: local-first)"
description="Prefer Claude over Ollama (default: local-first)"
) )
# Ollama Configuration (local - primary backend) # Ollama Configuration (local - primary backend)
OLLAMA_HOST: HttpUrl = Field( OLLAMA_HOST: HttpUrl = Field(default="http://localhost:11434", description="Ollama server URL")
default="http://localhost:11434", OLLAMA_DEFAULT_MODEL: str = Field(default="gemma4:e2b", description="Default Ollama model")
description="Ollama server URL" OLLAMA_TIMEOUT: int = Field(default=120, description="Ollama request timeout in seconds")
)
OLLAMA_DEFAULT_MODEL: str = Field(
default="gemma4:e2b",
description="Default Ollama model"
)
OLLAMA_TIMEOUT: int = Field(
default=120,
description="Ollama request timeout in seconds"
)
STEWARD_TIMEOUT: int = Field( STEWARD_TIMEOUT: int = Field(
default=60, default=60, description="Steward analysis timeout in seconds (gemma4 needs ~35s warm)"
description="Steward analysis timeout in seconds (gemma4 needs ~35s warm)"
) )
STREAM_TIMEOUT: int = Field( STREAM_TIMEOUT: int = Field(
default=20, default=20, description="Timeout for each streaming turn in seconds"
description="Timeout for each streaming turn in seconds"
) )
# SearXNG Configuration # SearXNG Configuration
SEARXNG_HOST: HttpUrl = Field( SEARXNG_HOST: HttpUrl = Field(
default="http://searxng:8080", default="http://searxng:8080",
description="SearXNG server URL (container name; internal port 8080)" description="SearXNG server URL (container name; internal port 8080)",
)
SEARXNG_TIMEOUT: int = Field(
default=30,
description="SearXNG request timeout in seconds"
) )
SEARXNG_TIMEOUT: int = Field(default=30, description="SearXNG request timeout in seconds")
# Redis Configuration # Redis Configuration
REDIS_HOST: str = Field( REDIS_HOST: str = Field(default="localhost", description="Redis server host")
default="localhost", REDIS_PORT: int = Field(default=6379, description="Redis server port")
description="Redis server host" REDIS_TIMEOUT: int = Field(default=5, description="Redis connection timeout in seconds")
)
REDIS_PORT: int = Field(
default=6379,
description="Redis server port"
)
REDIS_TIMEOUT: int = Field(
default=5,
description="Redis connection timeout in seconds"
)
# Library-Desk Configuration (The Librarian backend) # Library-Desk Configuration (The Librarian backend)
LIBRARIAN_TIMEOUT: int = Field( LIBRARIAN_TIMEOUT: int = Field(
default=180, default=180, description="Total time budget for a librarian delegation in seconds"
description="Total time budget for a librarian delegation in seconds"
) )
LIBRARY_DESK_HOST: HttpUrl = Field( LIBRARY_DESK_HOST: HttpUrl = Field(
default="http://library-desk:8089", default="http://library-desk:8089",
description="Library-Desk API URL (container name; internal port 8089)" description="Library-Desk API URL (container name; internal port 8089)",
) )
LIBRARY_DESK_API_KEY: str = Field( LIBRARY_DESK_API_KEY: str = Field(
default="", default="", description="API key for Library-Desk authentication"
description="API key for Library-Desk authentication"
) )
LIBRARY_DESK_TIMEOUT: int = Field( LIBRARY_DESK_TIMEOUT: int = Field(
default=60, default=60, description="Library-Desk request timeout in seconds"
description="Library-Desk request timeout in seconds"
) )
# Core-API Configuration (The Housekeeper backend) # Core-API Configuration (The Housekeeper backend)
CORE_API_HOST: HttpUrl = Field( CORE_API_HOST: HttpUrl = Field(
default="http://core-api:8083", default="http://core-api:8083",
description="Core-API URL for Home Assistant integration (container name; internal port 8083)" description="Core-API URL for Home Assistant integration (container name; internal port 8083)",
)
CORE_API_KEY: str = Field(
default="",
description="API key for Core-API authentication"
)
CORE_API_TIMEOUT: int = Field(
default=30,
description="Core-API request timeout in seconds"
) )
CORE_API_KEY: str = Field(default="", description="API key for Core-API authentication")
CORE_API_TIMEOUT: int = Field(default=30, description="Core-API request timeout in seconds")
# Qdrant Configuration (Memory vector storage) # Qdrant Configuration (Memory vector storage)
QDRANT_HOST: str = Field( QDRANT_HOST: str = Field(default="localhost", description="Qdrant server host")
default="localhost", QDRANT_PORT: int = Field(default=6333, description="Qdrant server port")
description="Qdrant server host"
)
QDRANT_PORT: int = Field(
default=6333,
description="Qdrant server port"
)
QDRANT_EMBEDDING_DIM: int = Field( QDRANT_EMBEDDING_DIM: int = Field(
default=768, default=768, description="Embedding dimension (768 for nomic-embed-text)"
description="Embedding dimension (768 for nomic-embed-text)"
) )
# Ollama Embedding Configuration # Ollama Embedding Configuration
OLLAMA_EMBEDDING_MODEL: str = Field( OLLAMA_EMBEDDING_MODEL: str = Field(
default="nomic-embed-text", default="nomic-embed-text", description="Ollama model for embeddings"
description="Ollama model for embeddings"
) )
# Redis Memory Database # Redis Memory Database
REDIS_MEMORY_DB: int = Field( REDIS_MEMORY_DB: int = Field(default=1, description="Redis database number for memory cache")
default=1, REDIS_MEMORY_TTL_HOURS: int = Field(default=24, description="TTL for session context in hours")
description="Redis database number for memory cache"
)
REDIS_MEMORY_TTL_HOURS: int = Field(
default=24,
description="TTL for session context in hours"
)
# Logging # Logging
LOG_LEVEL: str | None = Field( LOG_LEVEL: str | None = Field(
default=None, default=None, description="Logging level (auto-set based on environment if not specified)"
description="Logging level (auto-set based on environment if not specified)"
) )
# User Configuration # User Configuration
DEFAULT_USER: str | None = Field( DEFAULT_USER: str | None = Field(
default=None, default=None,
description="Default user for single-user setup (auto-set based on environment if not specified)" description="Default user for single-user setup (auto-set based on environment if not specified)",
) )
# CORS # CORS
CORS_ORIGINS: list[str] = Field( CORS_ORIGINS: list[str] = Field(default=["*"], description="Allowed CORS origins")
default=["*"],
description="Allowed CORS origins"
)
CORS_ALLOW_CREDENTIALS: bool = True CORS_ALLOW_CREDENTIALS: bool = True
CORS_ALLOW_METHODS: list[str] = ["*"] CORS_ALLOW_METHODS: list[str] = ["*"]
CORS_ALLOW_HEADERS: list[str] = ["*"] CORS_ALLOW_HEADERS: list[str] = ["*"]
@@ -235,8 +185,7 @@ class Config(BaseSettings):
if ( if (
self.ENVIRONMENT != Environment.PRODUCTION self.ENVIRONMENT != Environment.PRODUCTION
and self.DEFAULT_USER is not None and self.DEFAULT_USER is not None
and sanitize_user_id(self.DEFAULT_USER) and sanitize_user_id(self.DEFAULT_USER) == sanitize_user_id(PRODUCTION_TENANT)
== sanitize_user_id(PRODUCTION_TENANT)
): ):
raise ValueError( raise ValueError(
f"Refusing to start: ENVIRONMENT={self.ENVIRONMENT.value} is " f"Refusing to start: ENVIRONMENT={self.ENVIRONMENT.value} is "
@@ -299,8 +248,7 @@ class Config(BaseSettings):
return self.DEFAULT_USER or PRODUCTION_TENANT return self.DEFAULT_USER or PRODUCTION_TENANT
if self.DEFAULT_USER is not None and ( if self.DEFAULT_USER is not None and (
self.DEFAULT_USER == TEST_TENANT self.DEFAULT_USER == TEST_TENANT or self.DEFAULT_USER.startswith(TEST_TENANT_PREFIX)
or self.DEFAULT_USER.startswith(TEST_TENANT_PREFIX)
): ):
return self.DEFAULT_USER return self.DEFAULT_USER
return TEST_TENANT return TEST_TENANT
+18 -8
View File
@@ -16,7 +16,9 @@ Usage:
from src.core.context import get_user from src.core.context import get_user
user = get_user() # Returns current request's user user = get_user() # Returns current request's user
""" """
from contextvars import ContextVar from contextvars import ContextVar
from types import TracebackType
def get_default_user() -> str: def get_default_user() -> str:
@@ -28,6 +30,7 @@ def get_default_user() -> str:
""" """
# Import here to avoid circular dependency # Import here to avoid circular dependency
from src.core.config import config from src.core.config import config
return config.effective_default_user return config.effective_default_user
@@ -36,9 +39,7 @@ def get_default_user() -> str:
# and resolve the real default in get_user() # and resolve the real default in get_user()
_USER_NOT_SET = "__user_not_set__" _USER_NOT_SET = "__user_not_set__"
current_user: ContextVar[str] = ContextVar("current_user", default=_USER_NOT_SET) current_user: ContextVar[str] = ContextVar("current_user", default=_USER_NOT_SET)
current_conversation: ContextVar[str | None] = ContextVar( current_conversation: ContextVar[str | None] = ContextVar("current_conversation", default=None)
"current_conversation", default=None
)
def apply_tenant_guard(user: str) -> str: def apply_tenant_guard(user: str) -> str:
@@ -60,9 +61,8 @@ def apply_tenant_guard(user: str) -> str:
from src.core.config import PRODUCTION_TENANT, TEST_TENANT, Environment, config from src.core.config import PRODUCTION_TENANT, TEST_TENANT, Environment, config
from src.core.multi_tenancy import sanitize_user_id from src.core.multi_tenancy import sanitize_user_id
if ( if config.ENVIRONMENT != Environment.PRODUCTION and sanitize_user_id(user) == sanitize_user_id(
config.ENVIRONMENT != Environment.PRODUCTION PRODUCTION_TENANT
and sanitize_user_id(user) == sanitize_user_id(PRODUCTION_TENANT)
): ):
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
@@ -143,7 +143,12 @@ class RequestContext:
self._conv_token = current_conversation.set(self.conversation_id) self._conv_token = current_conversation.set(self.conversation_id)
return self return self
async def __aexit__(self, exc_type, exc_val, exc_tb) -> None: async def __aexit__(
self,
exc_type: type[BaseException] | None,
exc_val: BaseException | None,
exc_tb: TracebackType | None,
) -> None:
"""Reset context variables on exit.""" """Reset context variables on exit."""
if self._user_token is not None: if self._user_token is not None:
current_user.reset(self._user_token) current_user.reset(self._user_token)
@@ -156,7 +161,12 @@ class RequestContext:
self._conv_token = current_conversation.set(self.conversation_id) self._conv_token = current_conversation.set(self.conversation_id)
return self return self
def __exit__(self, exc_type, exc_val, exc_tb) -> None: def __exit__(
self,
exc_type: type[BaseException] | None,
exc_val: BaseException | None,
exc_tb: TracebackType | None,
) -> None:
"""Sync context manager exit.""" """Sync context manager exit."""
if self._user_token is not None: if self._user_token is not None:
current_user.reset(self._user_token) current_user.reset(self._user_token)
+8 -2
View File
@@ -8,7 +8,8 @@ Provides async embedding operations via Ollama API:
Adapted from library-desk patterns. Adapted from library-desk patterns.
""" """
from typing import Optional
from types import TracebackType
import httpx import httpx
@@ -73,7 +74,12 @@ class OllamaEmbeddingClient:
await self._get_client() await self._get_client()
return self return self
async def __aexit__(self, exc_type, exc_val, exc_tb) -> None: async def __aexit__(
self,
exc_type: type[BaseException] | None,
exc_val: BaseException | None,
exc_tb: TracebackType | None,
) -> None:
"""Async context manager exit.""" """Async context manager exit."""
await self.close() await self.close()
+6 -5
View File
@@ -2,12 +2,13 @@
Global exception definitions. Global exception definitions.
Domain-specific exceptions should be in their respective modules. Domain-specific exceptions should be in their respective modules.
""" """
from typing import Any from typing import Any
class AppException(Exception): class AppException(Exception):
"""Base exception for all application errors.""" """Base exception for all application errors."""
def __init__( def __init__(
self, self,
message: str = "An error occurred", message: str = "An error occurred",
@@ -22,26 +23,26 @@ class AppException(Exception):
class OllamaConnectionError(AppException): class OllamaConnectionError(AppException):
"""Raised when cannot connect to Ollama service.""" """Raised when cannot connect to Ollama service."""
def __init__(self, message: str = "Cannot connect to Ollama service"): def __init__(self, message: str = "Cannot connect to Ollama service"):
super().__init__(message=message, status_code=503) super().__init__(message=message, status_code=503)
class OllamaTimeoutError(AppException): class OllamaTimeoutError(AppException):
"""Raised when Ollama request times out.""" """Raised when Ollama request times out."""
def __init__(self, message: str = "Ollama request timed out"): def __init__(self, message: str = "Ollama request timed out"):
super().__init__(message=message, status_code=504) super().__init__(message=message, status_code=504)
class ModelNotFoundError(AppException): class ModelNotFoundError(AppException):
"""Raised when requested model is not available.""" """Raised when requested model is not available."""
def __init__(self, model_name: str): def __init__(self, model_name: str):
super().__init__( super().__init__(
message=f"Model '{model_name}' not found", message=f"Model '{model_name}' not found",
status_code=404, status_code=404,
details={"model": model_name} details={"model": model_name},
) )
+9 -9
View File
@@ -5,10 +5,10 @@ Provides centralized registry of household members (agents) with their
capabilities and tools. Supports two-tier abstraction: executive summaries capabilities and tools. Supports two-tier abstraction: executive summaries
for coordination and full toolsets for execution. for coordination and full toolsets for execution.
""" """
from typing import Any, Optional
from typing import Any
from pydantic import BaseModel, ConfigDict from pydantic import BaseModel, ConfigDict
from pydantic_ai import Agent
from .logging_config import get_logger from .logging_config import get_logger
@@ -22,6 +22,7 @@ class HouseholdCapability(BaseModel):
This is what the Steward and Butler see for coordination. This is what the Steward and Butler see for coordination.
High-level description without implementation details. High-level description without implementation details.
""" """
name: str # Unique identifier: "tatlock_core", "librarian", "developer" name: str # Unique identifier: "tatlock_core", "librarian", "developer"
role: str # Display name: "Butler's Core Tools", "The Librarian" role: str # Display name: "Butler's Core Tools", "The Librarian"
category: str # "core", "research", "technical", "automation" category: str # "core", "research", "technical", "automation"
@@ -38,11 +39,12 @@ class HouseholdMember(BaseModel):
Contains both the executive summary (for coordination) and Contains both the executive summary (for coordination) and
implementation details (tools/agent). implementation details (tools/agent).
""" """
model_config = ConfigDict(arbitrary_types_allowed=True) model_config = ConfigDict(arbitrary_types_allowed=True)
capability: HouseholdCapability capability: HouseholdCapability
tools: list[Any] # PydanticAI tool definitions (any type since Tool is a dataclass) tools: list[Any] # PydanticAI tool definitions (any type since Tool is a dataclass)
agent: Optional[Any] = None # For expert agents (Phase 4) agent: Any | None = None # For expert agents (Phase 4)
class HouseholdRegistry: class HouseholdRegistry:
@@ -55,7 +57,7 @@ class HouseholdRegistry:
3. Agent delegation (Phase 4) 3. Agent delegation (Phase 4)
""" """
def __init__(self): def __init__(self) -> None:
"""Initialize empty registry.""" """Initialize empty registry."""
self._members: dict[str, HouseholdMember] = {} self._members: dict[str, HouseholdMember] = {}
logger.info("household_registry_initialized") logger.info("household_registry_initialized")
@@ -65,7 +67,7 @@ class HouseholdRegistry:
name: str, name: str,
capability: HouseholdCapability, capability: HouseholdCapability,
tools: list[Any], tools: list[Any],
agent: Optional[Any] = None, agent: Any | None = None,
) -> None: ) -> None:
""" """
Register a household member. Register a household member.
@@ -95,9 +97,7 @@ class HouseholdRegistry:
... ) ... )
""" """
if name != capability.name: if name != capability.name:
raise ValueError( raise ValueError(f"Name mismatch: '{name}' != '{capability.name}'")
f"Name mismatch: '{name}' != '{capability.name}'"
)
self._members[name] = HouseholdMember( self._members[name] = HouseholdMember(
capability=capability, capability=capability,
@@ -132,7 +132,7 @@ class HouseholdRegistry:
role=member.capability.role, role=member.capability.role,
) )
def get_member(self, name: str) -> Optional[HouseholdMember]: def get_member(self, name: str) -> HouseholdMember | None:
""" """
Get full household member specification. Get full household member specification.
+33 -13
View File
@@ -4,12 +4,14 @@ Structured logging configuration using structlog.
Deeply integrates with FastAPI/uvicorn's built-in logging to provide Deeply integrates with FastAPI/uvicorn's built-in logging to provide
seamless structured logs across the entire application stack. seamless structured logs across the entire application stack.
""" """
import logging import logging
import logging.config import logging.config
import sys import sys
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from datetime import datetime, timezone from datetime import UTC, datetime
from typing import Any, AsyncIterator from typing import Any
import structlog import structlog
from structlog.types import EventDict, Processor from structlog.types import EventDict, Processor
@@ -19,7 +21,7 @@ from .config import config
def add_timestamp(logger: Any, method_name: str, event_dict: EventDict) -> EventDict: def add_timestamp(logger: Any, method_name: str, event_dict: EventDict) -> EventDict:
"""Add ISO 8601 timestamp to log entries.""" """Add ISO 8601 timestamp to log entries."""
event_dict["timestamp"] = datetime.now(timezone.utc).isoformat() event_dict["timestamp"] = datetime.now(UTC).isoformat()
return event_dict return event_dict
@@ -41,11 +43,28 @@ def extract_from_record(logger: Any, method_name: str, event_dict: EventDict) ->
# Extract custom fields from record # Extract custom fields from record
for key, value in record.__dict__.items(): for key, value in record.__dict__.items():
if key not in { if key not in {
"name", "msg", "args", "created", "filename", "funcName", "name",
"levelname", "levelno", "lineno", "module", "msecs", "msg",
"message", "pathname", "process", "processName", "relativeCreated", "args",
"thread", "threadName", "exc_info", "exc_text", "stack_info", "created",
"taskName" "filename",
"funcName",
"levelname",
"levelno",
"lineno",
"module",
"msecs",
"message",
"pathname",
"process",
"processName",
"relativeCreated",
"thread",
"threadName",
"exc_info",
"exc_text",
"stack_info",
"taskName",
}: }:
event_dict[key] = value event_dict[key] = value
@@ -164,7 +183,7 @@ def get_logger(name: str) -> structlog.stdlib.BoundLogger:
async def log_operation( async def log_operation(
operation: str, operation: str,
initial_context: dict[str, Any] | None = None, initial_context: dict[str, Any] | None = None,
logger_name: str = "tatlock.operations" logger_name: str = "tatlock.operations",
) -> AsyncIterator[dict[str, Any]]: ) -> AsyncIterator[dict[str, Any]]:
""" """
Context manager for automatic operation timing and logging. Context manager for automatic operation timing and logging.
@@ -187,21 +206,21 @@ async def log_operation(
context = initial_context or {} context = initial_context or {}
context["operation"] = operation context["operation"] = operation
start_time = datetime.now(timezone.utc) start_time = datetime.now(UTC)
logger.info("operation_started", **context) logger.info("operation_started", **context)
try: try:
yield context yield context
# Success case # Success case
duration = (datetime.now(timezone.utc) - start_time).total_seconds() duration = (datetime.now(UTC) - start_time).total_seconds()
context["duration_seconds"] = duration context["duration_seconds"] = duration
context["success"] = True context["success"] = True
logger.info("operation_completed", **context) logger.info("operation_completed", **context)
except Exception as e: except Exception as e:
# Error case # Error case
duration = (datetime.now(timezone.utc) - start_time).total_seconds() duration = (datetime.now(UTC) - start_time).total_seconds()
context["duration_seconds"] = duration context["duration_seconds"] = duration
context["success"] = False context["success"] = False
context["error"] = str(e) context["error"] = str(e)
@@ -228,7 +247,8 @@ def get_uvicorn_log_config() -> dict[str, Any]:
"()": structlog.stdlib.ProcessorFormatter, "()": structlog.stdlib.ProcessorFormatter,
"processors": [ "processors": [
structlog.stdlib.ProcessorFormatter.remove_processors_meta, structlog.stdlib.ProcessorFormatter.remove_processors_meta,
structlog.processors.JSONRenderer() if config.log_format == "json" structlog.processors.JSONRenderer()
if config.log_format == "json"
else structlog.dev.ConsoleRenderer(colors=True), else structlog.dev.ConsoleRenderer(colors=True),
], ],
}, },
+2 -1
View File
@@ -8,6 +8,7 @@ Provides short-term memory storage with TTL:
Uses Redis DB 1. Uses Redis DB 1.
""" """
import json import json
from typing import Any from typing import Any
@@ -15,7 +16,7 @@ import redis.asyncio as redis
from .config import config from .config import config
from .logging_config import get_logger from .logging_config import get_logger
from .multi_tenancy import get_session_key, get_entities_key from .multi_tenancy import get_entities_key, get_session_key
logger = get_logger(__name__) logger = get_logger(__name__)
+16 -14
View File
@@ -21,14 +21,14 @@ Usage:
# Get session context # Get session context
ctx = await memory_service.get_session_context(conversation_id) ctx = await memory_service.get_session_context(conversation_id)
""" """
from datetime import datetime, timezone
from datetime import UTC, datetime
from enum import Enum from enum import Enum
from typing import Any from typing import Any
from pydantic import BaseModel, Field from pydantic import BaseModel, Field
from .config import config from .context import get_conversation_id, get_user
from .context import get_user, get_conversation_id
from .embeddings import get_embedding_client from .embeddings import get_embedding_client
from .logging_config import get_logger from .logging_config import get_logger
from .memory_cache import get_memory_cache from .memory_cache import get_memory_cache
@@ -40,22 +40,24 @@ logger = get_logger(__name__)
class MemoryType(str, Enum): class MemoryType(str, Enum):
"""Types of memories stored in Qdrant.""" """Types of memories stored in Qdrant."""
USER_PROFILE = "user_profile" # Name, location, timezone
PREFERENCE = "preference" # Units, language, theme USER_PROFILE = "user_profile" # Name, location, timezone
LEARNED_FACT = "learned_fact" # "My car is a Tesla" PREFERENCE = "preference" # Units, language, theme
LEARNED_FACT = "learned_fact" # "My car is a Tesla"
class MemoryRecord(BaseModel): class MemoryRecord(BaseModel):
"""A memory record stored in Qdrant.""" """A memory record stored in Qdrant."""
id: str id: str
type: MemoryType type: MemoryType
key: str # e.g., "location", "timezone", "car" key: str # e.g., "location", "timezone", "car"
value: str # The actual content value: str # The actual content
keywords: list[str] = Field(default_factory=list) keywords: list[str] = Field(default_factory=list)
importance: float = 0.5 # 0.0 - 1.0 importance: float = 0.5 # 0.0 - 1.0
source: str = "explicit" # "explicit" | "inferred" | "conversation" source: str = "explicit" # "explicit" | "inferred" | "conversation"
created_at: str = Field(default_factory=lambda: datetime.now(timezone.utc).isoformat()) created_at: str = Field(default_factory=lambda: datetime.now(UTC).isoformat())
updated_at: str = Field(default_factory=lambda: datetime.now(timezone.utc).isoformat()) updated_at: str = Field(default_factory=lambda: datetime.now(UTC).isoformat())
class MemoryService: class MemoryService:
@@ -72,7 +74,7 @@ class MemoryService:
- Semantic recall: "What did I mention about X?" Use Memory Agent - Semantic recall: "What did I mention about X?" Use Memory Agent
""" """
def __init__(self): def __init__(self) -> None:
"""Initialize memory service with lazy client loading.""" """Initialize memory service with lazy client loading."""
self._qdrant = None self._qdrant = None
self._embedding = None self._embedding = None
@@ -528,7 +530,7 @@ class MemoryService:
"keywords": keywords, "keywords": keywords,
"importance": importance, "importance": importance,
"source": source, "source": source,
"updated_at": datetime.now(timezone.utc).isoformat(), "updated_at": datetime.now(UTC).isoformat(),
} }
result = await self.qdrant.upsert_memory( result = await self.qdrant.upsert_memory(
+6 -5
View File
@@ -2,6 +2,7 @@
Custom Pydantic base models for consistent serialization. Custom Pydantic base models for consistent serialization.
Following best practice of having a global base model. Following best practice of having a global base model.
""" """
from datetime import datetime from datetime import datetime
from typing import Any from typing import Any
@@ -17,12 +18,13 @@ def datetime_to_iso_str(dt: datetime) -> str:
class CustomBaseModel(BaseModel): class CustomBaseModel(BaseModel):
""" """
Custom base model with consistent configuration. Custom base model with consistent configuration.
All domain models should inherit from this for: All domain models should inherit from this for:
- Consistent JSON serialization - Consistent JSON serialization
- Timezone-aware datetime handling - Timezone-aware datetime handling
- Alias population support - Alias population support
""" """
model_config = ConfigDict( model_config = ConfigDict(
json_encoders={datetime: datetime_to_iso_str}, json_encoders={datetime: datetime_to_iso_str},
populate_by_name=True, populate_by_name=True,
@@ -30,14 +32,13 @@ class CustomBaseModel(BaseModel):
validate_assignment=True, validate_assignment=True,
arbitrary_types_allowed=True, arbitrary_types_allowed=True,
) )
def serializable_dict(self, **kwargs: Any) -> dict[str, Any]: def serializable_dict(self, **kwargs: Any) -> dict[str, Any]:
""" """
Return dict with only JSON-serializable fields. Return dict with only JSON-serializable fields.
Useful for logging and debugging. Useful for logging and debugging.
""" """
return jsonable_encoder( return jsonable_encoder(
self.model_dump(**kwargs), self.model_dump(**kwargs), custom_encoder={datetime: datetime_to_iso_str}
custom_encoder={datetime: datetime_to_iso_str}
) )
+5 -4
View File
@@ -7,6 +7,7 @@ Provides utilities for user namespace management across:
Adapted from library-desk patterns. Adapted from library-desk patterns.
""" """
import re import re
@@ -39,13 +40,13 @@ def sanitize_user_id(user_id: str) -> str:
sanitized = sanitized.replace(".", "_") sanitized = sanitized.replace(".", "_")
# Replace any non-alphanumeric characters with underscores # Replace any non-alphanumeric characters with underscores
sanitized = re.sub(r'[^a-z0-9_]', '_', sanitized) sanitized = re.sub(r"[^a-z0-9_]", "_", sanitized)
# Remove consecutive underscores # Remove consecutive underscores
sanitized = re.sub(r'_+', '_', sanitized) sanitized = re.sub(r"_+", "_", sanitized)
# Remove leading/trailing underscores # Remove leading/trailing underscores
sanitized = sanitized.strip('_') sanitized = sanitized.strip("_")
return sanitized return sanitized
@@ -141,7 +142,7 @@ def validate_user_id(user_id: str) -> bool:
return False return False
# Must contain at least one alphanumeric character # Must contain at least one alphanumeric character
if not re.search(r'[a-zA-Z0-9]', user_id): if not re.search(r"[a-zA-Z0-9]", user_id):
return False return False
return True return True
+14 -12
View File
@@ -3,15 +3,16 @@ Request preprocessing pipeline.
Analyzes requests via the Steward and creates scoped toolsets for Tatlock. Analyzes requests via the Steward and creates scoped toolsets for Tatlock.
""" """
from dataclasses import dataclass from dataclasses import dataclass
from datetime import datetime from datetime import datetime
from typing import Any, Optional from typing import Any
from src.agents.steward import analyze_request, format_steward_note from src.agents.steward import analyze_request, format_steward_note
from src.agents.steward.schemas import StewardRecommendation from src.agents.steward.schemas import StewardRecommendation
from src.core.household_registry import get_household_registry from src.core.household_registry import get_household_registry
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
from src.core.tracing import trace_span, SpanType from src.core.tracing import SpanType, trace_span
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -45,6 +46,7 @@ class EnrichedRequest:
recommendation: Full Steward recommendation recommendation: Full Steward recommendation
steward_reasoning: Plain text reasoning for streaming to user steward_reasoning: Plain text reasoning for streaming to user
""" """
original_request: str original_request: str
steward_note: str steward_note: str
scoped_tools: list[Any] # PydanticAI tool definitions scoped_tools: list[Any] # PydanticAI tool definitions
@@ -55,7 +57,7 @@ class EnrichedRequest:
async def preprocess_request( async def preprocess_request(
user_request: str, user_request: str,
conversation_history: list[dict], conversation_history: list[dict],
conversation_id: Optional[str] = None, conversation_id: str | None = None,
) -> EnrichedRequest: ) -> EnrichedRequest:
""" """
Analyze request via Steward and prepare scoped context for Tatlock. Analyze request via Steward and prepare scoped context for Tatlock.
@@ -111,12 +113,14 @@ async def preprocess_request(
# Update span with results # Update span with results
if span: if span:
span.metadata.update({ span.metadata.update(
"recommended_capabilities": recommendation.recommended_capabilities, {
"complexity": recommendation.estimated_complexity, "recommended_capabilities": recommendation.recommended_capabilities,
"has_memory_context": bool(recommendation.memory_context), "complexity": recommendation.estimated_complexity,
"has_conversation_context": recommendation.conversation_context.has_previous_context, "has_memory_context": bool(recommendation.memory_context),
}) "has_conversation_context": recommendation.conversation_context.has_previous_context,
}
)
span.details["reasoning"] = recommendation.reasoning span.details["reasoning"] = recommendation.reasoning
if recommendation.enriched_query: if recommendation.enriched_query:
span.details["enriched_query"] = recommendation.enriched_query span.details["enriched_query"] = recommendation.enriched_query
@@ -128,9 +132,7 @@ async def preprocess_request(
# Uses agent-as-tool pattern: expert agents get delegation wrappers, # Uses agent-as-tool pattern: expert agents get delegation wrappers,
# core tools are returned directly # core tools are returned directly
registry = get_household_registry() registry = get_household_registry()
scoped_tools = registry.get_delegation_tools( scoped_tools = registry.get_delegation_tools(recommendation.recommended_capabilities)
recommendation.recommended_capabilities
)
logger.info( logger.info(
"preprocessing_complete", "preprocessing_complete",
+2 -1
View File
@@ -8,8 +8,9 @@ Provides async operations for storing and retrieving memory embeddings:
Adapted from library-desk patterns. Adapted from library-desk patterns.
""" """
from typing import Any from typing import Any
from uuid import uuid4, uuid5, NAMESPACE_DNS from uuid import NAMESPACE_DNS, uuid4, uuid5
from qdrant_client import QdrantClient from qdrant_client import QdrantClient
from qdrant_client.http import models as qdrant_models from qdrant_client.http import models as qdrant_models
+3 -2
View File
@@ -1,6 +1,7 @@
""" """
Core router for health and root endpoints. Core router for health and root endpoints.
""" """
import logging import logging
from fastapi import APIRouter from fastapi import APIRouter
@@ -16,7 +17,7 @@ router = APIRouter(tags=["core"])
async def health_check() -> dict[str, str]: async def health_check() -> dict[str, str]:
""" """
Health check endpoint. Health check endpoint.
Returns: Returns:
Health status Health status
""" """
@@ -27,7 +28,7 @@ async def health_check() -> dict[str, str]:
async def root() -> dict[str, str]: async def root() -> dict[str, str]:
""" """
Root endpoint. Root endpoint.
Returns: Returns:
API information API information
""" """
+3 -2
View File
@@ -5,6 +5,7 @@ Handles initialization of household registry and other startup tasks.
This module should be called during application startup to register This module should be called during application startup to register
all household members. all household members.
""" """
from src.agents.biographer import register_biographer from src.agents.biographer import register_biographer
from src.agents.housekeeper import register_housekeeper from src.agents.housekeeper import register_housekeeper
from src.agents.librarian import register_librarian from src.agents.librarian import register_librarian
@@ -46,7 +47,7 @@ def log_tenant_guard() -> None:
) )
def register_household_members(): def register_household_members() -> None:
""" """
Register all household members with the registry. Register all household members with the registry.
@@ -112,7 +113,7 @@ def register_household_members():
) )
async def initialize_application(): async def initialize_application() -> None:
""" """
Initialize the application. Initialize the application.
+8 -21
View File
@@ -4,7 +4,6 @@ Tool call tracking.
Tracks which tools are recommended by the Steward versus which tools Tracks which tools are recommended by the Steward versus which tools
are actually used by Tatlock for debugging and analysis. are actually used by Tatlock for debugging and analysis.
""" """
from typing import Optional
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
@@ -19,11 +18,7 @@ class ToolCallTracker:
to measure recommendation accuracy. to measure recommendation accuracy.
""" """
def __init__( def __init__(self, recommended_capabilities: list[str], conversation_id: str | None = None):
self,
recommended_capabilities: list[str],
conversation_id: Optional[str] = None
):
""" """
Initialize tool call tracker. Initialize tool call tracker.
@@ -51,11 +46,11 @@ class ToolCallTracker:
return tool_name.replace("delegate_to_", "") return tool_name.replace("delegate_to_", "")
return tool_name return tool_name
def log_call(self, message: str): def log_call(self, message: str) -> None:
"""Log a tool call message (for UI display).""" """Log a tool call message (for UI display)."""
logger.debug("tool_call_message", message=message) logger.debug("tool_call_message", message=message)
async def track_call(self, tool_name: str, duration: float): async def track_call(self, tool_name: str, duration: float) -> None:
""" """
Record a tool call with timing. Record a tool call with timing.
@@ -87,7 +82,7 @@ class ToolCallTracker:
was_recommended=was_recommended, was_recommended=was_recommended,
) )
async def finalize(self): async def finalize(self) -> None:
""" """
Finalize tracking and log unused recommended tools. Finalize tracking and log unused recommended tools.
@@ -95,9 +90,7 @@ class ToolCallTracker:
tools that were recommended but never used. tools that were recommended but never used.
""" """
# Normalize actual tool names to capabilities for comparison # Normalize actual tool names to capabilities for comparison
used_capabilities = { used_capabilities = {self._extract_capability(tool) for tool in self.actual_calls.keys()}
self._extract_capability(tool) for tool in self.actual_calls.keys()
}
# Find tools that were recommended but not used # Find tools that were recommended but not used
unused_tools = self.recommended_capabilities - used_capabilities unused_tools = self.recommended_capabilities - used_capabilities
@@ -128,9 +121,7 @@ class ToolCallTracker:
""" """
total_calls = sum(len(durations) for durations in self.actual_calls.values()) total_calls = sum(len(durations) for durations in self.actual_calls.values())
# Normalize actual tool names to capabilities for comparison # Normalize actual tool names to capabilities for comparison
used_capabilities = { used_capabilities = {self._extract_capability(tool) for tool in self.actual_calls.keys()}
self._extract_capability(tool) for tool in self.actual_calls.keys()
}
unused = self.recommended_capabilities - used_capabilities unused = self.recommended_capabilities - used_capabilities
return { return {
@@ -139,12 +130,8 @@ class ToolCallTracker:
"tools_unused": list(unused), "tools_unused": list(unused),
"total_calls": total_calls, "total_calls": total_calls,
"accuracy": { "accuracy": {
"recommended_and_used": len( "recommended_and_used": len(self.recommended_capabilities & used_capabilities),
self.recommended_capabilities & used_capabilities
),
"recommended_but_unused": len(unused), "recommended_but_unused": len(unused),
"not_recommended_but_used": len( "not_recommended_but_used": len(used_capabilities - self.recommended_capabilities),
used_capabilities - self.recommended_capabilities
),
}, },
} }
+23 -14
View File
@@ -9,16 +9,16 @@ Enable via DEBUG=true environment variable.
Traces are written to logs/traces/{trace_id}.json Traces are written to logs/traces/{trace_id}.json
View with logs/traces/viewer.html View with logs/traces/viewer.html
""" """
from contextlib import asynccontextmanager
from contextvars import ContextVar
from dataclasses import dataclass, field
from datetime import datetime, timezone
from enum import Enum
from pathlib import Path
from typing import Any
import json import json
import secrets import secrets
from contextlib import asynccontextmanager
from contextvars import ContextVar
from dataclasses import dataclass, field
from datetime import UTC, datetime
from enum import Enum
from pathlib import Path
from typing import Any
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
@@ -27,6 +27,7 @@ logger = get_logger(__name__)
class SpanType(str, Enum): class SpanType(str, Enum):
"""Types of traced operations.""" """Types of traced operations."""
ROUTER = "router" ROUTER = "router"
STEWARD = "steward" STEWARD = "steward"
TATLOCK = "tatlock" TATLOCK = "tatlock"
@@ -36,6 +37,7 @@ class SpanType(str, Enum):
class SpanStatus(str, Enum): class SpanStatus(str, Enum):
"""Span completion status.""" """Span completion status."""
OK = "ok" OK = "ok"
ERROR = "error" ERROR = "error"
@@ -43,6 +45,7 @@ class SpanStatus(str, Enum):
@dataclass @dataclass
class Span: class Span:
"""A single traced operation.""" """A single traced operation."""
span_id: str span_id: str
name: str name: str
type: SpanType type: SpanType
@@ -88,6 +91,7 @@ class Span:
@dataclass @dataclass
class Trace: class Trace:
"""Complete trace of a request.""" """Complete trace of a request."""
trace_id: str trace_id: str
conversation_id: str | None conversation_id: str | None
user: str user: str
@@ -116,7 +120,9 @@ class Trace:
"conversation_id": self.conversation_id, "conversation_id": self.conversation_id,
"user": self.user, "user": self.user,
"timestamp": self.timestamp.isoformat(), "timestamp": self.timestamp.isoformat(),
"total_duration_ms": round(self.total_duration_ms, 2) if self.total_duration_ms else None, "total_duration_ms": round(self.total_duration_ms, 2)
if self.total_duration_ms
else None,
"status": self.status, "status": self.status,
"request": self.request, "request": self.request,
"response": self.response, "response": self.response,
@@ -132,6 +138,7 @@ _current_span: ContextVar[Span | None] = ContextVar("current_span", default=None
def tracing_enabled() -> bool: def tracing_enabled() -> bool:
"""Check if tracing is enabled (requires DEBUG=true).""" """Check if tracing is enabled (requires DEBUG=true)."""
from src.core.config import config from src.core.config import config
return config.DEBUG return config.DEBUG
@@ -163,7 +170,7 @@ def start_trace(
trace_id=_generate_id("trace_"), trace_id=_generate_id("trace_"),
conversation_id=conversation_id, conversation_id=conversation_id,
user=user, user=user,
timestamp=datetime.now(timezone.utc), timestamp=datetime.now(UTC),
request=request, request=request,
) )
_current_trace.set(trace) _current_trace.set(trace)
@@ -209,7 +216,7 @@ def start_span(
span_id=_generate_id("span_"), span_id=_generate_id("span_"),
name=name, name=name,
type=span_type, type=span_type,
start_time=datetime.now(timezone.utc), start_time=datetime.now(UTC),
parent_id=parent.span_id if parent else None, parent_id=parent.span_id if parent else None,
metadata=metadata or {}, metadata=metadata or {},
details=details or {}, details=details or {},
@@ -254,7 +261,7 @@ def end_span(
if not span: if not span:
return return
span.end_time = datetime.now(timezone.utc) span.end_time = datetime.now(UTC)
span.status = status span.status = status
if error: if error:
span.error = error span.error = error
@@ -405,7 +412,7 @@ def add_tool_spans_from_messages(messages: list[Any], parent_span: Span | None =
if isinstance(part, ToolCallPart): if isinstance(part, ToolCallPart):
tool_calls[part.tool_call_id] = { tool_calls[part.tool_call_id] = {
"name": part.tool_name, "name": part.tool_name,
"args": part.args if hasattr(part, 'args') else {}, "args": part.args if hasattr(part, "args") else {},
} }
elif isinstance(msg, ModelRequest): elif isinstance(msg, ModelRequest):
for part in msg.parts: for part in msg.parts:
@@ -418,7 +425,7 @@ def add_tool_spans_from_messages(messages: list[Any], parent_span: Span | None =
name=tool_info["name"], name=tool_info["name"],
type=SpanType.TOOL, type=SpanType.TOOL,
start_time=parent_span.start_time, # Approximate start_time=parent_span.start_time, # Approximate
end_time=parent_span.end_time or datetime.now(timezone.utc), end_time=parent_span.end_time or datetime.now(UTC),
parent_id=parent_span.span_id, parent_id=parent_span.span_id,
status=SpanStatus.OK, status=SpanStatus.OK,
metadata={ metadata={
@@ -427,7 +434,9 @@ def add_tool_spans_from_messages(messages: list[Any], parent_span: Span | None =
}, },
details={ details={
"args": tool_info.get("args", {}), "args": tool_info.get("args", {}),
"result": part.content[:2000] if isinstance(part.content, str) else str(part.content)[:2000], "result": part.content[:2000]
if isinstance(part.content, str)
else str(part.content)[:2000],
}, },
) )
parent_span.children.append(span.span_id) parent_span.children.append(span.span_id)
+20 -14
View File
@@ -4,7 +4,10 @@ Trace viewer router.
Serves the trace viewer UI and trace files when tracing is enabled. Serves the trace viewer UI and trace files when tracing is enabled.
Only available when DEBUG=true. Only available when DEBUG=true.
""" """
from datetime import UTC
from pathlib import Path from pathlib import Path
from typing import Any
from fastapi import APIRouter, HTTPException from fastapi import APIRouter, HTTPException
from fastapi.responses import HTMLResponse, JSONResponse from fastapi.responses import HTMLResponse, JSONResponse
@@ -66,12 +69,12 @@ async def list_traces(
return {"traces": [], "total": 0} return {"traces": [], "total": 0}
import json import json
from datetime import datetime, timezone, timedelta from datetime import datetime, timedelta
# Calculate cutoff time if filtering by time # Calculate cutoff time if filtering by time
cutoff_time = None cutoff_time = None
if since_minutes: if since_minutes:
cutoff_time = datetime.now(timezone.utc) - timedelta(minutes=since_minutes) cutoff_time = datetime.now(UTC) - timedelta(minutes=since_minutes)
# Get all trace files, sorted by modification time (newest first) # Get all trace files, sorted by modification time (newest first)
trace_files = sorted( trace_files = sorted(
@@ -80,7 +83,7 @@ async def list_traces(
reverse=True, reverse=True,
) )
traces = [] traces: list[dict[str, Any]] = []
for path in trace_files: for path in trace_files:
if len(traces) >= limit: if len(traces) >= limit:
break break
@@ -93,7 +96,7 @@ async def list_traces(
trace_timestamp = data.get("timestamp") trace_timestamp = data.get("timestamp")
if cutoff_time and trace_timestamp: if cutoff_time and trace_timestamp:
try: try:
ts = datetime.fromisoformat(trace_timestamp.replace('Z', '+00:00')) ts = datetime.fromisoformat(trace_timestamp.replace("Z", "+00:00"))
if ts < cutoff_time: if ts < cutoff_time:
continue continue
except (ValueError, TypeError): except (ValueError, TypeError):
@@ -109,15 +112,17 @@ async def list_traces(
if search and search.lower() not in request_preview.lower(): if search and search.lower() not in request_preview.lower():
continue continue
traces.append({ traces.append(
"trace_id": data.get("trace_id"), {
"timestamp": trace_timestamp, "trace_id": data.get("trace_id"),
"user": data.get("user"), "timestamp": trace_timestamp,
"status": trace_status, "user": data.get("user"),
"total_duration_ms": data.get("total_duration_ms"), "status": trace_status,
"span_count": len(data.get("spans", [])), "total_duration_ms": data.get("total_duration_ms"),
"request_preview": request_preview[:100], "span_count": len(data.get("spans", [])),
}) "request_preview": request_preview[:100],
}
)
except Exception as e: except Exception as e:
logger.warning("trace_list_parse_error", path=str(path), error=str(e)) logger.warning("trace_list_parse_error", path=str(path), error=str(e))
@@ -145,9 +150,10 @@ async def get_trace(trace_id: str):
try: try:
import json import json
with open(trace_path) as f: with open(trace_path) as f:
data = json.load(f) data = json.load(f)
return JSONResponse(content=data) return JSONResponse(content=data)
except Exception as e: except Exception as e:
logger.error("trace_read_error", trace_id=trace_id, error=str(e)) logger.error("trace_read_error", trace_id=trace_id, error=str(e))
raise HTTPException(status_code=500, detail="Failed to read trace") raise HTTPException(status_code=500, detail="Failed to read trace") from e
+8 -7
View File
@@ -9,8 +9,9 @@ Main responsibilities:
- Router registration - Router registration
- Lifecycle management - Lifecycle management
""" """
from collections.abc import AsyncGenerator
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from typing import AsyncGenerator
from fastapi import FastAPI, Request, status from fastapi import FastAPI, Request, status
from fastapi.exceptions import RequestValidationError from fastapi.exceptions import RequestValidationError
@@ -64,7 +65,7 @@ async def lifespan(app: FastAPI) -> AsyncGenerator[None, None]:
def create_application() -> FastAPI: def create_application() -> FastAPI:
""" """
Application factory. Application factory.
Creates and configures the FastAPI application. Creates and configures the FastAPI application.
Following best practice of using factory pattern. Following best practice of using factory pattern.
""" """
@@ -75,7 +76,7 @@ def create_application() -> FastAPI:
lifespan=lifespan, lifespan=lifespan,
debug=config.DEBUG, debug=config.DEBUG,
) )
# Add middleware # Add middleware
application.add_middleware( application.add_middleware(
CORSMiddleware, CORSMiddleware,
@@ -84,10 +85,10 @@ def create_application() -> FastAPI:
allow_methods=config.CORS_ALLOW_METHODS, allow_methods=config.CORS_ALLOW_METHODS,
allow_headers=config.CORS_ALLOW_HEADERS, allow_headers=config.CORS_ALLOW_HEADERS,
) )
# Register exception handlers # Register exception handlers
register_exception_handlers(application) register_exception_handlers(application)
# Include routers # Include routers
application.include_router(core_router) # Health and root endpoints application.include_router(core_router) # Health and root endpoints
application.include_router(chat_router, prefix=config.API_PREFIX) application.include_router(chat_router, prefix=config.API_PREFIX)
@@ -105,10 +106,10 @@ def create_application() -> FastAPI:
def register_exception_handlers(application: FastAPI) -> None: def register_exception_handlers(application: FastAPI) -> None:
""" """
Register global exception handlers. Register global exception handlers.
Provides consistent error responses compatible with OpenAI API. Provides consistent error responses compatible with OpenAI API.
""" """
@application.exception_handler(AppException) @application.exception_handler(AppException)
async def app_exception_handler( async def app_exception_handler(
request: Request, request: Request,
+3 -2
View File
@@ -2,6 +2,7 @@
Models router. Models router.
OpenAI-compatible /v1/models endpoint. OpenAI-compatible /v1/models endpoint.
""" """
import logging import logging
from fastapi import APIRouter from fastapi import APIRouter
@@ -18,9 +19,9 @@ router = APIRouter(prefix="/models", tags=["models"])
async def list_models() -> ModelsResponse: async def list_models() -> ModelsResponse:
""" """
List available models (OpenAI-compatible). List available models (OpenAI-compatible).
Currently returns mock model list. Currently returns mock model list.
Returns: Returns:
List of available models List of available models
""" """
+3
View File
@@ -1,11 +1,13 @@
""" """
OpenAI-compatible models schemas. OpenAI-compatible models schemas.
""" """
from src.core.models import CustomBaseModel from src.core.models import CustomBaseModel
class Model(CustomBaseModel): class Model(CustomBaseModel):
"""OpenAI-compatible model object.""" """OpenAI-compatible model object."""
id: str id: str
object: str = "model" object: str = "model"
created: int created: int
@@ -14,5 +16,6 @@ class Model(CustomBaseModel):
class ModelsResponse(CustomBaseModel): class ModelsResponse(CustomBaseModel):
"""OpenAI-compatible models list response.""" """OpenAI-compatible models list response."""
object: str = "list" object: str = "list"
data: list[Model] data: list[Model]
+33 -30
View File
@@ -2,8 +2,10 @@
Ollama HTTP client. Ollama HTTP client.
Handles all communication with the Ollama service. Handles all communication with the Ollama service.
""" """
import logging import logging
from typing import Any, AsyncGenerator from collections.abc import AsyncGenerator
from typing import Any
import httpx import httpx
from httpx import ConnectError, TimeoutException from httpx import ConnectError, TimeoutException
@@ -22,14 +24,14 @@ logger = logging.getLogger(__name__)
class OllamaClient: class OllamaClient:
""" """
Async client for Ollama API. Async client for Ollama API.
Follows best practice of using async for I/O operations. Follows best practice of using async for I/O operations.
""" """
def __init__(self, base_url: str | None = None, timeout: int | None = None): def __init__(self, base_url: str | None = None, timeout: int | None = None):
""" """
Initialize Ollama client. Initialize Ollama client.
Args: Args:
base_url: Ollama server URL (defaults to config) base_url: Ollama server URL (defaults to config)
timeout: Request timeout in seconds (defaults to config) timeout: Request timeout in seconds (defaults to config)
@@ -37,7 +39,7 @@ class OllamaClient:
self.base_url = base_url or str(config.OLLAMA_HOST) self.base_url = base_url or str(config.OLLAMA_HOST)
self.timeout = timeout or config.OLLAMA_TIMEOUT self.timeout = timeout or config.OLLAMA_TIMEOUT
self._client: httpx.AsyncClient | None = None self._client: httpx.AsyncClient | None = None
async def __aenter__(self) -> "OllamaClient": async def __aenter__(self) -> "OllamaClient":
"""Async context manager entry.""" """Async context manager entry."""
self._client = httpx.AsyncClient( self._client = httpx.AsyncClient(
@@ -45,32 +47,32 @@ class OllamaClient:
timeout=self.timeout, timeout=self.timeout,
) )
return self return self
async def __aexit__(self, *args: Any) -> None: async def __aexit__(self, *args: Any) -> None:
"""Async context manager exit.""" """Async context manager exit."""
if self._client: if self._client:
await self._client.aclose() await self._client.aclose()
async def chat( async def chat(
self, self,
request: OllamaChatRequest, request: OllamaChatRequest,
) -> OllamaChatResponse: ) -> OllamaChatResponse:
""" """
Send chat request to Ollama (non-streaming). Send chat request to Ollama (non-streaming).
Args: Args:
request: Chat request with model and messages request: Chat request with model and messages
Returns: Returns:
Complete chat response Complete chat response
Raises: Raises:
OllamaConnectionError: Cannot connect to Ollama OllamaConnectionError: Cannot connect to Ollama
OllamaTimeoutError: Request timed out OllamaTimeoutError: Request timed out
""" """
if not self._client: if not self._client:
raise RuntimeError("Client not initialized. Use async with context.") raise RuntimeError("Client not initialized. Use async with context.")
try: try:
response = await self._client.post( response = await self._client.post(
"/api/chat", "/api/chat",
@@ -78,38 +80,38 @@ class OllamaClient:
) )
response.raise_for_status() response.raise_for_status()
return OllamaChatResponse(**response.json()) return OllamaChatResponse(**response.json())
except ConnectError as e: except ConnectError as e:
logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}") logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}")
raise OllamaConnectionError() from e raise OllamaConnectionError() from e
except TimeoutException as e: except TimeoutException as e:
logger.error(f"Ollama request timed out after {self.timeout}s: {e}") logger.error(f"Ollama request timed out after {self.timeout}s: {e}")
raise OllamaTimeoutError() from e raise OllamaTimeoutError() from e
async def chat_stream( async def chat_stream(
self, self,
request: OllamaChatRequest, request: OllamaChatRequest,
) -> AsyncGenerator[dict[str, Any], None]: ) -> AsyncGenerator[dict[str, Any], None]:
""" """
Send streaming chat request to Ollama. Send streaming chat request to Ollama.
Args: Args:
request: Chat request with stream=True request: Chat request with stream=True
Yields: Yields:
Streaming response chunks Streaming response chunks
Raises: Raises:
OllamaConnectionError: Cannot connect to Ollama OllamaConnectionError: Cannot connect to Ollama
OllamaTimeoutError: Request timed out OllamaTimeoutError: Request timed out
""" """
if not self._client: if not self._client:
raise RuntimeError("Client not initialized. Use async with context.") raise RuntimeError("Client not initialized. Use async with context.")
# Ensure streaming is enabled # Ensure streaming is enabled
request.stream = True request.stream = True
try: try:
async with self._client.stream( async with self._client.stream(
"POST", "POST",
@@ -120,48 +122,49 @@ class OllamaClient:
async for line in response.aiter_lines(): async for line in response.aiter_lines():
if line.strip(): if line.strip():
import json import json
yield json.loads(line) yield json.loads(line)
except ConnectError as e: except ConnectError as e:
logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}") logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}")
raise OllamaConnectionError() from e raise OllamaConnectionError() from e
except TimeoutException as e: except TimeoutException as e:
logger.error(f"Ollama request timed out after {self.timeout}s: {e}") logger.error(f"Ollama request timed out after {self.timeout}s: {e}")
raise OllamaTimeoutError() from e raise OllamaTimeoutError() from e
async def list_models(self) -> OllamaModelsResponse: async def list_models(self) -> OllamaModelsResponse:
""" """
List available models from Ollama. List available models from Ollama.
Returns: Returns:
List of available models List of available models
Raises: Raises:
OllamaConnectionError: Cannot connect to Ollama OllamaConnectionError: Cannot connect to Ollama
""" """
if not self._client: if not self._client:
raise RuntimeError("Client not initialized. Use async with context.") raise RuntimeError("Client not initialized. Use async with context.")
try: try:
response = await self._client.get("/api/tags") response = await self._client.get("/api/tags")
response.raise_for_status() response.raise_for_status()
return OllamaModelsResponse(**response.json()) return OllamaModelsResponse(**response.json())
except ConnectError as e: except ConnectError as e:
logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}") logger.error(f"Cannot connect to Ollama at {self.base_url}: {e}")
raise OllamaConnectionError() from e raise OllamaConnectionError() from e
async def health_check(self) -> bool: async def health_check(self) -> bool:
""" """
Check if Ollama service is healthy. Check if Ollama service is healthy.
Returns: Returns:
True if healthy, False otherwise True if healthy, False otherwise
""" """
if not self._client: if not self._client:
raise RuntimeError("Client not initialized. Use async with context.") raise RuntimeError("Client not initialized. Use async with context.")
try: try:
response = await self._client.get("/") response = await self._client.get("/")
return response.status_code == 200 return response.status_code == 200
@@ -174,7 +177,7 @@ class OllamaClient:
async def get_ollama_client() -> AsyncGenerator[OllamaClient, None]: async def get_ollama_client() -> AsyncGenerator[OllamaClient, None]:
""" """
FastAPI dependency to provide Ollama client. FastAPI dependency to provide Ollama client.
Follows best practice of dependency injection. Follows best practice of dependency injection.
""" """
async with OllamaClient() as client: async with OllamaClient() as client:
+1
View File
@@ -5,6 +5,7 @@ Ollama's OpenAI-compatible API rejects messages with `content: null`,
which PydanticAI sends for assistant messages that only contain tool calls. which PydanticAI sends for assistant messages that only contain tool calls.
This provider sanitizes messages to use empty strings instead of null. This provider sanitizes messages to use empty strings instead of null.
""" """
from typing import Any from typing import Any
from openai import AsyncOpenAI from openai import AsyncOpenAI
+6
View File
@@ -2,6 +2,7 @@
Ollama API schemas. Ollama API schemas.
Internal models for Ollama API communication. Internal models for Ollama API communication.
""" """
from typing import Any from typing import Any
from src.core.models import CustomBaseModel from src.core.models import CustomBaseModel
@@ -9,12 +10,14 @@ from src.core.models import CustomBaseModel
class OllamaMessage(CustomBaseModel): class OllamaMessage(CustomBaseModel):
"""Message format for Ollama API.""" """Message format for Ollama API."""
role: str role: str
content: str content: str
class OllamaChatRequest(CustomBaseModel): class OllamaChatRequest(CustomBaseModel):
"""Chat request to Ollama API.""" """Chat request to Ollama API."""
model: str model: str
messages: list[OllamaMessage] messages: list[OllamaMessage]
stream: bool = False stream: bool = False
@@ -23,6 +26,7 @@ class OllamaChatRequest(CustomBaseModel):
class OllamaChatResponse(CustomBaseModel): class OllamaChatResponse(CustomBaseModel):
"""Chat response from Ollama API.""" """Chat response from Ollama API."""
model: str model: str
created_at: str created_at: str
message: OllamaMessage message: OllamaMessage
@@ -31,6 +35,7 @@ class OllamaChatResponse(CustomBaseModel):
class OllamaModelInfo(CustomBaseModel): class OllamaModelInfo(CustomBaseModel):
"""Model information from Ollama.""" """Model information from Ollama."""
name: str name: str
modified_at: str modified_at: str
size: int size: int
@@ -39,4 +44,5 @@ class OllamaModelInfo(CustomBaseModel):
class OllamaModelsResponse(CustomBaseModel): class OllamaModelsResponse(CustomBaseModel):
"""Response from Ollama models list endpoint.""" """Response from Ollama models list endpoint."""
models: list[OllamaModelInfo] models: list[OllamaModelInfo]
+4 -4
View File
@@ -9,12 +9,12 @@ This module implements the OpenAI Responses API format:
""" """
from src.responses.schemas import ( from src.responses.schemas import (
FunctionCallOutputItem,
MessageOutputItem,
OutputItem,
ReasoningOutputItem,
Response, Response,
ResponseRequest, ResponseRequest,
OutputItem,
MessageOutputItem,
ReasoningOutputItem,
FunctionCallOutputItem,
) )
__all__ = [ __all__ = [
+6 -17
View File
@@ -6,6 +6,7 @@ Handles:
- Context trimming to fit model limits - Context trimming to fit model limits
- Reserve tokens for output generation - Reserve tokens for output generation
""" """
from typing import Any from typing import Any
@@ -59,11 +60,7 @@ class ContextWindow:
# Approximate: 4 characters per token # Approximate: 4 characters per token
return total_chars // 4 return total_chars // 4
async def trim_to_fit( async def trim_to_fit(self, items: list[Any], reserve_tokens: int = 512) -> list[Any]:
self,
items: list[Any],
reserve_tokens: int = 512
) -> list[Any]:
""" """
Trim items to fit within context window. Trim items to fit within context window.
@@ -86,7 +83,7 @@ class ContextWindow:
return [] return []
# Start from most recent, work backwards # Start from most recent, work backwards
kept_items = [] kept_items: list[Any] = []
current_tokens = 0 current_tokens = 0
for item in reversed(items): for item in reversed(items):
@@ -102,11 +99,7 @@ class ContextWindow:
return kept_items return kept_items
async def fits_in_context( async def fits_in_context(self, items: list[Any], reserve_tokens: int = 512) -> bool:
self,
items: list[Any],
reserve_tokens: int = 512
) -> bool:
""" """
Check if items fit within context window. Check if items fit within context window.
@@ -121,11 +114,7 @@ class ContextWindow:
available_tokens = self.max_tokens - reserve_tokens available_tokens = self.max_tokens - reserve_tokens
return total_tokens <= available_tokens return total_tokens <= available_tokens
async def get_usage_stats( async def get_usage_stats(self, items: list[Any], reserve_tokens: int = 512) -> dict:
self,
items: list[Any],
reserve_tokens: int = 512
) -> dict:
""" """
Get context window usage statistics. Get context window usage statistics.
@@ -153,7 +142,7 @@ class ContextWindow:
"reserved_tokens": reserve_tokens, "reserved_tokens": reserve_tokens,
"available_tokens": available_tokens, "available_tokens": available_tokens,
"usage_percent": round(usage_percent, 2), "usage_percent": round(usage_percent, 2),
"fits": total_tokens <= available_tokens "fits": total_tokens <= available_tokens,
} }
# ======================================================================== # ========================================================================
+9 -16
View File
@@ -6,8 +6,8 @@ Supports hybrid approach:
- Optional conversation_id in metadata for server-side grouping - Optional conversation_id in metadata for server-side grouping
- Server can augment with vector memories (future) - Server can augment with vector memories (future)
""" """
import hashlib import hashlib
from typing import Dict, List
from src.responses.schemas import Response, ResponseRequest from src.responses.schemas import Response, ResponseRequest
@@ -36,7 +36,7 @@ class ConversationHistory:
Args: Args:
max_turns: Maximum number of response turns to keep per conversation max_turns: Maximum number of response turns to keep per conversation
""" """
self._conversations: Dict[str, List[Response]] = {} self._conversations: dict[str, list[Response]] = {}
self._max_turns = max_turns self._max_turns = max_turns
async def get_conversation_id(self, request: ResponseRequest) -> str: async def get_conversation_id(self, request: ResponseRequest) -> str:
@@ -61,11 +61,7 @@ class ConversationHistory:
first_msg = str(request.input[0]) if request.input else "" first_msg = str(request.input[0]) if request.input else ""
return hashlib.sha256(first_msg.encode()).hexdigest()[:16] return hashlib.sha256(first_msg.encode()).hexdigest()[:16]
async def add_response( async def add_response(self, conversation_id: str, response: Response) -> None:
self,
conversation_id: str,
response: Response
) -> None:
""" """
Add response to conversation history. Add response to conversation history.
@@ -81,7 +77,7 @@ class ConversationHistory:
# Trim old turns to stay within limit # Trim old turns to stay within limit
await self._trim_history(conversation_id) await self._trim_history(conversation_id)
async def get_history(self, conversation_id: str) -> List[Response]: async def get_history(self, conversation_id: str) -> list[Response]:
""" """
Retrieve conversation history. Retrieve conversation history.
@@ -125,20 +121,17 @@ class ConversationHistory:
conversation_id: Conversation identifier conversation_id: Conversation identifier
""" """
if len(self._conversations[conversation_id]) > self._max_turns: if len(self._conversations[conversation_id]) > self._max_turns:
self._conversations[conversation_id] = ( self._conversations[conversation_id] = self._conversations[conversation_id][
self._conversations[conversation_id][-self._max_turns:] -self._max_turns :
) ]
# ======================================================================== # ========================================================================
# Future: Vector Memory Integration # Future: Vector Memory Integration
# ======================================================================== # ========================================================================
async def get_relevant_memories( async def get_relevant_memories(
self, self, conversation_id: str, query: str, limit: int = 5
conversation_id: str, ) -> list[dict]:
query: str,
limit: int = 5
) -> List[dict]:
""" """
Retrieve relevant memories from vector store. Retrieve relevant memories from vector store.
+9 -12
View File
@@ -7,10 +7,10 @@ OpenAI-compatible /v1/responses endpoint with streaming support.
from fastapi import APIRouter, HTTPException from fastapi import APIRouter, HTTPException
from sse_starlette.sse import EventSourceResponse from sse_starlette.sse import EventSourceResponse
from src.responses import service from src.core.exceptions import AppException, ModelNotFoundError
from src.responses.schemas import ResponseRequest, Response
from src.core.exceptions import ModelNotFoundError, AppException
from src.core.logging_config import get_logger from src.core.logging_config import get_logger
from src.responses import service
from src.responses.schemas import Response, ResponseRequest
logger = get_logger(__name__) logger = get_logger(__name__)
@@ -58,14 +58,11 @@ async def create_response(
if use_steward: if use_steward:
logger.info("Streaming with Steward preprocessing for Tatlock request") logger.info("Streaming with Steward preprocessing for Tatlock request")
from src.responses.streaming import StreamingCoordinator from src.responses.streaming import StreamingCoordinator
coordinator = StreamingCoordinator() coordinator = StreamingCoordinator()
return EventSourceResponse( return EventSourceResponse(coordinator.stream_response_with_steward(request))
coordinator.stream_response_with_steward(request)
)
else: else:
return EventSourceResponse( return EventSourceResponse(service.create_response_stream(request))
service.create_response_stream(request)
)
# Non-streaming response # Non-streaming response
if use_steward: if use_steward:
@@ -76,12 +73,12 @@ async def create_response(
except ModelNotFoundError as e: except ModelNotFoundError as e:
logger.error(f"Model not found: {e}") logger.error(f"Model not found: {e}")
raise HTTPException(status_code=404, detail=str(e)) raise HTTPException(status_code=404, detail=str(e)) from e
except AppException as e: except AppException as e:
logger.error(f"Application error: {e}") logger.error(f"Application error: {e}")
raise HTTPException(status_code=e.status_code, detail=e.message) raise HTTPException(status_code=e.status_code, detail=e.message) from e
except Exception as e: except Exception as e:
logger.error(f"Unexpected error: {e}", exc_info=True) logger.error(f"Unexpected error: {e}", exc_info=True)
raise HTTPException(status_code=500, detail="Internal server error") raise HTTPException(status_code=500, detail="Internal server error") from e
+38 -46
View File
@@ -8,18 +8,20 @@ OpenAI Responses API format with support for:
- Streaming and non-streaming modes - Streaming and non-streaming modes
""" """
from typing import Literal, Any from typing import Any, Literal
from pydantic import Field, field_validator from pydantic import Field, field_validator
from src.core.models import CustomBaseModel from src.core.models import CustomBaseModel
# ============================================================================ # ============================================================================
# Output Item Schemas (appear in response.output array) # Output Item Schemas (appear in response.output array)
# ============================================================================ # ============================================================================
class OutputTextContent(CustomBaseModel): class OutputTextContent(CustomBaseModel):
"""Text content in message output.""" """Text content in message output."""
type: Literal["output_text"] = "output_text" type: Literal["output_text"] = "output_text"
text: str text: str
annotations: list[dict] = Field(default_factory=list) annotations: list[dict] = Field(default_factory=list)
@@ -31,6 +33,7 @@ class MessageOutputItem(CustomBaseModel):
Represents the assistant's final response message. Represents the assistant's final response message.
""" """
type: Literal["message"] = "message" type: Literal["message"] = "message"
id: str id: str
role: Literal["assistant"] = "assistant" role: Literal["assistant"] = "assistant"
@@ -45,6 +48,7 @@ class ReasoningOutputItem(CustomBaseModel):
Represents the model's thinking/reasoning process. Represents the model's thinking/reasoning process.
Displayed separately from the final answer. Displayed separately from the final answer.
""" """
type: Literal["reasoning"] = "reasoning" type: Literal["reasoning"] = "reasoning"
id: str id: str
summary: list[str] # List of reasoning steps summary: list[str] # List of reasoning steps
@@ -57,6 +61,7 @@ class FunctionCallOutputItem(CustomBaseModel):
Represents a tool/function that the model wants to execute. Represents a tool/function that the model wants to execute.
""" """
type: Literal["function_call"] = "function_call" type: Literal["function_call"] = "function_call"
id: str id: str
name: str name: str
@@ -66,15 +71,17 @@ class FunctionCallOutputItem(CustomBaseModel):
# Union type for all output items # Union type for all output items
# Type: ignore because Pydantic handles union types specially # Type: ignore because Pydantic handles union types specially
OutputItem = MessageOutputItem | ReasoningOutputItem | FunctionCallOutputItem # type: ignore OutputItem = MessageOutputItem | ReasoningOutputItem | FunctionCallOutputItem
# ============================================================================ # ============================================================================
# Usage Tracking # Usage Tracking
# ============================================================================ # ============================================================================
class ResponseUsage(CustomBaseModel): class ResponseUsage(CustomBaseModel):
"""Token usage statistics for the response.""" """Token usage statistics for the response."""
input_tokens: int input_tokens: int
output_tokens: int output_tokens: int
reasoning_tokens: int = 0 reasoning_tokens: int = 0
@@ -85,8 +92,10 @@ class ResponseUsage(CustomBaseModel):
# Request Schema # Request Schema
# ============================================================================ # ============================================================================
class Tool(CustomBaseModel): class Tool(CustomBaseModel):
"""Tool/function definition.""" """Tool/function definition."""
name: str name: str
description: str description: str
parameters: dict[str, Any] parameters: dict[str, Any]
@@ -94,6 +103,7 @@ class Tool(CustomBaseModel):
class ReasoningConfig(CustomBaseModel): class ReasoningConfig(CustomBaseModel):
"""Reasoning configuration.""" """Reasoning configuration."""
effort: Literal["none", "minimal", "low", "medium", "high", "xhigh"] = "medium" effort: Literal["none", "minimal", "low", "medium", "high", "xhigh"] = "medium"
summary: Literal["auto", "off"] = "auto" summary: Literal["auto", "off"] = "auto"
@@ -104,46 +114,25 @@ class ResponseRequest(CustomBaseModel):
OpenAI Responses API format with optional extensions. OpenAI Responses API format with optional extensions.
""" """
model: str = Field(description="Model ID to use") model: str = Field(description="Model ID to use")
input: list[dict] = Field( input: list[dict] = Field(description="Input messages or previous responses")
description="Input messages or previous responses"
)
reasoning: dict[str, Any] | None = Field( reasoning: dict[str, Any] | None = Field(
default=None, default=None, description="Reasoning configuration: {effort: 'medium', summary: 'auto'}"
description="Reasoning configuration: {effort: 'medium', summary: 'auto'}"
)
tools: list[dict] | None = Field(
default=None,
description="Available tools/functions"
) )
tools: list[dict] | None = Field(default=None, description="Available tools/functions")
metadata: dict[str, Any] | None = Field( metadata: dict[str, Any] | None = Field(
default=None, default=None, description="Custom metadata (e.g., conversation_id for server-side tracking)"
description="Custom metadata (e.g., conversation_id for server-side tracking)"
)
stream: bool = Field(
default=False,
description="Enable streaming mode"
)
max_output_tokens: int | None = Field(
default=None,
description="Maximum tokens to generate"
)
temperature: float = Field(
default=1.0,
ge=0.0,
le=2.0,
description="Sampling temperature"
)
stop: list[str] | None = Field(
default=None,
description="Stop sequences"
) )
stream: bool = Field(default=False, description="Enable streaming mode")
max_output_tokens: int | None = Field(default=None, description="Maximum tokens to generate")
temperature: float = Field(default=1.0, ge=0.0, le=2.0, description="Sampling temperature")
stop: list[str] | None = Field(default=None, description="Stop sequences")
user: str | None = Field( user: str | None = Field(
default=None, default=None, description="Unique identifier for end-user (OpenAI standard)"
description="Unique identifier for end-user (OpenAI standard)"
) )
@field_validator('reasoning') @field_validator("reasoning")
@classmethod @classmethod
def validate_reasoning(cls, v: dict[str, Any] | None) -> dict[str, Any] | None: def validate_reasoning(cls, v: dict[str, Any] | None) -> dict[str, Any] | None:
""" """
@@ -154,21 +143,21 @@ class ResponseRequest(CustomBaseModel):
- summary must be 'auto' or 'off' - summary must be 'auto' or 'off'
""" """
if v is not None: if v is not None:
if 'effort' in v: if "effort" in v:
allowed_efforts = ['none', 'minimal', 'low', 'medium', 'high', 'xhigh'] allowed_efforts = ["none", "minimal", "low", "medium", "high", "xhigh"]
if v['effort'] not in allowed_efforts: if v["effort"] not in allowed_efforts:
raise ValueError( raise ValueError(
f"reasoning.effort must be one of {allowed_efforts}, got '{v['effort']}'" f"reasoning.effort must be one of {allowed_efforts}, got '{v['effort']}'"
) )
if 'summary' in v: if "summary" in v:
allowed_summaries = ['auto', 'off'] allowed_summaries = ["auto", "off"]
if v['summary'] not in allowed_summaries: if v["summary"] not in allowed_summaries:
raise ValueError( raise ValueError(
f"reasoning.summary must be one of {allowed_summaries}, got '{v['summary']}'" f"reasoning.summary must be one of {allowed_summaries}, got '{v['summary']}'"
) )
return v return v
@field_validator('max_output_tokens') @field_validator("max_output_tokens")
@classmethod @classmethod
def validate_max_output_tokens(cls, v: int | None) -> int | None: def validate_max_output_tokens(cls, v: int | None) -> int | None:
""" """
@@ -180,7 +169,7 @@ class ResponseRequest(CustomBaseModel):
raise ValueError(f"max_output_tokens must be positive, got {v}") raise ValueError(f"max_output_tokens must be positive, got {v}")
return v return v
@field_validator('stop') @field_validator("stop")
@classmethod @classmethod
def validate_stop_sequences(cls, v: list[str] | None) -> list[str] | None: def validate_stop_sequences(cls, v: list[str] | None) -> list[str] | None:
""" """
@@ -203,20 +192,20 @@ class ResponseRequest(CustomBaseModel):
# Response Schema # Response Schema
# ============================================================================ # ============================================================================
class Response(CustomBaseModel): class Response(CustomBaseModel):
""" """
Complete response object. Complete response object.
Contains output array with reasoning, function calls, and messages. Contains output array with reasoning, function calls, and messages.
""" """
id: str = Field(description="Unique response ID") id: str = Field(description="Unique response ID")
object: Literal["response"] = "response" object: Literal["response"] = "response"
created_at: int = Field(description="Unix timestamp") created_at: int = Field(description="Unix timestamp")
model: str = Field(description="Model used") model: str = Field(description="Model used")
status: Literal["completed", "in_progress", "failed", "cancelled"] status: Literal["completed", "in_progress", "failed", "cancelled"]
output: list[OutputItem] = Field( output: list[OutputItem] = Field(description="Output items (reasoning, function_call, message)")
description="Output items (reasoning, function_call, message)"
)
usage: ResponseUsage = Field(description="Token usage statistics") usage: ResponseUsage = Field(description="Token usage statistics")
@@ -224,8 +213,10 @@ class Response(CustomBaseModel):
# Error Schema # Error Schema
# ============================================================================ # ============================================================================
class ErrorDetail(CustomBaseModel): class ErrorDetail(CustomBaseModel):
"""Error detail object.""" """Error detail object."""
type: str type: str
message: str message: str
code: int | None = None code: int | None = None
@@ -233,4 +224,5 @@ class ErrorDetail(CustomBaseModel):
class ErrorResponse(CustomBaseModel): class ErrorResponse(CustomBaseModel):
"""Error response format.""" """Error response format."""
error: ErrorDetail error: ErrorDetail
+67 -64
View File
@@ -52,9 +52,9 @@ def _extract_response_preview(response: Response) -> str:
"""Extract response preview text for tracing.""" """Extract response preview text for tracing."""
if response.output: if response.output:
for item in response.output: for item in response.output:
if hasattr(item, 'content'): if hasattr(item, "content"):
for content in item.content: for content in item.content:
if hasattr(content, 'text'): if hasattr(content, "text"):
return content.text[:200] return content.text[:200]
return "" return ""
@@ -79,10 +79,12 @@ async def _execute_single_delegation(
result summary is a curated user-safe sentence. result summary is a curated user-safe sentence.
""" """
import time import time
start_time = time.time() start_time = time.time()
if agent_name == "biographer": if agent_name == "biographer":
from src.agents.delegation import delegate_to_biographer from src.agents.delegation import delegate_to_biographer
result = await delegate_to_biographer(task=task, context=context) result = await delegate_to_biographer(task=task, context=context)
duration = time.time() - start_time duration = time.time() - start_time
await tracker.track_call("delegate_to_biographer", duration) await tracker.track_call("delegate_to_biographer", duration)
@@ -90,6 +92,7 @@ async def _execute_single_delegation(
elif agent_name == "librarian": elif agent_name == "librarian":
from src.agents.delegation import delegate_to_librarian from src.agents.delegation import delegate_to_librarian
result = await delegate_to_librarian(task=task, context=context) result = await delegate_to_librarian(task=task, context=context)
duration = time.time() - start_time duration = time.time() - start_time
await tracker.track_call("delegate_to_librarian", duration) await tracker.track_call("delegate_to_librarian", duration)
@@ -97,6 +100,7 @@ async def _execute_single_delegation(
elif agent_name == "housekeeper": elif agent_name == "housekeeper":
from src.agents.delegation import delegate_to_housekeeper from src.agents.delegation import delegate_to_housekeeper
result = await delegate_to_housekeeper(task=task, context=context) result = await delegate_to_housekeeper(task=task, context=context)
duration = time.time() - start_time duration = time.time() - start_time
await tracker.track_call("delegate_to_housekeeper", duration) await tracker.track_call("delegate_to_housekeeper", duration)
@@ -107,9 +111,7 @@ async def _execute_single_delegation(
async def _handle_text_delegation( async def _handle_text_delegation(
response: str, response: str, tracker: "ToolCallTracker", conversation_id: str
tracker: "ToolCallTracker",
conversation_id: str
) -> str: ) -> str:
""" """
Handle text-based delegation fallback. Handle text-based delegation fallback.
@@ -173,8 +175,7 @@ async def _handle_text_delegation(
conversation_id=conversation_id, conversation_id=conversation_id,
) )
tasks = [ tasks = [
_execute_single_delegation(agent.lower(), task, tracker) _execute_single_delegation(agent.lower(), task, tracker) for agent, task in matches
for agent, task in matches
] ]
results = await asyncio.gather(*tasks, return_exceptions=True) results = await asyncio.gather(*tasks, return_exceptions=True)
@@ -188,10 +189,7 @@ async def _handle_text_delegation(
got=len(results), got=len(results),
conversation_id=conversation_id, conversation_id=conversation_id,
) )
return ( return "I apologize, sir. I was unable to complete the " "requested delegations."
"I apologize, sir. I was unable to complete the "
"requested delegations."
)
# Combine results (failures carry curated user-safe sentences) # Combine results (failures carry curated user-safe sentences)
summaries = [] summaries = []
@@ -205,8 +203,7 @@ async def _handle_text_delegation(
conversation_id=conversation_id, conversation_id=conversation_id,
) )
summaries.append( summaries.append(
f"**{agent_name}**: " f"**{agent_name}**: " f"{get_think_message(agent_name, task, 'error')}"
f"{get_think_message(agent_name, task, 'error')}"
) )
else: else:
_, output, _ = item _, output, _ = item
@@ -226,9 +223,7 @@ async def _handle_text_delegation(
conversation_id=conversation_id, conversation_id=conversation_id,
) )
try: try:
_, output, _ = await _execute_single_delegation( _, output, _ = await _execute_single_delegation(agent_name, task, tracker)
agent_name, task, tracker
)
summaries.append(output) summaries.append(output)
except Exception as e: except Exception as e:
logger.error( logger.error(
@@ -421,7 +416,7 @@ def _calculate_usage(input_messages: list[dict], output_items: list) -> Response
elif isinstance(item, FunctionCallOutputItem): elif isinstance(item, FunctionCallOutputItem):
func_text = item.arguments func_text = item.arguments
output_tokens += len(func_text) // 4 output_tokens += len(func_text) // 4
elif hasattr(item, 'type'): elif hasattr(item, "type"):
# Agent OutputItem objects (backward compatibility) # Agent OutputItem objects (backward compatibility)
if item.type == "reasoning": if item.type == "reasoning":
reasoning_text = " ".join(item.data.get("summary", [])) reasoning_text = " ".join(item.data.get("summary", []))
@@ -439,7 +434,7 @@ def _calculate_usage(input_messages: list[dict], output_items: list) -> Response
input_tokens=input_tokens, input_tokens=input_tokens,
output_tokens=output_tokens, output_tokens=output_tokens,
reasoning_tokens=reasoning_tokens, reasoning_tokens=reasoning_tokens,
total_tokens=total_tokens total_tokens=total_tokens,
) )
@@ -527,7 +522,7 @@ async def create_response(request: ResponseRequest) -> Response:
model=request.model, model=request.model,
status="completed", status="completed",
output=converted_items, output=converted_items,
usage=usage usage=usage,
) )
# Track conversation history (for analytics and future vector memory) # Track conversation history (for analytics and future vector memory)
@@ -640,12 +635,15 @@ async def create_response_with_steward(request: ResponseRequest) -> Response:
# If Steward recommends ONLY delegation agents (biographer/librarian/housekeeper), # If Steward recommends ONLY delegation agents (biographer/librarian/housekeeper),
# we still use two-phase but delegate directly in Phase 1 # we still use two-phase but delegate directly in Phase 1
delegation_agents = {"biographer", "librarian", "housekeeper"} delegation_agents = {"biographer", "librarian", "housekeeper"}
delegation_only = all( delegation_only = (
cap in delegation_agents all(
for cap in enriched.recommendation.recommended_capabilities cap in delegation_agents for cap in enriched.recommendation.recommended_capabilities
) and enriched.recommendation.recommended_capabilities )
and enriched.recommendation.recommended_capabilities
)
from src.agents.tatlock import TatlockAgent from src.agents.tatlock import TatlockAgent
tatlock = TatlockAgent() tatlock = TatlockAgent()
# Use enriched query (with location/timezone context) if available # Use enriched query (with location/timezone context) if available
@@ -677,7 +675,9 @@ async def create_response_with_steward(request: ResponseRequest) -> Response:
) )
# Add text delegation results to expert_results # Add text delegation results to expert_results
if text_delegation_results != orchestration_results["raw_output"]: if text_delegation_results != orchestration_results["raw_output"]:
orchestration_results["expert_results"]["text_delegation"] = text_delegation_results orchestration_results["expert_results"]["text_delegation"] = (
text_delegation_results
)
# Phase 2: Synthesize butler-toned response from all results # Phase 2: Synthesize butler-toned response from all results
tatlock_response = await tatlock.synthesize_from_results( tatlock_response = await tatlock.synthesize_from_results(
@@ -694,23 +694,25 @@ async def create_response_with_steward(request: ResponseRequest) -> Response:
# Add Steward reasoning as reasoning output # Add Steward reasoning as reasoning output
if enriched.steward_reasoning: if enriched.steward_reasoning:
output_items.append(ReasoningOutputItem( output_items.append(
id=f"rs_{generate_id()}", ReasoningOutputItem(
summary=[enriched.steward_reasoning], id=f"rs_{generate_id()}",
status="completed" summary=[enriched.steward_reasoning],
)) status="completed",
)
)
# Add Tatlock's message # Add Tatlock's message
output_items.append(MessageOutputItem( output_items.append(
id=f"msg_{generate_id()}", MessageOutputItem(
role="assistant", id=f"msg_{generate_id()}",
content=[OutputTextContent( role="assistant",
type="output_text", content=[
text=tatlock_response, OutputTextContent(type="output_text", text=tatlock_response, annotations=[])
annotations=[] ],
)], status="completed",
status="completed" )
)) )
# Calculate usage (approximate) # Calculate usage (approximate)
usage = _calculate_usage(request.input, output_items) usage = _calculate_usage(request.input, output_items)
@@ -721,7 +723,7 @@ async def create_response_with_steward(request: ResponseRequest) -> Response:
model=request.model, model=request.model,
status="completed", status="completed",
output=output_items, output=output_items,
usage=usage usage=usage,
) )
# Track conversation history # Track conversation history
@@ -752,9 +754,7 @@ async def create_response_with_steward(request: ResponseRequest) -> Response:
raise raise
async def create_response_stream( async def create_response_stream(request: ResponseRequest) -> AsyncGenerator[dict, None]:
request: ResponseRequest
) -> AsyncGenerator[dict, None]:
""" """
Create streaming response. Create streaming response.
@@ -774,10 +774,7 @@ async def create_response_stream(
coordinator = StreamingCoordinator() coordinator = StreamingCoordinator()
async for event in coordinator.stream_response(request): async for event in coordinator.stream_response(request):
yield { yield {"event": event.event, "data": event.model_dump_json()}
"event": event.event,
"data": event.model_dump_json()
}
async def get_conversation_history(conversation_id: str) -> list[Response]: async def get_conversation_history(conversation_id: str) -> list[Response]:
@@ -802,7 +799,7 @@ async def get_conversation_stats() -> dict:
""" """
return { return {
"total_conversations": await _conversation_history.get_conversation_count(), "total_conversations": await _conversation_history.get_conversation_count(),
"max_turns_per_conversation": _conversation_history._max_turns "max_turns_per_conversation": _conversation_history._max_turns,
} }
@@ -843,23 +840,29 @@ def _convert_output_items(items: list) -> list:
for item in items: for item in items:
if item.type == "message": if item.type == "message":
converted.append(MessageOutputItem( converted.append(
id=item.id, MessageOutputItem(
content=[OutputTextContent(**c) for c in item.data["content"]], id=item.id,
status=item.data.get("status", "completed") content=[OutputTextContent(**c) for c in item.data["content"]],
)) status=item.data.get("status", "completed"),
)
)
elif item.type == "reasoning": elif item.type == "reasoning":
converted.append(ReasoningOutputItem( converted.append(
id=item.id, ReasoningOutputItem(
summary=item.data["summary"], id=item.id,
status=item.data.get("status", "completed") summary=item.data["summary"],
)) status=item.data.get("status", "completed"),
)
)
elif item.type == "function_call": elif item.type == "function_call":
converted.append(FunctionCallOutputItem( converted.append(
id=item.id, FunctionCallOutputItem(
name=item.data["name"], id=item.id,
arguments=item.data["arguments"], name=item.data["name"],
status=item.data.get("status", "completed") arguments=item.data["arguments"],
)) status=item.data.get("status", "completed"),
)
)
return converted return converted
+85 -85
View File
@@ -30,8 +30,10 @@ logger = get_logger(__name__)
# Stream Event Types # Stream Event Types
# ============================================================================ # ============================================================================
class StreamEventType(str, Enum): class StreamEventType(str, Enum):
"""Streaming event types for Responses API.""" """Streaming event types for Responses API."""
REASONING_SUMMARY_DELTA = "response.reasoning_summary_text.delta" REASONING_SUMMARY_DELTA = "response.reasoning_summary_text.delta"
REASONING_SUMMARY_DONE = "response.reasoning_summary_text.done" REASONING_SUMMARY_DONE = "response.reasoning_summary_text.done"
OUTPUT_TEXT_DELTA = "response.output_text.delta" OUTPUT_TEXT_DELTA = "response.output_text.delta"
@@ -46,30 +48,38 @@ class StreamEventType(str, Enum):
# Stream Event Schemas # Stream Event Schemas
# ============================================================================ # ============================================================================
class ReasoningSummaryDelta(CustomBaseModel): class ReasoningSummaryDelta(CustomBaseModel):
"""Reasoning summary text delta event.""" """Reasoning summary text delta event."""
event: Literal[StreamEventType.REASONING_SUMMARY_DELTA] = StreamEventType.REASONING_SUMMARY_DELTA
event: Literal[StreamEventType.REASONING_SUMMARY_DELTA] = (
StreamEventType.REASONING_SUMMARY_DELTA
)
delta: str delta: str
class ReasoningSummaryDone(CustomBaseModel): class ReasoningSummaryDone(CustomBaseModel):
"""Reasoning summary completion event.""" """Reasoning summary completion event."""
event: Literal[StreamEventType.REASONING_SUMMARY_DONE] = StreamEventType.REASONING_SUMMARY_DONE event: Literal[StreamEventType.REASONING_SUMMARY_DONE] = StreamEventType.REASONING_SUMMARY_DONE
class OutputTextDelta(CustomBaseModel): class OutputTextDelta(CustomBaseModel):
"""Output text delta event.""" """Output text delta event."""
event: Literal[StreamEventType.OUTPUT_TEXT_DELTA] = StreamEventType.OUTPUT_TEXT_DELTA event: Literal[StreamEventType.OUTPUT_TEXT_DELTA] = StreamEventType.OUTPUT_TEXT_DELTA
delta: str delta: str
class OutputTextDone(CustomBaseModel): class OutputTextDone(CustomBaseModel):
"""Output text completion event.""" """Output text completion event."""
event: Literal[StreamEventType.OUTPUT_TEXT_DONE] = StreamEventType.OUTPUT_TEXT_DONE event: Literal[StreamEventType.OUTPUT_TEXT_DONE] = StreamEventType.OUTPUT_TEXT_DONE
class FunctionCallDelta(CustomBaseModel): class FunctionCallDelta(CustomBaseModel):
"""Function call arguments delta event.""" """Function call arguments delta event."""
event: Literal[StreamEventType.FUNCTION_CALL_DELTA] = StreamEventType.FUNCTION_CALL_DELTA event: Literal[StreamEventType.FUNCTION_CALL_DELTA] = StreamEventType.FUNCTION_CALL_DELTA
delta: str delta: str
name: str | None = None # Only in first chunk name: str | None = None # Only in first chunk
@@ -77,31 +87,34 @@ class FunctionCallDelta(CustomBaseModel):
class FunctionCallDone(CustomBaseModel): class FunctionCallDone(CustomBaseModel):
"""Function call completion event.""" """Function call completion event."""
event: Literal[StreamEventType.FUNCTION_CALL_DONE] = StreamEventType.FUNCTION_CALL_DONE event: Literal[StreamEventType.FUNCTION_CALL_DONE] = StreamEventType.FUNCTION_CALL_DONE
class ResponseDone(CustomBaseModel): class ResponseDone(CustomBaseModel):
"""Response completion event with full response.""" """Response completion event with full response."""
event: Literal[StreamEventType.RESPONSE_DONE] = StreamEventType.RESPONSE_DONE event: Literal[StreamEventType.RESPONSE_DONE] = StreamEventType.RESPONSE_DONE
response: Response response: Response
class ErrorEvent(CustomBaseModel): class ErrorEvent(CustomBaseModel):
"""Error event.""" """Error event."""
event: Literal[StreamEventType.ERROR] = StreamEventType.ERROR event: Literal[StreamEventType.ERROR] = StreamEventType.ERROR
error: dict error: dict
# Union type for all stream events # Union type for all stream events
StreamEvent = ( StreamEvent = (
ReasoningSummaryDelta | ReasoningSummaryDelta
ReasoningSummaryDone | | ReasoningSummaryDone
OutputTextDelta | | OutputTextDelta
OutputTextDone | | OutputTextDone
FunctionCallDelta | | FunctionCallDelta
FunctionCallDone | | FunctionCallDone
ResponseDone | | ResponseDone
ErrorEvent | ErrorEvent
) )
@@ -109,6 +122,7 @@ StreamEvent = (
# Streaming Coordinator # Streaming Coordinator
# ============================================================================ # ============================================================================
class StreamingCoordinator: class StreamingCoordinator:
""" """
Coordinates streaming from agents to SSE format. Coordinates streaming from agents to SSE format.
@@ -123,7 +137,7 @@ class StreamingCoordinator:
async def stream_response_with_steward( async def stream_response_with_steward(
self, self,
request: "ResponseRequest" # type: ignore # Forward reference request: "ResponseRequest", # Forward reference
) -> AsyncGenerator[StreamEvent, None]: ) -> AsyncGenerator[StreamEvent, None]:
""" """
Stream response with Steward preprocessing and two-phase Tatlock execution. Stream response with Steward preprocessing and two-phase Tatlock execution.
@@ -181,10 +195,13 @@ class StreamingCoordinator:
# Check if direct delegation is recommended # Check if direct delegation is recommended
delegation_agents = {"biographer", "librarian", "housekeeper"} delegation_agents = {"biographer", "librarian", "housekeeper"}
delegation_only = all( delegation_only = (
cap in delegation_agents all(
for cap in enriched.recommendation.recommended_capabilities cap in delegation_agents
) and enriched.recommendation.recommended_capabilities for cap in enriched.recommendation.recommended_capabilities
)
and enriched.recommendation.recommended_capabilities
)
tatlock = TatlockAgent() tatlock = TatlockAgent()
@@ -222,7 +239,7 @@ class StreamingCoordinator:
# Stream the synthesized response # Stream the synthesized response
chunk_size = 50 chunk_size = 50
for i in range(0, len(tatlock_response), chunk_size): for i in range(0, len(tatlock_response), chunk_size):
yield OutputTextDelta(delta=tatlock_response[i:i + chunk_size]) yield OutputTextDelta(delta=tatlock_response[i : i + chunk_size])
await asyncio.sleep(0.02) await asyncio.sleep(0.02)
yield OutputTextDone() yield OutputTextDone()
@@ -231,12 +248,10 @@ class StreamingCoordinator:
message_item = MessageOutputItem( message_item = MessageOutputItem(
id=f"msg_{generate_id()}", id=f"msg_{generate_id()}",
role="assistant", role="assistant",
content=[OutputTextContent( content=[
type="output_text", OutputTextContent(type="output_text", text=tatlock_response, annotations=[])
text=tatlock_response, ],
annotations=[] status="completed",
)],
status="completed"
) )
output_items.append(message_item) output_items.append(message_item)
@@ -252,7 +267,7 @@ class StreamingCoordinator:
model=request.model, model=request.model,
status="completed", status="completed",
output=output_items, output=output_items,
usage=usage usage=usage,
) )
# Track conversation history # Track conversation history
@@ -267,8 +282,8 @@ class StreamingCoordinator:
async def _stream_direct_delegation( async def _stream_direct_delegation(
self, self,
user_message: str, user_message: str,
recommendation: "StewardRecommendation", # type: ignore recommendation: "StewardRecommendation",
tracker: "ToolCallTracker", # type: ignore tracker: "ToolCallTracker",
conversation_id: str, conversation_id: str,
conversation_history: list | None = None, conversation_history: list | None = None,
results: dict | None = None, results: dict | None = None,
@@ -319,17 +334,11 @@ class StreamingCoordinator:
try: try:
# Execute delegation # Execute delegation
if agent == "librarian": if agent == "librarian":
result = await delegate_to_librarian( result = await delegate_to_librarian(task=user_message, context=context)
task=user_message, context=context
)
elif agent == "biographer": elif agent == "biographer":
result = await delegate_to_biographer( result = await delegate_to_biographer(task=user_message, context=context)
task=user_message, context=context
)
elif agent == "housekeeper": elif agent == "housekeeper":
result = await delegate_to_housekeeper( result = await delegate_to_housekeeper(task=user_message, context=context)
task=user_message, context=context
)
else: else:
result = None result = None
@@ -366,17 +375,19 @@ class StreamingCoordinator:
yield ReasoningSummaryDone() yield ReasoningSummaryDone()
if results is not None: if results is not None:
results.update({ results.update(
"tools_called": tools_called, {
"expert_results": expert_results, "tools_called": tools_called,
"tool_outputs": {}, "expert_results": expert_results,
"raw_output": "", "tool_outputs": {},
"think_messages": think_messages, "raw_output": "",
}) "think_messages": think_messages,
}
)
async def stream_response( async def stream_response(
self, self,
request: "ResponseRequest" # type: ignore # Forward reference request: "ResponseRequest", # Forward reference
) -> AsyncGenerator[StreamEvent, None]: ) -> AsyncGenerator[StreamEvent, None]:
""" """
Coordinate streaming from agent to SSE events. Coordinate streaming from agent to SSE events.
@@ -435,18 +446,13 @@ class StreamingCoordinator:
elif item.type == "function_call": elif item.type == "function_call":
# Stream function call arguments # Stream function call arguments
# First chunk includes name # First chunk includes name
yield FunctionCallDelta( yield FunctionCallDelta(name=item.data["name"], delta="")
name=item.data["name"],
delta=""
)
# Stream arguments in chunks # Stream arguments in chunks
args = item.data["arguments"] args = item.data["arguments"]
chunk_size = 20 chunk_size = 20
for i in range(0, len(args), chunk_size): for i in range(0, len(args), chunk_size):
yield FunctionCallDelta( yield FunctionCallDelta(delta=args[i : i + chunk_size])
delta=args[i:i+chunk_size]
)
await asyncio.sleep(0.03) await asyncio.sleep(0.03)
yield FunctionCallDone() yield FunctionCallDone()
@@ -458,7 +464,7 @@ class StreamingCoordinator:
# Only stream the NEW text (delta) since last update # Only stream the NEW text (delta) since last update
if current_text.startswith(last_message_text): if current_text.startswith(last_message_text):
# Extract only the new portion # Extract only the new portion
delta_text = current_text[len(last_message_text):] delta_text = current_text[len(last_message_text) :]
if delta_text: if delta_text:
# Stream the delta text in chunks while preserving formatting # Stream the delta text in chunks while preserving formatting
@@ -466,17 +472,16 @@ class StreamingCoordinator:
chunk_size = 50 # characters per chunk chunk_size = 50 # characters per chunk
for i in range(0, len(delta_text), chunk_size): for i in range(0, len(delta_text), chunk_size):
chunk = delta_text[i:i+chunk_size] chunk = delta_text[i : i + chunk_size]
# Check stop sequences on full accumulated text # Check stop sequences on full accumulated text
stop_found, text_before_stop = self._check_stop_sequence( stop_found, text_before_stop = self._check_stop_sequence(
current_text, current_text, request.stop
request.stop
) )
if stop_found: if stop_found:
# Only emit remaining delta before stop # Only emit remaining delta before stop
remaining = text_before_stop[len(last_message_text):] remaining = text_before_stop[len(last_message_text) :]
if remaining: if remaining:
yield OutputTextDelta(delta=remaining) yield OutputTextDelta(delta=remaining)
yield OutputTextDone() yield OutputTextDone()
@@ -508,11 +513,12 @@ class StreamingCoordinator:
model=request.model, model=request.model,
status="completed", status="completed",
output=self._convert_output_items(output_items), output=self._convert_output_items(output_items),
usage=usage usage=usage,
) )
# Track conversation history (import here to avoid circular dependency) # Track conversation history (import here to avoid circular dependency)
from src.responses.service import _conversation_history from src.responses.service import _conversation_history
conversation_id = await _conversation_history.get_conversation_id(request) conversation_id = await _conversation_history.get_conversation_id(request)
await _conversation_history.add_response(conversation_id, final_response) await _conversation_history.add_response(conversation_id, final_response)
@@ -534,24 +540,30 @@ class StreamingCoordinator:
converted = [] converted = []
for item in items: for item in items:
if item.type == "message": if item.type == "message":
converted.append(MessageOutputItem( converted.append(
id=item.id, MessageOutputItem(
content=[OutputTextContent(**c) for c in item.data["content"]], id=item.id,
status=item.data.get("status", "completed") content=[OutputTextContent(**c) for c in item.data["content"]],
)) status=item.data.get("status", "completed"),
)
)
elif item.type == "reasoning": elif item.type == "reasoning":
converted.append(ReasoningOutputItem( converted.append(
id=item.id, ReasoningOutputItem(
summary=item.data["summary"], id=item.id,
status=item.data.get("status", "completed") summary=item.data["summary"],
)) status=item.data.get("status", "completed"),
)
)
elif item.type == "function_call": elif item.type == "function_call":
converted.append(FunctionCallOutputItem( converted.append(
id=item.id, FunctionCallOutputItem(
name=item.data["name"], id=item.id,
arguments=item.data["arguments"], name=item.data["name"],
status=item.data.get("status", "completed") arguments=item.data["arguments"],
)) status=item.data.get("status", "completed"),
)
)
return converted return converted
@@ -576,18 +588,10 @@ class StreamingCoordinator:
error_type = "internal_error" error_type = "internal_error"
code = 500 code = 500
return ErrorEvent( return ErrorEvent(error={"type": error_type, "message": str(error), "code": code})
error={
"type": error_type,
"message": str(error),
"code": code
}
)
def _check_stop_sequence( def _check_stop_sequence(
self, self, accumulated_text: str, stop_sequences: list[str] | None
accumulated_text: str,
stop_sequences: list[str] | None
) -> tuple[bool, str]: ) -> tuple[bool, str]:
""" """
Check if any stop sequence is encountered. Check if any stop sequence is encountered.
@@ -624,11 +628,7 @@ class StreamingCoordinator:
""" """
return len(text) // 4 return len(text) // 4
def _check_max_tokens( def _check_max_tokens(self, current_tokens: int, max_tokens: int | None) -> bool:
self,
current_tokens: int,
max_tokens: int | None
) -> bool:
""" """
Check if max tokens limit reached. Check if max tokens limit reached.
+3 -4
View File
@@ -2,9 +2,10 @@
Tests for Biographer capability registration. Tests for Biographer capability registration.
""" """
import pytest
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
import pytest
from src.agents.biographer.capability import ( from src.agents.biographer.capability import (
BIOGRAPHER_CAPABILITY, BIOGRAPHER_CAPABILITY,
get_biographer_capability, get_biographer_capability,
@@ -73,9 +74,7 @@ class TestBiographerRegistration:
"src.agents.biographer.capability.get_household_registry", "src.agents.biographer.capability.get_household_registry",
return_value=mock_registry, return_value=mock_registry,
): ):
with patch( with patch("src.agents.biographer.capability.get_biographer_agent") as mock_get_agent:
"src.agents.biographer.capability.get_biographer_agent"
) as mock_get_agent:
mock_agent = MagicMock() mock_agent = MagicMock()
mock_get_agent.return_value = mock_agent mock_get_agent.return_value = mock_agent
+3 -4
View File
@@ -2,9 +2,10 @@
Tests for Housekeeper capability registration. Tests for Housekeeper capability registration.
""" """
import pytest
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
import pytest
from src.agents.housekeeper.capability import ( from src.agents.housekeeper.capability import (
HOUSEKEEPER_CAPABILITY, HOUSEKEEPER_CAPABILITY,
get_housekeeper_capability, get_housekeeper_capability,
@@ -74,9 +75,7 @@ class TestHousekeeperRegistration:
"src.agents.housekeeper.capability.get_household_registry", "src.agents.housekeeper.capability.get_household_registry",
return_value=mock_registry, return_value=mock_registry,
): ):
with patch( with patch("src.agents.housekeeper.capability.get_housekeeper_agent") as mock_get_agent:
"src.agents.housekeeper.capability.get_housekeeper_agent"
) as mock_get_agent:
mock_agent = MagicMock() mock_agent = MagicMock()
mock_get_agent.return_value = mock_agent mock_get_agent.return_value = mock_agent
+2 -1
View File
@@ -2,9 +2,10 @@
Tests for the Core-API HTTP client. Tests for the Core-API HTTP client.
""" """
import pytest
from unittest.mock import AsyncMock, MagicMock from unittest.mock import AsyncMock, MagicMock
import httpx import httpx
import pytest
from src.agents.housekeeper.client import ( from src.agents.housekeeper.client import (
Area, Area,
+3 -4
View File
@@ -2,9 +2,10 @@
Tests for Librarian capability registration. Tests for Librarian capability registration.
""" """
import pytest
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
import pytest
from src.agents.librarian.capability import ( from src.agents.librarian.capability import (
LIBRARIAN_CAPABILITY, LIBRARIAN_CAPABILITY,
get_librarian_capability, get_librarian_capability,
@@ -67,9 +68,7 @@ class TestLibrarianRegistration:
"src.agents.librarian.capability.get_household_registry", "src.agents.librarian.capability.get_household_registry",
return_value=mock_registry, return_value=mock_registry,
): ):
with patch( with patch("src.agents.librarian.capability.get_librarian_agent") as mock_get_agent:
"src.agents.librarian.capability.get_librarian_agent"
) as mock_get_agent:
mock_agent = MagicMock() mock_agent = MagicMock()
mock_get_agent.return_value = mock_agent mock_get_agent.return_value = mock_agent
+8 -24
View File
@@ -137,9 +137,7 @@ class TestHybridSearch:
assert "docker" in result.keywords assert "docker" in result.keywords
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_hybrid_search_empty_results( async def test_hybrid_search_empty_results(self, client_with_mock, mock_httpx_client):
self, client_with_mock, mock_httpx_client
):
"""Test hybrid search with no results.""" """Test hybrid search with no results."""
mock_response = MagicMock() mock_response = MagicMock()
mock_response.json.return_value = { mock_response.json.return_value = {
@@ -445,9 +443,7 @@ class TestUpdateWikiPage:
assert page.tags == ["projects", "devops"] assert page.tags == ["projects", "devops"]
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_update_wiki_page_multiple_fields( async def test_update_wiki_page_multiple_fields(self, client_with_mock, mock_httpx_client):
self, client_with_mock, mock_httpx_client
):
"""Test updating multiple fields at once.""" """Test updating multiple fields at once."""
mock_response = MagicMock() mock_response = MagicMock()
mock_response.json.return_value = { mock_response.json.return_value = {
@@ -622,15 +618,11 @@ class TestExplicitUserContract:
@pytest.mark.asyncio @pytest.mark.asyncio
@pytest.mark.parametrize("method_name,kwargs", TENANT_SCOPED_METHODS) @pytest.mark.parametrize("method_name,kwargs", TENANT_SCOPED_METHODS)
async def test_user_from_context_is_sent_on_the_wire( async def test_user_from_context_is_sent_on_the_wire(self, method_name, kwargs):
self, method_name, kwargs
):
"""With no explicit user, the context user is resolved and sent.""" """With no explicit user, the context user is resolved and sent."""
client, mock_httpx = self._wire_client() client, mock_httpx = self._wire_client()
with patch( with patch("src.agents.librarian.client.get_user", return_value="llm_tester"):
"src.agents.librarian.client.get_user", return_value="llm_tester"
):
await getattr(client, method_name)(**kwargs) await getattr(client, method_name)(**kwargs)
assert self._sent_user(mock_httpx) == "llm_tester" assert self._sent_user(mock_httpx) == "llm_tester"
@@ -647,9 +639,7 @@ class TestExplicitUserContract:
@pytest.mark.asyncio @pytest.mark.asyncio
@pytest.mark.parametrize("method_name,kwargs", TENANT_SCOPED_METHODS) @pytest.mark.parametrize("method_name,kwargs", TENANT_SCOPED_METHODS)
async def test_empty_context_user_fails_before_any_request( async def test_empty_context_user_fails_before_any_request(self, method_name, kwargs):
self, method_name, kwargs
):
"""An empty resolved user raises before any bytes hit the wire.""" """An empty resolved user raises before any bytes hit the wire."""
client, mock_httpx = self._wire_client() client, mock_httpx = self._wire_client()
@@ -698,9 +688,7 @@ class TestExplicitUserContract:
from src.core import config as config_module from src.core import config as config_module
from src.core.config import Environment from src.core.config import Environment
monkeypatch.setattr( monkeypatch.setattr(config_module.config, "ENVIRONMENT", Environment.DEVELOPMENT)
config_module.config, "ENVIRONMENT", Environment.DEVELOPMENT
)
client, mock_httpx = self._wire_client() client, mock_httpx = self._wire_client()
await getattr(client, method_name)(user=explicit_user, **kwargs) await getattr(client, method_name)(user=explicit_user, **kwargs)
@@ -708,16 +696,12 @@ class TestExplicitUserContract:
assert self._sent_user(mock_httpx) == "llm_tester" assert self._sent_user(mock_httpx) == "llm_tester"
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_explicit_production_tenant_passes_through_in_prod( async def test_explicit_production_tenant_passes_through_in_prod(self, monkeypatch):
self, monkeypatch
):
"""In production the production tenant is sent unchanged.""" """In production the production tenant is sent unchanged."""
from src.core import config as config_module from src.core import config as config_module
from src.core.config import Environment from src.core.config import Environment
monkeypatch.setattr( monkeypatch.setattr(config_module.config, "ENVIRONMENT", Environment.PRODUCTION)
config_module.config, "ENVIRONMENT", Environment.PRODUCTION
)
client, mock_httpx = self._wire_client() client, mock_httpx = self._wire_client()
await client.hybrid_search("q", user="jpmschweitzer") await client.hybrid_search("q", user="jpmschweitzer")
+5 -15
View File
@@ -91,9 +91,7 @@ class TestBoundedRetries:
mock_httpx.post.side_effect = httpx.ConnectError("Connection refused") mock_httpx.post.side_effect = httpx.ConnectError("Connection refused")
with pytest.raises(httpx.ConnectError): with pytest.raises(httpx.ConnectError):
await client_with_mock.create_wiki_page( await client_with_mock.create_wiki_page(title="T", path="/t", content="c", user="u")
title="T", path="/t", content="c", user="u"
)
assert mock_httpx.post.call_count == 1 assert mock_httpx.post.call_count == 1
@@ -103,9 +101,7 @@ class TestBoundedRetries:
mock_httpx.post.side_effect = httpx.ConnectError("Connection refused") mock_httpx.post.side_effect = httpx.ConnectError("Connection refused")
with pytest.raises(httpx.ConnectError): with pytest.raises(httpx.ConnectError):
await client_with_mock.smart_create_wiki_page( await client_with_mock.smart_create_wiki_page(topic="T", tags=["x"], user="u")
topic="T", tags=["x"], user="u"
)
assert mock_httpx.post.call_count == 1 assert mock_httpx.post.call_count == 1
@@ -184,9 +180,7 @@ class TestModelRetryEscalation:
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_read_tool_raises_model_retry_on_transport_error(self): async def test_read_tool_raises_model_retry_on_transport_error(self):
mock_client = AsyncMock() mock_client = AsyncMock()
mock_client.hybrid_search.side_effect = httpx.ConnectError( mock_client.hybrid_search.side_effect = httpx.ConnectError("Connection refused")
"Connection refused"
)
with self._patched_client(mock_client): with self._patched_client(mock_client):
with pytest.raises(ModelRetry): with pytest.raises(ModelRetry):
@@ -219,14 +213,10 @@ class TestModelRetryEscalation:
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_write_tool_never_raises_model_retry(self): async def test_write_tool_never_raises_model_retry(self):
mock_client = AsyncMock() mock_client = AsyncMock()
mock_client.create_wiki_page.side_effect = httpx.ConnectError( mock_client.create_wiki_page.side_effect = httpx.ConnectError("Connection refused")
"Connection refused"
)
with self._patched_client(mock_client): with self._patched_client(mock_client):
result = await create_wiki_page( result = await create_wiki_page(title="T", path="/t", content="c", tags=["x"])
title="T", path="/t", content="c", tags=["x"]
)
assert "unable" in result assert "unable" in result
assert "Connection refused" not in result assert "Connection refused" not in result
+12 -28
View File
@@ -83,9 +83,7 @@ class TestHybridRAGContract:
assert source in SOURCE_ICONS, f"no icon for source '{source}'" assert source in SOURCE_ICONS, f"no icon for source '{source}'"
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_context_maps_to_formatted_context( async def test_context_maps_to_formatted_context(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""Top-level 'context' field maps to formatted_context.""" """Top-level 'context' field maps to formatted_context."""
response = await client_with_recorded_response.hybrid_search( response = await client_with_recorded_response.hybrid_search(
"home server infrastructure", user="testuser" "home server infrastructure", user="testuser"
@@ -94,9 +92,7 @@ class TestHybridRAGContract:
assert response.formatted_context != "" assert response.formatted_context != ""
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_keywords_and_synonyms_from_dict( async def test_keywords_and_synonyms_from_dict(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""keywords is a dict: core_keywords + nested synonyms map.""" """keywords is a dict: core_keywords + nested synonyms map."""
response = await client_with_recorded_response.hybrid_search( response = await client_with_recorded_response.hybrid_search(
"home server infrastructure", user="testuser" "home server infrastructure", user="testuser"
@@ -124,9 +120,7 @@ class TestHybridRAGContract:
assert len(response.related_dossiers) == len(set(response.related_dossiers)) assert len(response.related_dossiers) == len(set(response.related_dossiers))
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_payload_never_sends_zero_limits( async def test_payload_never_sends_zero_limits(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""The live service 422s on limits < 1; disabled legs use enable_* flags.""" """The live service 422s on limits < 1; disabled legs use enable_* flags."""
await client_with_recorded_response.hybrid_search( await client_with_recorded_response.hybrid_search(
"home server infrastructure", "home server infrastructure",
@@ -151,24 +145,18 @@ class TestHybridRAGContract:
assert config["enable_volatile"] is False assert config["enable_volatile"] is False
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_user_always_sent_as_query_param( async def test_user_always_sent_as_query_param(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""The tenant is always sent explicitly - library-desk is removing """The tenant is always sent explicitly - library-desk is removing
its server-side default, so a missing user would 422.""" its server-side default, so a missing user would 422."""
await client_with_recorded_response.hybrid_search( await client_with_recorded_response.hybrid_search(
"home server infrastructure", user="testuser" "home server infrastructure", user="testuser"
) )
params = client_with_recorded_response._client.post.call_args.kwargs[ params = client_with_recorded_response._client.post.call_args.kwargs["params"]
"params"
]
assert params["user"] == "testuser" assert params["user"] == "testuser"
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_source_counts_and_timing_parsed( async def test_source_counts_and_timing_parsed(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""source_counts and timing map into the response model.""" """source_counts and timing map into the response model."""
response = await client_with_recorded_response.hybrid_search( response = await client_with_recorded_response.hybrid_search(
"home server infrastructure", user="testuser" "home server infrastructure", user="testuser"
@@ -178,9 +166,7 @@ class TestHybridRAGContract:
assert response.timing.get("total_ms", 0) > 0 assert response.timing.get("total_ms", 0) > 0
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_source_status_absent_is_tolerated( async def test_source_status_absent_is_tolerated(self, client_with_recorded_response):
self, client_with_recorded_response
):
"""Recorded response predates source_status/degraded - defaults apply.""" """Recorded response predates source_status/degraded - defaults apply."""
response = await client_with_recorded_response.hybrid_search( response = await client_with_recorded_response.hybrid_search(
"home server infrastructure", user="testuser" "home server infrastructure", user="testuser"
@@ -218,7 +204,9 @@ class TestHybridRAGContract:
assert response.source_status["volatile"] == "disabled" assert response.source_status["volatile"] == "disabled"
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_tool_renders_no_unknown_results(self, client_with_recorded_response, monkeypatch): async def test_tool_renders_no_unknown_results(
self, client_with_recorded_response, monkeypatch
):
"""The hybrid_search tool renders real sources and non-zero scores.""" """The hybrid_search tool renders real sources and non-zero scores."""
class _Factory: class _Factory:
@@ -231,9 +219,7 @@ class TestHybridRAGContract:
async def __aexit__(self, *args): async def __aexit__(self, *args):
return None return None
monkeypatch.setattr( monkeypatch.setattr("src.agents.librarian.tools.LibraryDeskClient", _Factory())
"src.agents.librarian.tools.LibraryDeskClient", _Factory()
)
output = await hybrid_search("home server infrastructure") output = await hybrid_search("home server infrastructure")
@@ -278,9 +264,7 @@ class TestCoverageNote:
def test_wiki_leg_absence_is_not_degradation(self): def test_wiki_leg_absence_is_not_degradation(self):
"""vector/graph missing from top-N counts is healthy ranking, not outage.""" """vector/graph missing from top-N counts is healthy ranking, not outage."""
response = self._response( response = self._response(source_counts={"web": 2, "documents": 1, "volatile": 1})
source_counts={"web": 2, "documents": 1, "volatile": 1}
)
note = _coverage_note( note = _coverage_note(
response, include_web=True, include_documents=True, include_volatile=True response, include_web=True, include_documents=True, include_volatile=True
) )
+1 -3
View File
@@ -35,9 +35,7 @@ def _nullable_anyof_paths(schema: object, path: str = "") -> list[str]:
class TestLibrarianToolSchemas: class TestLibrarianToolSchemas:
"""All registered librarian tools emit Ollama-safe parameter schemas.""" """All registered librarian tools emit Ollama-safe parameter schemas."""
@pytest.mark.parametrize( @pytest.mark.parametrize("tool_func", LIBRARIAN_TOOLS, ids=lambda f: f.__name__)
"tool_func", LIBRARIAN_TOOLS, ids=lambda f: f.__name__
)
def test_no_nullable_anyof_in_schema(self, tool_func): def test_no_nullable_anyof_in_schema(self, tool_func):
schema = Tool(tool_func).function_schema.json_schema schema = Tool(tool_func).function_schema.json_schema

Some files were not shown because too many files have changed in this diff Show More