A live capture against claude 2.1.150 (both --print and the interactive stream-json transport clide uses) confirms --include-partial-messages emits the in-progress reply as stream_event envelopes wrapping Anthropic streaming deltas — NOT assistant events with partial:true, which is what T-168 assumed, so that handler never fired and streaming was inert. Replace it: accumulate content_block_delta text per message id (tracked from message_start, since deltas carry no id) and emit a placeholder under a stable partial-<id> uuid the controller upserts in place; the matching single-text-block assistant event reuses that uuid to finalize, while tool_use / thinking blocks keep their own uuids and append in order. Tests rewritten against the captured shape; spike doc records it. T-184. Co-Authored-By: Claude <noreply@anthropic.com>
16 KiB
Spike: Claude Code stream-json control protocol (T-165 / T-166)
Pinned to: claude 2.1.150. Undocumented, version-drifting internal contracts
(per D-75/D-77) — re-validate on a CC bump, keyed off the transcript/event version
field and claude_code_version in the init event.
Method: drove the real claude binary as a subprocess over stdin/stdout
(docs/spikes/ driver scripts were throwaway; the captured logs are the source of
truth) and read the shipped implementation directly — the CLI is a ~238 MB node
SEA at ~/.local/share/claude/versions/2.1.150; strings on it exposes the zod
schemas and the request/response builders. Every shape below was confirmed live
(a real prompt round-tripped) unless marked otherwise. Cost a handful of paid turns.
0. TL;DR for the implementation
Spawn:
claude --input-format stream-json --output-format stream-json --verbose \
--permission-prompt-tool stdio [ --resume <id> | --session-id <id> ]
--permission-prompt-tool stdiois mandatory to receive permission prompts. Without it, any tool that resolves to "ask" is auto-denied (no prompt reaches the client) — confirmed: aWritecame back aspermission_denials+ anis_errortool_result "you haven't granted it yet", never a control request.stdiois a hidden value (not in--help); the SDK uses it (f.sdkUrl?"stdio":…).- Read stdout as line-delimited JSON. Most lines are normal stream events
(
system/assistant/user/result/rate_limit_event); some arecontrol_requests you must answer or Claude hangs. - Send a prompt as
{"type":"user","message":{"role":"user","content":"…"}}. stream-json does not echo your user messages back unless--replay-user-messages— local-echo them yourself.
1. Output event stream (stdout)
One JSON object per line. Types seen:
system/subtype:"hook_started"|"hook_progress"|"hook_response"— session hooks (e.g. SessionStart). Informational; ignore for rendering.system/subtype:"init"— the config goldmine. Carriessession_id,cwd,model(claude-opus-4-7[1m]),permissionMode,tools[],mcp_servers[],slash_commands[],skills[],agents[],output_style,claude_code_version,apiKeySource,plugins[]. (This is a strong source forClaudeConfig/ status.)assistant—{message:{model,id,content:[…],usage:{…}}, uuid, session_id, request_id}. Content blocks are emitted as separateassistantevents sharing onemessage.id(e.g. atextblock, then atool_useblock). Block shapes are identical to the transcript:text/thinking(+signature) /tool_use(id,name,input).message.usagecarriesinput_tokens+cache_read_input_tokens+cache_creation_input_tokens(→ context-token count).user— two flavours: (a) tool results Claude received —message.content:[{type:"tool_result",tool_use_id,content,is_error}]plus a top-leveltool_use_result; (b) harness-injected user messages — a skill load (Skilltool → text begins"Base directory for this skill:"), a slash-command expansion, or a system reminder. Injected ones carryisSynthetic: trueat the top level (the transcript usesisMetainstead). Verified by boundary test: the inject only appears once theSkilltool is actually invoked (it's auto-allowed, no prompt); if Claude just runs a command inferred from the slash text, no inject is emitted. clide flags these (UserMessage.injected) and renders them as a muted, collapsed "context" card, not a blue "you" message.result— terminal turn summary:result(final text),usage,total_cost_usd,permission_denials[],num_turns,modelUsage.<model>.contextWindow(e.g.1000000) andmaxOutputTokens— i.e. the context-window size IS exposed here (relevant to the T-158 budget gap; the remaining-budget % still is not).rate_limit_event—{rate_limit_info:{status,resetsAt,rateLimitType,…}}.stream_event— only with--include-partial-messages(T-184). Wraps the raw Anthropic streaming deltas as they arrive:{type:"stream_event", event:{…}}whereevent.typeismessage_start(carriesmessage.id) →content_block_start→content_block_delta(delta:{type:"text_delta",text}/thinking_delta/input_json_delta, withindex) →content_block_stop→message_delta→message_stop. The full per-blockassistantevent still arrives interleaved. Verified against 2.1.150 in BOTH--printand the interactive--input-format stream-jsontransport — the flag's--helpclaim that it "only works with --print" is wrong; it streams in interactive mode too.content_block_deltaevents do NOT carry a message id, so track the id from the precedingmessage_start. (This is the real shape; the earlier guess ofassistant+partial:truewas wrong —StreamJsonSessionstreams text fromstream_eventand finalises via the matchingassistantevent.)
The existing parseTranscriptChunk (transcript_reader.dart) parses the
assistant/user events as-is (same message.content shapes; missing
uuid/timestamp default harmlessly). Permission mode comes from init, not a
permission-mode record.
2. Control protocol (bidirectional, same stdin/stdout)
Envelope
Claude → client:
{"type":"control_request","request_id":"<uuid>","request":{"subtype":"…", …}}
Client → Claude (the reply):
{"type":"control_response","response":{"subtype":"success","request_id":"<uuid>","response":{…}}}
On failure use {"subtype":"error","request_id":"…","error":"…"}. request_id
must echo exactly. An unanswered control_request hangs the turn.
can_use_tool (the permission ask) — confirmed live
Request:
{"type":"control_request","request_id":"70c8…","request":{
"subtype":"can_use_tool","tool_name":"Write","display_name":"Write",
"input":{"file_path":"…","content":"…"},
"description":"banana.txt",
"permission_suggestions":[{"type":"setMode","mode":"acceptEdits","destination":"session"}],
"tool_use_id":"toolu_…"}}
ALLOW — updatedInput is required (a bare {"behavior":"allow"} is rejected
with a ZodError: "updatedInput: expected record, received undefined"). Echo the
input back unchanged to allow as-is, or modify it to alter the call:
{"type":"control_response","response":{"subtype":"success","request_id":"70c8…",
"response":{"behavior":"allow","updatedInput":{ …the tool input… }}}}
DENY — message is required:
{"type":"control_response","response":{"subtype":"success","request_id":"70c8…",
"response":{"behavior":"deny","message":"User declined."}}}
(Decision zod union also allows optional updatedPermissions on allow.)
AskUserQuestion — confirmed live (the non-obvious one)
AskUserQuestion is not answered with a tool_result (that path is rejected as a
dismissal — it even shows up in permission_denials). It is permission-gated and
answered through the same can_use_tool channel. Its input is exactly clide's
own AskUserQuestion shape:
{questions:[{question,header,multiSelect,options:[{label,description}]}]}.
Return the user's choice by injecting an answers map into updatedInput:
answers : record(question-text → chosen-label) (multi-select = comma-separated labels).
{"behavior":"allow","updatedInput":{
"questions":[ …echoed… ],
"answers":{"Do you prefer cats or dogs?":"Dogs"}}}
Confirmed: Claude then reported "You chose Dogs"; the resulting tool_use_result was
{questions:[…],answers:{"Do you prefer cats or dogs?":"Dogs"}}. Leaving answers
empty makes Claude say "your selection didn't come through".
Other control_request subtypes (from the binary's zod schemas)
Client→Claude requests it accepts: initialize, interrupt,
set_permission_mode {mode}, mcp_message {server_name,message}.
Claude→client requests you may receive: can_use_tool, hook_callback {callback_id,input,tool_use_id?}. Answer every inbound control_request (even
unknown subtypes — reply error "Unsupported…") or the turn stalls. There is also a
control_cancel_request for in-flight cancellation.
initialize handshake — optional, but a config goldmine
Sending it first is not what enables permissions (the stdio flag does that —
verified: initialize-then-Write was still auto-denied). But the response is rich:
// → {"type":"control_request","request_id":"init-1","request":{"subtype":"initialize","hooks":{},"sdkMcpServers":[]}}
// ← response.response = {commands:[{name,description,argumentHint,aliases?}…],
// agents:[…], models:[…], output_style, available_output_styles,
// account:{email,organization,subscriptionType,apiProvider}, pid}
commands[] carries descriptions + argumentHints the init event's bare
slash_commands[] lacks — better source for the typeahead (T-152) and ClaudeConfig.
3. Implications for the tickets
- T-165: spawn with the flags in §0; feed
parseTranscriptChunkfrom theassistant/userevents; pullpermissionMode/modelfrominit; local-echo user sends. Isolate all protocol framing in one module (drift-containment, D-77). - T-166: wire
can_use_tool→ native prompt card (allow/deny, usingdisplay_name/description/permission_suggestions); AskUserQuestion → option picker, returningupdatedInput.answers. Answer every control_request. - ClaudeConfig (D-76): prefer the
initializeresponse'scommands[](has descriptions) over theinitslash_commands[].
4. MCP vs the control channel — which layer does what
These are complementary, not competing — a recurring point of confusion:
- The conversation + permissions + AskUserQuestion ride the stream-json control
channel (
can_use_toolover stdin/stdout). This is the SDK-blessed transport and whatcanUseToolmaps to. T-165/T-166 use it. MCP is not involved here. - MCP is how you give Claude extra tools/capabilities. That's exactly what the
VSCode/JetBrains extensions do: they host an MCP server exposing IDE features to
Claude (
mcp__ide__getDiagnostics,mcp__ide__executeCode, openDiff) — they do not route the conversation or permissions through MCP; permissions still go through the CLI/control channel. D-77's clide-hosted MCP broker (T-170) is the same idea: provide team messaging / task-sync tools to the agents. - One overlap exists:
--permission-prompt-toolaccepts either the specialstdiovalue (control channel — what we use) or an MCP tool name (Claude calls your MCP tool to get the allow/deny). So permissions could be routed through MCP — but it's strictly more machinery thanstdio, needs a server, and loses the structuredpermission_suggestions/display_namethe control request carries. Recommendation: keep permissions onstdio; reserve MCP for capability/tool provision (IDE context later, the team broker in T-170).
5. Resilience — the stdio channel is brittle; the fallback menu (no empty slate)
Decision (this spike): carry the conversation + permissions + AskUserQuestion on
the stdio control channel (--permission-prompt-tool stdio + can_use_tool).
Rationale: most direct — no extra process, no loopback "networking", and Claude hands
us structured prompt metadata for free. But it is an undocumented internal contract;
Anthropic can change or remove it on any version bump. This section exists so that
if it shifts we start from a researched menu, not a blank page.
How we'd detect a break (canaries):
- Pin
claude_code_version(here 2.1.150); diff it on every CC upgrade. - Symptoms of regression: tools auto-denied despite
--permission-prompt-tool stdio;can_use_toolrequests missing/renamed;ZodErrortool_results rejecting ourcontrol_response(e.g. theupdatedInput-required quirk changing). - The captured spike logs (this doc's source) double as regression fixtures — wire them into the parser/control-handler tests so a CC bump that breaks shapes goes red.
Fallback menu, in rough preference order if stdio is withdrawn:
- MCP permission tool —
--permission-prompt-tool mcp__clide__approveagainst the clide-hosted MCP server (T-170 builds that server anyway). Claude calls our MCP tool to get the allow/deny. More layers, but MCP is a stable, public protocol — the strongest fallback precisely because the broker will already exist. - Permission modes — degrade to
--permission-mode acceptEdits/dontAsk/bypassPermissionsfor a reduced-fidelity UX (no per-call prompt) as a stopgap. Public, stable flags; buys time to build a better path. - Agent SDK — if Anthropic ships/keeps a stable public
canUseToolSDK surface (TS/Python), shell out to or port it. Supported, but a heavier dep + language seam. - ACP (Agent Client Protocol) — if the ecosystem converges on ACP
(
session/request_permission, Zed's path) as the blessed editor-integration standard, adopt its adapter. Open standard; would reframe the whole transport.
Containment: all protocol framing lives behind one module (per D-77), so swapping transports is a one-seam change. Re-capture fixtures per pinned version.
6. Hosting an in-process MCP server over the control channel (T-170 — VERIFIED)
Verified live against 2.1.150 (2026-05-25): clide can host an MCP server whose tools
claude calls, entirely over the stream-json control channel — no subprocess, no
--mcp-config file, no socket. This is the cleanest fit for the single-process
guardrail and is what the team broker (T-170) is built on.
Registration — initialize handshake only (no --mcp-config needed). List the
server name(s) in the initialize control_request's sdkMcpServers:
// → {"type":"control_request","request_id":"init-1","request":{
// "subtype":"initialize","hooks":{},"sdkMcpServers":["clide-team"]}}
A run with no --mcp-config flag at all but this handshake worked end-to-end — the
flag is not required for SDK (in-process) servers. The server name surfaces tools to the
model as mcp__<server>__<tool> (e.g. mcp__clide-team__ping).
Handshake claude then drives (inbound mcp_message control_requests). For each,
the message is a JSON-RPC object; reply with a control_response carrying the JSON-RPC
result under response.response.mcp_response:
// ← {"type":"control_request","request_id":"<rid>","request":{
// "subtype":"mcp_message","server_name":"clide-team",
// "message":{"method":"initialize","params":{"protocolVersion":"2025-11-25",…},"jsonrpc":"2.0","id":0}}}
// → {"type":"control_response","response":{"subtype":"success","request_id":"<rid>",
// "response":{"mcp_response":{"jsonrpc":"2.0","id":0,"result":{
// "protocolVersion":"2025-11-25","capabilities":{"tools":{"listChanged":false}},
// "serverInfo":{"name":"clide-team","version":"0.0.1"}}}}}}
Sequence observed: initialize → notifications/initialized (no id; still answer it)
→ tools/list → (model calls a tool) → tools/call. The tools/call message:
{"method":"tools/call","params":{"name":"ping","arguments":{…},"_meta":{"claudecode/toolUseId":…,"progressToken":…}},"jsonrpc":"2.0","id":2};
answer with mcp_response.result = {"content":[{"type":"text","text":…}],"isError":false}.
SDK MCP tool calls ARE permission-gated. Before the tools/call, claude sends a
normal can_use_tool for mcp__clide-team__ping (with permission_suggestions →
addRules). So the broker's tools flow through the same allow/deny path as any tool —
no special-casing needed; the existing can_use_tool handler covers them.
Provenance: two live capture runs (init-strings, no --mcp-config; and mcpconfig)
both completed the full round-trip returning pong-from-clide. The raw logs aren't
committed (the initialize response embeds account email/org); the contract is instead
pinned in the transport tests (MCP server hosting (T-170)).
Not done / open
hook_callbackround-trip not exercised live (shape from the binary only).- Image/file paste intake over stream-json
contentblocks not tested here. - Remaining-usage budget % still not exposed (only
contextWindowsize, inresult).