diff --git a/build/icon.ico b/build/icon.ico index e399ce0..f42794d 100644 Binary files a/build/icon.ico and b/build/icon.ico differ diff --git a/build/icon.png b/build/icon.png index a7a22a7..ba9c2f4 100644 Binary files a/build/icon.png and b/build/icon.png differ diff --git a/build/icons/256x256.png b/build/icons/256x256.png index 60f9e6d..eca494e 100644 Binary files a/build/icons/256x256.png and b/build/icons/256x256.png differ diff --git a/build/icons/512x512.png b/build/icons/512x512.png index 807d687..2f47e20 100644 Binary files a/build/icons/512x512.png and b/build/icons/512x512.png differ diff --git a/docs/agent-activity-overlay.md b/docs/agent-activity-overlay.md index 088f3bf..96d8c46 100644 --- a/docs/agent-activity-overlay.md +++ b/docs/agent-activity-overlay.md @@ -54,7 +54,35 @@ This projection changes only CrewCode visualization. Tool availability and Claud Human-input request cards are independent of this preference. Approvals, questions, editor requests, and notifications must always render because the provider may be paused waiting for the response. Request rendering therefore takes precedence over both the Todo preference and a previously dismissed todo card. -CrewCoder-mode `crewcoder_clarify` / `crewcoder_propose_plan` is a session workflow gate, not a tool-permission pause. After those tools settle, the overlay shows a dedicated clarification or **Approve plan** card. Approve sends `/approve-plan` as a normal prompt (or follow-up if the turn is still running). It must never reuse Allow/Deny on a permission card — `/approve` is still only for pending tool-call grants. A later user message hides the card: `/approve-plan` or a short CrewCoder approval continues implementation, and a revision such as `yes, but also add logging` waits for the next `crewcoder_propose_plan`. Answering a clarification is not plan approval. The Todo preference must not hide this card. +CrewCoder-mode `crewcoder_clarify` / `crewcoder_propose_plan` is a session workflow gate, not a tool-permission pause. After those tools settle, the overlay shows a dedicated clarification or **Approve plan** card. Approve sends `/approve-plan` as a normal prompt (or follow-up if the turn is still running). It must never reuse Allow/Deny on a permission card — `/approve` is still only for pending tool-call grants. A later user message hides the card: `/approve-plan` or a short CrewCoder approval continues implementation, and a revision such as `yes, but also add logging` waits for the next `crewcoder_propose_plan`. Answering a clarification is not plan approval. The Todo preference must not hide this card. The clarification card has its own reply box and **send reply** button; the reply is sent as an ordinary prompt (follow-up while running) through the same surface's send path as Approve. Surfaces that cannot send a prompt fall back to "Reply in the composer". + +## Provider question cards + +Structured provider questions pause the turn and render in the shared `AgentRequestCard` (inline chat overlay, Crew lanes/timeline, supervisor, Mission Control, menulet). The card layout follows the question shape: + +| Shape | Card | +| --- | --- | +| options only (`kind: 'select'`) | one button per option; a click answers immediately. 2–4 short description-free options (Yes/No) sit on one row | +| options + free text (`kind: 'prompt'` with options) | option buttons plus an "or type your own answer…" input and **send reply** | +| free text only (`prompt` / `editor`) | input (Enter) or textarea (Ctrl/Cmd+Enter) plus **send reply** | +| multi-select (`multiple: true`) | toggle buttons answered with `optionIds` in option order, plus optional typed text | +| secret (`secret: true`) | masked input | + +**send reply** stays disabled until there is a non-empty answer. After a click the card locks (`sending…`) and unlocks only on an observed transport failure; success is confirmed by `user_request_resolved`. Cancel is an explicit cancel, never an empty answer. + +Provider mapping — only structured, provider-issued questions reach this card: + +| Provider | Native question | Free text | +| --- | --- | --- | +| Claude | `AskUserQuestion` via `canUseTool`, one card per question | always, matching Claude's built-in "Other"; `custom: false` / `allowFreeform: false` opts out | +| Codex | app-server `item/tool/requestUserInput` (legacy `tool/requestUserInput` alias), one card per question | when the question sets `isOther` or has no options | +| OpenCode | `question` SSE request | unless `custom: false` | +| Pi | extension UI `select` / `input` / `editor` | per method | +| CrewCoder | `crewcoder_clarify` workflow card (above) | reply box | + +Codex answers are returned as `{ answers: { [questionId]: { answers: string[] } } }`. A cancelled, empty, or out-of-contract answer (typed text where Codex offered no "Other") fails the whole request with an explicit JSON-RPC error; Codex never receives a fabricated or empty answer. When Codex settles a request itself (`serverRequest/resolved`, e.g. non-blocking questions) or the app-server exits, the bridge aborts the pending `RequestUserFn` signal: the promise settles as `cancel` and the card is retracted with `user_request_resolved`. + +Plain-text questions at the end of an agent reply ("Should I continue?") are not requests: the turn has already ended and nothing is paused, so they are answered in the composer. CrewCode must not guess questions from assistant prose or fabricate a request card for them. Codex MCP `mcpServer/elicitation/request` forms are not yet mapped. ## Surfaces diff --git a/docs/context-usage-meter.md b/docs/context-usage-meter.md new file mode 100644 index 0000000..8126cd4 --- /dev/null +++ b/docs/context-usage-meter.md @@ -0,0 +1,105 @@ +# Context usage meter + +The chat meter divides the latest observed context occupancy by the active +context window. The numerator is the provider's current prompt/context size, +not cumulative billing tokens for the entire conversation. + +Clicking the context pill opens a compact usage summary. When the provider +supplies token details, **Token logs** opens a scrollable right sidebar. The +sidebar can be closed with its close button, the backdrop, or Escape. Claude +provides context categories; Codex, OpenCode, and Grok provide request or turn +token counters, while CrewCoder provides session counters and its latest context +input reading. These rows can overlap and must not be added together. +If a provider reports token counts before a context window is known, the usage +strip shows a direct **Token logs** button and the sidebar marks context usage +as unavailable instead of showing a fabricated percentage. + +Codex token logs show the latest app-server request's input and output, plus +cached input and reasoning output when reported. OpenCode logs show its latest +assistant message's input, output, reasoning, cache read, and cache write +counts. Grok logs show latest-call input/output and, when present, cache read, +reasoning, and cumulative turn input/output. These providers do not expose +Claude-style system/tool/message context categories through these usage events; +the sidebar labels their rows as provider token counts rather than claiming +they sum to context occupancy. + +CrewCoder token logs use the ACP `session/prompt` result's namespaced +`crewcoder/usage` summary, falling back to its top-level `usage` mirror. They +show the latest context input (`lastInputTokens`) separately from cumulative +session input, output, total, cached input, cache write, and reasoning counts +when those fields are reported. The session counters can span multiple models +and requests and do not sum to the current context reading. If CrewCoder omits +`lastInputTokens`, cumulative session totals remain visible in Token logs but +do not become a context occupancy estimate. + +## Window sources + +Use these sources in order: + +1. A full window reported for the active session or turn by the provider bridge. +2. A full window from the provider's model catalog, when it supplies one. +3. A static model-family fallback, when no provider value is available. + +Codex app-server usage notifications include `modelContextWindow`. For known +models, this can be smaller than the model's full window because it is the +effective request prompt budget. CrewCode shows the known full model capacity +as the meter denominator and labels the reported smaller value separately as +**Codex prompt budget**. For unknown models, the reported value remains the +only available window and drives the meter. A reported value larger than model +metadata also takes precedence. The full model window can come from provider +catalog metadata or CrewCode's static model fallback; the latter can be stale +if a custom Codex configuration changes capacity. Compaction-drop inference +uses Codex's reported prompt budget as its threshold denominator even when the +displayed model window is larger. +The latest request's input tokens are also treated as an authoritative context +reading, so a decrease after native compaction is not replaced by an old floor. +The app-server `model/list` response discovers models but does not document a +context-window field, so model discovery alone cannot supply the Codex meter. + +Claude's SDK `getContextUsage()` supplies the active `maxTokens` (or +`rawMaxTokens`) window. The meter sums its active categories and excludes +deferred categories, unused space, and compaction headroom using each SDK row's +`kind`, with legacy category handling for older Claude binaries. CrewCode +requests `detail: 'full'`, since `summary` uses local estimates rather than +the token-counted `/context` breakdown. The categorized SDK reading is the +source of occupancy; result-message billing usage is separate and aggregates +multiple requests. The static model table can +still supply a window but cannot supply occupancy. Claude's `--help` does not +provide a model-window catalog. + +CrewCode asks `getContextUsage()` after each assembled assistant message while +the SDK query is active. It retries at the result only if no valid reading has +arrived, because a result-time call can race query shutdown and a full reading +can involve token-counting API requests. If no control reading succeeds for that turn, the latest +individual assistant message's input and cache token counts provide a bounded +request-context fallback. Only when neither current-turn source exists does +CrewCode reuse an earlier measured reading; with no reading at all, the meter +stays unknown. Result-message usage remains billing-only because it aggregates +multiple requests in one turn. + +CrewCoder ACP usage can also include a live `contextWindow` and +`lastInputTokens`; its bridge passes those through to the meter. Providers +that report no usable window can show no percentage if no catalog or static +fallback exists. A fallback percentage is an estimate, not a provider reading. + +The latest usage snapshot is persisted for resumed chats. A new provider +report replaces its window; occupancy normalization and compaction handling +are described in [Conversation storage](conversation-storage.md). + +## Automatic compaction + +Claude's `compact_boundary`, Codex app-server's `contextCompaction` completion, +CrewCoder's `_crewcoder/compaction_update`, and OpenCode's matching +`session.compacted` SSE event are native completion signals. They clear the +previous context reading in the main process, its persisted snapshot, and the +latest visible usage strip. The meter remains unknown until a new provider +usage reading arrives; a pre-compaction reading must not be replayed at turn +end. Historical provider token counters remain in Token logs, while old +Claude context categories are cleared. OpenCode's event is accepted only for +the active session. + +When no native boundary is observed, a high-occupancy to large-drop usage +reading can detect compaction for Codex, CrewCoder, OpenCode, and Grok. The +smaller observed reading becomes the new context baseline. Providers without +an observed boundary or trustworthy absolute context reading are not inferred +from silence or a completed turn. diff --git a/docs/conversation-storage.md b/docs/conversation-storage.md index ea226d5..213f948 100644 --- a/docs/conversation-storage.md +++ b/docs/conversation-storage.md @@ -85,6 +85,17 @@ and appends a visible compact-summary card, but keeps the full rich display transcript, provider session id, and live bridge. If CrewCoder reports a skipped small-session compact, CrewCode leaves replay and usage state unchanged. +Provider-native automatic compaction is also observable. Direct Codex maps its +app-server `contextCompaction` item lifecycle (plus the legacy +`thread/compacted` completion) to the same meter. CrewCoder forwards nested +Codex and Claude native boundaries through `_crewcoder/compaction_update`. +Claude maps `compact_boundary`, and OpenCode maps the selected session's +`session.compacted` SSE event. A native completion clears the previous live +occupancy, its persisted snapshot, and the latest visible usage strip until +new usage arrives. Native boundaries win; live-occupancy drop inference is only +a fallback for Codex, CrewCoder, OpenCode, or Grok turns where no native +boundary was observed, so a single compaction never creates duplicate cards. + ### `summary-reset` flow (pi / hermes / older CrewCoder) Native-session providers keep their context server-side and expose no compaction RPC, so we cannot shrink it directly. Instead: diff --git a/docs/crewcoder-provider.md b/docs/crewcoder-provider.md index 2a6d948..5fe762b 100644 --- a/docs/crewcoder-provider.md +++ b/docs/crewcoder-provider.md @@ -110,9 +110,16 @@ bridge, or seed a replacement session after a successful native compact. CrewCoder's compaction update is an additive namespaced ACP extension carrying started/completed/failed status, automatic intent, progress, and a human-readable -message. The bridge treats it as authoritative and does not also infer -compaction from the later context-token drop. Automatic updates omit the summary -body; the compacted summary remains only in CrewCoder's durable session. +message. It covers both CrewCoder's durable-session compaction and compaction +reported by a nested native provider. Codex app-server `contextCompaction` item +lifecycle and `thread/compacted` notifications therefore drive the existing +CrewCode loading bar through CrewCoder; Claude `compact_boundary` drives the +completed state. The bridge treats a native update as authoritative and does not +also infer compaction from the later context-token drop. If a CrewCoder provider +does not expose a native boundary, a verified high-water-to-large-drop occupancy +change produces one after-the-fact detected notification instead of staying +silent. Automatic updates omit the summary body; the compacted summary remains +only in CrewCoder's durable session. Host-requested `session/compact` returns the authoritative summary and includes it on the completed update. CrewCode replaces only its provider replay shard with that summary and appends the visible compact-summary card; it deliberately @@ -141,6 +148,17 @@ bridge registration; the next composer submission uses normal missing-bridge rec attempting to write to closed stdin and surfacing `crewcoder acp: process not writable`. When automatic compaction is off, the user explicitly runs `/compact` before continuing; this policy does not affect Pi or other providers. +For CrewCoder's built-in Codex provider, the nested app-server receives the +same resolved context policy as the outer loop. Sol, Terra, and Luna GPT-5.6 +sessions therefore use a 1.05M context and a 630k normal auto-compaction limit, +instead of app-server independently compacting near its smaller default. These +overrides are process-scoped and do not modify the user's Codex configuration. +Every CrewCoder built-in model declares a context window. Claude SDK-native +auto-compaction is disabled so CrewCoder remains the only owner of its durable +compaction boundary. ACP providers such as Grok can report the active window at +runtime; that value supersedes static metadata and recalculates CrewCoder's +percentage threshold for subsequent checks. + A prompt has a ten-minute **inactivity** watchdog rather than a wall-clock turn limit. Every matching ACP update or agent request resets it, and time awaiting a Build permission decision is excluded. If CrewCoder becomes genuinely silent, @@ -149,9 +167,20 @@ settle before emitting the timeout and `turn_end`. If cancellation itself remain unresponsive, CrewCode terminates that bridge so its replacement starts cleanly; a second prompt can never overlap the abandoned CrewCoder turn. -Usage prefers `_meta["crewcoder/usage"]`: `lastInputTokens` is the live -`contextTokens` value and `contextWindow` is the context limit. Top-level usage -is only the compatibility fallback. +Usage prefers `_meta["crewcoder/usage"]`: `lastInputTokens` is an authoritative +live `contextTokens` measurement and `contextWindow` is the registered full model +limit (1.05M for the Codex 5.6 family), not Codex app-server's smaller effective +per-request prompt budget. CrewCode accepts measured drops after provider-native +compaction instead of applying its generic monotonic resume floor. The top-level +mirror is the compatibility fallback and preserves the same live fields. Usage is scoped to +one ACP prompt; an errored or metadata-free prompt never inherits the preceding +turn's snapshot. +The chat's Token logs sidebar shows `lastInputTokens` as the latest context +input and labels input, output, total, cache, and reasoning counters from the +same ACP summary as cumulative session counts. These counters can overlap and +must not be added together as context occupancy. If `lastInputTokens` is absent, +the cumulative counters remain available in Token logs without a context +percentage. ACP `tool_call` updates carry a category `kind` (`read`, `edit`, `think`, …), a human `title`, and authoritative CrewCoder tool identity in diff --git a/docs/desktop-web-continuity.md b/docs/desktop-web-continuity.md index 0f74656..5aceea2 100644 --- a/docs/desktop-web-continuity.md +++ b/docs/desktop-web-continuity.md @@ -69,8 +69,8 @@ a rejected Promise where they require a disposer function. ## Source of truth On first enable, missing Brain runtime data is seeded from Electron `userData` into -`~/.crewcode/brain/runtime`. Existing Brain files always win for workspaces, keys, and -replay; they are never overwritten. Transcript shards are merged instead: a newer +`~/.crewcode/brain/runtime`. Existing Brain files win for keys and replay; they are +never overwritten. Transcript shards are merged instead: a newer desktop copy is folded into the Brain shard by message identity so work done locally before attachment is not stuck on the first seed snapshot. The seed includes registered workspaces, provider-native resume IDs, provider keys, rich transcripts, @@ -79,6 +79,15 @@ gets a non-destructive `web:` alias so the first Brain-backed prompt ca continue the desktop conversation. Provider-native resume IDs remain keyed by both session and provider. +Every explicit enable also reconciles the current desktop workspace registry into the +Brain before attachment: current desktop entries and ordering win for matching ids or +paths, while genuine browser-created workspaces are retained. Immediately after the +Brain attaches, but before Electron reloads, the foreground renderer sends its exact +current chat/session and workspace-tab catalogue through the owner-loopback desktop +control. This enable-time handoff supersedes an older persisted authority marker, so a +previous Brain run cannot replace the chats, tabs, workspace, and active selections the +owner was using when they selected **Enable**. + After attachment, the Brain store is authoritative for: - registered, Brain-authorized workspaces; diff --git a/docs/execution-modes.md b/docs/execution-modes.md index 64aa90e..331c53e 100644 --- a/docs/execution-modes.md +++ b/docs/execution-modes.md @@ -3,7 +3,8 @@ CrewCode's composer mode (`ModeLevel`) is a per-session gate with two halves: 1. **An optional prompt preamble** injected once at session start - (`buildModePreamble` in `src/renderer/src/hooks/chat-session-send.ts`). + (`buildModePreamble` in `src/renderer/src/hooks/chat-session-send.ts`), and + once more on each mid-session mode switch (`buildModeSwitchPreamble`). 2. **A real permission policy** applied per provider inside each bridge. The preamble is advisory; the bridge policy is the enforcement. Users edit mode @@ -52,7 +53,14 @@ to true for new and legacy sessions, and is copied when a session is duplicated. before the first send. The toggle locks after visible history or the delivery marker proves startup context was committed; provider context cannot be revoked. - Prompt edits apply only when a session has not sent its startup context. - Existing/restored sessions must not receive the prompt again. + Existing/restored sessions must not receive the startup prompt again. +- Switching mode mid-session sends the new mode's prompt exactly once, on the + next send, behind a short `` notice that the previous mode's + instructions no longer apply. Without it the earlier preamble stays in the + provider transcript as the only mode contract, so an Ask to Build switch + could leave the agent refusing to implement. It never repeats on later turns + in the same mode, and a restored session whose delivery marker is missing is + seeded silently rather than treated as a switch. - Disabled mode prompts leave provider-native/default system context in place. Skills, attachments, handoff packets, and delegation context still use their normal send paths. diff --git a/docs/plugins.md b/docs/plugins.md index 9718997..464fcfb 100644 --- a/docs/plugins.md +++ b/docs/plugins.md @@ -346,6 +346,10 @@ In the dedicated Plugins page you can install a public Git repository, check Git Plugin iframes also have a reload button. The browser API reports uncaught errors and unhandled promise rejections back to CrewCode so runtime failures are visible in the panel host and Settings debug log. +### Theme tokens in plugin panels + +Plugin iframes cannot inherit CSS variables from the CrewCode renderer. The tab host sends a `crewcode:theme` message after frame load and whenever the host's theme classes or inline token overrides change. The payload contains a fixed allowlist of computed semantic tokens (`--background`, `--foreground`, `--card`, `--card-foreground`, `--muted`, `--muted-foreground`, `--primary`, `--primary-foreground`, `--border`, `--input`, `--success`, `--warning`, `--destructive`, `--radius`, `--radius-lg`, and `--font-family-sans`) plus a `dark` or `light` mode. Panels should accept this message only from `window.parent`, apply the provided values to their own document, and provide CSS fallbacks for older CrewCode builds. + ## Security model ```txt diff --git a/docs/system-tray.md b/docs/system-tray.md index bb9e44a..740b655 100644 --- a/docs/system-tray.md +++ b/docs/system-tray.md @@ -20,6 +20,14 @@ Disabling the setting removes the tray icon immediately and restores normal close behavior. On macOS the Dock icon remains visible; CrewCode does not switch to an accessory-only application policy. +Desktop packaging and the Windows/Linux tray use generated icons from +`icon-logo-dark.png` (the white mark) on the dark brand background. Regenerate +`build/icon.png`, `build/icon.ico`, `build/icons/*.png`, and both public PWA icon +sets with `python3 scripts/generate-favicons.py --app-only` (requires Pillow). +These are fixed icons, not automatic light/dark variants. Rebuild/reinstall a +packaged app to update its launcher icon; fully quit and restart development +Electron to reload its window/tray icon. + The preference is renderer-persisted like other CrewCode settings and projected to the Electron main process over the narrow `tray:configure` IPC method. Web, Hub, and headless runtimes do not expose or emulate a system tray. diff --git a/package-lock.json b/package-lock.json index 3b71f82..7d47bd2 100644 --- a/package-lock.json +++ b/package-lock.json @@ -10,7 +10,7 @@ "hasInstallScript": true, "license": "Apache-2.0", "dependencies": { - "@anthropic-ai/claude-agent-sdk": "^0.3.179", + "@anthropic-ai/claude-agent-sdk": "^0.3.282", "@codemirror/autocomplete": "file:packages/crew-codemirror/autocomplete", "@codemirror/commands": "file:packages/crew-codemirror/commands", "@codemirror/lang-css": "file:packages/crew-codemirror/lang-css", @@ -98,22 +98,22 @@ } }, "node_modules/@anthropic-ai/claude-agent-sdk": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk/-/claude-agent-sdk-0.3.179.tgz", - "integrity": "sha512-BlZULVW61ZCcho6mb026QoOQbGZnfxbsEtcoNaW5lF0sIEu7D+/nnQjHSMK8buS/h7HVmIx2RtGbHVX1y4V0uQ==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk/-/claude-agent-sdk-0.3.283.tgz", + "integrity": "sha512-KB+mqU5JLbH2sztlSQeCOu71bK6padYAha3uacBzxFSOVfuRTywYzvsC9P+qV6gXmPXcu98FaPqQv6vBF9j8hA==", "license": "SEE LICENSE IN README.md", "engines": { "node": ">=18.0.0" }, "optionalDependencies": { - "@anthropic-ai/claude-agent-sdk-darwin-arm64": "0.3.179", - "@anthropic-ai/claude-agent-sdk-darwin-x64": "0.3.179", - "@anthropic-ai/claude-agent-sdk-linux-arm64": "0.3.179", - "@anthropic-ai/claude-agent-sdk-linux-arm64-musl": "0.3.179", - "@anthropic-ai/claude-agent-sdk-linux-x64": "0.3.179", - "@anthropic-ai/claude-agent-sdk-linux-x64-musl": "0.3.179", - "@anthropic-ai/claude-agent-sdk-win32-arm64": "0.3.179", - "@anthropic-ai/claude-agent-sdk-win32-x64": "0.3.179" + "@anthropic-ai/claude-agent-sdk-darwin-arm64": "0.3.283", + "@anthropic-ai/claude-agent-sdk-darwin-x64": "0.3.283", + "@anthropic-ai/claude-agent-sdk-linux-arm64": "0.3.283", + "@anthropic-ai/claude-agent-sdk-linux-arm64-musl": "0.3.283", + "@anthropic-ai/claude-agent-sdk-linux-x64": "0.3.283", + "@anthropic-ai/claude-agent-sdk-linux-x64-musl": "0.3.283", + "@anthropic-ai/claude-agent-sdk-win32-arm64": "0.3.283", + "@anthropic-ai/claude-agent-sdk-win32-x64": "0.3.283" }, "peerDependencies": { "@anthropic-ai/sdk": ">=0.93.0", @@ -122,9 +122,9 @@ } }, "node_modules/@anthropic-ai/claude-agent-sdk-darwin-arm64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-darwin-arm64/-/claude-agent-sdk-darwin-arm64-0.3.179.tgz", - "integrity": "sha512-2QoEl7p+RTkF4ARZ40N9OWWi+at8Iy9ZZg3UeP6sWoGrM0GL6NJSv8pbYUdoSRySbNeZshqWar5DF41265J5Sg==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-darwin-arm64/-/claude-agent-sdk-darwin-arm64-0.3.283.tgz", + "integrity": "sha512-UQkROekjufppyB/qrsrU81sM0fcYNWJBEITrGp7NLbOocJZBUR+uVMIx/UHVpe6j81trXPIeRWyfJWgZI17Kxg==", "cpu": [ "arm64" ], @@ -135,9 +135,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-darwin-x64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-darwin-x64/-/claude-agent-sdk-darwin-x64-0.3.179.tgz", - "integrity": "sha512-U5683Hi4uP5Wkg9IhHayqx8co9rgeZaAJcutLAgJiAN9ZnL128+WT0K4DKaBVuMbeeoORodVs9IJgM38lBxUxg==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-darwin-x64/-/claude-agent-sdk-darwin-x64-0.3.283.tgz", + "integrity": "sha512-WkwVppmX0cg1DnXTT9od82pmqM4OK9H3O6MIXoX89SWYty0OpMxsB37263C1tiqk4WkWCs2jwxRXhsoO62ArGQ==", "cpu": [ "x64" ], @@ -148,9 +148,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-linux-arm64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-arm64/-/claude-agent-sdk-linux-arm64-0.3.179.tgz", - "integrity": "sha512-Io9X2xF5J3cHR20b0oflLe0sLHyT/KFMCA7sVJhwuSvcqeCHqXTxn6sG+4lArD9wNf+RtLVKgSvOt5Oz4j7mew==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-arm64/-/claude-agent-sdk-linux-arm64-0.3.283.tgz", + "integrity": "sha512-47IEX/XWw4DUIzPPC85qiASO90rpz7E+2gWr4AKOz5ZAa7ZqEX0PYqPIcpHEM7WpEREfmXCMZgef2bwim24u8A==", "cpu": [ "arm64" ], @@ -161,9 +161,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-linux-arm64-musl": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-arm64-musl/-/claude-agent-sdk-linux-arm64-musl-0.3.179.tgz", - "integrity": "sha512-ehqWR9bV0dJONkoKleKg/LJIGDVevLGLPVIQpol+Bqrgyv05DE1hx68g6mBnfp9IEKo2BMZCba3aHKhkHlZvnQ==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-arm64-musl/-/claude-agent-sdk-linux-arm64-musl-0.3.283.tgz", + "integrity": "sha512-BRlnlh5fsRMoTtjDSGxZL6TathGRgSCMpJk2kdy9Wtz8l6qDjgkLuQilzTF12CwFIn+yf1Fqfrk+Njw+0EsL3A==", "cpu": [ "arm64" ], @@ -174,9 +174,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-linux-x64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-x64/-/claude-agent-sdk-linux-x64-0.3.179.tgz", - "integrity": "sha512-Ck2pLG2GJ5lYULK0Rb6FvAUhkLSvIFgJ77MsbVhvBJHG3C31Fuk6FnkXEL614WSZZdI1ST0Y7s7wc9yLI2kmiQ==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-x64/-/claude-agent-sdk-linux-x64-0.3.283.tgz", + "integrity": "sha512-cE5AebMvTlq7t7Oc0FmP5LzMOEtJ3F8XRGxkc3kTGr4aNcQgXplyYNH2yu1teW1yInuMXHzlCeFRdrmxpe2sRg==", "cpu": [ "x64" ], @@ -187,9 +187,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-linux-x64-musl": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-x64-musl/-/claude-agent-sdk-linux-x64-musl-0.3.179.tgz", - "integrity": "sha512-Q+clijh01m+Gg/tSNjHQWfPFRvXSVMIQJZqu6v28vFduaucL7ZL+p6o6VOSdlgLjhqtIFFjOO37aL+VkvEu9MQ==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-linux-x64-musl/-/claude-agent-sdk-linux-x64-musl-0.3.283.tgz", + "integrity": "sha512-T0T8mR7MSe7bVI96tmeK7DgUy2xhG8CZD1i0nCC2C05sSmkDDn/OYgNqpGPxJbo8Ic0Psxd7TOmngQYqPKWtpA==", "cpu": [ "x64" ], @@ -200,9 +200,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-win32-arm64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-win32-arm64/-/claude-agent-sdk-win32-arm64-0.3.179.tgz", - "integrity": "sha512-7wc4j3vDE/TiuiCEPV024U/TDNKXUJ89V0N9WteYJ57YrIUm0RA51WiKCrEmSxSqstQ7Slq26e0gWGHRf7ljDQ==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-win32-arm64/-/claude-agent-sdk-win32-arm64-0.3.283.tgz", + "integrity": "sha512-hZjyeIgZpALMvYq2QfyPzl4AZzOVRUoU8NOB82Z3cVXLcEvQJXJN03Hp3XNsuBq5YMUoSUJAffMannhM7UqVmg==", "cpu": [ "arm64" ], @@ -213,9 +213,9 @@ ] }, "node_modules/@anthropic-ai/claude-agent-sdk-win32-x64": { - "version": "0.3.179", - "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-win32-x64/-/claude-agent-sdk-win32-x64-0.3.179.tgz", - "integrity": "sha512-gFad1AVUZOXbEyS9lf3Pr0LDXrLuO6limZoZoUZQmjhjy0LnudnTU8P/VbCY0AHmT46YMZ/ieisFzeq8X5jTcg==", + "version": "0.3.283", + "resolved": "https://registry.npmjs.org/@anthropic-ai/claude-agent-sdk-win32-x64/-/claude-agent-sdk-win32-x64-0.3.283.tgz", + "integrity": "sha512-h5eZxFxk6f1LPSuUOoBH9fslhL9ncywdZN5Qw9XmWuijHYTfevXuoza5nm6N98O+hmSykJ8NjC3/muhONDvzBg==", "cpu": [ "x64" ], diff --git a/package.json b/package.json index 1bf6f78..99f24bf 100644 --- a/package.json +++ b/package.json @@ -146,7 +146,7 @@ "serve": "npm run build && node bin/crewcode-server.mjs serve", "hub": "npm run build && node bin/crewcode-server.mjs hub", "hub:mobile": "npm run build && node bin/crewcode-server.mjs hub mobile --tailscale", - "enroll": "npm run build && node bin/crewcode-server.mjs enroll", + "enroll": "node bin/crewcode-server.mjs enroll", "brain": "npm run build && node bin/crewcode-server.mjs brain", "prepack": "npm run build", "typecheck": "tsc -p tsconfig.node.json --noEmit && tsc -p tsconfig.web.json --noEmit && tsc -p packages/crewcode-plugin-api/tsconfig.json --noEmit", @@ -185,7 +185,7 @@ "vitest": "^4.1.11" }, "dependencies": { - "@anthropic-ai/claude-agent-sdk": "^0.3.179", + "@anthropic-ai/claude-agent-sdk": "^0.3.282", "@codemirror/autocomplete": "file:packages/crew-codemirror/autocomplete", "@codemirror/commands": "file:packages/crew-codemirror/commands", "@codemirror/lang-css": "file:packages/crew-codemirror/lang-css", diff --git a/public/icons/icon-192.png b/public/icons/icon-192.png index 93b8f09..2d0c4b0 100644 Binary files a/public/icons/icon-192.png and b/public/icons/icon-192.png differ diff --git a/public/icons/icon-512.png b/public/icons/icon-512.png index 807d687..02e8d9f 100644 Binary files a/public/icons/icon-512.png and b/public/icons/icon-512.png differ diff --git a/scripts/generate-favicons.py b/scripts/generate-favicons.py index bd2c08e..671c9a5 100644 --- a/scripts/generate-favicons.py +++ b/scripts/generate-favicons.py @@ -21,6 +21,7 @@ Usage: python3 scripts/generate-favicons.py """ +import argparse import base64 from io import BytesIO from pathlib import Path @@ -46,15 +47,25 @@ def load_trimmed(name: str) -> Image.Image: return im.crop(bbox) -def render(mark: Image.Image, size: int, margin: float, bg=None) -> Image.Image: +def render(mark: Image.Image, size: int, margin: float, bg=None, corner_radius: float = 0) -> Image.Image: """Fit the mark inside a square canvas, preserving aspect ratio.""" inner = max(1, int(size * (1 - 2 * margin))) w, h = mark.size scale = inner / max(w, h) - resized = mark.resize((max(1, round(w * scale)), max(1, round(h * scale))), Image.LANCZOS) + resized = mark.resize((max(1, round(w * scale)), max(1, round(h * scale))), Image.Resampling.LANCZOS) canvas = Image.new("RGBA", (size, size), bg if bg else (0, 0, 0, 0)) canvas.alpha_composite(resized, ((size - resized.width) // 2, (size - resized.height) // 2)) + if corner_radius > 0: + # Launcher and tray surfaces need transparent corners so the rounded + # brand tile does not appear as a square against the desktop chrome. + from PIL import ImageDraw + + mask = Image.new("L", (size, size), 0) + ImageDraw.Draw(mask).rounded_rectangle( + (0, 0, size - 1, size - 1), radius=round(size * corner_radius), fill=255 + ) + canvas.putalpha(mask) return canvas @@ -79,9 +90,27 @@ def dual_mode_svg(light: Image.Image, dark: Image.Image, size: int = 64) -> str: def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--app-only", action="store_true") + args = parser.parse_args() light = load_trimmed("icon-logo-light.png") # black mark dark = load_trimmed("icon-logo-dark.png") # white mark + if args.app_only: + build = REPO / "build" + render(dark, 1024, 0.12, BRAND_BG, 0.22).save(build / "icon.png") + render(dark, 256, 0.12, BRAND_BG, 0.22).save( + build / "icon.ico", + sizes=[(16, 16), (24, 24), (32, 32), (48, 48), (64, 64), (128, 128), (256, 256)], + ) + for size in (256, 512): + render(dark, size, 0.12, BRAND_BG, 0.22).save(build / "icons" / f"{size}x{size}.png") + for out in (REPO / "public/icons", REPO / "src/renderer/public/icons"): + for size in (192, 512): + render(dark, size, 0.18, BRAND_BG, 0.22).save(out / f"icon-{size}.png") + print("wrote desktop, tray, and PWA icons") + return + for out in TARGETS: out.mkdir(parents=True, exist_ok=True) diff --git a/src/main/agents/bridge-service.test.ts b/src/main/agents/bridge-service.test.ts index 8413b5b..734f63a 100644 --- a/src/main/agents/bridge-service.test.ts +++ b/src/main/agents/bridge-service.test.ts @@ -115,6 +115,37 @@ describe('AgentBridgeService', () => { } }) + it('withdraws a pending question card when the provider resolves it itself', async () => { + const controller = new AbortController() + let answer: Promise | undefined + const factory = vi.fn(async (_path, opts, _emit, requestUser) => { + answer = requestUser({ kind: 'prompt', title: 'Branch name?' }, controller.signal) + return { + bridgeId: opts.bridgeId, + pid: null, + prompt: async () => ({ ok: true }), + abort: async () => undefined, + stop: async () => undefined, + } + }) + const service = new AgentBridgeService(() => '/bin/fake-agent', factory) + const events: BridgeEvent[] = [] + service.subscribe(event => events.push(event)) + await service.start({ bridgeId: 'question-bridge', provider: 'codex', cwd: '/tmp' }) + + const asked = events.find(event => event.type === 'user_request') + expect(asked).toBeDefined() + controller.abort() + + // Settles as cancel (never an answer) and retracts the card. + await expect(answer).resolves.toMatchObject({ action: 'cancel' }) + expect(events).toContainEqual({ + type: 'user_request_resolved', + bridgeId: 'question-bridge', + requestId: asked?.type === 'user_request' ? asked.request.requestId : '', + }) + }) + it('rejects unknown and plugin providers before resolving a path', async () => { const resolvePath = vi.fn(() => '/bin/agent') const service = new AgentBridgeService(resolvePath) diff --git a/src/main/agents/bridge-service.ts b/src/main/agents/bridge-service.ts index d4578a3..a68e864 100644 --- a/src/main/agents/bridge-service.ts +++ b/src/main/agents/bridge-service.ts @@ -227,7 +227,7 @@ export class AgentBridgeService { } for (const listener of this.listeners) listener(event) } - const requestUser: RequestUserFn = request => this.requestUser(opts.bridgeId, request) + const requestUser: RequestUserFn = (request, signal) => this.requestUser(opts.bridgeId, request, signal) try { const bridge = await this.create(path, opts, emit, requestUser) const hasLocalHistory = !!opts.conversationKey && loadConversation(opts.conversationKey).length > 0 @@ -535,7 +535,8 @@ export class AgentBridgeService { } } - private requestUser(bridgeId: string, request: Omit): Promise { + private requestUser(bridgeId: string, request: Omit, signal?: AbortSignal): Promise { + if (signal?.aborted) return Promise.resolve({ requestId: `${bridgeId}:withdrawn`, action: 'cancel' }) const entry = this.bridges.get(bridgeId) const prepared = this.turnPermissionGrants.prepareRequest( bridgeId, @@ -550,6 +551,14 @@ export class AgentBridgeService { return new Promise(resolve => { this.pendingRequests.set(requestId, { bridgeId, request: payload, resolve }) for (const listener of this.listeners) listener({ type: 'user_request', request: payload }) + // Provider-side withdrawal: settle as cancel and retract the card. + signal?.addEventListener('abort', () => { + const pending = this.pendingRequests.get(requestId) + if (!pending) return + this.pendingRequests.delete(requestId) + pending.resolve({ requestId, action: 'cancel' }) + for (const listener of this.listeners) listener({ type: 'user_request_resolved', bridgeId, requestId }) + }, { once: true }) }) } diff --git a/src/main/agents/bridge-types.ts b/src/main/agents/bridge-types.ts index c1bf890..ae8f173 100644 --- a/src/main/agents/bridge-types.ts +++ b/src/main/agents/bridge-types.ts @@ -26,9 +26,8 @@ export interface CompactionStatus { reason?: string } -// One slice of the context window (system prompt, tools, MCP, memory files, -// messages, …). Only providers that report a breakdown populate it (claude via -// getContextUsage); the UI shows it so "what's using my context" is inspectable. +// One observed token-log row. Claude reports context categories; other +// providers can report request/turn counters, which may overlap. export interface ContextCategory { name: string tokens: number @@ -41,9 +40,14 @@ export interface TurnUsage { totalTokens?: number contextTokens?: number contextWindow?: number + /** Provider-reported effective prompt budget when smaller than model capacity. */ + promptBudgetTokens?: number + /** The provider measured current live occupancy, so a lower value is meaningful. */ + contextIsAuthoritative?: boolean model?: string compaction?: CompactionStatus contextBreakdown?: ContextCategory[] + contextBreakdownSource?: 'context' | 'usage' } export interface AgentUserRequest { @@ -61,6 +65,10 @@ export interface AgentUserRequest { source?: string /** Main-issued capability: this Build permission can grant the remaining turn. */ allowAllForTurn?: boolean + /** Question options are toggles; the answer is `optionIds` (plus optional typed text). */ + multiple?: boolean + /** Provider marked the answer sensitive; the card masks the typed input. */ + secret?: boolean } export interface AgentUserResponse { @@ -68,9 +76,19 @@ export interface AgentUserResponse { action: 'accept' | 'accept_for_turn' | 'decline' | 'submit' | 'cancel' value?: string optionId?: string + /** Toggled options for `multiple` question requests, in option order. */ + optionIds?: string[] } -export type RequestUserFn = (request: Omit) => Promise +/** + * Ask the human. `signal` lets a bridge withdraw a pending card when the + * provider resolved the request itself (e.g. a non-blocking Codex question); + * the promise then settles as `cancel`, never as an answer. + */ +export type RequestUserFn = ( + request: Omit, + signal?: AbortSignal, +) => Promise export type BridgeEvent = | { type: 'ready'; bridgeId: string } diff --git a/src/main/agents/claude-bridge.test.ts b/src/main/agents/claude-bridge.test.ts index 8ef63bc..e4b23a8 100644 --- a/src/main/agents/claude-bridge.test.ts +++ b/src/main/agents/claude-bridge.test.ts @@ -257,7 +257,7 @@ describe('claude bridge mode options', () => { } expect(claudeAskUserQuestionRequest(input)).toMatchObject({ - kind: 'select', + kind: 'prompt', title: 'Which approach?', options: [{ id: 'Safe', label: 'Safe' }, { id: 'Fast', label: 'Fast' }], }) @@ -269,6 +269,39 @@ describe('claude bridge mode options', () => { }) }) + it('offers option buttons plus a typed "Other" answer by default, matching Claude', () => { + const input = { questions: [{ question: 'Proceed?', options: [{ label: 'Yes' }, { label: 'No' }] }] } + const request = claudeAskUserQuestionRequest(input) + expect(request).toMatchObject({ kind: 'prompt', multiple: false, placeholder: 'or type your own answer…' }) + expect(request?.options).toHaveLength(2) + expect(answerClaudeAskUserQuestionInput(input, { + requestId: 'r1', action: 'submit', value: 'Only after tests pass', + })).toMatchObject({ answers: { 'Proceed?': 'Only after tests pass' } }) + }) + + it('keeps option-only questions when the input explicitly disables free text', () => { + expect(claudeAskUserQuestionRequest({ + questions: [{ question: 'Proceed?', custom: false, options: [{ label: 'Yes' }, { label: 'No' }] }], + })).toMatchObject({ kind: 'select', placeholder: undefined }) + }) + + it('renders multi-select options as toggles and answers with the chosen labels', () => { + const input = { + questions: [{ + question: 'Which checks?', + multiSelect: true, + options: [{ label: 'Lint' }, { label: 'Tests' }, { label: 'Types' }], + }], + } + expect(claudeAskUserQuestionRequest(input)).toMatchObject({ + multiple: true, + options: [{ id: 'Lint' }, { id: 'Tests' }, { id: 'Types' }], + }) + expect(answerClaudeAskUserQuestionInput(input, { + requestId: 'r1', action: 'submit', optionIds: ['Lint', 'Types'], value: 'e2e', + })).toMatchObject({ answers: { 'Which checks?': 'Lint, Types, e2e' } }) + }) + it('answers every Claude question through canUseTool without a permission fallthrough', async () => { const requestUser = vi.fn() .mockResolvedValueOnce({ requestId: 'r1', action: 'submit', optionId: 'Safe' }) @@ -291,7 +324,7 @@ describe('claude bridge mode options', () => { ) expect(requestUser).toHaveBeenCalledTimes(2) - expect(requestUser.mock.calls.map(([request]) => request.kind)).toEqual(['select', 'select']) + expect(requestUser.mock.calls.map(([request]) => request.kind)).toEqual(['prompt', 'prompt']) expect(result).toEqual({ behavior: 'allow', updatedInput: { @@ -626,11 +659,11 @@ describe('claude bridge mode options', () => { expect(result).toEqual({ behavior: 'allow', updatedInput: { command: 'git push --force' } }) }) - it('uses Claude SDK context usage instead of aggregate billing tokens', async () => { + it('uses active Claude SDK categories instead of aggregate billing tokens', async () => { queryMock.mockImplementation(() => ({ close: vi.fn(), getContextUsage: vi.fn().mockResolvedValue({ - categories: [], + categories: [{ name: 'Messages', tokens: 32_175, color: '#fff' }], totalTokens: 322_175, maxTokens: 1_000_000, rawMaxTokens: 1_000_000, @@ -665,7 +698,7 @@ describe('claude bridge mode options', () => { usage: expect.objectContaining({ inputTokens: 12, outputTokens: 3, - contextTokens: 322_175, + contextTokens: 32_175, contextWindow: 1_000_000, model: 'claude-opus-4-8', }), @@ -681,7 +714,7 @@ describe('claude bridge mode options', () => { // racing the per-turn query shutdown). Falling back to billing math makes // the ctx meter bounce; the bridge must reuse the previous SDK reading. const goodContext = { - categories: [], + categories: [{ name: 'Messages', tokens: 100_000, color: '#fff' }], totalTokens: 100_000, maxTokens: 1_000_000, rawMaxTokens: 1_000_000, @@ -753,11 +786,10 @@ describe('claude bridge mode options', () => { { name: 'System prompt', tokens: 3_000, color: '#fff' }, { name: 'System tools', tokens: 14_000, color: '#fff' }, { name: 'Messages', tokens: 8_000, color: '#fff' }, - { name: 'Autocompact buffer', tokens: 45_000, color: '#ccc' }, + { name: 'Auto-compaction buffer', tokens: 45_000, color: '#ccc' }, { name: 'Free space', tokens: 930_000, color: '#ccc' }, ], - // Claude derives totalTokens from cumulative API usage, so it can - // exceed the window on a resumed thread. + // Deliberately inconsistent totals exercise category classification. totalTokens: 1_240_000, maxTokens: 1_000_000, rawMaxTokens: 1_000_000, @@ -792,7 +824,7 @@ describe('claude bridge mode options', () => { queryMock.mockImplementation(() => ({ close: vi.fn(), getContextUsage: vi.fn().mockResolvedValue({ - categories: [], + categories: [{ name: 'Messages', tokens: 4_015_012, color: '#fff' }], totalTokens: 4_015_012, maxTokens: 1_000_000, rawMaxTokens: 1_000_000, @@ -823,7 +855,7 @@ describe('claude bridge mode options', () => { queryMock.mockImplementation(() => ({ close: vi.fn(), getContextUsage: vi.fn().mockResolvedValue({ - categories: [], + categories: [{ name: 'Messages', tokens: 322_175, color: '#fff' }], totalTokens: 322_175, maxTokens: 1_000_000, rawMaxTokens: 1_000_000, @@ -966,7 +998,7 @@ describe('claude bridge mode options', () => { })) }) - it('falls back to Claude request context when live context usage is unavailable', async () => { + it('omits a context estimate when live context usage is unavailable', async () => { queryMock.mockImplementation(() => ({ close: vi.fn(), getContextUsage: vi.fn().mockRejectedValue(new Error('control channel closed')), @@ -995,11 +1027,142 @@ describe('claude bridge mode options', () => { inputTokens: 12, outputTokens: 3, totalTokens: 15, - contextTokens: 15_015, contextWindow: 500_000, model: 'claude-opus-4-8', }), })) + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(turnEnd.usage.contextTokens).toBeUndefined() + }) + + it('does not turn SDK totals without active categories into context usage', async () => { + queryMock.mockImplementation(() => ({ + close: vi.fn(), + getContextUsage: vi.fn().mockResolvedValue({ + categories: [], + totalTokens: 900_000, + maxTokens: 1_000_000, + rawMaxTokens: 1_000_000, + percentage: 90, + gridRows: [], + model: 'claude-opus-4-8', + memoryFiles: [], + mcpTools: [], + agents: [], + }), + async *[Symbol.asyncIterator]() { + yield { type: 'result', subtype: 'success', usage: { input_tokens: 12, output_tokens: 3 } } + }, + })) + const emit = vi.fn() + const bridge = await createClaudeBridge('/bin/claude', { + bridgeId: 'b1', provider: 'claude', cwd: '/repo', mode: 'build', model: 'claude-opus-4-8', + }, emit) + + await bridge.prompt('run it') + + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(turnEnd.usage.contextTokens).toBeUndefined() + }) + + it('keeps a live context reading captured before the result closes the control channel', async () => { + const getContextUsage = vi.fn() + .mockResolvedValueOnce({ + categories: [{ name: 'Messages', tokens: 42_000, color: '#fff' }], + totalTokens: 42_000, maxTokens: 200_000, rawMaxTokens: 200_000, + percentage: 21, gridRows: [], model: 'claude-haiku-4-5', + memoryFiles: [], mcpTools: [], agents: [], + }) + .mockRejectedValue(new Error('control channel closed')) + queryMock.mockImplementation(() => ({ + close: vi.fn(), + getContextUsage, + async *[Symbol.asyncIterator]() { + yield { type: 'assistant', message: { + model: 'claude-haiku-4-5', + content: [{ type: 'text', text: 'done' }], + usage: { input_tokens: 10, cache_read_input_tokens: 20_000, output_tokens: 5 }, + } } + yield { type: 'result', subtype: 'success', usage: { + input_tokens: 100, output_tokens: 20, cache_read_input_tokens: 800_000, + } } + }, + })) + const emit = vi.fn() + const bridge = await createClaudeBridge('/bin/claude', { + bridgeId: 'b1', provider: 'claude', cwd: '/repo', mode: 'build', model: 'claude-haiku-4-5', + }, emit) + + await bridge.prompt('run it') + + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(turnEnd.usage).toMatchObject({ contextTokens: 42_000, contextWindow: 200_000 }) + expect(getContextUsage).toHaveBeenCalledTimes(1) + expect(getContextUsage).toHaveBeenCalledWith({ detail: 'full' }) + }) + + it('uses SDK category kinds instead of English labels for occupancy', async () => { + queryMock.mockImplementation(() => ({ + close: vi.fn(), + getContextUsage: vi.fn().mockResolvedValue({ + categories: [ + { name: 'Free space', kind: 'used', tokens: 18_000, color: '#fff' }, + { name: 'Messages', kind: 'free', tokens: 170_000, color: '#fff' }, + { name: 'Tools', kind: 'buffer', tokens: 10_000, color: '#fff' }, + { name: 'System', kind: 'deferred', tokens: 2_000, color: '#fff' }, + ], + totalTokens: 200_000, maxTokens: 200_000, rawMaxTokens: 200_000, + percentage: 100, gridRows: [], model: 'claude-haiku-4-5', + memoryFiles: [], mcpTools: [], agents: [], + }), + async *[Symbol.asyncIterator]() { + yield { type: 'result', subtype: 'success', usage: { input_tokens: 200, output_tokens: 10 } } + }, + })) + const emit = vi.fn() + const bridge = await createClaudeBridge('/bin/claude', { + bridgeId: 'b1', provider: 'claude', cwd: '/repo', mode: 'build', model: 'claude-haiku-4-5', + }, emit) + + await bridge.prompt('run it') + + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(turnEnd.usage).toMatchObject({ contextTokens: 18_000, contextWindow: 200_000 }) + }) + + it('uses the latest individual assistant request when every context control call fails', async () => { + queryMock.mockImplementation(() => ({ + close: vi.fn(), + getContextUsage: vi.fn().mockRejectedValue(new Error('control channel unavailable')), + async *[Symbol.asyncIterator]() { + yield { type: 'assistant', message: { + model: 'claude-haiku-4-5', content: [], + usage: { input_tokens: 100, cache_read_input_tokens: 20_000, output_tokens: 30 }, + } } + yield { type: 'assistant', message: { + model: 'claude-haiku-4-5', content: [], + usage: { input_tokens: 200, cache_read_input_tokens: 30_000, output_tokens: 40 }, + } } + yield { type: 'result', subtype: 'success', usage: { + input_tokens: 300, output_tokens: 70, cache_read_input_tokens: 950_000, + } } + }, + })) + const emit = vi.fn() + const bridge = await createClaudeBridge('/bin/claude', { + bridgeId: 'b1', provider: 'claude', cwd: '/repo', mode: 'build', model: 'claude-haiku-4-5', + }, emit) + + await bridge.prompt('run it') + + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(turnEnd.usage).toMatchObject({ + contextTokens: 30_240, + contextWindow: 200_000, + contextIsAuthoritative: true, + inputTokens: 300, + outputTokens: 70, + }) }) it('emits compaction events from Claude compact boundaries', async () => { @@ -1028,6 +1191,39 @@ describe('claude bridge mode options', () => { provider: 'claude', beforeTokens: 151_000, afterTokens: 24_000, + resetContext: true, + })) + }) + + it('does not reuse a pre-compaction context reading when the post-boundary control call fails', async () => { + const getContextUsage = vi.fn() + .mockResolvedValueOnce({ + categories: [{ name: 'Messages', kind: 'used', tokens: 150_000 }], + maxTokens: 200_000, + }) + .mockRejectedValueOnce(new Error('control channel closed')) + queryMock.mockImplementation(() => ({ + close: vi.fn(), + getContextUsage, + async *[Symbol.asyncIterator]() { + yield { type: 'assistant', message: { + model: 'claude-haiku-4-5', content: [], + usage: { input_tokens: 150_000, output_tokens: 100 }, + } } + yield { type: 'system', subtype: 'compact_boundary', compact_metadata: { trigger: 'auto' } } + yield { type: 'result', subtype: 'success', usage: { input_tokens: 150_000, output_tokens: 100 } } + }, })) + const emit = vi.fn() + const bridge = await createClaudeBridge('/bin/claude', { + bridgeId: 'b1', provider: 'claude', cwd: '/repo', mode: 'build', model: 'claude-haiku-4-5', + }, emit) + + await bridge.prompt('run it') + + const turnEnd = emit.mock.calls.map(([event]) => event).find(event => event.type === 'turn_end') + expect(getContextUsage).toHaveBeenCalledTimes(2) + expect(turnEnd.usage?.contextTokens).toBeUndefined() + expect(turnEnd.usage?.contextBreakdown).toBeUndefined() }) }) diff --git a/src/main/agents/claude-bridge.ts b/src/main/agents/claude-bridge.ts index eb7bcbc..7b03c79 100644 --- a/src/main/agents/claude-bridge.ts +++ b/src/main/agents/claude-bridge.ts @@ -1,6 +1,6 @@ import { query, resolveSettings, type CanUseTool, type PermissionMode, type PermissionResult, type Query, type SDKControlGetContextUsageResponse, type SDKMessage, type Settings } from '@anthropic-ai/claude-agent-sdk' import type { AgentBridge, AgentUserRequest, AgentUserResponse, BridgeStartOpts, ContextCategory, EmitFn, ModeLevel, PromptOptions, RequestUserFn, TurnUsage } from './bridge-types' -import { buildUsage, contextWindowFor } from './model-context' +import { contextWindowFor } from './model-context' import { tripwireForToolCall, extractShellCommand } from './dangerous-command' // Claude Code is driven through the official Agent SDK's `query()` rather than a @@ -105,6 +105,7 @@ interface ClaudeAskUserQuestionRequest { detail?: string options?: ClaudeQuestionOption[] placeholder?: string + multiple: boolean optionsById: Record } @@ -169,28 +170,38 @@ export function claudeAskUserQuestionRequest(input: Record, ind const optionsById = Object.fromEntries(options.map(option => [option.id, option])) const header = stringValue(question.header) const progress = questions.length > 1 ? `Question ${index + 1} of ${questions.length}.` : undefined - const multiple = question.multiple === true || question.multiSelect === true - const allowsCustom = question.custom === true || question.allowFreeform === true - const detail = multiple && options.length > 0 - ? options.map(option => `- ${option.label}${option.description ? ` — ${option.description}` : ''}`).join('\n') - : undefined + const multiple = (question.multiple === true || question.multiSelect === true) && options.length > 0 + // Claude's AskUserQuestion contract always offers "Other" free text, so a + // typed answer is valid unless the input explicitly opts out. + const allowsCustom = question.custom !== false && question.allowFreeform !== false return { questionText, - kind: multiple ? 'editor' : allowsCustom || options.length === 0 ? 'prompt' : 'select', + kind: allowsCustom || options.length === 0 ? 'prompt' : 'select', title: questionText, message: [header ? `Claude asks: ${header}` : undefined, progress].filter(Boolean).join(' ') || undefined, - detail, - options: multiple ? undefined : options, - placeholder: allowsCustom ? 'reply to Claude…' : undefined, + options: options.length > 0 ? options : undefined, + placeholder: allowsCustom ? (options.length > 0 ? 'or type your own answer…' : 'reply to Claude…') : undefined, + multiple, optionsById, } } +function claudeAnswerText(request: ClaudeAskUserQuestionRequest, response: AgentUserResponse): { answer: string; selected?: ClaudeQuestionOption } { + const typed = response.value?.trim() ?? '' + if (request.multiple) { + const labels = (response.optionIds ?? []) + .map(id => request.optionsById[id]?.label) + .filter((label): label is string => !!label) + return { answer: [...labels, ...(typed ? [typed] : [])].join(', ') } + } + const selected = response.optionId ? request.optionsById[response.optionId] : undefined + return { answer: selected?.label ?? (typed || response.optionId || ''), selected } +} + export function answerClaudeAskUserQuestionInput(input: Record, response: AgentUserResponse, index = 0): Record { const request = claudeAskUserQuestionRequest(input, index) if (!request) return input - const selected = response.optionId ? request.optionsById[response.optionId] : undefined - const answer = selected?.label ?? response.value ?? response.optionId ?? '' + const { answer, selected } = claudeAnswerText(request, response) const answers = objectValue(input.answers) ?? {} const updated: Record = { ...input, @@ -206,64 +217,84 @@ export function answerClaudeAskUserQuestionInput(input: Record, return updated } -// Result usage is aggregate turn/API billing data. It is only a fallback for the -// ctx pill; Claude SDK getContextUsage() is the authoritative context gauge. +// Result usage is aggregate turn/API billing data. Cache reads and repeated API +// calls make it unsuitable as a live context gauge. function usageFromResult(usage: unknown, model: string | undefined): TurnUsage | undefined { const u = usage as Record | undefined if (!u || typeof u !== 'object') return undefined const input = numberValue(u.input_tokens) const output = numberValue(u.output_tokens) - const cacheCreation = numberValue(u.cache_creation_input_tokens) ?? 0 - const cacheRead = numberValue(u.cache_read_input_tokens) ?? 0 - const contextWindow = contextWindowFor(model) - const withCache = (input ?? 0) + cacheCreation + cacheRead + (output ?? 0) - const withoutCacheRead = (input ?? 0) + cacheCreation + (output ?? 0) - // Fallback only: keep ctx-pop useful on SDK/control failures without letting - // repeated cache-read billing become 4M context. - const contextTokens = contextWindow && withCache > contextWindow * 1.1 ? withoutCacheRead : withCache - return buildUsage({ - inputTokens: input, - outputTokens: output, - contextTokens, - contextWindow, + if (input === undefined && output === undefined) return undefined + return { + inputTokens: input, + outputTokens: output, + totalTokens: (input ?? 0) + (output ?? 0), + contextWindow: contextWindowFor(model), model, - }) + } +} + +// Each assembled assistant message corresponds to one Claude API request. Its +// input/cache counts describe that request's prompt, unlike the result message +// which aggregates all requests in the turn. +function contextFromAssistantRequest(message: Record, model: string | undefined): TurnUsage | undefined { + const usage = objectValue(message.usage) + if (!usage) return undefined + const fields = ['input_tokens', 'cache_creation_input_tokens', 'cache_read_input_tokens'] as const + const counts = fields.map(field => numberValue(usage[field])) + if (counts.every(count => count === undefined)) return undefined + const input = counts.reduce((sum, count) => sum + Math.max(0, count ?? 0), 0) + const output = Math.max(0, numberValue(usage.output_tokens) ?? 0) + return { + contextTokens: input + output, + contextWindow: contextWindowFor(model), + contextIsAuthoritative: true, + model, + contextBreakdown: [{ name: 'Latest Claude request (context control unavailable)', tokens: input + output }], + contextBreakdownSource: 'usage', + } } // Claude's category list is built to fill the /context grid, so it ends with // synthetic *capacity* rows — the unused remainder ("Free space") and the // reserved compaction headroom — sized so every non-deferred category sums to // exactly maxTokens. Counting them as usage pins the meter at 100% forever. -const CLAUDE_CAPACITY_CATEGORIES = new Set(['free space', 'autocompact buffer', 'compact buffer']) +const CLAUDE_CAPACITY_CATEGORIES = new Set([ + 'freespace', 'autocompactbuffer', 'autocompactionbuffer', 'compactbuffer', 'compactionbuffer', +]) function isClaudeCapacityCategory(name: unknown): boolean { - const label = stringValue(name)?.trim().toLowerCase() + const label = stringValue(name)?.toLowerCase().replace(/[^a-z]/g, '') return label !== undefined && CLAUDE_CAPACITY_CATEGORIES.has(label) } +function claudeCategoryKind(category: SDKControlGetContextUsageResponse['categories'][number]): 'used' | 'free' | 'buffer' | 'deferred' { + // New SDK/CLI responses classify rows directly. Keep the older response + // shape usable while a user's installed Claude binary catches up. + const kind = (category as { kind?: unknown }).kind + if (kind === 'used' || kind === 'free' || kind === 'buffer' || kind === 'deferred') return kind + if (kind !== undefined) return 'deferred' + if (category.isDeferred) return 'deferred' + return isClaudeCapacityCategory(category.name) ? 'buffer' : 'used' +} + function claudeCategoryTokenSums(context: SDKControlGetContextUsageResponse): { active: number; deferred: number } { return context.categories.reduce((sum, category) => { - if (isClaudeCapacityCategory(category.name)) return sum + const kind = claudeCategoryKind(category) const tokens = numberValue(category.tokens) ?? 0 - if (category.isDeferred) sum.deferred += tokens - else sum.active += tokens + if (kind === 'deferred') sum.deferred += tokens + else if (kind === 'used') sum.active += tokens return sum }, { active: 0, deferred: 0 }) } function activeClaudeContextTokens(context: SDKControlGetContextUsageResponse): number | undefined { - const total = numberValue(context.totalTokens) + if (!Array.isArray(context.categories) || context.categories.length === 0) return undefined const max = numberValue(context.maxTokens) ?? numberValue(context.rawMaxTokens) const { active } = claudeCategoryTokenSums(context) - // Claude can report large deferred/reserved categories in totalTokens. Those - // are useful for Claude's own compaction logic, but they make CrewCode's live - // chat meter look full after a tiny prompt. - const used = active > 0 && (total === undefined || active < total) ? active : total - if (used === undefined) return undefined - // totalTokens can exceed the window (Claude derives it from cumulative API - // usage, which double-counts cache reads across a resumed thread). Live - // context physically cannot, so cap it instead of rendering >=100%. - return max !== undefined && max > 0 ? Math.min(used, max) : used + // The SDK marks exactly which categories occupy the /context window. Use + // those rows directly so free space, buffer, and deferred rows stay out. + return max !== undefined && max > 0 ? Math.min(active, max) : active } function pushContextRow(rows: ContextCategory[], name: string, tokens: unknown, deferred?: boolean): void { @@ -290,9 +321,10 @@ function claudeContextBreakdown(context: SDKControlGetContextUsageResponse): Con for (const category of context.categories) { // Capacity rows are headroom, not consumption — show them flagged so the // list still reconciles against maxTokens without reading as usage. - const capacity = isClaudeCapacityCategory(category.name) + const kind = claudeCategoryKind(category) + const capacity = kind === 'free' || kind === 'buffer' const name = stringValue(category.name) ?? 'unknown' - pushContextRow(rows, capacity ? `${name} (unused headroom)` : name, category.tokens, capacity || category.isDeferred === true) + pushContextRow(rows, capacity ? `${name} (unused headroom)` : name, category.tokens, capacity || kind === 'deferred') } // The top-level categories can say "tools" or "system" without explaining why @@ -380,14 +412,16 @@ function applyClaudeContextUsage(usage: TurnUsage | undefined, context: SDKContr ...(liveContextTokens !== undefined ? { contextTokens: liveContextTokens } : {}), ...(contextWindow !== undefined ? { contextWindow } : {}), ...(model ? { model } : {}), - ...(breakdown.length > 0 ? { contextBreakdown: breakdown } : {}), + ...(breakdown.length > 0 ? { contextBreakdown: breakdown, contextBreakdownSource: 'context' as const } : {}), } } async function readClaudeContextUsage(q: Query): Promise { if (typeof q.getContextUsage !== 'function') return undefined try { - return await q.getContextUsage() + // Full uses the same counted category breakdown as /context. Summary is + // faster but explicitly approximate, so it cannot drive this meter. + return await q.getContextUsage({ detail: 'full' }) } catch { return undefined } @@ -434,20 +468,28 @@ export async function createClaudeBridge( } followUpQueue.length = 0 } - // The ctx gauge must not bounce: getContextUsage() can fail right after a - // turn's result (the control channel races the per-turn query shutdown), and - // the billing-based fallback swings wildly with per-turn API-call counts. - // Cache the last good SDK reading and reuse it so the meter only moves on - // real SDK data; billing math is the fallback only before any reading exists. + // getContextUsage() can fail after a turn's result when the control channel + // races query shutdown. Keep a valid reading for the next turn only as a + // last resort after trying this turn's control and per-request measurements. let lastContextUsage: SDKControlGetContextUsageResponse | undefined - async function applyStableContextUsage(usage: TurnUsage | undefined, q: Query): Promise { + async function readMeasuredContext(q: Query): Promise { const context = await readClaudeContextUsage(q) - if (context) { + if (context && activeClaudeContextTokens(context) !== undefined) { lastContextUsage = context - return applyClaudeContextUsage(usage, context) + return context } - return applyClaudeContextUsage(usage, lastContextUsage, true) + return undefined + } + + function resolveTurnContextUsage( + billing: TurnUsage | undefined, + measured: SDKControlGetContextUsageResponse | undefined, + request: TurnUsage | undefined, + ): TurnUsage | undefined { + if (measured) return applyClaudeContextUsage(billing, measured) + if (request) return { ...(billing ?? {}), ...request } + return applyClaudeContextUsage(billing, lastContextUsage, true) } queueMicrotask(() => emit({ type: 'ready', bridgeId: opts.bridgeId })) @@ -597,6 +639,7 @@ export async function createClaudeBridge( detail: questionRequest.detail, options: questionRequest.options, placeholder: questionRequest.placeholder, + multiple: questionRequest.multiple || undefined, source: 'claude', }) if (response.action === 'cancel' || response.action === 'decline') { @@ -802,6 +845,7 @@ export async function createClaudeBridge( if (message.subtype === 'compact_boundary') { const metadata = (message as { compact_metadata?: Record }).compact_metadata ?? {} const automatic = stringValue(metadata.trigger) === 'auto' + lastContextUsage = undefined emit({ type: 'compaction_event', bridgeId: opts.bridgeId, @@ -811,6 +855,7 @@ export async function createClaudeBridge( provider: 'claude', beforeTokens: numberValue(metadata.pre_tokens), afterTokens: numberValue(metadata.post_tokens), + resetContext: true, message: automatic ? 'Claude auto-compacted context. Continue the conversation normally.' : 'Claude compacted context. Continue the conversation normally.', @@ -871,6 +916,8 @@ export async function createClaudeBridge( const turnId = startTurn() const state: TurnState = { sawStreamEvent: false, emittedAnyText: false, streamBlockTypeByIndex: {}, textByBlock: {}, thinkingByBlock: {}, streamedText: '', streamedThinking: '', thinkingTokens: 0, thinkingStatusActive: false } let lastUsage: TurnUsage | undefined + let turnContextUsage: SDKControlGetContextUsageResponse | undefined + let requestContextUsage: TurnUsage | undefined let turnSucceeded = false const env: Record = {} @@ -917,11 +964,27 @@ export async function createClaudeBridge( for await (const message of q) { const usage = handleMessage(message, turnId, state) + if (message.type === 'system' && message.subtype === 'compact_boundary') { + // Pre-boundary measurements belong to the old prompt. Only a new + // assistant reading (or a result-time control read) can refill it. + turnContextUsage = undefined + requestContextUsage = undefined + } if (message.type === 'result' && message.subtype === 'success' && !message.is_error) turnSucceeded = true - if (usage) lastUsage = await applyStableContextUsage(usage, q) + if (message.type === 'assistant') { + const assistant = message.message as unknown as Record + const actualModel = stringValue(assistant.model) + if (actualModel) sessionModel = actualModel + requestContextUsage = contextFromAssistantRequest(assistant, sessionModel) ?? requestContextUsage + turnContextUsage = await readMeasuredContext(q) ?? turnContextUsage + } + if (usage) lastUsage = usage + if (message.type === 'result' && !turnContextUsage) { + turnContextUsage = await readMeasuredContext(q) ?? turnContextUsage + } } - if (!lastUsage?.contextTokens) lastUsage = await applyStableContextUsage(lastUsage, q) + lastUsage = resolveTurnContextUsage(lastUsage, turnContextUsage, requestContextUsage) if (aborted) { endTurn('aborted', lastUsage); return { ok: false, error: 'aborted' } } endTurn(turnSucceeded ? 'success' : 'error', lastUsage) diff --git a/src/main/agents/codex-bridge.test.ts b/src/main/agents/codex-bridge.test.ts index 3dd0957..5f3dbd0 100644 --- a/src/main/agents/codex-bridge.test.ts +++ b/src/main/agents/codex-bridge.test.ts @@ -1,6 +1,7 @@ import { describe, expect, it } from 'vitest' -import { CODEX_COMPACT_METHOD, codexApprovalDecisionForMode, getModeConfig, mapCodexEffort, usageFromCodexTokenUsage } from './codex-bridge' +import { CODEX_COMPACT_METHOD, codexApprovalDecisionForMode, codexNativeCompactionStatus, getModeConfig, mapCodexEffort, usageFromCodexTokenUsage } from './codex-bridge' +import { normalizeContextUsage } from './compaction-meter' describe('codex bridge mode config', () => { it.each([ @@ -46,6 +47,13 @@ describe('codex bridge mode config', () => { expect(CODEX_COMPACT_METHOD).toBe('thread/compact/start') }) + it('recognizes the native Codex compaction lifecycle and legacy completion', () => { + expect(codexNativeCompactionStatus('item/started', { type: 'contextCompaction', id: 'compact-1' })).toBe('started') + expect(codexNativeCompactionStatus('item/completed', { type: 'contextCompaction', id: 'compact-1' })).toBe('completed') + expect(codexNativeCompactionStatus('thread/compacted', undefined)).toBe('completed') + expect(codexNativeCompactionStatus('item/started', { type: 'agentMessage', id: 'message-1' })).toBeNull() + }) + it('passes native reasoning effort through without downgrading xhigh', () => { expect(mapCodexEffort('off')).toBeUndefined() expect(mapCodexEffort('low')).toBe('low') @@ -68,18 +76,17 @@ describe('codex bridge mode config', () => { outputTokens: 400, totalTokens: 12_400, contextTokens: 12_400, - // 200K documented window wins over the 258,400 prompt budget Codex reports. - contextWindow: 200_000, + contextWindow: 258_400, model: 'gpt-5.4-mini', }) }) - it('uses the documented Codex GPT-5.5 context window instead of the prompt budget', () => { + it('separates Codex prompt budget from known model capacity', () => { const usage = usageFromCodexTokenUsage({ tokenUsage: { modelContextWindow: 272_000, total: { inputTokens: 120_000, outputTokens: 7_000, totalTokens: 127_000 }, - last: { inputTokens: 18_000, outputTokens: 600, totalTokens: 18_600 }, + last: { inputTokens: 18_000, outputTokens: 600, cachedInputTokens: 4_000, reasoningOutputTokens: 250, totalTokens: 18_600 }, }, }, 'gpt-5.5') @@ -88,11 +95,20 @@ describe('codex bridge mode config', () => { outputTokens: 600, contextTokens: 18_600, contextWindow: 400_000, + promptBudgetTokens: 272_000, + contextIsAuthoritative: true, + contextBreakdownSource: 'usage', + contextBreakdown: [ + { name: 'Latest request input', tokens: 18_000 }, + { name: 'Latest request output', tokens: 600 }, + { name: 'Cached input (included in input)', tokens: 4_000 }, + { name: 'Reasoning output (included in output)', tokens: 250 }, + ], model: 'gpt-5.5', }) }) - it('caps displayed Codex context usage at the selected model window', () => { + it('caps displayed Codex context usage at the model window', () => { const usage = usageFromCodexTokenUsage({ tokenUsage: { modelContextWindow: 272_000, @@ -106,7 +122,56 @@ describe('codex bridge mode config', () => { outputTokens: 963, contextTokens: 400_000, contextWindow: 400_000, + promptBudgetTokens: 272_000, model: 'gpt-5.5', }) }) + + it('falls back to model metadata when the app-server omits its window', () => { + const usage = usageFromCodexTokenUsage({ + tokenUsage: { last: { inputTokens: 18_000, outputTokens: 600 } }, + }, 'gpt-5.5') + + expect(usage?.contextWindow).toBe(400_000) + }) + + it.each(['gpt-6.1-sol', 'gpt-6-sol', 'gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna'])( + 'shows the full 1.05M window for %s separately from Codex\'s smaller prompt budget', model => { + const usage = usageFromCodexTokenUsage({ + tokenUsage: { + modelContextWindow: 828_400, + last: { inputTokens: 18_000, outputTokens: 592 }, + }, + }, model) + + expect(usage).toMatchObject({ + contextTokens: 18_592, + contextWindow: 1_050_000, + promptBudgetTokens: 828_400, + contextIsAuthoritative: true, + }) + }, + ) + + it('uses the reported budget as the window when full model capacity is unknown', () => { + const usage = usageFromCodexTokenUsage({ + tokenUsage: { modelContextWindow: 828_400, last: { inputTokens: 18_000, outputTokens: 592 } }, + }, 'custom-unknown-model') + + expect(usage?.contextWindow).toBe(828_400) + expect(usage?.promptBudgetTokens).toBeUndefined() + }) + + it('keeps a smaller native reading after Codex compacts context', () => { + const previous = usageFromCodexTokenUsage({ tokenUsage: { + modelContextWindow: 272_000, + last: { inputTokens: 180_000, outputTokens: 2_000 }, + } }, 'gpt-5.5') + const next = usageFromCodexTokenUsage({ tokenUsage: { + modelContextWindow: 272_000, + last: { inputTokens: 28_000, outputTokens: 500 }, + } }, 'gpt-5.5') + + expect(normalizeContextUsage(previous, next, { provider: 'codex' })?.contextTokens).toBe(28_500) + }) }) diff --git a/src/main/agents/codex-bridge.ts b/src/main/agents/codex-bridge.ts index 5340415..802bb9d 100644 --- a/src/main/agents/codex-bridge.ts +++ b/src/main/agents/codex-bridge.ts @@ -2,6 +2,13 @@ import { spawnAgentProcess } from './agent-spawn' import type { AgentBridge, BridgeStartOpts, EmitFn, ModeLevel, RequestUserFn, TurnUsage } from './bridge-types' import { buildUsage, contextWindowFor } from './model-context' import { tripwireForToolCall } from './dangerous-command' +import { + CODEX_REQUEST_USER_INPUT_METHODS, + codexQuestionAnswer, + codexQuestionRequest, + parseCodexUserInputQuestions, + type CodexAnswers, +} from './codex-user-input' // Codex app-server JSON-RPC over stdio. // Protocol: newline-delimited JSON (`"jsonrpc": "2.0"` header omitted on the wire). @@ -68,12 +75,21 @@ function pickNum(obj: Record | undefined, ...keys: string[]): n return undefined } -// Codex reports `modelContextWindow` as its per-request prompt budget, which is -// smaller than the model's real window, so the ctx meter under-reports if we -// trust it. Prefer a documented window when we have one; fall back to the wire -// value for models we don't know. -function codexContextWindowFor(model: string | undefined, reported: number | undefined): number | undefined { - return contextWindowFor(model) ?? reported +// Codex's usage notification reports an effective per-request prompt budget. +// The CLI's model window can be larger because it includes reserved capacity. +// Keep both values so the meter names the full model capacity while exposing +// the tighter active budget instead of presenting it as the model's size. +function codexContextWindows(model: string | undefined, reported: number | undefined): { contextWindow?: number; promptBudgetTokens?: number } { + const modelWindow = contextWindowFor(model) + const promptBudgetTokens = reported !== undefined && reported > 0 ? reported : undefined + return { + contextWindow: modelWindow !== undefined && promptBudgetTokens !== undefined + ? Math.max(modelWindow, promptBudgetTokens) + : modelWindow ?? promptBudgetTokens, + ...(promptBudgetTokens !== undefined && modelWindow !== undefined && modelWindow > promptBudgetTokens + ? { promptBudgetTokens } + : {}), + } } function codexContextTokensFor(inputTokens: number | undefined, outputTokens: number | undefined, contextWindow: number | undefined): number | undefined { @@ -88,7 +104,7 @@ export function usageFromCodexTokenUsage(params: Record, model: if (!last) return undefined const inputTokens = pickNum(last, 'inputTokens', 'input_tokens') const outputTokens = pickNum(last, 'outputTokens', 'output_tokens') - const contextWindow = codexContextWindowFor(model, pickNum(tu, 'modelContextWindow', 'model_context_window', 'contextWindow')) + const { contextWindow, promptBudgetTokens } = codexContextWindows(model, pickNum(tu, 'modelContextWindow', 'model_context_window', 'contextWindow')) const contextTokens = codexContextTokensFor(inputTokens, outputTokens, contextWindow) const usage = buildUsage({ inputTokens, @@ -99,7 +115,32 @@ export function usageFromCodexTokenUsage(params: Record, model: contextWindow, model, }) - return usage + if (!usage) return undefined + const rows: NonNullable = [] + const add = (name: string, value: number | undefined) => { + if (value !== undefined && Number.isFinite(value) && value > 0) rows.push({ name, tokens: value }) + } + add('Latest request input', inputTokens) + add('Latest request output', outputTokens) + add('Cached input (included in input)', pickNum(last, 'cachedInputTokens', 'cached_input_tokens')) + add('Reasoning output (included in output)', pickNum(last, 'reasoningOutputTokens', 'reasoning_output_tokens')) + return { + ...usage, + ...(promptBudgetTokens !== undefined ? { promptBudgetTokens } : {}), + ...(inputTokens !== undefined ? { contextIsAuthoritative: true } : {}), + ...(rows.length > 0 ? { contextBreakdown: rows, contextBreakdownSource: 'usage' as const } : {}), + } +} + +export function codexNativeCompactionStatus( + method: string, + item: Record | undefined, +): 'started' | 'completed' | null { + if (method === 'thread/compacted') return 'completed' + if (item?.type !== 'contextCompaction') return null + if (method === 'item/started') return 'started' + if (method === 'item/completed') return 'completed' + return null } // Codex app-server accepts these native effort values. Keep unsupported values @@ -155,6 +196,9 @@ export async function createCodexBridge( const toolStartedItems = new Set() let activePlanToolCallId: string | null = null let activePlanArgs: { plan: unknown[] } | null = null + // Open question cards keyed by the server request id, so `serverRequest/resolved` + // (Codex settled it itself) or process exit can withdraw the card. + const openQuestionRequests = new Map() function emitReady() { emit({ type: 'ready', bridgeId: opts.bridgeId }) @@ -186,6 +230,40 @@ export async function createCodexBridge( send({ id, result }) } + function respondError(id: number, message: string): void { + send({ id, error: { code: -32000, message } }) + } + + // One overlay card per question, in order. Any unanswered question fails the + // whole request explicitly — Codex never receives a fabricated/empty answer. + async function handleUserInputRequest(msg: JsonRpcServerRequest): Promise { + const questions = parseCodexUserInputQuestions(msg.params) + if (!requestUser || questions.length === 0) { + respondError(msg.id, requestUser ? 'request_user_input had no questions' : 'no interactive user is attached') + return + } + const controller = new AbortController() + openQuestionRequests.set(msg.id, controller) + const turnId = typeof msg.params.turnId === 'string' ? msg.params.turnId : currentTurnId ?? undefined + try { + const answers: CodexAnswers = {} + for (let index = 0; index < questions.length; index++) { + const question = questions[index] + const response = await requestUser(codexQuestionRequest(question, index, questions.length, turnId), controller.signal) + if (controller.signal.aborted) return + const answer = codexQuestionAnswer(question, response) + if (!answer) { + respondError(msg.id, 'user cancelled request_user_input') + return + } + answers[question.id] = { answers: answer } + } + respond(msg.id, { answers }) + } finally { + openQuestionRequests.delete(msg.id) + } + } + function handleResponse(msg: JsonRpcResponseOk | JsonRpcResponseErr) { const p = pending.get(msg.id) if (!p) return @@ -205,6 +283,42 @@ export async function createCodexBridge( // notification (not turn/completed), so we hold the latest snapshot and flush // it when the turn ends. let lastUsage: TurnUsage | undefined + let nativeCompactionActive = false + let nativeCompactionCompletedTurnId: string | undefined + + function emitNativeCompaction(status: 'started' | 'completed' | 'failed', notificationTurnId?: string) { + const turnId = currentTurnId ?? notificationTurnId + const completionKey = notificationTurnId ?? currentTurnId ?? undefined + if (status === 'started') { + if (nativeCompactionActive) return + nativeCompactionActive = true + } else if (status === 'completed') { + if (!nativeCompactionActive && completionKey && nativeCompactionCompletedTurnId === completionKey) return + nativeCompactionActive = false + nativeCompactionCompletedTurnId = completionKey + // A usage notification from before compaction must not be replayed by + // turn_end when the app-server has not reported the new prompt yet. + lastUsage = undefined + } else { + if (!nativeCompactionActive) return + nativeCompactionActive = false + } + emit({ + type: 'compaction_event', + bridgeId: opts.bridgeId, + ...(turnId ? { turnId } : {}), + status, + automatic: true, + provider: 'codex', + percent: status === 'completed' ? 100 : undefined, + message: status === 'started' + ? 'Codex is compacting its native context…' + : status === 'completed' + ? 'Codex compacted its native context. Continuing normally.' + : 'Codex native context compaction stopped before completion.', + ...(status === 'completed' ? { resetContext: true } : {}), + }) + } function endTurn(usage?: TurnUsage) { if (!currentTurnId) return @@ -409,6 +523,19 @@ export async function createCodexBridge( } function handleNotification(msg: JsonRpcNotification) { + const notificationItem = msg.params.item as Record | undefined + const compactionStatus = codexNativeCompactionStatus(msg.method, notificationItem) + if (compactionStatus) { + emitNativeCompaction(compactionStatus, typeof msg.params.turnId === 'string' ? msg.params.turnId : undefined) + return + } + if (msg.method === 'serverRequest/resolved') { + // Codex settled a request without us (e.g. non-blocking question). + // Withdraw the card rather than leave a stale "agent is waiting" prompt. + const id = Number(msg.params.requestId) + openQuestionRequests.get(id)?.abort() + return + } switch (msg.method) { case 'thread/started': emitReady() @@ -425,7 +552,12 @@ export async function createCodexBridge( return case 'turn/completed': + if (nativeCompactionActive) emitNativeCompaction('failed') + endTurn() + return + case 'turn/failed': + emitNativeCompaction('failed') endTurn() return @@ -544,22 +676,8 @@ export async function createCodexBridge( respond(msg.id, { decision: response.action === 'accept' || response.action === 'submit' ? 'accept' : 'decline' }) return } - if (msg.method === 'tool/requestUserInput') { - if (!requestUser) { - respond(msg.id, { decision: 'decline' }) - return - } - const response = await requestUser({ - kind: 'prompt', - turnId: currentTurnId ?? undefined, - title: typeof msg.params.title === 'string' ? msg.params.title : 'Agent Questions', - message: typeof msg.params.prompt === 'string' ? msg.params.prompt : typeof msg.params.message === 'string' ? msg.params.message : undefined, - detail: requestDetail(msg.params), - source: 'codex', - }) - respond(msg.id, response.action === 'submit' || response.action === 'accept' - ? { input: response.value ?? '', value: response.value ?? '', decision: 'accept' } - : { decision: 'decline' }) + if (CODEX_REQUEST_USER_INPUT_METHODS.has(msg.method)) { + await handleUserInputRequest(msg) return } if (msg.method === 'account/chatgptAuthTokens/refresh') { @@ -592,6 +710,9 @@ export async function createCodexBridge( }) proc.on('close', code => { + // The process that asked can no longer receive an answer. + for (const controller of openQuestionRequests.values()) controller.abort() + openQuestionRequests.clear() endTurn() // Surface the process's stderr tail so a nonzero exit (127 = binary not on // the host PATH, etc.) is legible instead of a bare exit code. diff --git a/src/main/agents/codex-user-input.test.ts b/src/main/agents/codex-user-input.test.ts new file mode 100644 index 0000000..c7dd2b5 --- /dev/null +++ b/src/main/agents/codex-user-input.test.ts @@ -0,0 +1,66 @@ +import { describe, expect, it } from 'vitest' + +import { + CODEX_REQUEST_USER_INPUT_METHODS, + codexQuestionAnswer, + codexQuestionRequest, + parseCodexUserInputQuestions, +} from './codex-user-input' + +// Shape from codex-cli 0.157.1 ToolRequestUserInputParams. +const params = { + threadId: 't1', + turnId: 'turn-1', + itemId: 'item-1', + isBlocking: true, + questions: [ + { id: 'confirm', header: 'Deploy', question: 'Deploy to staging?', options: [{ label: 'Yes', description: '' }, { label: 'No', description: '' }] }, + { id: 'branch', header: 'Branch', question: 'Which branch name?', isOther: true, options: [{ label: 'main', description: 'default' }] }, + { id: 'token', header: 'Token', question: 'Paste the API token', isSecret: true, options: null }, + ], +} + +describe('codex request_user_input', () => { + it('listens on the real v2 method and keeps the legacy alias', () => { + expect(CODEX_REQUEST_USER_INPUT_METHODS.has('item/tool/requestUserInput')).toBe(true) + expect(CODEX_REQUEST_USER_INPUT_METHODS.has('tool/requestUserInput')).toBe(true) + }) + + it('parses every question with stable ids', () => { + const questions = parseCodexUserInputQuestions(params) + expect(questions.map(q => q.id)).toEqual(['confirm', 'branch', 'token']) + expect(questions[0].options).toEqual([{ label: 'Yes' }, { label: 'No' }]) + }) + + it('maps option-only questions to button cards and "other" questions to input cards', () => { + const [confirm, branch, token] = parseCodexUserInputQuestions(params) + expect(codexQuestionRequest(confirm, 0, 3, 'turn-1')).toMatchObject({ + kind: 'select', + title: 'Deploy to staging?', + message: 'Codex asks: Deploy Question 1 of 3.', + options: [{ id: 'opt-0', label: 'Yes' }, { id: 'opt-1', label: 'No' }], + turnId: 'turn-1', + source: 'codex', + }) + expect(codexQuestionRequest(branch, 1, 3)).toMatchObject({ kind: 'prompt', placeholder: 'reply to Codex…' }) + expect(codexQuestionRequest(token, 2, 3)).toMatchObject({ kind: 'prompt', secret: true, options: undefined }) + }) + + it('answers with the option label or typed text', () => { + const [confirm, branch] = parseCodexUserInputQuestions(params) + expect(codexQuestionAnswer(confirm, { requestId: 'r', action: 'submit', optionId: 'opt-1' })).toEqual(['No']) + expect(codexQuestionAnswer(branch, { requestId: 'r', action: 'submit', value: ' feature/x ' })).toEqual(['feature/x']) + }) + + it('never turns a cancel, empty reply, or disallowed free text into an answer', () => { + const [confirm, branch] = parseCodexUserInputQuestions(params) + expect(codexQuestionAnswer(confirm, { requestId: 'r', action: 'cancel' })).toBeNull() + expect(codexQuestionAnswer(branch, { requestId: 'r', action: 'submit', value: ' ' })).toBeNull() + expect(codexQuestionAnswer(confirm, { requestId: 'r', action: 'submit', value: 'maybe' })).toBeNull() + }) + + it('accepts the legacy single-prompt shape as a free-text question', () => { + const [legacy] = parseCodexUserInputQuestions({ title: 'Input', prompt: 'Name the file' }) + expect(legacy).toMatchObject({ id: 'input', question: 'Name the file', isOther: true }) + }) +}) diff --git a/src/main/agents/codex-user-input.ts b/src/main/agents/codex-user-input.ts new file mode 100644 index 0000000..5f3d4fb --- /dev/null +++ b/src/main/agents/codex-user-input.ts @@ -0,0 +1,116 @@ +import type { AgentUserRequest, AgentUserResponse } from './bridge-types' + +/** + * Codex app-server `item/tool/requestUserInput` (ToolRequestUserInputParams, + * verified against codex-cli 0.157.1 `app-server generate-json-schema`). + * Each question carries a stable `id`; the response maps id -> answers[]. + */ +export const CODEX_REQUEST_USER_INPUT_METHODS: ReadonlySet = new Set([ + 'item/tool/requestUserInput', + // Pre-v2 name CrewCode originally listened for; kept so older app-servers + // still reach the card instead of the unknown-request fallback. + 'tool/requestUserInput', +]) + +export interface CodexUserInputOption { + label: string + description?: string +} + +export interface CodexUserInputQuestion { + id: string + header: string + question: string + /** Codex's "Other" affordance: free-text answers are accepted. */ + isOther: boolean + isSecret: boolean + options: CodexUserInputOption[] +} + +type CodexQuestionRequest = Omit + +function record(value: unknown): Record | null { + return value && typeof value === 'object' && !Array.isArray(value) ? value as Record : null +} + +function text(value: unknown): string { + return typeof value === 'string' ? value.trim() : '' +} + +function parseOptions(value: unknown): CodexUserInputOption[] { + if (!Array.isArray(value)) return [] + return value.flatMap(item => { + const row = record(item) + const label = text(row?.label) + if (!label) return [] + const description = text(row?.description) + return [{ label, ...(description ? { description } : {}) }] + }) +} + +/** Parse the v2 `questions` array; falls back to the legacy single-prompt shape. */ +export function parseCodexUserInputQuestions(params: Record): CodexUserInputQuestion[] { + if (Array.isArray(params.questions)) { + return params.questions.flatMap((item, index) => { + const row = record(item) + const question = text(row?.question) + if (!row || !question) return [] + return [{ + id: text(row.id) || `q${index + 1}`, + header: text(row.header), + question, + isOther: row.isOther === true, + isSecret: row.isSecret === true, + options: parseOptions(row.options), + }] + }) + } + const legacy = text(params.prompt) || text(params.message) || text(params.title) + if (!legacy) return [] + return [{ id: 'input', header: text(params.title), question: legacy, isOther: true, isSecret: false, options: [] }] +} + +function optionId(index: number): string { + return `opt-${index}` +} + +/** One overlay card per question: option buttons, plus free text when Codex allows it. */ +export function codexQuestionRequest(question: CodexUserInputQuestion, index: number, total: number, turnId?: string): CodexQuestionRequest { + const freeText = question.isOther || question.options.length === 0 + const message = [ + question.header ? `Codex asks: ${question.header}` : undefined, + total > 1 ? `Question ${index + 1} of ${total}.` : undefined, + ].filter(Boolean).join(' ') + return { + kind: freeText ? 'prompt' : 'select', + turnId, + title: question.question, + message: message || undefined, + options: question.options.length > 0 + ? question.options.map((option, i) => ({ id: optionId(i), label: option.label, description: option.description })) + : undefined, + placeholder: freeText ? 'reply to Codex…' : undefined, + secret: question.isSecret || undefined, + source: 'codex', + } +} + +/** + * Map a card response to Codex answers. `null` means the human did not answer + * (cancel/decline or an empty submission) — never send that as an answer. + */ +export function codexQuestionAnswer(question: CodexUserInputQuestion, response: AgentUserResponse): string[] | null { + if (response.action !== 'submit' && response.action !== 'accept') return null + if (response.optionId) { + const match = /^opt-(\d+)$/.exec(response.optionId) + const option = match ? question.options[Number(match[1])] : undefined + if (option) return [option.label] + } + const typed = response.value?.trim() ?? '' + if (!typed) return null + // A typed answer is only valid where Codex offered free text. + if (!question.isOther && question.options.length > 0) return null + return [typed] +} + +export type CodexAnswers = Record diff --git a/src/main/agents/compaction-meter.test.ts b/src/main/agents/compaction-meter.test.ts index db00997..cc02124 100644 --- a/src/main/agents/compaction-meter.test.ts +++ b/src/main/agents/compaction-meter.test.ts @@ -1,6 +1,6 @@ import { describe, expect, it } from 'vitest' -import { autoCompactionSignalForProvider, detectAutoCompaction, normalizeContextUsage, compactionStrategy } from './compaction-meter' +import { autoCompactionSignalForProvider, detectAutoCompaction, normalizeContextUsage, compactionStrategy, shouldInferAutoCompactionForProvider } from './compaction-meter' describe('compactionStrategy', () => { const base = { httpOnly: false, nativeResume: false, hasConversationKey: true } @@ -33,9 +33,10 @@ describe('compactionStrategy', () => { describe('autoCompactionSignalForProvider', () => { it.each([ ['claude', 'native'], - ['crewcoder', 'native'], - ['codex', 'usage'], + ['crewcoder', 'native-or-usage'], + ['codex', 'native-or-usage'], ['opencode', 'usage'], + ['grok', 'usage'], ['pi', 'none'], ['hermes', 'none'], ['ollama', 'none'], @@ -61,6 +62,26 @@ describe('detectAutoCompaction', () => { it('ignores ordinary context changes', () => { expect(detectAutoCompaction(before, { contextTokens: 160_000, contextWindow: 200_000 }, 'usage_update')).toBeNull() }) + + it('uses Codex prompt budget to recognize compaction below 80% of full model capacity', () => { + const codexBefore = { contextTokens: 680_000, contextWindow: 1_050_000, promptBudgetTokens: 828_400 } + const codexAfter = { contextTokens: 35_000, contextWindow: 1_050_000, promptBudgetTokens: 828_400 } + expect(detectAutoCompaction(codexBefore, codexAfter, 'turn_end')).toBe('detected') + }) +}) + +describe('shouldInferAutoCompactionForProvider', () => { + it('uses occupancy fallback for CrewCoder and Codex only when no native boundary arrived', () => { + expect(shouldInferAutoCompactionForProvider('crewcoder', false)).toBe(true) + expect(shouldInferAutoCompactionForProvider('codex', false)).toBe(true) + expect(shouldInferAutoCompactionForProvider('crewcoder', true)).toBe(false) + expect(shouldInferAutoCompactionForProvider('codex', true)).toBe(false) + }) + + it('keeps native-only and unobservable providers out of inference', () => { + expect(shouldInferAutoCompactionForProvider('claude', false)).toBe(false) + expect(shouldInferAutoCompactionForProvider('pi', false)).toBe(false) + }) }) describe('normalizeContextUsage', () => { @@ -126,6 +147,19 @@ describe('normalizeContextUsage', () => { expect(usage?.compaction).toBeUndefined() }) + it('accepts a lower provider-authoritative live context after native compaction', () => { + expect(normalizeContextUsage( + { contextTokens: 220_000, contextWindow: 1_050_000 }, + { contextTokens: 24_000, contextWindow: 1_050_000, outputTokens: 500, contextIsAuthoritative: true }, + { provider: 'crewcoder' }, + )).toEqual({ + contextTokens: 24_000, + contextWindow: 1_050_000, + outputTokens: 500, + contextIsAuthoritative: true, + }) + }) + it('does not pin authoritative Claude SDK context to a stale full baseline', () => { const usage = normalizeContextUsage( { contextTokens: 1_000_000, contextWindow: 1_000_000, model: 'claude-opus-4-8' }, diff --git a/src/main/agents/compaction-meter.ts b/src/main/agents/compaction-meter.ts index 00bddb2..709f97e 100644 --- a/src/main/agents/compaction-meter.ts +++ b/src/main/agents/compaction-meter.ts @@ -74,14 +74,15 @@ export function compactionStrategy(opts: { export function isLikelyAutoCompaction(previous: TurnUsage | undefined, next: TurnUsage | undefined, threshold = DEFAULT_THRESHOLD_PERCENT): boolean { if (!previous?.contextTokens || !next?.contextTokens || !previous.contextWindow) return false - const previousPercent = (previous.contextTokens / previous.contextWindow) * 100 + const effectiveLimit = previous.promptBudgetTokens ?? previous.contextWindow + const previousPercent = (previous.contextTokens / effectiveLimit) * 100 // Providers rarely announce auto-compaction consistently; a large context drop // after warning-level usage is the safest cross-provider signal we can infer. return previousPercent >= threshold && next.contextTokens < previous.contextTokens * 0.6 } export type AutoCompactionDetection = 'started' | 'detected' | null -export type AutoCompactionSignal = 'native' | 'usage' | 'none' +export type AutoCompactionSignal = 'native' | 'native-or-usage' | 'usage' | 'none' /** * Provider-specific auto-compaction observability. `native` bridges emit their @@ -92,16 +93,24 @@ export type AutoCompactionSignal = 'native' | 'usage' | 'none' export function autoCompactionSignalForProvider(provider: string): AutoCompactionSignal { switch (provider.toLowerCase()) { case 'claude': - case 'crewcoder': return 'native' + case 'crewcoder': case 'codex': + return 'native-or-usage' case 'opencode': + case 'grok': return 'usage' default: return 'none' } } +export function shouldInferAutoCompactionForProvider(provider: string, nativeCompactionObserved: boolean): boolean { + if (nativeCompactionObserved) return false + const signal = autoCompactionSignalForProvider(provider) + return signal === 'usage' || signal === 'native-or-usage' +} + /** * A mid-turn drop can truthfully open a loading meter until turn_end. Providers * that only publish usage at turn_end can only be reported after the fact. @@ -139,14 +148,15 @@ export function normalizeContextUsage(previous: TurnUsage | undefined, next: Tur } const contextTokens = num(usage.contextTokens) - const isAuthoritativeClaudeContext = provider === 'claude' && usage.contextBreakdown !== undefined + const isAuthoritativeContext = usage.contextIsAuthoritative === true + || (provider === 'claude' && usage.contextBreakdown !== undefined) if ( previousUsage?.contextTokens !== undefined && contextTokens !== undefined && contextTokens < previousUsage.contextTokens && !isLikelyAutoCompaction(previousUsage, usage) - && !isAuthoritativeClaudeContext + && !isAuthoritativeContext ) { // A dip below the prior baseline means the provider under-reported the // absolute context (e.g. a fresh resume, or claude's active-only category diff --git a/src/main/agents/crewcoder-bridge.test.ts b/src/main/agents/crewcoder-bridge.test.ts index 149250b..d2fa786 100644 --- a/src/main/agents/crewcoder-bridge.test.ts +++ b/src/main/agents/crewcoder-bridge.test.ts @@ -64,7 +64,7 @@ class FakeAcpProcess extends EventEmitter { } } -function crewCoderAcpHarness(options: { directoryError?: string; compact?: boolean } = {}): { proc: FakeAcpProcess; sent: Array> } { +function crewCoderAcpHarness(options: { directoryError?: string; compact?: boolean; promptResults?: Array> } = {}): { proc: FakeAcpProcess; sent: Array> } { const proc = new FakeAcpProcess() const sent: Array> = [] let input = '' @@ -96,6 +96,14 @@ function crewCoderAcpHarness(options: { directoryError?: string; compact?: boole proc.stdout.write(`${JSON.stringify({ jsonrpc: '2.0', id: message.id, result: { queued: true } })}\n`) } else if (message.method === 'session/set_reasoning_effort') { proc.stdout.write(`${JSON.stringify({ jsonrpc: '2.0', id: message.id, result: {} })}\n`) + } else if (message.method === 'session/prompt') { + if (options.promptResults) { + proc.stdout.write(`${JSON.stringify({ + jsonrpc: '2.0', + id: message.id, + result: options.promptResults.shift() ?? { stopReason: 'end_turn' }, + })}\n`) + } } else if (message.method === 'session/compact') { proc.stdout.write(`${JSON.stringify({ jsonrpc: '2.0', @@ -656,7 +664,15 @@ describe('CrewCoder ACP usage', () => { totalTokens: 1801, contextTokens: 1200, contextWindow: 200_000, + contextIsAuthoritative: true, model: 'codex:gpt-5.6-sol', + contextBreakdown: [ + { name: 'Latest context input', tokens: 1200 }, + { name: 'Session input', tokens: 1234 }, + { name: 'Session output', tokens: 567 }, + { name: 'Session total', tokens: 1801 }, + ], + contextBreakdownSource: 'usage', }) }) @@ -667,9 +683,105 @@ describe('CrewCoder ACP usage', () => { inputTokens: 40, outputTokens: 2, totalTokens: 42, - contextTokens: 42, + contextTokens: undefined, contextWindow: undefined, model: 'custom:model', + contextBreakdown: [ + { name: 'Session input', tokens: 40 }, + { name: 'Session output', tokens: 2 }, + { name: 'Session total', tokens: 42 }, + ], + contextBreakdownSource: 'usage', }) }) + + it('uses top-level live occupancy when an ACP transport strips namespaced metadata', () => { + expect(crewCoderUsageFromPromptResult({ + usage: { + inputTokens: 90_000, + outputTokens: 700, + totalTokens: 90_700, + lastInputTokens: 42_000, + contextWindow: 1_050_000, + }, + }, 'codex:gpt-5.6-sol')).toEqual({ + inputTokens: 90_000, + outputTokens: 700, + totalTokens: 90_700, + contextTokens: 42_000, + contextWindow: 1_050_000, + contextIsAuthoritative: true, + model: 'codex:gpt-5.6-sol', + contextBreakdown: [ + { name: 'Latest context input', tokens: 42_000 }, + { name: 'Session input', tokens: 90_000 }, + { name: 'Session output', tokens: 700 }, + { name: 'Session total', tokens: 90_700 }, + ], + contextBreakdownSource: 'usage', + }) + }) + + it('shows observed cache and reasoning counters without treating them as context categories', () => { + const usage = crewCoderUsageFromPromptResult({ + _meta: { + 'crewcoder/usage': { + inputTokens: 100, + outputTokens: 20, + totalTokens: 120, + lastInputTokens: 75, + cachedInputTokens: 30, + cacheWriteTokens: 10, + reasoningTokens: 5, + }, + }, + }, 'codex:gpt-5.6-sol') + + expect(usage?.contextBreakdownSource).toBe('usage') + expect(usage?.contextBreakdown).toEqual([ + { name: 'Latest context input', tokens: 75 }, + { name: 'Session input', tokens: 100 }, + { name: 'Session output', tokens: 20 }, + { name: 'Session total', tokens: 120 }, + { name: 'Session cached input', tokens: 30 }, + { name: 'Session cache write', tokens: 10 }, + { name: 'Session reasoning output', tokens: 5 }, + ]) + }) + + it('keeps context unknown when only a cumulative session total is reported', () => { + expect(crewCoderUsageFromPromptResult({ usage: { totalTokens: 42 } }, 'custom:model')).toEqual({ + inputTokens: undefined, + outputTokens: undefined, + totalTokens: 42, + contextTokens: undefined, + contextWindow: undefined, + model: 'custom:model', + contextBreakdown: [{ name: 'Session total', tokens: 42 }], + contextBreakdownSource: 'usage', + }) + }) + + it('does not reuse a preceding prompt usage snapshot when the next result omits usage', async () => { + const harness = crewCoderAcpHarness({ + promptResults: [ + { usage: { inputTokens: 42_000, outputTokens: 700, lastInputTokens: 42_000, contextWindow: 1_050_000 } }, + { stopReason: 'end_turn' }, + ], + }) + spawnAgentProcess.mockResolvedValue({ proc: harness.proc, dir: '/repo', remote: false }) + const events: BridgeEvent[] = [] + const bridge = await createCrewCoderBridge('crewcoder', { + bridgeId: 'bridge-usage', provider: 'crewcoder', cwd: '/repo', mode: 'build', + }, event => events.push(event)) + + await bridge.prompt('first') + await vi.waitFor(() => expect(events.filter(event => event.type === 'turn_end')).toHaveLength(1)) + await bridge.prompt('second') + await vi.waitFor(() => expect(events.filter(event => event.type === 'turn_end')).toHaveLength(2)) + + const turns = events.filter(event => event.type === 'turn_end') + expect(turns[0]).toMatchObject({ usage: { contextTokens: 42_000, contextWindow: 1_050_000 } }) + expect(turns[1]).toMatchObject({ usage: undefined }) + }) }) diff --git a/src/main/agents/crewcoder-bridge.ts b/src/main/agents/crewcoder-bridge.ts index 09d1abb..f2dc0b1 100644 --- a/src/main/agents/crewcoder-bridge.ts +++ b/src/main/agents/crewcoder-bridge.ts @@ -11,7 +11,7 @@ import type { RequestUserFn, TurnUsage, } from './bridge-types' -import { buildUsage } from './model-context' +import { contextWindowFor } from './model-context' import { enrichUsageContextWindow } from './openrouter-model-context' import { tripwireForToolCall } from './dangerous-command' import { crewCoderApprovalForProfile, normalizeCrewCoderMode } from '../../shared/crewcoder-types' @@ -343,16 +343,36 @@ export function crewCoderUsageFromPromptResult(result: unknown, model?: string): const source = rich ?? topLevel if (!source) return undefined - const usage = buildUsage({ - inputTokens: source.inputTokens ?? source.input_tokens, - outputTokens: source.outputTokens ?? source.output_tokens, - contextTokens: rich?.lastInputTokens ?? rich?.last_input_tokens, - contextWindow: rich?.contextWindow ?? rich?.context_window, - model, - }) - if (!usage) return undefined + const liveContextTokens = finiteNumber(source.lastInputTokens ?? source.last_input_tokens) + const inputTokens = finiteNumber(source.inputTokens ?? source.input_tokens) + const outputTokens = finiteNumber(source.outputTokens ?? source.output_tokens) const explicitTotal = finiteNumber(source.totalTokens ?? source.total_tokens) - return explicitTotal === undefined ? usage : { ...usage, totalTokens: explicitTotal } + if (inputTokens === undefined && outputTokens === undefined && explicitTotal === undefined && liveContextTokens === undefined) return undefined + const totalTokens = explicitTotal ?? (inputTokens !== undefined || outputTokens !== undefined + ? (inputTokens ?? 0) + (outputTokens ?? 0) + : undefined) + const rows: NonNullable = [] + const add = (name: string, value: unknown) => { + const tokens = finiteNumber(value) + if (tokens !== undefined && tokens > 0) rows.push({ name, tokens }) + } + add('Latest context input', liveContextTokens) + add('Session input', source.inputTokens ?? source.input_tokens) + add('Session output', source.outputTokens ?? source.output_tokens) + add('Session total', explicitTotal) + add('Session cached input', source.cachedInputTokens ?? source.cached_input_tokens) + add('Session cache write', source.cacheWriteTokens ?? source.cache_write_tokens) + add('Session reasoning output', source.reasoningTokens ?? source.reasoning_tokens) + return { + inputTokens, + outputTokens, + totalTokens, + contextTokens: liveContextTokens, + contextWindow: finiteNumber(source.contextWindow ?? source.context_window) ?? contextWindowFor(model), + model, + ...(liveContextTokens === undefined ? {} : { contextIsAuthoritative: true }), + ...(rows.length > 0 ? { contextBreakdown: rows, contextBreakdownSource: 'usage' as const } : {}), + } } export function writeBlocked(opts: Pick): boolean { @@ -901,6 +921,9 @@ export async function createCrewCoderBridge( } if (currentTurnId) return { ok: false, error: 'crewcoder acp: a turn is already running' } if (proc.stdin.destroyed || !proc.stdin.writable) return { ok: false, error: 'crewcoder acp: process not writable' } + // Usage belongs to one ACP prompt. Never attach the preceding turn's + // snapshot when this prompt fails or returns without usage metadata. + lastUsage = undefined startTurn() void (async () => { try { diff --git a/src/main/agents/grok-bridge.test.ts b/src/main/agents/grok-bridge.test.ts index d3db968..6568f90 100644 --- a/src/main/agents/grok-bridge.test.ts +++ b/src/main/agents/grok-bridge.test.ts @@ -226,6 +226,15 @@ describe('grokUsageFromPromptResult', () => { expect(usage?.outputTokens).toBe(32) expect(usage?.contextTokens).toBe(12298) expect(usage?.contextWindow).toBe(500000) + expect(usage?.contextBreakdownSource).toBe('usage') + expect(usage?.contextBreakdown).toEqual([ + { name: 'Latest call input', tokens: 12266 }, + { name: 'Latest call output', tokens: 32 }, + { name: 'Cache read (reported separately)', tokens: 128 }, + { name: 'Reasoning (reported separately)', tokens: 20 }, + { name: 'Turn total input', tokens: 24440 }, + { name: 'Turn total output', tokens: 100 }, + ]) }) it('reports the cumulative total when present', () => { @@ -515,6 +524,10 @@ describe('grok bridge vendor notification channel', () => { expect(usage).toBeDefined() expect(usage && 'usage' in usage ? usage.usage.inputTokens : undefined).toBe(12046) expect(usage && 'usage' in usage ? usage.usage.contextWindow : undefined).toBe(500000) + expect(usage && 'usage' in usage ? usage.usage.contextBreakdown : undefined).toEqual([ + { name: 'Latest call input', tokens: 12046 }, + { name: 'Latest call output', tokens: 68 }, + ]) }) it('ignores vendor chrome notifications', async () => { diff --git a/src/main/agents/grok-bridge.ts b/src/main/agents/grok-bridge.ts index 84282ef..d6e01fa 100644 --- a/src/main/agents/grok-bridge.ts +++ b/src/main/agents/grok-bridge.ts @@ -252,6 +252,27 @@ export function grokSelectedOption( // Usage // --------------------------------------------------------------------------- +function grokTokenRows(values: { + latestInput?: number + latestOutput?: number + cacheRead?: number + reasoning?: number + turnInput?: number + turnOutput?: number +}): NonNullable { + const rows: NonNullable = [] + const add = (name: string, value: number | undefined) => { + if (value !== undefined && Number.isFinite(value) && value > 0) rows.push({ name, tokens: value }) + } + add('Latest call input', values.latestInput) + add('Latest call output', values.latestOutput) + add('Cache read (reported separately)', values.cacheRead) + add('Reasoning (reported separately)', values.reasoning) + add('Turn total input', values.turnInput) + add('Turn total output', values.turnOutput) + return rows +} + /** * Per-turn usage rides on the `session/prompt` response `_meta`. `inputTokens` * there is the last model call's input, i.e. live context occupancy — the same @@ -278,7 +299,19 @@ export function grokUsageFromPromptResult( }) if (!usage) return undefined const explicitTotal = finiteNumber(cumulative?.totalTokens) - return explicitTotal === undefined ? usage : { ...usage, totalTokens: explicitTotal } + const rows = grokTokenRows({ + latestInput: finiteNumber(meta.inputTokens), + latestOutput: finiteNumber(meta.outputTokens), + cacheRead: finiteNumber(meta.cachedReadTokens), + reasoning: finiteNumber(meta.reasoningTokens), + turnInput: finiteNumber(cumulative?.inputTokens), + turnOutput: finiteNumber(cumulative?.outputTokens), + }) + return { + ...usage, + ...(explicitTotal === undefined ? {} : { totalTokens: explicitTotal }), + ...(rows.length > 0 ? { contextBreakdown: rows, contextBreakdownSource: 'usage' as const } : {}), + } } /** Context window for the active model, read off initialize/session-new `_meta`. */ @@ -703,8 +736,14 @@ export async function createGrokBridge( model: opts.model, }) if (!built || !currentTurnId) return - lastUsage = built - emit({ type: 'usage_update', bridgeId: opts.bridgeId, turnId: currentTurnId, usage: built }) + const rows = grokTokenRows({ + latestInput: finiteNumber(usage.input_tokens), + latestOutput: finiteNumber(usage.output_tokens), + cacheRead: finiteNumber(usage.cached_read_tokens), + reasoning: finiteNumber(usage.reasoning_tokens), + }) + lastUsage = rows.length > 0 ? { ...built, contextBreakdown: rows, contextBreakdownSource: 'usage' } : built + emit({ type: 'usage_update', bridgeId: opts.bridgeId, turnId: currentTurnId, usage: lastUsage }) } function resolveClientPath(value: string): string { diff --git a/src/main/agents/index.ts b/src/main/agents/index.ts index baebd83..2e531ca 100644 --- a/src/main/agents/index.ts +++ b/src/main/agents/index.ts @@ -20,7 +20,7 @@ import { requiredPermissionsForPluginAgentRuntime } from '../plugin-contract' import { parseProviderPayload } from './plugin-provider-payload' import type { AgentBridge, AgentUserRequest, AgentUserResponse, BridgeEvent, BridgeStartOpts, EmitFn, HandoffPromptOptions, PromptOptions, RequestUserFn, TurnUsage } from './bridge-types' import { HTTP_ONLY_PROVIDERS, API_KEY_PROVIDERS } from './bridge-types' -import { autoCompactionSignalForProvider, detectAutoCompaction, normalizeContextUsage, compactionStrategy } from './compaction-meter' +import { detectAutoCompaction, normalizeContextUsage, compactionStrategy, shouldInferAutoCompactionForProvider } from './compaction-meter' import { TurnPermissionGrantStore } from './turn-permission-grants' import { authorityOf, custodyJournal, scopeKeyFor, scopeViolation } from './custody' import { @@ -74,6 +74,10 @@ interface BridgeEntry { // Set after a mid-turn context drop opens the inferred auto-compaction meter; // turn_end closes that same meter. pendingAutoCompaction: boolean + // Native provider boundaries win over usage-drop inference. Reset for every + // real turn so providers with incomplete native telemetry can still use the + // observed occupancy fallback on a later turn. + nativeCompactionObservedThisTurn: boolean // True while a summary-reset compaction is in flight: the agent is producing a // structured summary that turn_end collapses into a fresh, small session. pendingSummaryReset: boolean @@ -885,7 +889,8 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi }) const win = BrowserWindow.fromWebContents(e.sender) - const requestUser: RequestUserFn = (request) => { + const requestUser: RequestUserFn = (request, signal) => { + if (signal?.aborted) return Promise.resolve({ requestId: `${opts.bridgeId}:withdrawn`, action: 'cancel' }) // Check custody before even an auto-response is prepared. Otherwise a // Full Access grant could approve a request after its scope disappeared. const liveEntry = bridges.get(opts.bridgeId) @@ -901,6 +906,17 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi return new Promise((resolve) => { pendingUserRequests.set(requestId, { bridgeId: opts.bridgeId, request: payload, resolve }) win?.webContents.send('bridge:event', { type: 'user_request', request: payload } satisfies BridgeEvent) + // The provider withdrew the request (it resolved it itself). Settle as + // cancel and retract the card; the answer was never observed. + signal?.addEventListener('abort', () => { + const pending = pendingUserRequests.get(requestId) + if (!pending) return + pendingUserRequests.delete(requestId) + pending.resolve({ requestId, action: 'cancel' }) + if (!win?.isDestroyed()) { + win?.webContents.send('bridge:event', { type: 'user_request_resolved', bridgeId: opts.bridgeId, requestId } satisfies BridgeEvent) + } + }, { once: true }) }) } const emit = (event: BridgeEvent) => { @@ -943,6 +959,10 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi // Keep the idle clock and running flag current for the sweep. const eventBridgeId = event.type === 'user_request' ? event.request.bridgeId : event.bridgeId const entry = bridges.get(eventBridgeId) + if (entry && event.type === 'compaction_event') { + entry.nativeCompactionObservedThisTurn = true + if (event.status === 'completed') entry.pendingAutoCompaction = false + } if (entry && event.type === 'compaction_event' && event.status === 'completed' && event.resetContext) { // Native compaction invalidates the old absolute occupancy immediately. // Do not display or persist a fabricated zero: leave context unknown until @@ -956,7 +976,7 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi // Only providers with a verified absolute-context usage contract use // inference. Claude emits authoritative compact_boundary events; Pi, // Hermes, HTTP providers, and plugins must not be guessed from CLI-ness. - const inferProviderCompaction = autoCompactionSignalForProvider(entry.provider) === 'usage' + const inferProviderCompaction = shouldInferAutoCompactionForProvider(entry.provider, entry.nativeCompactionObservedThisTurn) if (detection && inferProviderCompaction && !entry.pendingManualCompaction && !entry.pendingSummaryReset) { entry.pendingAutoCompaction = detection === 'started' win?.webContents.send('bridge:event', { @@ -1021,6 +1041,7 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi delete entry.assistantTextByTurn[event.turnId] } if (event.type === 'turn_start') { + entry.nativeCompactionObservedThisTurn = false entry.running = true entry.userInitiatedStop = false // This path handles provider-internal queued follow-ups. Direct @@ -1162,6 +1183,7 @@ export function registerAgentBridgeIpc(resolveAgentPath: AgentPathResolver): voi lastUsage: sessionKey ? getUsageSnapshot(sessionKey) : undefined, pendingManualCompaction: false, pendingAutoCompaction: false, + nativeCompactionObservedThisTurn: false, pendingSummaryReset: false, custodyScopeKey, custodyEnabled, diff --git a/src/main/agents/model-context.test.ts b/src/main/agents/model-context.test.ts index b284069..1eae818 100644 --- a/src/main/agents/model-context.test.ts +++ b/src/main/agents/model-context.test.ts @@ -22,14 +22,14 @@ describe('contextWindowFor', () => { expect(contextWindowFor('gpt-test-1m')).toBe(1_000_000) }) - it('keeps exact CrewCode provider windows ahead of catalog aliases', () => { + it('prefers provider catalog windows over exact static fallbacks', () => { // Register a value that differs from the exact rule, otherwise this asserts // nothing about precedence. registerContextWindow('openai/gpt-5.5', 1_050_000) expect(registeredContextWindowFor('openai/gpt-5.5')).toBe(1_050_000) - expect(contextWindowFor('openai/gpt-5.5')).toBe(400_000) - expect(contextWindowFor('gpt-5.5')).toBe(400_000) + expect(contextWindowFor('openai/gpt-5.5')).toBe(1_050_000) + expect(contextWindowFor('gpt-5.5')).toBe(1_050_000) }) it('matches display names to provider catalog ids without guessing by family', () => { diff --git a/src/main/agents/model-context.ts b/src/main/agents/model-context.ts index 59313d0..91f9d5d 100644 --- a/src/main/agents/model-context.ts +++ b/src/main/agents/model-context.ts @@ -1,8 +1,7 @@ import type { TurnUsage } from './bridge-types' -// Per-model context-window sizes (in tokens). Providers rarely report the -// window on the wire, so the bubble's "context usage" first checks live model -// catalog metadata and only then falls back to conservative family rules. +// Per-model context-window fallbacks (in tokens). Bridges prefer any active +// provider report; catalog metadata takes precedence over these static rules. interface WindowRule { match: RegExp @@ -55,10 +54,13 @@ export function registeredContextWindowFor(model: string | undefined): number | return undefined } -// Exact CrewCode provider ids win over catalog aliases; broad family rules stay -// fallbacks so provider metadata can refine non-owned model variants. +// Static values are fallbacks when no provider catalog window was observed. const EXACT_RULES: WindowRule[] = [ // The 5.6 family carries the large window; 5.5 and 5.4 do not. + { match: /^(?:openai\/)?gpt-6\.1-sol$/i, window: 1_050_000 }, + { match: /^(?:openai\/)?gpt-6-sol$/i, window: 1_050_000 }, + { match: /^(?:openai\/)?gpt-6-luna$/i, window: 1_050_000 }, + { match: /^(?:openai\/)?gpt-6-astra$/i, window: 1_050_000 }, { match: /^(?:openai\/)?gpt-5\.6-sol$/i, window: 1_050_000 }, { match: /^(?:openai\/)?gpt-5\.6-terra$/i, window: 1_050_000 }, { match: /^(?:openai\/)?gpt-5\.6-luna$/i, window: 1_050_000 }, @@ -70,12 +72,26 @@ const EXACT_RULES: WindowRule[] = [ // Ordered most-specific → least-specific; first hit wins. const RULES: WindowRule[] = [ + { match: /gpt-6\.1-sol/i, window: 1_050_000 }, +{ match: /gpt-6-sol/i, window: 1_050_000 }, + { match: /gpt-6-luna/i, window: 1_050_000 }, + { match: /gpt-6-astra/i, window: 1_050_000 }, + { match: /gpt-5\.6-sol/i, window: 1_050_000 }, + { match: /gpt-5\.6-terra/i, window: 1_050_000 }, + { match: /gpt-5\.6-luna/i, window: 1_050_000 }, + { match: /gpt-5\.6/i, window: 1_050_000 }, + { match: /gpt-5\.5/i, window: 400_000 }, + { match: /gpt-5\.4/i, window: 400_000 }, + { match: /gpt-5\.4-mini/i, window: 200_000 }, + // Claude Code windows vary by family; Opus exposes a larger 1M context. { match: /claude-opus-5/i, window: 1_000_000 }, + { match: /claude-opus-5-5/i, window: 1_000_000 }, + { match: /claude-sonnet-5-5/i, window: 1_000_000 }, { match: /claude-sonnet-5/i, window: 1_000_000 }, - { match: /claude-fable-5/i, window: 500_000 }, - { match: /claude-opus-4.8/i, window: 500_000 }, - { match: /claude-sonnet-4.6/i, window: 500_000 }, + { match: /claude-fable-5\.1/i, window: 1_000_000 }, + { match: /claude-opus-4\.8/i, window: 500_000 }, + { match: /claude-sonnet-4\.6/i, window: 500_000 }, { match: /claude-haiku-4-5/i, window: 200_000 }, // Google Gemini — 1M (pro variants go to 2M but 1M is the safe floor). { match: /gemini-1\.5-pro/i, window: 128_000 }, @@ -95,11 +111,11 @@ const RULES: WindowRule[] = [ */ export function contextWindowFor(model: string | undefined): number | undefined { if (!model) return undefined + const registered = registeredContextWindowFor(model) + if (registered) return registered for (const rule of EXACT_RULES) { if (rule.match.test(model)) return rule.window } - const registered = registeredContextWindowFor(model) - if (registered) return registered // Family rules are written hyphenated, but ids arrive as display names too // ("Claude Opus 4.8 (latest)"), so match the slug as well or those render no // context percentage at all. diff --git a/src/main/agents/model-detect.ts b/src/main/agents/model-detect.ts index b8e0329..823cfa9 100644 --- a/src/main/agents/model-detect.ts +++ b/src/main/agents/model-detect.ts @@ -280,15 +280,18 @@ function detectClaude(claudePath: string): DetectedModel[] { // Avoid scraping the executable: it includes legacy internal model ids that // are not the current /model picker choices. const models: DetectedModel[] = [ - { id: 'claude-opus-5', label: 'Claude Opus 5 (latest)', provider: 'anthropic', contextWindow: 1_000_000 }, - { id: 'claude-sonnet-5', label: 'Claude Sonnet 5 (latest)', provider: 'anthropic', contextWindow: 1_000_000 }, - { id: 'claude-haiku-4-5', label: 'Claude Haiku 4.5 (latest)', provider: 'anthropic', contextWindow: 200_000 }, + { id: 'claude-opus-5-5', label: 'Claude Opus 5.5 (latest)', provider: 'anthropic', contextWindow: 1_000_000 }, + { id: 'claude-sonnet-5-5', label: 'Claude Sonnet 5.5 (latest)', provider: 'anthropic', contextWindow: 1_000_000 }, + { id: 'claude-haiku-4-5', label: 'Claude Haiku 4.5 (latest)', provider: 'anthropic', contextWindow: 200_000 }, + { id: 'claude-opus-5', label: 'Claude Opus 5', provider: 'anthropic', contextWindow: 1_000_000 }, + { id: 'claude-sonnet-5', label: 'Claude Sonnet 5', provider: 'anthropic', contextWindow: 1_000_000 }, + { id: 'claude-opus-4-8', label: 'Claude Opus 4.8', provider: 'anthropic', contextWindow: 500_000 }, { id: 'claude-sonnet-4-6', label: 'Claude Opus 4.6', provider: 'anthropic', contextWindow: 500_000 }, ] if (advertisedAliases.has('fable')) { - models.push({ id: 'claude-fable-5', label: 'Fable 5', provider: 'anthropic', contextWindow: 500_000 }) + models.push({ id: 'claude-fable-5-1', label: 'Fable 5.1', provider: 'anthropic', contextWindow: 1_000_000 }) } return models diff --git a/src/main/agents/opencode-bridge.test.ts b/src/main/agents/opencode-bridge.test.ts index 16ccc99..8e2c8e9 100644 --- a/src/main/agents/opencode-bridge.test.ts +++ b/src/main/agents/opencode-bridge.test.ts @@ -1,6 +1,22 @@ import { describe, expect, it } from 'vitest' -import { answerOpencodeQuestion, buildOpencodePromptBody, prepareOpencodeQuestionRequest, usageFromOpencodeMessageInfo } from './opencode-bridge' +import { answerOpencodeQuestion, buildOpencodePromptBody, opencodeCompactionEvent, prepareOpencodeQuestionRequest, usageFromOpencodeMessageInfo } from './opencode-bridge' + +describe('opencode compaction', () => { + it('resets context only for the active session completion', () => { + expect(opencodeCompactionEvent({ sessionID: 'other' }, 'active', 'bridge', 'turn')).toBeNull() + expect(opencodeCompactionEvent({ sessionID: 'active' }, 'active', 'bridge', 'turn')).toEqual({ + type: 'compaction_event', + bridgeId: 'bridge', + turnId: 'turn', + status: 'completed', + automatic: true, + provider: 'opencode', + message: 'OpenCode auto-compacted context. Continue the conversation normally.', + resetContext: true, + }) + }) +}) describe('opencode question requests', () => { it('maps OpenCode question events to AgentRequestCard select requests', () => { @@ -38,6 +54,20 @@ describe('opencode question requests', () => { expect(prepared?.request.options).toHaveLength(1) expect(answerOpencodeQuestion(prepared!, { requestId: 'r', action: 'submit', value: 'feature/bar' })).toEqual(['feature/bar']) }) + + it('renders multiple-choice OpenCode questions as toggles answered by option ids', () => { + const prepared = prepareOpencodeQuestionRequest({ + question: 'Which checks?', + multiple: true, + options: [{ label: 'Lint, strict', description: '' }, { label: 'Tests', description: '' }], + }) + + expect(prepared?.request).toMatchObject({ multiple: true, options: [{ label: 'Lint, strict' }, { label: 'Tests' }] }) + // Toggled labels are passed whole, so a comma inside a label is not split. + expect(answerOpencodeQuestion(prepared!, { + requestId: 'r', action: 'submit', optionIds: [prepared!.request.options![0].id, prepared!.request.options![1].id], + })).toEqual(['Lint, strict', 'Tests']) + }) }) describe('opencode prompt body', () => { @@ -85,6 +115,14 @@ describe('opencode prompt body', () => { outputTokens: 700, totalTokens: 18_700, contextTokens: 18_500, + contextBreakdownSource: 'usage', + contextBreakdown: [ + { name: 'Latest request input', tokens: 18_000 }, + { name: 'Latest request output', tokens: 500 }, + { name: 'Reasoning output (reported separately)', tokens: 200 }, + { name: 'Cache read (reported separately)', tokens: 4_000_000 }, + { name: 'Cache write (reported separately)', tokens: 20_000 }, + ], model: 'opencode/claude-opus-4-8', }) }) diff --git a/src/main/agents/opencode-bridge.ts b/src/main/agents/opencode-bridge.ts index edd64c6..5320ed0 100644 --- a/src/main/agents/opencode-bridge.ts +++ b/src/main/agents/opencode-bridge.ts @@ -2,7 +2,7 @@ import { spawnAgentProcess } from './agent-spawn' import { isRemoteRoot, parseRemoteTarget } from '../remote/ssh-target' import { forwardRemotePort } from '../remote/ssh-pool' import type { ForwardHandle } from '../remote/ssh-pool' -import type { AgentBridge, AgentUserRequest, AgentUserResponse, BridgeStartOpts, EmitFn, ModeLevel, RequestUserFn, TurnUsage } from './bridge-types' +import type { AgentBridge, AgentUserRequest, AgentUserResponse, BridgeEvent, BridgeStartOpts, EmitFn, ModeLevel, RequestUserFn, TurnUsage } from './bridge-types' import { buildUsage } from './model-context' import { enrichUsageContextWindow } from './openrouter-model-context' @@ -115,13 +115,41 @@ export function usageFromOpencodeMessageInfo(info: { const reasoning = typeof t.reasoning === 'number' && Number.isFinite(t.reasoning) ? t.reasoning : 0 const outputTokens = outputRaw === undefined ? (reasoning || undefined) : outputRaw + reasoning const contextTokens = input === undefined ? undefined : input + (outputRaw ?? 0) - return buildUsage({ + const usage = buildUsage({ inputTokens: input, outputTokens, // Cache reads are billing/cache-hit accounting, not live prompt size. contextTokens, model: info.modelID ?? fallbackModel, }) + if (!usage) return undefined + const rows: NonNullable = [] + const add = (name: string, value: number | undefined) => { + if (value !== undefined && Number.isFinite(value) && value > 0) rows.push({ name, tokens: value }) + } + add('Latest request input', input) + add('Latest request output', outputRaw) + add('Reasoning output (reported separately)', reasoning) + add('Cache read (reported separately)', t.cache?.read) + add('Cache write (reported separately)', t.cache?.write) + return rows.length > 0 + ? { ...usage, contextBreakdown: rows, contextBreakdownSource: 'usage' } + : usage +} + +/** OpenCode's SSE completion belongs only to the selected session. */ +export function opencodeCompactionEvent(properties: Record, sessionId: string, bridgeId: string, turnId: string | null): BridgeEvent | null { + if (properties.sessionID !== sessionId) return null + return { + type: 'compaction_event', + bridgeId, + ...(turnId ? { turnId } : {}), + status: 'completed', + automatic: true, + provider: 'opencode', + message: 'OpenCode auto-compacted context. Continue the conversation normally.', + resetContext: true, + } } async function pickFreePort(): Promise { @@ -191,24 +219,21 @@ export function prepareOpencodeQuestionRequest(question: OpencodeQuestionInfo, i const multiple = question.multiple === true const allowsCustom = question.custom !== false const header = stringValue(question.header)?.trim() + const toggles = multiple && options.length > 0 const notes = [ header ? `OpenCode asks: ${header}` : undefined, total > 1 ? `Question ${index + 1} of ${total}.` : undefined, - multiple ? 'Select all that apply. Type multiple answers separated by commas or new lines.' : undefined, + multiple ? 'Select all that apply.' : undefined, ].filter(Boolean).join(' ') - const detail = multiple && options.length > 0 - ? options.map(option => `- ${option.label}${option.description ? ` — ${option.description}` : ''}`).join('\n') - : undefined - return { request: { - kind: multiple ? 'editor' : allowsCustom || options.length === 0 ? 'prompt' : 'select', + kind: allowsCustom || options.length === 0 ? 'prompt' : 'select', title, message: notes || undefined, - detail, - options: multiple ? undefined : options, - placeholder: allowsCustom ? 'reply to OpenCode…' : undefined, + options: options.length > 0 ? options : undefined, + placeholder: allowsCustom ? (options.length > 0 ? 'or type your own answer…' : 'reply to OpenCode…') : undefined, + multiple: toggles || undefined, source: 'opencode', }, optionsById, @@ -217,6 +242,12 @@ export function prepareOpencodeQuestionRequest(question: OpencodeQuestionInfo, i } export function answerOpencodeQuestion(prepared: PreparedOpencodeQuestion, response: AgentUserResponse): string[] { + if (prepared.multiple && response.optionIds) { + const toggled = response.optionIds + .map(id => prepared.optionsById[id]?.label) + .filter((label): label is string => !!label) + return [...toggled, ...splitTypedAnswers(response.value, true)] + } const selected = response.optionId ? prepared.optionsById[response.optionId] : undefined return selected ? [selected.label] : splitTypedAnswers(response.value ?? response.optionId, prepared.multiple) } @@ -658,6 +689,14 @@ export async function createOpencodeBridge( case 'session.idle': void endTurn() return + case 'session.compacted': { + const compacted = opencodeCompactionEvent(ev.properties, sessionId, opts.bridgeId, currentTurnId) + if (compacted) { + lastUsage = undefined + emit(compacted) + } + return + } case 'question.asked': void handleQuestionAsked(ev.properties as OpencodeQuestionRequest) return diff --git a/src/main/brain-desktop-service.test.ts b/src/main/brain-desktop-service.test.ts index 13cedb1..c251d7a 100644 --- a/src/main/brain-desktop-service.test.ts +++ b/src/main/brain-desktop-service.test.ts @@ -9,7 +9,7 @@ import { removeBrainDesktopConnection, writeBrainDesktopConnection, } from './brain-desktop-rendezvous' -import { BrainDesktopService, seedBrainRuntime } from './brain-desktop-service' +import { BrainDesktopService, mergeDesktopWorkspacesIntoBrainRuntime, seedBrainRuntime } from './brain-desktop-service' import { createMachineIdentity, machineCredentialPath, writeMachineCredential } from './hub-machine-enrollment' function fixture(): { root: string; desktop: string; brain: string } { @@ -105,6 +105,29 @@ describe('desktop Brain lifecycle', () => { } finally { rmSync(root, { recursive: true, force: true }) } }) + it('puts current desktop workspaces first while retaining browser-created workspaces on enable', () => { + const { root, desktop, brain } = fixture() + try { + mkdirSync(join(brain, 'runtime'), { recursive: true }) + writeFileSync(join(desktop, 'workspaces.json'), JSON.stringify({ workspaces: [ + { id: 'shared', path: '/shared', name: 'Current desktop name' }, + { id: 'desktop-new', path: '/desktop-new', name: 'Recent desktop workspace' }, + ] })) + writeFileSync(join(brain, 'runtime', 'workspaces.json'), JSON.stringify({ workspaces: [ + { id: 'shared', path: '/shared', name: 'Old Brain name' }, + { id: 'web-new', path: '/web-new', name: 'Browser workspace' }, + ] })) + + mergeDesktopWorkspacesIntoBrainRuntime(desktop, brain) + + expect(JSON.parse(readFileSync(join(brain, 'runtime', 'workspaces.json'), 'utf8')).workspaces).toEqual([ + { id: 'shared', path: '/shared', name: 'Current desktop name' }, + { id: 'desktop-new', path: '/desktop-new', name: 'Recent desktop workspace' }, + { id: 'web-new', path: '/web-new', name: 'Browser workspace' }, + ]) + } finally { rmSync(root, { recursive: true, force: true }) } + }) + it('starts, probes, and explicitly stops the detached Brain', async () => { const { root, desktop, brain } = fixture() const descriptor = connection() diff --git a/src/main/brain-desktop-service.ts b/src/main/brain-desktop-service.ts index 3d1021e..89af5bc 100644 --- a/src/main/brain-desktop-service.ts +++ b/src/main/brain-desktop-service.ts @@ -1,6 +1,6 @@ import { spawn, type ChildProcess } from 'child_process' import { createHash } from 'crypto' -import { copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, statSync, writeFileSync } from 'fs' +import { copyFileSync, cpSync, existsSync, mkdirSync, readFileSync, readdirSync, renameSync, statSync, writeFileSync } from 'fs' import { join } from 'path' import WebSocket from 'ws' import type { BrainDesktopConnection, BrainDesktopStatus } from '../shared/brain-desktop-types' @@ -125,6 +125,48 @@ export function seedBrainRuntime(desktopDataDir: string, brainDataDir: string): seedBrainConversationAliases(desktopDataDir, runtimeDir) } +interface StoredWorkspaceRow { + id: string + path: string + [key: string]: unknown +} + +function readWorkspaceRows(path: string): StoredWorkspaceRow[] { + if (!existsSync(path)) return [] + try { + const parsed = JSON.parse(readFileSync(path, 'utf8')) as { workspaces?: unknown } + if (!Array.isArray(parsed.workspaces)) return [] + return parsed.workspaces.filter((row): row is StoredWorkspaceRow => { + if (!row || typeof row !== 'object') return false + const candidate = row as { id?: unknown; path?: unknown } + return typeof candidate.id === 'string' && candidate.id.length > 0 + && typeof candidate.path === 'string' && candidate.path.length > 0 + }) + } catch { + return [] + } +} + +/** On explicit enable, prefer the current desktop registry while retaining + * workspaces that were genuinely added through a prior browser session. */ +export function mergeDesktopWorkspacesIntoBrainRuntime(desktopDataDir: string, brainDataDir: string): void { + const desktopRows = readWorkspaceRows(join(desktopDataDir, 'workspaces.json')) + if (desktopRows.length === 0) return + const runtimeDir = join(brainDataDir, 'runtime') + const brainPath = join(runtimeDir, 'workspaces.json') + const brainRows = readWorkspaceRows(brainPath) + const desktopIds = new Set(desktopRows.map(row => row.id)) + const desktopPaths = new Set(desktopRows.map(row => row.path)) + const merged = [ + ...desktopRows, + ...brainRows.filter(row => !desktopIds.has(row.id) && !desktopPaths.has(row.path)), + ] + mkdirSync(runtimeDir, { recursive: true, mode: 0o700 }) + const temporary = `${brainPath}.${process.pid}.${Date.now().toString(36)}.tmp` + writeFileSync(temporary, `${JSON.stringify({ workspaces: merged }, null, 2)}\n`, { encoding: 'utf8', mode: 0o600, flag: 'wx' }) + renameSync(temporary, brainPath) +} + export class BrainDesktopService { private readonly dataDir: string private readonly connectionPath: string @@ -200,7 +242,10 @@ export class BrainDesktopService { async setEnabled(enabled: boolean): Promise { writeBrainDesktopEnabled(this.preferencesPath, enabled) - if (enabled) await this.start() + if (enabled) { + mergeDesktopWorkspacesIntoBrainRuntime(this.options.desktopDataDir, this.dataDir) + await this.start() + } return this.status() } diff --git a/src/preload/index.ts b/src/preload/index.ts index 1b56827..a7ce42d 100644 --- a/src/preload/index.ts +++ b/src/preload/index.ts @@ -208,6 +208,7 @@ contextBridge.exposeInMainWorld('electronAPI', { action: 'accept' | 'accept_for_turn' | 'decline' | 'submit' | 'cancel' value?: string optionId?: string + optionIds?: string[] }) => ipcRenderer.invoke('bridge:respondUserRequest', response), bridgeSetMode: (bridgeId: string, mode: 'ask' | 'plan' | 'build' | 'full'): void => diff --git a/src/renderer/public/icons/icon-192.png b/src/renderer/public/icons/icon-192.png index 93b8f09..2d0c4b0 100644 Binary files a/src/renderer/public/icons/icon-192.png and b/src/renderer/public/icons/icon-192.png differ diff --git a/src/renderer/public/icons/icon-512.png b/src/renderer/public/icons/icon-512.png index 807d687..02e8d9f 100644 Binary files a/src/renderer/public/icons/icon-512.png and b/src/renderer/public/icons/icon-512.png differ diff --git a/src/renderer/src/components/chat/SoloChatView.tsx b/src/renderer/src/components/chat/SoloChatView.tsx index a9b25ff..f6ebbe8 100644 --- a/src/renderer/src/components/chat/SoloChatView.tsx +++ b/src/renderer/src/components/chat/SoloChatView.tsx @@ -397,6 +397,7 @@ export function SoloChatView(props: SoloChatViewProps) { onRespond={onAgentRequestResponse} planGate={planGate} onApprovePlan={() => onRunCommand?.(CREWCODER_APPROVE_PLAN_PROMPT)} + onReplyPlan={onRunCommand ? text => onRunCommand(text) : undefined} /> )} diff --git a/src/renderer/src/components/crew/CrewTimeline.tsx b/src/renderer/src/components/crew/CrewTimeline.tsx index feaf33b..efc2f96 100644 --- a/src/renderer/src/components/crew/CrewTimeline.tsx +++ b/src/renderer/src/components/crew/CrewTimeline.tsx @@ -235,6 +235,7 @@ export function CrewTimeline({ onRespond={onAgentRequestResponse} planGate={planGate} onApprovePlan={() => onSendToLane(lane.laneId, CREWCODER_APPROVE_PLAN_PROMPT)} + onReplyPlan={text => onSendToLane(lane.laneId, text)} /> )} diff --git a/src/renderer/src/components/crew/LaneColumn.tsx b/src/renderer/src/components/crew/LaneColumn.tsx index 97ec157..3fe99ae 100644 --- a/src/renderer/src/components/crew/LaneColumn.tsx +++ b/src/renderer/src/components/crew/LaneColumn.tsx @@ -238,6 +238,7 @@ export function LaneColumn({ onRespond={onAgentRequestResponse} planGate={planGate} onApprovePlan={() => onSend(CREWCODER_APPROVE_PLAN_PROMPT)} + onReplyPlan={text => onSend(text)} /> )} diff --git a/src/renderer/src/components/crew/SupervisorSidebar.tsx b/src/renderer/src/components/crew/SupervisorSidebar.tsx index b4075e8..db2a819 100644 --- a/src/renderer/src/components/crew/SupervisorSidebar.tsx +++ b/src/renderer/src/components/crew/SupervisorSidebar.tsx @@ -230,6 +230,7 @@ export function SupervisorSidebar({ onRespond={onAgentRequestResponse} planGate={planGate} onApprovePlan={() => onSend(CREWCODER_APPROVE_PLAN_PROMPT)} + onReplyPlan={text => onSend(text)} /> )} diff --git a/src/renderer/src/components/plugins/PluginTabHost.tsx b/src/renderer/src/components/plugins/PluginTabHost.tsx index 93f579f..acd6136 100644 --- a/src/renderer/src/components/plugins/PluginTabHost.tsx +++ b/src/renderer/src/components/plugins/PluginTabHost.tsx @@ -9,6 +9,25 @@ interface PluginTabHostProps { workspace: Workspace | null } +// Explicit token bridge: plugin iframes have a separate document and cannot +// inherit the renderer's CSS variables across the origin boundary. +const PLUGIN_THEME_TOKENS = [ + '--background', '--foreground', '--card', '--card-foreground', + '--muted', '--muted-foreground', '--primary', '--primary-foreground', + '--border', '--input', '--success', '--warning', '--destructive', + '--radius', '--radius-lg', '--font-family-sans', +] as const + +function pluginTheme() { + const style = getComputedStyle(document.body) + const tokens: Record = {} + for (const name of PLUGIN_THEME_TOKENS) tokens[name] = style.getPropertyValue(name).trim() + return { + tokens, + mode: document.body.classList.contains('dark') || document.body.dataset.theme === 'dark' ? 'dark' : 'light', + } +} + type PluginFrameMessage = | { type: 'crewcode:request' @@ -112,8 +131,24 @@ export function PluginTabHost({ tab, workspace }: PluginTabHostProps) { permissions: resolved.permissions, openContext: tab.pluginOpenContext ?? { source: 'restored-tab' }, }, '*') + iframeRef.current?.contentWindow?.postMessage({ type: 'crewcode:theme', ...pluginTheme() }, '*') } + useEffect(() => { + if (!resolved?.ok || typeof MutationObserver === 'undefined') return + let frame = 0 + const postTheme = () => { + cancelAnimationFrame(frame) + frame = requestAnimationFrame(() => { + iframeRef.current?.contentWindow?.postMessage({ type: 'crewcode:theme', ...pluginTheme() }, '*') + }) + } + const observer = new MutationObserver(postTheme) + observer.observe(document.body, { attributes: true, attributeFilter: ['class', 'style', 'data-theme'] }) + observer.observe(document.documentElement, { attributes: true, attributeFilter: ['class', 'style', 'data-theme'] }) + return () => { observer.disconnect(); cancelAnimationFrame(frame) } + }, [resolved]) + if (error) { return (
diff --git a/src/renderer/src/components/settings/SettingsScreen.tsx b/src/renderer/src/components/settings/SettingsScreen.tsx index d472137..5689dfb 100644 --- a/src/renderer/src/components/settings/SettingsScreen.tsx +++ b/src/renderer/src/components/settings/SettingsScreen.tsx @@ -45,6 +45,7 @@ import { type LocalVoiceDevice, } from '../../../../shared/voice-types' import { getCrewCodeClient, getCrewCodeRuntime } from '../../runtime/crewcode-client' +import { seedCurrentDesktopCatalogue } from '../../runtime/continuity-state' import { BrainAuthorizationSection } from './BrainAuthorizationSection' import { useAppBuildInfo } from '../../hooks/useAppBuildInfo' import type { EditorThemeId } from '../../../../shared/editor-theme-types' @@ -516,7 +517,21 @@ function BrainContinuitySection() { try { const next = await window.electronAPI!.brainDesktopSetEnabled(true) setStatus(next) - if (next.attached) window.location.reload() + if (next.attached) { + const seedCatalogue = window.electronAPI?.continuityDesktopSeed + if (!seedCatalogue) throw new Error('Desktop catalogue handoff is unavailable; Brain was not attached.') + try { + await seedCurrentDesktopCatalogue(seedCatalogue) + } catch (cause) { + let stopped = false + try { + setStatus(await window.electronAPI!.brainDesktopStop()) + stopped = true + } catch { /* preserve the handoff failure below */ } + throw new Error(`Could not hand off the current desktop chats and tabs: ${(cause as Error).message}.${stopped ? ' Background Brain was stopped; your desktop state is unchanged.' : ' Background Brain could not be stopped; stop it before retrying.'}`) + } + window.location.reload() + } } catch (cause) { setError((cause as Error).message) } finally { setBusy(false) } } diff --git a/src/renderer/src/components/thread/AgentActivityOverlay.tsx b/src/renderer/src/components/thread/AgentActivityOverlay.tsx index c51f0de..b0f5bec 100644 --- a/src/renderer/src/components/thread/AgentActivityOverlay.tsx +++ b/src/renderer/src/components/thread/AgentActivityOverlay.tsx @@ -30,6 +30,8 @@ interface AgentActivityOverlayProps { planGate?: CrewCoderPlanGate | null /** Sends `/approve-plan` as a prompt or follow-up. */ onApprovePlan?: () => void + /** Sends a clarify answer as an ordinary user prompt (never an approval). */ + onReplyPlan?: (text: string) => void /** Optional dismissal for completed inline activity. */ onDismiss?: () => void } @@ -48,7 +50,7 @@ export function shouldShowAgentActivity( * Converted from the Tailwind design in `.design/AgentActivityOverlay.tsx`. * This inline variant keeps chat history stable while still showing live task progress. */ -export function AgentActivityOverlay({ todos, isStreaming, request, onRespond, planGate, onApprovePlan, onDismiss }: AgentActivityOverlayProps) { +export function AgentActivityOverlay({ todos, isStreaming, request, onRespond, planGate, onApprovePlan, onReplyPlan, onDismiss }: AgentActivityOverlayProps) { const { state: settings } = useSettings() const [isExpanded, setIsExpanded] = useState(isStreaming) const [dismissed, setDismissed] = useState(false) @@ -92,6 +94,7 @@ export function AgentActivityOverlay({ todos, isStreaming, request, onRespond, p return (
{ @@ -99,6 +102,11 @@ export function AgentActivityOverlay({ todos, isStreaming, request, onRespond, p setPlanSent(true) onApprovePlan() }} + onReply={onReplyPlan ? (text) => { + if (planSent) return + setPlanSent(true) + onReplyPlan(text) + } : undefined} />
) @@ -250,10 +258,18 @@ interface CrewCoderPlanGateCardProps { gate: CrewCoderPlanGate sent: boolean onApprove: () => void + /** Present when the host surface can send a prompt for this chat/lane. */ + onReply?: (text: string) => void } -function CrewCoderPlanGateCard({ gate, sent, onApprove }: CrewCoderPlanGateCardProps) { +function CrewCoderPlanGateCard({ gate, sent, onApprove, onReply }: CrewCoderPlanGateCardProps) { const awaitingApproval = gate.phase === 'awaiting_approval' + const [reply, setReply] = useState('') + const trimmed = reply.trim() + const sendReply = (): void => { + if (!onReply || sent || !trimmed) return + onReply(trimmed) + } return (
@@ -277,7 +293,32 @@ function CrewCoderPlanGateCard({ gate, sent, onApprove }: CrewCoderPlanGateCardP
  • {question}
  • ))} -
    Reply in the composer. Answering a question is not plan approval.
    + {onReply ? ( +
    +