Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file modified build/icon.ico
Binary file not shown.
Binary file modified build/icon.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified build/icons/256x256.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified build/icons/512x512.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
30 changes: 29 additions & 1 deletion docs/agent-activity-overlay.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,35 @@ This projection changes only CrewCode visualization. Tool availability and Claud

Human-input request cards are independent of this preference. Approvals, questions, editor requests, and notifications must always render because the provider may be paused waiting for the response. Request rendering therefore takes precedence over both the Todo preference and a previously dismissed todo card.

CrewCoder-mode `crewcoder_clarify` / `crewcoder_propose_plan` is a session workflow gate, not a tool-permission pause. After those tools settle, the overlay shows a dedicated clarification or **Approve plan** card. Approve sends `/approve-plan` as a normal prompt (or follow-up if the turn is still running). It must never reuse Allow/Deny on a permission card — `/approve` is still only for pending tool-call grants. A later user message hides the card: `/approve-plan` or a short CrewCoder approval continues implementation, and a revision such as `yes, but also add logging` waits for the next `crewcoder_propose_plan`. Answering a clarification is not plan approval. The Todo preference must not hide this card.
CrewCoder-mode `crewcoder_clarify` / `crewcoder_propose_plan` is a session workflow gate, not a tool-permission pause. After those tools settle, the overlay shows a dedicated clarification or **Approve plan** card. Approve sends `/approve-plan` as a normal prompt (or follow-up if the turn is still running). It must never reuse Allow/Deny on a permission card — `/approve` is still only for pending tool-call grants. A later user message hides the card: `/approve-plan` or a short CrewCoder approval continues implementation, and a revision such as `yes, but also add logging` waits for the next `crewcoder_propose_plan`. Answering a clarification is not plan approval. The Todo preference must not hide this card. The clarification card has its own reply box and **send reply** button; the reply is sent as an ordinary prompt (follow-up while running) through the same surface's send path as Approve. Surfaces that cannot send a prompt fall back to "Reply in the composer".

## Provider question cards

Structured provider questions pause the turn and render in the shared `AgentRequestCard` (inline chat overlay, Crew lanes/timeline, supervisor, Mission Control, menulet). The card layout follows the question shape:

| Shape | Card |
| --- | --- |
| options only (`kind: 'select'`) | one button per option; a click answers immediately. 2–4 short description-free options (Yes/No) sit on one row |
| options + free text (`kind: 'prompt'` with options) | option buttons plus an "or type your own answer…" input and **send reply** |
| free text only (`prompt` / `editor`) | input (Enter) or textarea (Ctrl/Cmd+Enter) plus **send reply** |
| multi-select (`multiple: true`) | toggle buttons answered with `optionIds` in option order, plus optional typed text |
| secret (`secret: true`) | masked input |

**send reply** stays disabled until there is a non-empty answer. After a click the card locks (`sending…`) and unlocks only on an observed transport failure; success is confirmed by `user_request_resolved`. Cancel is an explicit cancel, never an empty answer.

Provider mapping — only structured, provider-issued questions reach this card:

| Provider | Native question | Free text |
| --- | --- | --- |
| Claude | `AskUserQuestion` via `canUseTool`, one card per question | always, matching Claude's built-in "Other"; `custom: false` / `allowFreeform: false` opts out |
| Codex | app-server `item/tool/requestUserInput` (legacy `tool/requestUserInput` alias), one card per question | when the question sets `isOther` or has no options |
| OpenCode | `question` SSE request | unless `custom: false` |
| Pi | extension UI `select` / `input` / `editor` | per method |
| CrewCoder | `crewcoder_clarify` workflow card (above) | reply box |

Codex answers are returned as `{ answers: { [questionId]: { answers: string[] } } }`. A cancelled, empty, or out-of-contract answer (typed text where Codex offered no "Other") fails the whole request with an explicit JSON-RPC error; Codex never receives a fabricated or empty answer. When Codex settles a request itself (`serverRequest/resolved`, e.g. non-blocking questions) or the app-server exits, the bridge aborts the pending `RequestUserFn` signal: the promise settles as `cancel` and the card is retracted with `user_request_resolved`.

Plain-text questions at the end of an agent reply ("Should I continue?") are not requests: the turn has already ended and nothing is paused, so they are answered in the composer. CrewCode must not guess questions from assistant prose or fabricate a request card for them. Codex MCP `mcpServer/elicitation/request` forms are not yet mapped.

## Surfaces

Expand Down
105 changes: 105 additions & 0 deletions docs/context-usage-meter.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
# Context usage meter

The chat meter divides the latest observed context occupancy by the active
context window. The numerator is the provider's current prompt/context size,
not cumulative billing tokens for the entire conversation.

Clicking the context pill opens a compact usage summary. When the provider
supplies token details, **Token logs** opens a scrollable right sidebar. The
sidebar can be closed with its close button, the backdrop, or Escape. Claude
provides context categories; Codex, OpenCode, and Grok provide request or turn
token counters, while CrewCoder provides session counters and its latest context
input reading. These rows can overlap and must not be added together.
If a provider reports token counts before a context window is known, the usage
strip shows a direct **Token logs** button and the sidebar marks context usage
as unavailable instead of showing a fabricated percentage.

Codex token logs show the latest app-server request's input and output, plus
cached input and reasoning output when reported. OpenCode logs show its latest
assistant message's input, output, reasoning, cache read, and cache write
counts. Grok logs show latest-call input/output and, when present, cache read,
reasoning, and cumulative turn input/output. These providers do not expose
Claude-style system/tool/message context categories through these usage events;
the sidebar labels their rows as provider token counts rather than claiming
they sum to context occupancy.

CrewCoder token logs use the ACP `session/prompt` result's namespaced
`crewcoder/usage` summary, falling back to its top-level `usage` mirror. They
show the latest context input (`lastInputTokens`) separately from cumulative
session input, output, total, cached input, cache write, and reasoning counts
when those fields are reported. The session counters can span multiple models
and requests and do not sum to the current context reading. If CrewCoder omits
`lastInputTokens`, cumulative session totals remain visible in Token logs but
do not become a context occupancy estimate.

## Window sources

Use these sources in order:

1. A full window reported for the active session or turn by the provider bridge.
2. A full window from the provider's model catalog, when it supplies one.
3. A static model-family fallback, when no provider value is available.

Codex app-server usage notifications include `modelContextWindow`. For known
models, this can be smaller than the model's full window because it is the
effective request prompt budget. CrewCode shows the known full model capacity
as the meter denominator and labels the reported smaller value separately as
**Codex prompt budget**. For unknown models, the reported value remains the
only available window and drives the meter. A reported value larger than model
metadata also takes precedence. The full model window can come from provider
catalog metadata or CrewCode's static model fallback; the latter can be stale
if a custom Codex configuration changes capacity. Compaction-drop inference
uses Codex's reported prompt budget as its threshold denominator even when the
displayed model window is larger.
The latest request's input tokens are also treated as an authoritative context
reading, so a decrease after native compaction is not replaced by an old floor.
The app-server `model/list` response discovers models but does not document a
context-window field, so model discovery alone cannot supply the Codex meter.

Claude's SDK `getContextUsage()` supplies the active `maxTokens` (or
`rawMaxTokens`) window. The meter sums its active categories and excludes
deferred categories, unused space, and compaction headroom using each SDK row's
`kind`, with legacy category handling for older Claude binaries. CrewCode
requests `detail: 'full'`, since `summary` uses local estimates rather than
the token-counted `/context` breakdown. The categorized SDK reading is the
source of occupancy; result-message billing usage is separate and aggregates
multiple requests. The static model table can
still supply a window but cannot supply occupancy. Claude's `--help` does not
provide a model-window catalog.

CrewCode asks `getContextUsage()` after each assembled assistant message while
the SDK query is active. It retries at the result only if no valid reading has
arrived, because a result-time call can race query shutdown and a full reading
can involve token-counting API requests. If no control reading succeeds for that turn, the latest
individual assistant message's input and cache token counts provide a bounded
request-context fallback. Only when neither current-turn source exists does
CrewCode reuse an earlier measured reading; with no reading at all, the meter
stays unknown. Result-message usage remains billing-only because it aggregates
multiple requests in one turn.

CrewCoder ACP usage can also include a live `contextWindow` and
`lastInputTokens`; its bridge passes those through to the meter. Providers
that report no usable window can show no percentage if no catalog or static
fallback exists. A fallback percentage is an estimate, not a provider reading.

The latest usage snapshot is persisted for resumed chats. A new provider
report replaces its window; occupancy normalization and compaction handling
are described in [Conversation storage](conversation-storage.md).

## Automatic compaction

Claude's `compact_boundary`, Codex app-server's `contextCompaction` completion,
CrewCoder's `_crewcoder/compaction_update`, and OpenCode's matching
`session.compacted` SSE event are native completion signals. They clear the
previous context reading in the main process, its persisted snapshot, and the
latest visible usage strip. The meter remains unknown until a new provider
usage reading arrives; a pre-compaction reading must not be replayed at turn
end. Historical provider token counters remain in Token logs, while old
Claude context categories are cleared. OpenCode's event is accepted only for
the active session.

When no native boundary is observed, a high-occupancy to large-drop usage
reading can detect compaction for Codex, CrewCoder, OpenCode, and Grok. The
smaller observed reading becomes the new context baseline. Providers without
an observed boundary or trustworthy absolute context reading are not inferred
from silence or a completed turn.
11 changes: 11 additions & 0 deletions docs/conversation-storage.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,17 @@ and appends a visible compact-summary card, but keeps the full rich display
transcript, provider session id, and live bridge. If CrewCoder reports a skipped
small-session compact, CrewCode leaves replay and usage state unchanged.

Provider-native automatic compaction is also observable. Direct Codex maps its
app-server `contextCompaction` item lifecycle (plus the legacy
`thread/compacted` completion) to the same meter. CrewCoder forwards nested
Codex and Claude native boundaries through `_crewcoder/compaction_update`.
Claude maps `compact_boundary`, and OpenCode maps the selected session's
`session.compacted` SSE event. A native completion clears the previous live
occupancy, its persisted snapshot, and the latest visible usage strip until
new usage arrives. Native boundaries win; live-occupancy drop inference is only
a fallback for Codex, CrewCoder, OpenCode, or Grok turns where no native
boundary was observed, so a single compaction never creates duplicate cards.

### `summary-reset` flow (pi / hermes / older CrewCoder)

Native-session providers keep their context server-side and expose no compaction RPC, so we cannot shrink it directly. Instead:
Expand Down
41 changes: 35 additions & 6 deletions docs/crewcoder-provider.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,9 +110,16 @@ bridge, or seed a replacement session after a successful native compact.

CrewCoder's compaction update is an additive namespaced ACP extension carrying
started/completed/failed status, automatic intent, progress, and a human-readable
message. The bridge treats it as authoritative and does not also infer
compaction from the later context-token drop. Automatic updates omit the summary
body; the compacted summary remains only in CrewCoder's durable session.
message. It covers both CrewCoder's durable-session compaction and compaction
reported by a nested native provider. Codex app-server `contextCompaction` item
lifecycle and `thread/compacted` notifications therefore drive the existing
CrewCode loading bar through CrewCoder; Claude `compact_boundary` drives the
completed state. The bridge treats a native update as authoritative and does not
also infer compaction from the later context-token drop. If a CrewCoder provider
does not expose a native boundary, a verified high-water-to-large-drop occupancy
change produces one after-the-fact detected notification instead of staying
silent. Automatic updates omit the summary body; the compacted summary remains
only in CrewCoder's durable session.
Host-requested `session/compact` returns the authoritative summary and includes
it on the completed update. CrewCode replaces only its provider replay shard
with that summary and appends the visible compact-summary card; it deliberately
Expand Down Expand Up @@ -141,6 +148,17 @@ bridge registration; the next composer submission uses normal missing-bridge rec
attempting to write to closed stdin and surfacing `crewcoder acp: process not writable`. When automatic compaction is off, the user explicitly
runs `/compact` before continuing; this policy does not affect Pi or other providers.

For CrewCoder's built-in Codex provider, the nested app-server receives the
same resolved context policy as the outer loop. Sol, Terra, and Luna GPT-5.6
sessions therefore use a 1.05M context and a 630k normal auto-compaction limit,
instead of app-server independently compacting near its smaller default. These
overrides are process-scoped and do not modify the user's Codex configuration.
Every CrewCoder built-in model declares a context window. Claude SDK-native
auto-compaction is disabled so CrewCoder remains the only owner of its durable
compaction boundary. ACP providers such as Grok can report the active window at
runtime; that value supersedes static metadata and recalculates CrewCoder's
percentage threshold for subsequent checks.

A prompt has a ten-minute **inactivity** watchdog rather than a wall-clock turn
limit. Every matching ACP update or agent request resets it, and time awaiting a
Build permission decision is excluded. If CrewCoder becomes genuinely silent,
Expand All @@ -149,9 +167,20 @@ settle before emitting the timeout and `turn_end`. If cancellation itself remain
unresponsive, CrewCode terminates that bridge so its replacement starts cleanly;
a second prompt can never overlap the abandoned CrewCoder turn.

Usage prefers `_meta["crewcoder/usage"]`: `lastInputTokens` is the live
`contextTokens` value and `contextWindow` is the context limit. Top-level usage
is only the compatibility fallback.
Usage prefers `_meta["crewcoder/usage"]`: `lastInputTokens` is an authoritative
live `contextTokens` measurement and `contextWindow` is the registered full model
limit (1.05M for the Codex 5.6 family), not Codex app-server's smaller effective
per-request prompt budget. CrewCode accepts measured drops after provider-native
compaction instead of applying its generic monotonic resume floor. The top-level
mirror is the compatibility fallback and preserves the same live fields. Usage is scoped to
one ACP prompt; an errored or metadata-free prompt never inherits the preceding
turn's snapshot.
The chat's Token logs sidebar shows `lastInputTokens` as the latest context
input and labels input, output, total, cache, and reasoning counters from the
same ACP summary as cumulative session counts. These counters can overlap and
must not be added together as context occupancy. If `lastInputTokens` is absent,
the cumulative counters remain available in Token logs without a context
percentage.

ACP `tool_call` updates carry a category `kind` (`read`, `edit`, `think`, …), a
human `title`, and authoritative CrewCoder tool identity in
Expand Down
13 changes: 11 additions & 2 deletions docs/desktop-web-continuity.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,8 +69,8 @@ a rejected Promise where they require a disposer function.
## Source of truth

On first enable, missing Brain runtime data is seeded from Electron `userData` into
`~/.crewcode/brain/runtime`. Existing Brain files always win for workspaces, keys, and
replay; they are never overwritten. Transcript shards are merged instead: a newer
`~/.crewcode/brain/runtime`. Existing Brain files win for keys and replay; they are
never overwritten. Transcript shards are merged instead: a newer
desktop copy is folded into the Brain shard by message identity so work done locally
before attachment is not stuck on the first seed snapshot. The seed includes
registered workspaces, provider-native resume IDs, provider keys, rich transcripts,
Expand All @@ -79,6 +79,15 @@ gets a non-destructive `web:<session>` alias so the first Brain-backed prompt ca
continue the desktop conversation. Provider-native resume IDs remain keyed by both
session and provider.

Every explicit enable also reconciles the current desktop workspace registry into the
Brain before attachment: current desktop entries and ordering win for matching ids or
paths, while genuine browser-created workspaces are retained. Immediately after the
Brain attaches, but before Electron reloads, the foreground renderer sends its exact
current chat/session and workspace-tab catalogue through the owner-loopback desktop
control. This enable-time handoff supersedes an older persisted authority marker, so a
previous Brain run cannot replace the chats, tabs, workspace, and active selections the
owner was using when they selected **Enable**.

After attachment, the Brain store is authoritative for:

- registered, Brain-authorized workspaces;
Expand Down
12 changes: 10 additions & 2 deletions docs/execution-modes.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
CrewCode's composer mode (`ModeLevel`) is a per-session gate with two halves:

1. **An optional prompt preamble** injected once at session start
(`buildModePreamble` in `src/renderer/src/hooks/chat-session-send.ts`).
(`buildModePreamble` in `src/renderer/src/hooks/chat-session-send.ts`), and
once more on each mid-session mode switch (`buildModeSwitchPreamble`).
2. **A real permission policy** applied per provider inside each bridge.

The preamble is advisory; the bridge policy is the enforcement. Users edit mode
Expand Down Expand Up @@ -52,7 +53,14 @@ to true for new and legacy sessions, and is copied when a session is duplicated.
before the first send. The toggle locks after visible history or the delivery
marker proves startup context was committed; provider context cannot be revoked.
- Prompt edits apply only when a session has not sent its startup context.
Existing/restored sessions must not receive the prompt again.
Existing/restored sessions must not receive the startup prompt again.
- Switching mode mid-session sends the new mode's prompt exactly once, on the
next send, behind a short `<system>` notice that the previous mode's
instructions no longer apply. Without it the earlier preamble stays in the
provider transcript as the only mode contract, so an Ask to Build switch
could leave the agent refusing to implement. It never repeats on later turns
in the same mode, and a restored session whose delivery marker is missing is
seeded silently rather than treated as a switch.
- Disabled mode prompts leave provider-native/default system context in place.
Skills, attachments, handoff packets, and delegation context still use their
normal send paths.
Expand Down
Loading
Loading