Skip to content

fix: stop rewriting Responses history by default - #6

Merged
voidsteed merged 1 commit into
masterfrom
fix/responses-context-trimming
Sep 17, 2026
Merged

voidsteed merged 1 commit into
masterfrom
fix/responses-context-trimming

Conversation

@voidsteed

Copy link
Copy Markdown
Owner

Fixes #5.

Every /v1/responses request was trimmed pre-flight against a byte ceiling derived from max_prompt_tokens, rewriting old input items into placeholders before Copilot ever saw the payload. The reporter is right that this is the wrong default, and measuring it turned up two separate problems.

The ratio was miscalibrated

CHARS_PER_TOKEN_ESTIMATE = 3.5 was inherited from context-manager.ts, where it measures unescaped content length. Here it measures serialized JSON — copying the constant silently changed its basis.

Measured against multi-turn payloads built from this repo's own source (user messages, function_call items, large function_call_output bodies, assistant replies):

JSON body bytes      : 424,786
content tokens       :  96,442   (o200k_base)
bytes / content token:    4.40
bytes / json token   :    3.70

So trimming engaged at roughly 77% of a model's real window, destroying context Copilot would have accepted:

max_prompt_tokens byte ceiling real tokens at ceiling % of window
272,000 924,000 ~209,800 77%
200,000 672,000 ~152,600 76%
128,000 420,000 ~95,400 74%

The rewrite broke prompt caching

dropOldInputItems replaces from the oldest item forward — precisely the stable prefix a cache is keyed on. Simulating a growing session against the real algorithm, measuring how many leading items stay byte-identical turn over turn:

turn | raw KB | sent KB | dropped | stable cacheable prefix
  70 |    834 |     834 |       0 | 277/281
  76 |    906 |     894 |       3 | 1/305   <-- collapse
  80 |    954 |     896 |      19 | 16/321

The reusable prefix collapses from 277 items to 1 on the turn trimming first engages, and is re-poisoned each time the boundary advances.

What changed

Default is now passthrough. Client history forwards verbatim; Copilot enforces its own token limits, so a rejection is upstream's real answer instead of a local guess. Only the ~5 MB Azure Front Door cliff is enforced pre-flight — that one is a verified wire limit, not an estimate.

Token-derived trimming stays available behind --responses-context-trim, with the ratio corrected to 4.3, and warns at startup when enabled.

The reactive path is repaired, since passthrough now depends on it being trustworthy:

  • Context-overflow errors returned a local bytes/4 estimate instead of upstream's reported counts. A genuine 300000 > 272000 was being rewritten as 16 + 1000 > 272000 — and Claude Code sizes its compaction pass off that gap. Upstream's numbers are now used when present, with the original body preserved in error.upstream_error.
  • isContextOverflow treated any 5xx on a >2 MB payload as context overflow, so a transient 503 made clients compact away history to work around what was really a Copilot outage. Now restricted to explicit overflow markers and plain-500 dead-zone hangs, with transient-outage phrasings excluded.

Notes

dropOldInputItems still drops oldest-first when it runs. I tried reversing it to preserve the cacheable prefix and reverted that — dropping from the tail discards the active request, which is worse than the cache miss it avoids. The real mitigation is that it no longer runs by default; the comment there now says so.

Four existing trimming tests needed fixture sizes scaled to still overflow the raised ceiling. They cover the same behavior but no longer exercise the default path — tests/create-responses-context.test.ts does that, with 5 new tests covering default passthrough (asserts all 8 turns survive), the transport ceiling still holding, upstream count surfacing, the 503 case, and the opt-in flag.

Caveat: neither 4.3 nor 4.4 is Copilot's actual prompt-token count — confirming that needs a live measurement against the real endpoint, which I haven't done. It matters less now that the ratio doesn't gate the default path.

bun run typecheck and bun test are green (363 tests); lint is clean on all changed files.

Ratio, cache-prefix collapse, and tokenizer cost (~39ms at 1MB, vs. the "300-1000ms" claimed in the context-manager.ts comment) were measured independently by two models reaching the same conclusions.

🤖 Generated with Claude Code

Every /v1/responses request was trimmed pre-flight against a byte ceiling
derived from max_prompt_tokens, rewriting old input items into placeholders
before Copilot ever saw the payload. Two problems with that:

The ratio was wrong. CHARS_PER_TOKEN_ESTIMATE = 3.5 came from
context-manager.ts, where it measures *unescaped content length*; here it
measures serialized JSON. Measured against multi-turn payloads built from this
repo's own source, the real figure is 4.4 bytes per content token (3.7 per
token of the JSON itself), so trimming engaged at roughly 77% of a model's
real window — destroying context Copilot would have accepted.

The rewrite broke prompt caching. dropOldInputItems replaces from the oldest
item forward, which is precisely the stable prefix a cache is keyed on. In a
simulated growing session the reusable prefix collapsed from 277 items to 1 on
the turn trimming first engaged, and was re-poisoned each time the boundary
advanced.

Default behavior is now to forward the client's history verbatim and let
Copilot enforce its own limits, so a rejection is upstream's real answer
instead of a local guess. The ~5 MB Azure Front Door transport cliff is still
enforced pre-flight — that one is a verified wire limit, not an estimate.
Token-derived trimming remains available behind --responses-context-trim, with
the ratio corrected to 4.3.

Also repairs the reactive path it now depends on:

- Context-overflow errors returned upstream's reported token counts instead of
  a local bytes/4 estimate. A genuine "300000 > 272000" was being rewritten as
  "16 + 1000 > 272000", and Claude Code sizes its compaction pass off that gap.
  The original upstream body is preserved in error.upstream_error.
- isContextOverflow treated any 5xx on a >2 MB payload as context overflow, so
  a transient 503 made clients compact away history to work around what was
  really a Copilot outage. Now restricted to explicit overflow markers and
  plain-500 dead-zone hangs, with transient-outage phrasings excluded.

Ratio, cache-prefix collapse, and tokenizer cost were measured independently by
two models reaching the same conclusions.

Fixes #5

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@voidsteed
voidsteed force-pushed the fix/responses-context-trimming branch from b828940 to 214704f Compare September 17, 2026 05:18
@voidsteed
voidsteed merged commit 868b5f3 into master Sep 17, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Responses API silently rewrites history and breaks prompt-cache continuity

1 participant