Skip to content

Mark prompt-cache breakpoints for Claude when the client marked none - #306

Merged
fylorn merged 1 commit into
mainfrom
feat/auto-cache-breakpoints
Oct 9, 2026
Merged

fylorn merged 1 commit into
mainfrom
feat/auto-cache-breakpoints

Conversation

@fylorn

@fylorn fylorn commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #305. Converted Codex requests to Claude carried no cache breakpoints, so Anthropic never cached them.

Why

Anthropic caches only what a request marks with cache_control, and Bedrock only what it marks with cachePoint. Clients in OpenAI or Gemini formats can't mark anything: Codex, Chat clients and Gemini clients rely on those providers caching a repeated prefix automatically. The IR only carried breakpoints that an Anthropic or Bedrock client had set, so a request converted to Claude had none. Every turn of an agent loop paid full input price for the whole conversation.

What changes

When the IR has no breakpoints, i.e. the client marked nothing:

Target Rule
Anthropic format (any Anthropic-format upstream; compatible endpoints accept cache_control, because Claude Code sends it to them) cache_control: {"type": "ephemeral"} on the last tool, the last system block, and the last block of the last two user messages
Bedrock, Claude model ids (anthropic.claude-…, us.anthropic.claude-…) {"cachePoint": {"type": "default"}} after the same four places
Bedrock, other models (Nova, Llama, application inference profile ARNs) nothing. Models that don't accept cachePoint reject the whole request, and an ARN doesn't say which model is behind it
Chat, Gemini, Responses nothing (the formats have no breakpoints)
  • At most four marks, Anthropic's limit, and Converse's as well. Default 5-minute TTL, never 1 hour.
  • Why the last two user turns: the earlier of the two is exactly where the previous request put its last mark. Each turn reads back what the turn before wrote, even when a step adds more blocks than Anthropic's ~20-block lookback (parallel tool calls). This is Anthropic's documented multi-turn pattern. Marks aren't part of the cached content, so the one the previous turn had further back can move.
  • Placement is on the encoded wire format, after merging consecutive same-role turns and putting tool results first, so it lands on what is actually sent. It is deterministic. Thinking blocks are never marked.
  • No minimum-length estimation: the API ignores a mark below a model's minimum cacheable length without failing.
  • Untouched: requests that carry their own breakpoints (Claude Code, other Anthropic-native clients, Bedrock clients with cachePoint) keep exactly those. Same-format passthrough stays byte-identical, since this lives only in the encoders.
  • Always on, no config key (owner's choice). Documented in docs/config.md and docs/config.zh-CN.md, in the providers section next to the description of what a forwarded request carries.

Cost accounting

No code change was needed; this verifies it end to end:

  • Usage parsing already reads Anthropic's cache_creation_input_tokens / cache_read_input_tokens (including cache_creation.ephemeral_1h_input_tokens) and Bedrock's cacheWriteInputTokens / cacheReadInputTokens.
  • The usage sniffer reads the upstream's own bytes, so converted requests are measured in the upstream's format.
  • The recorder prices with cost_of: input, output, cache_read, and cache_write at the 5-minute or 1-hour write rate of the price sheet resolved for that upstream and model.
  • New gateway e2e test: a Codex request to a mock Claude that reports cache_creation_input_tokens: 2000 and cache_read_input_tokens: 9000 produces RequestFinished.usage = {input: 30, cache_write: 2000, cache_read: 9000, output: 12}.
  • New recorder test: cache writes and reads are charged at the cache prices. Sonnet 4.5: 1000 in, 500 out, 100k read, 10k written gives $0.078, and the cache saving is recorded as reads saved minus the write premium.

Tests

  • tw-dialect/tests/auto_cache.rs (new), built on Codex-shaped Responses Lite requests:
    • The four expected marks: last tool, last system block, previous user turn, last user turn.
    • Three consecutive turns: every request has a mark on the block where the previous request put its last one, and everything up to that block is byte-identical once marks are removed. Tools and system are identical including marks.
    • A Claude Code request converted to Anthropic or Bedrock keeps exactly its 2 marks, including its 1-hour TTL.
    • A one-message Chat request gets one mark; Chat, Gemini and Responses targets get none.
    • Bedrock Claude gets the same four cachePoints; Nova, Llama and an ARN get none.
  • tw-gateway/tests/conversion.rs: the cache-usage e2e test above, plus an Anthropic request without marks passed straight through byte for byte.
  • tw-store recorder: cache pricing.
  • Adjusted: the developer-message prefix test from Codex compaction on converted routes; mid-conversation developer messages stay in place #305 compares messages without marks, because marks move each turn. The DeepSeek Harness Chat-to-Anthropic test now expects the marks.

Checks: cargo fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings, cargo test --workspace --no-fail-fast (3147 passed, 0 failed, 8 ignored), scripts/smoke.sh (79 passed, 0 failed).

🤖 Generated with Claude Code

Anthropic caches only what a request marks with `cache_control`, and Bedrock
only what it marks with `cachePoint`. Clients in OpenAI or Gemini formats
(Codex, Chat clients) cannot mark anything, because those providers cache a
repeated prefix on their own. So a converted request to Claude was never
cached, and every turn paid full price for the whole conversation.

When the IR has no breakpoints, the Anthropic encoder (any Anthropic-format
upstream) and the Bedrock encoder (Claude model ids only) now mark, with the
default 5-minute TTL:

- the end of the tools
- the end of the system prompt
- the end of the last user turn and of the one before it, at most four in
  total

The earlier user mark is the previous request's last breakpoint, so each
turn reads back what the previous one wrote even when a step adds more
blocks than Anthropic's ~20-block lookback. Thinking blocks are never
marked. Requests that carry their own breakpoints (Claude Code) keep exactly
those, and same-format passthrough is untouched.

Cost accounting needed no change: usage parsing already reads
`cache_creation_input_tokens` / `cache_read_input_tokens` and Bedrock's
`cacheWriteInputTokens` / `cacheReadInputTokens`, the sniffer reads the
upstream's own bytes for converted requests, and the recorder charges cache
writes and reads at the price table's cache rates. New tests cover both.

Documented in docs/config.md and docs/config.zh-CN.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit 34ac244 into main Oct 9, 2026
5 checks passed
@fylorn
fylorn deleted the feat/auto-cache-breakpoints branch October 9, 2026 10:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant