Skip to content

Responses conversion: keep the tools Codex declares in input, and every item it sends - #304

Merged
fylorn merged 1 commit into
mainfrom
fix/codex-responses-items
Oct 9, 2026
Merged

fylorn merged 1 commit into
mainfrom
fix/codex-responses-items

Conversation

@fylorn

@fylorn fylorn commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

A Codex user routed through the gateway to a non-Responses upstream (Anthropic / Chat / Gemini / Bedrock) got "the workspace read/write and command execution tools are not mounted"; the traffic record listed input.additional_tools, text.verbosity and more as dropped.

Root cause

With Responses Lite (model_info.use_responses_lite), Codex sends no top-level tools. It puts every tool into an input item {"type": "additional_tools", "role": "developer", "tools": [...]} (client.rs#L972-L987, models.rs#L1016-L1023), built by create_tools_json_for_responses_lite, which groups functions and freeform tools into the functions namespace. responses::request::decode_request read only top-level tools, and unknown input items fell into other => dropped.path("input.{other}") — so the converted request reached the upstream with no tools at all.

What changes

  • additional_tools entries go through the same tool decoding as top-level tools, namespaces included. Order is deterministic: top-level tools first, then additional_tools / tool_search_output in input order; a later declaration with the same name replaces the earlier one in place — Codex's incremental tool updates say exactly that (top_level_tools.rs#L266), and keeping positions keeps the tool list (the start of the prompt cache) stable.
  • Names: tools in Codex's default functions namespace keep their plain names (exec_command, apply_patch) — Codex treats functions.x and x as the same tool (tool_name.rs#L7, #L39). Other namespaces join the way Codex does (mcp__codex_apps__calendar + _create_event → mcp__codex_apps__calendar_create_event, otherwise ns__name) and names over 64 characters (the OpenAI Chat / Gemini / Bedrock limit) are cut with a stable FNV hash. ClientShape.namespaced maps every flattened name back, so calls go back to Codex with namespace + name, as function_call or custom_tool_call.
  • Namespace descriptions (for MCP namespaces these are the server instructions, rmcp_client.rs#L854) go into the system prompt; Codex's filler "Tools in the X namespace." does not.
  • Client-executed tool_search becomes a function tool; when the upstream calls it, Codex receives a tool_search_call (execution: client, object arguments, tsc_ id, no argument deltas), which is what Codex dispatches (router.rs#L275-L294). The tools found in a tool_search_output become callable.
  • New ClientShape.tool_search, Request.verbosity / ir::Verbosity, Feature::Verbosity, responses::request::{flat_tool_name, freeform_tools, TOOL_SEARCH} — all additive (the enterprise edition builds Request with ..Default::default() and doesn't touch ClientShape or Feature).
  • tw-store's transcript reader uses the same flattened names, reads freeform tools from additional_tools (so a wrapped apply_patch from another format is unwrapped), and doesn't show the declaration item as a turn.
  • No changes in tw-gateway sources (the new gateway tests live in tests/conversion.rs).

Every ResponseItem Codex can send (models.rs#L1015-L1259)

Item Before Now
additional_tools dropped (input.additional_tools) each tool decoded like top-level tools
message (+ phase) system/developer → system prompt, others a turn same; phase has no equivalent, the text is kept, not reported
agent_message dropped user turn with its input_text (sender/task are in the text, inter_agent_message.rs); encrypted_content reported dropped; caller.rs covers the text for content filtering
reasoning thinking + carried signature unchanged
function_call / custom_tool_call (+ outputs) call / result, ns__name call / result; namespace: "functions" maps to the plain name
local_shell_call dropped → its function_call_output was an orphaned result (Anthropic rejects the request) local_shell call with action as input, paired with its output (normalize.rs#L104); without call_id reported dropped
tool_search_call dropped tool_search call (needs call_id)
tool_search_output dropped result naming the tools now available; those tools are added to the tool list
configuration_update dropped last one overrides reasoning.effort — Codex pins the request-level effort for cache reasons and records mid-session changes this way (reasoning_effort.rs); a model-specific effort is reported dropped
web_search_call, image_generation_call dropped still dropped and reported: server-side tool traces; their results are in the following assistant text
compaction, item_reference refused unchanged
context_compaction dropped refused when it carries encrypted_content, otherwise reported dropped
compaction_trigger dropped → upstream answered normally → Codex failed with "expected exactly one compaction output item" (compact_remote_v2.rs#L480) refused up front with a clear message

Top-level fields:

Field Decision
text.verbosity (and Chat verbosity) carried in the IR; written as verbosity (Chat) / text.verbosity (Responses) when the upstream model is GPT-5 or later (ir::Verbosity::understood_by), otherwise reported dropped — other models answer 400 or ignore it
service_tier reported dropped unless auto / default (no equivalent tier elsewhere)
include not reported: converted reasoning items always carry encrypted_content
store, prompt_cache_key, stream_options, reasoning.context OpenAI-side storage / transport hints, nothing reaches the model; not reported
client_metadata Codex session metadata — never forwarded to another vendor; not reported

Typical Codex Responses Lite request, dropped list

  • Before: input.additional_tools, text.verbosity, plus input.local_shell_call / input.tool_search_call / input.tool_search_output / input.configuration_update / input.agent_message when present — and no tools upstream.
  • Now (Anthropic): input.reasoning (OpenAI-encrypted reasoning, unchanged), tools.custom.format (apply_patch's lark grammar, unchanged), text.verbosity; Gemini/Bedrock also parallel_tool_calls; Chat to a GPT-5 model drops no verbosity. Every tool arrives.

Tests

  • tw-dialect/tests/codex_lite.rs (new): a Codex-shaped Lite request (built from Codex's serde types, cf. responses_lite.rs) with exec_command, write_stdin, freeform apply_patch, update_plan, an MCP namespace, web.run, tool_search, plus history with local_shell_call, a tool search and an effort change — to Anthropic, Chat, Gemini and Bedrock: every tool present (read back with each format's own decoder), history calls paired with results, exact dropped list; the upstream calling exec_command, apply_patch (wrapped), _create_event and tool_search comes back as function_call + namespace: functions, custom_tool_call with the raw patch, function_call + namespace: mcp__codex_apps__calendar, and tool_search_call — whole, streamed (7-byte chunks, added/done items readable by Codex's types, no deltas for the search) and stream-collected; strip_carried leaves the Lite body alone.
  • Unit tests: Lite decoding (tools, shapes, system notes, effort override, paired history, agent message), mixed top-level + additional_tools with redefinitions, name flattening and the 64-char cut, refusals (compaction_trigger, encrypted context_compaction), verbosity by model.
  • caller.rs cross-check extended with the new items (agent message text is caller text, namespace notes are system prompt).
  • Gateway e2e (tw-gateway/tests/conversion.rs): a Lite request reaches a mock Claude with every tool, the stream comes back with Codex-shaped calls, the reported dropped list has no tool declarations; a Lite request to an OpenAI Responses upstream arrives byte for byte.
  • tw-store: a Lite request's transcript reads its tools from input and unwraps the converted apply_patch.

Checks (rebased on fc8838e): cargo fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings, cargo test --workspace --no-fail-fast (3121 passed, 8 ignored; the one failure locally is tw-watch's timing-sensitive fd_budget FSEvents test, which this PR doesn't touch and which passes on its own), scripts/smoke.sh (79 passed, 0 failed).

🤖 Generated with Claude Code

…ry item it sends

Codex with Responses Lite (`use_responses_lite`) sends no top-level `tools`:
every tool goes into an `additional_tools` input item, with functions and
freeform tools grouped in the `functions` namespace. The Responses decoder
read only top-level `tools` and dropped unknown input items, so a request
converted for an Anthropic, Chat, Gemini or Bedrock upstream reached it with
no tools at all.

- `additional_tools` entries decode like top-level tools (namespaces
  included). Later declarations replace earlier ones in place, the way
  Codex's incremental tool updates say they should.
- Tools in the default `functions` namespace keep their plain names; other
  namespaced names follow Codex's own joining and are cut to 64 characters
  with a stable hash. Namespace descriptions (MCP server instructions) go
  into the system prompt.
- Client-executed `tool_search` becomes a function tool; an upstream call to
  it goes back to Codex as a `tool_search_call`. `tool_search_call` /
  `tool_search_output` history becomes a call and result, and the tools it
  found become callable.
- `local_shell_call` history becomes a `local_shell` call paired with its
  `function_call_output` (an orphaned result is rejected by Anthropic).
- `agent_message` becomes a user turn; `configuration_update` (effort
  changed mid-conversation) overrides the request's reasoning effort.
- `compaction_trigger` and encrypted `context_compaction` are refused like
  `compaction`; `service_tier` other than auto/default is reported dropped.
- `text.verbosity` / Chat `verbosity` are carried in the IR and written for
  GPT-5-family upstream models; elsewhere they are reported dropped.
- The transcript reader uses the same tool names and reads freeform tools
  from `additional_tools`.

Same-format passthrough is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn force-pushed the fix/codex-responses-items branch from 5a91dee to 1bcba6c Compare October 9, 2026 08:40
@fylorn
fylorn merged commit f9b6a25 into main Oct 9, 2026
5 checks passed
@fylorn
fylorn deleted the fix/codex-responses-items branch October 9, 2026 08:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant