Repository navigation
Responses conversion: keep the tools Codex declares in input, and every item it sends - #304
Merged
Merged
Conversation
…ry item it sends Codex with Responses Lite (`use_responses_lite`) sends no top-level `tools`: every tool goes into an `additional_tools` input item, with functions and freeform tools grouped in the `functions` namespace. The Responses decoder read only top-level `tools` and dropped unknown input items, so a request converted for an Anthropic, Chat, Gemini or Bedrock upstream reached it with no tools at all. - `additional_tools` entries decode like top-level tools (namespaces included). Later declarations replace earlier ones in place, the way Codex's incremental tool updates say they should. - Tools in the default `functions` namespace keep their plain names; other namespaced names follow Codex's own joining and are cut to 64 characters with a stable hash. Namespace descriptions (MCP server instructions) go into the system prompt. - Client-executed `tool_search` becomes a function tool; an upstream call to it goes back to Codex as a `tool_search_call`. `tool_search_call` / `tool_search_output` history becomes a call and result, and the tools it found become callable. - `local_shell_call` history becomes a `local_shell` call paired with its `function_call_output` (an orphaned result is rejected by Anthropic). - `agent_message` becomes a user turn; `configuration_update` (effort changed mid-conversation) overrides the request's reasoning effort. - `compaction_trigger` and encrypted `context_compaction` are refused like `compaction`; `service_tier` other than auto/default is reported dropped. - `text.verbosity` / Chat `verbosity` are carried in the IR and written for GPT-5-family upstream models; elsewhere they are reported dropped. - The transcript reader uses the same tool names and reads freeform tools from `additional_tools`. Same-format passthrough is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fylorn
force-pushed
the
fix/codex-responses-items
branch
from
October 9, 2026 08:40
5a91dee to
1bcba6c
Compare
This was referenced Oct 9, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A Codex user routed through the gateway to a non-Responses upstream (Anthropic / Chat / Gemini / Bedrock) got "the workspace read/write and command execution tools are not mounted"; the traffic record listed
input.additional_tools,text.verbosityand more as dropped.Root cause
With Responses Lite (
model_info.use_responses_lite), Codex sends no top-leveltools. It puts every tool into an input item{"type": "additional_tools", "role": "developer", "tools": [...]}(client.rs#L972-L987, models.rs#L1016-L1023), built bycreate_tools_json_for_responses_lite, which groups functions and freeform tools into thefunctionsnamespace.responses::request::decode_requestread only top-leveltools, and unknown input items fell intoother => dropped.path("input.{other}")— so the converted request reached the upstream with no tools at all.What changes
additional_toolsentries go through the same tool decoding as top-leveltools, namespaces included. Order is deterministic: top-leveltoolsfirst, thenadditional_tools/tool_search_outputin input order; a later declaration with the same name replaces the earlier one in place — Codex's incremental tool updates say exactly that (top_level_tools.rs#L266), and keeping positions keeps the tool list (the start of the prompt cache) stable.functionsnamespace keep their plain names (exec_command,apply_patch) — Codex treatsfunctions.xandxas the same tool (tool_name.rs#L7, #L39). Other namespaces join the way Codex does (mcp__codex_apps__calendar+_create_event→mcp__codex_apps__calendar_create_event, otherwisens__name) and names over 64 characters (the OpenAI Chat / Gemini / Bedrock limit) are cut with a stable FNV hash.ClientShape.namespacedmaps every flattened name back, so calls go back to Codex withnamespace+name, asfunction_callorcustom_tool_call.tool_searchbecomes a function tool; when the upstream calls it, Codex receives atool_search_call(execution: client, objectarguments,tsc_id, no argument deltas), which is what Codex dispatches (router.rs#L275-L294). The tools found in atool_search_outputbecome callable.ClientShape.tool_search,Request.verbosity/ir::Verbosity,Feature::Verbosity,responses::request::{flat_tool_name, freeform_tools, TOOL_SEARCH}— all additive (the enterprise edition buildsRequestwith..Default::default()and doesn't touchClientShapeorFeature).tw-store's transcript reader uses the same flattened names, reads freeform tools fromadditional_tools(so a wrappedapply_patchfrom another format is unwrapped), and doesn't show the declaration item as a turn.tw-gatewaysources (the new gateway tests live intests/conversion.rs).Every
ResponseItemCodex can send (models.rs#L1015-L1259)additional_toolsinput.additional_tools)toolsmessage(+phase)phasehas no equivalent, the text is kept, not reportedagent_messageinput_text(sender/task are in the text, inter_agent_message.rs);encrypted_contentreported dropped;caller.rscovers the text for content filteringreasoningfunction_call/custom_tool_call(+ outputs)ns__namenamespace: "functions"maps to the plain namelocal_shell_callfunction_call_outputwas an orphaned result (Anthropic rejects the request)local_shellcall withactionas input, paired with its output (normalize.rs#L104); withoutcall_idreported droppedtool_search_calltool_searchcall (needscall_id)tool_search_outputconfiguration_updatereasoning.effort— Codex pins the request-level effort for cache reasons and records mid-session changes this way (reasoning_effort.rs); a model-specific effort is reported droppedweb_search_call,image_generation_callcompaction,item_referencecontext_compactionencrypted_content, otherwise reported droppedcompaction_triggerTop-level fields:
text.verbosity(and Chatverbosity)verbosity(Chat) /text.verbosity(Responses) when the upstream model is GPT-5 or later (ir::Verbosity::understood_by), otherwise reported dropped — other models answer 400 or ignore itservice_tierauto/default(no equivalent tier elsewhere)includeencrypted_contentstore,prompt_cache_key,stream_options,reasoning.contextclient_metadataTypical Codex Responses Lite request, dropped list
input.additional_tools,text.verbosity, plusinput.local_shell_call/input.tool_search_call/input.tool_search_output/input.configuration_update/input.agent_messagewhen present — and no tools upstream.input.reasoning(OpenAI-encrypted reasoning, unchanged),tools.custom.format(apply_patch's lark grammar, unchanged),text.verbosity; Gemini/Bedrock alsoparallel_tool_calls; Chat to a GPT-5 model drops no verbosity. Every tool arrives.Tests
tw-dialect/tests/codex_lite.rs(new): a Codex-shaped Lite request (built from Codex's serde types, cf. responses_lite.rs) withexec_command,write_stdin, freeformapply_patch,update_plan, an MCP namespace,web.run,tool_search, plus history withlocal_shell_call, a tool search and an effort change — to Anthropic, Chat, Gemini and Bedrock: every tool present (read back with each format's own decoder), history calls paired with results, exact dropped list; the upstream callingexec_command,apply_patch(wrapped),_create_eventandtool_searchcomes back asfunction_call+namespace: functions,custom_tool_callwith the raw patch,function_call+namespace: mcp__codex_apps__calendar, andtool_search_call— whole, streamed (7-byte chunks, added/done items readable by Codex's types, no deltas for the search) and stream-collected;strip_carriedleaves the Lite body alone.additional_toolswith redefinitions, name flattening and the 64-char cut, refusals (compaction_trigger, encryptedcontext_compaction), verbosity by model.caller.rscross-check extended with the new items (agent message text is caller text, namespace notes are system prompt).tw-gateway/tests/conversion.rs): a Lite request reaches a mock Claude with every tool, the stream comes back with Codex-shaped calls, the reported dropped list has no tool declarations; a Lite request to an OpenAI Responses upstream arrives byte for byte.tw-store: a Lite request's transcript reads its tools from input and unwraps the convertedapply_patch.Checks (rebased on fc8838e):
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warnings,cargo test --workspace --no-fail-fast(3121 passed, 8 ignored; the one failure locally istw-watch's timing-sensitivefd_budgetFSEvents test, which this PR doesn't touch and which passes on its own),scripts/smoke.sh(79 passed, 0 failed).🤖 Generated with Claude Code