Learning AI Quality 返回 KuthorX Blog II博客首页

第 15 章

第 3 周:Codex 怎么定义、分发和执行工具?

读 OpenAI Codex 源码里的工具系统和 MCP client:ToolSpec 和 ToolExecutor 怎么把定义和实现绑在一起,schema 只保留哪些字段,apply_patch 为什么用 Lark 语法而不用 JSON,一次调用经过哪些关卡,哪些工具能并行,输出怎么保头保尾截断,MCP server 怎么接进来(命名空间、超时、审批)。用一个本地假模型服务实测畸形工具调用,不调真实模型。一段讲解视频,一个工具调用分发模拟器。

第 2 周我们手写的 agent loop 里,工具就是 TOOLS 列表加 TOOL_IMPLS 字典:模型说调哪个,就从字典里取函数执行,出错就回一个 is_error。第 11 章又写了一个 MCP server(bughunt),让任何 Agent 都能插上用。这一章看一个生产级 Agent 是怎么做同一件事的:Codex 怎么把工具定义发给模型、模型的调用怎么一路分发到执行、哪些能并行、输出太长怎么办、外接的 MCP server 怎么接进来。最后把第 2 周的 bughunt 挂进 Codex,用一个假的模型服务发一组畸形调用,看 Codex 实际回给模型的是什么。

源码基于 openai/codex commit 7993248 (2026-10-01),下文路径都相对 codex-rs/。实验用本机 codex-cli 0.159.2(macOS),模型换成本地假服务,不调真实模型、不消耗额度。

讲解视频

互动演示

上面是工具调用分发模拟器:选一个模型发来的调用(正常的、缺参数的、补丁格式错的、未知工具、MCP 参数截断、MCP 超时……),单步看它在哪一关停下、回填给模型的是什么。可以切换运行方式(codex exec 还是交互式)、MCP 审批模式、超时和模型的截断策略,截断数字按源码公式实时计算,错误文字取自本机实测。下面是并行门时间线:按顺序添加几个调用,看读写锁怎么排队。页面底部有自动判分的练习。

互动演示:工具调用分发模拟器 在新标签页打开

一张图

模型的流式输出  response.output_item.done
        │  function_call / custom_tool_call
        ▼
 ToolRouter::build_tool_call     → ToolCall { tool_name, payload }
        │  tokio::spawn,按出现顺序放进 FuturesOrdered
        ▼
 ① 并行门    可并行 → 读锁;其余(含未知工具)→ 写锁
 ② 查注册表  找不到 → "unsupported call: <名字>"
 ③ payload 类型   function 工具收到 custom 调用 → Fatal
 ④ PreToolUse hook   可拦截、可改写参数
 ⑤ handler   解析参数 → 审批 → 执行
 ⑥ PostToolUse hook
 ⑦ 格式化 + 截断 → function_call_output / custom_tool_call_output
        │  按调用顺序写进历史
        ▼
 下一次请求的 input

无论在哪一关停下,这个 call_id 都会有一条 output 回给模型,turn 继续。下面一层层拆开。

工具 = spec + handler

每个工具都实现 ToolExecutor 这个 trait( tools/src/tool_executor.rs:106-130 ,节选):

pub trait ToolExecutor<Invocation>: Send + Sync {
    /// The concrete tool name handled by this runtime instance.
    fn tool_name(&self) -> ToolName;

    fn spec(&self) -> ToolSpec;

    /// The preferred exposure before the host applies step-specific policy.
    fn exposure(&self) -> ToolExposure {
        ToolExposure::Direct
    }
    // ...
    fn supports_parallel_tool_calls(&self) -> bool {
        false
    }

    /// Handles one invocation without retaining capabilities borrowed by the host.
    fn handle<'a>(&'a self, invocation: Invocation) -> ToolExecutorFuture<'a>
    where
        Invocation: 'a;
}

不懂 Rust 也能读:trait 相当于 Python 的抽象基类,fn 是方法,&self 是 self,-> X 是返回类型。带函数体的方法(exposure、supports_parallel_tool_calls)是默认实现,具体工具可以覆盖。

  • tool_name():注册表里的键,由"命名空间 + 名字"组成。
  • spec():发给模型的工具定义。
  • exposure():默认 Direct,直接出现在请求的 tools 里;还可以是 Deferred(藏在 tool_search 后面,模型搜到才加载)、Hidden 等,一共 6 种(另外三种是 DeferredModelOnly、DirectModelOnly、CodeModeOnly, tools/src/tool_executor.rs:49-99 )。
  • supports_parallel_tool_calls():默认 false。想和别的调用同时跑,工具得自己声明。
  • handle():真正执行,返回一个异步的 future。

和第 2 周对比:我们的 TOOLS 和 TOOL_IMPLS 是两份东西,靠名字字符串对上,加了定义忘了加实现,只有模型真调到时才 KeyError。Codex 把两者放在同一个对象上,注册时就配好了。“默认不能并行"也是一个设计选择:新工具不用考虑并发安全,出错的代价是慢,而不是两个调用同时改坏一个文件。

ToolSpec 是一个枚举( tools/src/tool_spec.rs:20-56 ),对应 Responses API 里五种工具:

变体序列化后的 type谁在用
Functionfunctionexec_command、view_image 等,参数是 JSON Schema
Namespacenamespace一组工具的容器。每个 MCP server 一个,例如 mcp__bughunt
Freeformcustomapply_patch:输入是纯文本,格式由 Lark 语法描述
ToolSearchtool_search搜索延迟加载的工具
WebSearchweb_search服务端托管的搜索,不经过本地分发

每个 turn 开始前,build_tool_router( core/src/tools/spec_plan.rs:123-188 )按固定顺序把工具装进注册表:内置工具 → MCP 工具(再按策略决定直接列出还是延迟)→ 扩展工具 → 动态工具 → 托管工具,最后生成发给模型的列表。

schema:只保留一个子集,strict 一律 false

Codex 的 JsonSchema 是个普通结构体( tools/src/json_schema/types.rs:35-75 ),只有 type、description、encrypted、enum、items、minItems、properties、required、additionalProperties、anyOf / oneOf / allOf、$ref / $defs / definitions 这些字段。外来的 schema(MCP server、动态工具)先经过 sanitize_json_schema( tools/src/json_schema.rs:71-80 ):const 改写成单值 enum,缺 type 的按出现的关键字推断。结构体里没有的字段,反序列化时直接丢掉。所以数组的 minItems 会保留,minimum、pattern、maxLength 会丢。源码测试里 {"minimum": 1} 解析完只剩 {"type": "number"}( tools/src/json_schema_tests.rs:198-214 )。MCP 工具的 schema 如果没写 properties,会补一个空对象( tools/src/mcp_tool.rs:45-55 )。

拿第 2 周的 bughunt 实测。Python SDK 生成的 submit_bug_report 输入 schema(title 字段是 SDK 自动加的):

{"type": "object", "title": "submit_bug_reportArguments",
 "properties": {
   "module":   {"title": "Module", "type": "string"},
   "title":    {"title": "Title", "type": "string"},
   "steps":    {"title": "Steps", "type": "array", "items": {"type": "string"}},
   "severity": {"enum": ["low", "medium", "high"], "title": "Severity", "type": "string"}},
 "required": ["module", "title", "steps", "severity"]}

Codex 发给模型的那一项(摘自 mock_model_tools 场景,即模型 slug 设成未知的 mock-model 时第一次请求里的 mcp__bughunt 命名空间,节选,省略 description。默认的 gpt-5.5 会把 MCP 工具延迟到 tool_search 后面,第一次请求里看不到,见下文"延迟加载”):

{"type": "function", "name": "submit_bug_report", "strict": false,
 "parameters": {"type": "object",
   "properties": {"module": {"type": "string"},
                  "severity": {"type": "string", "enum": ["low", "medium", "high"]},
                  "steps": {"type": "array", "items": {"type": "string"}},
                  "title": {"type": "string"}},
   "required": ["module", "title", "steps", "severity"]}}

三处变化:title 被丢掉(list_injected_bugs 参数里的 "default": null 也一样);属性按字母序排了(properties 是 BTreeMap,有序映射);strict 是 false。最后一点写死在代码里:MCP 和动态工具转换成 Responses API 工具时一律 strict: false( tools/src/responses_api.rs:164-173 ),内置的 exec_command 也是 strict: false( core/src/tools/handlers/shell_spec.rs:106 )。

这意味着两件事。第一,服务端不保证模型给的参数符合 schema;第二,Codex 本地也不按 schema 校验 MCP 参数,只做 JSON 解析。实测模型把 severity 写成 "critical",Codex 原样转发,是 bughunt 的 pydantic 拒绝的。所以第 9 章「模型怎么’调用’一个函数?」说的"边界处一律校验",在 Codex 体系里落在 MCP server 一侧:minimum、pattern 这类约束模型根本看不到,server 必须自己校验,并把错误写到模型能看懂。

schema 太大还会被压缩:单个 MCP 工具的输入 schema 超过 5,000 字节( tools/src/json_schema/compaction.rs:15 ,可用 tool_input_schema_max_bytes 按 server 调整, config/src/mcp_types.rs:252-255 )时,依次做四轮越来越有损的压缩:去掉描述、去掉 $defs、折叠深层对象、修剪组合关键字(anyOf 等),每轮之前先看是否已经够小( tools/src/json_schema/compaction.rs:18-37 )。源码注释写明这是 “best-effort rather than a hard cap”,四轮做完仍可能超限。

内置工具

本机 0.159.2 用 gpt-5.5 这个模型 slug 指向假服务,第一次请求里实际发出的工具:

工具类型能否并行做什么
exec_commandfunction是在 PTY 里跑命令。必填 cmd,可选 workdir、tty、yield_time_ms、max_output_tokens 等;没跑完的返回 session ID
write_stdinfunction是往还在跑的 session 写输入、取新输出
apply_patchcustom否改文件
view_imagefunction是把本地图片放进上下文
list_mcp_resources 等 3 个function是读 MCP resource
request_user_inputfunction否向用户提问
get_goal / create_goal / update_goalfunction否目标管理
tool_searchtool_search是搜索延迟加载的工具
web_searchweb_search—服务端执行

“能否并行"一栏来自各个 handler 的 supports_parallel_tool_calls,例如 exec_command 在 core/src/tools/handlers/unified_exec/exec_command.rs:142-144 。没有单独的老 shell 工具了。工具清单随模型元数据和 feature flag 变化:把模型 slug 换成一个内置目录里没有的名字,请求里就没有 apply_patch 和 tool_search,多了 multi_agent_v1 命名空间。不同版本、不同模型请以实际请求为准。

exec_command 的 schema 写在 core/src/tools/handlers/shell_spec.rs:24-115 :required: ["cmd"],additionalProperties: false,max_output_tokens 的描述是 “Output token budget. Defaults to 10000 tokens; larger requests may be capped by policy."( core/src/tools/handlers/shell_spec.rs:61 ,实测请求体里也是这句)。

apply_patch:不用 JSON,用语法

apply_patch 是一个 custom 工具( core/src/tools/handlers/apply_patch_spec.rs:5-28 )。它的 format 是 {"type": "grammar", "syntax": "lark", "definition": <语法全文>},描述里写着 “This is a FREEFORM tool, so do not wrap the patch in JSON."。语法全文( core/assets/tools/apply_patch.lark:1-19 ):

start: begin_patch hunk+ end_patch
begin_patch: "*** Begin Patch" LF
end_patch: "*** End Patch" LF?

hunk: add_hunk | delete_hunk | update_hunk
add_hunk: "*** Add File: " filename LF add_line+
delete_hunk: "*** Delete File: " filename LF
update_hunk: "*** Update File: " filename LF change_move? change?

filename: /(.+)/
add_line: "+" /(.*)/ LF -> line

change_move: "*** Move to: " filename LF
change: (change_context | change_line)+ eof_line?
change_context: ("@@" | "@@ " /(.+)/) LF
change_line: ("+" | "-" | " ") /(.*)/ LF
eof_line: "*** End of File" LF

%import common.LF

读法:一个补丁以 *** Begin Patch 开头、*** End Patch 结尾,中间是一个或多个 hunk。hunk 有三种:新建文件(之后每行以 + 开头)、删除文件、修改文件(可选改名,然后是若干段改动;每段以 @@ 加一行定位上下文开头,再跟 + / - / 空格开头的行)。例如:

*** Begin Patch
*** Update File: src/cart.py
@@ def checkout(cart):
-    if cart.qty >= 0:
+    if cart.qty > 0:
         submit(cart)
*** End Patch

为什么不用 JSON 参数:补丁里全是换行、引号、反斜杠,塞进 JSON 字符串要多一层转义,模型很容易在转义上出错。OpenAI 的函数调用文档把 custom 工具描述为输入可以是不受约束的自由文本,也可以用 Lark 或正则语法约束输出格式( OpenAI 函数调用指南 、 API 参考 CustomToolInputFormat )。也就是说格式由 API 一侧约束;Codex 本地仍然再解析一遍。所以假服务直接塞一个坏补丁(绕过了语法),得到的是本地解析器的报错:

apply_patch verification failed: invalid patch: The first line of the patch must be '*** Begin Patch'

apply_patch 的 handler 只接受 custom 形态的调用( core/src/tools/handlers/apply_patch.rs:407-409 )。如果有个 provider 不支持 custom 工具、模型用 function_call 的形态调它,会怎样?下一节的"Fatal”。

分发:一次调用经过哪些关卡

流里每出现一个完整的输出项,ToolRouter::build_tool_call( core/src/tools/router.rs:248-300 )把它变成内部的 ToolCall:function_call 变成 ToolPayload::Function { arguments }(arguments 此时还是字符串),custom_tool_call 变成 ToolPayload::Custom { input },namespace 和 name 合成 ToolName。然后 ToolCallRuntime 把它 spawn 成一个异步任务( core/src/tools/parallel.rs:196-245 ),任务里依次:

  1. 并行门:可并行的拿读锁,其余拿写锁(下一节细讲)。注意它在查注册表之前。查"能否并行"时工具还不存在,tool_supports_parallel 的 unwrap_or(false) 让未知工具也拿写锁( core/src/tools/router.rs:237-241 )。
  2. 查注册表( core/src/tools/registry.rs:551-571 ):找不到就返回 RespondToModel,文字由 unsupported_tool_call_message 生成:custom 调用是 unsupported custom tool call: X,其余是 unsupported call: X( core/src/tools/registry.rs:852-857 )。
  3. payload 类型检查( core/src/tools/registry.rs:584-600 ):matches_kind 不通过就是 Fatal("tool X invoked with incompatible payload")。默认的 matches_kind 只接受 Function 和 ToolSearch( core/src/tools/registry.rs:85-90 )。
  4. PreToolUse hook( core/src/tools/registry.rs:602-654 ):用户配的 hook 可以拦下这次调用(回一条说明给模型),也可以改写参数。
  5. handler:解析参数、审批、执行。参数解析失败是 RespondToModel("failed to parse function arguments: ...")( core/src/tools/handlers/mod.rs:86-93 )。
  6. PostToolUse hook( core/src/tools/registry.rs:709-771 ):成功的结果还能被 hook 拦下或附加反馈。
  7. 格式化、截断、回填:结果变成 function_call_output 或 custom_tool_call_output。所有调用任务放在一个 FuturesOrdered 里( core/src/session/turn.rs:2589 ),所以即使并行执行,写回历史的顺序也和模型发出的顺序一致。

错误分两类,但模型只看到文字

工具出错时返回 FunctionCallError( tools/src/function_call_error.rs:4-10 ):

pub enum FunctionCallError {
    #[error("{0}")]
    RespondToModel(String),
    #[error("Fatal error: {0}")]
    Fatal(String),
}

RespondToModel 的意思是"把这段文字当作工具输出还给模型,让它自己改”;Fatal 是"不该发生的内部错误”。#[error(...)] 是错误转成字符串时的格式。

RespondToModel 和其他非 Fatal 错误,在 failure_response( core/src/tools/parallel.rs:303-328 )里变成一条普通的 output,内部标记 success: false。但这个标记不发给模型( protocol/src/models.rs:2175-2183 、 protocol/src/models.rs:2253-2263 ):

/// `body` serializes directly as the wire value for `function_call_output.output`.
/// `success` remains internal metadata for downstream handling.
pub struct FunctionCallOutputPayload {
    pub body: FunctionCallOutputBody,
    pub success: Option<bool>,
}
impl Serialize for FunctionCallOutputPayload {
    fn serialize<S>(&self, serializer: S) -> Result<S::Ok, S::Error>
    // ...
        match &self.body {
            FunctionCallOutputBody::Text(content) => serializer.serialize_str(content),
            FunctionCallOutputBody::ContentItems(items) => items.serialize(serializer),
        }

自定义的 Serialize 只输出 body。对比第 2 周:Claude API 的 tool_result 有 is_error 字段,MCP 的工具结果有 isError 字段;到了 Codex 发给 Responses API 的请求里,两者都只剩文字。实测 bughunt 返回 isError: true 的结果,模型看到的是 Error executing tool submit_bug_report: ... 这段文本,没有任何标志。错误文字写得清不清楚,直接决定模型能不能自己改对。

Fatal 并不会终止 turn

直觉上 Fatal 应该让整个 turn 失败。handle_output_item_done 里确实留了一个分支:如果 build_tool_call 返回 Fatal,就转成 CodexErr::Fatal 结束 turn( core/src/stream_events_utils.rs:427-429 )。但这个 commit 里 build_tool_call 只会返回 RespondToModel,不会返回 Fatal( core/src/tools/router.rs:248-300 )。真正会出现的 Fatal 来自 registry 里的类型检查,它发生在已经 spawn 出去的工具任务里,走的是另一条路。(第 14 章讲 agent loop 时把 Fatal 归入"终止",说的是采样请求出错时 CodexErr 的重试分类,和这里工具任务返回的 FunctionCallError::Fatal 不是一回事。)turn 收尾时,drain_in_flight 逐个等待在途的工具任务( core/src/session/turn.rs:2472-2500 ):

            Err(err) => {
                error_or_panic(format!("in-flight tool future failed during drain: {err}"));
            }

error_or_panic( core/src/util.rs:81-87 )在 debug 构建(cfg!(debug_assertions))里 panic,在 release 构建里只记一条 error 日志。这次调用没有产出 output,发下一次请求前,历史规范化给缺 output 的调用补一条内容为 "aborted" 的 output( core/src/context_manager/normalize.rs:51-67 )。

实测(release 版 0.159.2):假模型用 function_call 形态调 apply_patch,日志里有 Fatal error: tool apply_patch invoked with incompatible payload,第二次请求里这个 call_id 的 output 是 aborted,turn 正常结束,codex 退出码 0。反过来用 custom_tool_call 调 exec_command 也一样:output 是 aborted,退出码 0。custom 调用缺 output 时,规范化阶段还会再调一次 error_or_panic( core/src/context_manager/normalize.rs:87-92 ),debug 构建在这里同样会 panic。

这对测试很重要:同一个用例,debug 构建会 panic,release 构建只是悄悄多一条 aborted。CI 里跑 Codex 的集成测试,要写清楚用的是哪种构建;断言"Fatal 会让 turn 失败"的测试在 release 下会失败。

并行:一把公平的读写锁

工具任务拿到这把锁才开始执行( core/src/tools/parallel.rs:205-209 ):

                let guard = if supports_parallel {
                    Either::Left(lock.read().await)
                } else {
                    Either::Right(lock.write().await)
                };

所有调用共用一把 RwLock( core/src/tools/parallel.rs:44-65 )。声明可并行的拿读锁,多个读锁可以同时持有;其余拿写锁,独占。Either 只是把两种锁守卫装进同一个变量。锁一直持有到 handler 执行完。

请求体里的 parallel_tool_calls 除了走 Responses Lite 的模型以外都是 true( core/src/client.rs:1001 ,实测也是 true),这只表示允许模型一次发多个调用;真正能不能同时跑,看这把锁。tokio 的 RwLock 是公平锁。Codex 的 Cargo.lock 锁定 tokio 1.52.3,它的 src/sync/rwlock.rs 类型文档写的是:“Fairness is ensured using a first-in, first-out queue for the tasks awaiting the lock; a read lock will not be given out until all write lock requests that were queued before it have been acquired and released."( tokio 1.52.3 src/sync/rwlock.rs:41-47 )。也就是说,排在写锁后面的读锁要等写锁用完。所以调用顺序会影响总耗时:按这个先进先出模型,exec_command(2 秒)、exec_command(2 秒)、apply_patch(1 秒)、exec_command(2 秒)要 5 秒;把 apply_patch 挪到最后只要 3 秒。apply_patch 执行太快不好实测,我用一个 2 秒的 MCP 调用(写锁)代替它跑了 lock_order 场景:exec、exec、replay_steps、exec 各 2 秒,两次请求间隔 6.12 秒;如果后面的读锁能插队,应该是 4 秒左右。

哪些工具声明了可并行:exec_command、write_stdin、view_image、tool_search、三个 MCP resource 工具。MCP 工具看配置和标注( core/src/tools/handlers/mcp.rs:148-159 ):

    fn supports_parallel_tool_calls(&self) -> bool {
        // Correctly implemented MCP servers should tolerate parallel calls to
        // tools that advertise themselves as read-only.
        self.tool_info.supports_parallel_tool_calls
            || self
                .tool_info
                .tool
                .annotations
                .as_ref()
                .and_then(|annotations| annotations.read_only_hint)
                .unwrap_or(false)
    }

即 server 配置了 supports_parallel_tool_calls = true( config/src/mcp_types.rs:248-250 ),或者工具标注了 readOnlyHint: true。注释的理由是:实现正确的 MCP server 应该能承受对只读工具的并行调用。

实测:假模型一次发三个 replay_steps(bughunt 里每次 await anyio.sleep(2))。

场景第 1 次 → 第 2 次请求间隔每条 output 的 Wall time
默认配置(MCP 工具拿写锁)6.06 秒2.0180 / 2.0089 / 2.0034 秒
supports_parallel_tool_calls = true2.02 秒2.0043 / 2.0035 / 2.0026 秒
三个 exec_command "sleep 2"2.09 秒1.8790 / 1.8727 / 1.8728 秒

串行时每条 Wall time 仍然只有 2 秒左右:它只量 handler 自己的执行时间,排队等锁的时间不算(源码也把拿到锁的时刻记作"执行开始”, core/src/tools/parallel.rs:211-214 )。只看 output 是看不出调用被串行化的,要看请求间隔。

还有一个副作用:MCP 的审批发生在 handler 里,也就是拿着锁的时候。交互模式下一个等用户点批准的 MCP 调用(写锁)会挡住它后面的所有调用。

输出截断:保头保尾,砍中间

exec_command 的输出预算取两者中较小的( core/src/tools/context.rs:481-489 ):

    fn model_output_policy(&self) -> TruncationPolicy {
        let requested_policy = TruncationPolicy::Tokens(resolve_max_tokens(self.max_output_tokens));
        if requested_policy.byte_budget() < self.truncation_policy.byte_budget() {
            requested_policy
        } else {
            self.truncation_policy
        }
    }

超出预算时,预算一半给开头、一半给结尾,中间换成一个标记( utils/string/src/truncate.rs:132-143 ):

fn split_budget(budget: usize) -> (usize, usize) {
    let left = budget / 2;
    (left, budget - left)
}

fn format_truncation_marker(use_tokens: bool, removed_count: u64) -> String {
    if use_tokens {
        format!("…{removed_count} tokens truncated…")
    } else {
        format!("…{removed_count} chars truncated…")
    }
}

全是 ASCII 时,输出 \(N\) 字节、预算 \(B\) 个 token:

$$\text{保留} = 4B\ \text{字节},\qquad \text{标记里的数} = \left\lceil \frac{N - 4B}{4} \right\rceil,\qquad \text{Original token count} = \left\lceil \frac{N}{4} \right\rceil$$

实测 python3 -c "print('x'*200000)",\(N = 200{,}001\)(含末尾换行),\(B = \min(10{,}000, 10{,}000)\):保留 40,000 字节,头尾各 20,000;标记 \(\lceil 160{,}001 / 4 \rceil = 40{,}001\);Original token count \(\lceil 200{,}001/4 \rceil = 50{,}001\)。回给模型的 output(中间的 x 省略):

Chunk ID: 8b4012
Wall time: 0.0000 seconds
Process exited with code 0
Original token count: 50001
Output:
Warning: truncated output (original token count: 50001)
Total output lines: 1

xxxx…(20,000 个 x)…40001 tokens truncated…xxxx(19,999 个 x 加换行)

整条 40,209 个字符(40,213 字节,两个省略号各占 3 字节)。头部的格式见 core/src/tools/context.rs:524-548 。如果模型元数据走兜底的 10,000 字节策略,预算就只有 10,000 字节,标记变成 …N chars truncated…。exec 进程那一侧另有 1 MiB 的缓冲上限( core/src/unified_exec/mod.rs:80 )。

MCP 工具的输出加一行 Wall time 头,按该工具配置的 output_token_limit 或模型策略截断,再乘 1.2 的"序列化余量"( core/src/tools/context.rs:184-210 、 utils/output-truncation/src/lib.rs:14-18 )。

和第 2 周对比:我们写的是 f.read()[:20_000],只留开头。测试命令的失败摘要、编译器最后一个报错都在结尾,只留开头会把最有用的部分砍掉。Codex 还在头部写明原始 token 数,模型知道自己没看全,可以换个命令(比如 tail)去看。

MCP client:连接、命名空间、超时、审批

配置和连接

每个 server 是 config.toml 里的一个 [mcp_servers.<名字>] 表( config/src/config_toml.rs:293-294 )。传输二选一( config/src/mcp_types.rs:612-646 ):写 command 是 stdio,Codex 把 server 当子进程拉起;写 url 是 Streamable HTTP。常用字段( config/src/mcp_types.rs:222-300 ):

字段默认作用
enabledtrue关掉就不连
requiredfalse为 true 时,server 起不来就让 codex exec 报错退出
startup_timeout_sec30启动握手加首次列工具的超时
tool_timeout_sec300单次工具调用的超时
enabled_tools / disabled_tools无工具白名单 / 黑名单
supports_parallel_tool_callsfalse这个 server 的工具都拿读锁
default_tools_approval_modeauto这个 server 工具的默认审批模式
[mcp_servers.<名字>.tools.<工具>]—按工具设 approval_mode、output_token_limit

两个默认超时是常量( codex-mcp/src/rmcp_client.rs:106-107 ),没配置时套用( codex-mcp/src/connection_manager.rs:358-365 )。调用方自己也带了超时时取两者较小的( codex-mcp/src/binding.rs:313-318 );不过模型发起的调用传的 requested_timeout 是 None( core/src/mcp_tool_call.rs:466 ),实际生效的就是 tool_timeout_sec。

命名空间

模型看到的不是一串扁平的工具名,而是每个 server 一个 namespace 工具,里面放这个 server 的全部工具( core/src/tools/handlers/mcp.rs:498-531 )。实测请求体里 bughunt 长这样(摘自 mock_model_tools 场景,即未知模型 slug 时的第一次请求;节选,省略各工具的 description):

{"type": "namespace", "name": "mcp__bughunt",
 "description": "BugHunt-Bench 评测环境:查询注入 Bug、提交 Bug 报告。",
 "tools": [{"type": "function", "name": "list_injected_bugs", "strict": false,
            "parameters": {"type": "object",
                           "properties": {"module": {"anyOf": [{"type": "string"}, {"type": "null"}]}}}},
           ...]}

命名空间的描述就是 server 初始化时给的 instructions。名字的规则( codex-mcp/src/tools.rs:105-214 、 codex-mcp/src/tools.rs:225-300 ):

  • server 名清洗成只含 [A-Za-z0-9_]( codex-mcp/src/mcp/mod.rs:578-593 ),加上前缀 mcp__。前缀由一个 non_prefixed_mcp_tool_names feature 控制,默认关闭,即默认带前缀( core/src/config/mod.rs:1890-1893 )。
  • 清洗后撞名的,加 _ 和 SHA1 的前 12 位十六进制;模型可见的名字最长 128。
  • 工具名本身不加前缀。调 server 时用的仍是原始工具名。

codex mcp add 对 server 名有更严的检查,只允许字母、数字和 - _ : @ / .( cli/src/mcp_cmd.rs:1148-1157 )。

一个容易看错的地方:日志和错误文字里把命名空间和工具名直接拼接( protocol/src/tool_name.rs:54-62 ),所以你会看到 mcp__bughuntlist_injected_bugs,中间没有分隔符。hook 里用的名字则是 mcp__bughunt__list_injected_bugs( core/src/tools/handlers/mcp.rs:120-137 )。

延迟加载

模型支持搜索工具时,MCP 工具默认是 Deferred:不出现在请求的 tools 里,而是在 tool_search 的描述中列出 server(“bughunt: BugHunt-Bench 评测环境…"),模型先搜,搜到再加载( core/src/mcp_tool_exposure.rs:85-89 、 core/src/tools/spec_plan.rs:653-655 )。实测 gpt-5.5 的第一次请求里确实没有 mcp__bughunt;换一个目录里没有的模型名,mcp__bughunt 就直接出现在 tools 里。注意分发阶段并不检查工具是不是已经被搜到:假模型直接调 mcp__bughunt 下的工具,照样执行。

调用路径上的错误

MCP handler 先把 arguments 解析成 JSON。解析失败不调 server,直接构造一个错误结果( core/src/mcp_tool_call.rs:145-160 ),模型看到 err: EOF while parsing a value at line 1 column 28。只做 JSON 解析,不按 schema 校验。

server 没起来(命令写错、依赖装不上)时,默认只是它的工具没注册,模型调用得到 unsupported call: mcp__bughuntlist_injected_bugs,和调一个不存在的工具没有区别,启动失败的原因只在 Codex 日志里。设了 required = true,codex exec 在发第一次请求前就失败退出( config/src/mcp_types.rs:238-242 ;注释还提到 startup_readiness = "catalog" 时有效的缓存目录也能满足启动)。

审批

每次 MCP 调用前,handler 判断要不要审批( core/src/mcp_tool_call.rs:2472-2503 ):

fn requires_mcp_tool_approval(annotations: Option<&ToolAnnotations>) -> bool {
    let destructive_hint = annotations.and_then(|annotations| annotations.destructive_hint);
    if destructive_hint == Some(true) {
        return true;
    }

    let read_only_hint = annotations
        .and_then(|annotations| annotations.read_only_hint)
        .unwrap_or(false);
    if read_only_hint {
        return false;
    }

    destructive_hint.unwrap_or(true)
        || annotations
            .and_then(|annotations| annotations.open_world_hint)
            .unwrap_or(true)
}

fn requires_mcp_tool_approval_for_mode(
    annotations: Option<&ToolAnnotations>,
    approval_mode: AppToolApproval,
) -> bool {
    match approval_mode {
        AppToolApproval::Auto => requires_mcp_tool_approval(annotations),
        AppToolApproval::Prompt => true,
        AppToolApproval::Writes => !annotations
            .and_then(|annotations| annotations.read_only_hint)
            .unwrap_or(false),
        AppToolApproval::Approve => false,
    }
}

Option 是"可能没有值”,unwrap_or(x) 是"没有值时当作 x"。读下来:默认的 auto 模式里,标了 destructiveHint: true 要审批,标了 readOnlyHint: true 免审批,什么都没标的工具当作有破坏性、要审批(缺省值都取了保守的 true)。prompt 总是审批,writes 只有只读工具免审批,approve 从不审批。

需要审批而审批策略是 never 时,直接拒绝( core/src/mcp_tool_call.rs:1626-1630 ):

    if *approval_policy == AskForApproval::Never {
        return ReviewDecision::denied(
            "MCP tool call requires approval, but approval policy is never",
        );
    }

而 codex exec 默认把审批策略设成 never( exec/src/lib.rs:574-576 ,注释里说如果审批人是自动审核会再重算)。两者叠起来就是 headless 跑评测时的坑:bughunt 的工具都没写 annotations,实测(codex exec -s workspace-write)模型调 list_injected_bugs 收到的是:

Wall time: 0.0017 seconds
Output:
MCP tool call requires approval, but approval policy is never

也有自动批准的情况,条件写在 codex-mcp/src/mcp/mod.rs:88-110 :工具的 approval_mode = "approve" 直接通过;否则审批策略必须是 never,并且 permission profile 是 Disabled(danger-full-access 映射到它, protocol/src/models.rs:288 )或 External,或者是 Managed 但可写全盘。打开 strict auto review 时这段自动批准整个跳过( core/src/mcp_tool_call.rs:1521-1531 )。为了跑评测放开整个磁盘不划算,更合适的做法是让 server 给只读工具标上 readOnlyHint(第 11 章的 bughunt 没标),或者在配置里对具体工具设 approval_mode = "approve"。

Codex 能不能当 MCP server?

在这个 commit 里不能。workspace 里没有 mcp-server 这个 crate,CLI 的子命令列表里只有 mcp,说明是 “Manage external MCP servers for Codex”( cli/src/main.rs:143-239 );本机 0.159.2 的 codex --help 也没有 mcp-server。对外的程序化集成走 codex app-server(JSON-RPC)。早期版本里的 codex mcp-server 是什么时候去掉的,本地是浅克隆看不到历史,没有核实。

和第 2 周对照

第 2 周(第 8、9、11 章)Codex
工具定义和实现TOOLS 列表 + TOOL_IMPLS 字典,靠名字对上同一个 ToolExecutor 对象的 spec() 和 handle()
参数校验自己用 jsonschema / pydantic 校验MCP 参数只解析 JSON,strict: false;校验交给 server
错误回传tool_result 带 is_error: true只有文字,success 是内部元数据
未知工具KeyError 被 except 接住,回 is_errorunsupported call: X,turn 继续
多个调用一个个串行执行读写锁:声明可并行的同时跑
输出截断read()[:20_000] 只留开头按 token 预算保头保尾,头部写原始 token 数
MCP自己写 server,用 Inspector 和 SDK client 调内置 client:命名空间、超时、审批、延迟加载

动手 1:把 bughunt 挂进 Codex

最小配置(路径换成你自己的,下同):

# ~/.codex/config.toml
[mcp_servers.bughunt]
command = "python3.13"
args = ["/path/to/week02_Agent原理/code/mcp_server_demo/bughunt_server.py"]
# 以下可选
startup_timeout_sec = 30          # 默认就是 30
tool_timeout_sec = 10             # 默认 300;replay_steps 要 2 秒
default_tools_approval_mode = "approve"   # headless 评测时免审批;或者按工具单独设
# supports_parallel_tool_calls = true     # 确认 server 能并发再开

[mcp_servers.bughunt.tools.submit_bug_report]
approval_mode = "prompt"          # 写操作仍然每次问

也可以用命令添加。为了不碰自己的 ~/.codex,下面用一个临时的 CODEX_HOME(这些命令都不调模型):

export CODEX_HOME=/tmp/laq15_home && mkdir -p $CODEX_HOME
codex mcp add bughunt -- python3.13 /path/to/bughunt_server.py
cat $CODEX_HOME/config.toml
codex mcp get bughunt --json

本机输出(路径已替换):

Added global MCP server 'bughunt'.
[mcp_servers.bughunt]
command = "python3.13"
args = ["/path/to/bughunt_server.py"]
{
  "name": "bughunt",
  "enabled": true,
  "disabled_reason": null,
  "transport": {
    "type": "stdio",
    "command": "python3.13",
    "args": [
      "/path/to/bughunt_server.py"
    ],
    "env": null,
    "env_vars": [],
    "cwd": null
  },
  "enabled_tools": null,
  "disabled_tools": null,
  "startup_timeout_sec": null,
  "tool_timeout_sec": null
}

几个配置错误的实际表现:

写法结果
codex mcp add "bug hunt" -- ...Error: invalid server name 'bug hunt' (use letters, numbers, '-', '_', ':', '@', '/', '.'),退出码 1
同时写 command 和 urlurl is not supported for stdio / in `mcp_servers.x`,配置加载失败
把 tool_timeout_sec 写成 tool_timeout = 5静默忽略,codex mcp get 显示 "tool_timeout_sec": null,超时仍是 300 秒
同上,加 codex exec --strict-configunknown configuration field `mcp_servers.bughunt.tool_timeout` ,退出码 1

字段拼错被静默忽略这一条,原因是原始配置结构体只在生成 JSON Schema 时声明了 deny_unknown_fields,serde 反序列化时并不拒绝未知字段( config/src/mcp_types.rs:365-435 );--strict-config 是专门加的开关( exec/src/cli.rs:20-22 )。CI 里跑 codex exec 时建议加上。

本机有一个环境问题:本机的 codex 启动器在 Rosetta(x86_64)下运行,子进程也是 x86_64,而 python3.13 里装的 pydantic_core 是 arm64 的,server 起不来。实验脚本里改用 command = "/usr/bin/arch"、args = ["-arm64", "python3.13", ...]。一般机器不需要这样写。

动手 2:畸形工具调用的测试用例

要测"模型发来一个畸形调用时 Codex 怎么处理",不需要真实模型,也不该用真实模型:你没法让模型稳定地发出一个截断的 JSON。做法和 Codex 自己的集成测试一样(见第 18 章「不调真实模型,怎么测一个 Agent?」):起一个假的 Responses API,按脚本回放 SSE 事件,然后断言 Codex 发来的第二次请求,里面每个 call_id 的 output 就是模型下一轮真正看到的东西。

脚本在学习目录的 week03_Codex源码/code/codex_tools_mock/run_codex_mock.py。核心是三步:

# 1. 模型的"回复"是写死的 SSE 事件:一个 turn 里发出若干工具调用
def function_call(call_id, name, arguments, namespace=None):
    item = {"type": "function_call", "call_id": call_id, "name": name, "arguments": arguments}
    if namespace:
        item["namespace"] = namespace
    return {"type": "response.output_item.done", "item": item}

# 2. 临时 CODEX_HOME 里把模型服务指向本地假服务(不会碰 ~/.codex)
#    [model_providers.mock] base_url = "http://127.0.0.1:<port>/v1"  wire_api = "responses"
#    然后:codex exec --skip-git-repo-check -s workspace-write --json go   (stdin 接 /dev/null)

# 3. 从第二次请求里取出每个 call_id 的 output
def outputs_of(request):
    result = {}
    for item in request["input"]:
        if item["type"] in ("function_call_output", "custom_tool_call_output"):
            out = item["output"]
            result[item["call_id"]] = out if isinstance(out, str) else "\n".join(
                part.get("text", "") for part in out)
    return result

运行 python3 run_codex_mock.py(跑全部 10 个场景;加参数只跑名字匹配的场景),完整输出在同目录的 transcript.txt,每次请求体存在 requests/ 里。用例和本机实测结果:

#畸形调用关注点模型看到的 output(实测)
1调不存在的 delete_all_bugs未知工具unsupported call: delete_all_bugs
2exec_command 参数写成 {"command": "echo hi"}缺必填字段failed to parse function arguments: missing field `cmd` at line 1 column 21
3list_injected_bugs 传 module: "payment"server 业务错误Error executing tool list_injected_bugs: 未知模块 'payment',可选:cart, login, search
4submit_bug_report 的 severity: "critical"枚举越界,Codex 不校验Error executing tool submit_bug_report: 1 validation error ... Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='critical', ...]
5MCP arguments 是截断的 JSON {"module": "cart", "title":JSON 解析失败,不调 servererr: EOF while parsing a value at line 1 column 28
6replay_steps(2 秒),tool_timeout_sec = 1超时tool call failed for `bughunt/replay_steps` … timed out awaiting tools/call after 1000ms
7exec_command 输出 200,001 字节截断头尾各 20,000,…40001 tokens truncated…
8apply_patch 补丁缺首行本地补丁解析apply_patch verification failed: invalid patch: The first line of the patch must be '*** Begin Patch'
9apply_patch 修改不存在的文件执行前校验apply_patch verification failed: Failed to read file to update .../notes/missing.md: No such file or directory (os error 2)
10用 function_call 形态调 apply_patch;用 custom_tool_call 形态调 exec_commandpayload 类型不匹配(Fatal)两个都是 aborted;turn 正常结束,退出码 0
11默认审批模式、headless审批MCP tool call requires approval, but approval policy is never
12server 启动失败(脚本路径不存在)默认非 requiredunsupported call: mcp__bughuntlist_injected_bugs
13同上,required = true启动即失败退出码 1,0 次请求,required MCP servers failed to initialize: bughunt: ...
14三个 replay_steps 并发并行门默认请求间隔 6.06 秒,supports_parallel_tool_calls = true 时 2.02 秒
15exec、exec、replay_steps、exec(各 2 秒)写锁后面的读锁会不会插队请求间隔 6.12 秒,没有插队

(MCP 场景里 3–6、14 配了 default_tools_approval_mode = "approve",否则都会停在第 11 行的审批拒绝上。第 3–6 行的 output 前面还有 Wall time: ... seconds / Output: 两行头。)

设计这组用例时值得注意的几点:

  1. 断言请求体,不断言模型的反应。被测对象是 Codex 的分发逻辑,模型只是输入源。断言写成"第二次请求里 call_id c5 的 output 包含 EOF while parsing",而不是"模型最后修好了参数"。
  2. 每个 call_id 都要有 output。这本身就是一条断言:少一条,下一次请求在真实 API 上就会被拒。
  3. 耗时类的断言看请求间隔,不看 Wall time。场景 14 里串行时每条 Wall time 也只有 2 秒。
  4. 别逐字比对会抖动的文字,固定构建类型。场景 6 超时文字里的时长是剩余超时:binding 层先算出截止时间,把剩下的时间传给 call_tool( codex-mcp/src/binding.rs:342-356 ),rmcp client 用这个时长构造 Timeout( rmcp-client/src/rmcp_client.rs:1493-1500 ),再按 "timed out awaiting {label} after {duration:.0?}" 格式化( rmcp-client/src/rmcp_client.rs:289-290 )。{:.0?} 对不到 1 秒的 Duration 用毫秒显示并四舍五入:999.99ms 显示成 1000ms,999.4ms 显示成 999ms(用 rustc 实测)。所以 1 秒超时会显示成 after 1000ms 或 after 999ms,而不是 after 1s,断言写成"包含 timed out awaiting tools/call“比逐字比对更稳。场景 10 在 debug 构建会 panic,CI 要写清楚用哪种构建。
  5. 配置错误也是用例。字段拼错(tool_timeout)、server 起不来、名字不合法,它们的表现差别很大,有的报错、有的静默。

常见错误说法

  • “Codex 会按工具的 schema 校验参数”:MCP 参数只做 JSON 解析,strict 一律 false;枚举越界照样发给 server。
  • “工具出错时,模型会收到一个 is_error 标志”:success 是内部元数据,不序列化。模型只看到文字。
  • “Fatal 错误会终止 turn”:工具任务返回的 Fatal 在 release 构建里只记日志,模型收到 aborted,turn 继续;只有 debug 构建会 panic。
  • “请求里 parallel_tool_calls: true,工具就会并行执行”:那只是允许模型一次发多个调用。本地有一把读写锁,默认不能并行,MCP 工具也一样。
  • “Codex 截断输出就是保留前 N 个字符”:保留开头和结尾各一半,中间换成 …N tokens truncated…,头部写原始 token 数。
  • “MCP 工具名是 mcp__server__tool,平铺在 tools 里”:现在是一个 mcp__<server> 命名空间工具,里面放原名的工具;支持搜索的模型还会把它延迟到 tool_search 后面。
  • "codex exec 不会审批,所以 MCP 工具直接执行”:exec 把审批策略设成 never,没标注的 MCP 工具在默认 auto 模式下需要审批,于是直接被拒(工具设了 approval_mode = "approve",或者用 danger-full-access 这类可写全盘的权限时例外)。
  • "codex mcp-server 可以把 Codex 当 MCP server 用":这个 commit 里没有这个子命令。

下一讲:工具输出、对话历史越积越多,上下文快满了,Codex 怎么办?