第 15 章
第 3 周:Codex 怎么定义、分发和执行工具?
读 OpenAI Codex 源码里的工具系统和 MCP client:ToolSpec 和 ToolExecutor 怎么把定义和实现绑在一起,schema 只保留哪些字段,apply_patch 为什么用 Lark 语法而不用 JSON,一次调用经过哪些关卡,哪些工具能并行,输出怎么保头保尾截断,MCP server 怎么接进来(命名空间、超时、审批)。用一个本地假模型服务实测畸形工具调用,不调真实模型。一段讲解视频,一个工具调用分发模拟器。
第 2 周我们手写的 agent loop 里,工具就是 TOOLS 列表加 TOOL_IMPLS 字典:模型说调哪个,就从字典里取函数执行,出错就回一个 is_error。第 11 章又写了一个 MCP server(bughunt),让任何 Agent 都能插上用。这一章看一个生产级 Agent 是怎么做同一件事的:Codex 怎么把工具定义发给模型、模型的调用怎么一路分发到执行、哪些能并行、输出太长怎么办、外接的 MCP server 怎么接进来。最后把第 2 周的 bughunt 挂进 Codex,用一个假的模型服务发一组畸形调用,看 Codex 实际回给模型的是什么。
源码基于 openai/codex commit
7993248
(2026-10-01),下文路径都相对 codex-rs/。实验用本机 codex-cli 0.159.2(macOS),模型换成本地假服务,不调真实模型、不消耗额度。
讲解视频
互动演示
上面是工具调用分发模拟器:选一个模型发来的调用(正常的、缺参数的、补丁格式错的、未知工具、MCP 参数截断、MCP 超时……),单步看它在哪一关停下、回填给模型的是什么。可以切换运行方式(codex exec 还是交互式)、MCP 审批模式、超时和模型的截断策略,截断数字按源码公式实时计算,错误文字取自本机实测。下面是并行门时间线:按顺序添加几个调用,看读写锁怎么排队。页面底部有自动判分的练习。
一张图
模型的流式输出 response.output_item.done
│ function_call / custom_tool_call
▼
ToolRouter::build_tool_call → ToolCall { tool_name, payload }
│ tokio::spawn,按出现顺序放进 FuturesOrdered
▼
① 并行门 可并行 → 读锁;其余(含未知工具)→ 写锁
② 查注册表 找不到 → "unsupported call: <名字>"
③ payload 类型 function 工具收到 custom 调用 → Fatal
④ PreToolUse hook 可拦截、可改写参数
⑤ handler 解析参数 → 审批 → 执行
⑥ PostToolUse hook
⑦ 格式化 + 截断 → function_call_output / custom_tool_call_output
│ 按调用顺序写进历史
▼
下一次请求的 input
无论在哪一关停下,这个 call_id 都会有一条 output 回给模型,turn 继续。下面一层层拆开。
工具 = spec + handler
每个工具都实现 ToolExecutor 这个 trait(
tools/src/tool_executor.rs:106-130
,节选):
pub trait ToolExecutor<Invocation>: Send + Sync {
/// The concrete tool name handled by this runtime instance.
fn tool_name(&self) -> ToolName;
fn spec(&self) -> ToolSpec;
/// The preferred exposure before the host applies step-specific policy.
fn exposure(&self) -> ToolExposure {
ToolExposure::Direct
}
// ...
fn supports_parallel_tool_calls(&self) -> bool {
false
}
/// Handles one invocation without retaining capabilities borrowed by the host.
fn handle<'a>(&'a self, invocation: Invocation) -> ToolExecutorFuture<'a>
where
Invocation: 'a;
}
不懂 Rust 也能读:trait 相当于 Python 的抽象基类,fn 是方法,&self 是 self,-> X 是返回类型。带函数体的方法(exposure、supports_parallel_tool_calls)是默认实现,具体工具可以覆盖。
tool_name():注册表里的键,由"命名空间 + 名字"组成。spec():发给模型的工具定义。exposure():默认Direct,直接出现在请求的tools里;还可以是Deferred(藏在tool_search后面,模型搜到才加载)、Hidden等,一共 6 种(另外三种是DeferredModelOnly、DirectModelOnly、CodeModeOnly,tools/src/tool_executor.rs:49-99)。supports_parallel_tool_calls():默认 false。想和别的调用同时跑,工具得自己声明。handle():真正执行,返回一个异步的 future。
和第 2 周对比:我们的 TOOLS 和 TOOL_IMPLS 是两份东西,靠名字字符串对上,加了定义忘了加实现,只有模型真调到时才 KeyError。Codex 把两者放在同一个对象上,注册时就配好了。“默认不能并行"也是一个设计选择:新工具不用考虑并发安全,出错的代价是慢,而不是两个调用同时改坏一个文件。
ToolSpec 是一个枚举(
tools/src/tool_spec.rs:20-56
),对应 Responses API 里五种工具:
| 变体 | 序列化后的 type | 谁在用 |
|---|---|---|
Function | function | exec_command、view_image 等,参数是 JSON Schema |
Namespace | namespace | 一组工具的容器。每个 MCP server 一个,例如 mcp__bughunt |
Freeform | custom | apply_patch:输入是纯文本,格式由 Lark 语法描述 |
ToolSearch | tool_search | 搜索延迟加载的工具 |
WebSearch | web_search | 服务端托管的搜索,不经过本地分发 |
每个 turn 开始前,build_tool_router(
core/src/tools/spec_plan.rs:123-188
)按固定顺序把工具装进注册表:内置工具 → MCP 工具(再按策略决定直接列出还是延迟)→ 扩展工具 → 动态工具 → 托管工具,最后生成发给模型的列表。
schema:只保留一个子集,strict 一律 false
Codex 的 JsonSchema 是个普通结构体(
tools/src/json_schema/types.rs:35-75
),只有 type、description、encrypted、enum、items、minItems、properties、required、additionalProperties、anyOf / oneOf / allOf、$ref / $defs / definitions 这些字段。外来的 schema(MCP server、动态工具)先经过 sanitize_json_schema(
tools/src/json_schema.rs:71-80
):const 改写成单值 enum,缺 type 的按出现的关键字推断。结构体里没有的字段,反序列化时直接丢掉。所以数组的 minItems 会保留,minimum、pattern、maxLength 会丢。源码测试里 {"minimum": 1} 解析完只剩 {"type": "number"}(
tools/src/json_schema_tests.rs:198-214
)。MCP 工具的 schema 如果没写 properties,会补一个空对象(
tools/src/mcp_tool.rs:45-55
)。
拿第 2 周的 bughunt 实测。Python SDK 生成的 submit_bug_report 输入 schema(title 字段是 SDK 自动加的):
{"type": "object", "title": "submit_bug_reportArguments",
"properties": {
"module": {"title": "Module", "type": "string"},
"title": {"title": "Title", "type": "string"},
"steps": {"title": "Steps", "type": "array", "items": {"type": "string"}},
"severity": {"enum": ["low", "medium", "high"], "title": "Severity", "type": "string"}},
"required": ["module", "title", "steps", "severity"]}
Codex 发给模型的那一项(摘自 mock_model_tools 场景,即模型 slug 设成未知的 mock-model 时第一次请求里的 mcp__bughunt 命名空间,节选,省略 description。默认的 gpt-5.5 会把 MCP 工具延迟到 tool_search 后面,第一次请求里看不到,见下文"延迟加载”):
{"type": "function", "name": "submit_bug_report", "strict": false,
"parameters": {"type": "object",
"properties": {"module": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"steps": {"type": "array", "items": {"type": "string"}},
"title": {"type": "string"}},
"required": ["module", "title", "steps", "severity"]}}
三处变化:title 被丢掉(list_injected_bugs 参数里的 "default": null 也一样);属性按字母序排了(properties 是 BTreeMap,有序映射);strict 是 false。最后一点写死在代码里:MCP 和动态工具转换成 Responses API 工具时一律 strict: false(
tools/src/responses_api.rs:164-173
),内置的 exec_command 也是 strict: false(
core/src/tools/handlers/shell_spec.rs:106
)。
这意味着两件事。第一,服务端不保证模型给的参数符合 schema;第二,Codex 本地也不按 schema 校验 MCP 参数,只做 JSON 解析。实测模型把 severity 写成 "critical",Codex 原样转发,是 bughunt 的 pydantic 拒绝的。所以第 9 章「模型怎么’调用’一个函数?」说的"边界处一律校验",在 Codex 体系里落在 MCP server 一侧:minimum、pattern 这类约束模型根本看不到,server 必须自己校验,并把错误写到模型能看懂。
schema 太大还会被压缩:单个 MCP 工具的输入 schema 超过 5,000 字节(
tools/src/json_schema/compaction.rs:15
,可用 tool_input_schema_max_bytes 按 server 调整,
config/src/mcp_types.rs:252-255
)时,依次做四轮越来越有损的压缩:去掉描述、去掉 $defs、折叠深层对象、修剪组合关键字(anyOf 等),每轮之前先看是否已经够小(
tools/src/json_schema/compaction.rs:18-37
)。源码注释写明这是 “best-effort rather than a hard cap”,四轮做完仍可能超限。
内置工具
本机 0.159.2 用 gpt-5.5 这个模型 slug 指向假服务,第一次请求里实际发出的工具:
| 工具 | 类型 | 能否并行 | 做什么 |
|---|---|---|---|
exec_command | function | 是 | 在 PTY 里跑命令。必填 cmd,可选 workdir、tty、yield_time_ms、max_output_tokens 等;没跑完的返回 session ID |
write_stdin | function | 是 | 往还在跑的 session 写输入、取新输出 |
apply_patch | custom | 否 | 改文件 |
view_image | function | 是 | 把本地图片放进上下文 |
list_mcp_resources 等 3 个 | function | 是 | 读 MCP resource |
request_user_input | function | 否 | 向用户提问 |
get_goal / create_goal / update_goal | function | 否 | 目标管理 |
tool_search | tool_search | 是 | 搜索延迟加载的工具 |
web_search | web_search | — | 服务端执行 |
“能否并行"一栏来自各个 handler 的 supports_parallel_tool_calls,例如 exec_command 在
core/src/tools/handlers/unified_exec/exec_command.rs:142-144
。没有单独的老 shell 工具了。工具清单随模型元数据和 feature flag 变化:把模型 slug 换成一个内置目录里没有的名字,请求里就没有 apply_patch 和 tool_search,多了 multi_agent_v1 命名空间。不同版本、不同模型请以实际请求为准。
exec_command 的 schema 写在
core/src/tools/handlers/shell_spec.rs:24-115
:required: ["cmd"],additionalProperties: false,max_output_tokens 的描述是 “Output token budget. Defaults to 10000 tokens; larger requests may be capped by policy."(
core/src/tools/handlers/shell_spec.rs:61
,实测请求体里也是这句)。
apply_patch:不用 JSON,用语法
apply_patch 是一个 custom 工具(
core/src/tools/handlers/apply_patch_spec.rs:5-28
)。它的 format 是 {"type": "grammar", "syntax": "lark", "definition": <语法全文>},描述里写着 “This is a FREEFORM tool, so do not wrap the patch in JSON."。语法全文(
core/assets/tools/apply_patch.lark:1-19
):
start: begin_patch hunk+ end_patch
begin_patch: "*** Begin Patch" LF
end_patch: "*** End Patch" LF?
hunk: add_hunk | delete_hunk | update_hunk
add_hunk: "*** Add File: " filename LF add_line+
delete_hunk: "*** Delete File: " filename LF
update_hunk: "*** Update File: " filename LF change_move? change?
filename: /(.+)/
add_line: "+" /(.*)/ LF -> line
change_move: "*** Move to: " filename LF
change: (change_context | change_line)+ eof_line?
change_context: ("@@" | "@@ " /(.+)/) LF
change_line: ("+" | "-" | " ") /(.*)/ LF
eof_line: "*** End of File" LF
%import common.LF
读法:一个补丁以 *** Begin Patch 开头、*** End Patch 结尾,中间是一个或多个 hunk。hunk 有三种:新建文件(之后每行以 + 开头)、删除文件、修改文件(可选改名,然后是若干段改动;每段以 @@ 加一行定位上下文开头,再跟 + / - / 空格开头的行)。例如:
*** Begin Patch
*** Update File: src/cart.py
@@ def checkout(cart):
- if cart.qty >= 0:
+ if cart.qty > 0:
submit(cart)
*** End Patch
为什么不用 JSON 参数:补丁里全是换行、引号、反斜杠,塞进 JSON 字符串要多一层转义,模型很容易在转义上出错。OpenAI 的函数调用文档把 custom 工具描述为输入可以是不受约束的自由文本,也可以用 Lark 或正则语法约束输出格式( OpenAI 函数调用指南 、 API 参考 CustomToolInputFormat )。也就是说格式由 API 一侧约束;Codex 本地仍然再解析一遍。所以假服务直接塞一个坏补丁(绕过了语法),得到的是本地解析器的报错:
apply_patch verification failed: invalid patch: The first line of the patch must be '*** Begin Patch'
apply_patch 的 handler 只接受 custom 形态的调用(
core/src/tools/handlers/apply_patch.rs:407-409
)。如果有个 provider 不支持 custom 工具、模型用 function_call 的形态调它,会怎样?下一节的"Fatal”。
分发:一次调用经过哪些关卡
流里每出现一个完整的输出项,ToolRouter::build_tool_call(
core/src/tools/router.rs:248-300
)把它变成内部的 ToolCall:function_call 变成 ToolPayload::Function { arguments }(arguments 此时还是字符串),custom_tool_call 变成 ToolPayload::Custom { input },namespace 和 name 合成 ToolName。然后 ToolCallRuntime 把它 spawn 成一个异步任务(
core/src/tools/parallel.rs:196-245
),任务里依次:
- 并行门:可并行的拿读锁,其余拿写锁(下一节细讲)。注意它在查注册表之前。查"能否并行"时工具还不存在,
tool_supports_parallel的unwrap_or(false)让未知工具也拿写锁(core/src/tools/router.rs:237-241)。 - 查注册表(
core/src/tools/registry.rs:551-571):找不到就返回RespondToModel,文字由unsupported_tool_call_message生成:custom 调用是unsupported custom tool call: X,其余是unsupported call: X(core/src/tools/registry.rs:852-857)。 - payload 类型检查(
core/src/tools/registry.rs:584-600):matches_kind不通过就是Fatal("tool X invoked with incompatible payload")。默认的matches_kind只接受 Function 和 ToolSearch(core/src/tools/registry.rs:85-90)。 - PreToolUse hook(
core/src/tools/registry.rs:602-654):用户配的 hook 可以拦下这次调用(回一条说明给模型),也可以改写参数。 - handler:解析参数、审批、执行。参数解析失败是
RespondToModel("failed to parse function arguments: ...")(core/src/tools/handlers/mod.rs:86-93)。 - PostToolUse hook(
core/src/tools/registry.rs:709-771):成功的结果还能被 hook 拦下或附加反馈。 - 格式化、截断、回填:结果变成
function_call_output或custom_tool_call_output。所有调用任务放在一个FuturesOrdered里(core/src/session/turn.rs:2589),所以即使并行执行,写回历史的顺序也和模型发出的顺序一致。
错误分两类,但模型只看到文字
工具出错时返回 FunctionCallError(
tools/src/function_call_error.rs:4-10
):
pub enum FunctionCallError {
#[error("{0}")]
RespondToModel(String),
#[error("Fatal error: {0}")]
Fatal(String),
}
RespondToModel 的意思是"把这段文字当作工具输出还给模型,让它自己改”;Fatal 是"不该发生的内部错误”。#[error(...)] 是错误转成字符串时的格式。
RespondToModel 和其他非 Fatal 错误,在 failure_response(
core/src/tools/parallel.rs:303-328
)里变成一条普通的 output,内部标记 success: false。但这个标记不发给模型(
protocol/src/models.rs:2175-2183
、
protocol/src/models.rs:2253-2263
):
/// `body` serializes directly as the wire value for `function_call_output.output`.
/// `success` remains internal metadata for downstream handling.
pub struct FunctionCallOutputPayload {
pub body: FunctionCallOutputBody,
pub success: Option<bool>,
}
impl Serialize for FunctionCallOutputPayload {
fn serialize<S>(&self, serializer: S) -> Result<S::Ok, S::Error>
// ...
match &self.body {
FunctionCallOutputBody::Text(content) => serializer.serialize_str(content),
FunctionCallOutputBody::ContentItems(items) => items.serialize(serializer),
}
自定义的 Serialize 只输出 body。对比第 2 周:Claude API 的 tool_result 有 is_error 字段,MCP 的工具结果有 isError 字段;到了 Codex 发给 Responses API 的请求里,两者都只剩文字。实测 bughunt 返回 isError: true 的结果,模型看到的是 Error executing tool submit_bug_report: ... 这段文本,没有任何标志。错误文字写得清不清楚,直接决定模型能不能自己改对。
Fatal 并不会终止 turn
直觉上 Fatal 应该让整个 turn 失败。handle_output_item_done 里确实留了一个分支:如果 build_tool_call 返回 Fatal,就转成 CodexErr::Fatal 结束 turn(
core/src/stream_events_utils.rs:427-429
)。但这个 commit 里 build_tool_call 只会返回 RespondToModel,不会返回 Fatal(
core/src/tools/router.rs:248-300
)。真正会出现的 Fatal 来自 registry 里的类型检查,它发生在已经 spawn 出去的工具任务里,走的是另一条路。(第 14 章讲 agent loop 时把 Fatal 归入"终止",说的是采样请求出错时 CodexErr 的重试分类,和这里工具任务返回的 FunctionCallError::Fatal 不是一回事。)turn 收尾时,drain_in_flight 逐个等待在途的工具任务(
core/src/session/turn.rs:2472-2500
):
Err(err) => {
error_or_panic(format!("in-flight tool future failed during drain: {err}"));
}
error_or_panic(
core/src/util.rs:81-87
)在 debug 构建(cfg!(debug_assertions))里 panic,在 release 构建里只记一条 error 日志。这次调用没有产出 output,发下一次请求前,历史规范化给缺 output 的调用补一条内容为 "aborted" 的 output(
core/src/context_manager/normalize.rs:51-67
)。
实测(release 版 0.159.2):假模型用 function_call 形态调 apply_patch,日志里有 Fatal error: tool apply_patch invoked with incompatible payload,第二次请求里这个 call_id 的 output 是 aborted,turn 正常结束,codex 退出码 0。反过来用 custom_tool_call 调 exec_command 也一样:output 是 aborted,退出码 0。custom 调用缺 output 时,规范化阶段还会再调一次 error_or_panic(
core/src/context_manager/normalize.rs:87-92
),debug 构建在这里同样会 panic。
这对测试很重要:同一个用例,debug 构建会 panic,release 构建只是悄悄多一条 aborted。CI 里跑 Codex 的集成测试,要写清楚用的是哪种构建;断言"Fatal 会让 turn 失败"的测试在 release 下会失败。
并行:一把公平的读写锁
工具任务拿到这把锁才开始执行(
core/src/tools/parallel.rs:205-209
):
let guard = if supports_parallel {
Either::Left(lock.read().await)
} else {
Either::Right(lock.write().await)
};
所有调用共用一把 RwLock(
core/src/tools/parallel.rs:44-65
)。声明可并行的拿读锁,多个读锁可以同时持有;其余拿写锁,独占。Either 只是把两种锁守卫装进同一个变量。锁一直持有到 handler 执行完。
请求体里的 parallel_tool_calls 除了走 Responses Lite 的模型以外都是 true(
core/src/client.rs:1001
,实测也是 true),这只表示允许模型一次发多个调用;真正能不能同时跑,看这把锁。tokio 的 RwLock 是公平锁。Codex 的 Cargo.lock 锁定 tokio 1.52.3,它的 src/sync/rwlock.rs 类型文档写的是:“Fairness is ensured using a first-in, first-out queue for the tasks awaiting the lock; a read lock will not be given out until all write lock requests that were queued before it have been acquired and released."(
tokio 1.52.3 src/sync/rwlock.rs:41-47
)。也就是说,排在写锁后面的读锁要等写锁用完。所以调用顺序会影响总耗时:按这个先进先出模型,exec_command(2 秒)、exec_command(2 秒)、apply_patch(1 秒)、exec_command(2 秒)要 5 秒;把 apply_patch 挪到最后只要 3 秒。apply_patch 执行太快不好实测,我用一个 2 秒的 MCP 调用(写锁)代替它跑了 lock_order 场景:exec、exec、replay_steps、exec 各 2 秒,两次请求间隔 6.12 秒;如果后面的读锁能插队,应该是 4 秒左右。
哪些工具声明了可并行:exec_command、write_stdin、view_image、tool_search、三个 MCP resource 工具。MCP 工具看配置和标注(
core/src/tools/handlers/mcp.rs:148-159
):
fn supports_parallel_tool_calls(&self) -> bool {
// Correctly implemented MCP servers should tolerate parallel calls to
// tools that advertise themselves as read-only.
self.tool_info.supports_parallel_tool_calls
|| self
.tool_info
.tool
.annotations
.as_ref()
.and_then(|annotations| annotations.read_only_hint)
.unwrap_or(false)
}
即 server 配置了 supports_parallel_tool_calls = true(
config/src/mcp_types.rs:248-250
),或者工具标注了 readOnlyHint: true。注释的理由是:实现正确的 MCP server 应该能承受对只读工具的并行调用。
实测:假模型一次发三个 replay_steps(bughunt 里每次 await anyio.sleep(2))。
| 场景 | 第 1 次 → 第 2 次请求间隔 | 每条 output 的 Wall time |
|---|---|---|
| 默认配置(MCP 工具拿写锁) | 6.06 秒 | 2.0180 / 2.0089 / 2.0034 秒 |
supports_parallel_tool_calls = true | 2.02 秒 | 2.0043 / 2.0035 / 2.0026 秒 |
三个 exec_command "sleep 2" | 2.09 秒 | 1.8790 / 1.8727 / 1.8728 秒 |
串行时每条 Wall time 仍然只有 2 秒左右:它只量 handler 自己的执行时间,排队等锁的时间不算(源码也把拿到锁的时刻记作"执行开始”,
core/src/tools/parallel.rs:211-214
)。只看 output 是看不出调用被串行化的,要看请求间隔。
还有一个副作用:MCP 的审批发生在 handler 里,也就是拿着锁的时候。交互模式下一个等用户点批准的 MCP 调用(写锁)会挡住它后面的所有调用。
输出截断:保头保尾,砍中间
exec_command 的输出预算取两者中较小的(
core/src/tools/context.rs:481-489
):
fn model_output_policy(&self) -> TruncationPolicy {
let requested_policy = TruncationPolicy::Tokens(resolve_max_tokens(self.max_output_tokens));
if requested_policy.byte_budget() < self.truncation_policy.byte_budget() {
requested_policy
} else {
self.truncation_policy
}
}
requested_policy:调用参数max_output_tokens,不传时默认 10,000(core/src/unified_exec/mod.rs:79、core/src/unified_exec/mod.rs:227-229)。self.truncation_policy:模型元数据里的截断策略。内置模型目录(models-manager/models.json)里每个模型都是 10,000 tokens;找不到模型元数据时的兜底是 10,000 字节(protocol/src/openai_models.rs:1088)。- 比较的是换算成字节后的预算。token 按 4 字节估算(
utils/string/src/truncate.rs:4)。
超出预算时,预算一半给开头、一半给结尾,中间换成一个标记(
utils/string/src/truncate.rs:132-143
):
fn split_budget(budget: usize) -> (usize, usize) {
let left = budget / 2;
(left, budget - left)
}
fn format_truncation_marker(use_tokens: bool, removed_count: u64) -> String {
if use_tokens {
format!("…{removed_count} tokens truncated…")
} else {
format!("…{removed_count} chars truncated…")
}
}
全是 ASCII 时,输出 \(N\) 字节、预算 \(B\) 个 token:
$$\text{保留} = 4B\ \text{字节},\qquad \text{标记里的数} = \left\lceil \frac{N - 4B}{4} \right\rceil,\qquad \text{Original token count} = \left\lceil \frac{N}{4} \right\rceil$$实测 python3 -c "print('x'*200000)",\(N = 200{,}001\)(含末尾换行),\(B = \min(10{,}000, 10{,}000)\):保留 40,000 字节,头尾各 20,000;标记 \(\lceil 160{,}001 / 4 \rceil = 40{,}001\);Original token count \(\lceil 200{,}001/4 \rceil = 50{,}001\)。回给模型的 output(中间的 x 省略):
Chunk ID: 8b4012
Wall time: 0.0000 seconds
Process exited with code 0
Original token count: 50001
Output:
Warning: truncated output (original token count: 50001)
Total output lines: 1
xxxx…(20,000 个 x)…40001 tokens truncated…xxxx(19,999 个 x 加换行)
整条 40,209 个字符(40,213 字节,两个省略号各占 3 字节)。头部的格式见
core/src/tools/context.rs:524-548
。如果模型元数据走兜底的 10,000 字节策略,预算就只有 10,000 字节,标记变成 …N chars truncated…。exec 进程那一侧另有 1 MiB 的缓冲上限(
core/src/unified_exec/mod.rs:80
)。
MCP 工具的输出加一行 Wall time 头,按该工具配置的 output_token_limit 或模型策略截断,再乘 1.2 的"序列化余量"(
core/src/tools/context.rs:184-210
、
utils/output-truncation/src/lib.rs:14-18
)。
和第 2 周对比:我们写的是 f.read()[:20_000],只留开头。测试命令的失败摘要、编译器最后一个报错都在结尾,只留开头会把最有用的部分砍掉。Codex 还在头部写明原始 token 数,模型知道自己没看全,可以换个命令(比如 tail)去看。
MCP client:连接、命名空间、超时、审批
配置和连接
每个 server 是 config.toml 里的一个 [mcp_servers.<名字>] 表(
config/src/config_toml.rs:293-294
)。传输二选一(
config/src/mcp_types.rs:612-646
):写 command 是 stdio,Codex 把 server 当子进程拉起;写 url 是 Streamable HTTP。常用字段(
config/src/mcp_types.rs:222-300
):
| 字段 | 默认 | 作用 |
|---|---|---|
enabled | true | 关掉就不连 |
required | false | 为 true 时,server 起不来就让 codex exec 报错退出 |
startup_timeout_sec | 30 | 启动握手加首次列工具的超时 |
tool_timeout_sec | 300 | 单次工具调用的超时 |
enabled_tools / disabled_tools | 无 | 工具白名单 / 黑名单 |
supports_parallel_tool_calls | false | 这个 server 的工具都拿读锁 |
default_tools_approval_mode | auto | 这个 server 工具的默认审批模式 |
[mcp_servers.<名字>.tools.<工具>] | — | 按工具设 approval_mode、output_token_limit |
两个默认超时是常量(
codex-mcp/src/rmcp_client.rs:106-107
),没配置时套用(
codex-mcp/src/connection_manager.rs:358-365
)。调用方自己也带了超时时取两者较小的(
codex-mcp/src/binding.rs:313-318
);不过模型发起的调用传的 requested_timeout 是 None(
core/src/mcp_tool_call.rs:466
),实际生效的就是 tool_timeout_sec。
命名空间
模型看到的不是一串扁平的工具名,而是每个 server 一个 namespace 工具,里面放这个 server 的全部工具(
core/src/tools/handlers/mcp.rs:498-531
)。实测请求体里 bughunt 长这样(摘自 mock_model_tools 场景,即未知模型 slug 时的第一次请求;节选,省略各工具的 description):
{"type": "namespace", "name": "mcp__bughunt",
"description": "BugHunt-Bench 评测环境:查询注入 Bug、提交 Bug 报告。",
"tools": [{"type": "function", "name": "list_injected_bugs", "strict": false,
"parameters": {"type": "object",
"properties": {"module": {"anyOf": [{"type": "string"}, {"type": "null"}]}}}},
...]}
命名空间的描述就是 server 初始化时给的 instructions。名字的规则(
codex-mcp/src/tools.rs:105-214
、
codex-mcp/src/tools.rs:225-300
):
- server 名清洗成只含
[A-Za-z0-9_](codex-mcp/src/mcp/mod.rs:578-593),加上前缀mcp__。前缀由一个non_prefixed_mcp_tool_namesfeature 控制,默认关闭,即默认带前缀(core/src/config/mod.rs:1890-1893)。 - 清洗后撞名的,加
_和 SHA1 的前 12 位十六进制;模型可见的名字最长 128。 - 工具名本身不加前缀。调 server 时用的仍是原始工具名。
codex mcp add 对 server 名有更严的检查,只允许字母、数字和 - _ : @ / .(
cli/src/mcp_cmd.rs:1148-1157
)。
一个容易看错的地方:日志和错误文字里把命名空间和工具名直接拼接(
protocol/src/tool_name.rs:54-62
),所以你会看到 mcp__bughuntlist_injected_bugs,中间没有分隔符。hook 里用的名字则是 mcp__bughunt__list_injected_bugs(
core/src/tools/handlers/mcp.rs:120-137
)。
延迟加载
模型支持搜索工具时,MCP 工具默认是 Deferred:不出现在请求的 tools 里,而是在 tool_search 的描述中列出 server(“bughunt: BugHunt-Bench 评测环境…"),模型先搜,搜到再加载(
core/src/mcp_tool_exposure.rs:85-89
、
core/src/tools/spec_plan.rs:653-655
)。实测 gpt-5.5 的第一次请求里确实没有 mcp__bughunt;换一个目录里没有的模型名,mcp__bughunt 就直接出现在 tools 里。注意分发阶段并不检查工具是不是已经被搜到:假模型直接调 mcp__bughunt 下的工具,照样执行。
调用路径上的错误
MCP handler 先把 arguments 解析成 JSON。解析失败不调 server,直接构造一个错误结果(
core/src/mcp_tool_call.rs:145-160
),模型看到 err: EOF while parsing a value at line 1 column 28。只做 JSON 解析,不按 schema 校验。
server 没起来(命令写错、依赖装不上)时,默认只是它的工具没注册,模型调用得到 unsupported call: mcp__bughuntlist_injected_bugs,和调一个不存在的工具没有区别,启动失败的原因只在 Codex 日志里。设了 required = true,codex exec 在发第一次请求前就失败退出(
config/src/mcp_types.rs:238-242
;注释还提到 startup_readiness = "catalog" 时有效的缓存目录也能满足启动)。
审批
每次 MCP 调用前,handler 判断要不要审批(
core/src/mcp_tool_call.rs:2472-2503
):
fn requires_mcp_tool_approval(annotations: Option<&ToolAnnotations>) -> bool {
let destructive_hint = annotations.and_then(|annotations| annotations.destructive_hint);
if destructive_hint == Some(true) {
return true;
}
let read_only_hint = annotations
.and_then(|annotations| annotations.read_only_hint)
.unwrap_or(false);
if read_only_hint {
return false;
}
destructive_hint.unwrap_or(true)
|| annotations
.and_then(|annotations| annotations.open_world_hint)
.unwrap_or(true)
}
fn requires_mcp_tool_approval_for_mode(
annotations: Option<&ToolAnnotations>,
approval_mode: AppToolApproval,
) -> bool {
match approval_mode {
AppToolApproval::Auto => requires_mcp_tool_approval(annotations),
AppToolApproval::Prompt => true,
AppToolApproval::Writes => !annotations
.and_then(|annotations| annotations.read_only_hint)
.unwrap_or(false),
AppToolApproval::Approve => false,
}
}
Option 是"可能没有值”,unwrap_or(x) 是"没有值时当作 x"。读下来:默认的 auto 模式里,标了 destructiveHint: true 要审批,标了 readOnlyHint: true 免审批,什么都没标的工具当作有破坏性、要审批(缺省值都取了保守的 true)。prompt 总是审批,writes 只有只读工具免审批,approve 从不审批。
需要审批而审批策略是 never 时,直接拒绝(
core/src/mcp_tool_call.rs:1626-1630
):
if *approval_policy == AskForApproval::Never {
return ReviewDecision::denied(
"MCP tool call requires approval, but approval policy is never",
);
}
而 codex exec 默认把审批策略设成 never(
exec/src/lib.rs:574-576
,注释里说如果审批人是自动审核会再重算)。两者叠起来就是 headless 跑评测时的坑:bughunt 的工具都没写 annotations,实测(codex exec -s workspace-write)模型调 list_injected_bugs 收到的是:
Wall time: 0.0017 seconds
Output:
MCP tool call requires approval, but approval policy is never
也有自动批准的情况,条件写在
codex-mcp/src/mcp/mod.rs:88-110
:工具的 approval_mode = "approve" 直接通过;否则审批策略必须是 never,并且 permission profile 是 Disabled(danger-full-access 映射到它,
protocol/src/models.rs:288
)或 External,或者是 Managed 但可写全盘。打开 strict auto review 时这段自动批准整个跳过(
core/src/mcp_tool_call.rs:1521-1531
)。为了跑评测放开整个磁盘不划算,更合适的做法是让 server 给只读工具标上 readOnlyHint(第 11 章的 bughunt 没标),或者在配置里对具体工具设 approval_mode = "approve"。
Codex 能不能当 MCP server?
在这个 commit 里不能。workspace 里没有 mcp-server 这个 crate,CLI 的子命令列表里只有 mcp,说明是 “Manage external MCP servers for Codex”(
cli/src/main.rs:143-239
);本机 0.159.2 的 codex --help 也没有 mcp-server。对外的程序化集成走 codex app-server(JSON-RPC)。早期版本里的 codex mcp-server 是什么时候去掉的,本地是浅克隆看不到历史,没有核实。
和第 2 周对照
| 第 2 周(第 8、9、11 章) | Codex | |
|---|---|---|
| 工具定义和实现 | TOOLS 列表 + TOOL_IMPLS 字典,靠名字对上 | 同一个 ToolExecutor 对象的 spec() 和 handle() |
| 参数校验 | 自己用 jsonschema / pydantic 校验 | MCP 参数只解析 JSON,strict: false;校验交给 server |
| 错误回传 | tool_result 带 is_error: true | 只有文字,success 是内部元数据 |
| 未知工具 | KeyError 被 except 接住,回 is_error | unsupported call: X,turn 继续 |
| 多个调用 | 一个个串行执行 | 读写锁:声明可并行的同时跑 |
| 输出截断 | read()[:20_000] 只留开头 | 按 token 预算保头保尾,头部写原始 token 数 |
| MCP | 自己写 server,用 Inspector 和 SDK client 调 | 内置 client:命名空间、超时、审批、延迟加载 |
动手 1:把 bughunt 挂进 Codex
最小配置(路径换成你自己的,下同):
# ~/.codex/config.toml
[mcp_servers.bughunt]
command = "python3.13"
args = ["/path/to/week02_Agent原理/code/mcp_server_demo/bughunt_server.py"]
# 以下可选
startup_timeout_sec = 30 # 默认就是 30
tool_timeout_sec = 10 # 默认 300;replay_steps 要 2 秒
default_tools_approval_mode = "approve" # headless 评测时免审批;或者按工具单独设
# supports_parallel_tool_calls = true # 确认 server 能并发再开
[mcp_servers.bughunt.tools.submit_bug_report]
approval_mode = "prompt" # 写操作仍然每次问
也可以用命令添加。为了不碰自己的 ~/.codex,下面用一个临时的 CODEX_HOME(这些命令都不调模型):
export CODEX_HOME=/tmp/laq15_home && mkdir -p $CODEX_HOME
codex mcp add bughunt -- python3.13 /path/to/bughunt_server.py
cat $CODEX_HOME/config.toml
codex mcp get bughunt --json
本机输出(路径已替换):
Added global MCP server 'bughunt'.
[mcp_servers.bughunt]
command = "python3.13"
args = ["/path/to/bughunt_server.py"]
{
"name": "bughunt",
"enabled": true,
"disabled_reason": null,
"transport": {
"type": "stdio",
"command": "python3.13",
"args": [
"/path/to/bughunt_server.py"
],
"env": null,
"env_vars": [],
"cwd": null
},
"enabled_tools": null,
"disabled_tools": null,
"startup_timeout_sec": null,
"tool_timeout_sec": null
}
几个配置错误的实际表现:
| 写法 | 结果 |
|---|---|
codex mcp add "bug hunt" -- ... | Error: invalid server name 'bug hunt' (use letters, numbers, '-', '_', ':', '@', '/', '.'),退出码 1 |
同时写 command 和 url | url is not supported for stdio / in `mcp_servers.x`,配置加载失败 |
把 tool_timeout_sec 写成 tool_timeout = 5 | 静默忽略,codex mcp get 显示 "tool_timeout_sec": null,超时仍是 300 秒 |
同上,加 codex exec --strict-config | unknown configuration field `mcp_servers.bughunt.tool_timeout` ,退出码 1 |
字段拼错被静默忽略这一条,原因是原始配置结构体只在生成 JSON Schema 时声明了 deny_unknown_fields,serde 反序列化时并不拒绝未知字段(
config/src/mcp_types.rs:365-435
);--strict-config 是专门加的开关(
exec/src/cli.rs:20-22
)。CI 里跑 codex exec 时建议加上。
本机有一个环境问题:本机的 codex 启动器在 Rosetta(x86_64)下运行,子进程也是 x86_64,而 python3.13 里装的 pydantic_core 是 arm64 的,server 起不来。实验脚本里改用 command = "/usr/bin/arch"、args = ["-arm64", "python3.13", ...]。一般机器不需要这样写。
动手 2:畸形工具调用的测试用例
要测"模型发来一个畸形调用时 Codex 怎么处理",不需要真实模型,也不该用真实模型:你没法让模型稳定地发出一个截断的 JSON。做法和 Codex 自己的集成测试一样(见第 18 章「不调真实模型,怎么测一个 Agent?」):起一个假的 Responses API,按脚本回放 SSE 事件,然后断言 Codex 发来的第二次请求,里面每个 call_id 的 output 就是模型下一轮真正看到的东西。
脚本在学习目录的 week03_Codex源码/code/codex_tools_mock/run_codex_mock.py。核心是三步:
# 1. 模型的"回复"是写死的 SSE 事件:一个 turn 里发出若干工具调用
def function_call(call_id, name, arguments, namespace=None):
item = {"type": "function_call", "call_id": call_id, "name": name, "arguments": arguments}
if namespace:
item["namespace"] = namespace
return {"type": "response.output_item.done", "item": item}
# 2. 临时 CODEX_HOME 里把模型服务指向本地假服务(不会碰 ~/.codex)
# [model_providers.mock] base_url = "http://127.0.0.1:<port>/v1" wire_api = "responses"
# 然后:codex exec --skip-git-repo-check -s workspace-write --json go (stdin 接 /dev/null)
# 3. 从第二次请求里取出每个 call_id 的 output
def outputs_of(request):
result = {}
for item in request["input"]:
if item["type"] in ("function_call_output", "custom_tool_call_output"):
out = item["output"]
result[item["call_id"]] = out if isinstance(out, str) else "\n".join(
part.get("text", "") for part in out)
return result
运行 python3 run_codex_mock.py(跑全部 10 个场景;加参数只跑名字匹配的场景),完整输出在同目录的 transcript.txt,每次请求体存在 requests/ 里。用例和本机实测结果:
| # | 畸形调用 | 关注点 | 模型看到的 output(实测) |
|---|---|---|---|
| 1 | 调不存在的 delete_all_bugs | 未知工具 | unsupported call: delete_all_bugs |
| 2 | exec_command 参数写成 {"command": "echo hi"} | 缺必填字段 | failed to parse function arguments: missing field `cmd` at line 1 column 21 |
| 3 | list_injected_bugs 传 module: "payment" | server 业务错误 | Error executing tool list_injected_bugs: 未知模块 'payment',可选:cart, login, search |
| 4 | submit_bug_report 的 severity: "critical" | 枚举越界,Codex 不校验 | Error executing tool submit_bug_report: 1 validation error ... Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='critical', ...] |
| 5 | MCP arguments 是截断的 JSON {"module": "cart", "title": | JSON 解析失败,不调 server | err: EOF while parsing a value at line 1 column 28 |
| 6 | replay_steps(2 秒),tool_timeout_sec = 1 | 超时 | tool call failed for `bughunt/replay_steps` … timed out awaiting tools/call after 1000ms |
| 7 | exec_command 输出 200,001 字节 | 截断 | 头尾各 20,000,…40001 tokens truncated… |
| 8 | apply_patch 补丁缺首行 | 本地补丁解析 | apply_patch verification failed: invalid patch: The first line of the patch must be '*** Begin Patch' |
| 9 | apply_patch 修改不存在的文件 | 执行前校验 | apply_patch verification failed: Failed to read file to update .../notes/missing.md: No such file or directory (os error 2) |
| 10 | 用 function_call 形态调 apply_patch;用 custom_tool_call 形态调 exec_command | payload 类型不匹配(Fatal) | 两个都是 aborted;turn 正常结束,退出码 0 |
| 11 | 默认审批模式、headless | 审批 | MCP tool call requires approval, but approval policy is never |
| 12 | server 启动失败(脚本路径不存在) | 默认非 required | unsupported call: mcp__bughuntlist_injected_bugs |
| 13 | 同上,required = true | 启动即失败 | 退出码 1,0 次请求,required MCP servers failed to initialize: bughunt: ... |
| 14 | 三个 replay_steps 并发 | 并行门 | 默认请求间隔 6.06 秒,supports_parallel_tool_calls = true 时 2.02 秒 |
| 15 | exec、exec、replay_steps、exec(各 2 秒) | 写锁后面的读锁会不会插队 | 请求间隔 6.12 秒,没有插队 |
(MCP 场景里 3–6、14 配了 default_tools_approval_mode = "approve",否则都会停在第 11 行的审批拒绝上。第 3–6 行的 output 前面还有 Wall time: ... seconds / Output: 两行头。)
设计这组用例时值得注意的几点:
- 断言请求体,不断言模型的反应。被测对象是 Codex 的分发逻辑,模型只是输入源。断言写成"第二次请求里 call_id
c5的 output 包含EOF while parsing",而不是"模型最后修好了参数"。 - 每个 call_id 都要有 output。这本身就是一条断言:少一条,下一次请求在真实 API 上就会被拒。
- 耗时类的断言看请求间隔,不看 Wall time。场景 14 里串行时每条 Wall time 也只有 2 秒。
- 别逐字比对会抖动的文字,固定构建类型。场景 6 超时文字里的时长是剩余超时:binding 层先算出截止时间,把剩下的时间传给
call_tool(codex-mcp/src/binding.rs:342-356),rmcp client 用这个时长构造Timeout(rmcp-client/src/rmcp_client.rs:1493-1500),再按"timed out awaiting {label} after {duration:.0?}"格式化(rmcp-client/src/rmcp_client.rs:289-290)。{:.0?}对不到 1 秒的 Duration 用毫秒显示并四舍五入:999.99ms 显示成1000ms,999.4ms 显示成999ms(用 rustc 实测)。所以 1 秒超时会显示成after 1000ms或after 999ms,而不是after 1s,断言写成"包含timed out awaiting tools/call“比逐字比对更稳。场景 10 在 debug 构建会 panic,CI 要写清楚用哪种构建。 - 配置错误也是用例。字段拼错(
tool_timeout)、server 起不来、名字不合法,它们的表现差别很大,有的报错、有的静默。
常见错误说法
- “Codex 会按工具的 schema 校验参数”:MCP 参数只做 JSON 解析,
strict一律 false;枚举越界照样发给 server。 - “工具出错时,模型会收到一个 is_error 标志”:
success是内部元数据,不序列化。模型只看到文字。 - “Fatal 错误会终止 turn”:工具任务返回的 Fatal 在 release 构建里只记日志,模型收到
aborted,turn 继续;只有 debug 构建会 panic。 - “请求里
parallel_tool_calls: true,工具就会并行执行”:那只是允许模型一次发多个调用。本地有一把读写锁,默认不能并行,MCP 工具也一样。 - “Codex 截断输出就是保留前 N 个字符”:保留开头和结尾各一半,中间换成
…N tokens truncated…,头部写原始 token 数。 - “MCP 工具名是
mcp__server__tool,平铺在 tools 里”:现在是一个mcp__<server>命名空间工具,里面放原名的工具;支持搜索的模型还会把它延迟到tool_search后面。 - "
codex exec不会审批,所以 MCP 工具直接执行”:exec 把审批策略设成 never,没标注的 MCP 工具在默认 auto 模式下需要审批,于是直接被拒(工具设了approval_mode = "approve",或者用danger-full-access这类可写全盘的权限时例外)。 - "
codex mcp-server可以把 Codex 当 MCP server 用":这个 commit 里没有这个子命令。
下一讲:工具输出、对话历史越积越多,上下文快满了,Codex 怎么办?