Automation

MCP Server

MCP Server 用于把外部工具和知识源开放给 Codex,优先从只读能力开始。

将 Codex 作为 MCP 服务器运行

您可以将 Codex 作为 MCP 服务器运行,并从其他 MCP 客户端(例如,使用 OpenAI 代理 SDK MCP 集成 构建的代理)连接它。

要将 Codex 作为 MCP 服务器启动,您可以使用以下命令:

codex mcp-server

您可以使用 Model Context Protocol 检查器 启动 Codex MCP 服务器:

npx @modelcontextprotocol/inspector codex mcp-server

发送工具/列表请求以查看两个工具:

codex:运行 Codex 会话。接受与 Codex 配置结构匹配的配置参数。 Codex 工具具有以下属性:

物业类型描述提示(必填)字符串初始用户提示启动 Codex 对话。 CXSYNC6 令牌字符串模型生成的 shell 命令的批准策略:不受信任、on-request 和从不。 base-instructionsbase-instructions字符串要使用的指令集而不是默认指令集。配置对象覆盖 $CODEX_HOME/config.toml 中内容的单独配置设置。西德字符串会话的工作目录。如果是相对的,则根据服务器进程的当前目录进行解析。 CXSYNC9令牌布尔是否在对话中包含计划工具。型号字符串型号名称的可选覆盖(例如,o3、o4-mini)。简介字符串配置文件名称; Codex 加载 $CODEX_HOME/profile-name.config.toml 以指定默认选项。沙箱字符串沙盒模式:read-only、workspace-write 或 danger-full-access。

codex-reply:通过提供线程 ID 和提示来继续 Codex 会话。 codex-reply 工具具有以下属性:

物业类型描述提示(必填)字符串下一个用户提示继续 Codex 对话。线程 ID(必填)字符串要继续的线程的 ID。 conversationId(已弃用)字符串已弃用 threadId 的别名(为了兼容性而保留)。

在工具/调用响应中使用 StructuredContent.threadId 中的 threadId。批准提示(exec/patch)还在其参数负载中包含 threadId。

响应负载示例:

{ "structuredContent": { "threadId": "019bbb20-bff6-7130-83aa-bf45ab33250e", "content": "`ls -lah` (or `ls -alh`) — long listing, includes dotfiles, human-readable sizes." }, "content": [ { "type": "text", "text": "`ls -lah` (or `ls -alh`) — long listing, includes dotfiles, human-readable sizes." } ] }

请注意,现代 MCP 客户端通常仅报告“结构化内容”作为工具调用的结果(如果存在),尽管 Codex MCP 服务器也会返回“内容”以利于旧版 MCP 客户端。

创建 multi-agent 工作流程

Codex CLI can do far more than run ad-hoc tasks. By exposing the CLI as a Model Context Protocol (MCP) server and orchestrating it with the OpenAI Agents SDK, you can create deterministic, reviewable workflows that scale from a single agent to a complete software delivery pipeline.

本指南将介绍 OpenAI Cookbook 中展示的相同工作流程。你会:

启动 Codex CLI 作为 long-running MCP 服务器,

构建一个专注的 single-agent 工作流程来生成可玩的浏览器游戏,以及

使用 hand-offs、护栏和完整跟踪来协调 multi-agent 团队,您可以事后查看。

开始之前,请确保您拥有:

Codex CLI installed locally so the codex command is available.

Python 3.10+ 与 pip。

如果您想运行上面的 MCP 检查器示例,请使用 Node.js 18+。

本地存储的 OpenAI API 密钥。您可以在 OpenAI 仪表板 中创建或管理密钥。

为指南创建一个工作目录并将 API 密钥添加到 .env 文件中:

mkdir codex-workflows cd codex-workflows printf "OPENAI_API_KEY=sk-..." > .env

安装依赖项

代理 SDK 处理 Codex、hand-offs 和跟踪之间的编排。安装最新的 SDK 软件包:

python -m venv .venv source .venv/bin/activate pip install --upgrade openai openai-agents python-dotenv

激活虚拟环境可以使 SDK 依赖项与系统的其余部分隔离。

将 Codex CLI 初始化为 MCP 服务器

首先将 Codex CLI 转换为代理 SDK 可以调用的 MCP 服务器。服务器公开两个工具(codex() 用于启动对话,codex-reply() 用于继续对话),并在多个代理轮流中保持 Codex 处于活动状态。

创建一个名为 codex_mcp.py 的文件并添加以下内容:

import asyncio from agents import Agent, Runner from agents.mcp import MCPServerStdio async def main() -> None: async with MCPServerStdio( name="Codex CLI", params={ "command": "codex", "args": ["mcp-server"], }, client_session_timeout_seconds=360000, ) as codex_mcp_server: print("Codex MCP server started.") # More logic coming in the next sections. return if __name__ == "__main__": asyncio.run(main())

运行一次脚本以验证 Codex 是否成功启动:

python codex_mcp.py

打印 Codex MCP 服务器启动后,脚本退出。在接下来的部分中,您将在更丰富的工作流程中重用相同的 MCP 服务器。

构建 single-agent 工作流程

让我们从一个使用 Codex MCP 来发布小型浏览器游戏的示例开始。工作流程依赖于两个代理:

游戏设计师:为游戏撰写简介。

游戏开发者:通过调用Codex MCP来实现游戏。

使用以下代码更新 codex_mcp.py。它保留上面的 MCP 服务器设置并添加两个代理。

import asyncio import os from dotenv import load_dotenv from agents import Agent, Runner, set_default_openai_api from agents.mcp import MCPServerStdio load_dotenv(override=True) set_default_openai_api(os.getenv("OPENAI_API_KEY")) async def main() -> None: async with MCPServerStdio( name="Codex CLI", params={ "command": "codex", "args": ["mcp-server"], }, client_session_timeout_seconds=360000, ) as codex_mcp_server: developer_agent = Agent( name="Game Developer", instructions=( "You are an expert in building simple games using basic html + css + javascript with no dependencies. " "Save your work in a file called index.html in the current directory. " "Always call codex with \"approval-policy\": \"never\" and \"sandbox\": \"workspace-write\"." ), mcp_servers=[codex_mcp_server], ) designer_agent = Agent( name="Game Designer", instructions=( "You are an indie game connoisseur. Come up with an idea for a single page html + css + javascript game that a developer could build in about 50 lines of code. " "Format your request as a 3 sentence design brief for a game developer and call the Game Developer coder with your idea." ), model="gpt-5", handoffs=[developer_agent], ) await Runner.run(designer_agent, "Implement a fun new game!") if __name__ == "__main__": asyncio.run(main())

执行脚本:

python codex_mcp.py

Codex will read the designer’s brief, create an index.html file, and write the full game to disk. Open the generated file in a browser to play the result. Every run produces a different design with unique play-style twists and polish.

扩展到 multi-agent 工作流程

现在将 single-agent 设置转变为精心策划的、可追踪的工作流程。系统添加:

项目经理:创建共享需求、协调 hand-offs 并加强防护。

设计人员、前端开发人员、服务器开发人员和测试人员:每个人都有范围指令和输出文件夹。

创建一个名为 multi_agent_workflow.py 的新文件:

import asyncio import os from dotenv import load_dotenv from agents import ( Agent, ModelSettings, Runner, WebSearchTool, set_default_openai_api, ) from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX from agents.mcp import MCPServerStdio from openai.types.shared import Reasoning load_dotenv(override=True) set_default_openai_api(os.getenv("OPENAI_API_KEY")) async def main() -> None: async with MCPServerStdio( name="Codex CLI", params={"command": "codex", "args": ["mcp"]}, client_session_timeout_seconds=360000, ) as codex_mcp_server: designer_agent = Agent( name="Designer", instructions=( f"""{RECOMMENDED_PROMPT_PREFIX}""" "You are the Designer.\n" "Your only source of truth is AGENT_TASKS.md and REQUIREMENTS.md from the Project Manager.\n" "Do not assume anything that is not written there.\n\n" "You may use the internet for additional guidance or research." "Deliverables (write to /design):\n" "- design_spec.md – a single page describing the UI/UX layout, main screens, and key visual notes as requested in AGENT_TASKS.md.\n" "- wireframe.md – a simple text or ASCII wireframe if specified.\n\n" "Keep the output short and implementation-friendly.\n" "When complete, handoff to the Project Manager with transfer_to_project_manager." "When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}." ), model="gpt-5", tools=[WebSearchTool()], mcp_servers=[codex_mcp_server], ) frontend_developer_agent = Agent( name="Frontend Developer", instructions=( f"""{RECOMMENDED_PROMPT_PREFIX}""" "You are the Frontend Developer.\n" "Read AGENT_TASKS.md and design_spec.md. Implement exactly what is described there.\n\n" "Deliverables (write to /frontend):\n" "- index.html – main page structure\n" "- styles.css or inline styles if specified\n" "- main.js or game.js if specified\n\n" "Follow the Designer’s DOM structure and any integration points given by the Project Manager.\n" "Do not add features or branding beyond the provided documents.\n\n" "When complete, handoff to the Project Manager with transfer_to_project_manager_agent." "When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}." ), model="gpt-5", mcp_servers=[codex_mcp_server], ) backend_developer_agent = Agent( name="Backend Developer", instructions=( f"""{RECOMMENDED_PROMPT_PREFIX}""" "You are the Backend Developer.\n" "Read AGENT_TASKS.md and REQUIREMENTS.md. Implement the backend endpoints described there.\n\n" "Deliverables (write to /backend):\n" "- package.json – include a start script if requested\n" "- server.js – implement the API endpoints and logic exactly as specified\n\n" "Keep the code as simple and readable as possible. No external database.\n\n" "When complete, handoff to the Project Manager with transfer_to_project_manager_agent." "When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}." ), model="gpt-5", mcp_servers=[codex_mcp_server], ) tester_agent = Agent( name="Tester", instructions=( f"""{RECOMMENDED_PROMPT_PREFIX}""" "You are the Tester.\n" "Read AGENT_TASKS.md and TEST.md. Verify that the outputs of the other roles meet the acceptance criteria.\n\n" "Deliverables (write to /tests):\n" "- TEST_PLAN.md – bullet list of manual checks or automated steps as requested\n" "- test.sh or a simple automated script if specified\n\n" "Keep it minimal and easy to run.\n\n" "When complete, handoff to the Project Manager with transfer_to_project_manager." "When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}." ), model="gpt-5", mcp_servers=[codex_mcp_server], ) project_manager_agent = Agent( name="Project Manager", instructions=( f"""{RECOMMENDED_PROMPT_PREFIX}""" """ You are the Project Manager. Objective: Convert the input task list into three project-root files the team will execute against. Deliverables (write in project root): - REQUIREMENTS.md: concise summary of product goals, target users, key features, and constraints. - TEST.md: tasks with [Owner] tags (Designer, Frontend, Backend, Tester) and clear acceptance criteria. - AGENT_TASKS.md: one section per role containing: - Project name - Required deliverables (exact file names and purpose) - Key technical notes and constraints Process: - Resolve ambiguities with minimal, reasonable assumptions. Be specific so each role can act without guessing. - Create files using Codex MCP with {"approval-policy":"never","sandbox":"workspace-write"}. - Do not create folders. Only create REQUIREMENTS.md, TEST.md, AGENT_TASKS.md. Handoffs (gated by required files): 1) After the three files above are created, hand off to the Designer with transfer_to_designer_agent and include REQUIREMENTS.md and AGENT_TASKS.md. 2) Wait for the Designer to produce /design/design_spec.md. Verify that file exists before proceeding. 3) When design_spec.md exists, hand off in parallel to both: - Frontend Developer with transfer_to_frontend_developer_agent (provide design_spec.md, REQUIREMENTS.md, AGENT_TASKS.md). - Backend Developer with transfer_to_backend_developer_agent (provide REQUIREMENTS.md, AGENT_TASKS.md). 4) Wait for Frontend to produce /frontend/index.html and Backend to produce /backend/server.js. Verify both files exist. 5) When both exist, hand off to the Tester with transfer_to_tester_agent and provide all prior artifacts and outputs. 6) Do not advance to the next handoff until the required files for that step are present. If something is missing, request the owning agent to supply it and re-check. PM Responsibilities: - Coordinate all roles, track file completion, and enforce the above gating checks. - Do NOT respond with status updates. Just handoff to the next agent until the project is complete. """ ), model="gpt-5", model_settings=ModelSettings( reasoning=Reasoning(effort="medium"), ), handoffs=[designer_agent, frontend_developer_agent, backend_developer_agent, tester_agent], mcp_servers=[codex_mcp_server], ) designer_agent.handoffs = [project_manager_agent] frontend_developer_agent.handoffs = [project_manager_agent] backend_developer_agent.handoffs = [project_manager_agent] tester_agent.handoffs = [project_manager_agent] task_list = """ Goal: Build a tiny browser game to showcase a multi-agent workflow. High-level requirements: - Single-screen game called "Bug Busters". - Player clicks a moving bug to earn points. - Game ends after 20 seconds and shows final score. - Optional: submit score to a simple backend and display a top-10 leaderboard. Roles: - Designer: create a one-page UI/UX spec and basic wireframe. - Frontend Developer: implement the page and game logic. - Backend Developer: implement a minimal API (GET /health, GET/POST /scores). - Tester: write a quick test plan and a simple script to verify core routes. Constraints: - No external database—memory storage is fine. - Keep everything readable for beginners; no frameworks required. - All outputs should be small files saved in clearly named folders. """ result = await Runner.run(project_manager_agent, task_list, max_turns=30) print(result.final_output) if __name__ == "__main__": asyncio.run(main())

运行脚本并观察生成的文件:

python multi_agent_workflow.py ls -R

项目经理代理编写 REQUIREMENTS.md、TEST.md 和 AGENT_TASKS.md,然后在设计器、前端、服务器和测试器代理之间协调 hand-offs。每个代理在将控制权交还给项目经理之前,都会将范围内的工件写入自己的文件夹中。

追踪工作流程

Codex automatically records traces that capture every prompt, tool call, and hand-off. After the multi-agent run completes, open the Traces dashboard to inspect the execution timeline.

high-level 跟踪突出显示了项目经理在继续操作之前如何验证 hand-offs。单击各个步骤可查看提示、Codex MCP 调用、写入的文件和执行持续时间。这些细节使得审核每个 hand-off 并了解工作流程如何依次演变变得简单。这些跟踪使得调试工作流故障、审计代理行为以及测量一段时间内的性能变得简单,而无需额外的仪器。

站内延伸阅读