Codex CLI as an MCP Tool Inside an Agents-SDK Loop

An outer-planner/inner-coder loop from an official OpenAI recipe: an Agents SDK orchestrator plans and verifies while Codex CLI, wrapped as an MCP server, performs one bounded code change per turn.

prompt
→ Claude
Wrap Codex CLI as an MCP server and drive it from an OpenAI Agents SDK orchestrator loop — the outer agent plans/verifies, the inner Codex call does one bounded code change per turn. Guardrails: Stop when the goal is verifiably met, or stop after 15 iterations, whichever comes first. Verify each pass by running the relevant tests or checks — self-reported success does not count. Keep changes minimal and never touch files outside the task’s scope.
claude-code

Implementation note

When to use: workflows needing separation between planning and coding — you want an orchestrator that decides what to do and verifies outcomes, with the code-touching capability isolated behind a tool boundary. From an official OpenAI recipe. How it works: Codex CLI is wrapped as an MCP server, and an OpenAI Agents SDK orchestrator drives it in a loop: the outer agent plans the work and verifies results, while each inner Codex call performs exactly one bounded code change per turn. The MCP boundary makes the coder a discrete, inspectable tool call rather than an open-ended session. Safety: the outer-planner, inner-coder split is the structural rail — the agent that changes code is not the agent that judges whether the change succeeded, and each code mutation is bounded to one tool invocation. Put iteration and budget caps in the orchestrator loop, since it, not Codex, controls how many turns run. Hardened 2026-07-27: explicit stop/cap/verification guardrails appended; regraded D→A.

Source: OpenAI Cookbook

More automation loops

Ship production-grade apps autonomously

Loop/ralphnew

Hand an idea to Claude Code; it authors specs, designs, builds, tests, secures, and ships until enterprise done or budget exhausted.

prompt
→ Claude
# dare-to-be-stupid — Design (v2, refined) > A Claude Code plugin. One command, /dare , hands an idea or PRD to an autonomous > loop that authors specs, designs, builds, tests, secures, ships, fixes, and iterates > until the app passes an enterprise-production definition of done — or the budget dies. > > Named for the Weird Al song. The joke is that it runs the Ralph Loop on purpose , > with --dangerously-skip-permissions , and narrates the whole thing in the voice of an > '80s Junkion. Pre-production only. Never points at anything with users. This is v2. It keeps the strong core of the original spec (external reviewer, ratchet, guard hook, Junkion style) and adds the three phases the original left thin relative to the actual goal: PRD authoring, a design phase, and a real enterprise DoD including security, CI, docs/observability, and design quality (with quality plugins auto-installed). --- ## 0. The premise, in one paragraph The User builds documentation-first: spec → system docs → API contracts → CLAUDE.md → code. dare-to-be-stupid is the deliberate inverse, packaged as comedy that also solves two real engineering problems. It is a real build , not a joke ar
automationhigh riskclaude-code

claude-progress.txt harness pattern (Anthropic)

Loop/ralph★ Anthropic

Anthropic's first-party file-as-memory harness for long-running agents: every fresh-context session recovers state from a progress file and the git log, does one unit of work, updates the file, commits, and exits.

prompt
→ Claude
Long-running agent harness: each fresh-context session starts by reading `claude-progress.txt` + git log to recover state, does one unit of work, updates the progress file, commits, exits. Initializer session sets up the file; coder sessions loop. Guardrails: Stop when the goal is verifiably met, or stop after 15 iterations, whichever comes first. Verify each pass by running the relevant tests or checks — self-reported success does not count. Keep changes minimal and never touch files outside the task’s scope.
automationmedium riskclaude-code

Complete all tasks in tasks.md

Loop/ralphnew

Work through your tasks.md file, completing each task until none remain or max iterations is reached.

prompt
→ Claude
/ralph-loop "Complete all tasks in tasks.md" --max-iterations 50
automationmedium riskclaude-code