Codex Iterative Repair Loop (JSON-Schema Review → Repair)

OpenAI's first-party loop recipe: a script alternates a Codex review pass that emits machine-readable findings with a repair pass fed those findings verbatim, looping until validation passes, attempts run out, progress stalls, or a decision needs human review.

prompt
→ Claude Code
Script alternates two `codex exec` calls: (a) review pass with a JSON schema output ("list remaining issues as machine-readable findings"), (b) repair pass fed those findings verbatim. Loop while findings remain, capped by max attempts. Stops for one of four reasons: validation passes, max attempts reached, remaining delta stops changing, or next decision needs human review.
claude-code

Implementation note

When to use: repair work you want structured rather than freeform — a codebase with known issues where you need machine-readable findings driving fixes, with clear stop conditions. This is OpenAI's first-party loop recipe. How it works: a script alternates two codex exec calls. The review pass emits remaining issues as machine-readable findings conforming to a JSON schema; the repair pass is fed those findings verbatim and fixes them. The loop continues while findings remain and halts for exactly one of four reasons: validation passes, max attempts are reached, the remaining delta stops changing between passes, or the next decision needs human review. Safety: the four enumerated stop conditions are the rail — stalled progress and needs-a-human are first-class exits rather than failure modes discovered later, and the attempt cap bounds spend. The JSON-schema findings double as an audit trail of what the loop believed was wrong at each pass.

Source: OpenAI Cookbook ↗graded C · 55/100 — how grades work →

More review loops

Dual-reviewer convergence gate before opening a PR

Loop/goalGitHubAnew

Two reviewers from different model families, each in a fresh context, must both pass an objective rubric before a PR is opened — the maker-checker split with real model diversity. Each failed round fixes only what was flagged and re-reviews with reviewers that have no memory of the last round, so nothing anchors on prior findings. Hard cap of three rounds, tests may never be edited to pass, and the loop ends by opening a PR rather than pushing. Adapted from ECC's /santa-loop (affaan-m/ECC, MIT), which auto-pushes on agreement and has no verifiable exit; this version adds the test-exit condition and the human merge gate.

prompt
→ Claude Code
/goal Ship the current diff only after two independent reviewers both PASS. Scope is `git diff --name-only HEAD` (or the path in $ARGUMENTS). First write a rubric with objective PASS/FAIL criteria: correctness, security (no secrets, injection, OWASP top 10), error handling, completeness, internal consistency, no regressions. Each round launch two reviewers in parallel with fresh context and no memory of earlier rounds: Reviewer A is a Claude subagent, Reviewer B is a different model via `codex exec --sandbox read-only` (fall back to a second Claude subagent and say so in the report). Both return a JSON verdict with per-criterion PASS/FAIL and critical issues. If either FAILs: fix only the flagged issues with minimal diffs, never modify the tests to make them pass, commit "fix: address review findings (round N)", run `npm test`, and re-review with fresh reviewers. Exit when both reviewers PASS and `npm test` exits 0 — then open a PR for me to merge; never push to main or merge. Stop after 3 rounds; if still failing, print the unresolved issues and escalate to me instead of shipping.

The nested perfect loop

Loop/loopGitHubB

A loop wrapping a goal wrapping a review: every 30 minutes, drive all PR review comments to resolved via /review, 10 turns max per pass.

prompt
→ Claude Code
/loop 30m /goal all PR review comments resolved via /review, stop after 10 turns

Fix code review blockers until convergence

Loop/goalGitHubBnew

Run a full review pass inline, apply validated blocking findings at their owning layer, re-run fresh until the gate clears or blockers plateau—capped at 3 rounds with escalation.

prompt
→ Claude Code
/goal why-review fix-loop convergence: repeatedly run the full-mode /why-review pass INLINE over {target}. After each review, apply only VALIDATED findings that block the current round (all open severities in round 1 — Round-1 LOW closure; CRITICAL/HIGH/MEDIUM from round 2) at their owning layer, then re-run a FRESH full review pass over the CHANGED target. If the current round's blocking findings >0 → apply fixes and run another round; if a fresh full review pass clears the current bar (round 1: zero open findings; round 2: zero CRITICAL/HIGH/MEDIUM, LOW deferred) → CONVERGED, clear the gate. Do NOT open another round for LOW-only findings from round 2. Cap at {N=3} review rounds; a failing test gate is outside the cap and the no-progress rule — keep fixing and re-running until the tests pass, never forcing green; if review blockers do not shrink across 2 consecutive rounds or increase, or the review budget (round 3) is spent with review blockers still open → STOP and escalate via AskUserQuestion. Never loop open-ended
reviewmedium risk