Testing loops

Loops that run your test suite, fix what fails, and exit only when everything is green.

17 graded loops in this category.

Hit acceptance criteria

Loop/goalXB

Drive a feature to done against explicit acceptance criteria: a working paginated endpoint, passing tests, clean lint, and a hard turn cap.

prompt
→ Claude Code
/goal the /users endpoint returns 200 with a paginated JSON body, all tests pass, and lint is clean — stop after 20 turns

Build a REST API with tests

Loop/ralphGitHubB

Run autonomous iterations to ship a complete REST API with full test coverage until completion is promised.

prompt
→ Claude Code
/ralph-loop "Build a REST API with tests" --max-iterations 30 --completion-promise "COMPLETE
testingmedium risk

Test all endpoints

Loop/goalGitHubB

Run tests against every API endpoint until coverage is complete or you hit 30 iterations.

prompt
→ Claude Code
/goal all endpoints tested or stop after 30 turns
testingmedium risk

# AGENTS.md - Codex Ralph Vault Loop ##…

Loop/ralphGitHubB

Community ralph loop for testing, sourced from github. Verified exit condition, evaluator-gated.

prompt
→ Claude Code
# AGENTS.md - Codex Ralph Vault Loop ## Mission codex-ralph-vault-loop is a Codex App/CLI native orchestration overlay for multi-agent engineering work. It keeps Codex main as the decision maker, uses external models only through MCP tools, verifies work through gates, and stores durable memory in the vault layer. ## Core Rules - Codex main decides. The primary Codex session owns final decisions, edits, synthesis, safety, and verification. - External models advise. Z.ai, MiniMax, and other non-OpenAI systems provide analysis or worker output only through MCP tools. - Gates verify. Tests, lint, security checks, scorecards, and migration checkpoints decide whether a phase can pass. - Vault remembers. Durable memory belongs in the approved Ralph/Codex memory paths, not in ad hoc repo files. - Do not bypass critical hooks. If prettier , gitleaks , semgrep , or pre-commit are missing from PATH , use the local machine binaries when present, install only with approval, or stop and report the blocker; do not use --no-verify to skip security or formatting gates unless the user explicitly orders that exact bypass. - Do not merge or close a PR until review feedback and automated Cap the run at 25 iterations; leave remaining work for the next session.
testinghigh risk

/ralph-loop "fix all tests" --max-iterations 5 (iterates until…

Loop/ralphGitHubB

Community ralph loop for testing, sourced from github. Verified exit condition, evaluator-gated.

prompt
→ Claude Code
/ralph-loop "fix all tests" --max-iterations 5 (iterates until passing
testingmedium risk

API contract test backfill

Loop/goallooprepoB

Generate a contract test for every documented endpoint in the OpenAPI spec so the spec and the implementation can never silently drift.

prompt
→ Claude Code
/goal every path in openapi.yaml has a contract test asserting its status codes and response schema — add tests for one untested endpoint per turn, run the suite, and fix either the spec or the handler when they disagree (tell me which you chose); stop when all paths are covered or after 15 turns

Backlog-clearing loop with a separate verifier

Loop/ralphXA

The four-settings loop template: a separate verifier model that never shares context with the writer, a hard stop rule, a state file re-read each cycle, and worktree isolation. Point it at a checkable backlog and let it run overnight.

prompt
→ Claude Code
GOAL: every test in [/tests/TARGET] passes, lint is clean, zero type errors. EACH CYCLE: 1. run the suite, read every failure 2. pick the single highest-impact failure 3. write the smallest change that fixes it 4. re-run tests + lint + type check VERIFY: a separate model instance checks the goal — never the writer. Verifier prompt: "You are a verifier. You did not write this code. GOAL: <the exact goal string>. Given the diff and the test output, answer ONLY: PASS — every condition in GOAL is objectively met, with evidence, or FAIL: <the specific condition not met, and the evidence>. Do not fix anything. If unsure, FAIL." STOP WHEN: verify passes, OR after 10 iterations, OR $5 spent, OR no progress in 2 attempts. ON BLOCKER: log it, skip to the next item, never halt the whole loop. STATE: append done / failed / next to a state file, re-read it at the top of every cycle. ISOLATION: one git worktree per subagent.
testinglow risk

Keep QA tests passing

Loop/loopGitHubB

Run pre-QA tests repeatedly, diagnose and fix failures, report only on blockers or green.

prompt
→ Claude Code
/loop 2m Read tasks/qa-config.md to get the pre-QA test command and log command. Run the pre-QA test command. If it fails or any infrastructure issue is detected: read logs using the log command, diagnose the root cause, fix it, and re-run. Report to the user only when something fails or when all checks pass clean. Do not interrupt QA for passing tests
testingmedium risk

Get the fix to score 90+

Loop/goalGitHubA

Run verify commands and benchmark checks until the target code scores 90 or higher on the benchmark, stopping after 5 attempts.

prompt
→ Claude Code
/goal the transcript reports a line matching SCORE: <n>/100 (threshold: 90, gate: pr) with n >= 90 for <slug> on branch <branch>, after <(bugfix) the REPRODUCE test is shown failing on unfixed code and passing on the final code, and> <verify commands, comma-separated> <(hot-path) and make bench-compare> all pass on the final code, or stop after 5 fix rounds
testingmedium risk

Implement spec.md with TDD

Loop/ralphGitHubB

Implement spec.md with test-driven development, loop until all tests pass and output COMPLETE.

prompt
→ Claude Code
/ralph-loop "Implement spec.md. TDD. Output <promise>COMPLETE</promise> when all tests pass." --max-iterations 30
testingmedium risk
Sponsored · ConnectMyEmail

Loops that read your inbox.

Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.

connectmyemail.com →

Ralph a test backlog

Iterate over a prioritized list of untested modules with fresh context each pass, writing real behavioral tests for one module at a time and banking lessons in a guardrails file.

prompt
→ Claude Code
/loop each iteration with fresh context: read .ralph/test-backlog.json and .ralph/guardrails.md, pick the top unfinished module, write behavioral tests for its public API (no snapshot-only tests), run the suite, and mark the module done only when its tests pass and coverage for it exceeds 80%; append any discovered testing gotcha (fixtures, mocking rules, async traps) to .ralph/guardrails.md; stop when the backlog is empty or after 25 turns

Make tests pass, gate frozen

Loop/goalGitHubA

Run verify.sh repeatedly until it exits 0, all acceptance gates hold, and test quality gates are met, or report after 20 turns.

prompt
→ Claude Code
/goal ./scripts/verify.sh exits 0, .ai/spec-tdd/state.json phase is done, frozen tests and acceptance gates are unchanged, no tests are skipped/weakened, and no TODO/stub/hardcoded test-only implementation remains; or stop after 20 turns with a clear blocked report
testingmedium risk

Fix interpreter matching test failures

Loop/loopGitHubB

Run the test suite repeatedly, fixing InterpreterMatchingServiceTests failures until all pass.

prompt
→ Claude Code
/loop --max-turns 25 until tests pass: fix InterpreterMatchingServiceTests failures
testinglow risk

Trim what loads before first paint

Reduce the data downloaded before the first screen appears, with tests and screenshots guarding behavior and appearance.

prompt
→ Claude Code
Reduce the data [web app] downloads before its first screen appears. First record passing tests, mobile and desktop screenshots, and compressed transferred bytes—the data actually downloaded. Use the build report only to suggest candidates. Defer, compress, or remove one item, then rebuild and rerun every check. Keep it only if tests pass, screenshots are pixel-identical, and bytes decrease; otherwise revert. Stop when no safe candidate remains, progress stalls, or approval is needed. Return measurements, changes, and untested states.

Cursor "Iterate Until Tests Pass, Never Touch the Tests"

Loop/ralphcommunityB

First-party Cursor guidance for the iterate-until-green loop, with the key anti-reward-hacking clause: the agent may never modify the tests it is trying to satisfy. Works in Cursor, Claude Code /goal, and Codex.

prompt
→ Claude Code
"Write code that makes these tests pass. Do NOT modify the tests. Keep iterating — run the suite, fix failures, run again — until all tests pass." (paraphrase of Cursor's official agent best-practices guidance)
testingmedium risk

Ship avatar endpoint to spec

Loop/goalGitHubB

Build and test a new /users/:id/avatar route until it passes tests, lint, and matches the OpenAPI schema.

prompt
→ Claude Code
/goal until grep -q "users/:id/avatar" src/routes/users.ts AND pnpm test -- tests/integration/users.avatar.spec.ts passes AND grep -q "/users/{id}/avatar" openapi.yaml AND pnpm lint passes, or stop after 15 turns
testingmedium risk

Stabilize flaky tests for good

Measure the flakiness, fix one root cause at a time, and stop after a defined streak of stable full-suite runs.

prompt
→ Claude Code
Run [test suite] [N] times under the same conditions and list tests whose result changes. Fix the most frequent flake at its root cause—shared state, timing, ordering, or an external dependency—never with a blind sleep or retry. Run that test [N] times, then rerun the full suite. Repeat until [N] consecutive full-suite runs pass, progress stalls, or approval is required. Return each flake, root cause, fix, evidence, and justified quarantine.
Testing agent loops for Claude Code & Codex | looprepo