Cursor "Iterate Until Tests Pass, Never Touch the Tests"

First-party Cursor guidance for the iterate-until-green loop, with the key anti-reward-hacking clause: the agent may never modify the tests it is trying to satisfy. Works in Cursor, Claude Code /goal, and Codex.

prompt
→ Claude Code
"Write code that makes these tests pass. Do NOT modify the tests. Keep iterating — run the suite, fix failures, run again — until all tests pass." (paraphrase of Cursor's official agent best-practices guidance)
claude-code

Implementation note

When to use: any test-driven agent session where the tests define done — you have a failing suite (TDD-style or a regression pile) and want the agent iterating until green without the classic cheat. How it works: the instruction is a plain contract: write code that makes these tests pass, do NOT modify the tests, and keep iterating — run the suite, fix failures, run again — until all tests pass. This is first-party Cursor guidance from their agent best-practices, and the pattern translates directly to Claude Code /goal and to Codex. Safety: the never-touch-the-tests clause is the entire anti-reward-hacking rail — without it, the cheapest path to green is editing an assertion, and agents find cheap paths. State it explicitly every time. The residual check is yours: confirm at the end that the test files are untouched (a quick git diff on the test paths) and that the implementing code is honest.

Source: Cursor blog ↗graded B · 80/100 — how grades work →

More testing loops

Build a REST API with tests

Loop/ralphGitHubB

Run autonomous iterations to ship a complete REST API with full test coverage until completion is promised.

prompt
→ Claude Code
/ralph-loop "Build a REST API with tests" --max-iterations 30 --completion-promise "COMPLETE
testingmedium risk

Test all endpoints

Loop/goalGitHubB

Run tests against every API endpoint until coverage is complete or you hit 30 iterations.

prompt
→ Claude Code
/goal all endpoints tested or stop after 30 turns
testingmedium risk

Hit acceptance criteria

Loop/goalXB

Drive a feature to done against explicit acceptance criteria: a working paginated endpoint, passing tests, clean lint, and a hard turn cap.

prompt
→ Claude Code
/goal the /users endpoint returns 200 with a paginated JSON body, all tests pass, and lint is clean — stop after 20 turns