← All guides

How to Grade and Test a Loop Library Loop

Forward Future's Loop Library is one of the best public sources of practical agent loops: dozens of community-published prompts with use-when notes, verification sections, and stopping conditions, plus the installable loopy skill that can find, adapt, craft, and run them from inside Claude Code (/loopy), Codex, or Cursor. looprepo imports from it under its MIT license, with attribution on every imported loop.

A library and a grader do different jobs, though. A library tells you what a loop says. A grade tells you whether it's safe and likely to work before you hand it to an agent running unattended. Loop Library's verification sections are written by the loop's author; nothing machine-checks that the bounds are actually there. That's the gap this workflow fills — and it works on any loop text, not just Loop Library's.

Step 1 — Get the loop text

Any of these works:

  • Copy the prompt from a Loop Library page (the Prompt section, verbatim).
  • Copy an entry the loopy skill saved into your project's LOOPS.md — the fenced
  • Ask /loopy to show you the loop it's about to run, and copy that.

Step 2 — Grade it

Paste the prompt into looprepo's review page. If your agent is connected to the looprepo MCP server, you can skip the browser: call loops_grade with the same text and get the identical result.

The grade is deterministic — the same text always produces the same score, because it comes from the same default-fail evaluator that gates every loop published in this directory, not from a model's opinion (Safety Score v1.0; the full rubric, weights, and its known limits are on the how-we-grade page). It checks for:

  • Exit conditions — does the loop say when it's done, in checkable terms?
  • Iteration budgets — is there a turn cap or run limit, so a stuck loop stops
  • Verification commands — does it verify results with something executable
  • Scoping — does it constrain what the agent may touch ("only edit
  • Human gates — does risky output route through a PR or review step?
  • Dangerous patterns — force-pushes, piped-to-shell downloads, recursive

A and B mean the loop is well-bounded as written. C means run it, but read the advice first. D means the loop is missing the bounds that make unattended runs safe — common and fixable, not a verdict on the author. F flags a dangerous pattern (a recursive delete, a force-push, a piped-to-shell download); none of the Loop Library snapshot below hit it.

Step 3 — Harden what's missing

Missing bounds are the norm, not the exception. We ran looprepo's grader over a full Loop Library catalog snapshot (70 loops, published 2026-06-26): 74% include a machine verification command — the catalog's editorial bar genuinely shows — but only 29% state an explicit exit condition, and 0% carry an iteration cap. None contained dangerous commands. In other words: well-verified loops that mostly don't say when to stop. That's exactly the profile the grade catches, on their catalog and on yours.

The review lists exactly what to add, with suggested wording. The usual fixes take a minute each:

  • No turn cap → append "stop after 8 iterations."
  • No exit condition → state the finish line as something checkable: "stop when
  • No verification → name the command that proves the work: the test suite, the
  • Unscoped → say what's off-limits: "never touch files outside docs/."

Agents can do the same programmatically with the loops_harden MCP tool, which returns a rewritten prompt plus the reasoning.

Step 4 — Save it back

If you use the loopy skill, save the hardened version to your project's LOOPS.md the way Loopy expects: the loop name, a one-sentence explanation, the exact prompt, and the save date — and, since your version is adapted from a published loop, the source loop's URL and its modified date. That keeps provenance intact in both directions: Loop Library credits the author, your LOOPS.md credits Loop Library, and if you later submit the hardened loop anywhere (including here — same evaluator, no separate review queue), the attribution travels with it.

Why this is a workflow, not a rivalry

looprepo's directory already carries dozens of Loop Library loops, imported under MIT with per-loop attribution — the two catalogs share a taxonomy and a conviction that loops should state their checks and stopping conditions. The difference is enforcement: a published verification section is a promise, a default-fail gate is a check. Use Loop Library (or the /loopy skill) to find and craft loops; use the grade to know what you're running; use the builder when you want the full harness — hooks, budgets, evaluator — generated around the loop you just hardened.

Ready to run one? Browse the loop directory →
All EVALUATION loops →

Draft a sprint plan from the backlog

Turn the open issue backlog into a proposed two-week sprint plan with estimates, a dependency ordering, and an explicit cut line, written as a document for the team to edit.

Open loop →

Fifteen-minute verification report while you work

A recurring six-phase check (build, types, lint, tests with coverage, secret and debug-log grep, diff size) that prints one fixed-format report every fifteen minutes so you catch drift before the PR, not in review. It is strictly read-only: the loop reports, you decide what to fix. It exits when the tests pass and the report reads READY twice running, escalates instead of retrying when the same check fails three times, and stops after eight checks no matter what. Adapted from ECC's verification-loop skill (affaan-m/ECC, MIT), whose 'continuous mode' names the cadence but has no turn cap, exit, or failure path; all three are added here.

Open loop →

Fix interpreter matching test failures

Run the test suite repeatedly, fixing InterpreterMatchingServiceTests failures until all pass.

Open loop →

Receipt & statement archiver

Scan the inbox for receipts, invoices, and statements and file each under a Receipts label — read-only on everything else, deleting nothing.

Open loop →