/loop cadence: on-demand queue. For each incoming question in the queue, a cheap model attempts an answer and self-checks against [CRITERIA]. Append (question, cheap-model verdict, pass/fail) to state-file escalation-log.md. Escalate to Fable ONLY where the cheap model logged a failure; Fable answers just those. Read + own-file writes only. Stop when the queue is empty; log errors and stop. Budget: cheap-first, cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Firecrawl MCP, diff tracked regulatory/authoritative sources [LIST] against the last snapshot. Append only material changes to state-file source-digest.md. Each round, write a plain-language 'what changed / so what / who it affects' for each material change; ignore cosmetic edits. Read + own-file writes only. Stop after the digest; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: nightly (synthesis monthly). Using Exa MCP + a Firecrawl monitor on tracked sources [LIST], have a cheap model log what materially changed to state-file intel-feed.md each night. Route to Fable only for the monthly synthesis: read the month's log and write ONE briefing of what changed and why it matters. Read + own-file writes only. Stop nightly after logging / monthly after the brief; log errors and stop. Budget: cheap-first, cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: daily or weekly. Using the Stripe API (read), maintain an aging ledger of unpaid/overdue invoices. Append (customer, invoice, days overdue, threshold hit) to state-file invoice-aging.md. When an invoice crosses a threshold [7/14/30d], draft the appropriate reminder in the right tone — DRAFT ONLY. Never send outbound, never touch charges or refunds. Stop after drafting due nudges; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: daily. Using the PostHog MCP and Stripe (read), check my key metrics [LIST] against their normal bands. Append daily readings to state-file kpi-watch.md. On a normal day, do nothing but log. When a metric breaks its band, pre-investigate (segment, correlate, likely cause) and draft an alert with the diagnosis. Never change data or send customer-facing messages. Stop after the check; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: on-demand or weekly. From my answer library [SOURCE] and the incoming RFP/proposal [DOC], draft the reusable ~80% (boilerplate, standard answers, past-response matches). Append (RFP, sections drafted, novel questions flagged) to state-file proposal-backlog.md. Each round, complete ONE proposal's reusable portion and list the new-20% questions a human must answer. Draft only; submit nothing. Stop after one; log and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Notion MCP (read), compare recent task/project records against the documented SOP pages [LINKS]. Append drift findings (step, written vs actual, evidence) to state-file sop-drift.md. Each round, draft ONE SOP update proposal for the biggest drift — propose only, edit no live SOP. Stop after one proposal; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: monthly. Using the QuickBooks API (read), categorize the period's transactions against historical patterns. Append matches + anomalies to state-file month-close.md. Each round, build/refresh the exception list (uncategorized, unusual, likely-miscoded) for a human to review. Never post, file, or reconcile anything in QuickBooks. Stop when the exception list is complete; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: daily. Using the Gmail MCP (read + draft), process each unread thread and classify it decide / delegate / defer / drop. Append (thread, classification, rationale) to state-file inbox-triage.md. For 'decide' and 'delegate' threads, save a draft reply — draft only, send nothing, archive nothing. Stop after the unread batch; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the PostHog MCP (read), find the funnel step with the steepest drop-off. Append (step, drop rate, hypotheses) to state-file drop-points.md. Each round, draft rewritten copy/microcopy for the single worst drop screen — draft only, ship nothing to production. Stop after one screen; log analytics errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.
/loop cadence: weekly. Using the app store review APIs [STORES], pull new reviews + support exports. Append (issue, frequency, severity, star-impact) to state-file review-roadmap.md. Each round, re-rank the backlog by pain and draft a one-paragraph problem statement for the top unaddressed item. Write to the file only; change no roadmap tool live. Stop after one; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Reddit API, HN Algolia, and Exa, sweep mentions of [BRAND / PRODUCT]. Append (source, mention, sentiment, feature ask/complaint) to state-file mention-radar.md. Each round, cluster and surface the single loudest theme, then draft an implementation plan for it — plan only, build nothing. Stop after one plan; log source errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. From my pasted/exported pre-sale questions [SOURCE], cluster recurring pre-purchase questions and objections. Append (question cluster, frequency, best current answer) to state-file presale-bank.md. Each round, write or improve ONE canonical answer for the most-asked unanswered question. Write to the bank file only; send nothing to any customer. Stop after one answer; log and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Zendesk MCP (read), pull tickets expressing a mismatch between expectation and reality ('I thought it...', 'the site said...'). Append (ticket theme, misread feature, suspected source page) to state-file expectation-gaps.md. Each round, trace the single most common gap back to the exact page sentence and draft a clarifying rewrite — draft only, edit nothing live. Stop after one; log errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: every [N] days. Using the Meta Marketing API (read), watch frequency, CTR decay, and CPA drift on active creatives. Append fatigue signals to state-file ad-fatigue.md. When a creative crosses [THRESHOLD], draft the next creative variant batch (hooks, angles, copy) — draft only, launch nothing, touch no budget. Stop after drafting one batch; log API errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: daily (morning). Using twitterapi.io to read my curated X list [LIST ID] and Typefully to hold drafts, rank the last 24h of posts by engagement-per-follower. Append top themes to state-file x-ideas.md. Each round, draft ONE post in MY voice [VOICE NOTES] on the strongest theme as a Typefully draft — never publish. Stop after one draft; log API errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Search Console MCP, mine my own impressions/clicks for queries with demand but weak coverage. Append candidate topics (query, intent, current page, gap) to state-file brief-backlog.md. Each round, expand the single highest-opportunity topic into ONE fully-specified brief (angle, outline, target terms, internal links). Draft only. Stop after one brief; log GSC errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using DataForSEO MCP for keyword/SERP data and a Firecrawl monitor on competitor blogs [LIST], detect new competitor articles and the terms they target. Append (competitor, URL, target terms, gap vs our coverage) to state-file competitor-content.md. Each round, draft ONE counter-move brief against the largest open gap. Read + draft only; publish nothing. Stop after one brief; log source errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
Every week, ask the AI engines the same 'best [category]' questions your buyers ask, log where your brand ranks, and track the trend over time instead of guessing.
/loop cadence: weekly. Using the Perplexity API, run my fixed list of 'best [category]' / buyer-intent queries [PASTE QUERIES]. Append each result (query, my rank/mention, competitors named, date) to state-file share-of-model.md as a time series. Each round, draft ONE action for the query where I dropped the most. Never message anyone or change any page. Stop after logging the week; if a query errors, log the failure and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
/loop cadence: weekly. Using the Search Console MCP, pull my programmatic/template page set and check each for thin content, near-duplicate bodies, and impressions-without-clicks. Append flags (URL, issue, evidence) to state-file pseo-quality.md. Each round, draft a fix or consolidation recommendation for the WORST page only — recommend, never edit or deindex. Stop after one recommendation; log GSC errors and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.
/loop cadence: weekly. Using the Exa MCP, search for buyer-intent questions in [MY CATEGORY] that AI answer engines field but my site [DOMAIN] does not rank for or answer. Append findings (question, current answer source, whether we cover it) to state-file answer-engine-gaps.md. Each round, draft ONE page/section outline to close the single biggest gap — draft only, publish nothing. Stop after drafting one gap; if Exa errors, log it and stop. Budget: cap $[X]/run. Hard cap: stop after 1 iteration per run.
Split a large mechanical job into 2–5 independent Codex lanes, each isolated in its own worktree with a frozen acceptance bar and binding judge, then block the final merge behind a full integration judge.
Dispatch a parallel Codex legion for a large mechanical job. Split the work into 2–5 genuinely independent lanes, one lane per piece.
First announce a muster table with:
- each lane,
- the exact files that lane may touch,
- the frozen acceptance check for that lane.
Do not proceed until I approve the split.
Before dispatch, the orchestrator must freeze and record each lane’s acceptance bar. After dispatch, each worker treats `.git` as read-only.
Each lane must run in its own git worktree with:
- a frozen acceptance bar recorded before code changes,
- a strictly disjoint may-touch manifest,
- its own sandbox,
- read-only `.git` state for the worker.
If any lane’s file footprint overlaps another lane, refuse the split and serialize the work instead.
When a lane finishes, run a fresh-context judge against that lane’s frozen bar. The judge must return binding PASS or FAIL. Allow at most 2 retries per lane; stop after 2 failed attempts, then escalate loudly.
Merge lanes in a fixed order. After merging, require a mandatory integration judge that reruns the full test suite across the combined result. Do not commit or merge unless the integration judge returns PASS.
Hard cap: 5 workers. If there are more than 5 pieces, run later waves. Never merge without the integration judge.
praetor is a Claude Code plugin that runs a plan → freeze acceptance bar → dispatch → independent fresh-context judge → resolve loop. Claude plans and judges, Codex executes; a FAIL from the judge cannot be overridden, with at most 2 retries before a loud takeover.
Plan the task and freeze the acceptance criteria in .codex/ACCEPTANCE.md before any work begins. Isolate on a throwaway branch, write a self-contained brief, then dispatch execution to Codex. When Codex finishes, spawn a fresh-context independent judge that runs every check in the frozen bar against the uncommitted working tree and returns a binding PASS or FAIL — a FAIL cannot be overridden. The judge never fixes anything and commits nothing; it touches manifest paths only. Resolve with at most 2 retries; on continued failure, hand back with a loud takeover. Commit only after the judge passes, then clean up and write the ledger. Iron laws: frozen bar before dispatch, binding judge, max 2 retries then loud takeover.
#!/usr/bin/env python3 """ RALPH Loop Runner for RefactorBench. Runs iterative agent retry loops with filesystem-based memory on a SINGLE task at a time. RALPH pattern: - Agent runs, does work, exits - Work persists on the filesystem (working directory retains all changes) - Fresh agent starts, reads progress.json from the working directory, continues - Repeats until tests pass or max iterations Usage: python3 ralph runner.py --repo django refactor --task add-log-parameter-get-resolver python3 ralph runner.py --repo django refactor --task add-log-parameter-get-resolver --chains 2 --iterations 3 python3 ralph runner.py --repo django refactor --task add-log-parameter-get-resolver --verbose """ import argparse import asyncio import json import os import re import shutil import sys import time from dataclasses import dataclass from datetime import datetime from pathlib import Path from refactor agent import get task info, run test, setup workdir from notebook import ( FileSnapshot, NotebookWriter, parse stream json, compute solution diff, compute diff stats, ) import ralph prompt builder BENCH ROOT = Path( file ).parent / ".refactorbench" # --------------
# Ralph  Ralph is a minimal, file‑based agent loop for autonomous coding. Each iteration starts fresh, reads the same on‑disk state, and commits work for one story at a time. ## How it works Ralph treats files and git as memory, not the model context: - PRD (JSON) defines stories, gates, and status - Loop executes one story per iteration - State persists in .ralph/  ## Global CLI (recommended) Install and run Ralph from anywhere: bash npm i -g @iannuttall/ralph ralph prd # launches an interactive prompt ralph build 1 # one Ralph run ### Template hierarchy Ralph will look for templates in this order: 1. .agents/ralph/ in the current project (if present) 2. Bundled defaults shipped with this repo State and logs always go to .ralph/ in the project. ### Install templates into a project (optional overrides) bash ralph install This creates .agents/ralph/ in the current repo so you can customize prompts and loop behavior. During install, you’ll be asked if you want to add the required skills. ### Install required skills (optional) bash ralph install --skills You’ll be prom Cap the run at 25 iterations; leave remaining work for the next session.
/loop run the repo's static analyzer (semgrep, CodeQL, or whatever is already configured) with the security ruleset; take ONE finding — highest severity first — and fix it minimally, then re-run the analyzer to verify the finding is gone and run the test suite. Never suppress or downgrade a rule to make a finding disappear; anything that needs a design change gets flagged for human review instead. Continue until the analyzer reports zero findings at high severity — stop after 10 turns and propose the fixes as one PR.
Pick the oldest untriaged issue, validate it against the current build, then label it, close if obsolete, or document repro steps—one per turn until the queue empties or 12 iterations pass.
/loop pick the single oldest untriaged issue in the tracker; reproduce or validate it against the current build, then either label it (area, priority, effort), close it with a polite explanation if it is obsolete, or write the missing repro steps. One issue per iteration, never close anything that still reproduces, and propose bulk closes for maintainer review instead of executing them. Continue until the untriaged queue hits zero — stop after 12 turns and report the triaged/closed/escalated counts.
/schedule every Friday at 2pm, draft release notes from the PRs merged since the last draft: group changes into features, fixes, and docs, write one plain-English line per change with the PR link, and save to releases/DRAFT.md. Draft only — a human edits and publishes; never post or send anywhere. Verify every PR link resolves and every merged PR since the last run is covered before saving. Stop after 1 pass per run.
Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.