claude-progress.txt harness pattern (Anthropic)

Anthropic's first-party file-as-memory harness for long-running agents: every fresh-context session recovers state from a progress file and the git log, does one unit of work, updates the file, commits, and exits.

prompt
→ Claude Code
Long-running agent harness: each fresh-context session starts by reading `claude-progress.txt` + git log to recover state, does one unit of work, updates the progress file, commits, exits. Initializer session sets up the file; coder sessions loop. Guardrails: Stop when the goal is verifiably met, or stop after 15 iterations, whichever comes first. Verify each pass by running the relevant tests or checks — self-reported success does not count. Keep changes minimal and never touch files outside the task’s scope.
claude-code

Implementation note

When to use: agent work too long for one session — multi-day builds where context will be lost repeatedly and the question becomes how each new session knows where things stand. This is Anthropic's first-party answer. How it works: every fresh-context session begins by reading claude-progress.txt plus the git log to recover state, does one unit of work, updates the progress file, commits, and exits. An initializer session sets up the progress file; coder sessions then loop the pattern indefinitely. The progress file is curated state — what is done, what is next, what to watch for — while git history is the ground truth it points into. Safety: the discipline lives in the exit ritual: a session that fails to update the progress file before exiting strands the next one, so treat update-then-commit as non-negotiable. One unit of work per session keeps commits reviewable and recovery cheap when an iteration goes sideways. Hardened 2026-07-27: explicit stop/cap/verification guardrails appended; regraded D→A.

Source: Anthropic Engineering ↗graded A · 95/100 — how grades work →

More automation loops

Ship PRD stories via dual-agent loop

Loop/ralphGitHubA

Ralph runs a generator and evaluator in tandem until all user stories pass acceptance criteria and browser tests.

prompt
→ Claude Code
# Ralph Harness — Agent Instructions ## Overview Ralph Harness is an autonomous AI agent loop that runs AI coding tools (Amp or Claude Code) repeatedly until all PRD items are complete. Each iteration is a fresh instance with clean context. Ralph supports two modes: | Mode | Architecture | When to use | |------|-------------|-------------| | simple | Single agent (self-implement, self-check) | Quick tasks, backend-only stories, well-defined small changes | | harness | Generator + Evaluator (dual-agent with contract) | UI-heavy features, complex stories, when quality is critical | ## Architecture: Harness Mode ralph.sh orchestrator │ ├── Planner (prd.json) │ Defines user stories, acceptance criteria, dependencies │ ├── Generator (generator-prompt.md) │ Drafts sprint contracts → Implements stories → Fixes based on feedback │ └── Evaluator (evaluator-prompt.md) Reviews contracts → Signs/locks → Tests in browser → Scores → Writes feedback ### Per-Story Flow 1. Contract Negotiation : Generator drafts contract.json → Evaluator reviews → Back-and-forth until Evaluator signs → Contract locked (immutable) 2. Build : Generator reads Hard cap: stop after 30 iterations even if PRD items remain.
automationhigh risk

Ship production-grade apps autonomously

Loop/ralphGitHubB

Hand an idea to Claude Code; it authors specs, designs, builds, tests, secures, and ships until enterprise done or budget exhausted.

prompt
→ Claude Code
# dare-to-be-stupid — Design (v2, refined) > A Claude Code plugin. One command, /dare , hands an idea or PRD to an autonomous > loop that authors specs, designs, builds, tests, secures, ships, fixes, and iterates > until the app passes an enterprise-production definition of done — or the budget dies. > > Named for the Weird Al song. The joke is that it runs the Ralph Loop on purpose , > with --dangerously-skip-permissions , and narrates the whole thing in the voice of an > '80s Junkion. Pre-production only. Never points at anything with users. This is v2. It keeps the strong core of the original spec (external reviewer, ratchet, guard hook, Junkion style) and adds the three phases the original left thin relative to the actual goal: PRD authoring, a design phase, and a real enterprise DoD including security, CI, docs/observability, and design quality (with quality plugins auto-installed). --- ## 0. The premise, in one paragraph The User builds documentation-first: spec → system docs → API contracts → CLAUDE.md → code. dare-to-be-stupid is the deliberate inverse, packaged as comedy that also solves two real engineering problems. It is a real build , not a joke ar

Set and ship autonomous goals

Loop/goalGitHubB

Run a persistent goal autonomously until completion, routing each task to the optimal model for cost and capability.

prompt
→ Claude Code
/goal <objective · /loop <task | Set a persistent goal · run autonomously until complete (≤25 turns) |
automationlow risk