/goal — run until a condition is true

Set a completion condition; a fast evaluator model checks the transcript after each turn and the agent keeps going until verified done. No built-in turn cap — put "stop after N turns" inside the condition itself.

/goal runs until a condition is true: you write the end state, an evaluator model checks the transcript after each turn, and the agent keeps working until the check passes. The goal examples below show what a verifiable end state actually looks like — “all tests pass and lint is clean”, “the PR is merged”, “the migration applies with zero errors” — because the difference between a good /goal and a runaway one is whether the condition can be checked, not vibed. There’s no built-in turn cap, so the strongest pattern here is putting “stop after N turns” inside the condition itself. Start with the goal-examples guide for the full pattern language, or /goal vs /loop vs Ralph if you’re choosing a mechanism.

Draft a sprint plan from the backlog

Loop/goallooprepoB

Turn the open issue backlog into a proposed two-week sprint plan with estimates, a dependency ordering, and an explicit cut line, written as a document for the team to edit.

prompt
→ Claude Code
/goal SPRINT-PLAN.md contains a proposed 2-week plan — read all open issues labeled `ready`, estimate each as S/M/L based on the code it touches, order them by dependency and value, draw a cut line at a realistic capacity, and list what falls below it with reasons; make no changes to the issues themselves; stop after 6 turns
planninglow risk
⧉ 1

N+1 query hunt

Loop/goallooprepoB

Instrument the test suite with query logging, find endpoints issuing N+1 queries, and fix them with eager loading or batching until the hot paths are clean.

prompt
→ Claude Code
/goal no endpoint in the integration test suite issues more than 10 SQL queries per request — enable query logging in the test environment, find the worst N+1 offender, fix it with eager loading or a batched query, verify the count dropped and tests still pass, then move to the next; stop after 12 turns
databasemedium risk

Dual-reviewer convergence gate before opening a PR

Loop/goalGitHubAnew

Two reviewers from different model families, each in a fresh context, must both pass an objective rubric before a PR is opened — the maker-checker split with real model diversity. Each failed round fixes only what was flagged and re-reviews with reviewers that have no memory of the last round, so nothing anchors on prior findings. Hard cap of three rounds, tests may never be edited to pass, and the loop ends by opening a PR rather than pushing. Adapted from ECC's /santa-loop (affaan-m/ECC, MIT), which auto-pushes on agreement and has no verifiable exit; this version adds the test-exit condition and the human merge gate.

prompt
→ Claude Code
/goal Ship the current diff only after two independent reviewers both PASS. Scope is `git diff --name-only HEAD` (or the path in $ARGUMENTS). First write a rubric with objective PASS/FAIL criteria: correctness, security (no secrets, injection, OWASP top 10), error handling, completeness, internal consistency, no regressions. Each round launch two reviewers in parallel with fresh context and no memory of earlier rounds: Reviewer A is a Claude subagent, Reviewer B is a different model via `codex exec --sandbox read-only` (fall back to a second Claude subagent and say so in the report). Both return a JSON verdict with per-criterion PASS/FAIL and critical issues. If either FAILs: fix only the flagged issues with minimal diffs, never modify the tests to make them pass, commit "fix: address review findings (round N)", run `npm test`, and re-review with fresh reviewers. Exit when both reviewers PASS and `npm test` exits 0 — then open a PR for me to merge; never push to main or merge. Stop after 3 rounds; if still failing, print the unresolved issues and escalate to me instead of shipping.

Hit acceptance criteria

Loop/goalXB

Drive a feature to done against explicit acceptance criteria: a working paginated endpoint, passing tests, clean lint, and a hard turn cap.

prompt
→ Claude Code
/goal the /users endpoint returns 200 with a paginated JSON body, all tests pass, and lint is clean — stop after 20 turns

Kill-criteria loop

Loop/goalXB

Make every live option declare the single fact that would disqualify it upfront, then hunt for that evidence.

prompt
→ Claude Code
/goal for each live option [LIST], state its ONE disqualifying condition upfront, then run an evidence hunt for that disqualifier. Done when: for every option, the state-file kill-criteria.md contains either PASTED evidence the disqualifier is true (kill it) or a documented search showing it isn't (keep it). No option stays undecided. Paste the evidence table as proof, or paste what's still missing and stop. Judge: a second smaller model checks each verdict cites evidence. Budget: cap $[X]. Hard cap: stop after 1 iteration per run.
researchlow risk

Execute roadmap phases to completion

Loop/goalGitHubB

Work through each phase in your roadmap, verify each one, run the final audit, and stop when all phases pass.

prompt
→ Claude Code
/goal "Execute all phases of <run-root /ROADMAP.md sequentially. Read <run-root /phases/phase-N.md for each phase; do the work; run mandatory commands; print SUPERGOAL PHASE VERIFY then SUPERGOAL PHASE DONE for each phase; follow the failure-recovery protocol in <run-root /PROTOCOL.md if any criterion fails. After the last phase, run the FINAL AUDIT in <run-root /PROTOCOL.md (re-verify against <run-root /ROADMAP.md; re-run aggregated mandatory commands; spot-check criteria; on gaps, write <run-root /phases/audit-fix-<round .md and execute inline). Only after AUDIT COMPLETE, print SUPERGOAL RUN COMPLETE. Done when SUPERGOAL RUN COMPLETE appears in the transcript with one SUPERGOAL PHASE DONE per phase, AUDIT COMPLETE printed before SUPERGOAL RUN COMPLETE, and no FAILURE HANDOFF or AUDIT HANDOFF this run

Pre-mortem loop

Loop/goalXB

Before committing, write the 'it's 12 months later and this failed' story so the failure modes are on the table now.

prompt
→ Claude Code
/goal for decision [X], write the pre-mortem: assume it's 12 months later and this failed, then enumerate the specific causes, ranked by likelihood, each with an early-warning signal and a mitigation. Done when: state-file premortem.md contains the failure narrative + a ranked cause table with signals and mitigations, and the top 3 causes each have a concrete owner/mitigation PASTED in. Paste the completed table as proof, or paste what's unfinished and stop. Budget: cap $[X]. Hard cap: stop after 1 iteration per run.
researchlow risk

Set and ship autonomous goals

Loop/goalGitHubB

Run a persistent goal autonomously until completion, routing each task to the optimal model for cost and capability.

prompt
→ Claude Code
/goal <objective · /loop <task | Set a persistent goal · run autonomously until complete (≤25 turns) |
automationlow risk

Test all endpoints

Loop/goalGitHubB

Run tests against every API endpoint until coverage is complete or you hit 30 iterations.

prompt
→ Claude Code
/goal all endpoints tested or stop after 30 turns
testingmedium risk

Get the build green

Loop/goalcommunityB

Run the build, fix the first error, and repeat until npm run build exits 0, with a 10-turn cap.

prompt
→ Claude Code
/goal `npm run build` exits 0 — run the build, fix the first error, repeat until it succeeds; stop after 10 turns
Sponsored · ConnectMyEmail

Loops that read your inbox.

Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.

connectmyemail.com →

Refactor auth module, pass tests

Loop/goalGitHubB

Refactor the authentication module iteratively until all tests pass, stopping after ten attempts.

prompt
→ Claude Code
/goal "Refactor auth module until all tests pass" --max-iterations 10
refactoringlow risk

Measure and ship performance gains

Loop/goalGitHubB

Discover and execute a benchmark optimization goal, freeze the protocol, compare candidates under identical conditions, and deploy only validated improvements with regression protection.

prompt
→ Claude Code
/goal Use the installed shape-goal and goal-engine skills to discover, approve, and complete this repository's next Measured Optimization / Benchmark objective. During shaping, load shape-goal's required-input specification for Measured Optimization / Benchmark; exhaustively inspect repository instructions, Git state and history, requirements, architecture, plans, tests and CI, runtime behavior, prior goal state, the project harness, and any connected authoritative sources before asking the user. Resolve every material input from evidence where possible; ask only unresolved owner decisions, one at a time with a recommended answer, and do not make production changes until the user approves a Goal Contract. Then hand off within this same goal to goal-engine to freeze the benchmark protocol, compare champion and challengers under identical conditions, and retain only meaningful improvements; apply relevant assurance overlays, repository-native verification, regression protection, independent review where warranted, durable progress state, and reusable closeout. Do not declare success when shaping is complete. Finish only when every approved acceptance and overlay gate passes with surfaced evidence and protected behavior has not regressed. Stop only for a contract-defined blocker, approval boundary, budget, material goal drift, or two consecutive no-progress cycles
performancemedium risk

Upgrade to current Node LTS

Loop/goallooprepoB

Move the project to the current Node LTS across .nvmrc, CI config, Dockerfiles, and engines, fixing deprecations until everything is green on the new runtime.

prompt
→ Claude Code
/goal the project runs on the current Node LTS — update .nvmrc, the engines field, CI workflow files, and any Dockerfile base images to the LTS version, then run install, build, lint, and the full test suite on it, fixing deprecation warnings and breakages one at a time; stop when all are green or after 15 turns

Fix code review blockers until convergence

Loop/goalGitHubBnew

Run a full review pass inline, apply validated blocking findings at their owning layer, re-run fresh until the gate clears or blockers plateau—capped at 3 rounds with escalation.

prompt
→ Claude Code
/goal why-review fix-loop convergence: repeatedly run the full-mode /why-review pass INLINE over {target}. After each review, apply only VALIDATED findings that block the current round (all open severities in round 1 — Round-1 LOW closure; CRITICAL/HIGH/MEDIUM from round 2) at their owning layer, then re-run a FRESH full review pass over the CHANGED target. If the current round's blocking findings >0 → apply fixes and run another round; if a fresh full review pass clears the current bar (round 1: zero open findings; round 2: zero CRITICAL/HIGH/MEDIUM, LOW deferred) → CONVERGED, clear the gate. Do NOT open another round for LOW-only findings from round 2. Cap at {N=3} review rounds; a failing test gate is outside the cap and the no-progress rule — keep fixing and re-running until the tests pass, never forcing green; if review blockers do not shrink across 2 consecutive rounds or increase, or the review budget (round 3) is spent with review blockers still open → STOP and escalate via AskUserQuestion. Never loop open-ended
reviewmedium risk

Ship GOALS.md phases 1-13

Loop/goalGitHubB

Implement each GOALS.md phase with tests and validation, committing and pushing after stable milestones until unblocked.

prompt
→ Claude Code
/goal Complete GOALS.md phases 1-13 in order. For each phase, implement the deliverables, add/update tests, run the common validation plus that phase's Automated QA, commit after the phase passes, and push after stable milestones. Preserve unrelated user changes. Stop only if blocked by missing credentials, external service access, or an explicit product decision that cannot be safely inferred Cap the run at 20 turns.
planninghigh risk

Apply database migrations cleanly

Loop/goalcommunityB

Run migrations, fix schema or SQL errors, and repeat until prisma migrate status reports clean, capped at 6 turns.

prompt
→ Claude Code
/goal all database migrations apply cleanly — run them, fix schema or SQL errors, repeat until `npx prisma migrate status` is clean; stop after 6 turns
databasemedium risk

Migrate an API import by import

Loop/goalXB

Sweep a codebase from a legacy API to its v2 replacement with tests and typecheck as the safety net, capped at 30 turns.

prompt
→ Claude Code
/goal every file importing from `./legacy-api` now imports from `./v2-api`, all tests pass, and `npm run typecheck` is clean — stop after 30 turns

Review and self-fix changes to clean

Loop/goalGitHubB

Review diffs iteratively, validate findings, self-fix each one, and restart until zero critical/high/medium issues remain.

prompt
→ Claude Code
/goal changes-review self-recursive loop: review the full diff → run /why-review --validate-findings on every finding → SELF-FIX each validated finding → restart /changes-review from Phase 0 over the WHOLE updated diff (combined with the prior fixes, not just the last fix) → loop until one complete review pass clears that round's bar (rounds 1-2: zero findings; round 3+: zero CRITICAL/HIGH/MEDIUM, a LOW-only round ENDS the loop with the LOWs recorded as deferred) → then run Phase 7.5: one standalone FULL-mode /why-review over the whole target+diff (NOT --validate-findings), fixing and re-running until it is clean → only then run the Phase 8 /docs-update. Do not stop while any validated finding is unfixed, any review pass is non-clean, or the holistic full-mode /why-review has unaddressed findings

Source-verified brief loop

Loop/goalXB

Write a one-page brief where every claim has three or more sources and every link is opened and confirmed to support the claim — the loop that catches hallucinated citations a single prompt never can.

prompt
→ Claude Code
/goal write a one-page brief on [TOPIC] where every claim has at least three sources and every link opens to a real page that supports the claim. Open each link to confirm it before you call it done. Replace any source that is dead or does not back up the claim. Done when every source check passes on every claim — maximum 30 iterations.
researchlow risk

Dead code elimination

Loop/goallooprepoB

Hunt down unreferenced exports, unused files, and unreachable branches, deleting them in small verified steps until the analyzer reports clean.

prompt
→ Claude Code
/goal `npx knip` reports no unused files or exports — remove one cluster of dead code at a time, run the full test suite and typecheck after each removal, and revert any deletion that breaks them; stop after 15 turns
Sponsored · ConnectMyEmail

Loops that read your inbox.

Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.

connectmyemail.com →

Secrets scan until clean

Loop/goallooprepoB

Run a secrets scanner over the working tree and drive the findings to zero: real secrets get flagged for rotation, false positives get baselined.

prompt
→ Claude Code
/goal `gitleaks detect --no-git` reports zero findings — for each finding, tell me whether it looks like a real credential (flag it for rotation and replace it with an env var lookup) or a false positive (add it to the baseline with a comment); never print the secret value itself; stop after 8 turns

Get the build green, every time

Loop/goalGitHubB

Work through each task in .toh/plan.md, fix blockers, until the build command succeeds or you hit 40 turns.

prompt
→ Claude Code
/goal every task in .toh/plan.md is checked and the build command exits 0 — or stop after 40 turns . A Haiku evaluator judges the condition FROM THE TRANSCRIPT — one more reason the QC gate quotes actual output: unquoted results are invisible to the evaluator. |

API contract test backfill

Loop/goallooprepoB

Generate a contract test for every documented endpoint in the OpenAPI spec so the spec and the implementation can never silently drift.

prompt
→ Claude Code
/goal every path in openapi.yaml has a contract test asserting its status codes and response schema — add tests for one untested endpoint per turn, run the suite, and fix either the spec or the handler when they disagree (tell me which you chose); stop when all paths are covered or after 15 turns

Run flow until gate or timeout

Loop/goalGitHubB

Execute a flow step repeatedly until it signals DONE or GATE, or halt after 40 turns to proceed.

prompt
→ Claude Code
/goal FLOW says DONE or GATE, or stop after 40 turns then /flow-next

Clean up the slop

Loop/goalcommunityA

Review your recent diff for debug code, dead branches, and bad names, then fix with minimal edits until lint and tests pass.

prompt
→ Claude Code
/goal the recent diff is clean and convention-aligned — review it for debug code, dead branches, and bad names, fix with minimal edits until `npm run lint && npm test` passes; stop after 4 turns

Ship infrastructure readiness gates

Loop/goalGitHubB

Shape and execute infrastructure readiness objectives by gathering evidence, seeking approval at decision boundaries, then validating provisioning, configuration, and rollback in non-production before production deployment.

prompt
→ Claude Code
/goal Use the installed shape-goal and goal-engine skills to discover, approve, and complete this repository's next Infrastructure / Deployment Readiness objective. During shaping, load shape-goal's required-input specification for Infrastructure / Deployment Readiness; exhaustively inspect repository instructions, Git state and history, infrastructure-as-code, environment and secret references, build artifacts, deployment workflows, health checks, runbooks, rollback paths, supported environments, prior incidents and goals, the project harness, and connected authoritative systems before asking the user. Resolve every material input from evidence where possible. Continue inside this /goal only when an already-approved Goal Contract or authoritative artifact resolves every owner decision. Otherwise create or resume SHAPING.md, save the unresolved decision and one recommended question, stop as Approval required, and tell the user to resume shape-goal outside /goal; do not ask the question or take another autonomous turn, and do not make production changes before approval. Then hand off within this same goal to goal-engine to reconcile infrastructure and application assumptions, validate provisioning and configuration in approved non-production or simulated environments, verify artifact provenance, migrations, smoke and health gates, observability, failure handling, and rollback, and remove verified readiness blockers; apply relevant assurance overlays, repository-native verification, regression protection, independent review where warranted, durable progress state, and reusable closeout. Do not declare success when shaping is complete. Finish only when every approved readiness gate passes with surfaced evidence, environment differences and residual risks are documented, rollback remains viable, and no production deployment or mutation has occurred without explicit authority. Stop only for a contract-defined blocker, approval boundary, budget, material goal drift, or two consecutive no-progress cycles

Lock down Supabase RLS policies

Loop/goalGitHubB

Replace overpermissive 'always true' policies with org-scoped RLS across six tables until security advisor clears all findings.

prompt
→ Claude Code
/goal In Supabase prod project udooysjajglluvuxkijp, replace each authenticated write <table> ALL policy on public.customers/orders/order items/quotes/quote items/products (currently USING + WITH CHECK both literally true) with an org/tenant-scoped USING + WITH CHECK, or drop the policy if the table is unused in RA. End state: get advisors(project id=udooysjajglluvuxkijp, type:security) returns 0 rls policy always true findings for those 6 tables. Or stop after 6 turns if the owning tenant column cannot be confirmed
securityhigh risk

Set "remove every TODO comment in src/ and…

Loop/goalGitHubB

Community goal loop for docs, sourced from github. Verified exit condition, evaluator-gated.

prompt
→ Claude Code
/goal set "remove every TODO comment in src/ and explain each removal" stop after 8 turns

Commit message hygiene on a branch

Loop/goallooprepoA

Rewrite the commit messages on your feature branch to conventional-commit format with meaningful bodies before opening the PR, leaving the code untouched.

prompt
→ Claude Code
/goal every commit on this branch (ahead of main) has a conventional-commit subject under 72 characters and a body explaining why — use interactive rebase to reword only (no code changes, no commits dropped), verify with `git log main..HEAD`, and confirm the diff against the original branch tip is empty; stop after 5 turns

Audit code diffs, ship clean

Loop/goalGitHubB

Run code review on each step's diff, stop when zero findings or after 3 rounds.

prompt
→ Claude Code
/goal /codex review reports zero real-or-regression findings on every step's diff (the verdict pasted in full each round); or stop after 3 rounds, reporting anything unresolved
Sponsored · ConnectMyEmail

Loops that read your inbox.

Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.

connectmyemail.com →

Work goals to completion

Loop/goalGitHubB

Process implementation goals in order from a GOALMAP, stopping to ask before high-risk steps, and record evidence recipes confirming each Done Means item.

prompt
→ Claude Code
/goal Read context/goals/GOALMAP.md, its product loop, global constraints, and global invariants. Work the goals one at a time in status order. For each goal, read the brief. If Interpretation Risk is High or any Stop / Ask Condition is met, stop and ask before proceeding; do not guess to maintain momentum. Otherwise do the work needed to satisfy Done Means while preserving Constraints and Invariants. Record Evidence as a reproducible recipe: exact command/workflow/artifact/source review plus expected result a reviewer can re-run or inspect. The Evidence Recipe must confirm each Done Means item, including the highest-risk one, not an easier adjacent claim. Update Status and Completion Notes as bookkeeping only. When all implementation goals are done, run the final validation goal and present its evidence recipe/results plus the status table for acceptance; do not self-close the map
planninghigh risk

Memory leak hunt

Loop/goallooprepoB

Drive a suspected memory leak to ground: reproduce growth under a repeated workload, capture heap snapshots, and fix the retention until memory stays flat.

prompt
→ Claude Code
/goal heap usage stays flat (within 5%) across 500 repetitions of the failing workload in the leak-repro script — capture heap snapshots before and after, identify what is being retained and by which reference chain, fix the leak, and re-run the repro to confirm; stop after 10 turns
debuggingmedium risk

Get the fix to score 90+

Loop/goalGitHubA

Run verify commands and benchmark checks until the target code scores 90 or higher on the benchmark, stopping after 5 attempts.

prompt
→ Claude Code
/goal the transcript reports a line matching SCORE: <n>/100 (threshold: 90, gate: pr) with n >= 90 for <slug> on branch <branch>, after <(bugfix) the REPRODUCE test is shown failing on unfixed code and passing on the final code, and> <verify commands, comma-separated> <(hot-path) and make bench-compare> all pass on the final code, or stop after 5 fix rounds
testingmedium risk

Complete goal with skill stack

Loop/goalGitHubB

Read goal.md, follow your skill guides, meet all acceptance criteria, and stop after 20 turns or verification passes.

prompt
→ Claude Code
/goal Read goal.md and follow CLAUDE.md plus .claude/skills/setgoal/SKILL.md, complete all acceptance criteria, include verifier PASS and command outputs in the transcript, stop after 20 turns if blocked
planninglow risk

Upgrade dependencies one at a time

Loop/goallooprepoB

Walk through outdated dependencies one package per turn, upgrading, running the full check suite, and pinning back anything that breaks.

prompt
→ Claude Code
/goal `npm outdated` lists no minor or patch updates — upgrade exactly one package per turn, run tests, lint, and build after each, commit if green, and pin the previous version with a note in UPGRADE-BLOCKERS.md if it fails; stop after 20 turns

Get lint and E2E tests passing

Loop/goalGitHubA

Execute work from PLAN.md until npm run lint and npm run test:e2e pass, scoping changes per AGENTS.md.

prompt
→ Claude Code
/goal Implement the work described in PLAN.md. Stop only when npm run lint and npm run test:e2e pass. Follow AGENTS.md, keep changes scoped, and report verification evidence Stop after 25 turns even if the goal is not reached.
cimedium risk

Make tests pass, gate frozen

Loop/goalGitHubA

Run verify.sh repeatedly until it exits 0, all acceptance gates hold, and test quality gates are met, or report after 20 turns.

prompt
→ Claude Code
/goal ./scripts/verify.sh exits 0, .ai/spec-tdd/state.json phase is done, frozen tests and acceptance gates are unchanged, no tests are skipped/weakened, and no TODO/stub/hardcoded test-only implementation remains; or stop after 20 turns with a clear blocked report
testingmedium risk

Index every foreign key

Loop/goallooprepoA

Query the catalog for unindexed foreign keys, create a migration per missing index, and verify each with a full test run until all FKs are covered or 6 turns complete.

prompt
→ Claude Code
/goal every foreign key in the schema has a covering index — query the catalog (pg_catalog or information_schema) to list FKs with no matching index, add one migration per missing index, and re-run the catalog check until it returns zero rows. After each index, apply the migration on a scratch database and run the full test suite to verify nothing regresses; stop after 6 turns. Only add indexes — never change constraints or table definitions — and propose the final migration set as a PR for review.
databasehigh risk

Changelog generation from commits

Loop/goallooprepoB

Turn the commit history since the last release tag into a human-readable, categorized CHANGELOG entry ready for the next version.

prompt
→ Claude Code
/goal CHANGELOG.md has a complete entry for the next release — read every commit since the last version tag, group changes into Added, Changed, Fixed, and Removed, write user-facing descriptions (not commit messages), link PR numbers, and flag anything that looks like a breaking change; stop after 5 turns

Run pilot verdicts until rejection

Loop/goalGitHubB

Keep executing /flow-next:pilot until it returns PILOT VERDICT=NO WORK or 20 turns elapse, routing deferrals to /flow-next:land and parking async questions.

prompt
→ Claude Code
/goal keep running /flow-next:pilot until it prints PILOT_VERDICT=NO_WORK, or stop after 20 turns - note PILOT_VERDICT=DEFERRED_TO_LAND is its own terminal (an all-done spec whose open PR land owns); route it to /flow-next:land, not a pilot re-run. In backlog mode the grammar also carries PILOT_VERDICT=ASKED <id> (<n>) - a durable park, not a stop: the loop simply continues to the next item next tick, and the human answers async in the spec / tracker
productmedium risk
Sponsored · ConnectMyEmail

Loops that read your inbox.

Email is the missing tool in your harness. ConnectMyEmail gives Claude Code and Codex a clean MCP into Gmail, Outlook, iCloud and IMAP — triage, drafts, follow-ups, on a loop.

connectmyemail.com →

Check active goals against rubric

Loop/goalGitHubB

Verify each goal in active.md has evidence attached, stopping after 25 turns or completion.

prompt
→ Claude Code
/goal all rubric items in .ultragoal/goals/active.md are checked with evidence, or stop after 25 turns
evaluationlow risk

Next.js 15 app deploys to a Vercel preview…

Loop/goalGitHubA

Community goal loop for devops, sourced from github. Verified exit condition, evaluator-gated.

prompt
→ Claude Code
/goal Next.js 15 app deploys to a Vercel preview URL returning 200, with brand tokens (colors + Plus Jakarta Sans/Inter/IBM Plex Mono) configured and Supabase magic-link auth gating /dashboard so logged-out users redirect to /login; you prove this by npm run build passing, the preview URL, and an incognito visit to /dashboard redirecting; do not add features beyond auth shell, do not change the locked stack; or stop after 100 turns

Ship avatar endpoint to spec

Loop/goalGitHubB

Build and test a new /users/:id/avatar route until it passes tests, lint, and matches the OpenAPI schema.

prompt
→ Claude Code
/goal until grep -q "users/:id/avatar" src/routes/users.ts AND pnpm test -- tests/integration/users.avatar.spec.ts passes AND grep -q "/users/{id}/avatar" openapi.yaml AND pnpm lint passes, or stop after 15 turns
testingmedium risk

Ship a PR until green

Loop/goalcommunityA

Implement a change, open the PR with gh, then keep fixing CI failures until every check passes, all in one goal loop.

prompt
→ Claude Code
/goal a PR is open for this change and every CI check passes — implement it, test locally, push, open the PR with `gh pr create`, then keep fixing failures (re-check with `gh pr checks`) until green; stop after 10 turns

Stabilize flaky tests for good

Measure the flakiness, fix one root cause at a time, and stop after a defined streak of stable full-suite runs.

prompt
→ Claude Code
Run [test suite] [N] times under the same conditions and list tests whose result changes. Fix the most frequent flake at its root cause—shared state, timing, ordering, or an external dependency—never with a blind sleep or retry. Run that test [N] times, then rerun the full suite. Repeat until [N] consecutive full-suite runs pass, progress stalls, or approval is required. Return each flake, root cause, fix, evidence, and justified quarantine.