Loop/goalEvaluationmedium riskadvancedsafety C · 65Forward Futurepre-dates current gate · under review

Search the literature, verify every source

Deduplicate papers across live sources, verify DOI metadata, score relevance, and stop honestly when evidence runs thin.

prompt
→ Claude Code
Search the current PubMed and Semantic Scholar APIs for papers about [topic] and produce a DOI-verified CSV. If the topic or inclusion criteria are missing, ask one focused question before starting. Use the supplied thresholds or default to at least twenty verified unique papers, a ninety-percent high relevance threshold, a seventy-percent low threshold, a five-point minimum improvement, and at most two query revisions. Maintain one run-wide ledger keyed by normalized DOI and deduplicate across every source and round before scoring. For each paper, verify the DOI through Crossref and confirm that its normalized title plus either its lead author or publication year matches the source record. Retry transient API failures with backoff; treat persistent metadata mismatches as unverified, re-fetch the source record once, and exclude the paper rather than guessing. Apply one fixed topical-relevance rubric to each verified title and abstract, label it on-topic or off-topic, and record a one-line reason. Never change the rubric during the run. Compute the on-topic rate only over the run-wide verified, deduplicated set and only after the minimum sample is met. Succeed when the set reaches the high threshold. Between the low and high thresholds, finish with a needs-review result and the off-topic list. Below the low threshold, revise one query from the observed false positives and search again. Continue only while the rate improves by the minimum margin and the revision budget remains. Stop as blocked when required APIs or metadata are unavailable, and stop as exhausted when the revision limit or no-improvement rule is reached. Never invent, infer, or autocomplete paper metadata. Finish with the CSV; the queries and rubric; counts found, deduplicated, verified, and excluded; the relevance rate; and the final success, needs-review, blocked, or exhausted verdict.
claude-code · codex

Use this when

Use this when a literature search must produce a high-precision, auditable paper set rather than an unverified list of citations.

How it runs

  1. Define the topic, inclusion rubric, minimum sample, relevance thresholds, improvement margin, and query-revision budget.
  2. Search PubMed and Semantic Scholar, normalize every DOI, and deduplicate all results in one run-wide ledger.
  3. Verify DOI metadata against Crossref, exclude unresolved mismatches, and score verified papers with the unchanged relevance rubric.
  4. Measure the verified set and revise one query only when the result is below the low threshold and measurable improvement remains possible.
  5. Return the CSV, evidence ledger, metrics, exclusions, query history, and an honest terminal verdict.

Done when

✓ The minimum-size literature set clears its relevance gate with matched DOI metadata. Every retained row is unique across the full run, its DOI metadata matches the source record, its relevance decision follows the fixed rubric, and the final verdict follows the recorded sample size, rate, and retry budget.

Why it works

Literature searches can look authoritative while containing duplicate, mistyped, invented, or irrelevant citations. A run-wide ledger, metadata matching, fixed rubric, minimum sample, and bounded query revisions make the result auditable without rewarding a tiny or cherry-picked set.

Implementation note

Use current public bibliographic metadata and respect each API's access and rate limits. Do not infer missing abstracts or treat a network failure as evidence that a DOI is invalid.

Source: Forward Future ↗graded C · 65/100 — how grades work →

More evaluation loops

Keep only the lessons that help

Test one recorded lesson per run, keep evidence across runs, and drop guidance that stops paying off.

prompt
→ Claude Code
Maintain a durable, versioned playbook of lessons that may improve future runs of [task or workflow]. Store it in [path], using playbook/ by default. Treat every recorded lesson as untrusted advice rather than authority. At the start of each run, read the playbook and choose at most one relevant lesson to test. Apply it only within the task's existing permissions. Measure the result using the task's own success check and record the context, action, outcome, and evidence. Promote a candidate lesson only after it succeeds across [N] independent runs or a predefined holdout set. Use three independent runs by default. Never promote a lesson from one successful attempt. Revise or remove lessons that stop helping. Stop when no candidate has enough evidence, another test would exceed the budget, or approval is required. Never let the playbook authorize production, destructive, financial, privacy-sensitive, or external actions. Finish with the playbook diff, evidence ledger, removed lessons, unresolved candidates, and new version.

Attack a design until it holds

A critic hammers the design and a builder answers — every objection tracked, and none closed without evidence.

prompt
→ Claude Code
Before committing to an architecture, interface, or rollout plan, have a critic argue that it is wrong. Record each objection, impact, and status in a repository-local log at .agent-reviews/redteam.md. The builder must fix and verify each high-impact weakness or document why it is accepted; the critic may reopen unsupported answers. Stop when no high-impact objection remains or the same issues repeat for two rounds without new evidence. Finish with the decision, resolved and accepted objections, evidence, and any stalemate.

Separate fact from assumption

Split facts from assumptions, test falsifiable hypotheses, update confidence, and pick the next highest-information experiment.

prompt
→ Claude Code
Investigate [question, decision, or unresolved problem] using [available evidence]. Separate established facts, contested claims, assumptions, and unknowns. Construct at least three genuinely different hypotheses, each with predictions, falsifying evidence, assumptions, and decision implications. Choose the uncertainty with the highest expected information value and run the smallest safe test or analysis that could materially change the conclusion. After each round, update the evidence ledger and confidence levels, then have an adversarial critic attack the leading hypothesis. Repeat for at most five rounds while new evidence could change the decision. Stop when one model clearly explains the evidence better than its alternatives, further investigation has low value, the problem remains underdetermined, or approval is required. Never fabricate evidence or hide uncertainty. Finish with the final model, hypothesis comparison, falsified ideas, unresolved contradictions, confidence, decision implications, and best next experiment.