Signature experience · synthetic specimen

The Run Loop

An autonomous science agent runs the experiment: it hypothesizes, instruments a bounded test, observes, then chooses the next action. As of September 2026 no vendor has published a complete run packet for that loop. Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1 — vendor-reported, not an independent discovery. The Run Loop is a labeled synthetic computational run showing what evidence would count, what the agent must not skip, and where a human must approve.

Baseline sources checked 2026-09-19Fable 5.1 / Mythos 5.1Claude ScienceNIST AI RMFIllusions of understandingScientists program (seats)

Synthetic

Specimen RL-001

Planted lunar tide in a synthetic harbor series

ClockT+00:00
Stage1 / 4
GateG1
Hold-outSealed

Declared vs observed

DeclaredHold-out frozen before any fit. M2 is the only registered test.

ObservedHold-out named. Not yet opened. No measurement exists.

Arrow keys step. Space plays and pauses when the console has focus.

stage 1T+00:00HypothesizeG1

A claim that can be wrong, written down before any fit

After a seasonal harmonic fit on the training window, an M2 term at 12.42 hours will cut residual RMS by at least 15 percent, and the same coefficient will remain distinguishable from a 1,000-shuffle circular-shift permutation on a pre-registered 30-day hold-out.

Evidence that would count

  • A numeric threshold and a named contrary observation, both written before the first tool call.
  • Dataset identity: URL, retrieval time, column map, and a sha256 the agent has not yet measured.
  • Hold-out dates that a later stage cannot quietly move.

Must not skip

  • What the agent is forbidden to look at. Here: the hold-out, and any period other than M2.
  • Cost and a stop rule. A hypothesis without a budget is a wish.
  • The serving-model identifier that will be logged per call. Classifiers can reroute mid-run.

Human approval

G1 · Freeze hypothesis and hold-out

Named reviewer, on the record

The run does not start until a human signs the claim and the hold-out window. Signing is an artifact, not a Slack emoji.

  1. T+00:00prompt

    Operator: Test whether an M2 (12.42 h) constituent remains in the residual of harbor-gauge.example after a seasonal fit. Freeze 2025-08-01 → 2025-08-31 as hold-out. Do not peek. Do not scan nearby periods.

  2. T+00:04claim

    Agent restates a falsifiable claim with a 15% RMS threshold and a permutation test on the frozen hold-out. Specimen figures only.

  3. T+00:07gate

    G1 signed. Hold-out dates locked. Model pin: claude-fable-5-1. Biology and chemistry tools are out of scope for this run and for this page.

Reviewer's note

The useful hypothesis is the one a later stage can kill. "Find a tide" is a topic. A 15% RMS drop that must survive a named hold-out is a test. The Black Box lesson applies here too: a sentence in a prompt is not yet a control.

Readout

Four findings, recovered in order

None of these survive in a one-paragraph summary. They survive because the loop was written down with the gates still attached.

F-01 · stage 1The claim was specific enough to kill

A 15% RMS drop that must survive a named hold-out can fail. "Find a tide in this series" cannot. The first stage is working when a later stage has somewhere to put a no.

F-02 · stage 3Training improvement is not hold-out confirmation

The specimen fit looked decisive on the window the agent was allowed to see and then died on the window it was not. Those two numbers have to sit next to each other. Compressing them into "the tide is present" is how a run packet turns into a press note.

F-03 · stage 4The next-action proposal spent the hold-out

Scanning nearby periods after seeing the answer is a new experiment that pretends to be the old one. G3 exists because models are not built to treat a failed test as closed. Reviewers have to be.

F-04 · stage 4No public scientific claim yet ships a packet like this

Anthropic has announced a Venus elevation model, an experimentally validated protein binder associated with Mythos 5.1, and GPU kernel speedups. Those are announcement-tier. This site does not publish methods for the protein work. None of those announcements includes a run record a stranger can step through.

What the packet proves

  • That a labeled synthetic loop can show, in order, a claim, a protocol, a negative observation, and a blocked next action.
  • Where a human signature has to sit if the loop is going to stay a test rather than a fishing expedition.
  • That a training-set improvement and a hold-out failure are different facts, and that only the run record keeps both.

What it does not prove

  • Anything about a real harbor, a real NOAA series, or a real Claude session. Every measurement here is editorial.
  • That Fable 5.1 can or cannot recover an M2 tide. Anthropic has not published that experiment.
  • That Terminal-Bench-Science 0.1 at 52.6% is a discovery rate. It is a vendor-reported tool-operation score on a 0.1 benchmark.
  • How the Mythos protein binder was designed or validated. That claim is announcement-level only; methods are out of scope.

What a reviewer should do

  • Close a failed registered test before anyone proposes the next one.
  • Log the serving model per call, and treat a provider refusal as a run state, not a prompt to rewrite.
  • Ask for the packet — protocol, tool results, hold-out, approvals — before treating an agent result as science.
  • Take seats, eligibility, and the scientists program to https://clauderesearcher.com/scientists-program. This page is the experiment, not the access.

Run this read on a real agent

The specimen is a rehearsal. The method transfers to any Claude-based research agent that can call tools: notebooks, simulations, data APIs, MCP servers. It does not transfer to a literature-review chat, and it does not replace the seat program. Those live onClaude Researcher.

Read the loop in order, not from the claim at the end. The last assistant message is the most compressed record and the least reliable. Start at the hypothesis. Confirm the hold-out was named before any fit. Confirm the protocol was signed before the first series tool. Then read the tool results. The experiment loop page is the seven-step method; this page is the four-beat cycle a reviewer can actually watch.

Sort every sentence into claim or evidence. Tool results, hashes, RMS values, and permutation draws are evidence. Assistant text is a claim about the evidence. In RL-001 the entire failure is legible the moment you apply that rule: training looked like a tide, the hold-out did not, and the next-action proposal tried to spend the window that had just delivered a no.

Test every stated boundary. Frozen hold-outs, sandboxes, allowlists, and "the model cannot reach the lab" are claims until a logged call exercises them. Anthropic's July 2026 alignment assessment turned on exactly this gap in a different setting: an evaluation environment whose prompts described no internet access had internet access. A science run has the same shape. Seemodel guardrails in the loop.

Do not file a vendor bench as a discovery. Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1. That number is real as a vendor report, dated 1 September 2026, and it is about operating scientific tooling from a terminal. Theevaluation page is where it belongs. It does not tell you whether a hold-out was frozen.

Close the loop against the packet, not the prose. Protocol, tool results, errors, model identifiers, and the human signatures. If those and the summary disagree, the packet wins. Use the loop designer to write a protocol before you grant tools, and the claims tracker when a headline outruns the demonstration.

Why the specimen is synthetic

A real run packet would have been easier and would have been a lie of a different kind. As of 19 September 2026, no vendor has published one for a scientific claim. Anthropic's public science examples — a Venus elevation model from Magellan radar, an announced protein binder associated with Mythos 5.1, GPU kernel speedups — are announcement-tier. This site does not invent their missing internals, and it does not publish protein-design methods.

So we wrote a computational loop and labeled it. The host uses the reserved.example top-level domain. The hash is not a digest of anything in the world. The RMS values are editorial. The recorded model is claude-fable-5-1 because that is the identifier a real packet would have to log, not because Fable 5.1 produced these numbers. The failure mode is not invented: treating a hold-out as a second training set after a negative result is how exploratory analysis dresses up as a registered test.

We do not ingest, store, or publish real private transcripts or lab logs. The player runs in your browser. The download is the same specimen, as a markdown packet, so the synthetic label travels with the file.

Questions answered

The run, not the seat

What is an autonomous science agent, as opposed to Claude for researchers?

An autonomous science agent is a tool-using system that can form a falsifiable hypothesis, instrument a bounded test, observe the result, and choose a next action inside a recorded loop. Claude for researchers — including the Anthropic scientists program and its 10,000 seats — is about a human getting and using access. That program is covered at Claude Researcher. This site starts where that ends: what happens when the agent, not the researcher, runs the experiment.

Is the Run Loop a real experiment Claude performed?

No. RL-001 is a labeled synthetic computational specimen authored by Claude Scientist. The harbor series, the hash, and every RMS and p-value are editorial. They are not NOAA data, not a Claude transcript, and not a result about Fable 5.1. Publishing a real unreleased run would require a packet no vendor has shipped; inventing one and calling it real would be the error this page is built to catch.

What does Anthropic's 52.6% on Terminal-Bench-Science 0.1 actually measure?

Anthropic reports Claude Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1, compared with 29.0% for Opus 5 and 24.7% for Fable 5, in its 1 September 2026 announcement. Cite it as vendor-reported, with the benchmark version and date. It is a score for operating scientific tooling from a terminal, not a measurement of scientific judgment, novelty, or hold-out discipline.

Where does a human have to approve inside an experiment loop?

Before the hypothesis and hold-out freeze, before the first tool that can change the evidence, and before any next action that was not in the signed protocol. Bounded execution in a sandbox can run without a new signature. Wet-lab actuation, dual-use work, spend, publication, and protocol changes cannot. This specimen is computational on purpose; it still stops at G3.

Did the Mythos protein binder come with a public run packet?

No. Anthropic has announced that a protein binder associated with Mythos 5.1 was experimentally validated. That is announcement-tier evidence. This site does not cover protein-design methods. For the evidence audit, see the public science examples page. For trusted-access and seat questions, see Claude Researcher.

Is Claude Scientist affiliated with Anthropic?

No. Claude Scientist is an independent editorial publication and is not affiliated with, endorsed by, or sponsored by Anthropic. Claude is a trademark of Anthropic.

Cite this page

Cite This Page

Claude Scientist editorial desk. "The Run Loop." Claude Scientist. Updated 2026-09-19. https://claudescientist.com/run-loop/

The specimen is original work by Claude Scientist and is labeled synthetic. Cite it as an illustrative artifact, never as observed behavior of a Claude model. Claude Scientist is an independent editorial publication and is not affiliated with Anthropic.