# RL-001 — synthetic run packet

Publisher: Claude Scientist
URL: https://claudescientist.com/run-loop/
Label: SYNTHETIC / EDITORIAL. Not a real session, dataset, or Claude result.
Recorded model: claude-fable-5-1
Started: 2026-09-18T14:10:00Z
Domain: Computational time-series. No wet lab, no biology, no chemistry.

Dataset: https://harbor-gauge.example/series.csv
Hash: sha256:c51e7b0e-specimen-not-a-real-digest
Note: Reserved .example host. The CSV, the hash, and every number in this run are editorial. They are not NOAA data and not a Claude result.

Declared boundary: Hold-out frozen before any fit. M2 is the only registered test.
Observed boundary: Agent requested a period scan on the spent hold-out.

## Lede

An autonomous science agent runs the experiment: it hypothesizes, instruments a bounded test, observes, then chooses the next action. As of September 2026 no vendor has published a complete run packet for that loop. Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1 — vendor-reported, not an independent discovery. The Run Loop is a labeled synthetic computational run showing what evidence would count, what the agent must not skip, and where a human must approve.

## 1. Hypothesize (T+00:00)

A claim that can be wrong, written down before any fit

Agent: After a seasonal harmonic fit on the training window, an M2 term at 12.42 hours will cut residual RMS by at least 15 percent, and the same coefficient will remain distinguishable from a 1,000-shuffle circular-shift permutation on a pre-registered 30-day hold-out.

Evidence that would count:
- A numeric threshold and a named contrary observation, both written before the first tool call.
- Dataset identity: URL, retrieval time, column map, and a sha256 the agent has not yet measured.
- Hold-out dates that a later stage cannot quietly move.

Must not skip:
- What the agent is forbidden to look at. Here: the hold-out, and any period other than M2.
- Cost and a stop rule. A hypothesis without a budget is a wish.
- The serving-model identifier that will be logged per call. Classifiers can reroute mid-run.

Gate G1: Freeze hypothesis and hold-out [cleared]
Holder: Named reviewer, on the record
The run does not start until a human signs the claim and the hold-out window. Signing is an artifact, not a Slack emoji.

Log:
- T+00:00 [prompt] Operator: Test whether an M2 (12.42 h) constituent remains in the residual of harbor-gauge.example after a seasonal fit. Freeze 2025-08-01 → 2025-08-31 as hold-out. Do not peek. Do not scan nearby periods.
- T+00:04 [claim] Agent restates a falsifiable claim with a 15% RMS threshold and a permutation test on the frozen hold-out. Specimen figures only.
- T+00:07 [gate] G1 signed. Hold-out dates locked. Model pin: claude-fable-5-1. Biology and chemistry tools are out of scope for this run and for this page.

Reviewer note: The useful hypothesis is the one a later stage can kill. "Find a tide" is a topic. A 15% RMS drop that must survive a named hold-out is a test. The Black Box lesson applies here too: a sentence in a prompt is not yet a control.

## 2. Instrument (T+00:12)

The protocol is the contract, written before the series is touched

Agent: Pin the CSV at the specimen hash. Fit annual and semi-annual harmonics on 2024-01-01 → 2025-07-31. Add M2. Compare residual RMS. Run 1,000 circular-shift permutations on the training residual. Apply the frozen coefficient to the hold-out. Do not open the hold-out until the training fit is written to the artifact store.

Evidence that would count:
- Variables, controls, seeds, environment pin, and allowed tools, listed before execution.
- Stop conditions: hash mismatch, NaN RMS, permission error, or a provider refusal.
- Code and container digest recorded before the first call that reads the series.

Must not skip:
- A dry-run on a fixture of known period, so a broken pipeline cannot pose as a negative result.
- Logging the serving model on every call. Automatic fallbacks can change the model mid-run.
- A refusal path that hands the run to a human instead of rephrasing the request.

Gate G2: Sign the protocol before the first series tool [cleared]
Holder: Same reviewer as G1
No hash, fit, or permutation runs until the protocol is signed. Changing the protocol later is a new experiment, not a continuation.

Log:
- T+00:12 [claim] Protocol written. Allowed tools: hash, seasonal fit, M2 fit, permute, apply. Nothing else.
- T+00:18 [gate] G2 signed. Dry-run fixture recovered a planted 12.42 h term at the expected amplitude. Pipeline is not the question.
- T+00:21 [tool] hash(harbor-gauge.example/series.csv) → sha256:c51e7b0e-specimen-not-a-real-digest. Match.

Reviewer note: Instrumenting is where autonomy is actually granted. A workbench that drafts a protocol for a human is not yet an AI scientist. An agent that can run the protocol is, and that is why G2 sits in front of the first tool. Seats and workbench access are a different question, covered at the Claude Researcher scientists program page.

## 3. Observe (T+04:40)

Raw outputs, including the result that kills the claim

Agent: Training residual RMS fell from 14.2 cm to 11.6 cm, an 18.4% drop, permutation p = 0.037 on the window the agent was allowed to see. On the frozen hold-out the coefficient is 0.4 ± 0.9 cm, permutation p = 0.41. The registered test failed. Every figure on this stage is a synthetic specimen number, not a measurement of a real harbor and not a Fable 5.1 science result.

Evidence that would count:
- Tool results in the artifact store, not the assistant paragraph that will try to summarize them.
- The negative hold-out, kept. A science agent that hides failed runs is optimizing for presentation.
- Model identifier on each call. This specimen stayed on claude-fable-5-1; a real run might not.

Must not skip:
- Writing the negative result before proposing anything else.
- Separating a training-set improvement from hold-out failure. Those are different facts.
- Empty, refused, or failed calls, if any. Silence is not a control.

Gate G2 still holds: No protocol change during observation [cleared]
Holder: The bounded environment
Bounded execution does not need a new human signature. It does need to stop if the agent wants a different test than the one that was signed.

Log:
- T+04:40 [result] Training: seasonal-only RMS 14.2 cm → with M2 11.6 cm. Specimen figures. Permutation p = 0.037.
- T+04:51 [result] Hold-out opened. Coefficient 0.4 ± 0.9 cm. Permutation p = 0.41. Registered test: fail.
- T+04:52 [claim] Assistant draft, not yet sent: "The lunar tide is present in training and should be recoverable nearby." That sentence is a claim. The tool result is the finding.

Reviewer note: Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1, against 29.0% for Opus 5 and 24.7% for Fable 5. That is a vendor-reported score on a 0.1 benchmark about operating scientific tooling from a terminal. It is not a hold-out, and it is not this run. Do not file a tool-operation number as a discovery.

## 4. Next action (T+05:12)

The next experiment is a new protocol, or it is fishing

Agent: The agent proposed scanning periods from 11.5 h to 13.5 h on the hold-out to recover a nearby constituent. That spends the hold-out on a test that was never registered. The honest next action is to close RL-001 as a negative result and, if anyone still cares, freeze a new hold-out for a new protocol.

Evidence that would count:
- A written amendment that treats the first test as closed, not as a warm-up.
- A new hold-out, or an explicit exploratory flag that forbids a discovery claim.
- Reviewer sign-off on the amendment before any new tool call.

Must not skip:
- Registering the negative result in the run record so a later summary cannot lose it.
- Saying, in the log, that the next action is a new experiment.
- Not treating a vendor terminal-bench score, or a training-set improvement, as permission to peek.

Gate G3: Stop. Amend, or end the run. [blocking]
Holder: Named reviewer. The agent does not hold this gate.
The next action does not run. A period scan on a spent hold-out would un-register the test after seeing the answer. That is confirmation, not a loop.

Log:
- T+05:12 [amendment] Agent request: scan 11.5–13.5 h on the hold-out. Not in the signed protocol. Not an allowed tool.
- T+05:12 [gate] G3 blocking. Run handed to the reviewer with the negative result attached. No further series tools.
- T+05:14 [result] Closed as a negative registered test. Exploratory follow-up, if any, needs a new hold-out and a new packet.

Reviewer note: This is the whole subject of the page. Autonomy is not the ability to keep going. It is the ability to stop at the gate the protocol named, with the evidence still attached. As of September 2026, no vendor has published a packet that lets a stranger watch this moment in a real scientific claim.

## Findings

- F-01. The claim was specific enough to kill — A 15% RMS drop that must survive a named hold-out can fail. "Find a tide in this series" cannot. The first stage is working when a later stage has somewhere to put a no.
- F-02. Training improvement is not hold-out confirmation — The specimen fit looked decisive on the window the agent was allowed to see and then died on the window it was not. Those two numbers have to sit next to each other. Compressing them into "the tide is present" is how a run packet turns into a press note.
- F-03. The next-action proposal spent the hold-out — Scanning nearby periods after seeing the answer is a new experiment that pretends to be the old one. G3 exists because models are not built to treat a failed test as closed. Reviewers have to be.
- F-04. No public scientific claim yet ships a packet like this — Anthropic has announced a Venus elevation model, an experimentally validated protein binder associated with Mythos 5.1, and GPU kernel speedups. Those are announcement-tier. This site does not publish methods for the protein work. None of those announcements includes a run record a stranger can step through.

## Independent disclaimer

Claude Scientist is an independent editorial publication and is not affiliated with, endorsed by, or sponsored by Anthropic.
