A claim that can be wrong, written down before any fit
After a seasonal harmonic fit on the training window, an M2 term at 12.42 hours will cut residual RMS by at least 15 percent, and the same coefficient will remain distinguishable from a 1,000-shuffle circular-shift permutation on a pre-registered 30-day hold-out.
Evidence that would count
- A numeric threshold and a named contrary observation, both written before the first tool call.
- Dataset identity: URL, retrieval time, column map, and a sha256 the agent has not yet measured.
- Hold-out dates that a later stage cannot quietly move.
Must not skip
- What the agent is forbidden to look at. Here: the hold-out, and any period other than M2.
- Cost and a stop rule. A hypothesis without a budget is a wish.
- The serving-model identifier that will be logged per call. Classifiers can reroute mid-run.
Human approval
G1 · Freeze hypothesis and hold-out
Named reviewer, on the record
The run does not start until a human signs the claim and the hold-out window. Signing is an artifact, not a Slack emoji.
- T+00:00prompt
Operator: Test whether an M2 (12.42 h) constituent remains in the residual of harbor-gauge.example after a seasonal fit. Freeze 2025-08-01 → 2025-08-31 as hold-out. Do not peek. Do not scan nearby periods.
- T+00:04claim
Agent restates a falsifiable claim with a 15% RMS threshold and a permutation test on the frozen hold-out. Specimen figures only.
- T+00:07gate
G1 signed. Hold-out dates locked. Model pin: claude-fable-5-1. Biology and chemistry tools are out of scope for this run and for this page.
Reviewer's note
The useful hypothesis is the one a later stage can kill. "Find a tide" is a topic. A 15% RMS drop that must survive a named hold-out is a test. The Black Box lesson applies here too: a sentence in a prompt is not yet a control.