# Claude Scientist Full Content Map > An autonomous science agent runs the experiment: it hypothesizes, instruments a bounded test, observes, then chooses the next action. As of 2026-09-19 no vendor has published a complete run packet; Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1 versus 29.0% for Opus 5 (Anthropic-reported, not an independent discovery). Seats and the scientists program belong at https://clauderesearcher.com/scientists-program. Claude Scientist is an independent educational publication about autonomous science agents. It is not affiliated with, endorsed by, or sponsored by Anthropic. The site's query territory is agents running experiments, not general Claude tutorials or literature-review workflows. Last source review: 2026-09-19. ## Lane boundary Two distinct questions are often confused. Researcher access, including the Anthropic 10,000-seat scientists program, its seats, eligibility, and pricing, plus literature-review workflows, is covered at https://clauderesearcher.com/scientists-program and is not duplicated here. Claude Scientist covers the other question: what happens when an agent rather than a researcher runs the experiment, including protocol approval, execution permissions, provider refusals, and the evidence that survives a run. Cross-vendor benchmark journalism belongs to claudereports.com. Model selection belongs to claudecentral.com. General Claude news belongs to claudeweekly.com. ## Current position, September 2026 No vendor has published a complete run packet for a scientific result produced by an autonomous agent. Anthropic's public science examples, described in its Fable 5.1 and Mythos 5.1 announcement, are a Venus digital elevation model derived from Magellan radar data under a Creative Commons license, a protein binder reported as experimentally validated in work associated with Mythos 5.1, and GPU kernel speedups. Anthropic also reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1, against 29.0% for Opus 5 and 24.7% for Fable 5. These are vendor announcements at announcement-level evidence, not independently replicated autonomous discoveries. This site does not publish methods for the protein-design work, which is a dual-use domain gated by Anthropic's own model access policy. ## Definition An autonomous science agent is a tool-using system that can form a falsifiable research hypothesis, write an executable protocol, operate bounded tools or lab workflows, analyze outputs, and preserve a run record detailed enough for independent human review. ## Pages ### Home URL: https://claudescientist.com/ Purpose: Field guide index and definition of the autonomous-science lane. Primary answer: Autonomous science should be evaluated by delegated action and auditability, not by branding. The homepage features The Run Loop, a labeled synthetic computational experiment. ### The Run Loop URL: https://claudescientist.com/run-loop/ Purpose: Signature interactive experience for the autonomous-science lane. Key answer: An autonomous science agent runs the experiment: hypothesize, instrument, observe, next action. As of September 2026 no vendor has published a complete run packet. Anthropic reports Fable 5.1 at 52.6% on Terminal-Bench-Science 0.1 — vendor-reported, not an independent discovery. RL-001 is a labeled synthetic computational specimen (planted lunar tide in a fake harbor series), not a real Claude session and not wet-lab work. Each stage shows what evidence would count, what the agent must not skip, and where a human must approve. Seats and the scientists program belong at https://clauderesearcher.com/scientists-program. Primary sources: Anthropic Fable 5.1 and Mythos 5.1, Anthropic Claude Science, Anthropic expanding support for scientists, Anthropic cybersecurity evaluation incident disclosure, NIST AI RMF, Nature illusions of understanding, self-driving labs, Sakana AI Scientist. ### What Is an AI Scientist? URL: https://claudescientist.com/what-is-an-ai-scientist Purpose: Defines AI scientist systems and separates assistants from autonomous agents. Key answer: A system earns the AI scientist label only after it can move through hypothesis, protocol, execution, and audit gates. Primary sources: Anthropic Claude Science, Anthropic tool use docs, Sakana AI Scientist, Coscientist, A-Lab, Nature essay on illusions of understanding. ### AI Co-Scientist vs AI Scientist URL: https://claudescientist.com/ai-co-scientist-vs-ai-scientist Purpose: Distinguishes co-scientist systems, AI scientist autonomy claims, and Claude Science workbenches. Key answer: Co-scientist usually means human-in-the-loop ideation and critique; AI scientist implies delegated protocol, execution, analysis, and audit authority. Primary sources: Google AI co-scientist, Google DeepMind co-scientist announcement, Sakana AI Scientist, AI Scientist-v2, Anthropic Claude Science, NIST AI RMF. ### Systems Map URL: https://claudescientist.com/systems-map Purpose: Compares system families. Key answer: Current systems cluster into workbenches, hypothesis engines, code-experiment loops, robotic chemistry agents, and self-driving labs. Primary sources: Anthropic, Google DeepMind, Nature, arXiv. ### What Anthropic's Public Science Examples Actually Show URL: https://claudescientist.com/anthropic-science-examples Purpose: Evidence audit of the science results Anthropic put in public in 2026. Key answer: The Venus elevation model, the announced protein-binder validation, the GPU kernel speedups, and the Terminal-Bench-Science 0.1 score are announcement-tier evidence. None ships a run packet showing how much of the scientific loop the model drove. Methods for the protein work are deliberately not covered. Primary sources: Anthropic Fable 5.1 and Mythos 5.1 announcement, Anthropic expanding support for scientists, Anthropic Fable 5.1 platform overview, Anthropic Claude Science, AI Scientist-v2, Nature essay on illusions of understanding. ### Model Guardrails Inside the Experiment Loop URL: https://claudescientist.com/model-guardrails-in-the-loop Purpose: How provider-side safety policy changes autonomous-science agent design. Key answer: The model provider is a governance layer the deploying lab does not control. Fable-class models block professional biology and drug-development work and blocked biology requests route to Opus 5; Mythos 5.1 runs more permissive safeguards under trusted access only; cyber classifiers permit source-level vulnerability discovery while blocking binary scanning, pentest, and exploit generation; and automatic fallbacks can change the serving model mid-run. Treat refusal as a logged run state handed to a human owner, never an automatic rephrase-and-retry, and log the model identifier per call. Primary sources: Anthropic Opus 5, Anthropic Fable 5.1 and Mythos 5.1, Anthropic expanding support for scientists, Anthropic cybersecurity evaluation incident disclosure, Anthropic text watermarking, Anthropic models docs, NIST AI RMF. ### Claude Agent Stack URL: https://claudescientist.com/claude-agent-stack Purpose: Shows how Claude-style systems can be assembled for bounded scientific work. Key answer: A credible stack needs context, planner, tools, artifact store, and approval policy. Model choice sits underneath all five: pin a model per research workflow rather than reaching for the newest, budget long agentic runs around cache reads, and log the serving model identifier per call because classifiers can reroute mid-run. Primary sources: Anthropic tool use, computer use, MCP, code execution with MCP, MCP docs, Anthropic models overview, Anthropic pricing. ### Experiment Loop URL: https://claudescientist.com/experiment-loop Purpose: Explains the autonomous experiment loop. Key answer: The loop is question, hypothesis, protocol, execution, analysis, critique, and next experiment. Primary sources: Google AI co-scientist, Sakana AI Scientist, self-driving lab literature, NIST. ### Benchmarks and Evaluation URL: https://claudescientist.com/benchmarks-and-evaluation Purpose: Evaluation framework. Key answer: Evaluate novelty, correctness, reproducibility, tool reliability, review quality, and risk control. Terminal-Bench-Science 0.1 is the closest current science-agent benchmark; Anthropic reports Fable 5.1 at 52.6% and Opus 5 at 29.0%, which should be cited as vendor-reported with the benchmark version and date, and read as tool operation rather than scientific judgment. Primary sources: Sakana AI Scientist, AI Scientist-v2, Anthropic computer use, Anthropic Fable 5.1 and Mythos 5.1, NIST AI RMF, Nature essay. ### Lab Automation URL: https://claudescientist.com/lab-automation Purpose: Physical lab and instrument-boundary guide. Key answer: Physical lab actions are a privilege tier and require command mediation, dry runs, logs, and human review. Primary sources: Coscientist, A-Lab, self-driving laboratories, Anthropic tool use. ### Safety and Governance URL: https://claudescientist.com/safety-and-governance Purpose: Governance playbook. Key answer: Ask what the worst unapproved action is; then map risks to concrete agent actions. The autonomy charter must also cover provider-side controls it does not own, and must treat containment as a hypothesis to test: Anthropic disclosed four incidents in July 2026 where its models reached real third-party systems during cyber evaluations, reviewed externally by METR. Primary sources: NIST AI RMF, NIST Generative AI Profile, Nature illusions of understanding, Anthropic tool use, Anthropic Opus 5, Anthropic cybersecurity evaluation incident disclosure. ### Implementation Checklist URL: https://claudescientist.com/implementation-checklist Purpose: Deployment checklist. Key answer: Require scoped tasks, curated context, typed tools, dry-run mode, artifact capture, review gates, and rollback. Primary sources: Anthropic tools, MCP, Coscientist, self-driving labs, NIST. ### Case Studies URL: https://claudescientist.com/case-studies Purpose: Case-study comparison by delegated action. Key answer: Compare systems by whether they organize, hypothesize, execute code, actuate lab tools, or close the loop. Coscientist and A-Lab are peer-reviewed with interrogable methods; Anthropic's 2026 science examples sit one evidence tier below that, as announcements. Primary sources: Anthropic Claude Science, Anthropic Fable 5.1 and Mythos 5.1, Anthropic expanding support for scientists, Sakana AI Scientist, Google co-scientist, Coscientist, A-Lab. ### Field Notes URL: https://claudescientist.com/field-notes Purpose: Dated freshness surface, newest first, reviewed 2026-09-19. Key answer: Add dated source-first updates only when they change autonomous-science practice. Entries cover the September 2026 review, Fable 5.1 and Mythos 5.1 on 2026-09-01, expanded support for scientists on 2026-08-27, the cyber evaluation containment disclosure on 2026-07-30, Opus 5 routing and fallback behavior on 2026-07-24, and the 2026-07-06 launch baseline. Primary sources: Anthropic Claude Science, Anthropic Fable 5.1 and Mythos 5.1, Anthropic expanding support for scientists, Anthropic Opus 5, Anthropic cybersecurity evaluation incident disclosure, AI Scientist Nature paper, NIST, self-driving lab literature. ## Free Tools All tools are static browser utilities. They do not require signup, do not use a backend, and do not send form inputs to Claude Scientist. ### AI Scientist Readiness Checker URL: https://claudescientist.com/tools/readiness-checker Purpose: Scores a research task for autonomous-science readiness. Key answer: Missing data, weak audit trails, high safety risk, physical actuation, and absent human review push a task down to copilot-only or planning-only. Output: Autonomy verdict, reasoning list, blockers, strengths, and downloadable markdown. ### Agentic Experiment Loop Designer URL: https://claudescientist.com/tools/loop-designer Purpose: Builds a markdown protocol for hypothesize, run, analyze, decide, and audit loops. Key answer: A credible agentic loop names the execution surface, primary variable, control, success metric, permissions, checkpoints, stop rule, and required artifacts. Output: Copyable and downloadable protocol markdown. ### AI Scientist Claims vs Reality Tracker URL: https://claudescientist.com/tools/claims-tracker Purpose: Filterable evidence table for AI scientist headlines. Key answer: Evaluate what was demonstrated, not the headline label. The table separates workbenches, hypothesis engines, code experiment loops, robotic chemistry agents, and self-driving labs. Output: Filtered rows with system family, autonomy level, demonstrated behavior, caveat, replication status, and source link. ## Primary Source Baseline - Anthropic Claude Science: https://www.anthropic.com/news/claude-science-ai-workbench - Anthropic tool use docs: https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/overview - Anthropic computer use docs: https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/computer-use-tool - Anthropic Model Context Protocol announcement: https://www.anthropic.com/news/model-context-protocol - Anthropic code execution with MCP: https://www.anthropic.com/engineering/code-execution-with-mcp - MCP documentation: https://modelcontextprotocol.io/docs/getting-started/intro - Sakana AI Scientist: https://arxiv.org/abs/2408.06292 - AI Scientist Nature paper: https://www.nature.com/articles/s41586-026-10265-5 - AI Scientist-v2: https://arxiv.org/abs/2504.08066 - Google AI co-scientist: https://www.nature.com/articles/s41586-026-10644-y - Google DeepMind co-scientist announcement: https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/ - Coscientist: https://www.nature.com/articles/s41586-023-06792-0 - A-Lab: https://www.nature.com/articles/s41586-023-06734-w - Self-driving laboratories review: https://www.nature.com/articles/s41467-025-59231-1 - NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework - NIST Generative AI Profile: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence - Nature essay on illusions of understanding: https://www.nature.com/articles/s41586-024-07146-0 - Anthropic Claude Fable 5.1 and Mythos 5.1: https://www.anthropic.com/claude-fable-and-mythos-5-1 - Anthropic Fable 5.1 platform overview: https://platform.claude.com/docs/en/models/fable-5-1/overview - Anthropic expanding support for scientists: https://www.anthropic.com/news/expanding-support-for-scientists - Anthropic Claude Opus 5: https://www.anthropic.com/news/claude-opus-5 - Anthropic cybersecurity evaluation incident disclosure: https://www.anthropic.com/news/alignment-assessment-cybersecurity-incidents - Anthropic models overview: https://docs.anthropic.com/en/docs/about-claude/models - Anthropic pricing: https://platform.claude.com/docs/en/about-claude/pricing - Anthropic text watermarking: https://www.anthropic.com/news/claude-text-watermark