AI Agents

Anthropic’s biology lab pairs Claude with human scientists

Anthropic has established a molecular biology laboratory where Claude helps investigate scientific questions, and humans carry out the experiments. It is a research operation, not an autonomous robot lab.

The company’s 23 September announcement says the research group formed in spring 2026. Claude searches DNA datasets, reviews literature and proposes hypotheses for scientists to assess.

What Claude has actually done

Anthropic reports a search involving roughly 950 agents, 21 hours and 210 million tokens. It identified array-associated reverse transcriptases, an enzyme system with CRISPR-like features whose function remains unknown.

The company says its laboratory works at biosafety levels 1 and 2, does not handle pathogens that infect humans, and leaves all physical lab work to human scientists.

Those disclosures do not establish that AI has invented a working gene editor. Identifying an intriguing system and demonstrating a useful function are different achievements.

Autonomy is not a single switch

For a research operator, the practical question is which decisions an agent can make without approval. Searching a database, selecting a candidate, authorising an experiment and accepting a result represent different kinds of authority. Describing all four as “autonomous research” hides the distinction that matters most.

Our reading is that the meaningful boundary lies between generating a proposal and approving action. A persuasive report should not silently acquire the status of an approved experiment. Faster hypothesis generation makes that boundary more important, not less.

What other researchers can use

Claude Science is separate from the physical laboratory. Anthropic currently lists the app as beta on macOS, Windows, and Linux for Pro, Max, Team, and Enterprise plans. An organisation administrator must enable access for Team and Enterprise users.

It is a workbench around Claude, not a new model. The app connects analysis tools, scientific databases and computing resources. Access to that software should not be confused with access to Anthropic’s laboratory.

The public documentation describes code running in a sandbox, with approval required before the app uses new folders, network hosts or remote jobs. These are documented controls for the customer-facing app, not independent verification of the configuration used inside Anthropic’s lab.

An audit trail is not independent validation

Anthropic says Claude Science preserves the code, software environment and conversation behind its outputs. Its background reviewer checks claims against the execution record, but the documentation explicitly warns that it does not rerun analyses and cannot eliminate errors.

That distinction matters. A result can be traceable yet wrong because its starting assumptions, data selection or interpretation were flawed. Keeping the analysis history helps someone inspect those decisions; it does not prove the biological conclusion.

For example, reproducing a chart from saved code checks the computational record. Repeating a biological result tests a different claim. A product demonstration that establishes the first should not be treated as evidence that the second has happened.

The US National Institutes of Health frames scientific rigour around controlled methods, unbiased analysis and transparent reporting. Applying that standard to an AI-assisted workflow means judging the evidence and experimental design, not the fluency of the generated explanation.

For broader model selection, our comparison of AI models for research addresses source handling and evidence quality. Choosing an assistant and validating a scientific finding remain separate decisions.

Expanded biology access has a data-retention trade-off

Separately, Anthropic’s Life Sciences Verification Program offers vetted organisations broader permissions for biology-related work. Its published process reviews research credentials, security standards and ethical oversight, with access tied to declared uses.

The programme monitors usage against those approved activities and requires 30-day data retention. Anthropic says that data is excluded from model training and cannot be accessed by its life sciences research teams. These are programme commitments, not an independent audit of the internal laboratory.

The phrase “runs on your infrastructure” also has important limits. Claude Science’s data documentation says model requests can include file contents, code output and connector results. Local storage therefore does not mean that everything stays on the local machine.

The same documentation says Claude Science does not offer zero-data-retention arrangements or coverage under Anthropic’s Business Associate Agreement, and instructs users to keep protected health information out. Research teams cannot infer permission to use sensitive data simply from an app’s or subscription’s availability.

The next test is useful findings, not more agents

The headline search figures are not, by themselves, a productivity benchmark. A meaningful comparison would include expert review time, rejected leads, follow-up work and the proportion of candidates that survive experimental testing. Otherwise, a system could appear productive by generating more work for scientists to reject.

Likewise, a token total is not an invoice. A defensible cost comparison needs the model and billing breakdown, computing costs, and human laboratory effort. Search speed alone cannot establish the cost of a validated discovery.

DIY AI’s view: The useful question is whether AI assistance improves the rate and cost of reliable findings while preserving accountable human decisions. More proposals help only when the evidence needed to accept or reject them keeps pace.

Written by Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Back to AI News