Business

Google DeepMind, Meta and Isomorphic Labs Join $1.8B Biohub AI Biology Initiative

Google DeepMind, Meta and Isomorphic Labs are collectively investing $300 million in Biohub’s Virtual Biology Initiative, expanding a five-year effort that now represents $1.8 billion in funding, existing data resources, computation and biological measurement technology.

The 7 October 2026 announcement confirms the investment alongside more than $500 million from the US Department of Energy and biomedical resources created through more than $500 million of previous National Institutes of Health investment. Biohub had already committed $500 million when it launched the initiative in April.

The headline number is substantial, but the more important development for AI builders is what the partners are trying to build: a large, standardised biological data layer designed to train predictive AI models. The initiative could eventually reduce one of the biggest constraints on AI biology: access to sufficiently broad, consistent, and experimentally useful training data.

The $1.8 billion figure is not simply $1.8 billion of new cash

The funding structure needs unpacking because treating the entire $1.8 billion as a new investment would overstate what was announced.

ParticipantContributionWhat it supports
Biohub$500 millionExisting five-year commitment, including technology development, internal data generation and external research funding
Google DeepMind, Isomorphic Labs and Meta$300 million collectivelyNew technologies and multimodal datasets for predictive biological models
US Department of EnergyMore than $500 million over five yearsMeasurement, modelling, computation and national laboratory infrastructure
National Institutes of HealthResources derived from more than $500 million of prior federal investmentExisting datasets, repositories and knowledge resources that will be standardised for AI training

That distinction changes how the initiative should be evaluated. This is partly a funding programme, but it is also an attempt to combine existing scientific infrastructure, previously funded datasets, new experiments and large-scale computing into a common AI-ready resource.

Biohub’s 7 October announcement says the participating organisations intend to create shared standards, common identifiers and a single access layer rather than simply publishing another collection of disconnected research datasets.

Open data will not necessarily mean equal access on day one

One less obvious detail is how access will work. Reuters reports that datasets financed by commercial partners can have embargo periods during which the companies funding them receive early access. Government-funded work running alongside those projects is expected to have no equivalent restriction.

That makes the initiative more complicated than an ordinary open-data project. A dataset can eventually become public while still giving the companies that helped finance its production an early competitive advantage.

For an AI lab, even a temporary head start can matter. Early access creates time to build ingestion pipelines, develop preprocessing methods, identify data-quality problems, design benchmarks, and start training before the same material becomes generally available. By the time competitors receive the data, early participants may already understand which parts are most useful.

This does not mean Google DeepMind, Meta or Isomorphic Labs will control the resulting datasets. It does mean that builders evaluating supposedly open AI resources should look beyond the licence and ask when different parties actually receive access.

The real bottleneck is experimental coverage, not another larger model

Biohub wants the initiative to contribute to predictive models that estimate how cells respond to interventions. The ambition is often described as creating a “virtual cell”, but that label can imply considerably more fidelity than current biological AI systems can guarantee.

A recurring concern among practitioners working around computational biology is that existing datasets cover only a tiny fraction of the combinations that matter. Cell type, tissue environment, disease state, genetic variation, chemical intervention, dosage and time can all alter a biological response. Increasing the number of recorded cells does not automatically solve that coverage problem if the experiments themselves remain concentrated around similar conditions.

This is why the most interesting part of the Biohub project may be its investment in measurement rather than AI architecture. The programme includes cryo-electron tomography, large-scale microscopy and technologies for perturbing and observing biology at molecular, cellular, tissue and organism levels.

That produces a different kind of AI scaling problem. Language models can consume enormous collections of existing text. A predictive biological system may require researchers to create much of the training evidence through physical experiments first.

Multimodal biological data creates a harder integration problem

The initiative also intends to combine multiple types of biological information rather than relying on a single measurement such as gene expression.

That should improve the potential usefulness of the resulting models, but it creates difficult engineering decisions. Imaging, transcriptomic, proteomic and perturbation datasets are generated at different resolutions, with different instruments, metadata conventions and error characteristics. Simply storing them in the same repository does not make them interoperable.

Standardisation work may therefore be as important as raw dataset size. Builders will need consistent identifiers, clearly documented experimental conditions, provenance records and mappings between measurement types. Missing metadata can make an otherwise enormous dataset unsuitable for answering a specific biological question.

This is similar to a principle we apply across the DIY AI data hub: the usefulness of a benchmark or training dataset depends on what was measured, how it was measured and what the data cannot tell you. Raw scale is a poor substitute for experimental coverage.

The initiative could lower one barrier while creating another

If Biohub succeeds in releasing large, well-documented datasets publicly, smaller research groups and AI companies could gain access to biological training material that would be extremely expensive to generate independently.

That would reduce the advantage held by organisations that can finance their own laboratories. It would not eliminate the advantage entirely.

Training, validating and repeatedly updating large biological models can still require substantial compute, specialised engineering and domain expertise. Organisations with early dataset access and established research infrastructure could therefore remain ahead even after the underlying data becomes public.

For Isomorphic Labs, which is focused on AI-driven drug discovery, the strategic relevance is particularly clear. Better perturbation and cellular-response data could support models used earlier in the drug discovery pipeline. Google DeepMind and Meta also gain exposure to a dataset-generation programme tackling a problem that cannot easily be solved by collecting more public internet data.

Five checks will determine whether the first Biohub datasets are genuinely useful

Biohub expects the first dataset from the expanded effort in roughly a year. When it appears, dataset size should be one of the last metrics examined, not the first.

  • Intervention coverage: how many genuinely different biological perturbations, cell states and environmental conditions are represented?
  • Modality alignment: can imaging, genomic, transcriptomic and other measurements reliably be connected at the sample or experiment level?
  • Provenance: can researchers trace individual measurements back to the experimental protocol, instrument and processing pipeline that produced them?
  • Evaluation separation: are there credible held-out experiments for testing whether models generalise beyond closely related training examples?
  • Access timing: which datasets are immediately public, which have commercial embargoes, and how long do those restrictions last?

Those checks matter because biological model evaluation is particularly vulnerable to misleading results when training and test conditions are too similar. A system that predicts another measurement from familiar cell types is not necessarily capable of predicting an unseen biological response.

The same distinction applies to general AI benchmarking. Our AI model comparison separates headline capability from the practical conditions under which models are tested. Biological AI will require an even stricter version of that discipline because an incorrect prediction can look plausible while failing to represent the underlying mechanism.

DIY AI view: the dataset could become more strategically important than any single model

The most significant part of this announcement is not that three large AI companies have written another sizeable cheque. It is that major commercial labs, US government agencies and biological research organisations are converging on the same constraint: AI biology needs better experimental data before model scaling alone can deliver reliable predictive systems.

Biohub’s attempt to create shared infrastructure could therefore have a longer useful life than the first models trained on it. Model architectures change quickly. A carefully designed dataset containing expensive biological measurements can remain valuable across several generations of models.

There are still reasons to be cautious. The project has targets rather than demonstrated outcomes. The “virtual cell” concept remains much broader than what existing systems can faithfully simulate, and temporary commercial access advantages complicate the open-science framing.

The milestone worth watching is not whether the partners can accumulate billions of cell measurements. It is whether the resulting datasets capture enough diverse interventions and biological contexts for a model to predict an experiment it has never effectively seen before. If it passes that test, the Virtual Biology Initiative could become important infrastructure for AI-driven biological research rather than simply another large AI funding programme.

Written by Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Back to AI News