✉ The Friday AI Brief: the week's 5 best AI stories, tools & comparisons — in your inbox every Friday morning.

Chan Zuckerberg Biohub logo

Biohub’s $1.8 Billion Bet: US Government, Google and Meta Join the Push for AI Biology Data

October 7, 2026 — San Francisco

The US government, Google, and Meta are all-in on a $1.8 billion effort to build the open datasets AI models need to decode biology — a project Biohub says could compress drug development timelines that currently take years.

Bottom line: The biggest bottleneck in AI-for-biology isn’t models anymore — it’s data. $1.8 billion from the US government, Meta, Google DeepMind, and Isomorphic Labs is now aimed squarely at fixing that, starting with datasets that will eventually be public.

What happened

Biohub — the philanthropic venture of Meta CEO Mark Zuckerberg and Dr. Priscilla Chan — announced Wednesday that the US government and tech heavyweights are joining its effort to build open datasets for training AI models on biological research, bringing total investment to $1.8 billion.

The commitments break down like this: Meta, Google DeepMind, and drug-discovery startup Isomorphic Labs are jointly investing $300 million. The Department of Energy is putting in more than $500 million over five years for laboratory measurement, modeling, and computation. And the National Institutes of Health will coordinate datasets built from more than $500 million in earlier federal funding, which Biohub will standardize for AI training.

That follows the $500 million Biohub itself put into the project in April.

What’s actually being built

The money funds the Virtual Biology Initiative, which aims to measure how cells respond to changes across far more conditions than scientists have studied so far — and use that data to build predictive models for biology.

Current cell datasets run to hundreds of millions of cells. An accurate predictive model needs billions, eventually trillions. Biohub’s goal is to close that gap: data from techniques like spatial transcriptomics (mapping molecular activity inside intact tissue) and screens recording how cells respond to environmental changes, much of it never generated in a coordinated way.

Biohub expects a first dataset within about a year and accurate predictive models within five years — work its head of science Alex Rives says would normally take decades.

The catch: it’s “open” with an embargo

The datasets will eventually be released publicly, but commercial funders get a head start. Rives described embargo periods during which funding companies can work with the data before it becomes a public scientific resource. Government-funded work runs in parallel with no such restrictions. Biohub says it plans to approach pharmaceutical companies and philanthropies next.

Why it matters

From a “everything AI, tested” angle: every headline model right now is trained on internet text. Biology is different — the data to train on simply doesn’t exist yet at the needed scale. This initiative treats AI-ready biology data as infrastructure, like a highway system, with governments and the biggest AI labs splitting the bill. If it works, the companies that funded it get the first crack at models that could design drugs in weeks instead of years — and the rest of the scientific world gets the data afterward. It’s also a telling sign of where the AI race is heading: not just bigger language models, but AI systems grounded in physical reality, starting with the cell.

Frequently asked questions

What is the Biohub Virtual Biology Initiative?
A $1.8 billion project to measure how cells respond to changes across many more conditions than science has studied, creating open datasets to train AI models that predict biological behavior.

Who is funding it?
Biohub itself ($500 million in April), Meta, Google DeepMind, and Isomorphic Labs ($300 million jointly), the Department of Energy (over $500 million across five years), and the NIH (coordinating $500 million+ in earlier federal funding).

When will the data be available?
The first dataset is expected in about a year; predictive models within five years. Commercial funders get embargoed early access before data goes fully public.

Why does AI need biological datasets?
Current cell datasets cover hundreds of millions of cells, but accurate predictive models require billions to trillions of data points — a gap no one has closed because the measurements have never been done in a coordinated way.

Sources: Reuters

Leave a Comment

Your email address will not be published. Required fields are marked *

Get the 5 best AI tools every week

Top AI news, tools, and prompts — one short email. Free, unsubscribe anytime.

Run a newsletter of your own? Monetize and grow it with SparkLoop →

Scroll to Top