October 7, 2026, (Inside AI) — The U.S. Department of Energy will commit more than $500 million over five years to a nonprofit effort aimed at building open biological datasets for training artificial intelligence models, part of a broader $1.8 billion push that now includes Meta Platforms, Google DeepMind, and drug discovery startup Isomorphic Labs.
The nonprofit, Biohub, announced the commitments on Wednesday. The three corporate partners are jointly investing $300 million. The National Institutes of Health will coordinate datasets and repositories built with more than $500 million in earlier federal funding, which Biohub will standardize for AI training.
The money funds the Virtual Biology Initiative, which seeks to measure how cells respond to changes across far more conditions than scientists have studied and use that data to build predictive models. The goal is to compress drug development timelines that currently stretch for years.
The effort arrives as AI labs race to apply machine learning to biology. Anthropic has expanded its biology work with a wet lab, and the OpenAI Foundation launched a grant program exceeding $125 million to fund biological and medical datasets. Biohub's initiative stands out for its scale and its mix of public and private funding.
Read: Anthropic CEO Dario Amodei Earned $18 Million in 2025, IPO Filing Shows
The commitments follow the $500 million that Biohub, a philanthropic venture of Meta CEO Mark Zuckerberg and his wife, Dr. Priscilla Chan, put into the project in April.
Chan framed the work as a shift in how biology advances.
"Biology has been just sort of a clever discovery-based science until this point," Chan said in an interview. "We have always held this as a community asset, not just for one group, so that it can build upon itself over time."
The datasets will eventually be released publicly, but companies that fund them get a head start, according to Biohub's head of science Alex Rives.
"With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," Rives said.
The arrangement is how Biohub draws private money into a project it describes as open science. Rives said government-funded work running in parallel will carry no such restrictions, and that Biohub plans to approach pharmaceutical companies and philanthropies next.
Read: AI's $30 Trillion Bet: Productivity Gains Remain Elusive
Current cell datasets run to hundreds of millions of cells, Rives said, while an accurate predictive model will require billions and eventually trillions. Biohub's goal is to close that gap.
"We need to capture the language of biology, we need to capture the language of the cell. And that doesn't exist today," he said.
The data will come from techniques including spatial transcriptomics, which maps molecular activity inside intact tissue, and screens that record how cells respond to changes in environment. Much of it has never been generated in a coordinated way.
Rives said the work would normally take decades and that the partners aim to compress it into five years, with a first dataset ready in about a year. He expects accurate predictive models within five years.
The initiative reflects a growing belief that AI can accelerate biological discovery by learning patterns from cellular data at scale. If successful, it could reduce the time and cost of bringing new drugs to market, a process that often takes a decade or more.
Biohub's model of blending federal funding, corporate investment, and open data release could serve as a template for other large-scale scientific AI projects. The first dataset is expected in about a year, with predictive models targeted within five years.