The bottleneck is not another bigger model

The bottleneck is not another bigger model

Xaira’s X-Cell conversation argues that virtual-cell models need causal perturbation data, not just larger architectures, and that generalization depends on the lab-data loop.

Bo Wang and Ci Chu's central claim in this Latent Space conversation is simple: a virtual-cell model cannot learn reliable counterfactual biology from descriptive data alone. If the question is "what happens when I turn this gene down?", the training set needs experiments in which a gene was actually perturbed and the resulting cellular state was measured. More parameters can improve a model, but they cannot manufacture the missing intervention. 1
Wang is Xaira Therapeutics' senior vice president and head of biomedical AI, and previously an associate professor at the University of Toronto. Chu leads AI-enabled discovery at Xaira and describes prior work at Insitro and Verily, with a focus on the high-throughput biology that produces data for the models. Their conversation is less a model launch pitch than an explanation of why the lab has to be built alongside the AI system. 2
正在加载内容卡片…

Observation can describe a cell without explaining it

The distinction becomes clear with a small thought experiment. Suppose genes A, B, and C rise and fall together in a large observational dataset. That pattern is compatible with A regulating B and C. It is also compatible with B regulating A, or with all three responding to a fourth cause. The same correlations can fit several causal stories.
Chu contrasts this with the large cell-by-gene datasets used to train earlier single-cell foundation models. Those datasets are valuable for descriptive work such as correcting batch effects across laboratories and technologies. But the models trained on them have not consistently beaten simple linear baselines on perturbation or counterfactual tasks. The problem is not that the data are poor in every respect. It is that observation does not tell the model which intervention produced a change. 1
That is why Xaira treats causal data generation as a first-class engineering problem. Its broader platform combines protein design, a virtual-cell system, and patient-representation models. The point is to connect models across the drug-discovery pipeline rather than let each one optimize a disconnected prediction task.

Perturb-seq turns biology into a training matrix

The data strategy centers on pooled Perturb-seq: high-throughput CRISPR perturbations paired with single-cell RNA sequencing. In plain terms, researchers reduce the expression of different genes across a large population of cells, use guide-RNA barcodes to identify which perturbation occurred in each cell, and then measure the expression of thousands of genes in that same cell.
That creates a two-dimensional training object. One axis is the intervention; the other is the cellular readout. A human cell may contain roughly 20,000 genes, and the experiment asks what happens across the rest of the system when one of them is turned down. The approach can scale many perturbations into one pooled experiment, rather than requiring a separate plate and workflow for every gene. 2
The scale is only useful if the measurements are consistent. Chu describes a quality-filtered dataset of more than 25 million cells, with many more cells handled before they passed the final filters. The team introduced chemical fixation and other industrialized steps so that cells could be processed across a long operating day without their changing state becoming an accidental batch effect. This is a reminder that the data advantage is partly a robotics and operations advantage: the lab protocol is part of the model's capability.
Xaira is also increasing biological context. The work began with easier-to-grow immortalized cell lines, then expanded to induced pluripotent stem cells, primary cells, and an experiment that differentiated one stem-cell population into ten cell types. More cells are not automatically more information. The team repeatedly returns to context, diversity, and information per dollar as the ingredients of generalization.

X-Cell changes the modeling assumptions too

X-Cell is a 4.9-billion-parameter diffusion language model for predicting cellular responses to perturbations. Its architecture is designed around an awkward fact about gene-expression data: expression values do not have the natural left-to-right order of a sentence. Autoregressive models can impose an order for convenience, but diffusion lets the system iteratively refine a noisy, incomplete expression state into a full prediction. 2
The model also incorporates several kinds of biological prior, including literature-derived gene information, protein-protein interaction networks, cancer-essentiality information, cell morphology, and embeddings from an earlier single-cell model. Wang's ranking of what matters most is revealing: data quality and scale come first, architecture next, and prior knowledge after that, with the contribution of a given prior varying by cell type.
That ordering pushes back on a familiar AI shortcut. Adding a more elaborate architecture is not a substitute for collecting the right evidence. The architecture matters because it can use the evidence more effectively, but the causal training matrix is what gives the model a chance to learn dynamics rather than merely reproduce a profile.

Generalization is the real benchmark

The strongest demonstrations in the conversation are held-out contexts. X-Cell was trained on resting T-cell data and then asked to predict perturbations in activated T cells, which it had not seen in that combination. It recovered expected biology and also identified candidate inactivators that appeared in the subsequent screen. The team also held out a cell type from a multi-cell-type experiment and tested transfer from a T-cell line to primary T cells collected from donors. 1
These are promising results, not a finished virtual patient. The guests stress that the point of the model is not to eliminate experiments. It is to use experiments where they are scalable, then use the learned system to generate better hypotheses in contexts that are expensive or impossible to screen exhaustively, such as organs, organoids, animals, and eventually patients.
That is also why they are cautious about benchmarks. Mean absolute error can reward an average prediction in sparse single-cell data even when the model misses the biological change that matters. They emphasize metrics that compare predicted gene-expression changes with the observed changes after a perturbation, and they still want more biological validation before treating the approach as settled.

The next missing measurement is time

The long-term ambition is broader than a better static cell profile. Chu wants measurements of proteins, their modifications, abundance, localization, and cell-to-cell context. Wang's wish is a technology that can measure the same cell repeatedly over time instead of killing it to sequence it. A model that could see how one cell evolves would move closer to the dynamic meaning of a virtual cell. 2
The practical lesson for AI-for-science builders is therefore unusually concrete. Before choosing a larger model, define the intervention you need to predict, build the experiment that identifies it, and measure whether the model generalizes to a context held out from training. In this field, the frontier is not just a parameter count. It is the tightness of the loop between a perturbation in the lab, a prediction in software, and a biological result that can prove the prediction wrong.

相似内容

  • 登录后可发表评论。
More from this channel