Table of Contents
Fetching ...

Artificial Intelligence Virtual Cells: From Measurements to Decisions across Modality, Scale, Dynamics, and Evaluation

Chengpeng Hu, Calvin Yu-Chian Chen

TL;DR

This paper outlines Artificial Intelligence Virtual Cells (AIVCs) and the Cell-State Latent (CSL) framework to model cell state across multimodal, multiscale measurements. It proposes an operator grammar—measurement, lift/project, and perturbation—and a decision-aligned evaluation blueprint that emphasizes function-space readouts such as pathway activity and spatial neighborhoods. The authors review progress in single-cell and spatial foundation models, highlight challenges in cross-modality unification, cross-scale transport, dynamics, and robust evaluation, and advocate operator-aware data design and leakage-resistant benchmarking. The work argues for auditable, like-for-like comparisons and outlines future directions toward perturbation-rich, spatially anchored models with clinical relevance and potential for counterfactual phenotyping.

Abstract

Artificial Intelligence Virtual Cells (AIVCs) aim to learn executable, decision-relevant models of cell state from multimodal, multiscale measurements. Recent studies have introduced single-cell and spatial foundation models, improved cross-modality alignment, scaled perturbation atlases, and explored pathway-level readouts. Nevertheless, although held-out validation is standard practice, evaluations remain predominantly within single datasets and settings; evidence indicates that transport across laboratories and platforms is often limited, that some data splits are vulnerable to leakage and coverage bias, and that dose, time and combination effects are not yet systematically handled. Cross-scale coupling also remains constrained, as anchors linking molecular, cellular and tissue levels are sparse, and alignment to scientific or clinical readouts varies across studies. We propose a model-agnostic Cell-State Latent (CSL) perspective that organizes learning via an operator grammar: measurement, lift/project for cross-scale coupling, and intervention for dosing and scheduling. This view motivates a decision-aligned evaluation blueprint across modality, scale, context and intervention, and emphasizes function-space readouts such as pathway activity, spatial neighborhoods and clinically relevant endpoints. We recommend operator-aware data design, leakage-resistant partitions, and transparent calibration and reporting to enable reproducible, like-for-like comparisons.

Artificial Intelligence Virtual Cells: From Measurements to Decisions across Modality, Scale, Dynamics, and Evaluation

TL;DR

This paper outlines Artificial Intelligence Virtual Cells (AIVCs) and the Cell-State Latent (CSL) framework to model cell state across multimodal, multiscale measurements. It proposes an operator grammar—measurement, lift/project, and perturbation—and a decision-aligned evaluation blueprint that emphasizes function-space readouts such as pathway activity and spatial neighborhoods. The authors review progress in single-cell and spatial foundation models, highlight challenges in cross-modality unification, cross-scale transport, dynamics, and robust evaluation, and advocate operator-aware data design and leakage-resistant benchmarking. The work argues for auditable, like-for-like comparisons and outlines future directions toward perturbation-rich, spatially anchored models with clinical relevance and potential for counterfactual phenotyping.

Abstract

Artificial Intelligence Virtual Cells (AIVCs) aim to learn executable, decision-relevant models of cell state from multimodal, multiscale measurements. Recent studies have introduced single-cell and spatial foundation models, improved cross-modality alignment, scaled perturbation atlases, and explored pathway-level readouts. Nevertheless, although held-out validation is standard practice, evaluations remain predominantly within single datasets and settings; evidence indicates that transport across laboratories and platforms is often limited, that some data splits are vulnerable to leakage and coverage bias, and that dose, time and combination effects are not yet systematically handled. Cross-scale coupling also remains constrained, as anchors linking molecular, cellular and tissue levels are sparse, and alignment to scientific or clinical readouts varies across studies. We propose a model-agnostic Cell-State Latent (CSL) perspective that organizes learning via an operator grammar: measurement, lift/project for cross-scale coupling, and intervention for dosing and scheduling. This view motivates a decision-aligned evaluation blueprint across modality, scale, context and intervention, and emphasizes function-space readouts such as pathway activity, spatial neighborhoods and clinically relevant endpoints. We recommend operator-aware data design, leakage-resistant partitions, and transparent calibration and reporting to enable reproducible, like-for-like comparisons.
Paper Structure (7 sections, 5 equations, 1 figure)

This paper contains 7 sections, 5 equations, 1 figure.

Figures (1)

  • Figure 1: Overview of the Artificial Intelligence Virtual Cell (AIVC) framework and the Cell-State Latent (CSL) perspective.(A) The conceptual architecture of the AIVC. The framework integrates heterogeneous multimodal and multiscale data with Biological Knowledge Priors and Dynamics and Perturbations to learn a unified Cell State Latent (CSL) representation. This shared latent space supports diverse downstream tasks categorized into Cellular Characterization, Dynamic Prediction, and Context Inference.(B) Core challenges for AIVCs.(C) A decision-aligned evaluation blueprint.