Identity & ontology
Are genes, cells, diseases, perturbations and assays named consistently?
- Input
- source identifiers + metadata
- Artifact
- canonical IDs, ontology terms and ambiguity receipts
- Gate
- No silent alias, species or cell-type coercion
Technology & Method Center · checked 2026-08-16
A model name is not a method decision. Start with the biological task, make the data and output contracts explicit, keep simple baselines permanent, and escalate the claim only when every required evaluation layer passes.
This registry is a technical decision aid. It does not claim that every model family was executed, that source-reported performance transfers to NMD-VCell, or that DMD candidate responses are calibrated.
Technical architecture
A model adapter is only one layer. Identity, task design, baselines, evaluation and outcome return determine whether its output is interpretable.
Are genes, cells, diseases, perturbations and assays named consistently?
Can a machine and a reviewer reconstruct what each matrix dimension means?
Which technical effects were measured, corrected or left unresolved?
What is predicted, and along which axis must it generalize?
Does the task require a complex model at all?
Can a model family consume this contract and return the required object?
Did the model recover perturbation-specific and biologically useful signal?
Can the result be reproduced, challenged and updated by an experiment?
Typed data contract
AnnData-shaped typed dataset. obsm embeddings are derived views and never replace the declared source matrix or provenance.
X = declared analysis matrix; raw counts are retained in a named layer when available
cell/sample identifier · donor or culture · batch/assay · disease · cell type/state · perturbation · dose · time · control class
stable gene identifier · symbol · feature selection state · reference genome
raw/counts · normalized · model input only when transformation is named
source accession · license · checksums · processing lineage · task eligibility · split hash
Baseline firewall
A complex model advances only when it adds perturbation-specific, distributional or calibrated value beyond simple controls on the frozen task.
Model family matrix
Local state is explicit on every card. “Source verified” means the method was checked in a primary or official source; it does not mean the method ran here.
Predict no change, a matched mean, a regularized linear response or a nearest observed neighbor.
Best-fit questionIs there learnable signal beyond systematic assay and context structure?
Disentangle basal cell state, perturbation and covariates in a latent representation, then compose an unseen condition.
Best-fit questionCan known factors be recombined across dose, time or cell context?
Propagate gene and perturbation information over co-expression, ontology or learned graphs.
Best-fit questionCan structured gene relationships improve unseen-gene or combination response prediction?
Pretrain token or rank-based cell representations at scale, then adapt embeddings or decoders to a downstream task.
Best-fit questionDoes broad pretraining improve a precisely frozen perturbation or cell-state task?
Represent a cell population as a set and condition one set of cells on another context or prompt.
Best-fit questionCan a model predict population transitions while using context examples at inference time?
Learn a transport map from an unpaired control population to a treated population.
Best-fit questionHow does a distribution of control cells move under treatment when cells are not paired?
Estimate interpretable perturbation effects with a probabilistic prior and calibrated uncertainty.
Best-fit questionWhich gene-level effects are supported, and where should the model abstain?
Learn a conditional generative process that samples heterogeneous post-perturbation cells.
Best-fit questionCan the full conditional response distribution be generated rather than only its mean?
Six-level evaluation ladder
Each level catches a different failure and states what it still cannot prove.
Is the average predicted profile numerically close?
Did the model recover the change caused by this perturbation rather than systematic variation?
Do predicted cells occupy the held-out treated distribution?
Are rare states, proportions and response modes preserved?
Does uncertainty increase where the task leaves the training support?
Does molecular fidelity translate into a reproducible, disease-relevant functional result?
Knowledge content system
Each track ends in an existing platform route, so content changes how the user inspects data, chooses a method or designs an experiment.
Read a matrix as a biological measurement contract, not an anonymous tensor.
Choose an architecture from its assumptions, input and output—not from its reputation.
Detect trivial mean predictors, leakage, mode collapse and overconfident extrapolation.
Convert uncertainty or abstention into a study whose outcome can update the evidence graph.
Source-to-claim matrix
Peer-reviewed primary studies and official specifications are separated from product/preprint claims. Contradictory results are retained because they define safer evaluation gates.
| Source | Evidence | Supports | Does not support | Local state |
|---|---|---|---|---|
| Systematic evaluation of perturbation-response predictors and simple matching baselinesSRC-SYSTEMA-2025 · 2025 | PEER REVIEWED PRIMARYGrade A | mandatory perturbed/matching mean baselines · perturbation-specific evaluation · warning that common metrics reward systematic variation | a universal winning architecture · local NMD-VCell performance | METHOD RULE ABSORBED |
| Benchmarking 27 perturbation-response methods across 29 datasetsSRC-NM-BENCHMARK-2025 · 2025 | PEER REVIEWED PRIMARYGrade A | task- and context-dependent evaluation · multiple datasets and metrics · importance of cellular context | one method as best for every generalization axis | METHOD RULE ABSORBED |
| Deep learning perturbation models do not consistently beat deliberate simple baselinesSRC-SIMPLE-BASELINES-2025 · 2025 | PEER REVIEWED PRIMARYGrade A | permanent simple baselines · architecture-neutral benchmarking | that deep learning can never be useful | METHOD RULE ABSORBED |
| CZI Virtual Cell Models benchmark ecosystemSRC-CZI-BENCHMARKS · 2026 | OFFICIAL DOCGrade B | standardized task definitions · shared metrics · package, CLI and no-code entry points | NMD-VCell task compatibility or performance | PRODUCT PATTERN ABSORBED |
| Arc State virtual-cell architectureSRC-ARC-STATE · 2025 | OFFICIAL PRODUCT PREPRINTGrade B | set-of-cells representation · separate state embedding and transition tasks · evaluation beyond mean expression | advancement beyond permanent baselines · validated DMD prediction | HISTORICAL RECEIPTS MIGRATED |
| Arc foundation model stack and in-context learningSRC-ARC-STACK · 2026 | OFFICIAL PRODUCT PREPRINTGrade B | cell-set prompting as a model pattern · source-reported large-scale pretraining | checkpoint benchmark after runtime timeout · NMD-VCell performance inheritance | LOCAL NVME CORE BASE FORWARD PASS NO CHECKPOINT |
| Pertpy perturbation-analysis frameworkSRC-PERTPY-2025 · 2025 | PEER REVIEWED PRIMARYGrade A | end-to-end perturbation analysis · MMD, energy and Wasserstein distances · typed analysis modules | a disease-response prediction result | METHOD REFERENCE |
| AnnData annotated matrix and on-disk specificationSRC-ANNDATA · 2026 | OFFICIAL DOCGrade A | X/obs/var/layers/obsm/uns data contract · typed annotated matrices | biological correctness of any uploaded object | DATA CONTRACT ABSORBED |
| CELLxGENE schema 5.2.0SRC-CELLXGENE-SCHEMA · 2026 | OFFICIAL DOCGrade A | ontology-backed biological and technical metadata · schema as a search and integration contract | automatic suitability for a perturbation task | ONTOLOGY PATTERN ABSORBED |
| GEARS graph neural network for perturbation predictionSRC-GEARS-2023 · 2023 | PEER REVIEWED PRIMARYGrade A | graph-informed gene and perturbation embeddings · structured combination prediction task | success on the local frozen task · reliable combinations from single perturbations alone | LOCALLY EXECUTED FAILED FIVE SPLITS |
| CellOT neural optimal transportSRC-CELLOT-2023 · 2023 | PEER REVIEWED PRIMARYGrade A | unpaired control-to-treated population mapping · distributional prediction | stable inference in sparse cell types · local DMD validation | SOURCE VERIFIED NOT RUN |
| Compositional perturbation autoencoderSRC-CPA-2023 · 2023 | PEER REVIEWED PRIMARYGrade A | factorized perturbation and covariate representations · dose/time/context composition | constant performance as unseen covariates accumulate · local NMD-VCell calibration | SOURCE VERIFIED NOT RUN |
| GPerturb probabilistic perturbation-effect modelSRC-GPERTURB-2025 · 2025 | PEER REVIEWED PRIMARYGrade A | sparse interpretable effects · uncertainty estimates | treating cells as independent biological replicates · clinical prediction | SOURCE VERIFIED NOT RUN |
| Open Problems perturbation prediction benchmarkSRC-OPENPROBLEMS-PERTURBATION · 2024 | OFFICIAL DOCGrade B | versioned task and metric metadata · automated score checking | generalization outside the declared competition task | BENCHMARK PATTERN ABSORBED |
Operating boundary
This registry is a technical decision aid. It does not claim that every model family was executed, that source-reported performance transfers to NMD-VCell, or that DMD candidate responses are calibrated.