NMD-VCell Neuromuscular Virtual Cell Research Platform Module: Registry · Evidence → perturbation → experiment → outcome NMD = neuromuscular disorders

Release 2026.08DMD context observedDMD candidate-conditioned prediction not yet eligible

View scientific status
Evidence freeze: 3 August 2026 Resource: v1.2.0-measured-dmd-evidence Schema: 1.1 Open release status →

Technology & Method Center · checked 2026-08-16

Understand the stack before trusting the prediction.

A model name is not a method decision. Start with the biological task, make the data and output contracts explicit, keep simple baselines permanent, and escalate the claim only when every required evaluation layer passes.

Current local decisionKeep ridge as a permanent same-assay baseline. TxPert: STOP_UNTIL_NEW_PREREGISTRATION_OR_UNTOUCHED_TRUTH. STATE: NO_ADVANCEMENT. Stack: NO_NMD_B1_BENCHMARK_NO_COMPATIBLE_PRETRAINED_GENETIC_CHECKPOINT. Tahoe-x1: DO_NOT_RUN_ON_LOG_NORMALIZED_INPUT; ACQUIRE_RAW_COUNTS_FIRST. SLIM: NOT_ELIGIBLE_FOR_EXECUTION. None is a locally validated NMD model.

This registry is a technical decision aid. It does not claim that every model family was executed, that source-reported performance transfers to NMD-VCell, or that DMD candidate responses are calibrated.

Technical architecture

Eight gates connect source data to a returned biological outcome.

A model adapter is only one layer. Identity, task design, baselines, evaluation and outcome return determine whether its output is interpretable.

Download architecture TSV →
TECH-0101

Identity & ontology

Are genes, cells, diseases, perturbations and assays named consistently?

Input
source identifiers + metadata
Artifact
canonical IDs, ontology terms and ambiguity receipts
Gate
No silent alias, species or cell-type coercion
Open local layer →
TECH-0202

Dataset contract

Can a machine and a reviewer reconstruct what each matrix dimension means?

Input
matrix + sample/cell + feature metadata
Artifact
typed AnnData-shaped Dataset Card with provenance
Gate
X, obs, var, layers and uns roles are explicit
Open local layer →
TECH-0303

QC & harmonization

Which technical effects were measured, corrected or left unresolved?

Input
raw/count layers + donor, batch and assay metadata
Artifact
QC receipt, excluded units and transformation lineage
Gate
Biological replicates remain the inferential unit
Open local layer →
TECH-0404

Task & split builder

What is predicted, and along which axis must it generalize?

Input
context × perturbation × time × output contract
Artifact
frozen estimand, split, leakage checks and holdout identity
Gate
Target, context and output are unseen exactly as declared
Open local layer →
TECH-0505

Permanent baseline lane

Does the task require a complex model at all?

Input
the identical training set, genes, split and metric code
Artifact
zero-change, mean, linear and compatible neighbor results
Gate
Advanced models cannot advance without same-coverage controls
Open local layer →
TECH-0606

Model adapter

Can a model family consume this contract and return the required object?

Input
frozen task + versioned model/configuration
Artifact
ModelRun with mean vector, population or uncertainty output
Gate
Architecture capability is not inferred from model name
Open local layer →
TECH-0707

Evaluation & gate

Did the model recover perturbation-specific and biologically useful signal?

Input
predictions + held-out truth + permanent baselines
Artifact
six-level scorecard, uncertainty and abstention decision
Gate
One aggregate expression metric cannot unlock a claim
Open local layer →
TECH-0808

Registry & outcome return

Can the result be reproduced, challenged and updated by an experiment?

Input
ModelRun, Study, Prediction and measured Outcome
Artifact
immutable receipt and evidence-graph return
Gate
Null, toxic and failed outcomes are retained
Open local layer →

Typed data contract

An input matrix must retain its biological meaning.

AnnData-shaped typed dataset. obsm embeddings are derived views and never replace the declared source matrix or provenance.

XDeclared matrix

X = declared analysis matrix; raw counts are retained in a named layer when available

obsBiological + technical context

cell/sample identifier · donor or culture · batch/assay · disease · cell type/state · perturbation · dose · time · control class

varStable feature identity

stable gene identifier · symbol · feature selection state · reference genome

layersTransformation lineage

raw/counts · normalized · model input only when transformation is named

unsProvenance + task

source accession · license · checksums · processing lineage · task eligibility · split hash

Baseline firewall

Complexity earns entry only by adding task-specific information.

A complex model advances only when it adds perturbation-specific, distributional or calibrated value beyond simple controls on the frozen task.

Permanent controlszero change · control/train mean · perturbed or matching mean where legal · ridge/linear · nearest neighbor where legal
Same-coverage ruleEvery comparator uses the same training objects, held-out units, feature set, seeds and metric implementation as the advanced model.
Stop ruleIf the task lacks matched truth, the correct result is an abstention plus a prospective study—not a proxy score.

Model family matrix

Choose the family from the output and failure mode—not prestige.

Local state is explicit on every card. “Source verified” means the method was checked in a primary or official source; it does not mean the method ran here.

Download model-family TSV →
MF-01LOCALLY EXECUTED LIMITED

Linear & non-parametric baselines

Predict no change, a matched mean, a regularized linear response or a nearest observed neighbor.

Best-fit questionIs there learnable signal beyond systematic assay and context structure?

Inputmatched feature space; frozen split; context labels for matching where permittedOutputaggregate response vector or neighbor-based reference
Strength
transparent · low variance · fast leakage diagnostic · permanent comparator
Failure modes
cannot represent complex interactions · matching can leak held-out context · good mean error may miss perturbation identity
Permanent baselines
zero change · train/control mean · perturbed or matching mean when legal · ridge · nearest neighbor when legal
Evaluate with
mean fidelity · perturbation-specific delta · split integrity

Next gateRetain ridge as a permanent same-assay comparator; do not extend it to DMD transfer.

MF-02SOURCE VERIFIED NOT RUN

Factorized latent generative models

Disentangle basal cell state, perturbation and covariates in a latent representation, then compose an unseen condition.

Best-fit questionCan known factors be recombined across dose, time or cell context?

Inputcell-level expression with perturbation, covariate, dose/time and control labelsOutputcounterfactual cells or an expected post-perturbation distribution
Strength
compositional representation · covariate conditioning · counterfactual generation
Failure modes
disentanglement is not guaranteed · performance declines as unseen covariates accumulate · batch may be encoded as biology
Permanent baselines
matching mean · context-conditioned ridge · nearest observed condition
Evaluate with
mean fidelity · delta recovery · distribution fidelity · OOD covariate stress test

Next gateRun only after a factorial task exposes which covariate combinations are genuinely unseen.

MF-03LOCALLY EXECUTED FAILED OR PENDING

Graph-informed perturbation models

Propagate gene and perturbation information over co-expression, ontology or learned graphs.

Best-fit questionCan structured gene relationships improve unseen-gene or combination response prediction?

Inputperturbation responses plus gene identities and a versioned graph or embedding sourceOutputpost-perturbation mean expression or effect vector
Strength
uses gene relationships · supports structured inductive bias · can represent combinations
Failure modes
graph mismatch · single-perturbation data may not identify combinations · relation priors can dominate sparse truth
Permanent baselines
additive single-perturbation baseline · ridge · matching mean
Evaluate with
held-out genes · held-out combinations · sign and interaction recovery · seed stability

Next gateGEARS remains failed on five frozen splits; TxPert remains calibration-pending. Neither unlocks DMD prediction.

MF-04LOCALLY EXECUTED FAILED

Foundation transformer & embedding models

Pretrain token or rank-based cell representations at scale, then adapt embeddings or decoders to a downstream task.

Best-fit questionDoes broad pretraining improve a precisely frozen perturbation or cell-state task?

Inputgene-aligned expression plus the exact tokenizer, vocabulary, checkpoint and adaptation recipeOutputcell embeddings, labels or decoded response vectors depending on the adapter
Strength
broad representation prior · transferable embeddings · large reference context
Failure modes
embedding quality is not perturbation accuracy · vocabulary/context mismatch · scale can obscure task leakage
Permanent baselines
PCA or linear embedding · ridge · train mean · task-specific shallow model
Evaluate with
task-level baseline comparison · OOD split · ablation of pretraining · reproduction receipt

Next gateThe local scGPT adapter lost to same-coverage baselines; a new checkpoint is a new frozen ModelRun, not an inherited upgrade.

MF-05COMPATIBILITY ROUTE EXECUTED NO ADVANCEMENT

Set-to-set & context-prompted models

Represent a cell population as a set and condition one set of cells on another context or prompt.

Best-fit questionCan a model predict population transitions while using context examples at inference time?

Inputcell sets, context labels, perturbation identity and a task-compatible feature universeOutputpredicted cell set, state embedding or transition distribution
Strength
population-native · context prompting · heterogeneity-aware representation
Failure modes
prompt leakage · set composition confounding · source-reported scale may not transfer to disease context
Permanent baselines
matching population · stratified mean · optimal-transport baseline
Evaluate with
distribution fidelity · composition recovery · prompt ablation · unseen-context holdout

Next gateSTATE: Do not rerun by reputation; require a corrected preregistration, untouched truth or materially different task-compatible checkpoint. Stack: Identify a task-compatible pretrained genetic checkpoint and preregister an untouched benchmark before any model-performance claim.

MF-06SOURCE VERIFIED NOT RUN

Optimal-transport population maps

Learn a transport map from an unpaired control population to a treated population.

Best-fit questionHow does a distribution of control cells move under treatment when cells are not paired?

Inputunpaired control and treated cell populations with shared features and sufficient state coverageOutputtransported cells or a treatment-conditioned population distribution
Strength
distributional output · unpaired design · population geometry
Failure modes
rare states are unstable · transport assumptions may not identify biology · composition shifts can mimic state transitions
Permanent baselines
identity map · mean shift · nearest-neighbor transport
Evaluate with
MMD or energy distance · Wasserstein distance · state composition · rare-state stratification

Next gateUse only for a population-output task with explicit rare-state and composition stress tests.

MF-07SOURCE VERIFIED NOT RUN

Probabilistic sparse-effect models

Estimate interpretable perturbation effects with a probabilistic prior and calibrated uncertainty.

Best-fit questionWhich gene-level effects are supported, and where should the model abstain?

Inputreplicated perturbation responses with biological units, covariates and stable featuresOutputeffect posterior, uncertainty interval and sparse active set
Strength
uncertainty-aware · interpretable effects · appropriate for sparse signals
Failure modes
prior sensitivity · poor scaling · cell-level pseudo-replication produces false confidence
Permanent baselines
regularized linear model · empirical Bayes shrinkage · no-effect model
Evaluate with
interval coverage · calibration · effect-sign recovery · donor-level resampling

Next gatePrioritize when replicated DMD perturbation outcomes exist and uncertainty is decision-critical.

MF-08WATCHLIST NOT RUN

Distributional flow & diffusion generators

Learn a conditional generative process that samples heterogeneous post-perturbation cells.

Best-fit questionCan the full conditional response distribution be generated rather than only its mean?

Inputlarge cell-level perturbation datasets with context, dose/time and robust controlsOutputsampled post-perturbation cell population
Strength
multimodal distributions · heterogeneity · sample-level counterfactuals
Failure modes
plausible-looking hallucinated states · mode collapse · weak calibration · high compute burden
Permanent baselines
matching population · CellOT or transport baseline · conditional Gaussian baseline
Evaluate with
distribution metrics · mode coverage · calibration · biological state validity · prospective function

Next gateDo not adopt until the population task, compute budget and prospective validation route are frozen.

Six-level evaluation ladder

A good mean profile is the first test, not the final conclusion.

Each level catches a different failure and states what it still cannot prove.

EVAL-0101

Mean fidelity

Is the average predicted profile numerically close?

MeasuresRMSE · MAE · mean correlation
Catchesgross reconstruction error
Cannot proveperturbation identity, heterogeneity or biological function
EVAL-0202

Perturbation-specific signal

Did the model recover the change caused by this perturbation rather than systematic variation?

Measuresdelta correlation/cosine with amplitude checks · signed DE recovery · precision/recall of responsive genes · rank recovery
Catchesmean predictors that ignore perturbation identity
Cannot provecell-population fidelity or function
EVAL-0303

Population distribution

Do predicted cells occupy the held-out treated distribution?

MeasuresMMD · energy distance · Wasserstein · classifier two-sample test
Catchescorrect mean with wrong spread or geometry
Cannot provecorrect cell-state composition or mechanism
EVAL-0404

State composition & heterogeneity

Are rare states, proportions and response modes preserved?

Measurescell-state proportion error · rare-state recall · mode coverage · stratified distribution metrics
Catchesmode collapse and composition confounding
Cannot proveuncertainty reliability or disease relevance
EVAL-0505

Calibration, OOD & abstention

Does uncertainty increase where the task leaves the training support?

Measuresinterval coverage · calibration error · selective risk · OOD-stratified performance · seed stability
Catchesconfident extrapolation and unstable wins
Cannot provetherapeutic efficacy or clinical utility
EVAL-0606

Biological & prospective validation

Does molecular fidelity translate into a reproducible, disease-relevant functional result?

Measuresindependent donor replication · prespecified functional endpoint · safety/toxicity · prospective outcome return
Catchesmolecular proxies that do not alter function
Cannot proveclinical benefit without a separate clinical design

Knowledge content system

Learn by scientific decision, mental model and failure mode.

Each track ends in an existing platform route, so content changes how the user inspects data, chooses a method or designs an experiment.

Download curriculum TSV →
KNOW-02Model mechanisms

Choose an architecture from its assumptions, input and output—not from its reputation.

  1. Mean vector or cell population?the output object determines the model familyFailure mode · calling an aggregate vector a virtual-cell population
  2. Where does context enter?covariate, graph, prompt and transport encode different assumptionsFailure mode · architecture-name inference
  3. What was actually executed?paper capability and local ModelRun are separate ledgersFailure mode · inheriting external performance
KNOW-03Evaluation without shortcuts

Detect trivial mean predictors, leakage, mode collapse and overconfident extrapolation.

  1. Why simple baselines are permanentsystematic variation can make trivial predictors look strongFailure mode · complexity bias
  2. Why one metric is insufficientmean, delta, distribution and function answer different questionsFailure mode · metric substitution
  3. How OOD axes change the claimunseen perturbation, donor and disease are different tasksFailure mode · generalization collapse

Source-to-claim matrix

Every external source states both what it supports and what it cannot support.

Peer-reviewed primary studies and official specifications are separated from product/preprint claims. Contradictory results are retained because they define safer evaluation gates.

Download source TSV →
SourceEvidenceSupportsDoes not supportLocal state
Systematic evaluation of perturbation-response predictors and simple matching baselinesSRC-SYSTEMA-2025 · 2025PEER REVIEWED PRIMARYGrade Amandatory perturbed/matching mean baselines · perturbation-specific evaluation · warning that common metrics reward systematic variationa universal winning architecture · local NMD-VCell performanceMETHOD RULE ABSORBED
Benchmarking 27 perturbation-response methods across 29 datasetsSRC-NM-BENCHMARK-2025 · 2025PEER REVIEWED PRIMARYGrade Atask- and context-dependent evaluation · multiple datasets and metrics · importance of cellular contextone method as best for every generalization axisMETHOD RULE ABSORBED
Deep learning perturbation models do not consistently beat deliberate simple baselinesSRC-SIMPLE-BASELINES-2025 · 2025PEER REVIEWED PRIMARYGrade Apermanent simple baselines · architecture-neutral benchmarkingthat deep learning can never be usefulMETHOD RULE ABSORBED
CZI Virtual Cell Models benchmark ecosystemSRC-CZI-BENCHMARKS · 2026OFFICIAL DOCGrade Bstandardized task definitions · shared metrics · package, CLI and no-code entry pointsNMD-VCell task compatibility or performancePRODUCT PATTERN ABSORBED
Arc State virtual-cell architectureSRC-ARC-STATE · 2025OFFICIAL PRODUCT PREPRINTGrade Bset-of-cells representation · separate state embedding and transition tasks · evaluation beyond mean expressionadvancement beyond permanent baselines · validated DMD predictionHISTORICAL RECEIPTS MIGRATED
Arc foundation model stack and in-context learningSRC-ARC-STACK · 2026OFFICIAL PRODUCT PREPRINTGrade Bcell-set prompting as a model pattern · source-reported large-scale pretrainingcheckpoint benchmark after runtime timeout · NMD-VCell performance inheritanceLOCAL NVME CORE BASE FORWARD PASS NO CHECKPOINT
Pertpy perturbation-analysis frameworkSRC-PERTPY-2025 · 2025PEER REVIEWED PRIMARYGrade Aend-to-end perturbation analysis · MMD, energy and Wasserstein distances · typed analysis modulesa disease-response prediction resultMETHOD REFERENCE
AnnData annotated matrix and on-disk specificationSRC-ANNDATA · 2026OFFICIAL DOCGrade AX/obs/var/layers/obsm/uns data contract · typed annotated matricesbiological correctness of any uploaded objectDATA CONTRACT ABSORBED
CELLxGENE schema 5.2.0SRC-CELLXGENE-SCHEMA · 2026OFFICIAL DOCGrade Aontology-backed biological and technical metadata · schema as a search and integration contractautomatic suitability for a perturbation taskONTOLOGY PATTERN ABSORBED
GEARS graph neural network for perturbation predictionSRC-GEARS-2023 · 2023PEER REVIEWED PRIMARYGrade Agraph-informed gene and perturbation embeddings · structured combination prediction tasksuccess on the local frozen task · reliable combinations from single perturbations aloneLOCALLY EXECUTED FAILED FIVE SPLITS
CellOT neural optimal transportSRC-CELLOT-2023 · 2023PEER REVIEWED PRIMARYGrade Aunpaired control-to-treated population mapping · distributional predictionstable inference in sparse cell types · local DMD validationSOURCE VERIFIED NOT RUN
Compositional perturbation autoencoderSRC-CPA-2023 · 2023PEER REVIEWED PRIMARYGrade Afactorized perturbation and covariate representations · dose/time/context compositionconstant performance as unseen covariates accumulate · local NMD-VCell calibrationSOURCE VERIFIED NOT RUN
GPerturb probabilistic perturbation-effect modelSRC-GPERTURB-2025 · 2025PEER REVIEWED PRIMARYGrade Asparse interpretable effects · uncertainty estimatestreating cells as independent biological replicates · clinical predictionSOURCE VERIFIED NOT RUN
Open Problems perturbation prediction benchmarkSRC-OPENPROBLEMS-PERTURBATION · 2024OFFICIAL DOCGrade Bversioned task and metric metadata · automated score checkinggeneralization outside the declared competition taskBENCHMARK PATTERN ABSORBED

Operating boundary

Paper capability, local execution and disease validation are three different claims.

This registry is a technical decision aid. It does not claim that every model family was executed, that source-reported performance transfers to NMD-VCell, or that DMD candidate responses are calibrated.