NMD-VCell Neuromuscular Virtual Cell Research Platform Module: Registry · Evidence → perturbation → experiment → outcome NMD = neuromuscular disorders

Release 2026.08DMD context observedDMD candidate-conditioned prediction not yet eligible

View scientific status
Evidence freeze: 3 August 2026 Resource: v1.2.0-measured-dmd-evidence Schema: 1.1 Open release status →

VCC 2026 readiness

← Benchmark center

NMD-VCell × VCC 2026 Generalization Readiness

Official task is now disclosed: unseen cell-context CRISPRi response prediction.

Arc's 20 August 2026 announcement defines a zero-shot task: predict CRISPRi knockdown responses in six cell lines whose perturbation responses are withheld. This board measures NMD-VCell's readiness gap without claiming VCC-compatible model performance.

Official task announced
Official taskUnseen-context CRISPRiSix cell lines · perturbation truth withheld
NMD-VCell status9/26 readiness checks available7 partial · 10 missing
Benchmarkable todayAdapter · baselines · abstention55-target aggregate diagnostic remains reference-only
Cannot claimZero-shot NMD performanceNo DMD candidate-conditioned perturbation truth
Official sourcearc-official-news-2026-08-20
Source URL
https://arcinstitute.org/news/virtual-cell-challenge-2026
Checked at
2026-09-08
Task status
ANNOUNCED_ZERO_SHOT_CONTEXT_GENERALIZATION
Task summary
Unseen cell-context CRISPRi response prediction; three validation cell lines and three final-test cell lines are held back as ground truth.
Stale after
14 days

NMD-VCell current status. Same-context aggregate benchmark infrastructure exists; cross-context transfer and DMD candidate-conditioned perturbation response remain unsupported.

Freshness gate. If this page is older than the freshness window, keep the official task fields visible but mark challenge-readiness as stale until the Arc source and challenge website are rechecked.

Update behavior. Bind task, data, split, metric and timeline fields only from official Arc or challenge-site sources; preserve prior social-announcement snapshots as history, not current task state.

baseline

Baseline layer

4/6 available

Require complex models to beat simple, interpretable comparators.

AVAILABLE
Zero-change / control-mean baselineRepeated-fold RMSE comparison against zero response is frozen.

Retain as a mandatory comparator in every future task.

AVAILABLE
Outer-training response meanLeakage-safe train-mean comparator is evaluated in every outer fold.

Keep fold-specific fitting and reporting.

AVAILABLE
Regularized linear modelThe current safe external model is a strongly shrunk linear/ridge estimator.

Preserve as the interpretable model floor.

AVAILABLE
Target-level pseudo-bulkCells are aggregated to target-level response objects before inference.

Publish the exact aggregation recipe when source-level raw objects are available.

MISSING
Nearest-neighbour perturbation transferNo frozen nearest-neighbour comparator is registered.

Add training-only perturbation similarity and a leakage-safe neighbour baseline.

MISSING
Public foundation-model comparatorNo public model output has been imported under the same split and metric contract.

Evaluate an openly reproducible model only after exact input and output parity is established.

hidden_generalization

Hidden generalization layer

1/6 available

Separate same-dataset interpolation from biologically meaningful transfer.

AVAILABLE
Same-context unseen perturbationTwenty deterministic repeated balanced five-fold realizations hold out entire HepG2 perturbation targets.

Keep this as the lowest generalization tier, not as disease transfer.

MISSING
Leave-one-donor-outThe current direct perturbation substrate does not expose multiple donors.

Acquire donor-resolved muscle/DMD perturbation data and lock donor-disjoint splits.

MISSING
Leave-one-dataset-outExternal muscle datasets have different roles and do not form a harmonized perturbation response panel.

Create accession-disjoint training and test objects after modality harmonization.

MISSING
Leave-one-disease-stage-outNo stage-resolved perturbation outcome matrix is available.

Define stage labels and hold out complete disease stages, never random cells.

PARTIAL
Cross-cell-type / context transferA 55-target Frangieh external aggregate-delta diagnostic is executed and concludes NO_DIRECTIONAL_TRANSFER_SUPPORT; it is not muscle or DMD truth.

Generate or import matched muscle-context perturbation outcomes.

PARTIAL
Perturbation-conditioned time courseGSE52529 supplies unperturbed 0/24/48/72-hour myogenic reference states only.

Sample the same candidate perturbations across matched differentiation timepoints.

multi_metric

Multi-metric layer

3/8 available

Prevent a single metric from hiding magnitude, direction or distribution failures.

AVAILABLE
Global expression errorRepeated-fold RMSE is frozen; MAE-delta is now executed on the 55-target external diagnostic beside identical zero and train-mean comparators.

Rebind MAE only after the official 2026 expression scale and prediction unit are known.

AVAILABLE
Perturbation direction recoveryRaw/residual cosine plus external delta Pearson and delta Spearman endpoints are reported; the external directional result remains negative.

Retain the current negative result: raw-direction support is mixed.

MISSING
DES / DEG recoveryNo VCC-compatible differential-expression score is computed.

Add up/down DEG precision-recall and threshold sensitivity using cell-eval-compatible outputs.

PARTIAL
PDS / perturbation discriminationRaw L1 retrieval, a scale sweep and truth-norm-matched sensitivity are executed on 55 aggregate external targets; this is not an official AnnData or challenge run.

Re-execute with official 2026 AnnData inputs and preserve raw plus scale-audit outputs.

MISSING
Single-cell distribution distanceThe current model predicts target-level aggregates, not cell distributions.

Gate Wasserstein/MMD or cell-eval distribution metrics on genuine cell-level predictions.

PARTIAL
Pathway / program recoveryFrozen pathway panels are descriptive overlaps, not prediction-versus-truth metrics.

Score prespecified program changes against held-out observed responses.

MISSING
Cell-state proportion recoveryNo generative cell-distribution output or matched proportion truth is available.

Evaluate only when model outputs represent cell-level distributions.

AVAILABLE
Uncertainty and abstentionHierarchical bootstrap intervals, strict gates and an abstention boundary are frozen.

Extend uncertainty to donor, dataset and disease-stage components.

disease_loop

Disease closure layer

1/6 available

Link expression predictions to independent NMD/DMD biological outcomes.

AVAILABLE
DMD observational contextFour frozen DMD evidence channels expose observed direction and source agreement.

Use as context qualification, never as perturbation outcome.

AVAILABLE_BOUNDED
Human-myoblast perturbation phenotypeStage B0 retains 250 GSE293514 fusion hits, 125 individually validated genes and a 395-gene state signature; 9/21 current candidates were assessed and none was a fusion hit.

Use the observed network as a healthy-myoblast reference, then add matched DMD/control molecular and functional perturbation outcome.

PARTIAL
DMD-correction transcriptomeGSE272233 is an orthogonal correction reference; candidates were not perturbed.

Use as a disease-state benchmark, not as candidate causal evidence.

PARTIAL
Prospective prediction registrySchemas and an append-only outcome board exist, but contain zero registered outcomes.

Register predictions before assays and bind them to immutable outcome definitions.

MISSING
Regeneration, fibrosis, inflammation and vascular endpointsNo unified candidate perturbation outcome panel measures these disease functions.

Prespecify one functional primary endpoint and bounded secondary endpoints per study.

MISSING
Motor function / drug or clinical responseNo validated link from current model outputs to clinical outcomes exists.

Keep clinical utility locked until independent prospective validation.

Why this upgrade is needed

Arc's 2025 postmortem reports that almost all submissions were worse than the baseline on MAE, while the broader Generalist evaluation used seven metrics. The public cell-eval suite can compare predicted and observed AnnData objects, but NMD-VCell cannot run equivalent distribution metrics until genuine cell-level predictions and matched truth exist.

Boundary. This readiness matrix measures the distance between benchmark readiness and disease-level virtual-cell validity. It does not claim VCC-compatible performance, zero-shot NMD generalization, DMD response prediction, treatment ranking or clinical validity.