NMD-VCell Research Workbench Module: Validate / registry · evidence-to-experiment workflow NMD = neuromuscular disorders
v1.0 candidate Frozen 25 Jul 2026 DOI pending
v1.0.0-database-resource Schema 1.1 Model ridge-safe-v2.3 Benchmark repeated-fold-v2.2 Build EA-20260729-15 v1.0 is a database and evidence-governance release; it is not a validated disease-prediction or clinical decision-support release.
Current evidence ceiling L2 observed HepG2 limited L3a context No independent DMD perturbation validation Open boundary

← Benchmark center

VCC 2026 readiness · implementation audit

Prepare the evaluation system before the new task is disclosed

Arc has announced a new prediction problem, a wider metric range and an August 20, 2026 launch. The task, cell context, perturbation modality and scoring formula are not yet public, so this board measures reusable evaluation infrastructure rather than claiming task compatibility.

Task details not announced
Launch20 Aug 2026Arc Virtual Cell Challenge · round two
Current audit9/26 available7 partial · 10 missing
Executed metric adapter55 external targetsAggregate delta only · source result remains negative
Priority before launchOfficial task bindingAnnData, scale, split and metric weights remain undisclosed
Official sourceofficial-organizer-social-announcement-2026-07
Source URL
https://www.linkedin.com/posts/arc-institute-org_start-assembling-your-team-because-the-2026-activity-7477802662269186048-SXyD
Checked at
2026-07-29
Task status
NOT_YET_ANNOUNCED
Stale after
7 days

Stale behavior. Display the last checked date and retain the NOT_YET_ANNOUNCED state until an official task specification is verified.

Update behavior. Bind task, data, split and metric fields only from a versioned official source; preserve the prior announcement snapshot.

baseline

Baseline layer

4/6 available

Require complex models to beat simple, interpretable comparators.

AVAILABLE
Zero-change / control-mean baselineRepeated-fold RMSE comparison against zero response is frozen.

Retain as a mandatory comparator in every future task.

AVAILABLE
Outer-training response meanLeakage-safe train-mean comparator is evaluated in every outer fold.

Keep fold-specific fitting and reporting.

AVAILABLE
Regularized linear modelThe current safe external model is a strongly shrunk linear/ridge estimator.

Preserve as the interpretable model floor.

AVAILABLE
Target-level pseudo-bulkCells are aggregated to target-level response objects before inference.

Publish the exact aggregation recipe when source-level raw objects are available.

MISSING
Nearest-neighbour perturbation transferNo frozen nearest-neighbour comparator is registered.

Add training-only perturbation similarity and a leakage-safe neighbour baseline.

MISSING
Public foundation-model comparatorNo public model output has been imported under the same split and metric contract.

Evaluate an openly reproducible model only after exact input and output parity is established.

hidden_generalization

Hidden generalization layer

1/6 available

Separate same-dataset interpolation from biologically meaningful transfer.

AVAILABLE
Same-context unseen perturbationTwenty deterministic repeated balanced five-fold realizations hold out entire HepG2 perturbation targets.

Keep this as the lowest generalization tier, not as disease transfer.

MISSING
Leave-one-donor-outThe current direct perturbation substrate does not expose multiple donors.

Acquire donor-resolved muscle/DMD perturbation data and lock donor-disjoint splits.

MISSING
Leave-one-dataset-outExternal muscle datasets have different roles and do not form a harmonized perturbation response panel.

Create accession-disjoint training and test objects after modality harmonization.

MISSING
Leave-one-disease-stage-outNo stage-resolved perturbation outcome matrix is available.

Define stage labels and hold out complete disease stages, never random cells.

PARTIAL
Cross-cell-type / context transferA 55-target Frangieh external aggregate-delta diagnostic is executed and concludes NO_DIRECTIONAL_TRANSFER_SUPPORT; it is not muscle or DMD truth.

Generate or import matched muscle-context perturbation outcomes.

PARTIAL
Perturbation-conditioned time courseGSE52529 supplies unperturbed 0/24/48/72-hour myogenic reference states only.

Sample the same candidate perturbations across matched differentiation timepoints.

multi_metric

Multi-metric layer

3/8 available

Prevent a single metric from hiding magnitude, direction or distribution failures.

AVAILABLE
Global expression errorRepeated-fold RMSE is frozen; MAE-delta is now executed on the 55-target external diagnostic beside identical zero and train-mean comparators.

Rebind MAE only after the official 2026 expression scale and prediction unit are known.

AVAILABLE
Perturbation direction recoveryRaw/residual cosine plus external delta Pearson and delta Spearman endpoints are reported; the external directional result remains negative.

Retain the current negative result: raw-direction support is mixed.

MISSING
DES / DEG recoveryNo VCC-compatible differential-expression score is computed.

Add up/down DEG precision-recall and threshold sensitivity using cell-eval-compatible outputs.

PARTIAL
PDS / perturbation discriminationRaw L1 retrieval, a scale sweep and truth-norm-matched sensitivity are executed on 55 aggregate external targets; this is not an official AnnData or challenge run.

Re-execute with official 2026 AnnData inputs and preserve raw plus scale-audit outputs.

MISSING
Single-cell distribution distanceThe current model predicts target-level aggregates, not cell distributions.

Gate Wasserstein/MMD or cell-eval distribution metrics on genuine cell-level predictions.

PARTIAL
Pathway / program recoveryFrozen pathway panels are descriptive overlaps, not prediction-versus-truth metrics.

Score prespecified program changes against held-out observed responses.

MISSING
Cell-state proportion recoveryNo generative cell-distribution output or matched proportion truth is available.

Evaluate only when model outputs represent cell-level distributions.

AVAILABLE
Uncertainty and abstentionHierarchical bootstrap intervals, strict gates and an abstention boundary are frozen.

Extend uncertainty to donor, dataset and disease-stage components.

disease_loop

Disease closure layer

1/6 available

Link expression predictions to independent NMD/DMD biological outcomes.

AVAILABLE
DMD observational contextFour frozen DMD evidence channels expose observed direction and source agreement.

Use as context qualification, never as perturbation truth.

PARTIAL
Human-myoblast perturbation phenotypeGSE293514 provides bounded KO/fusion evidence for 9 of 21 candidates.

Expand candidate coverage and measure multiple muscle functions.

PARTIAL
DMD-correction transcriptomeGSE272233 is an orthogonal correction reference; candidates were not perturbed.

Use as a disease-state benchmark, not as candidate causal evidence.

PARTIAL
Prospective prediction registrySchemas and an append-only outcome board exist, but contain zero registered outcomes.

Register predictions before assays and bind them to immutable outcome definitions.

MISSING
Regeneration, fibrosis, inflammation and vascular endpointsNo unified candidate perturbation outcome panel measures these disease functions.

Prespecify one functional primary endpoint and bounded secondary endpoints per study.

MISSING
Motor function / drug or clinical responseNo validated link from current model outputs to clinical outcomes exists.

Keep clinical utility locked until independent prospective validation.

Why this upgrade is needed

Arc's 2025 postmortem reports that almost all submissions were worse than the baseline on MAE, while the broader Generalist evaluation used seven metrics. The public cell-eval suite can compare predicted and observed AnnData objects, but NMD-VCell cannot run equivalent distribution metrics until genuine cell-level predictions and matched truth exist.

Boundary. This readiness matrix is an implementation audit and competition-preparation plan. It does not claim VCC 2026 task compatibility before the rules are released, nor does it convert same-HepG2 benchmark support into muscle, DMD, treatment or clinical validity.