Release 2026.08DMD context observedDMD candidate-conditioned prediction not yet eligible
View scientific status
Evidence freeze: 3 August 2026Resource: v1.2.0-measured-dmd-evidenceSchema: 1.1Open release status →
Model checkBenchmark center
How good is the current model?Four benchmark questions, four independent routes.
Four cards separate what the model can do from what has not yet been tested.Results inside the original dataset, transfer to external data and future out-of-distribution tests are reported separately so their evidence boundaries remain clear.
Product contract · BENCHMARK-CARD-1.0
Benchmark Card: On which sealed truth and permanent baselines did a model advance or stop?
Average prediction errorSlightly better than a simple baseline
Direction of changeNot reliable
External datasetDid not transfer
DMD predictionNot tested
What this means: the model is useful as a cautious same-dataset baseline, not as a validated predictor of DMD muscle response.
Generalization at a glanceExecution stops before unseen family, donor, state and DMD disease-line testsThese lanes are separate tasks, not a cumulative model score.
Every model family faces the same frozen task contract.
No single aggregate score
A model only enters a cell when a reproducible run exists for that exact task. Empty cells remain visible; they are not converted into simulated performance.
Task
Zero baseline
Train mean
Ridge
Advanced adapter
Current verdict
G0 · same-context holdout
Executed
Executed
Executed · limited
Not run
Small RMSE gain; strict gate 0/16
G1 · unseen perturbation
Reference
Reference
Partial · negative
Not run
No directional-transfer support
G2–G5 · family, donor, state, laboratory
Registered
Registered
Registered
Registered
Not executed
G6 · independent DMD line
Locked
Locked
Locked
Locked
No disease-context outcomes
G7 · KO, OE and combinations
Additive null
Locked
Locked
Locked
No learned interaction result
Reading rule: “not run”, “registered” and “locked” describe missing evaluation—not model failure. Advanced models must beat the permanent simple baselines on prospective holdouts before they can support a stronger claim.
NMD Benchmark Suite
Eight tasks; each has its own truth, baseline, metric and blocker.
Only an additive no-interaction preview is available
No observed candidate doubles or matched DMD drug screensOpen task route →
New · v1.3 cross-disease state audit
DMD, FSHD and SMA share one descriptive stress-direction signal.
Sample-aware · not candidate efficacy
The frozen module tables show Stress_response above the within-dataset reference in all three disease contexts. Leave-one-disease-out direction checks match in 3/3 held-out contexts. This is a state-context audit, not a pooled disease score, target-response prediction or therapeutic ranking.
Module
DMD
FSHD
SMA
Interpretation
Stress response
+1.18 · UP
+0.12 · UP
+0.54 · UP
Common module; 3/3 descriptive held-out direction matches
Myogenesis
+1.15 · UP
−0.22 · DOWN
NA
Context-specific; not pooled
Membrane repair
+0.63 · UP
+0.00 · UP
NA
Context-specific; not pooled
OXPHOS
−0.37 · DOWN
NA
−0.23 · DOWN
Context-specific; not pooled
Evidence boundary: DMD sample-level units are 3 case versus 2 control; FSHD and SMA use 2 lines per genotype. The audit adds no candidate-level DMD perturbation outcome: 0/21. It does not unlock a DMD virtual-cell prediction.
Expression error alone is not enough—and raw PDS is not enough.
The Arc Virtual Cell Challenge separates differential-expression recovery, perturbation discrimination and global expression error. NMD-VCell expands this into a seven-family metric firewall and requires PDS scale diagnostics before interpretation.
DESDifferential-expression recoveryDoes the prediction recover perturbation-responsive genes?PDS+Identity with scale safeguardsRaw, norm-matched and amplitude-sensitivity variants stay together.MAEGlobal expression errorIs overall expression close—without hiding a trivial mean predictor?DISTPopulation geometryDo predicted cells occupy the observed distribution?CALCalibration and abstentionDoes uncertainty mean what the model says it means?DMDFunction, replication and safetyDoes molecular fidelity translate into a reproducible disease-relevant result?