NMD-VCell Neuromuscular Virtual Cell Research Platform Module: Registry · Evidence → perturbation → experiment → outcome NMD = neuromuscular disorders

Release 2026.08DMD context observedDMD candidate-conditioned prediction not yet eligible

View scientific status
Evidence freeze: 3 August 2026 Resource: v1.2.0-measured-dmd-evidence Schema: 1.1 Open release status →

Core HepG2 benchmark

Core benchmark · same-context only

HepG2 repeated-fold model audit

This route contains only the frozen same-HepG2 benchmark. External-transfer and future-challenge readiness have separate denominators and routes.

← Benchmark center

Four benchmark questions

What does the safe model actually support?

Plain-language result: The current model slightly improves average prediction error inside the HepG2 dataset, but it does not reliably recover response direction and did not transfer to an external perturbation dataset.

Average prediction errorSlightly better than a simple baseline
Direction of changeNot reliable
External datasetDid not transfer
DMD predictionNot tested

What this means: the model is useful as a cautious same-dataset baseline, not as a validated predictor of DMD muscle response.

Beats zero baseline?

Yes · repeated RMSE support

RMSE improvement (baseline − model) 0.006292 (95% CI 0.005240 to 0.007352). Positive means lower model RMSE than the baseline.

Beats train mean?

Yes · modestly

RMSE improvement (baseline − model) 0.000466 (95% CI 0.000225 to 0.000702). Positive means lower model RMSE; this is a small same-dataset improvement.

Raw direction robust?

No · mixed

0.001224 (95% CI -0.000216 to 0.002498). The interval crosses zero; raw-direction recovery is not supported.

Residual structure?

Reproducible within HepG2

0.045281 (95% CI 0.029650 to 0.061178). Positive within this dataset; it does not establish disease-context validity.

One presentation directionAll displayed effects use positive = favourable. RMSE differences are sign-flipped from the frozen audit objects only for presentation; raw source values remain unchanged. Mean across 20 repeated balanced-fold realizations; 2,160 held-out targets per realization; hierarchical bootstrap over realizations and target modules (10,000 replicates). Source: typed summary object and table. Benchmark release repeated-fold-v2.2.

Visual benchmark interpretation

Five panels for the four benchmark questions

The figure below translates the frozen benchmark objects into a journal-style summary: RMSE support, raw-direction uncertainty, strict-versus-legacy gate behaviour, absolute error scale and target-level direction counts.

Five-panel Benchmark interpretation figure summarising RMSE effects, direction support, strict gates, absolute RMSE scale and target distribution.
Panel a and b use positive = favourable for interpretation. The underlying JSON and TSV audit files remain unchanged; RMSE signs are flipped only for display. α=1,000 in 26/100 folds · α=10,000 in 74/100 folds.
Regularisation diagnostic

α=1,000 in 26/100 folds · α=10,000 in 74/100 folds

The expanded inner-CV grid reached 100,000,000 and selected its finite upper boundary in 0/100 folds. The selected values still indicate strong shrinkage close to the outer-training response mean, not robust disease-context prediction. Open the diagnostic object.

Strict versus legacy gates

0/16 strict · 16/16 legacy

The 16 outcomes are the frozen direct-head model/split/endpoint combinations. The strict composite requires paired RMSE support, multiplicity control and raw-direction support; the legacy rule accepted the broader RMSE comparator and is retained only as a historical sensitivity.

Absolute scale and target distribution

Model RMSE

0.117507

Across the 20 frozen safe_external_v1 realizations.

Train-mean RMSE

0.117973

Relative improvement: 0.395%.

Zero-change RMSE

0.123799

Relative improvement: 5.082%.

Target-level direction

1128 improved · 1032 worsened

Across 2,160 targets; 0 were exactly unchanged.

Audit all 16 frozen strict outcomes
Outcome IDModelSplitEndpointFavourable-direction effect95% CIMultiplicity-adjusted PStrict gate
STRICT-01 ridge_v1 val rmse gain vs zero 0.005559 -0.004011 to 0.015838 Not available for validation split FAIL
STRICT-02 ridge_v1 val rmse gain vs train mean 0.007689 0.000406 to 0.01488 Not available for validation split FAIL
STRICT-03 ridge_v1 test rmse gain vs zero 0.006045 -0.00107 to 0.013744 0.641587 FAIL
STRICT-04 ridge_v1 test rmse gain vs train mean 0.006371 -0.001359 to 0.013912 0.641587 FAIL
STRICT-05 mlp_v1 val rmse gain vs zero 0.00665 -0.00271 to 0.016749 Not available for validation split FAIL
STRICT-06 mlp_v1 val rmse gain vs train mean 0.00878 0.000834 to 0.016048 Not available for validation split FAIL
STRICT-07 mlp_v1 test rmse gain vs zero 0.002712 -0.004185 to 0.009659 1 FAIL
STRICT-08 mlp_v1 test rmse gain vs train mean 0.003038 -0.00432 to 0.00985 1 FAIL
STRICT-09 ridge_v2 val rmse gain vs zero 0.002517 -0.011617 to 0.014142 Not available for validation split FAIL
STRICT-10 ridge_v2 val rmse gain vs train mean 0.004647 -0.004459 to 0.012929 Not available for validation split FAIL
STRICT-11 ridge_v2 test rmse gain vs zero 0.006808 0.000178 to 0.014113 0.48959 FAIL
STRICT-12 ridge_v2 test rmse gain vs train mean 0.007134 -0.000712 to 0.014433 0.48959 FAIL
STRICT-13 mlp_v2 val rmse gain vs zero 0.00357 -0.006906 to 0.013639 Not available for validation split FAIL
STRICT-14 mlp_v2 val rmse gain vs train mean 0.0057 -0.001864 to 0.012786 Not available for validation split FAIL
STRICT-15 mlp_v2 test rmse gain vs zero 0.002641 -0.00457 to 0.01017 1 FAIL
STRICT-16 mlp_v2 test rmse gain vs train mean 0.002968 -0.00416 to 0.009724 1 FAIL

Download the typed outcome records. Validation-split rows predate the frozen primary multiplicity family and therefore retain a null adjusted-P field rather than an invented value.

Experimental repeatability context

No formal noise ceiling is available

Descriptive split-half medians are 0.331 (cell split) and 0.333 (batch split); guide-pair median cosine is 0.006 across 133 auditable targets. These are reliability diagnostics, not an experimental ceiling. Independent biological replicates are still required before reporting a fraction of ceiling achieved.

Read the reliability boundary

Biological meaning: the model captures a small amount of target-specific error reduction within the processed HepG2 dataset. It does not mean: robust perturbation direction, muscle validation or therapeutic efficacy.