Core HepG2 benchmark
Core benchmark · same-context only
HepG2 repeated-fold model audit
This route contains only the frozen same-HepG2 benchmark. External-transfer and future-challenge readiness have separate denominators and routes.
Four benchmark questions
What does the safe model actually support?
Plain-language result: The current model slightly improves average prediction error inside the HepG2 dataset, but it does not reliably recover response direction and did not transfer to an external perturbation dataset.
What this means: the model is useful as a cautious same-dataset baseline, not as a validated predictor of DMD muscle response.
Yes · repeated RMSE support
RMSE improvement (baseline − model) 0.006292 (95% CI 0.005240 to 0.007352). Positive means lower model RMSE than the baseline.
Yes · modestly
RMSE improvement (baseline − model) 0.000466 (95% CI 0.000225 to 0.000702). Positive means lower model RMSE; this is a small same-dataset improvement.
No · mixed
0.001224 (95% CI -0.000216 to 0.002498). The interval crosses zero; raw-direction recovery is not supported.
Reproducible within HepG2
0.045281 (95% CI 0.029650 to 0.061178). Positive within this dataset; it does not establish disease-context validity.
Visual benchmark interpretation
Five panels for the four benchmark questions
The figure below translates the frozen benchmark objects into a journal-style summary: RMSE support, raw-direction uncertainty, strict-versus-legacy gate behaviour, absolute error scale and target-level direction counts.
α=1,000 in 26/100 folds · α=10,000 in 74/100 folds
The expanded inner-CV grid reached 100,000,000 and selected its finite upper boundary in 0/100 folds. The selected values still indicate strong shrinkage close to the outer-training response mean, not robust disease-context prediction. Open the diagnostic object.
0/16 strict · 16/16 legacy
The 16 outcomes are the frozen direct-head model/split/endpoint combinations. The strict composite requires paired RMSE support, multiplicity control and raw-direction support; the legacy rule accepted the broader RMSE comparator and is retained only as a historical sensitivity.
Absolute scale and target distribution
0.117507
Across the 20 frozen safe_external_v1 realizations.
0.117973
Relative improvement: 0.395%.
0.123799
Relative improvement: 5.082%.
1128 improved · 1032 worsened
Across 2,160 targets; 0 were exactly unchanged.
Audit all 16 frozen strict outcomes
| Outcome ID | Model | Split | Endpoint | Favourable-direction effect | 95% CI | Multiplicity-adjusted P | Strict gate |
|---|---|---|---|---|---|---|---|
STRICT-01 |
ridge_v1 | val | rmse gain vs zero | 0.005559 | -0.004011 to 0.015838 | Not available for validation split | FAIL |
STRICT-02 |
ridge_v1 | val | rmse gain vs train mean | 0.007689 | 0.000406 to 0.01488 | Not available for validation split | FAIL |
STRICT-03 |
ridge_v1 | test | rmse gain vs zero | 0.006045 | -0.00107 to 0.013744 | 0.641587 | FAIL |
STRICT-04 |
ridge_v1 | test | rmse gain vs train mean | 0.006371 | -0.001359 to 0.013912 | 0.641587 | FAIL |
STRICT-05 |
mlp_v1 | val | rmse gain vs zero | 0.00665 | -0.00271 to 0.016749 | Not available for validation split | FAIL |
STRICT-06 |
mlp_v1 | val | rmse gain vs train mean | 0.00878 | 0.000834 to 0.016048 | Not available for validation split | FAIL |
STRICT-07 |
mlp_v1 | test | rmse gain vs zero | 0.002712 | -0.004185 to 0.009659 | 1 | FAIL |
STRICT-08 |
mlp_v1 | test | rmse gain vs train mean | 0.003038 | -0.00432 to 0.00985 | 1 | FAIL |
STRICT-09 |
ridge_v2 | val | rmse gain vs zero | 0.002517 | -0.011617 to 0.014142 | Not available for validation split | FAIL |
STRICT-10 |
ridge_v2 | val | rmse gain vs train mean | 0.004647 | -0.004459 to 0.012929 | Not available for validation split | FAIL |
STRICT-11 |
ridge_v2 | test | rmse gain vs zero | 0.006808 | 0.000178 to 0.014113 | 0.48959 | FAIL |
STRICT-12 |
ridge_v2 | test | rmse gain vs train mean | 0.007134 | -0.000712 to 0.014433 | 0.48959 | FAIL |
STRICT-13 |
mlp_v2 | val | rmse gain vs zero | 0.00357 | -0.006906 to 0.013639 | Not available for validation split | FAIL |
STRICT-14 |
mlp_v2 | val | rmse gain vs train mean | 0.0057 | -0.001864 to 0.012786 | Not available for validation split | FAIL |
STRICT-15 |
mlp_v2 | test | rmse gain vs zero | 0.002641 | -0.00457 to 0.01017 | 1 | FAIL |
STRICT-16 |
mlp_v2 | test | rmse gain vs train mean | 0.002968 | -0.00416 to 0.009724 | 1 | FAIL |
Download the typed outcome records. Validation-split rows predate the frozen primary multiplicity family and therefore retain a null adjusted-P field rather than an invented value.
No formal noise ceiling is available
Descriptive split-half medians are 0.331 (cell split) and 0.333 (batch split); guide-pair median cosine is 0.006 across 133 auditable targets. These are reliability diagnostics, not an experimental ceiling. Independent biological replicates are still required before reporting a fraction of ceiling achieved.
Read the reliability boundaryBiological meaning: the model captures a small amount of target-specific error reduction within the processed HepG2 dataset. It does not mean: robust perturbation direction, muscle validation or therapeutic efficacy.