# Figure legends — v21 evidence upgrade

## Figure 3 | Unique-target cross-fitting reveals modest predictive signal

**a,** Historical repeated holdouts tested 1,646 of 2,160 eligible HepG2 perturbation targets, reused 787 targets and never tested 514. The replacement five-fold cross-fitting design assigns all 2,160 targets to one non-overlapping outer fold (432 targets per fold) and uses 32 leakage-controlled external or curated pre-model features; DepMap features and HepG2-wide outcome summaries are excluded. **b,** Mean target-level paired effects with bootstrap 95% confidence intervals for the safe model and three negative controls. Negative RMSE differences and positive cosine differences favor the model. Filled points meet the intended within-model Holm-adjusted directional test. The outcome-coordinate permutation is off-scale for raw cosine and is annotated at its mean value. **c,d,** Target-level effects across frozen tertiles of the primary reliability measure (720 targets per tertile), with bootstrap 95% confidence intervals. Filled points indicate stratum-level support after within-stratifier Holm correction; pₕ denotes the Holm-adjusted continuous trend test. The safe model reduces mean RMSE by 5.10% relative to zero and 0.41% relative to the train-mean baseline, with a raw cosine gain of 0.00164. Reliability strongly tracks the versus-zero RMSE and perturbation-specific residual cosine, but not improvement over the train mean or raw directional gain. Together, these data support a statistically detectable but practically modest full-out-of-fold benchmark signal, not strong directional prediction.

## Supplementary Figure S13 | Sampling, coverage and sensitivity audit

**a,** Target accounting for historical repeated holdouts and final unique-target cross-fitting. **b,** Bootstrap 95% confidence intervals for the safe-model raw cosine difference in the superseded unbalanced fold implementation and the corrected balanced implementation; the change in directional support is retained as an implementation-sensitivity audit. **c,** scGPT and GEARS availability across the same 2,160-target denominator. **d,** Standardized mean differences between covered and uncovered targets for prespecified pre-model variables, with bootstrap 95% confidence intervals. Filled points and asterisks mark signals detected after within-model Holm correction. No audited pre-model selection variable is detected for scGPT, whereas GEARS coverage is enriched for DMD-consensus and DepMap availability. **e,** Pearson correlations between four reliability definitions and target-level model effects; asterisks mark Holm-adjusted trend p<0.05. The guide-pair analysis is exploratory because only 133 targets are available. **f,** Target-level effects against the primary reliability measure. Black points show means across 20 equal-frequency bins and dashed lines show the frozen tertile cutpoints. These audits establish that baseline choice, coverage and reliability materially constrain interpretation.
