# Sealed perturbation-response transfer exposes cell-state limits in a neuromuscular virtual-cell framework

**Version:** v09.38 Figure 1/2/3/5/6-synced Science Advances manuscript candidate, 2026-08-20  
**Article type:** Research Article  
**Evidence state:** sealed diagonal response-transfer holdout failed; Figures 1, 3 and 5 redrawn; Figures 2 and 6 light-polished; Figure 1-6 callouts, captions and submission figure files synchronized; public release and DOI deposition pending  
**Claim ceiling:** evidence-governance framework, bounded perturbation-signature retrieval, and rejected cross-cell-line diagonal response-transfer claim under a sealed Replogle K562-to-RPE1 holdout; no DMD wet-lab therapeutic efficacy, safety, clinical utility, completed prospective wet-lab validation, or post-hoc replacement transfer claim.

## Authors

[Author list, affiliations and corresponding-author details required.]

## Abstract

Computational disease resources increasingly connect patient atlases, perturbation screens, and virtual-cell models, but claims can drift across biological context, endpoint, and inferential unit faster than they can be experimentally validated. We built NMD-VCell as an evidence-to-experiment framework that keeps each prediction linked to its disease context, assay endpoint, unit structure, claim boundary, abstention reason, and next experiment. Duchenne muscular dystrophy (DMD) was used as the stress test: 21 inherited candidate hypotheses had source-context perturbation evidence and computational DMD projections, but 0 of 21 had candidate-level responses measured in replicated DMD muscle. To test whether source-context perturbation responses could be promoted across cell context, we locked a diagonal gene-wise shrinkage model and evaluated it once on an external Replogle K562-to-RPE1 essential-perturbation holdout. Before holdout evaluation, the denominator, comparators, and primary criterion were frozen. The locked model failed both primary gates across 373 perturbations and 1,999 scored genes: source identity outperformed diagonal shrinkage (paired median comparator-minus-model angular-risk difference, -0.175; 95% CI, -0.212 to -0.137; one-sided sign-flip p = 1.0), and target-context mean response outperformed it by a larger margin (-0.421; 95% CI, -0.518 to -0.329; p = 1.0). Vector diagnostics showed broad directional and magnitude failure, and post-hoc Reactome anatomy localized excess error to cell-cycle, mitotic-spindle, DNA-replication, and translation programs. These results support a stricter conclusion than positive transfer: source-derived gene-wise shrinkage can fail under cell-context shift when target responses depend on coordinated cell-state programs. NMD-VCell therefore functions as a claim-governed framework for preserving missing experiments, negative evidence, and transferable endpoints without converting them into unsupported therapeutic or cross-context response claims.

## Introduction

Public transcriptomic atlases, perturbation screens and machine-learning models now make it possible to connect disease-associated cell states with experimentally actionable hypotheses. The most consequential errors often occur at the interfaces between these resource types. A perturbation measured in a transformed cell line can be displayed beside a disease signature as if it were measured in disease-relevant muscle; a pathway-overlap graph can be read as a causal regulatory model; thousands of cells can be treated as thousands of independent donors; and a missing validation experiment can disappear behind a continuous ranking.

This problem is acute across neuromuscular diseases because the relevant biological logic is not uniform. DMD is anchored by dystrophin loss, with locus correction as a mechanism-specific reference axis rather than a universal therapeutic-response vector. Facioscapulohumeral muscular dystrophy (FSHD) contains a sparse DUX4-driven transcriptional state. Myotonic dystrophy type 1 (DM1) is organized around a toxic DMPK repeat and event-level splice dysregulation. Spinal muscular atrophy (SMA) combines SMN restoration with residual neural and muscle phenotypes. A shared interface is useful, but a shared therapeutic score would flatten the intervention, endpoint and experimental unit that make each disease interpretable.

Open Targets, CZ CELLxGENE Discover and scPerturb support target-disease evidence integration, single-cell distribution and perturbation-data harmonization [@buniello2025opentargets; @abdulla2025cellxgene; @peidli2024scperturb]. Models such as CPA, GEARS, scGPT and Geneformer provide increasingly general prediction machinery [@lotfollahi2023cpa; @roohani2024gears; @cui2024scgpt; @theodoris2023geneformer]. Recent benchmarking has emphasized that interpolation, perturbation-specific direction and cross-study generalization are different tasks [@vinas2026systema; @wei2026benchmark]. Disease-focused use therefore requires a decision object that records what was measured, in which model and biological unit, which transformation was applied, what conclusion is prohibited and which minimum experiment would support, revise or retire the hypothesis.

NMD-VCell implements this decision object and stress-tests it in DMD. The DMD candidate set was not selected de novo by this study; it is an inherited 21-gene checkpoint with visible provenance limits. All 21 candidates have source-context perturbation measurements and computational DMD projections, but none has a direct response measurement in replicated DMD muscle. The manuscript therefore asks a narrower and more falsifiable question: which parts of perturbation evidence can be transferred across datasets, and where does a sealed external holdout force the claim to stop?

This paper makes three verifiable contributions. First, it defines an evidence contract that keeps disease context, endpoint, unit structure, lifecycle state, and claim ceiling attached to each neuromuscular prediction object. Second, it shows that DMD candidate evidence remains blocked at the missing DMD-muscle perturbation layer despite source perturbation and graph-projection support. Third, it locks and tests a cross-cell-line perturbation-response transfer claim in an external Replogle K562-to-RPE1 holdout. The locked diagonal gene-wise shrinkage model fails against both frozen primary comparators, and the failure is concentrated in coordinated proliferation and mitotic programs. The resulting contribution is not a positive transfer claim, but a boundary condition for virtual-cell evidence transfer.

## Results

### NMD-VCell makes the missing experiment the primary output

NMD-VCell separates question identity from evidence metadata (Figure 1a). A research question fixes disease, species, tissue or cell state, perturbation modality and direction, dose when available, time, assay endpoint and comparator. Evidence metadata then record origin, context match, observation and inferential units, lifecycle, outcome state, permitted inference, prohibited interpretation and minimum next experiment. Missing target-context measurement is not encoded as a zero biological effect.

The interface begins with the user decision rather than dataset inventory. Four disease routes share one evidence schema without sharing a therapeutic score (Figure 1b). The current workbench exposes 37 disease objects and 21 DMD decision contracts, distinct from the released gene-record search layer. Each contract states why a gene remains in the candidate set, what can be decided now, which conclusion is unavailable, the highest missing evidence layer, the next calibration action and the blocked escalation action.

Across the DMD registry, present actions comprise computational replications, risk or context reviews, expression checks, muscle assays and one external-audit hold (Figure 1c). The lifecycle contains 21 draft Study Cards, one checksum-addressed local ADAM10 Study/Prediction/Endpoint/Outcome-access lock pending public DOI deposit and zero completed outcomes. These denominators prevent interface completeness from being mistaken for experimental completion.

### DMD reference axes bound transfer before candidate claims

GSE233606 was used to construct engineered-DMD-versus-healthy and patient-DMD-versus-healthy gene effects during human myogenic differentiation [@mozin2024dystrophin]. Across 23,324 genes detected in at least 1% of cells, engineered-DMD and patient-DMD effects had Pearson correlation 0.615, Spearman correlation 0.450, cosine similarity 0.615 and directional agreement 68.3% (Figure 2a; Supplementary Table S1). Because each condition is represented by one library, this axis describes a measured within-experiment relationship rather than population-level replication. The selected-gene intersection is shown separately and is not used to change the primary denominator (Figure 2b).

GSE272233 supplied disease and DMD-locus-correction contrasts in dup2, dup2-9 and dup8-9 backgrounds [@lemoine2024correction] (Figure 2c). Cosine similarity to the negative disease-associated expression axis was 0.139 for dup2, 0.136 for dup2-9 and 0.419 for dup8-9. Although 171 genes reversed in at least two backgrounds, only eight reversed in all three (Figure 2d). Direct repair of the disease gene therefore supplies a mechanism-specific reference rather than a universal therapeutic-response vector.

GSE277637 was retained as an organoid transfer stress test [@kindler2025organoid]. Median transfer to the GSE233606 patient axis was 0.038 at gene level and 0.030 at Reactome pathway level across three patient lines compared with one shared control. The weak and line-dependent result is negative transfer evidence, not a failed target claim.

### Candidate provenance and source measurement precede assignment

The DMD candidate set begins at a declared provenance checkpoint (Figure 3a). A historical report records a 58-candidate source queue, and retained code applies positive projected-opposition, direction-risk, broad-essentiality, dependency and priority-tier exclusions. The resulting 21-row evidence vector is available, but the original 58-row TSV and generated v3 manifest are absent. The formal provenance exception starts the auditable lineage at the available 21-row checkpoint; the 21 genes are therefore an inherited hypothesis set, not a new genome-wide discovery.

GSE264667/TRADE supplied direct CRISPRi measurements for all 21 candidates in HepG2 [@nadig2025trade]. Individual candidates had 81 to 269 observed cells, giving 2,559 candidate-labelled cells against 4,976 shared controls. All candidates met the frozen cell-count coverage floor (Figure 3b), but only seven passed the frozen 1000-repeat balanced-resampling criterion. Fourteen candidates had an available target coordinate with a measured negative target-gene delta, and 12 met the frozen quantitative engagement margin of target detected plus delta <= -0.10.

The independent GSE293514 healthy-myoblast screen measured 7,197 screen rows, reported 250 fusion-defect hits and individually validated 125 genes [@zhang2026myoblast]. Nine of the current 21 candidates were covered by this screen and 0/9 were fusion hits; the other 12 were not covered and therefore cannot be treated as measured nulls. ADAM10, CALR and MPHOSPH6 are all assessed non-hits in this healthy-myoblast fusion endpoint. This is endpoint discordance, not evidence of no DMD effect.

### The computational bridge is diagnosable but not therapeutic validation

Reactome release 97 supplied pathway membership [@ragueneau2026reactome]. The 21 candidates were each tested against 1,378 pathways, giving 28,938 candidate-pathway rows. Of these, 1,830 pairs had candidate-wise cameraPR false-discovery rate below 0.05, and 1,782 retained the same non-zero direction in the direct estimate and both deterministic halves. The shared control pool makes candidate rows correlated, so this is a within-candidate competitive statistic rather than an experiment-wide false-discovery claim.

After the pre-graph size rule excluded 18 broad pathways, the overlap graph contained 1,360 pathway nodes and 11,336 undirected edges. In five-fold pathway holdout, median Spearman correlation was 0.767 and directional agreement was 0.826; degree-preserving rewired-null sensitivity gave median Spearman correlations of 0.456 to 0.477. The graph reconstructs non-random structure within the same annotation system, but this remains a topology diagnostic rather than external biological validation.

Candidate projections were retained separately for three disease source groups: GSE233606 engineered/patient myogenic disease, GSE277637 organoid disease and GSE272233 disease. Across candidate, pathway and source, 3,210 of 86,814 evaluable candidate-pathway-source rows (3.7%) showed projected opposition to the corresponding disease-associated expression axis; none is an observed candidate response in DMD muscle. ADAM10 showed positive RWR projection in all three disease groups; CALR and MPHOSPH6 were positive in two of three. Correction alignment is stored as a mechanism-specific annotation, not as an assignment rule.

All 21 candidates passed the topology applicability prerequisite. Under the executable candidate-assignment layer, seven passed repeated-resampling measurement stability, 16 bulk skeletal-muscle expression feasibility, 14 cross-context dependency caution, 11 RWR disease projection and 12 the quantitative target-engagement margin. The repeated-resampling and target-engagement-margin counts are source-bound to the locked v08.5 rule-freeze table, whereas the muscle-expression, dependency-caution and DMD-source-projection counts are source-bound to the current candidate network summary; these are dependent audit filters, not independent validation streams. ADAM10, CALR and MPHOSPH6 pass all five local audit rules and form the five-rule-concordant pilot stratum (Figure 3c and Figure 3d); the other 18 receive reason-coded abstentions and remain predefined non-advancement strata, not negative predictions. The evidence ladder remains 21/21 source-context responses measured, 21/21 pathway-cascade objects generated, 0/21 candidate responses measured in DMD muscle and 0/21 functionally validated candidates (Figure 1d).

### A sealed Replogle holdout rejects the locked diagonal-shrinkage transfer claim

We next asked whether the locked diagonal gene-wise shrinkage model could transfer perturbation responses from Replogle K562 essential perturbations to Replogle RPE1 essential perturbations under a sealed external-holdout criterion. Before inspecting the RPE1 holdout outcomes, we froze the denominator, comparator set, and primary decision rule. The denominator contained 373 matched perturbations and 1,999 scored genes. The two primary comparators were source identity and the target-context mean response; both primary comparisons required a positive paired median angular-risk difference, a lower 95% bootstrap confidence bound above zero, and one-sided Holm-controlled sign-flip support.

The locked model failed both primary gates (Figure 4a). Against source identity, the paired median comparator-minus-diagonal angular-risk difference was negative (Δ = -0.175, 95% CI -0.212 to -0.137, one-sided sign-flip p = 1.0), indicating that source identity had lower angular risk than the diagonal predictor. Against the target-context mean response, the failure was larger (Δ = -0.421, 95% CI -0.518 to -0.329, p = 1.0). Because the primary endpoint was frozen before holdout evaluation, no secondary analysis can convert this result into support for the locked cross-cell-line response-transfer claim. The appropriate interpretation is therefore rejection of that prespecified claim.

### Vector geometry shows broad directional failure rather than a marginal threshold miss

We then examined whether the failed primary result reflected a narrow decision-threshold artifact or a broader geometric mismatch. Across the 373 perturbations, the diagonal model had a median angular risk of 1.094, worse than target-context mean response (0.686), scalar-calibrated source transfer (0.691), source identity (0.956), and the zero vector (1.000; Figure 4b). The same ranking was visible in best-model counts: target mean was best for 156 perturbations, scalar transfer for 118, diagonal shrinkage for 73, source identity for 25, and the zero vector for 1. Thus, diagonal shrinkage was not merely displaced by a single strong comparator; it lost to both target-context averaging and scalar-calibrated source transfer.

Perturbation-level paired differences showed that this underperformance was widespread (Figure 4c). Diagonal shrinkage had higher angular risk than the target mean response for 283 of 373 perturbations and higher risk than scalar-calibrated source transfer for 287 of 373 perturbations. The vector diagnostics further explained the failure mode: diagonal predictions were directionally anti-aligned at the median (median cosine = -0.094) and strongly compressed in magnitude relative to the target response (median norm ratio to target = 0.158). This pattern is consistent with a source-to-target transfer model that suppresses gene-level effects but does not recover the target cell-state response direction.

### Reactome module anatomy localizes the failure to cell-state programs

Having rejected the primary transfer claim, we performed a post-hoc explanatory analysis to ask whether the excess error was biologically structured. This analysis did not alter the frozen v09.17 endpoint. We scored each gene by diagonal excess absolute error relative to scalar-calibrated source transfer and tested Reactome pathways with 5-500 overlapping scored genes. Across 954 tested pathways, 170 were significant after BH correction and 105 remained significant under BY sensitivity correction. The strongest pathway anchors were cell-cycle and mitotic programs, including Cell Cycle (170 overlapping genes, Cliff's δ = 0.593, BH q = 7.56 × 10^-35), Cell Cycle, Mitotic (142 genes, δ = 0.608, BH q = 3.00 × 10^-31), and Cell Cycle Checkpoints (73 genes, δ = 0.632, BH q = 6.93 × 10^-18).

The same signal appeared in seven predeclared Reactome module families tested after the sealed failure but before module computation (Figure 4d). Four families passed both BH and Holm correction: cell-cycle checkpoints, mitotic spindle/chromosome, DNA replication/repair, and translation/ribosome. The strongest family was mitotic spindle/chromosome (143 genes, Cliff's δ = 0.605, BH q = 2.53 × 10^-33, Holm-adjusted p = 4.34 × 10^-33). When angular risk was recomputed after restricting perturbation vectors to module genes, diagonal shrinkage remained worse inside this family (median risk = 1.408) than scalar transfer (0.395) or target mean response (0.411; Figure 4e). Leading genes in the mitotic module included PTTG1, CKS1B, CENPF, TYMS, AURKA, TOP2A, MAD2L1, and TPX2 (Figure 4f).

### The external holdout defines a boundary condition for perturbation transfer

Together, these results define a boundary condition rather than a successful-transfer result. The sealed primary analysis rejected the locked diagonal-shrinkage transfer claim, and the post-hoc mechanism layer showed that the failure was concentrated in coherent cell-state programs rather than distributed uniformly across the transcriptome. The data therefore support a more conservative but stronger claim: source-derived gene-wise shrinkage can fail under cell-context shift when target responses depend on coordinated proliferation and mitotic programs. This framing preserves the integrity of the sealed holdout while turning the negative result into a mechanistically interpretable stress test for virtual-cell transfer.

### Disease modules preserve negative evidence that a shared score would erase

The disease-route module view makes reuse of the evidence contract explicit without biological flattening (Figure 5a). FSHD retains patient-state heterogeneity separately from causal DUX4 induction (Figure 5b). DM1 shows negligible global expression alignment between repeat excision and patient disease but convergence at event-level splicing genes (Figure 5c). SMA does not show global ASO expression reversal, while treated muscle retains directional OXPHOS, denervation and fibrosis programs and KIF5A decreases after SMN loss in two neuronal contexts (Figure 5d). Across all routes, null and residual results determine the next experiment and remain database content rather than a pooled therapeutic score (Figure 5a).

### Single-cell extensions refine localization without raising the claim ceiling

The FSHD and SMA extensions retained 40,289 and 37,157 cells, respectively, under locally frozen analysis plans with no public timestamped registration (Figure 6a). The prespecified FSHD DUX4-target genotype-by-injury effect was 0.392, and its confidence interval crossed zero (Figure 6b). After eight SMA technical libraries were collapsed to four biological lines, Motor_neuron_like showed the largest marker-defined composition difference, with SMA-minus-control effect -0.339 (Figure 6c). SMA candidate and module effects retained pairwise ranges rather than calibrated small-n claims (Figure 6d). Both extensions remain directional small-n evidence and promote zero candidates to validated targets.

## Discussion

NMD-VCell addresses a practical gap between a disease atlas and a perturbation predictor. Its contribution is not another therapeutic target score, but a claim-governed evidence-transfer framework that keeps question identity, measurement context, unit structure, provenance, computation, negative evidence, and missing experiments attached to one another. The central empirical lesson is that this governance layer can force a model claim to fail cleanly rather than letting a neighboring endpoint substitute for it.

The DMD use case shows why this matters. Source perturbation, pathway propagation, and disease-axis comparison yield a local five-rule-concordant pilot stratum, including ADAM10, CALR, and MPHOSPH6, but the evidence ladder still stops at 0 of 21 direct DMD candidate responses. The 3-of-21 assignment rate is conditional on an inherited, historically prefiltered set and cannot be interpreted as prospective enrichment, genome-wide prioritization performance, or final therapeutic ranking. NMD-VCell is useful precisely because it preserves that blocked state instead of smoothing it into an apparent target list.

The sealed Replogle holdout clarifies the same principle at the model level. A locked diagonal gene-wise shrinkage model did not provide evidence for source-to-target response transfer from K562 to RPE1. It failed against source identity and against the target-context mean response under the frozen primary decision rule. Because the denominator, comparators, and criterion were locked before holdout evaluation, post-hoc analyses cannot convert this result into positive transfer. The failed primary endpoint is therefore not a weakness to hide; it is the stress test that defines the boundary of the claim.

The failure was nevertheless biologically informative. Vector diagnostics showed that diagonal predictions were compressed in magnitude and anti-aligned at the median, indicating a broad geometric mismatch rather than a marginal threshold miss. Reactome analysis localized excess diagonal error to cell-cycle, mitotic-spindle, DNA-replication, and translation programs. These modules should not be read as a replacement endpoint or a causal mechanism proof. Their value is explanatory: they show that the model fails most strongly where target responses depend on coordinated cell-state programs rather than isolated gene-wise scaling.

This distinction separates three evidence layers that are often blurred. Retrieval of matched perturbation identity can be a useful bounded endpoint. Organization of disease and perturbation evidence can guide the next experiment. Prediction of target-context perturbation-response vectors is a stricter claim and can fail even when the first two layers remain useful. For Science Advances, the stronger story is not that the model succeeded everywhere, but that a sealed audit identified exactly where success must not be claimed.

The next evidence transition requires executable public registration rather than interface expansion. Before any DMD panel outcome is generated, all 21 candidate strata should be frozen as advance-eligible or abstention strata, with CRISPR direction, primary estimand, common endpoints, margins, and outcome-classification rules locked. In parallel, future perturbation-response transfer models should be evaluated under the same sealed-comparator discipline, with target-context mean and source-identity baselines retained as hard comparators. Until that outcome layer exists, NMD-VCell should be regarded as an evidence-transfer and experiment-design framework, not a validated therapeutic prediction system.

## Limitations

The sealed response-transfer holdout used Replogle K562-to-RPE1 essential perturbations rather than DMD myogenic perturbations. It therefore tests a cross-cell-context transfer claim in a public perturbation setting, not DMD therapeutic response.

The rejected model was a locked diagonal gene-wise shrinkage model. Its failure does not refute all perturbation-response models, but it does reject this prespecified source-derived shrinkage claim under this denominator, comparator set, and criterion.

The Reactome module analysis was performed after the sealed primary failure. Although the module-family tests were structured and multiplicity-controlled, they are explanatory anatomy rather than a replacement endpoint.

The DMD candidate layer remains unvalidated experimentally. All 21 current candidate responses were measured in source contexts, and 0 of 21 have direct candidate perturbation responses in replicated DMD muscle.

The inherited 21-candidate set has incomplete historical replay provenance because the original 58-row queue and generated v3 manifest are absent. The formal provenance exception prevents an end-to-end replay claim, and the current analysis cannot estimate unbiased genome-wide discovery performance.

Several source datasets have limited biological replication. GSE233606 contains one library per condition; GSE277637 contains three DMD lines and one shared unrelated control; the HepG2 source object lacks biological-replicate metadata; and both analysis-plan-frozen single-cell extensions contain only two biological lines per genotype. Cells, deterministic halves and technical libraries do not repair these design limits.

Reactome edges encode shared genes rather than regulation or temporal causality. Pathway holdout and rewired-null results remain internal annotation-system diagnostics. Stable assignment across three numerical graph models does not establish external graph biology or intervention effect size in DMD muscle.

Public DOI deposition, public repository release, named-creator metadata, license approval, final figure production and journal-compliance checks remain incomplete. The submission package therefore remains a working draft until the external DOI and metadata gates close.

## Materials and methods

### Evidence contract and resource architecture

Each research object records a stable identifier, source accession, question identity, experimental context, biological and technical units, effect definition, uncertainty where estimable, provenance, allowed use, prohibited interpretation, claim ceiling and minimum next experiment. Prospective evidence uses an explicit claim graph: Study defines Prediction; Prediction binds to Estimand; Estimand uses Endpoint; Study produces ObservedOutcome; ObservedOutcome evaluates Prediction; PredictionEvaluation updates Claim and ClaimCeiling. One Study can measure multiple endpoints, and transcriptomic, functional, viability and morphology outcomes are not forced into a single value.

The website implements Diseases, Virtual Cell Studio, Browse, Gene compare, Design, Models, Registry and Data & API modules. Modules define research handoffs rather than rankings. The disease API uses a common JSON shape for DMD, FSHD, DM1 and SMA while retaining distinct anchors, units and claim ceilings. The cross-disease score field is null by design.

### Data acquisition, integrity and unit structure

Official GEO and source-study objects were acquired with accession, URL, file size, gzip-integrity and SHA-256 records. GENCODE release 50 was the transcript-mapping authority for bulk modules [@mudge2025gencode]. Public datasets and permitted roles are listed in the supplementary tables. Observation units, technical units, biological units, inferential units and repeated-measure units were recorded separately when available. Cells localize states and estimate library-level means but do not replace donors, lines, cultures or independently prepared biological replicates.

### DMD reference axes

GSE233606 was used to construct engineered-DMD-versus-healthy and patient-DMD-versus-healthy gene effects during human myogenic differentiation. Agreement was evaluated over the common unselected gene universe by Pearson and Spearman correlation, cosine similarity and directional agreement. GSE272233 supplied disease and DMD-locus-correction contrasts in dup2, dup2-9 and dup8-9 backgrounds. GSE277637 was analyzed by organoid line as a transfer stress test. GSE293514 supplied an observed healthy-myoblast fusion-perturbation layer; its phenotype is relevant to muscle applicability but is neither DMD-specific nor a transcriptomic reversal assay.

### Candidate source responses and local audit rules

GSE264667/TRADE supplied direct CRISPRi measurements for all 21 candidates in HepG2. Candidate-labelled cells were compared with a common control pool, and deterministic halves assessed internal gene- and pathway-direction stability. Reactome release 97 supplied pathway membership. Candidate-versus-control gene statistics entered competitive cameraPR testing. A candidate-pathway pair was retained as a measured source seed only when candidate-family false-discovery rate was below 0.05 and the direct, split-A and split-B effects had the same non-zero direction. Pathways formed an overlap graph, with cosine-normalized shared-gene edges and up to 12 eligible neighbours per node. Random-walk propagation with restart 0.35 produced a topology-cascade object. Graph edges represent annotation overlap, not regulation or temporal causality.

The current five-rule architecture was finalized retrospectively during methodological audit and before any DMD candidate-level experimental outcome existed. It was not preregistered before analysis of the retrospective source data and is not a final therapeutic gate system. Candidate assignment used source-measurement adequacy, bulk skeletal-muscle expression feasibility, cross-context dependency caution, projected opposition to the disease-associated expression axis in at least two of three DMD source groups and observed source-context target engagement. These are five dependent local audit rules, not five independent evidence streams.

### Frozen perturbation-signature retrieval endpoint

The v09 validation sequence separated development failure, endpoint conversion and independent validation. The initial point-loss metric was evaluated only as a development candidate and was not advanced because it failed the prespecified paired win-rate criterion. The replacement endpoint was a retrieval endpoint aligned with candidate prioritization: for each perturbation, split-A information was used to rank the matching split-B perturbation against eligible alternatives by residualized angular risk.

The response panel, split rule and endpoint were frozen before independent RPE1 target noncontrol expression access. The validation used the 2,000-gene frozen Replogle response panel and deterministic cell split `sha256(cell_id|NMDVCELL-CAL-V09.1) first-byte parity`. RPE1 perturbation deltas used same-split, gem_group-weighted matched CONTROL baselines. Eligible perturbations required at least 50 total target cells, at least 5 cells in each split, matched control support in observed gem_group/split strata and at least 300 eligible perturbations overall. The retrieval validation endpoint ranked each perturbation's matching split-B signature among all eligible RPE1 perturbations by source-half residualized angular risk. No second retrieval metric, endpoint change, threshold relaxation or panel change was run after target noncontrol expression access.

### Sealed Replogle K562-to-RPE1 response-transfer holdout

The diagonal source-to-target response-transfer model, denominator, comparator set, and primary criterion were locked before the Replogle RPE1 target outcomes were evaluated. The denominator contained 373 matched perturbations and 1,999 scored genes. The tested model applied gene-wise diagonal shrinkage estimated from source-context information. The frozen primary comparators were source identity and target-context mean response. The primary criterion required a positive paired median comparator-minus-model angular-risk difference, a lower 95% bootstrap confidence bound above zero, and one-sided Holm-controlled sign-flip support for both primary comparisons. No comparator replacement, denominator revision, response-panel change, threshold relaxation, or secondary-rescue criterion was permitted after holdout evaluation. Vector diagnostics and Reactome module analyses were performed only after the sealed primary endpoint failed and were interpreted as explanatory anatomy.

### Disease modules and single-cell extensions

FSHD patient-state evidence uses GSE143452 [@jiang2020fshd2], and causal DUX4 induction uses GSE205421 [@brennan2022dux4mapk]. DM1 uses GSE127296 [@andre2019dm1isogenic] for repeat-bearing versus repeat-excised clones and GSE201255 [@hale2023congenitaldm1] for independent adult and congenital biopsy records. SMA uses GSE290979 [@faravelli2026smaorganoid], GSE252128 [@grandi2024smamuscle] and GSE302774 [@akiyama2026kif5a]. Molecular, morphology, function and patient outcomes are not collapsed into a single treatment-response measure.

The FSHD and SMA single-cell analysis plans, seed and thresholds were frozen before complete matrices were read, but no public timestamped registration exists. GSE303359 contains two FSHD and two healthy myoblast lines, each measured at baseline and after 6-hour injury. GSE290980 contains two SMA and two control donor-derived organoid lines, each represented by two technical libraries. Libraries were merged to biological lines before composition, gene and module effects were calculated. With two lines per genotype in both extensions, estimates are reported as directional small-n evidence.

### Reproducibility and cross-surface validation

The release stores source manifests, frozen tables, figure-table traces, key numerical-claim traces, model-run records and package checksums. Local manuscript, API and interface objects identify build EA-20260817-57; independent production verification remains a release gate. Website inventory counts are not interpreted as independent validation counts. Six historical ModelRuns are recorded, with zero calibrated DMD candidate-response runs.

## Data availability

The live resource is available at [https://nmdvcell.com/resource/diseases/](https://nmdvcell.com/resource/diseases/). The module meaning map is available at [https://nmdvcell.com/resource/meaning-map/](https://nmdvcell.com/resource/meaning-map/), the machine-readable semantic contract at [https://nmdvcell.com/resource/api/v1.1/module_semantic_contract.json](https://nmdvcell.com/resource/api/v1.1/module_semantic_contract.json), the evidence-governance schema at [https://nmdvcell.com/resource/api/v1.1/evidence_governance_schema.json](https://nmdvcell.com/resource/api/v1.1/evidence_governance_schema.json), the disease manifest at [https://nmdvcell.com/resource/api/v2/diseases/manifest.json](https://nmdvcell.com/resource/api/v2/diseases/manifest.json), the cross-surface consistency contract at [https://nmdvcell.com/resource/api/v1.1/research_consistency_contract.json](https://nmdvcell.com/resource/api/v1.1/research_consistency_contract.json), and downloads at [https://nmdvcell.com/resource/downloads/multidisease-v02/](https://nmdvcell.com/resource/downloads/multidisease-v02/).

The v09.5 endpoint-lock evidence snapshot is locally packaged as `release_snapshots/nmd-vcell-v09.5-sa-endpoint-lock-20260819.tar.gz` with SHA256 `00e59c7df45d15c37591de4b990058d58b1da91174708c768dbcc5fe3c468815`. The website synchronization object is `deploy/nmd-vcell-site/public/resource/api/v1.1/sa_endpoint_lock_v095.json`. Public DOI deposition, repository release, named-creator metadata and license approval remain required before submission. The prospective DMD candidate Perturb-seq outcome dataset does not yet exist.

## Code availability

Analysis and build scripts are included in the project workspace and curated package with file-level checksums. They implement official-file acquisition, biological-unit-aware analysis, serialization repair, figure generation, website consistency checks, executable-rule synchronization, endpoint-lock construction and document rendering. A public versioned repository and DOI archive are required before submission. The missing historical 58-row DMD queue prevents end-to-end replay of the original 58-to-21 selection.

## Ethics statement

This study is a secondary computational analysis of public, de-identified data and did not recruit participants or generate new human biospecimens. Ethical approvals, consent and data-use restrictions belong to the original studies and require final author verification. The proposed DMD Perturb-seq experiment has not been performed and would require all applicable institutional approvals.

## Author contributions

[Complete using the CRediT taxonomy after the author list is fixed.]

## Acknowledgements

[Funding sources, institutional support and data-generator acknowledgements required.]

## Funding

[Author completion required.]

## Competing interests

[Author declaration required.]

## AI use disclosure

Generative AI tools supported drafting, code organization, audit-table assembly and review-package generation. All claims, citations and computational outputs require named-author verification, and the authors retain responsibility for the final work.

## Figure captions

### Figure 1

Figure file: `figures/Figure_1_evidence_governance_v09_38.pdf` (matched PNG/SVG/TIFF also provided in the v09.38 packet).

**Figure 1. Evidence objects preserve context, unit structure and claim boundaries.**  
**a,** A research question identity, an evidence-state record and a claim contract are represented as separable objects. Transformations may change representation but cannot silently change evidence origin, context match, biological or inferential unit, lifecycle state or outcome state. **b,** DMD, FSHD, DM1 and SMA share the evidence-contract schema while retaining disease-specific biological anchors; no pooled cross-disease therapeutic score is released. **c,** Lifecycle states make prospective work auditable without converting local locks or pending DOI deposition into completed outcomes. **d,** The DMD evidence ladder stops before direct disease-context candidate validation: 21 of 21 candidates have source-context responses and 21 of 21 have pathway-cascade objects, but 0 of 21 have direct DMD candidate responses and 0 of 21 are functionally validated. **Guardrail.** This figure defines the evidence-governance object and DMD claim ceiling; it does not establish therapeutic efficacy, safety, clinical utility or candidate validation.

### Figure 2

Figure file: `figures/Figure_2_dmd_axes_v09_38.pdf` (matched PNG/SVG/TIFF also provided in the v09.38 packet).

**Figure 2. Measured DMD axes define references, not candidate outcomes.**  
**a,** The unselected 23,324-gene engineered-versus-patient DMD comparison is the primary denominator and is shown without differential-expression selection. **b,** The selected significant-gene intersection shows higher apparent directional agreement and is displayed separately as a conditioning-sensitive contrast. **c,** DMD-locus correction alignment varies across dup2, dup2-9 and dup8-9 backgrounds. **d,** Significant-gene reversal is genotype dependent; 171 genes reverse in at least two backgrounds and only eight reverse in all three. These measured DMD axes are reference evidence, not candidate-level DMD perturbation outcomes.

### Figure 3

**Figure 3. Candidate provenance and source-bound assignment audit define a bounded pilot stratum without DMD validation.**  
**a,** The DMD candidate lineage starts at the first checksum-addressed recoverable checkpoint: a historical 58-candidate source queue is reported, but the original source TSV and full replay manifest are unavailable, so the retained 21-row vector is treated as an inherited hypothesis set rather than a new genome-wide discovery result.  
**b,** Coverage and endpoint evidence remain distinct: 21 of 21 candidates have source-context responses and pass the HepG2 coverage floor; 9 of 21 are covered by an independent healthy-myoblast fusion screen, 0 of 9 are fusion-defect hits, and 12 of 21 are not measured in that endpoint.  
**c,** Five dependent local audit rules summarize current assignment. Seven of 21 candidates pass repeated-resampling stability, 16 of 21 pass bulk skeletal-muscle expression feasibility, 14 of 21 pass cross-context dependency caution, 11 of 21 pass DMD-source projection, and 12 of 21 pass target-engagement margin. The repeated-resampling and target-engagement-margin counts are taken from the locked v08.5 rule-freeze table, whereas the muscle-expression, dependency-caution and DMD-source-projection counts are taken from the current candidate network summary; these are dependent audit filters, not independent validation streams. ADAM10, CALR and MPHOSPH6 pass all five and form the current 3-of-21 pilot stratum; the remaining 18 receive reason-coded abstentions.  
**d,** Direct-seed, no-propagation and RWR graph models are numerically distinct but converge on the same current pilot set.  
**Guardrail.** This figure shows provenance recovery, source-context adequacy and conditional assignment robustness. It does not show candidate responses measured in DMD muscle, therapeutic efficacy, safety, clinical utility or functional validation.


### Figure 4

Figure file: `figures/Figure_4_boundary_holdout_v09_38.pdf` (matched PNG/SVG/TIFF also provided in the v09.38 packet).

**Figure 4. A sealed Replogle holdout exposes a cell-state boundary for source-to-target perturbation transfer.**  
**A,** Frozen primary comparisons for the locked diagonal shrinkage model in Replogle K562 essential perturbations transferred to Replogle RPE1 essential perturbations. Points show paired median comparator-minus-diagonal angular-risk differences; bars show bootstrap 95% confidence intervals. The shaded region marks the prespecified pass direction. Both comparisons fail. **B,** Median angular risk across reference models. Lower values are better; labels report median risk and the number of perturbations for which each model was best. **C,** Perturbation-level paired risk differences show that diagonal shrinkage was broadly worse than target-context and scalar-calibrated comparators. Points mark medians and thick bars mark interquartile ranges. **D,** Post-hoc Reactome module-family enrichment for gene-level diagonal excess error relative to scalar calibration. Points show Cliff's δ; size indicates scored genes; color indicates -log10 BH q; black outlines indicate Holm-significant families. **E,** Median angular risk recomputed after restricting perturbation vectors to module genes. **F,** Leading mitotic/chromosome module genes ranked by diagonal excess error relative to scalar calibration. **Guardrail.** Panels D-F are explanatory mechanism anatomy after the sealed primary failure; they are not a replacement primary endpoint.

### Figure 5

Figure file: `figures/Figure_5_disease_module_matrix_v09_38.pdf` (matched PNG/SVG/TIFF also provided in the v09.38 packet).

**Figure 5. Disease modules preserve negative and residual evidence instead of collapsing routes into a shared score.**  
**a,** FSHD, DM1, SMA muscle and SMA neuronal routes are represented with the same evidence-preservation grammar: measured axis, retained negative or residual result, claim ceiling and next experiment. FSHD separates patient-state and DUX4-induction evidence; DM1 separates global-expression non-reversal from splice-event convergence; SMA separates weak global ASO reversal from residual muscle programs and neuronal KIF5A response. **b,** FSHD target-high nuclei remain a sample-dependent state-localization signal and do not constitute donor replication. **c,** DM1 retains negligible global expression alignment between repeat excision and patient disease while preserving event-level splice convergence as the interpretable endpoint. **d,** SMA preserves weak global ASO reversal, residual OXPHOS/denervation/fibrosis muscle programs and KIF5A decreases after SMN loss as separate measured questions. **Guardrail.** The modules share a claim-governed evidence contract, not a pooled therapeutic score. Negative and residual results define next experiments and do not establish therapeutic efficacy, safety, clinical utility or validated targets.

### Figure 6

Figure file: `figures/Figure_6_single_cell_extensions_v09_38.pdf` (matched PNG/SVG/TIFF also provided in the v09.38 packet).

**Figure 6. Single-cell extensions refine localization without raising the claim ceiling.**  
**a,** GSE303359 and GSE290980 analysis plans were frozen locally before complete-matrix access but lack public timestamped registration. **b,** FSHD genotype-by-injury module intervals cross zero, including the prespecified DUX4-target module. **c,** SMA marker-state composition is shown after eight technical libraries collapse to four biological lines. **d,** SMA gene and module effects retain pairwise ranges rather than calibrated small-n claims. Both analyses remain directional at two lines per genotype and promote zero targets.

## References

::: {#refs}
:::
