# NMD-VCell: a bounded evidence resource for perturbation-response evaluation and DMD-guided hypothesis triage

> v1.4-v2.4r1 integrity-corrected author-review candidate, 2026-07-22. Statistical results and Figure 1–6/S1–S15 legends are aligned to the isolated v2.4 evidence package. A unified isolated technical candidate with 113 non-empty SQLite tables has been built and independently verified; the formal public homepage is not updated, no DOI has been minted, and author/institutional metadata remain pending.

Running title: NMD-VCell evidence resource

Correspondence: MAINTAINER_CONTACT_PENDING

Public URL: https://nmd-vcell.sema3c.chatgpt.site/resource/

Archive DOI: DOI_PENDING

Submission category: regular NAR Data Resources and Analyses

## Abstract

Perturbation-response predictions can appear useful because of weak baselines, restricted coverage or outcome-proxy inputs. NMD-VCell is a versioned evidence resource for a bounded Duchenne muscular dystrophy (DMD)-guided use case, linking heterogeneous priors, a 2,160-target HepG2 benchmark, audit outputs, external context and candidate actions. Across 20 five-fold realizations, a strongly shrunk 32-input ridge candidate achieved mean RMSE 0.117671 versus 0.117973 for the training-response mean and 0.123799 for zero change. Corrected two-stage intervals supported both RMSE differences, but response-module block intervals crossed zero for both cosine endpoints; median target benefit over the training mean was 0.000055 and 51.99% improved. No adjusted reliability association met the audit criterion, and historical candidate scores were component-dependent. A complete scan of the GSE293514 deposited matrix conserved all 1,294,825,793 coordinates and showed that column 62,711 contains 5,202 nonzero records plus a terminal zero; unresolved gene identity keeps biological transfer locked. External muscle and myoblast analyses are context filters, not perturbational validation. The resource supports transparent same-dataset evaluation and experiment triage, not muscle response, therapeutic efficacy or clinical utility.

Keywords: perturbation response; evidence provenance; unique-target cross-fitting; coverage-aware benchmarking; Duchenne muscular dystrophy; hypothesis triage

## Introduction

Single-cell perturbation experiments and predictive models create opportunities to organize experiments at genome scale [@perturbseq_2022; @gears_2023; @scgpt_2024]. Their evaluation is nevertheless vulnerable to design choices that are easy to hide in a final performance table. A model can improve error mainly by reproducing a common response, cover only a selected fraction of targets, reuse the same targets across computational splits, or incorporate variables derived from related outcomes. Recent benchmark work has reinforced the importance of strong linear baselines and explicit target coverage [@ahlmann_eltze_linear_2025].

Disease-guided applications introduce another problem: observational patient signatures, cell-line perturbations, tissue expression, genetic dependency and functional screens have different experimental units. DMD is a severe neuromuscular disorder with substantial unmet need [@dmd_primer_2021], but a DMD-associated expression contrast is not a perturbation response, a cancer-cell dependency is not muscle safety, and expression in skeletal muscle is not evidence that a perturbation has the predicted direction.

NMD-VCell was developed as an evidence resource rather than a therapeutic prediction system. It joins heterogeneous inputs while preserving their provenance, denominator, analytical role and prohibited interpretations. The v2.3 statistical audit repaired the primary shrinkage analysis, uncertainty calculation, feature terminology, reliability interpretation, candidate display and external myoblast-screen statistics. The present v2.4 transition adds a full public-identity and deposited-matrix structure audit for GSE293514. It preserves the separation between current analyses and historical model comparisons and does not promote technical reconstruction to biological validation.

The resource uses a project-defined five-level claim ladder. L1 denotes presence and provenance of a resource record; L2 denotes descriptive or same-dataset association; L3 denotes independent contextual concordance; L4 denotes direct perturbational validation in a relevant biological system; and L5 denotes therapeutic or clinical evidence. These labels are an NMD-VCell governance device, not a standard evidence-grading system. The present resource supports L1–L2 inspection and selected L3 context statements. It does not support L4–L5 claims.

## Materials and methods

### Evidence architecture and source roles

Four evidence classes were retained separately (Figure 1). First, DMD-prior tables contain source-specific observational effects and rule variants. Second, a processed HepG2 Perturb-seq object derived from GSE264667 supplies the perturbation benchmark [@nadig_trade_2025]. Third, external context tables include GTEx v8 skeletal-muscle expression [@gtex_2020], patient-muscle single-nucleus data associated with PRJNA772047 [@scripture_adams_muscle_2022], DMD skeletal-muscle organoids from GSE277637 [@kindler_organoid_2025], and a human-myoblast CRISPR screen from GSE293514 [@zhang_myoblast_fusion_2026]. Fourth, candidate rows expose source observations, cautions and next actions. Cross-layer joins do not convert one class into validation of another.

The processed HepG2 resource contains 2,393 perturbation labels plus one control label. A target was eligible for the primary benchmark when it met the frozen inclusion rules, producing 2,160 targets (Supplementary Figure S1). The response array contains 2,000 gene coordinates. The ordered response-gene list has SHA-256 `bdea16949964be5ee62c3ea067c3f69cd0c7465fc5863384b9960caba1f090b8`. The source H5AD, original delta store and the exact HVG-selection code, parameters and seed are absent. Consequently, selection independence of the 2,000-gene response panel cannot be verified. The v2.3 repair starts from this locked response array and does not claim to reconstruct its gene-selection history.

### DMD-prior metric audit

Four DMD source definitions were evaluated on their native standardized scales. Pairwise Pearson correlation, Spearman correlation and cosine similarity were independently recomputed from the source-specific gene table. Gene-wise spread was summarized only for genes available in at least three source definitions. Because sources share genes, systems and processing dependencies and lack a common set of independent study-level standard errors, they were not treated as mutually independent study estimates and no pooled meta-analytic effect was calculated.

### Response-independent input candidate

The primary predictor comprises 16 source values and 16 paired missingness indicators. The 16 values capture DMD-consensus summaries, source-availability fields and disease-priority annotations. The 32 frozen columns were regenerated from their declared source tables within a maximum absolute difference of 2.35 × 10−7. For every outer and inner training set, missing values were imputed using training-set statistics and continuous variables were standardized using training-set parameters.

The feature set is called `response_independent_v1_candidate` (Supplementary Figure S8). The word candidate is essential: timestamps and audited code paths show no use of the locked HepG2 response array in predictor construction, but they do not independently prove acquisition chronology. Manual disease-priority fields are absent for 2,157 of 2,160 eligible targets, their ten-gene registry is marked `to_verify`, and two substring-derived flags are subjective and can overlap. Historical DepMap/Chronos [@depmap_2017; @chronos_2021] and dataset-summary predictors were excluded from this primary set. NicheNet-derived provenance fields [@nichenet_2020] are inputs or context only and do not establish active signaling.

### Balanced unique-target cross-fitting and shrinkage

Twenty prespecified realizations were used. Each realization partitions all 2,160 targets into five non-overlapping outer folds of 432 targets, so every eligible target receives exactly one out-of-fold prediction per realization. The earlier target-reuse design is retained as a historical record (Supplementary Figure S2). Within each outer-training set, inner cross-validation selected ridge regularization from `10^-2` through `10^8`; an explicit training-response-mean predictor was evaluated alongside the finite grid. Preprocessing and regularization selection were repeated within the applicable training data.

Effective degrees of freedom were computed from the singular values of the standardized training design and selected ridge penalty. Prediction distance from the outer-training response mean was summarized as RMSE. A zero-change predictor and the fold-specific training-response mean were the two primary baselines. The model output is a target-by-response-gene delta matrix.

Historical repeated holdouts and external-model outputs were retained as versioned sensitivity objects. GEARS [@gears_2023], scGPT [@scgpt_2024], STATE [@state_2025] and Linear/PCA records differ in software version, target vocabulary, feature budget and coverage. Comparisons were restricted to declared target sets and endpoints. No pooled architecture rank was calculated.

### Metrics and corrected uncertainty

Primary endpoints were RMSE difference versus zero, RMSE difference versus the fold-specific training-response mean, raw cosine difference versus the training mean, and perturbation-specific residual cosine. Negative RMSE differences and positive cosine differences are favourable. Raw cosine was not designated as a headline success endpoint because common-response prediction can dominate it.

The corrected two-stage resampling procedure samples 20 fold realizations with replacement, samples targets with replacement inside every selected realization, computes a mean within each selected realization, and then averages the 20 means. The former procedure that sampled one realization per bootstrap replicate is retained only as a single-realization mixture sensitivity and is not described as an interval for the cross-realization mean.

Target dependence was assessed by ordering the 2,160 response profiles using average-linkage clustering of 25 principal components of cosine-normalized responses and dividing the ordering into 54 non-overlapping empirical modules of 40 targets. Bootstrap draws sampled whole modules and fold realizations. These are data-adaptive response modules, not curated pathways or independent biological replicates.

### Practical-gain and coordinate-corruption audits

For every target, benefit over the training-response mean was averaged across 20 realizations. We reported the median, interquartile range, fraction improved, and concentration of positive and net gain. A previously used global output-coordinate permutation was reconstructed. Inverting each saved permutation reproduced the primary ridge prediction within numerical precision. That operation is therefore treated as coordinate-label corruption, not an independent negative model, and is excluded from the main model-comparison axis.

### Reliability-association audit

Four reliability definitions were associated with four outcomes using unadjusted, sampling/scale-adjusted and full-adjusted linear models with HC3 standard errors. The measurement inventory and replicate definitions are reported in Supplementary Figure S3. The full sensitivity specification included log cell count, zero-baseline error, training-mean residual response norm, target-delta variance, a same-dataset target-expression proxy and missingness, batch count, and leave-HepG2-out DepMap median gene effect and missingness. P values were Holm-adjusted within the 16-test full-model family. A result required the prespecified direction, a 95% interval excluding zero and a Holm-adjusted P value below 0.05. Guide-pair reliability was limited to 133 targets. High variance inflation, including a maximum VIF of 113.4, requires descriptive rather than causal interpretation.

### Candidate-score dependency and action display

The historical 21-gene composite was reconstructed from its released formula and compared with its source values. Component dependence was assessed with Spearman correlations. Leave-one-block-out analyses removed each evidence block in turn, and nine deterministic DMD-source variants quantified rank sensitivity. These variants are sensitivity scenarios, not independent datasets; their percentile ranges are not confidence intervals.

The historical composite remains downloadable for version reproduction but no longer orders the reader-facing display. Candidates are ordered by a source-stored next-action rule and gene. Four actions are shown: verify expression, review risk/context, perform a muscle-relevant assay, or perform computational replication. Evidence axes, missing fields and cautions remain visible separately. Historical Priority A/B/C labels are preserved as released labels only and are not used as therapeutic grades.

### External applicability analyses

GTEx skeletal-muscle TPM thresholds were used only to indicate expression visibility. External patient-muscle and organoid/model concordance was evaluated with Pearson, Spearman and cosine metrics and bootstrap intervals. These observational contexts are not guide-resolved perturbation experiments.

For GSE293514, the published gene-level Supplementary Data 5 contains 7,197 rows and 250 hits meeting FDR <0.1 and P <0.0035 [@zhang_myoblast_fusion_2026]. The published `pos|lfc` field was labelled positive-selection MAGeCK LFC (killer-selected versus control); it is not a direct continuous measurement of fusion efficiency. Prespecified 2 × 2 tables were analysed using conditional exact odds ratios and two-sided Fisher exact tests. P values were Holm-adjusted across the four strata. No table contained a zero cell, so no continuity correction was required for the primary estimates.

### GSE293514 public-identity and deposited-matrix audit

The reconstruction preflight was separated from the published gene-level myoblast-screen analysis. Official Supplementary Data 2 contains 21,396 library rows. After collapsing duplicate oligo rows and applying identity-safe exclusions, the full public technical reference contained 20,492 unique unambiguous 20-nucleotide sequences: 19,592 target sequences and 900 non-targeting-control sequences. The frozen 1,592-sequence candidate and full public reference were compared against the same 520,596,853 guide reads. Exact-reference counts, observed sequences and observed genes were treated as technical identity metrics only. Because the author's exact `guidesnew2.csv`, context-column semantics and cell-keyed metadata are unavailable, neither reference was declared author-pipeline equivalent and no guide-resolved expression endpoint was opened.

The public split-pipe 1.1.2 source and the authors' stated Ensembl release-109 input were used to reconstruct a deterministic lexical `gene_id` order containing 62,710 rows. This was not silently attached to the 62,711-column deposited matrix. The 3,974,545,794-byte compressed matrix object was SHA-256 frozen after bounded reacquisition of corrupted byte windows and gzip EOF validation. A complete MatrixMarket coordinate scan compared declared and observed dimensions and record counts, summarized columns 62,710 and 62,711, and retained the terminal coordinate. Structural observations were permitted; gene identity, insertion position, expression phenotype and biological transfer were prespecified as blocked without the ordered author `all_genes.csv` or hash-frozen runtime `gene_info` (Figure 6c; Supplementary Figure S15).

### Resource implementation and integrity checks

The audited local resource candidate contains a no-login static browser, JSON records, a portable SQLite database, a data dictionary, download manifests and checksum-addressed release files (Figure 1; Supplementary Figure S12). Generator checks, cross-layer checks and upstream statistical checks are separate test families. A separate implementation means separate code, not an external laboratory or third-party organization. The isolated `/evidence-transition` route, matrix data card, gate table and downloads passed 3/3 route tests. The unified isolated v2.4r1 technical candidate synchronizes the corrected manuscript evidence, 113-table SQLite resource, versioned APIs and manifests and has passed generator and independent verification. It has not been deployed. DOI deposit, licensing, maintainer contact and final author/institutional metadata remain external release gates.

## Results

### Heterogeneous evidence is retained without converting context into validation

The resource connects perturbation, DMD-prior, external-context and candidate-action records while retaining their native units (Figure 1). The unified isolated database candidate contains 113 non-empty tables, 1,110 dictionary rows, 201 downloads and 632 release-manifest files, but it remains a non-public review candidate. The project-defined claim ladder makes the ceiling explicit: searchable provenance, descriptive association and bounded cross-context annotation are supported; direct muscle perturbation response, therapeutic efficacy and clinical utility are not.

Four DMD inputs differed in coverage and analytical role (Figure 2; Supplementary Figure S7). Independent recomputation confirmed that Pearson, Spearman and cosine were calculated from their declared definitions; the largest absolute Pearson–cosine difference across source pairs was 0.0054. A total of 13,559 genes were available in at least three source definitions, with substantial gene-wise spread and directional conflict. These results justify displaying source heterogeneity rather than collapsing four dependent inputs into a pooled disease effect.

### Expanded regularization removes the boundary artifact but reveals near-mean prediction

Across 100 outer folds, alpha 1,000 was selected 26 times and alpha 10,000 was selected 74 times; the upper finite-grid boundary was selected in 0/100 folds (Figure 3; Supplementary Figure S14). Ridge was preferred to the explicit training-mean-only comparator in inner cross-validation in all folds. The selected fits nevertheless had mean effective degrees of freedom 4.08 and mean prediction distance from the outer-training response mean of RMSE 0.00573. Thus, the former claim of uniform boundary selection at alpha 1,000 is withdrawn, while the substantive interpretation remains one of strong shrinkage close to the common response.

Across 20 realizations, mean RMSE was 0.117671 for the candidate, 0.117973 for the training-response mean and 0.123799 for zero change. Aggregate relative improvements were 0.256% and 4.95%, respectively. Corrected two-stage mean intervals were −0.006366 to −0.005885 for RMSE difference versus zero and −0.000370 to −0.000222 versus the training mean. The corresponding raw-cosine difference was 0.000819 (0.000619 to 0.001026), and residual cosine was 0.029681 (0.025826 to 0.033568).

Dependence-preserving response-module resampling narrowed the supported claim. Module-block intervals remained below zero for RMSE differences versus zero (−0.011896 to −0.000282) and the training mean (−0.000549 to −0.000028), but crossed zero for raw-cosine difference (−0.000031 to 0.001749) and residual cosine (−0.009053 to 0.070318). The corrected result supports a small same-dataset mean RMSE improvement, not robust directional prediction.

Practical gain was heterogeneous (Figure 3e). Median target benefit over the training mean was 0.000055 (interquartile range −0.001053 to 0.001175); 51.99% of targets improved and 48.01% worsened. The top decile accounted for 63.84% of gross positive gain and 247.63% of net gain, showing that losses outside the leading decile offset much of the aggregate benefit. The coordinate-permutation reconstruction further showed that the prior negative-control output was a label-corruption transformation of the primary prediction, not an independent model.

### Adjusted reliability does not support the previous moderation claim

Unadjusted reliability patterns were sensitive to outcome and specification. Under the full HC3 adjustment, 0/16 reliability-by-outcome tests met the prespecified support criterion (Figure 3f; Supplementary Figure S13). This included the composite weight, split-half cosine, batch-split cosine and the 133-target guide-pair subset. The previous reliability-moderation interpretation is therefore withdrawn. Because several covariates were highly collinear, these analyses are reported as sensitivity adjustments and not causal estimates.

### Historical comparator results remain coverage- and version-specific

Historical model records differed by more than tenfold in evaluated-target coverage (Figure 4; Supplementary Figure S4; Supplementary Figure S5; Supplementary Figure S6). In the current complete-denominator audit, scGPT produced eligible outputs for 232 of 2,160 targets and GEARS for 2,143. Historical five-split summaries reused targets and were not biological replicates. Same-target contrasts remain useful as descriptive, version-specific checks, but unequal target vocabularies, feature budgets and implementations prevent a controlled architecture leaderboard. Ridge is shown as a reference, not as a universal winner.

### Candidate display now separates action from a dependent historical score

The historical composite was reproduced to a maximum absolute error of 3.55 × 10−15, but its components were not independent (maximum absolute off-diagonal component Spearman correlation 0.378). Removing the observed-counteralignment block produced a maximum rank shift of 10, and 11 genes moved by at least three positions (Figure 5; Supplementary Figure S9). The leading historical rows therefore do not justify a calibrated target rank.

Reader-facing output now assigns four genes to expression verification, seven to risk/context review, three to muscle-relevant assay and seven to computational replication. DMD-prior counteralignment, source agreement, source-variant sensitivity, skeletal-muscle expression and external cautions are displayed as distinct fields. The four historical A, eight B and nine C labels remain version-reproduction metadata. None denotes treatment probability, efficacy, safety or validation.

### External analyses provide context and risk checks, not muscle perturbation validation

GTEx skeletal-muscle expression was matched for 2,361 of 2,393 perturbation targets; 2,253, 2,133 and 1,805 exceeded 0.1, 1 and 5 TPM, respectively (Figure 6a; Supplementary Figure S10). Among 21 candidates, the corresponding counts were 20, 16 and eight. These data can flag weak expression context, but cannot establish perturbation direction.

In the myoblast screen, 9/273 genes in the unstratified queue and 173/1,302 comparison genes were published hits. The conditional exact odds ratio was 0.223 (95% CI 0.099–0.440; Holm-adjusted P=8.56 × 10−7). Among non-broad-essential genes, the conditional exact odds ratio was 0.586 (95% CI 0.214–1.548; Holm-adjusted P=0.796). The unstratified depletion did not remain statistically supported after essentiality restriction. The screen is therefore a human-myoblast risk/context stress test, not evidence that queue membership improves fusion.

Cross-context expression concordance was weak and metric-dependent (Figure 6d; Supplementary Figure S11). Patient-muscle Pearson concordance was 0.012 (95% bootstrap CI −0.010 to 0.036; Holm-adjusted P=0.194). Organoid/model Pearson concordance was 0.043 (95% CI 0.008 to 0.076; Holm-adjusted P=0.006), whereas rank-based direction was opposite. The perturbation object directly covered 8/64 neuromuscular target labels, 6/64 at the operational 20-cell threshold and 0/5 DMD/BMD targets. This is an applicability boundary, not external perturbational validation.

### Public reconstruction narrows but does not close the GSE293514 transfer gate

Across all 520,596,853 guide reads, the 20,492-sequence full public reference added 183,832 exact matches relative to the frozen 1,592-sequence candidate, equal to 0.03531 percentage points of audited reads. It added 948 observed sequences and 878 observed target genes, while every candidate-sequence count and the non-targeting-control total were conserved. This small technical increment did not establish equivalence to the missing author guide table and did not justify opening a second biological cell-assignment analysis.

The full deposited-matrix scan conserved 1,294,825,793 of 1,294,825,793 declared coordinates in an object with 2,782,263 rows and 62,711 columns (Figure 6c; Supplementary Figure S15). Column 62,711 contained 5,203 records: 5,202 nonzero records with value sum 5,265 plus the terminal zero coordinate `2782263 62711 0`. The first last-column record occurred at row 10. This falsifies the working hypothesis that the final column is a pure zero-valued split-pipe cap. It does not identify the unmatched feature: the public-input reconstruction has 62,710 ordered genes, and an unknown earlier insertion could shift all later positions. Nine structural gates and 14/14 independent checks passed, but exact author-runtime equivalence, exact gene-to-column attachment and biological-transfer release remained blocked.

### The resource is queryable, while user performance and release completion remain unclaimed

The browser and downloads are designed to support gene-level evidence inspection, coverage-aware model comparison, candidate filtering and identification of unsupported claim levels. No completed independent-user study is available, so task success, usability and interpretation accuracy are not claimed. Automated verification passed 39/39 generator, 149/149 cross-layer and 120/120 upstream checks, while only 15/19 release gates are currently satisfied (Supplementary Figure S12). The isolated evidence-transition route additionally passed 3/3 route tests, and its matrix structure layer passed 9/12 gates with 14/14 independent checks. The three blocked structure gates concern missing identity artifacts and biological-transfer eligibility; the remaining release gates concern external release state and author/institutional decisions rather than hidden statistical success.

## Discussion

NMD-VCell's contribution is an evidence architecture that keeps denominators, baselines, provenance, uncertainty and prohibited interpretations attached to reusable records. The v2.3 repair changes the central computational conclusion. Expanding the alpha grid removes an artificial boundary result, but the selected ridge fits still have approximately four effective degrees of freedom and stay close to the fold-specific response mean. Corrected uncertainty supports a mean RMSE reduction of only 0.256% over that strong comparator, and nearly half of targets worsen. Dependence-preserving intervals do not establish either directional endpoint.

This conclusion is narrower and more useful than a generic model-performance claim. It shows that a statistically stable aggregate error difference can coexist with negligible median target benefit, concentrated gains and weak directional evidence. It also separates two questions: improvement over zero captures common-response prediction, whereas improvement over the training mean measures additional target-specific information. Only the latter is an appropriate estimate of the incremental signal highlighted here.

The audit also clarifies what provenance can and cannot establish. The 32 predictor columns are reproducible from declared, response-independent source tables and are processed within training folds. They cannot be described as unconditionally outcome-independent, because acquisition chronology and subjective registry fields remain imperfect. More importantly, the 2,000-gene response panel cannot be reconstructed from the available source H5AD and selection log. That unresolved gap prevents a complete end-to-end independence claim even though the audited predictor code path does not consume held-out responses.

Reliability and candidate ranking required similar restraint. The former descriptive reliability pattern disappears under a full, highly collinear adjustment family and therefore cannot support stratified validation. The historical candidate composite is exactly reproducible but mixes dependent axes and changes materially when a dominant block is removed. Exposing next actions and individual evidence fields is more faithful to the data than presenting a single therapeutic-looking order.

External muscle data define where the resource may fail. GTEx expression, patient muscle, organoids and a myoblast functional screen add relevant context, but none supplies the missing guide-resolved perturbation transcriptome in an independently replicated DMD-relevant muscle system. Exact analysis further shows that the apparent fusion-screen association is sensitive to essentiality restriction. The correct role of these sources is exclusion, caution and experimental design.

The GSE293514 transition audit makes the distinction between structural and semantic reproducibility concrete. A complete coordinate scan can prove matrix dimensions, record conservation, last-column occupancy and a writer-compatible terminal cap. It cannot name a column when the ordered author feature table is absent. Treating the one-column discrepancy as an ignorable placeholder would have converted a falsified implementation hypothesis into a biological annotation. The retained block therefore protects downstream transfer analyses from an unquantified position shift rather than treating technical completeness as biological validation.

The project-defined L1–L5 ladder turns these limitations into machine-readable governance. The current resource may be useful for finding evidence, checking coverage and selecting the next experiment, but it is not a virtual clinical model. Advancing to L4 requires prospectively frozen, guide-resolved perturbation data in a relevant human muscle system with independent biological replication. L5 additionally requires appropriately designed therapeutic and clinical evidence.

## Data availability and maintenance

The public URL continues to serve the earlier locked homepage and is not the authority for the v2.3 statistical repair or v2.4 matrix-structure transition. The unified isolated v2.4r1 technical candidate now contains 113 non-empty SQLite tables, 1,110 dictionary rows, 201 downloads and 632 resource-manifest rows and has passed generator and independent verification. It remains undeployed. The isolated v2.4 figure review package contains 21 figures in PNG, PDF, SVG, EPS and 600-dpi TIFF formats, source-data manifests and validation records. The matrix body was scanned remotely and was not copied into the local review package. These are author-review artefacts, not a DOI-frozen public release.

Before public release and submission, the verified technical candidate must receive author-approved license, maintainer contact, author, affiliation, funding and conflict metadata; a repository/archive DOI must be minted; the same frozen bytes must be deployed; and normal and private-browser checks must be completed. The final submitted snapshot should be immutable and accompanied by checksums. Later scientific changes should receive a new version rather than overwrite the submitted snapshot.

## Limitations

The primary benchmark is internal to one processed HepG2 Perturb-seq dataset. Balanced cross-fitting prevents outer-fold target overlap within each realization, but 20 realizations remain algorithmic reassignments of the same 2,160 targets. Target deltas are estimated from cells, and guide-pair reliability is available for only 133 targets. The response-module bootstrap preserves empirical similarity but its modules are data-adaptive, not curated pathways or biological replicates.

The original response-gene selection is not reproducible from the files currently present. Comparator coverage and feature budgets differ. DMD source effects share genes and processing dependencies. External muscle contexts are observational or use different functional endpoints. The GSE293514 public-input reconstruction yields 62,710 ordered candidate genes for a 62,711-column matrix; the missing author `all_genes.csv` or runtime `gene_info` prevents exact attachment. No expression phenotype was opened from that matrix. Candidate actions are retrospective rules over incomplete, partially dependent evidence. The resource has no completed independent-user evaluation, prospective lockbox, muscle perturbation validation, therapeutic study or clinical assessment.

Release limitations are also material. The unified v2.4r1 technical candidate is built and verified but remains isolated; the formal public site has not been updated, and the DOI, license, maintainer contact and final author metadata are unresolved. These items retain overall submission status as not ready despite completion of the revised figures and documents.

## Ethics declaration

This computational resource recruited no new participants and generated no new human-subject data. Public datasets were reused under their source-study access terms. Ethical approvals, consent procedures and data-use restrictions belong to the original studies and require final author verification. NMD-VCell is not intended for diagnosis, prognosis, treatment selection or clinical decision-making.

## Author contributions

[To complete using the CRediT taxonomy after the final author list is fixed.]

## Acknowledgements

[To complete.]

## Funding

[To complete.]

## Conflict of interest

[To complete.]

## AI use disclosure

Generative AI tools were used for drafting, code organization, audit-table assembly and review-package generation. All computational outputs in this candidate were subjected to scripted integrity checks. Final scientific claims, provenance, citations and wording remain blocked pending independent verification and approval by the named authors; this candidate must not be submitted before that attestation is completed.

## References

Reference metadata are maintained in `references_verified.bib`, with claim contexts and verification status in `22_reference_verification_matrix.tsv`. DOI, URL, accession and policy fields require final author verification before submission.
