NMD-VCell Research Workbench Module: Data & API · evidence-to-experiment workflow NMD = neuromuscular disorders
v1.0 candidate Frozen 25 Jul 2026 DOI pending
v1.0.0-database-resource Schema 1.1 Model ridge-safe-v2.3 Benchmark repeated-fold-v2.2 Build EA-20260729-15 v1.0 is a database and evidence-governance release; it is not a validated disease-prediction or clinical decision-support release.
Current evidence ceiling L2 observed HepG2 limited L3a context No independent DMD perturbation validation Open boundary

Reconstruction-oriented methods

Data, transformations, validation and inference

Claim boundary. The frozen model evaluates held-out target identities within the processed HepG2 dataset. It is not a DMD-muscle predictor.
Gene/API objects · PASS Study Card contracts · PASS Exact response-panel rebuild · BLOCKED Source-balanced prior · NOT COMPUTED Machine-readable status

Data sources and roles

SourceSpecies/contextRoleCurrent boundary
HepG2 TRADE perturbation matrixHuman HepG2; CRISPR perturbation2,160-target response benchmarkThe public release contains processed aggregates and derived response objects. A primary public accession, source H5AD, original delta store and exact HVG-selection log are not available in the frozen public layer; end-to-end response-panel selection cannot therefore be independently reconstructed.
GSE277637Human DMD and control muscle single-cellDMD context rowsObservational disease contrast, not perturbation validation
GSE270868Human DMD and control muscle single-cellDMD context rowsObservational disease contrast, not perturbation validation
SEMA3C source-state analysesHuman DMD muscle contextsDMD DE and regulatory-context rowsHeterogeneous source; no independent source-level standard errors
GSE293514Human myoblast fusion screenL3a external-context screeningNot the same perturbation replication
GTEx v8 skeletal muscleHuman tissue expressionExpression feasibility contextExpression does not establish function

Exact source files, hashes, inclusion tables and provenance paths are in the download ledger. Where a public accession or licence is unavailable, the release preserves that gap instead of inventing metadata.

Response construction

  1. Use the already frozen 2,000-coordinate response array; the public package cannot verify where or when those HVGs were selected because the source H5AD and selection log are absent.
  2. Aggregate cells to one target-level perturbation response; cells are observational support and the perturbation target is the inferential unit.
  3. Eligibility is exactly 2,160 non-targeting-excluded targets with at least 20 cells, fixed before fold assignment. Fewer than 20 cells is the frozen low-support exclusion threshold for this benchmark.
  4. Partition every realization into five non-overlapping 432-target outer folds, balanced by outcome-blind cell-count quintile and external-feature availability.
  5. Fit imputation, standardization and α selection on the applicable training data only.
  6. Compare every out-of-fold prediction with zero change and the outer-training response mean.

Feature construction

The safe model uses 16 DMD/context or curated priority values and 16 paired missingness indicators. “Response-independent” refers to the audited build path; acquisition chronology is not independently proven. DepMap and dataset-derived response summaries are excluded.

Historical-name note. Frozen feature identifiers that contain older terms such as rescue or priority are retained byte-for-byte for reproducibility. Their current semantics are counteralignment or historical input provenance; they are not therapeutic-rescue labels or current ranks.

List all 32 frozen feature columns
FeatureSource groupLeakage classification
n_evidence_rowsdmd_consensusexternal_or_curated_no_target_response
n_evidence_rows_missingdmd_consensusexternal_or_curated_no_target_response
n_significant_rowsdmd_consensusexternal_or_curated_no_target_response
n_significant_rows_missingdmd_consensusexternal_or_curated_no_target_response
n_contextsdmd_consensusexternal_or_curated_no_target_response
n_contexts_missingdmd_consensusexternal_or_curated_no_target_response
n_sourcesdmd_consensusexternal_or_curated_no_target_response
n_sources_missingdmd_consensusexternal_or_curated_no_target_response
median_effect_log2fcdmd_consensusexternal_or_curated_no_target_response
median_effect_log2fc_missingdmd_consensusexternal_or_curated_no_target_response
mean_effect_log2fcdmd_consensusexternal_or_curated_no_target_response
mean_effect_log2fc_missingdmd_consensusexternal_or_curated_no_target_response
max_abs_effect_log2fcdmd_consensusexternal_or_curated_no_target_response
max_abs_effect_log2fc_missingdmd_consensusexternal_or_curated_no_target_response
sign_consistencydmd_consensusexternal_or_curated_no_target_response
sign_consistency_missingdmd_consensusexternal_or_curated_no_target_response
consensus_signed_scoredmd_consensusexternal_or_curated_no_target_response
consensus_signed_score_missingdmd_consensusexternal_or_curated_no_target_response
mean_abs_signed_scoredmd_consensusexternal_or_curated_no_target_response
mean_abs_signed_score_missingdmd_consensusexternal_or_curated_no_target_response
dmd_direction_updmd_consensusexternal_or_curated_no_target_response
dmd_direction_up_missingdmd_consensusexternal_or_curated_no_target_response
dmd_direction_downdmd_consensusexternal_or_curated_no_target_response
dmd_direction_down_missingdmd_consensusexternal_or_curated_no_target_response
is_disease_priority_genepriority_v1external_or_curated_no_target_response
is_disease_priority_gene_missingpriority_v1external_or_curated_no_target_response
disease_priority_inversepriority_v1external_or_curated_no_target_response
disease_priority_inverse_missingpriority_v1external_or_curated_no_target_response
fit_highpriority_v1external_or_curated_no_target_response
fit_high_missingpriority_v1external_or_curated_no_target_response
fit_mediumpriority_v1external_or_curated_no_target_response
fit_medium_missingpriority_v1external_or_curated_no_target_response

Full provenance remains available in the feature table and leakage ledger.

Cross-validation and regularisation

Estimands and uncertainty

RMSE is calculated for each target over its 2,000-dimensional response vector. Raw cosine compares the predicted and observed target vectors. Residual cosine first subtracts the applicable outer-training mean response from both vectors. Primary contrasts are paired target-level differences versus the declared baseline.

The joint hierarchical bootstrap resamples realizations first and target modules second (10,000 replicates; seed 20260737). The public layer does not currently expose the target-module membership/count table required to rerun that exact second stage; the reported interval is auditable but not independently rebuildable from zero until that object is released.

The strict 16-outcome family requires paired RMSE support, multiplicity control and raw-direction support. All 16 rows—including model, split, endpoint, estimate, interval, adjusted P when defined and gate—are exposed on the Benchmark page and in typed JSON. The legacy 16/16 result is historical only.

Software and exact reconstruction

Commands, code paths, source hashes, environment records and expected outputs are distributed through the two release layers. Start with the release manifest, frozen protocol, model card and schema 1.1 dictionary. Package versions not present in a locked source object are explicitly not asserted here.