Model capability
See what NMD-VCell can estimate today—and what evidence unlocks the next level. Capabilities are organized by biological transfer, not model reputation. Open the Research view when you need exact ModelRuns, baselines and stopping decisions.
Executed evidence 6 frozen runs
Model decisions One bounded baseline · negative controls retained
Next value transition DMD truth for calibration
Question answered What actually ran on a frozen task, against which baseline, with which split, metric, failure reason and permitted use?
Scope & evidence unit Frozen DMD/assay benchmark runs and model audits are available. FSHD, DM1 and SMA do not currently have calibrated NMD-VCell model runs. One ModelRun bound to a task, data revision, split, baseline, seed, metrics and gate decision.
Produces Model cards, a six-run execution ledger, benchmark contracts, negative results and explicit readiness gates.
Cannot establish Published external performance is not inherited; no calibrated DMD or multi-disease predictor is released.
Next handoff Carry only a frozen, gate-passing run into prospective registration and outcome comparison. Continue →
Model ability map
What can this model layer answer now—and what unlocks the next answer? The registry now behaves like a scientific decision map: every model layer is tied to answerable questions, non-answerable questions, the dataset it needs and the outcome that would make the next claim meaningful.
Open capability map JSON →
Layer 1 · CAP-MODEL-SAME-ASSAY-RESPONSE Available as a bounded technical comparator
Same-assay response baseline
Can answer now Whether the released HepG2 CRISPRi ridge baseline improves mean response error within its own processed assay context.
→ Next unlock Task-bound response vectors with the same frozen feature space, split rule and leakage controls.
Can answer now Whether the released HepG2 CRISPRi ridge baseline improves mean response error within its own processed assay context. Which permanent controls a future perturbation model must beat before it earns a stronger local claim.
Cannot answer alone DMD muscle response, pathway reversal, patient response or treatment simulation. Directionally reliable biology outside the same processed HepG2 task.
Needs dataset A compatible same-assay holdout for continued technical benchmarking. A harmonized external cell-level perturbation outcome before any transfer claim.
Needs outcome Task-bound response vectors with the same frozen feature space, split rule and leakage controls. For DMD relevance, matched disease-context perturbation outcomes must be registered separately.
Next decision Retain as the permanent comparator every new architecture must clear.
Layer 2 · CAP-MODEL-EXTERNAL-TRANSFER-COMPARATOR Executed evidence shows where complex models did not advance
External transfer and advanced comparators
Can answer now Which frozen advanced-comparator runs failed, stopped or remained negative against the released baselines.
→ Next unlock Prospective transfer outcomes returned under the same evaluator and permanent baseline policy.
Can answer now Which frozen advanced-comparator runs failed, stopped or remained negative against the released baselines. Which provenance, seed and feature-coverage gaps prevent a model-family reputation from becoming local evidence.
Cannot answer alone General superiority of GEARS, scGPT, TxPert, MORPH or any watched external model inside NMD-VCell. A valid DMD or muscle-context transfer claim from aggregate HepG2-only evidence.
Needs dataset Harmonized cell-level perturbation outcomes with exact gene-feature alignment and reusable adapters. External holdouts with preregistered task, split, baseline stack and failure-preserving ModelRun receipts.
Needs outcome Prospective transfer outcomes returned under the same evaluator and permanent baseline policy. Feature-coverage, seed, code, container and data-digest receipts that make reruns auditable.
Next decision Use these runs as a value-producing negative-control library, then bind any new comparator to a fresh frozen task.
Layer 3 · CAP-MODEL-DMD-PERTURBATION-RESPONSE Ready as an evaluation contract; awaiting disease-context truth
DMD perturbation response
Can answer now Which exact evidence object is missing before NMD-VCell can train or release a disease-conditioned response model.
→ Next unlock Replicated molecular plus fusion or viability endpoint returned through Study → Prediction → Outcome objects.
Can answer now Which exact evidence object is missing before NMD-VCell can train or release a disease-conditioned response model. How a candidate should move from Evidence Card to Study Card to registered Prediction and returned Outcome.
Cannot answer alone Candidate-conditioned DMD molecular response, myogenic functional effect, toxicity or calibrated uncertainty. Gene-to-drug translation in DMD without matched genetic and chemical perturbation screens.
Needs dataset Matched DMD and control myogenic perturbation dataset with donor, state, time, perturbation modality and target-engagement fields. For the gene–chemical bridge, paired genetic and compound screens in the same relevant muscle-state space.
Needs outcome Replicated molecular plus fusion or viability endpoint returned through Study → Prediction → Outcome objects. Donor- and context-disjoint holdouts with calibration, abstention and toxicity-miss audits.
Next decision Convert the strongest candidate handoffs into a small registered DMD outcome pilot instead of emitting a premature prediction.
Layer 4 · CAP-MODEL-PATIENT-FUNCTIONAL-GENERALIZATION Research roadmap with governance requirements
Patient and functional generalization
Can answer now What patient-linked validation, privacy and prospective evaluation would require before trajectory modeling becomes meaningful.
→ Next unlock Prospective functional outcomes that can test calibration, safety, subgroup performance and clinical utility.
Can answer now What patient-linked validation, privacy and prospective evaluation would require before trajectory modeling becomes meaningful. Which current artifacts can prepare the path: dataset registry, outcome pilot, distribution contract and ModelRun ledger.
Cannot answer alone Patient-specific progression, individual treatment response, clinical decision support or therapeutic utility. Functional recovery claims without longitudinal patient-linked outcomes and independent clinical governance.
Needs dataset Longitudinal patient-linked cell-state, intervention and phenotype datasets with auditable consent and privacy boundaries. Site-, donor- and time-disjoint cohorts connected to molecular and functional readouts.
Needs outcome Prospective functional outcomes that can test calibration, safety, subgroup performance and clinical utility. A governance-approved endpoint definition before any patient-facing interpretation is exposed.
Next decision Keep this lane as the north-star validation program while near-term work focuses on DMD perturbation truth.
Capability ladder
Scientific capability first; model details second. Each step states the strongest current use and the evidence needed next.
Open truth-generation study → Level 1 · Same-assay response Available as a bounded baseline HepG2 aggregate mean response only
Level 2 · External-assay transfer Tested; transfer not supported Useful negative benchmark retained
Level 3 · Muscle-context transfer Needs matched perturbation truth Truth-generation study is the next action
Level 4 · DMD perturbation response Activates after calibration Requires qualifying DMD outcomes
Level 5 · Patient-level generalization Not established Requires donor-disjoint validation
Level 6 · Functional outcome Needs prospective functional truth Molecular and phenotype endpoints must return together
Product contract · MODEL-CARD-1.0
Model Card: What did this exact version receive, return, beat or fail? Version + provider Task input → output Training context Split + leakage Baselines + metrics Run receipt + failure
ModelRun evidence ledger. Completed negative and blocked runs remain first-class evidence. GEARS, scGPT, TxPert and MORPH are preserved with their frozen task, baseline comparison, gate decision and provenance gaps; none is promoted into a DMD prediction claim. Download the run ledger .
Shared reading grammar
Evidence metadata fields—never one confidence score. Every state uses text plus a shape or border. Missing evidence and a equivalent-null result are different states.
Origin ● Observed△ Computational□ Missing
Context relation D Target contextM Same tissueL Locus correction↗ External cell
Unit structure n Donor / linec Culturep Pathway? Not estimable
Lifecycle ◇ Draft◆ Frozen▣ Registered● Released
Outcome / absence ↑ Supportive0 Measured null? Inconclusive□ Not measured
Scientific model view
Gate matrix: baseline, threshold, uncertainty and decision stay together. A completed run is evidence of execution, not proof of disease validity. Each row keeps the comparator and advancement rule beside the result.
Download run ledger →
Executed locally 6 frozen ModelRuns One limited same-assay baseline; all advanced comparators failed or stopped.
External reference only 6 watched model families Published architectures and source claims are not inherited as local performance.
Future architecture 3 locked model contracts No DMD transition, gene–chemical bridge or patient trajectory model has qualifying truth.
Baseline-to-threshold view Bars are task-specific and must not be compared across metrics. Arrows state whether lower or higher is better.
Ridge 0.117507
Train mean 0.117973
Zero 0.123799
Small mean-error gain; direction unsupported.
TxPert · cosine ↑ × FailedModel 0.346196
Threshold 0.373307
One seed completed; confirmatory seeds remain locked.
Model 0.014162
Train mean 0.010558
Zero 0.014293
16-condition validation failed; the test partition stayed sealed.
GEARS × × × × × 5/5 splits failed advancementscGPT × × × × × 5/5 splits failed advancement
□ Not measuredDMD candidate response remains outside every executed run. This is a missing outcome set, not a measured null response.
Decisive local results
Complex models did not earn advancement on the released frozen tasks. These values are deliberately comparator-first. They summarize the gate, not a universal model rank.
6 completed · 5 stopped or failed · 1 limited baseline
6 completed runs
5 no-advance decisions
5 + 5 GEARS and scGPT frozen seeds
Awaiting truth calibrated DMD outputs
Ridge 0.117507 RMSE vs train mean 0.117973
Limited same-context baseline GEARS No split advanced vs ridge won all 5 same-coverage comparisons
Fail · negative control scGPT No split advanced vs higher RMSE than ridge in all 5 splits
Fail · restricted coverage TxPert 0.346196 cosine vs train mean 0.373307
Fail · confirmatory seeds locked MORPH +34.13% MSE vs worse than train mean; won 2/16 conditions
No-go · test partition sealed
Frozen execution history
ModelRun ledger: positive, negative and stopped runs stay visible.
Run Role Task Decision state Rule Baseline result
MRUN-RIDGE-SAFE-2.3-G0-REPEATED-FOLDSame-context ridge residual baseline
PERMANENT BASELINE
G0
RETAINED LIMITED BASELINE
LIMITED_PASS_SAME_CONTEXT_ONLY
Small mean-RMSE improvement over train mean; raw response direction remains unsupported.
MRUN-TRANSFER-DIAGNOSTIC-1.0-G1External perturbation transfer diagnostic
TRANSFER DIAGNOSTIC
G1
UNSUPPORTED TRANSFER
FAIL
The current response substrate did not transfer directionally to the external aggregate.
MRUN-GEARS-0.1.2-FIVE-SEED-20260713GEARS
OFFICIAL NEGATIVE CONTROL
G0
OFFICIAL NEGATIVE CONTROL
FAIL
Failed paired RMSE and raw directional support in all five splits; ridge won every direct same-coverage comparison after Holm correction.
MRUN-SCGPT-0.2.5-FIVE-SEED-20260714scGPT
OFFICIAL NEGATIVE CONTROL
G0
OFFICIAL NEGATIVE CONTROL
FAIL
Failed paired RMSE and raw directional support in all five splits; RMSE was higher than same-coverage ridge in every split.
MRUN-TXPERT-CONFIG-GAT-SEED-20260712TxPert public-STRING config-gat
PREREGISTERED ADVANCED COMPARATOR
G0
ADVANCEMENT GATE FAILED
FAIL
Improved RMSE and retrieval over simple controls but missed the preregistered median-delta-cosine gate against train mean.
MRUN-MORPH-DEPMAP25Q3-VALIDATION-20260722MORPH DepMap-25Q3 validation pilot
VALIDATION ONLY NEGATIVE CONTROL
G0-VALIDATION-PILOT
SCIENTIFIC NO GO TEST SEALED
NO_GO
Reduced MSE by 0.92% versus zero but was 34.13% worse than train mean; beat train mean on only 2 of 16 validation conditions.
MRUN-RIDGE-SAFE-2.3-G0-REPEATED-FOLDLIMITED_PASS_SAME_CONTEXT_ONLY
Same-context ridge residual baseline
Role PERMANENT BASELINE
Task G0
Decision RETAINED LIMITED BASELINE
Small mean-RMSE improvement over train mean; raw response direction remains unsupported.
Open ModelRun release →
MRUN-TRANSFER-DIAGNOSTIC-1.0-G1FAIL
External perturbation transfer diagnostic
Role TRANSFER DIAGNOSTIC
Task G1
Decision UNSUPPORTED TRANSFER
The current response substrate did not transfer directionally to the external aggregate.
Open ModelRun release →
MRUN-GEARS-0.1.2-FIVE-SEED-20260713FAIL
GEARS
Role OFFICIAL NEGATIVE CONTROL
Task G0
Decision OFFICIAL NEGATIVE CONTROL
Failed paired RMSE and raw directional support in all five splits; ridge won every direct same-coverage comparison after Holm correction.
Open ModelRun release →
MRUN-SCGPT-0.2.5-FIVE-SEED-20260714FAIL
scGPT
Role OFFICIAL NEGATIVE CONTROL
Task G0
Decision OFFICIAL NEGATIVE CONTROL
Failed paired RMSE and raw directional support in all five splits; RMSE was higher than same-coverage ridge in every split.
Open ModelRun release →
MRUN-TXPERT-CONFIG-GAT-SEED-20260712FAIL
TxPert public-STRING config-gat
Role PREREGISTERED ADVANCED COMPARATOR
Task G0
Decision ADVANCEMENT GATE FAILED
Improved RMSE and retrieval over simple controls but missed the preregistered median-delta-cosine gate against train mean.
Open ModelRun release →
MRUN-MORPH-DEPMAP25Q3-VALIDATION-20260722NO_GO
MORPH DepMap-25Q3 validation pilot
Role VALIDATION ONLY NEGATIVE CONTROL
Task G0-VALIDATION-PILOT
Decision SCIENTIFIC NO GO TEST SEALED
Reduced MSE by 0.92% versus zero but was 34.13% worse than train mean; beat train mean on only 2 of 16 validation conditions.
Open ModelRun release →
Migration rule: missing commit, container or data digests remain explicit null provenance fields. A missing field does not erase a completed run, and a completed run does not imply DMD validity.
NMDVCELL-RIDGE-SAFE-2.3 Executed · limited
Same-context perturbation baseline
Input Training folds of processed HepG2 perturbation-response deltas
→ Output Held-out 2,000-feature mean response vector
Training context One processed HepG2 CRISPRi assay context
Held-out task Repeated target-level balanced folds · G0
Permanent baselines zero change · training-response mean · ridge residual
Primary evaluation RMSE · raw cosine · residual cosine
Current result Small average-error improvement; response direction is not reliable. Same-assay technical baseline and benchmark control only
What unlocks the next claim Beat simple baselines on harmonized external cell-level perturbation outcome before any transport claim.
NMDVCELL-TRANSFER-DIAGNOSTIC-1.0 Executed · unsupported
External perturbation transfer diagnostic
Input Frozen 55-target external aggregate diagnostic
→ Output Directional-transfer assessment
Training context HepG2 response substrate
Held-out task External aggregate target set · G1
Permanent baselines zero change · training-response mean · ridge
Primary evaluation aggregate directional-agreement diagnostic
Current result The current model did not transfer directionally. Method diagnostic; not a cell-population or DMD benchmark
What unlocks the next claim Add harmonized cell-level outcome, exact feature alignment and a preregistered external estimator.
NMDVCELL-DMD-TRANSITION-FUTURE Not trained
Disease-conditioned state-transition model
Input Required matched DMD/control myogenic perturbations with donor, state, time and function
→ Output Desired molecular response, state transition, functional effect, toxicity and calibrated uncertainty
Training context No qualifying disease-relevant perturbation training set
Held-out task Planned unseen donor · state · laboratory · disease line · G2–G6
Permanent baselines zero change · train mean · ridge · nearest-neighbour transfer
Primary evaluation DES · PDS · MAE · calibration · AUPRC · hit rate · replication · abstention
Current result No calibrated DMD state-transition prediction is emitted. Architecture and evaluation contract only
What unlocks the next claim Return independent DMD perturbation outcomes through frozen Study, Prediction and Outcome objects.
NMDVCELL-GENE-CHEMICAL-BRIDGE-FUTURE Not trained
Gene–chemical bridge model
Input Required matched genetic and chemical screens in relevant muscle or DMD states
→ Output Desired cross-modality response, mechanism concordance and uncertainty
Training context No matched DMD genetic–chemical screen
Held-out task Planned unseen compound · gene · donor · disease state
Permanent baselines nearest-neighbour · pathway mean · additive transfer
Primary evaluation retrieval · response similarity · calibration · prospective hit rate
Current result No drug-response or gene-to-compound translation is emitted. Future design contract only
What unlocks the next claim Import traceable dose, time, target-engagement and phenotype truth for both modalities.
NMDVCELL-PATIENT-TRAJECTORY-FUTURE Not available
Patient trajectory model
Input Required longitudinal patient-linked cell state, intervention and outcome data
→ Output Desired patient-specific trajectory and intervention response with uncertainty
Training context No patient-linked longitudinal perturbation trajectory
Held-out task Required prospective patient and site holdout
Permanent baselines natural-history and population-level reference models
Primary evaluation prospective calibration · safety · subgroup performance · clinical utility
Current result No patient-specific trajectory or treatment simulation is permitted. Research roadmap only; no clinical decision support
What unlocks the next claim Requires longitudinal data, prospective evaluation, privacy safeguards and clinical governance.
Global model watch · primary sources
Published performance is not inherited; local runs are linked separately. Sources checked 30 Jul 2026
Arc Virtual Cell Initiative STATE v1 State embedding plus context-conditioned state transition
Reported scope 167M observational and more than 100M perturbational cells across 70 human contexts
NMD-VCell state External architecture reference · not executed in this release Open primary source → Nature Methods scGPT Generative pretraining for single-cell tasks including perturbation response
Reported scope More than 33M cells
NMD-VCell state Local five-seed benchmark completed · official negative control · no DMD claim Open primary source → Nature Geneformer Context-aware attention model for network biology
Reported scope Approximately 30M single-cell transcriptomes
NMD-VCell state External backbone candidate · no NMD-VCell benchmark run Open primary source → arXiv preprint · March 2026 SCALE Conditional population transport for perturbation prediction
Reported scope Preprint reports evaluation on Tahoe-100M
NMD-VCell state External preprint watch · not independently reproduced here Open primary source → arXiv preprint · April 2026 PRiMeFlow End-to-end flow matching in gene-expression space for genetic and chemical perturbations
Reported scope Preprint reports distribution-level evaluation and the method behind a 2025 VCC Generalist Prize entry
NMD-VCell state External preprint watch · author-reported performance only · not reproduced here Open primary source → arXiv preprint · March 2026 Lingshu-Cell Masked discrete diffusion for perturbation-conditioned single-cell population generation
Reported scope Preprint reports whole-transcriptome population generation and evaluation across perturbation datasets
NMD-VCell state External preprint watch · author-reported performance only · not reproduced here Open primary source →
Evaluation rule: foundation-model embeddings and complex architectures must beat zero, train-mean, ridge and nearest-neighbour baselines on the same prospective holdout. External publications do not unlock a local disease claim.
Next output contract
Move from one average vector to a governed cell-population prediction. A future model must emit generated cells, state proportions, pseudobulk effects, uncertainty and an abstention state. NMD-VCell now defines that contract without pretending the locked model already exists.
Population-prediction status Not available yet No DMD cell-level perturbation outcome
Inspect contract and metric firewall
One result · three reading levels
Choose the explanation that matches your task.
Researcher Biologist Public
Researcher Ridge, transfer diagnostics, GEARS, scGPT, TxPert and MORPH now have explicit run records. Their negative or stopped gates are retained, while disease-response and patient-level estimands still have no qualifying DMD perturbation outcome.
Biologist Several models were genuinely tested, and the simple baseline often remained stronger. None has shown that it can predict what happens after perturbation in DMD muscle.
Public The project keeps unsuccessful model tests instead of hiding them. Those tests improve the research process, but they are not a disease prediction.