Translational validity
A translation score is not a verdict on whether a study is good. It asks how directly the design supports an inference about humans. A connectome may be excellent basic science while scoring modestly as a treatment model; mutation-specific editing can score higher when the causal nucleotide, target tissue, and endpoint are concrete. This rubric weighs conserved mechanism, model fidelity, robustness across strains and laboratories, intervention feasibility, and convergent human evidence.
The literature does not support one universal mouse-to-human success rate. A 2024 systematic review covering 122 reviews, 54 diseases, and 367 interventions estimated that 50% progressed to any human study, 40% to a randomized trial, and 5% to regulatory approval; it also reported 86% concordance between positive animal and human-study results. These numbers coexist because approval is a late, multi-causal endpoint shaped by manufacturing, commercial decisions, trial design, safety, and efficacy.
Translation varies by domain. Monogenic disorders with an exact variant and measurable target engagement often bridge more cleanly than heterogeneous psychiatric or neurodegenerative syndromes. Conditioned freezing cannot reproduce autobiographical trauma, forced-swim immobility cannot instantiate depression, and Alzheimer’s transgenic mice express selected pathologies on compressed timescales. The honest claim is often “this tests a mechanism,” not “this models the disease.”
Standard strains, pathogen-free housing, one sex, and one age improve control while shrinking the biology sampled. C57, DBA, and BTBR differences can reshape conclusions; microbiota and immune experience matter too. Multi-site replication, heterogenized cohorts, preregistration, blinding, null-result reporting, and triangulation with human genetics, biomarkers, organoids, cohorts, and trials build a stronger bridge.
Model fidelity also depends on endpoint and exposure. A drug dose that produces target engagement in a young inbred mouse may not reproduce pharmacokinetics in an older patient taking several medicines. Behavioral readouts can be especially vulnerable to handling, lighting, time of day, experimenter sex, and prior housing. Conversely, some mouse discoveries translate as biological principles even when the original intervention never does: immune checkpoints, gene targeting, and spatial coding became influential because multiple kinds of evidence converged around them. The dossier therefore scores the claim attached to the listed experiment, not the eventual importance of the entire field.
Scores should be read comparatively and locally. Studies #37 and #43 sit near the top because cross-species spatial coding and correction of a monogenic mutation have unusually strong bridges. Universe 25 sits at the bottom because its most famous human analogy is unsupported, not because the enclosure observations are unreal. Recent aging interventions score cautiously because lifespan, frailty, epigenetic clocks, and tissue function are different endpoints. No entry receives a nine or ten: even the best animal model remains a model, and direct human evidence always carries its own limitations.
For readers, the practical rule is simple: translate mechanisms before metaphors, compare doses and endpoints explicitly, and demand human convergence before turning an animal effect into health advice.
| 50% | progressed to any human study |
|---|---|
| 40% | progressed to an RCT |
| 5% | obtained regulatory approval |
| 86% | positive animal–human concordance reported |
2024 PLOS Biology meta-analysis ↗ NIH rigor working group ↗ Wildling mice overview ↗