Research dossier · edition 01

Keep the wonder.
Show the receipts.

A field-by-field audit of fifty rodent stories: what the paper actually did, what remains uncertain, and how far each result can travel toward humans.

00 / Editorial contract

Uncertainty is part of the record.

This is a living reference database, not a finished bibliography. Every supplied claim received an identity check. A verified paper is labeled peer-reviewed; preprints and research-program summaries stay distinct; unresolved claims remain visible and marked unverified.

The 1–10 translation score measures the bridge to humans, not scientific quality. Scores are editorial judgments with a visible rationale.

Read the scoring rubric ↓

01—50 / Source audit

The evidence ledger

A—D / Across the fifty

The claims behind the claims

A

Translational validity

A translation score is not a verdict on whether a study is good. It asks how directly the design supports an inference about humans. A connectome may be excellent basic science while scoring modestly as a treatment model; mutation-specific editing can score higher when the causal nucleotide, target tissue, and endpoint are concrete. This rubric weighs conserved mechanism, model fidelity, robustness across strains and laboratories, intervention feasibility, and convergent human evidence.

The literature does not support one universal mouse-to-human success rate. A 2024 systematic review covering 122 reviews, 54 diseases, and 367 interventions estimated that 50% progressed to any human study, 40% to a randomized trial, and 5% to regulatory approval; it also reported 86% concordance between positive animal and human-study results. These numbers coexist because approval is a late, multi-causal endpoint shaped by manufacturing, commercial decisions, trial design, safety, and efficacy.

Translation varies by domain. Monogenic disorders with an exact variant and measurable target engagement often bridge more cleanly than heterogeneous psychiatric or neurodegenerative syndromes. Conditioned freezing cannot reproduce autobiographical trauma, forced-swim immobility cannot instantiate depression, and Alzheimer’s transgenic mice express selected pathologies on compressed timescales. The honest claim is often “this tests a mechanism,” not “this models the disease.”

Standard strains, pathogen-free housing, one sex, and one age improve control while shrinking the biology sampled. C57, DBA, and BTBR differences can reshape conclusions; microbiota and immune experience matter too. Multi-site replication, heterogenized cohorts, preregistration, blinding, null-result reporting, and triangulation with human genetics, biomarkers, organoids, cohorts, and trials build a stronger bridge.

Model fidelity also depends on endpoint and exposure. A drug dose that produces target engagement in a young inbred mouse may not reproduce pharmacokinetics in an older patient taking several medicines. Behavioral readouts can be especially vulnerable to handling, lighting, time of day, experimenter sex, and prior housing. Conversely, some mouse discoveries translate as biological principles even when the original intervention never does: immune checkpoints, gene targeting, and spatial coding became influential because multiple kinds of evidence converged around them. The dossier therefore scores the claim attached to the listed experiment, not the eventual importance of the entire field.

Scores should be read comparatively and locally. Studies #37 and #43 sit near the top because cross-species spatial coding and correction of a monogenic mutation have unusually strong bridges. Universe 25 sits at the bottom because its most famous human analogy is unsupported, not because the enclosure observations are unreal. Recent aging interventions score cautiously because lifespan, frailty, epigenetic clocks, and tissue function are different endpoints. No entry receives a nine or ten: even the best animal model remains a model, and direct human evidence always carries its own limitations.

For readers, the practical rule is simple: translate mechanisms before metaphors, compare doses and endpoints explicitly, and demand human convergence before turning an animal effect into health advice.

50%progressed to any human study
40%progressed to an RCT
5%obtained regulatory approval
86%positive animal–human concordance reported

2024 PLOS Biology meta-analysis ↗ NIH rigor working group ↗ Wildling mice overview ↗

B

Ethics of animal research

Scientific importance does not erase animal experience. This collection includes anesthesia, fear conditioning, isolation, social defeat, forced swimming, injury, surgery, neurodegeneration, genetic disease, and terminal tissue analysis. Ethical cost belongs beside methods and findings. Readers should know whether distress was required, whether humane endpoints were used, whether a non-animal system could answer the question, and whether information gained was proportionate to harm.

The governing framework is the 3Rs: replace animals with non-animal or less sentient approaches where possible; reduce numbers while retaining valid power; and refine housing, handling, procedures, and endpoints. Reduction cannot mean underpowered experiments that waste animals. Refinement includes social and environmental needs, not only analgesia. US research commonly receives Institutional Animal Care and Use Committee review, but compliance is a floor rather than proof that a design is necessary.

The forced-swim test shows why interpretation and welfare are inseparable. If immobility is labeled “despair,” the procedure appears to model a profound disorder. If it is an acute coping response that happens to be drug-sensitive, both the claim and ethical justification change. Similar caution applies to social defeat, pain-directed helping, and Universe 25: vivid human words can obscure what was measured.

Replacement is advancing where human-derived biology can answer species-specific questions. Organoids, organs-on-chips, advanced cultures, computational toxicology, and carefully governed human data can complement or replace some tests. In April 2025 the FDA announced a roadmap to reduce, refine, or potentially replace animal testing for selected products with new-approach methods, including AI and organoid toxicity tests. Ethical scrutiny is not anti-science; it asks whether knowledge is reliable and important enough to justify its cost.

Good ethical reporting also improves reproducibility. ARRIVE-style reporting of randomization, blinding, exclusions, sample-size rationale, housing, enrichment, sex, age, strain, anesthesia, analgesia, and euthanasia lets readers evaluate both welfare and bias. A study cannot be meaningfully replicated when basic animal conditions are missing. Pain or stress may also become an uncontrolled experimental variable, changing immunity, metabolism, learning, and social behavior. Refinement can therefore improve both the life of the animal and the validity of the result.

Historical studies require context without moral insulation. Standards, oversight, and available alternatives differed, but a retrospective account can still describe harms plainly. Contemporary frontier work deserves equally close scrutiny: elegant wireless devices may reduce tethering while still requiring viral injection; prenatal editing may promise early prevention while raising consent and germline questions; organoids may reduce animal use while creating new governance issues of their own. The ethical section should evolve with the science, documenting replacements that become validated and retiring animal procedures when they no longer add necessary information.

Transparency also lets readers disagree intelligently: the same evidence can support different moral judgments, but hidden procedures, missing harms, and euphemistic language prevent any serious judgment at all.

3Rsreplace · reduce · refine
2025FDA roadmap announced
n < 10small-sample flag threshold in this dossier

NC3Rs: the 3Rs ↗ FDA roadmap ↗ NIH IACUC tutorial ↗

C

Failures, reinterpretations & replication

Replicated is rarely binary. A laboratory may reproduce an assay without confirming its interpretation; another may confirm a mechanism in a different task without repeating the protocol. Forced-swim immobility is reproducible while its label as behavioral despair is disputed. Place and grid coding are robust across laboratories and species. A 2025 circuit paper can be peer reviewed yet preliminary because other teams have not had time to test it.

Universe 25 illustrates narrative drift. Calhoun documented a real colony trajectory in an engineered, non-dispersing enclosure. The popular conclusion that density inevitably causes human collapse is not a replication—it is unsupported extrapolation. Rat Park is also a series of rat studies, not the polished internet story of two bottles and one decisive result. Enrichment and isolation affect drug taking, but taste, dose, withdrawal, and housing complicate the exact replication history.

Every 2025–2026 entry here is flagged preliminary unless independent replication was found. #3 may duplicate the revival paper. #36’s supplied numbers did not match a 2025 Allen simulation; its current source is a 2019 antecedent. #42’s exact Cas3 safety and 80% reduction claim remains unverified. #49’s group-versus-alone percentages could not be traced. Keeping these slots visible is more honest than silently substituting vaguely related papers.

Small groups, single strains, sex-restricted cohorts, flexible endpoints, and publication bias inflate certainty. Preregistration, blinded allocation, justified sample sizes, shared data, reported exclusions, multi-lab reproduction, and biological heterogeneity are practical responses. A dramatic unresolved claim belongs on a candidate list, not in finished book copy.

Four labels are useful in practice. Direct replication repeats the core protocol and endpoint. Conceptual replication tests the same proposition through a different design. Robustness asks whether the effect survives reasonable changes in strain, sex, age, facility, analyst, or apparatus. Generalization asks whether it appears in another species or human dataset. These are not interchangeable. A large open dataset such as MICrONS invites computational reuse, while a destructive anatomical reconstruction cannot simply be rerun like a behavioral assay. A historical research program such as place cells accumulates confidence through many converging methods rather than one exact repetition.

Absence of a located replication is not evidence of failure, especially for papers published months ago. The interface therefore says “not located in this audit” and dates the snapshot. Equally, a failed replication does not automatically erase an effect: power, protocol drift, biological heterogeneity, and selective original reporting all need examination. The book’s source notes should record null and contradictory evidence with the same prominence as supporting studies, and revisions should retain a public change log. That prevents a corrected chapter from making yesterday’s certainty disappear without explanation.

That discipline protects wonder rather than diminishing it: a finding that survives precise restatement, adversarial testing, and time is ultimately more interesting than a dramatic claim sustained by repetition alone.

4entries currently marked unverified
2025–26preliminary absent independent replication
n=5small groups in parts of study #40

Universe 25 original ↗ Forced-swim commentary ↗ NIH rigor recommendations ↗

D

What the original fifty leave out

The list is strongest as a contemporary tour of social neuroscience, aging, regeneration, and intervention technology. It is not yet a defensible ranking of the fifty most important rodent studies. Several slots divide one program: revival work occupies #1 and possibly #3; REM work #12–13; blood and parabiosis several aging entries; Acomys #31–34. Other slots are methods, clusters, or unverified claims. “Fifty stories” is more accurate than an objective historical ranking.

Major candidates are missing. The ob/ob mouse and leptin transformed energy-balance biology. The first transgenic and knockout mice changed causal genetics, and OncoMouse made engineered cancer models scientifically and legally consequential. Circadian mutants and clock genes connected behavior to molecular feedback loops. These advances likely changed medicine more broadly than several unresolved 2025 claims.

Cancer and immunology also need more space. Syngeneic and engineered tumors contributed to immune-checkpoint concepts, CAR-T strategies, and combination therapies. Bone-marrow transplantation, monoclonal antibodies, and histocompatibility research relied on mice. Infectious disease, vaccines, metabolism, and development are underrepresented. The neuroscience-heavy balance is an editorial voice, not neutral history.

Species labels matter: learned helplessness began with dogs and rats; place-cell and brain-interface landmarks were rats; Meaney’s maternal-care epigenetics and much of Panksepp’s play work were rats. Revise transparently: define whether each unit is a paper, program, or chapter; require a primary record; score influence, novelty, consequence, replication, ethics, and narrative distinctness; publish replacements. Immediate candidates for #3, #36, #42, and #49 are leptin, knockout mice, circadian genes, and a cancer-immunotherapy landmark.

A replacement shortlist should also protect narrative diversity. Leptin represents physiology and molecular discovery; knockout mice represent a platform that changed experimental causality; clock genes connect molecular feedback to whole-animal behavior; cancer immunotherapy illustrates both clinical impact and translational limits. Other candidates include coat-color transplantation and histocompatibility, monoclonal antibodies, the nude mouse in oncology, embryonic stem cells, Cre–lox conditional genetics, germ-free models, and humanized immune mice. Each offers a different way that a mouse changed what scientists could ask, not merely another surprising phenotype.

Editorially, the cleanest architecture may use fifty chapters but allow each chapter to contain more than one paper. A “discovery program” entry could place the foundational study, decisive replication, human evidence, and later correction on one timeline. That would solve several current duplications without discarding recent work. It would also make room for an explicit “what they got wrong” element inside every chapter rather than isolating failure at the end. The final table should expose why each candidate earned a slot—historical influence, clinical consequence, conceptual novelty, reproducibility, ethical significance, or visual narrative power—so readers can distinguish importance from recency.

This approach would make the final selection arguable in the best sense: readers could see the values behind it, inspect the evidence, and propose better substitutions without rebuilding the project from scratch.

4immediate unresolved-slot replacements
4 slotsAcomys currently receives #31–34
1994leptin/ob gene landmark
1989gene-targeted knockout era

Nobel: gene targeting ↗ Rockefeller: leptin ↗ Nobel: circadian mechanisms ↗

Scoring / living method

How the bridge is graded