When Models Lie
Why Promising Preclinical Results Fail in Humans
Drug development is built on models.
Cell lines help us test mechanism. Animal studies provide an initial view of efficacy, pharmacokinetics, and safety. Computational models generate hypotheses and refine candidate selection. Each of these tools plays an essential role in reducing uncertainty before a compound reaches human trials.
Yet despite decades of progress, the vast majority of drug candidates that appear promising in preclinical development never become approved therapies. Depending on the therapeutic area, fewer than 10 percent of compounds entering clinical testing ultimately reach the market. And for areas such as oncology, success rates are often substantially lower. The principal reason is relatively straightforward: biological systems that approximate human disease do not fully replicate human biology. Or, to put it another way-
Models are extraordinarily useful. They are also simplifications.
When those simplifications fail to capture the relevant aspects of target biology, tissue context, or species-specific toxicity, the resulting data can be highly convincing and profoundly misleading.

Why Good Models
Still Fail
A model does not need to be “wrong” to be misleading. It only needs to omit a critical biological variable. Common sources of translational failure include:
Species-Specific Biology
Targets may differ in expression patterns, receptor structure, downstream signaling, or tissue distribution between preclinical species used for efficacy and safety studies, and in humans.
Incomplete Disease Representation
Many disease models capture selected features of pathology rather than the full cellular complexity, microenvironmental barriers, and heterogeneity observed in actual patients.
Hidden Off-Target Effects
Unexpected binding in human tissues may not be evident in animal or cell-based studies, particularly when target expression density differs across species.
Differences in Tissue Context
Cells behave differently depending on tissue architecture, extracellular matrix interactions, and the local disease microenvironment.
Inadequate Biomarker Strategy
A compound may engage with the intended target, but only in a specific, biologically defined subset of patients or disease states.
These limitations do not diminish the value of modelswhen used wisely. However, they underscore the critical importance of stress-testing and validating key preclinical assumptions in the most relevant human biological context available: intact human disease tissue.
Case Studies in Translational Failure
Here we have collected seven of the most instructive examples of translational failure in clinical and preclinical drug development. These case studies explore what happens when preclinical models fail to predict human outcomes and demonstrate how human tissue based validation could have provided the decision-grade safety and efficacy intelligence needed to avert these outcomes.



Part I:
Safety Failures
What happens when the Model Cleared a Compound That Should Not Have Been Used in a Human
TGN1412

Safe in Monkeys, Catastrophic in Humans
TGN1412 was a CD28 superagonist monoclonal antibody intended to selectively expand regulatory T cells for autoimmune disease and leukemia. Preclinical toxicology studies in cynomolgus monkeys showed no major safety concerns, even at doses approximately 500 times higher than those administered to human volunteers.
However, in the 2006 first-in-human study, all six healthy volunteers developed severe, life-threatening cytokine release syndrome (CRS) and multi-organ failure within hours of receiving a sub-microgram dose.
What the Model Missed
Cynomolgus macaques lack CD28 expression on their CD4+ effector memory T cells (CD4em). In humans, this cell population expresses CD28 at high levels and was the primary driver of the pro-inflammatory cytokine storm.
Human Tissue Validation Opportunity
Map CD28 expression across human lymphoid tissue compartments (spleen, lymph nodes, tonsil). Perform comparative species IHC to detect the CD4em target expression mismatch.


Fialuridine

Human-Specific Mitochondrial Toxicity
Fialuridine (FIAU) was developed as an oral nucleoside analogue for the treatment of chronic hepatitis B. Extensive preclinical toxicology studies conducted in mice, rats, dogs, and primates showed no signs of liver toxicity. Yet during a Phase II clinical trial, 7 of 15 participants developed acute liver failure, lactic acidosis, and pancreatitis. Five patients died, and two required emergency liver transplantation.
What the Model Missed
Human hepatocytes express the nucleoside transporter hENT1 on both the plasma membrane and the inner mitochondrial membrane. This dual localization is human-specific; animal hepatocytes lack mitochondrial hENT1. FIAU accumulated in human mitochondria, inhibiting polymerase gamma (pol-γ) and depleting mitochondrial DNA (mtDNA), causing cellular collapse.
Human Tissue Validation Opportunity
Characterization of hENT1 expression and subcellular localization in human vs. safety species liver. Quantification of mitochondrial morphological health in human liver tissue sections under ex vivo drug challenge.
Metrifonate
![1_AdobeStock_662350748-[Converted]_edite](https://static.wixstatic.com/media/6a8cb2_e667e590f6e3411f964833cffc992a57~mv2.png/v1/fill/w_84,h_84,al_c,q_85,usm_0.66_1.00_0.01,enc_avif,quality_auto/1_AdobeStock_662350748-%5BConverted%5D_edite.png)
Unpredicted Neuromuscular Toxicity
Originally used as an antiparasitic, metrifonate was repurposed and advanced into clinical trials as an acetylcholinesterase (AChE) inhibitor for Alzheimer’s disease. Preclinical models demonstrated strong efficacy signals and an acceptable tolerability profile, suggesting it could safely improve cognitive function. During clinical trials, however, human participants experienced severe respiratory paralysis due to profound neuromuscular toxicity.
What the Model Missed
Animal models failed to capture the delicate threshold of human neuromuscular junction sensitivity. The mismatch between the dose required for central nervous system efficacy and the dose that caused peripheral toxicity was entirely masked in preclinical testing.
Human Tissue Validation Opportunity
-
Validate drug target distribution and AChE expression density at the human neuromuscular junction.
-
Spatially map functional impact and toxicity thresholds in human peripheral tissues.

Thalidomide
A Tragic Lesson in Species Differences

Introduced in the late 1950s, thalidomide was marketed as a highly effective, non-barbiturate sedative and was heavily prescribed to pregnant women to alleviate morning sickness. Standard preclinical toxicology screening at the time—conducted primarily in rodents—revealed no significant teratogenic effects. The compound appeared so benign that it was sold over-the-counter in several countries.
The clinical reality of this drug resulted in one of the darkest chapters in medical history. An estimated 20,000 to 30,000 infants were born with severe congenital abnormalities, most notably phocomelia (malformation of the limbs), alongside widespread rates of miscarriage.
Subsequent investigations revealed that mice and rats metabolize thalidomide differently than humans. Rodent models did not generate the specific teratogenic metabolites responsible for the developmental toxicity. Furthermore, the embryonic pathways disrupted by the drug are highly species-dependent. Only later, when tested in specific breeds of rabbits and non-human primates, could researchers replicate the teratogenicity observed in humans.
What the Model Missed
Rodent metabolism and embryonic development pathways did not reflect the specific enzymatic breakdown and biological pathways active during human fetal development. The model demonstrated safety for mice, not humans.
Human Tissue Validation Opportunity
While developmental toxicology presents unique ethical and technical challenges for direct human tissue testing, the underlying principle remains paramount. Assessing context-specific target and pathway activity in human embryonic stem cell systems or highly specialized human in vitro models is necessary when species concordance cannot be assumed.
Part II:
Translation in Efficacy
When Models Produce Convincing but Ultimately Useless Efficacy Data
The Mouse Didn't Lie...
we just asked it the wrong questions
While severe toxicity causes spectacular and highly public clinical failures, lack of efficacy is the silent killer of the pharmaceutical industry. A drug that is perfectly safe but fundamentally ineffective is just an expensive placebo - one that costs hundreds of millions of dollars to discover. The root of this problem lies in the mirage of the preclinical model.
In an engineered animal model, a drug might easily reach its target, bind perfectly, and reverse the modeled disease state. However, proving efficacy in a mouse does not confirm Target Reality in a human.
In actual human patients, target expression is often highly heterogeneous, the tissue microenvironment presents unforeseen physical and biochemical barriers to engagement, or the targeted pathway simply doesn't drive the disease the way it did in the model.
Without validating true target engagement in pathologically relevant human tissue before entering the clinic, sponsors are essentially gambling that human biology will mimic a mouse.
_edited.jpg)
Tarenflurbil
![1_AdobeStock_662350748-[Converted]_edite](https://static.wixstatic.com/media/6a8cb2_e667e590f6e3411f964833cffc992a57~mv2.png/v1/fill/w_84,h_84,al_c,q_85,usm_0.66_1.00_0.01,enc_avif,quality_auto/1_AdobeStock_662350748-%5BConverted%5D_edite.png)
When Transgenic Mice Overstate Efficacy
Tarenflurbil was developed as a selective amyloid-beta 42 lowering agent. In transgenic mouse models engineered to overproduce human amyloid, the drug performed exactly as intended, reducing amyloid burden and improving cognitive metrics. Despite these highly convincing preclinical efficacy signals, the drug completely failed to demonstrate clinical benefit in a massive Phase III trial, leading to the abandonment of the program.
What the Model Missed
Transgenic mice simulate specific pathological features of Alzheimer's, but they do not replicate the complex, decades-long pathophysiology of the human disease. More critically, the preclinical models did not accurately predict whether the drug was achieving sufficient target engagement within the human brain to alter the disease course.
Human Tissue Validation Opportunity
-
Confirm drug-target binding directly in diseased human neural tissue.
-
Utilize spatial validation to ensure the mechanism of action aligns with true human pathology.
.png)

.png)

Phenserine
![1_AdobeStock_662350748-[Converted]_edite](https://static.wixstatic.com/media/6a8cb2_2e5265e784e3472ab3a96b13a283667b~mv2.png/v1/fill/w_84,h_84,al_c,q_85,usm_0.66_1.00_0.01,enc_avif,quality_auto/1_AdobeStock_662350748-%5BConverted%5D_edite.png)
When Early Signals Fade
Phenserine was designed to inhibit both acetylcholinesterase and the production of amyloid precursor protein. Following robust efficacy in animal models, it showed positive early signals in initial human testing. However, in Phase III clinical trials, it failed to separate from placebo with statistical significance.
What the Model Missed
The failure highlighted the gap between basic mechanistic activity in a model and meaningful clinical efficacy in a heterogeneous human population. The models lacked the complexity required to stratify which patients possessed the specific tissue-level pathology most likely to respond to the intervention.
Human Tissue Validation Opportunity
Establish human tissue-based proof of mechanism prior to large-scale trials. Develop biomarker-driven patient stratification strategies based on human tissue profiling.
The Pattern Behind the Failures
Bridging the Translational Gap with Tissue Insights
Whether the focus is immunology, neurology, or cardiovascular medicine, decades of clinical attrition reveal that failure is rarely random. Most failures cluster in familiar patterns around several missing base validations.
Human tissue-based studies do not replace standard animal models or formal toxicology; they complement them by stress-testing those assumptions. A practical translational framework from preclinical promise to clinical reality relies on moving beyond approximations to definitively answer four questions before clinical exposure.
At Offspring Biosciences, we have built our Tissue Insights™ Platform to address these exact decision points:
Module 1: Target Validation
Right Target
Is the target causally relevant and accessible in human disease tissue?
Module 2: Antibody Selection & Optimization
Does the candidate engage the target in the complex tissue environment?
Module 3: Efficacy & Proof of Mechanism
Does the drug physically bind and trigger the intended biology?
Module 4: Preclinical Safety
Right Drug
Pillar 2: Binding
Right Target / Right Tissue Pillars 1-3
Right Safety
Does the candidate exhibit on or off-target binding that may indicate possible toxicity liabilities in human organs?
Module 5: Clinical Biomarkers & Patient Stratification
Can we define a specific responder population for Phase II Enrichment?
Right Patient
SOCA Enrichment
Each of these modules reduces a distinct source of translational risk, replacing hope with validated data and helping sponsors significantly reduce uncertainty before committing to expensive, late-stage clinical trials.
For a deeper dive into how Offspring Biosciences addresses these critical hurdles and a look into how to implement these steps into your pipeline, see our comprehensive white paper, “Operationalizing Translational Success”
Moving Beyond Approximations
Preclinical models remain indispensable. They help researchers narrow hypotheses, optimize chemistry, identify safety concerns, and build confidence in a development strategy.
But every model is an approximation.
The most costly failures in drug development often occur when promising model data are interpreted as definitive proof rather than evidence supporting a hypothesis. Across all the case studies explored above, failure was not random. It consistently clustered around the missing human tissue validations shown on the previous page.
The question for translational teams is no longer whether a mouse model worked. The question is whether the underlying assumptions have been stress-tested in the specific human tissue where the therapy is intended to act. By anchoring translational workflows in human tissue validation early in development, sponsors can definitively answer that question—ensuring that when a program advances, it does so based on human reality, not a model's simplification.
That distinction can determine whether a program advances with confidence, or becomes another example of When Models Lie.


