Skip to content
Eonix
Menu

How to read a longevity study

Six questions that separate a result worth acting on from a headline, and the published evidence for why each one matters.

Published 10 min read

In short

Most longevity claims rest on one of three things: a study in mice, a measurement that stands in for living longer, or a single result nobody has repeated.

Six questions settle what a study can support: what was measured, in whom, whether assignment was random, how large the effect was, whether it has been replicated, and whether it has been peer reviewed.

None of these require statistical training. They require finding the answers in the paper, which is usually where they are, and which is usually what coverage of the paper leaves out.

Six sieve screens standing in a row, each with a finer mesh than the one before; a dense cloud of small particles enters the first and only three particles continue past the sixth.
Six questions, each a finer filter than the last. Most claims are held at the first.Illustration

A longevity claim reaches most people after passing through several hands. A research group publishes a paper, a press office writes a release, an outlet writes a headline, and a supplement company writes a product page. Each step shortens the sentence and strengthens it. By the end, a compound that changed a blood marker in forty people over three months has become a compound that slows aging.

The original paper usually states its own limits plainly. What follows is a way to find them.

What was actually measured?

Start here, because this question disposes of more overstated claims than the other five together.

Very little longevity research measures how long anyone lived. Doing so would take a human lifetime, and no funder or participant will wait. So studies measure something faster and treat it as a stand-in: a marker in blood, a score for inflammation, an estimate of biological age, a measure of grip strength. This substitute is called a surrogate endpoint, and the US Food and Drug Administration defines it as a trial endpoint used in place of a direct measure of how a patient feels, functions or survives.

The regulator also grades these measurements, and the grade is worth knowing because it is a statement about evidence rather than about biology. A validated surrogate endpoint is backed by a clear mechanistic rationale and by clinical data giving strong evidence that moving it predicts a specific benefit. A reasonably likely surrogate endpoint has the rationale but not enough clinical data to show that it predicts anything at all.

Nearly every marker in longevity research sits in the second category. That is not a scandal, and it is not a reason to dismiss the field. It is the reason a sentence like “improved markers of aging” and a sentence like “extended life” are different sentences, and why a study supporting the first does not support the second.

The strictest measure available is all-cause mortality, which counts deaths from any cause whatever. It is strict precisely because it cannot be improved by shifting deaths from one column to another. It is also rare in this field, because detecting a change in it requires very many people followed for very many years.

Was it people, or was it mice?

The second question is who or what the study was done in, and the honest answer is often an animal.

Animal work is not a lesser kind of science. It is where nearly every human treatment begins. It is also where most candidate treatments end, and the size of that attrition is documented. A 2010 review in PLOS Medicine by van der Worp and colleagues, examining how animal findings fare afterwards, reported that about one-third of studies translated at the level of human randomised trials, and about one-tenth of the interventions were subsequently approved for use in patients.

The same review gives a concrete illustration from stroke research: roughly 500 neuroprotective strategies showed promise in animals, and of these only aspirin and thrombolysis proved effective in human trials.

The reasons are specific rather than mysterious, and they are worth knowing because they apply directly to aging research. Laboratory animals are typically young and healthy, while the patients a treatment is intended for are old and have several conditions at once. Doses and timing that are straightforward in a cage are often impossible in a clinic. And the outcome measured in the animal is frequently not the outcome that matters to a person: the review notes studies measuring the size of a lesion in animals where the human trials measured how well patients could function.

A mouse result is a reason to run a human trial. It is not a small version of one.

Was assignment random, or was it observed?

The third question separates two kinds of study that are routinely reported in the same language.

In a randomized controlled trial, chance decides who receives the treatment. This matters because people who choose a treatment differ from people who do not, in ways nobody can fully enumerate. Someone who takes a longevity supplement is, on average, also someone who exercises, sleeps, earns and sees doctors differently. An observational study can adjust for the differences it measured. It cannot adjust for the ones it did not.

Two safeguards usually accompany randomization and are worth looking for by name. Allocation concealment means the person enrolling a participant could not know which group came next, so the sickest could not be quietly steered elsewhere. Blinded outcome assessment means whoever judged the result did not know which group they were judging.

These are worth looking for because they are often missing. The PLOS Medicine review found that only about one-third of animal studies reported randomization, and fewer still reported allocation concealment or blinded outcome assessment. It also found that sample size calculations were reported in 0 to 3 percent of studies — meaning that in nearly all of them, nobody stated in advance how many animals would be needed to detect the effect they were looking for.

How large was the effect?

The fourth question is about size, and it is where percentages do their damage.

Results are usually reported as a hazard ratio or a relative reduction. A hazard ratio of 0.8 means the event occurred 20 percent less often in the treated group. It does not say how often the event occurred at all. Twenty percent less of something that happens to two people in a thousand each year is one person in a thousand. Both statements are true; only one of them sounds like a breakthrough.

There is a subtler problem underneath. A single hazard ratio compresses an entire follow-up period into one number, and the standard method for producing it assumes the gap between groups stayed roughly constant throughout. When that assumption fails, the statistical test can miss a real effect.

This is not a theoretical worry. In 2024, researchers reanalysed two decades of mouse lifespan data from the National Institute on Aging’s Interventions Testing Program. The log-rank test, which had produced the programme’s published results, assumes a uniformly reduced risk of death across the lifespan and is relatively insensitive to any compound whose effect is not uniform. Running the same data through the Gehan test, which is more sensitive to differences earlier in life, identified five additional compounds, among them metformin and enalapril.

No new mice were involved. The animals, the doses and the deaths were identical. What changed was which question the statistics were asked.

Has anyone repeated it?

The fifth question is replication, and in this field there is an unusually clear example of what it looks like when it is taken seriously.

The Interventions Testing Program was set up by the National Institute on Aging in 2004 to test compounds proposed by the wider research community. Each candidate is tested in parallel at three independent sites — the Jackson Laboratory, the University of Michigan, and the University of Texas Health Science Center at San Antonio — in male and female mice of a genetically heterogeneous stock called UM-HET3, using identical standardised protocols. The genetic mixture is deliberate: it greatly reduces the chance that a result holds only for one inbred strain.

The output of that design is instructive. A 2017 programme update reported 53 lifespan experiments involving 30 test agents over the first eleven years, with significant effects on longevity, in one or both sexes, published for six of them. A 2025 review covering two decades put the totals at 54 agents in more than 30,000 mice, and singled out how often the effect appeared in one sex and not the other.

Two things follow. The first is how much testing it takes to produce a short list. The second is that “worked in mice” and “worked in mice, repeated at three sites, in both sexes, on a standard protocol” are claims of very different strength, and both are reported with the same four words.

Has it been peer reviewed?

The sixth question is the quickest to answer. If the link goes to bioRxiv or medRxiv, the work is a preprint: posted publicly by its authors, and not yet checked by anyone outside the group that produced it.

medRxiv says so itself, and in unusually direct terms. Manuscripts posted there are not certified by scientific peer review, edited, or typeset. Readers are warned that the articles have not been finalised by their authors, might contain errors, and report information that has not yet been accepted or endorsed by the scientific or medical community. The site asks journalists to say so when reporting on them, and states that preprints should not be used to directly inform clinical decision-making.

None of this makes preprints worthless. Posting early is how a field moves quickly, and peer review is itself an imperfect filter. But a preprint is a result in public, not a result confirmed, and the distinction disappears in most coverage.

What each kind of study can support

Kind of study What it can establish What it cannot
Cell or tissue work That a mechanism exists and can be altered Anything about a whole organism
Animal lifespan study That a compound changed survival in that species, strain and protocol That the same holds in humans, or in another strain
Early or uncontrolled human study Tolerability, dose, and whether a marker moves Whether the marker’s movement is worth having
Randomized controlled trial That the treatment caused the change in what it measured Effects beyond its own endpoint, duration and population
Meta-analysis of trials What the pooled evidence shows, including the absence of an effect More than the trials it pools: weak inputs give a precise weak answer

The last line deserves emphasis. A meta-analysis is the strongest kind of study in this table, and it is a summary of other studies, not a replacement for them. Pooling several small, brief trials of a marker produces a confident statement about that marker in those conditions, and nothing more.

The short version

Six questions, in the order in which they save the most time:

  1. What was measured, and is it healthspan, lifespan, or a marker standing in for one of them?
  2. Was the work done in people, or in animals?
  3. Was assignment random, or were existing habits observed?
  4. How large was the effect in absolute terms, not as a percentage?
  5. Has it been repeated by anyone else?
  6. Has it been peer reviewed, or is it a preprint?

A claim that survives all six is worth attention. Most do not get past the first.

Questions

Does a study in mice mean nothing for people?
It means less than coverage of it usually implies. Work in animals is where nearly every human treatment begins, and it is also where most of them end. A review in PLOS Medicine found that about one-third of highly cited animal studies went on to be tested in human randomised trials, and about one-tenth of the interventions were eventually approved for patients.
Why do so many longevity studies measure blood markers instead of lifespan?
Because measuring a human lifespan takes a human lifetime. A marker that moves in months is the only practical option, so it is used as a substitute. The substitution is legitimate research practice; what is not legitimate is reporting the marker as though it were the lifespan.
Is a preprint worthless?
No, but it has not been checked by anyone outside the research group. medRxiv states that its manuscripts are not certified by peer review and warns that they might contain errors and should not be used to directly inform clinical decisions. A preprint is a result in public, not a result confirmed.
If a trial found no effect, does that settle the question?
Only for what that trial measured, in the people it enrolled, over the time it ran. A trial too small or too short to detect a difference will not find one whether or not a difference exists, which is why the number of participants and the length of follow-up belong in any summary of a null result.
What is the single most useful question to ask?
What exactly was measured. Almost every overstated longevity claim survives on the gap between the thing that was measured and the thing the reader assumes was measured.

Sources

  1. Can animal models of disease reliably inform human studies? PLOS Medicine, 2010
  2. Surrogate endpoint resources for drug and biologic development, US Food and Drug Administration
  3. The NIA Interventions Testing Program: an update, Innovation in Aging, 2017
  4. The Gehan test identifies life-extending compounds overlooked by the log-rank test in the NIA Interventions Testing Program, GeroScience, 2024
  5. Sex as a major determinant of pro-longevity drug efficacy: a review of two decades of the NIA Interventions Testing Program, Journals of Gerontology Series A, 2025
  6. medRxiv frequently asked questions

More in Research Watch