Skip to content
Eonix
Menu

Do biological age tests actually work?

What the reliability studies found when they measured the same sample twice, and then measured the same person twice.

Published 6 min read

In short

Epigenetic clocks work well on populations and poorly on individuals, and the tests sold to consumers are individual measurements.

Technical noise alone produced deviations of up to nine years between repeat readings of the same sample for six prominent clocks, before any biology was involved.

A 2026 study of 18 aging biomarkers found most were technically reproducible but only low to moderately stable across repeated samples from the same person taken hours apart, before and after meals or under stress.

One sample vial on a bench, connected by five tubes to five identical instrument boxes, whose five needles point in five different directions on blank dials.
One sample, five instruments, five answers. The spread is in the measurement, not in the person.Illustration

The pitch is straightforward: a saliva or blood sample, a few weeks, and a number telling you how old your body is as distinct from how old you are. The number is real in the sense that a real measurement produced it. Whether it is about you is the question the research has spent the last several years answering, and the answer is more interesting than either the marketing or the dismissals.

What the test is measuring

Cells carry chemical tags on their DNA, and the pattern of those tags shifts with age. The measurement is DNA methylation, the model that turns it into a number is an epigenetic clock, and the number is called biological age.

An important detail disappears at this point in almost every product description: dozens of these clocks exist, and they were trained to predict different things. The earliest were built to guess chronological age from a sample. Later ones, including PhenoAge and GrimAge, were trained against health measures or time to death. Newer ones estimate a pace of aging rather than a level.

Those are not variations on one measurement. A person who scores young on a clock trained to predict chronological age is someone the model found hard to place. A person who scores young on a clock trained to predict mortality is someone whose methylation pattern resembles that of people who lived longer. Two tests, both labelled biological age, answering different questions.

Does the same sample give the same answer?

This is the easier of the two reliability questions, and for years the answer was uncomfortable.

In 2022, Higgins-Chen and colleagues reported in Nature Aging that epigenetic clock data can be surprisingly unreliable: technical noise alone produced deviations of up to nine years between replicates for six prominent clocks. The same tube of blood, split and run twice, could come back nearly a decade apart.

Nine years is not a rounding error. It is larger than any effect a lifestyle change is plausibly going to produce, which means that for a time the instrument was less precise than the thing it was being used to detect.

The same paper offered a fix. By computing principal components from the underlying methylation data before predicting age, the authors produced retrained versions of six clocks whose replicates mostly agreed within 1.5 years. This is real progress, and the better commercial tests use methods of this kind. The improvement is worth asking about by name, because it distinguishes a current test from one built on the older approach.

Does the same person give the same answer?

This is the harder question, and it is where the field’s most useful recent finding sits.

A 2026 study in Aging Cell evaluated 18 DNA methylation-based aging biomarkers for both kinds of reliability. Technically, most performed well on replicate assays, though several proved sensitive to mundane experimental factors such as the position of a sample on a slide and the protocol used to extract DNA. Principal-component clocks, particularly PCGrimAge and SystemsAge, held up across all the conditions tested.

Biological reliability was another matter. Assessed across repeated samples collected within short intervals under varying conditions — before and after meals, under acute stress, across different environmental exposures — stability was substantially lower, with most clocks showing only low to moderate stability. Adjusting for the mix of immune cells in the sample made it worse.

The authors added the observation that turns this from a technical footnote into something a buyer should know: technical reproducibility did not predict biological reliability. A test can be impeccably consistent about a sample and still be inconsistent about the person, and being told the first tells you nothing about the second.

Population science, personal product

The gap between what these models are good at and what they are sold for has now been argued explicitly in the literature.

A 2025 review in Epigenomics set out the case. Epigenetic clocks have been instrumental at the population level, the authors write, revealing how disease risk emerges from behavioural, environmental and psychosocial factors and how some interventions may alter those trajectories. Given that success, it is reasonable to assume they might work as individual-level biomarkers. Their contention is that they do not: technical and biological properties of the algorithms prohibit their current use at the individual level, and the clocks fail to meet the standards commonly applied to established clinical biomarkers.

A separate 2025 review in Clinical Epigenetics raises a second limit that consumer tests rarely mention. Participants in epigenetic research are overwhelmingly of European ancestry, and it is unclear whether these clocks deliver equitable results when applied to other populations. A model is reliable on the kind of data it was trained on; nobody has established how far outside that these clocks remain meaningful.

What a result can and cannot support

The claim Does the evidence support it?
“Your biological age is 41.3” No. The precision is an artefact of the arithmetic, not of the measurement
“This population ages faster than that one” Yes. This is what the clocks were built for and where they have worked
“Your result improved, so the supplement is working” No. Repeated measures of the same person shift with meals, stress and sample handling
“This test is more reliable than older ones” Possibly, if it uses principal-component methods — worth asking which clock it runs
“Your result means you should change your treatment” No. No regulator has approved these tests for clinical decisions

Everything a clock produces is a surrogate endpoint: a stand-in for lifespan or healthspan, used because the real thing takes a lifetime to observe. The US Food and Drug Administration distinguishes surrogates validated by clinical data from those merely supported by mechanistic reasoning. No epigenetic clock is in the first category.

The honest use

If the appeal is curiosity, a test is a reasonable purchase and an interesting one. If the intention is to track change, the requirements are strict: the same test, the same laboratory, the same time of day, the same conditions, repeated over years rather than months, and read as a trend rather than as a series of verdicts. Anything less is measuring the noise the reliability studies described.

What a result cannot do is tell you how long you will live, or settle whether something you are taking is working. The research community that built these clocks is clear about this in print. The market that sells them is not.

Questions

Are biological age tests a scam?
No. The underlying science is real and has produced genuine findings about populations. The problem is the product: a measurement that is informative across thousands of people is not automatically informative about one, and that is the use being sold.
How far off can a result be?
A 2022 study in Nature Aging found that technical noise alone produced deviations of up to nine years between repeat readings of the same sample for six prominent clocks. Newer principal-component versions brought most replicates within about 1.5 years.
Does eating or sleeping badly change my result?
The evidence says it can. A 2026 study measuring the same people repeatedly within short intervals, including before and after meals and under acute stress, found most clocks only low to moderately stable across those conditions.
Do different tests agree with each other?
Often they do not, because they were built to predict different things. Some estimate chronological age, some estimate time to death, some estimate the rate of decline. Comparing two results labelled 'biological age' can mean comparing two different quantities.
Is there any use in taking one?
As a matter of curiosity, and possibly as a trend if the same test is repeated under the same conditions over years. A single result is not a basis for a health decision, and no regulator has approved these tests for that purpose.

Sources

  1. A computational solution for bolstering reliability of epigenetic clocks, Nature Aging, 2022
  2. Biological versus technical reliability of epigenetic clocks and implications for disease prognosis and intervention response, Aging Cell, 2026
  3. From population science to the clinic? Limits of epigenetic clocks as personal biomarkers, Epigenomics, 2025
  4. What's counted counts: the implications of underrepresentation for the application of epigenetic clocks in diverse populations, Clinical Epigenetics, 2025
  5. Surrogate endpoint resources for drug and biologic development, US Food and Drug Administration

More in Measures