When is a test result worth acting on?
Four questions that decide whether a number from a longevity panel should change anything, and why most results fail the first one.
Published 5 min read
In short
A test result is worth acting on when it is reliable across repeat measurements, when it predicts something that matters, and when a different result would lead to a different decision.
Most longevity panel measurements fail the first condition or the last, and the two failures look identical on a printed report.
The distinction regulators draw between a validated surrogate and a reasonably likely one is the same distinction, written in the language of drug approval.

A longevity panel returns numbers. Some are familiar clinical measures with decades of use behind them. Others are newer: an estimate of biological age, a pace-of-aging score, an inflammation index. On the page they look alike, and the report gives no hint that they differ in how much they can support.
Four questions separate them, in the order that saves the most effort.
Would the same measurement repeat?
This comes first because nothing survives its failure. If a number moves on its own, no interpretation of one reading is safe, however sophisticated the interpretation.
The question has two halves that are easy to conflate, and the research on epigenetic clocks has made the distinction unusually concrete.
Technical reliability asks whether the same sample gives the same answer twice. In 2022, a study in Nature Aging reported that technical noise alone produced deviations of up to nine years between replicates for six prominent clocks — before any biology was involved. The same paper’s principal-component method brought most replicates within about 1.5 years.
Biological reliability asks something harder: whether the same person gives the same answer twice. A 2026 study in Aging Cell evaluated 18 methylation-based aging biomarkers across repeated samples taken within short intervals under varying conditions — before and after meals, under acute stress, across different exposures — and found most showed only low to moderate stability.
The finding worth carrying away is the relationship between the two: technical reproducibility did not predict biological reliability. A laboratory can be precise about a tube of blood and imprecise about the person it came from, and the certificate will describe only the first.
Does it predict what it stands for?
Every measurement on a longevity panel is a surrogate endpoint: a stand-in for something that cannot be measured in the time available. The US Food and Drug Administration defines it as a trial endpoint used as a substitute for a direct measure of how a patient feels, functions or survives.
The regulator then grades these substitutes, and the grade is the answer to this question. A validated surrogate endpoint is supported by a clear mechanistic rationale and by clinical data providing strong evidence that an effect on it predicts a specific clinical benefit. A reasonably likely surrogate endpoint has the mechanistic or epidemiological rationale, but not enough clinical data to show that it predicts anything.
Applied to a consumer panel, the question becomes concrete: has anyone shown that changing this number changes an outcome? For most markers in this field, the answer is that nobody has looked, which is a different statement from “no” and should not be reported as either.
Is it valid for someone like me?
A model performs on the kind of data it was trained on. Moving outside that is not a technicality.
This is the burden of a 2025 review in Epigenomics, which argued that fundamental technical and biological properties of epigenetic clocks prohibit their current use at the individual level, and that the clocks fail to meet common standards for clinical utility compared with established biomarkers. The reviewers’ case is not that the science is bad — it is that a tool built and validated on populations is being handed to individuals without the validation that use would require.
What would change?
The last question is the most practical and the least asked: if the number came back different, what would be done differently?
Where an answer exists, the test is doing work. A cholesterol result changes a treatment decision. A blood glucose result changes a diagnosis. Where no answer exists — where the same advice about sleep, exercise and diet follows from any possible reading — the test is producing a number rather than a decision.
This is not an argument that such tests are worthless. Curiosity is a legitimate reason to buy something, and tracking a measure over years under identical conditions is a defensible project. It is an argument against the specific confusion these reports invite: that a result which cannot change a decision is nonetheless telling you what to do.
The four questions
| Question | What a good answer looks like |
|---|---|
| Would it repeat? | The provider can name the method and its replicate agreement |
| Does it predict what it stands for? | Published clinical data, not a mechanism story |
| Is it valid for someone like me? | The populations it was built on are stated |
| What would change? | A specific decision that depends on the result |
A measurement that passes all four is worth acting on. A measurement that fails the first is not worth interpreting at all, and a measurement that fails only the last is worth having — as information, not as instruction. Establishing which is which requires a randomized controlled trial somewhere in the chain, and for most of this panel there is not one yet.
Questions
- What is the first thing to check about a test result?
- Whether the same measurement, repeated, would give the same answer. If a result moves by several years depending on the day or the laboratory, no interpretation of a single reading is safe.
- What is the difference between technical and biological reliability?
- Technical reliability asks whether the same sample gives the same answer twice. Biological reliability asks whether the same person does. A 2026 study found that the first does not predict the second.
- Why does it matter whether a marker is validated?
- Because validation is the evidence that moving the marker moves the outcome. The FDA reserves the term for surrogates supported by clinical data showing exactly that; markers with only a mechanistic rationale are graded lower.
- What is the last question?
- What would change. If the same advice follows from any possible result, the test is not informing a decision, whatever else it is doing.
Sources
- Surrogate endpoint resources for drug and biologic development, US Food and Drug Administration
- Biological versus technical reliability of epigenetic clocks and implications for disease prognosis and intervention response, Aging Cell, 2026
- From population science to the clinic? Limits of epigenetic clocks as personal biomarkers, Epigenomics, 2025
- A computational solution for bolstering reliability of epigenetic clocks, Nature Aging, 2022