AI-powered health forecasting

Companies increasingly sell software that reads your blood tests, scans or watch data and tells you which illnesses you are likely to get. A few of these tools genuinely do one narrow job well, but almost none of them has ever been shown to make anyone healthier.
AI health forecasting is real in a handful of narrow places — mammography reading, diabetic-retinopathy screening, polyp detection, low-ejection-fraction detection from an ECG — and is essentially unevidenced everywhere it is sold to consumers. The reason is a single distinction that marketing routinely erases: a high AUROC on a stored dataset means the model ranks patients correctly on data it never influenced. It is not evidence that using the model changes what happens to anyone. Roughly 1,500 AI-enabled devices hold FDA marketing authorisation, and there are about 41 randomised trials of machine-learning interventions in the whole of health care.
- Accuracy is not benefit: of 81 deep-learning-versus-clinician studies, only 9 were prospective and only 6 ran in a real clinical setting (Nagendran 2020, PMID 32213531).
- Colonoscopy AI is the most randomised application in medicine — 44 RCTs, adenoma detection up ~20% relative, and no increase at all in advanced neoplasia (Soleymanjahi 2024, PMID 39531400).
- No aging clock — epigenetic, retinal, brain, ECG or proteomic — is a validated treatment target, a validated surrogate endpoint, or an FDA-cleared diagnostic; the field's own consensus body says so in print (Moqri 2023, PMID 37657418).
- Machine learning frequently does not beat logistic regression on tabular clinical data: across 145 head-to-head comparisons at low risk of bias, the difference in logit(AUC) was 0.00 (95% CI −0.18 to 0.18) (Christodoulou 2019, PMID 30763612).
- Algorithmic bias is measured, not hypothetical: correcting one widely deployed care-management algorithm would have raised the share of Black patients receiving extra help from 17.7% to 46.5% (Obermeyer 2019, PMID 31649194).
The four levels of evidence, and why almost nothing reaches level 4
AUROC — the area under the receiver-operating curve — is the probability that a model ranks a random case above a random control. It is the number in nearly every AI health headline, and it answers a narrow question: does the model sort people correctly on data it has already been shown? It says nothing about whether the number displayed to you is true, whether anyone acts on it, or whether acting on it does more good than harm.
Four distinct levels of evidence sit behind the phrase clinically validated, and marketing routinely presents the first as though it were the fourth.
What an AI health claim can actually rest on
| Level | What it means | How common in AI health |
|---|---|---|
| 1. Retrospective AUROC | The model scores well on a stored dataset it never influenced. Says nothing about whether using it changes anything. | Nearly all published AI health research, and effectively all consumer longevity AI |
| 2. Prospective accuracy | The model runs live on consecutive real patients; accuracy is measured against a reference standard collected independently. Still no outcome. | Uncommon |
| 3. Randomised process outcome | A randomised trial shows the AI changes what clinicians do or find — detection rate, diagnosis rate, workload. | Rare: a few dozen trials in total, most of them in endoscopy |
| 4. Randomised patient outcome | A randomised trial shows the AI changes whether people live longer, have fewer strokes, or get fewer cancers. | Essentially non-existent in this field |
Where AI genuinely works — and where detection stops short of benefit
Imaging is the strongest area, and it is worth being precise about why. Mammography screening has a high-quality existing comparator (two radiologists reading every film), a hard downstream endpoint (interval cancer), and national programmes willing to randomise. That combination is rare, and it is what produced the best AI evidence in medicine.
MASAI randomised 105,934 Swedish women to AI-supported reading or standard double reading. Cancer detection rose from 5.0 to 6.4 per 1,000, a ratio of 1.29 (95% CI 1.09–1.51), with screen-reading workload down 44%. When the primary endpoint finally reported, interval cancer was 1.55 versus 1.76 per 1,000 — a proportion ratio of 0.88 (0.65–1.18). That is non-inferior, not superior. Nobody has shown that AI-supported mammography reduces breast cancer mortality, and given the trial sizes required, nobody is likely to.
Below the imaging tier, a consistent pattern appears: AI reliably finds more of something, and finding more of it has repeatedly failed to help. Smartwatch and patch screening doubles-to-triples atrial fibrillation detection. Three randomised trials with hard endpoints — LOOP, STROKESTOP and GUARD-AF — have not demonstrated a stroke reduction. The LOOP authors put it in their own abstract: not all atrial fibrillation is worth screening for, and not all screen-detected atrial fibrillation merits anticoagulation.
The evidence ledger
Ranked on: the highest tier of evidence that genuinely exists for each approach. These tiers were written for drugs, so read them as follows for software: approved means a real FDA marketing authorisation (510(k), De Novo or PMA) exists for the specific output being sold; preclinical means the work is real but exists only as models fitted to stored human datasets, with no clinical deployment evidence; marketed-unproven means the product is sold to the public with no regulatory review of the claim and no controlled evidence of benefit. Note that clearance and benefit are different axes: several approved-tier tools here have randomised evidence that they change detection and no evidence at all that they change outcomes.
Mammography AI (Transpara, Lunit INSIGHT MMG)
ApprovedReads screening mammograms as a triage tool and second reader alongside radiologists.
The best-evidenced AI in medicine. Detection ratio 1.29 and 44% less reading work in MASAI; interval cancer non-inferior, not reduced. Pooled across MASAI, ScreenTrustCAD and PRAIM (597,419 examinations) the benefit is +0.9 cancers per 1,000 screens, with the confidence interval touching zero. Real and useful. Not AI detects cancer years earlier.
FDA 510(k) cleared (K181704 → K241831; K211678). Randomised (MASAI) and prospective paired-reader (ScreenTrustCAD) evidence.NCT04838756; PMID 41620232Autonomous diabetic-retinopathy screening (IDx-DR / LumineticsCore)
ApprovedGrades fundus photographs for more-than-mild diabetic retinopathy without a clinician reading the image.
Sensitivity 87.2% (81.8–91.2), specificity 90.7% (88.3–92.7) in 900 patients, in primary-care offices rather than eye clinics, against independent reading-centre grading. The cleanest example in this dossier of AI done properly. Competitors cleared on the same product code: EyeArt (K200667) and AEYE-DS (K221183).
FDA De Novo DEN180001, 11 April 2018 — the first autonomous AI diagnostic authorised in any field. Prospective pivotal trial in primary care.NCT02963441; PMID 31304320Colonoscopy computer-aided detection (CADe)
ApprovedFlags polyps in the endoscopist's live field of view.
Adenoma detection rate ratio 1.21 (1.15–1.28) — robust and consistent. But advanced neoplasia per colonoscopy showed no difference at all (incidence rate difference 0.01, −0.01 to 0.02), the increase was driven largely by diminutive lesions, and there were roughly two extra non-neoplastic polypectomies per ten procedures. The most randomised AI application in medicine, and level 3 did not become level 4.
FDA cleared and CE-marked. 44 randomised trials, 36,201 patients.PMID 39531400; PMID 37639719AI-ECG low ejection fraction (Anumana, Eko, Tempus)
ApprovedFlags probable left-ventricular ejection fraction at or below 50% from a standard 12-lead ECG.
Genuinely accurate — AUC 0.93 for EF ≤35% in the development work — and genuinely randomised. Note what it predicts: your ejection fraction now, not your future risk. And note the absolute effect in EAGLE: new low-EF diagnoses rose from 1.6% to 2.1%.
FDA 510(k) cleared as notification software (K232699, K233409, K250119, K250652). Cluster-randomised trial (EAGLE, 22,641 patients).NCT04000087; PMID 30617318Lung nodule malignancy scoring (Optellum Virtual Nodule Clinic)
ApprovedScores the malignancy risk of a lung nodule a human has already found on CT.
A risk-stratification aid with no trial showing it changes stage at diagnosis or mortality. Worth knowing how the underlying research behaves: in the landmark Google NLST model, performance beat all six radiologists when no prior CT was available and was merely on par with them when a prior scan existed — which is the clinically realistic condition.
FDA 510(k) cleared (K202300, 2021). Retrospective validation only; no RCT.K202300; PMID 31110349Smartwatch irregular-rhythm notification
ApprovedDetects a pulse pattern suggestive of atrial fibrillation and sends a notification.
Detection works; benefit does not follow. Only 34% (Apple) and 32.2% (Fitbit) of alerts were confirmed as AF on a follow-up patch. Taking a 10-second watch ECG after an alert has a diagnostic yield of 7.6%, versus 60.8% for a four-week patch. And across LOOP, STROKESTOP and GUARD-AF, finding more AF has not been shown to prevent strokes. USPSTF still rates the evidence for AF screening in adults 50 and over as insufficient — a Grade I statement.
FDA De Novo DEN180042 (Apple). Pragmatic studies in 419,297 (Apple Heart) and 455,699 (Fitbit Heart) participants.NCT03335800; NCT04380415; PMID 35076659Consumer hypertension and sleep-apnea notifications
ApprovedFlags multi-week wearable patterns suggestive of hypertension or moderate-to-severe sleep apnea.
The clearance summaries are more honest than the marketing. For hypertension, the FDA summary reports sensitivity 41.2% (37.2–45.3) and specificity 92.3% in 1,863 analysed participants — it misses roughly 59% of hypertension and 46% of stage 2 hypertension, and its own indication forbids using it for surveillance or to track treatment. For sleep apnea, sensitivity 66.3%. Both are legitimate as a nudge toward a proper test. Neither is a measurement.
FDA 510(k) cleared (Apple hypertension K250507, Sep 2025; Apple sleep apnea K240929, predicate Samsung DEN230041). Sponsor validation only.K250507; K240929Sepsis prediction inside the electronic health record (Epic Sepsis Model)
Sold · unprovenContinuously scores inpatients for probable sepsis onset.
AUC 0.63 across 38,455 hospitalisations, missing 67% of sepsis patients while alerting on 18% of all admissions. It identified only 7% of septic patients beyond what timely clinical practice already caught. The authors call widespread adoption despite this performance a fundamental concern. The regulatory gap that let it deploy unexamined is the point.
Not FDA-regulated — it shipped as EHR functionality to hundreds of US hospitals. External validation exists, and it failed.PMID 34152373Consumer epigenetic biological-age tests
Sold · unprovenReturn a single biological age number from DNA methylation in blood, saliva or a cheek swab.
The underlying clocks are real, replicated mortality associations. The consumer product is a precision claim the assay cannot support: technical noise produces deviations of up to 9 years between replicate runs of the same sample for six prominent clocks, and oral-versus-blood tissue differences reach almost 30 years in the same person for some clocks. Most direct-to-consumer tests use saliva or cheek swabs.
Laboratory-developed tests under CLIA only. FDA's LDT rule was vacated on 31 March 2025 and formally rescinded on 19 September 2025, so no regulator reviews the clinical validity of the age claim.PMID 36277076; PMID 39780748Retinal, brain, ECG and proteomic age gaps
Preclinical onlyEstimate an organ or whole-body age from a retinal photo, MRI, ECG or plasma proteome, then report the gap versus your birthday.
Statistically real at population scale and much weaker than they sound individually. Retinal age gap: all-cause mortality HR 1.02 per year, confidence interval touching 1.00, with no significant association with cardiovascular or cancer mortality. Brain age: mortality HR 1.061 per year — but when entered alongside simple grey-matter and CSF volumes it was no longer significant (p=0.12), so the deep-learning step added nothing over a volume measurement.
Research tools. No FDA clearance for any of these outputs; device-name searches of FDA's 510(k) database return no cleared device that outputs a biological or physiological age.PMID 35042683; PMID 28439103Wellness scores: readiness, recovery, strain, sleep score, HRV, wearable age
Sold · unprovenComposite daily scores derived largely from heart rate, HRV and estimated sleep stages.
Consumer wearables assign the correct sleep stage 50–65% of the time (Cohen's kappa 0.20–0.52), and are biased toward calling everything sleep: wake-detection specificity runs 0.18–0.54. Every deep-sleep minute count and every REM-derived recovery score is built on that. FDA's Whoop warning letter also established on the record that a not-for-medical-use disclaimer does not rescue a product designed to produce a clinical estimate.
Unregulated general wellness. No FDA clearance was identified for Oura, Garmin or Ultrahuman in FDA's database; Whoop's only clearance is an ECG feature (K243236).PMID 36016077; PMID 33378539General-purpose LLM health chatbots
Sold · unprovenAnswer health questions, triage symptoms, and suggest what to do next in natural language.
The models are not the weak link — the handoff is. In a preregistered randomised study, models tested alone identified the condition in 94.9% of cases; 1,298 members of the public using those same models identified conditions in under 34.5% — no better than control. On 2,400 real patient cases rather than vignettes, state-of-the-art models performed significantly worse than physicians and were sensitive to the order in which information was given.
No general-purpose LLM chatbot is FDA-cleared or authorised as a medical device, and FDA's own device list does not yet identify which authorised devices contain LLM functionality.PMID 41663592; PMID 38965432Multi-omics panels and whole-body MRI screening services
Sold · unprovenBundle hundreds of biomarkers, genomics and a whole-body scan into a personalised health forecast.
Pooled across 12 studies and 5,373 asymptomatic subjects, whole-body MRI produced critical or indeterminate incidental findings in 32.1%, of which only 12.6% were ever verified. One vendor's own database of 21,651 asymptomatic scans found incidental pancreatic cysts in 7.0% of everyone scanned, of which about 99% were below the size that would trigger intervention — yet all enter surveillance. The American College of Radiology states there is not sufficient evidence to recommend total-body screening in people without symptoms, risk factors or family history.
CLIA-only laboratory-developed tests plus general-wellness enforcement discretion. Scanner and image-enhancement components may be cleared; the screening service is not.PMID 30932247; PMID 41801199
Cleared medical device vs consumer wellness app: they are not the same object
The single most useful skill in this field is telling apart four regulatory objects that sound identical in marketing copy: an FDA-authorised medical device, a laboratory-developed test running under CLIA, a general-wellness product operating under enforcement discretion, and software with no regulatory relationship at all.
Even inside the cleared category the evidence bar is low. Of ~1,524 authorisations, 96.2% went through 510(k), which asks whether a device is substantially equivalent to a predicate already on the market — not whether it improves outcomes. In the 2024 cohort, only 29.2% reported both sensitivity and specificity, only 15.5% reported race or ethnicity data, and only 16.7% shipped with a Predetermined Change Control Plan, meaning most cleared models are frozen at clearance and cannot be updated as populations drift.
What is cleared, what is proven, and what each tool actually predicts
| Tool or claim | Regulatory status | Prospective evidence? | What it actually predicts |
|---|---|---|---|
| Mammography AI (Transpara, Lunit) | FDA 510(k) cleared; CE-marked | Yes — randomised (MASAI, NCT04838756, n=105,934) and prospective paired-reader (ScreenTrustCAD, NCT04778670, n=55,581) | Presence of screen-detectable breast cancer on this mammogram |
| IDx-DR / LumineticsCore | FDA De Novo DEN180001 (Apr 2018) | Yes — prospective pivotal trial in primary care, NCT02963441, n=900 | More-than-mild diabetic retinopathy or macular oedema from fundus photos |
| Colonoscopy CADe | FDA cleared; CE-marked | Yes — 44 randomised trials, 36,201 patients | Presence of a polyp in the current field of view. No increase in advanced neoplasia |
| AI-ECG low ejection fraction | FDA 510(k) cleared (K232699, K250652) | Yes — cluster RCT (EAGLE, NCT04000087, n=22,641) | Current LV ejection fraction ≤50% — not future risk |
| Optellum Virtual Nodule Clinic | FDA 510(k) cleared (K202300) | No — retrospective validation only | Malignancy risk of an already-detected lung nodule |
| Smartwatch irregular-rhythm notification | FDA De Novo DEN180042 | Yes — pragmatic studies, n=419,297 and n=455,699 | Probable atrial fibrillation now. 34% / 32% of alerts confirm AF |
| AF screening as stroke prevention | Not a device claim | Yes — three RCTs with hard endpoints (LOOP, STROKESTOP, GUARD-AF) | No demonstrated stroke reduction. LOOP HR 0.80 (0.61–1.05); GUARD-AF HR 1.10 (0.69–1.75) |
| Apple Hypertension Notification | FDA 510(k) K250507 (Sep 2025) | Sponsor validation only, n=1,863 | Patterns suggestive of hypertension. Sensitivity 41.2% |
| Epic Sepsis Model | Not FDA-regulated — deployed as EHR functionality | External validation only — and it failed | Nominally sepsis onset; actually performed at AUC 0.63, missing 67% of cases |
| Epigenetic clocks (Horvath, PhenoAge, GrimAge, DunedinPACE) | No FDA clearance for any clinical claim; sold as LDTs or wellness | No — association studies only | Chronological age, or mortality risk, depending on the training target. Not a modifiable target |
| Retinal, brain and proteomic age gaps | Research tools; no clearance | No — cohort associations, largely within single datasets | Mortality and morbidity risk at population level |
| Consumer biological-age tests | CLIA-only LDTs; the FDA LDT rule was vacated 31 Mar 2025, so no FDA review of the age claim | None | Nothing validated. The same sample run twice can differ by years |
| HRV, readiness, recovery, strain, sleep scores, wearable age | Unregulated general wellness | None | Nothing validated. Sleep stage is correct 50–65% of the time |
| OTC continuous glucose monitors in non-diabetics | Cleared for glucose measurement only (Stelo K234070, Lingo K233655, Libre Rio K233861) | No outcome RCT in people without diabetes | Interstitial glucose, reading roughly 0.9 mmol/L above capillary sampling |
| General-purpose LLM chatbots | No FDA-authorised device is identified as LLM-powered | Benchmarks plus three RCTs — in which the human-plus-model arm did not beat the model alone | Benchmark answers, not clinical outcomes |
| Digital twin (Unlearn.AI / PROCOVA) | EMA qualification opinion — for clinical-trial efficiency; described by EMA as a special case of ANCOVA | Not applicable — it is a statistical method | Trial sample-size reduction. Not personalised treatment |
| Whole-body MRI screening services | Scanner and image-enhancement components cleared; the screening service is not | Single-arm observational only (NCT06212479, completes 2037) | Incidental findings. 32.1% pooled incidentaloma rate; ACR says the evidence is insufficient |
The regulatory vocabulary, decoded
| The phrase | What it actually attaches to |
|---|---|
| FDA-cleared AI (whole-body MRI services) | An image-enhancement algorithm, not the screening service being sold |
| FDA-cleared (Whoop) | A single ECG feature, K243236. Recovery, Strain, Whoop Age and Healthspan are unregulated |
| EMA-qualified digital twin | A special case of ANCOVA for reducing clinical-trial sample size. FDA declined it as already covered by existing covariate-adjustment guidance |
| CLIA-certified (every biological-age and multi-omics test) | The laboratory measures the analyte reproducibly. It says nothing about clinical validity — and since 31 Mar 2025 no US regulator reviews that claim |
| FDA-registered / FDA-listed | Facility registration. Meaningless for the claim being made |
| Clinically validated | Usually a retrospective correlation study, often authored by the company selling the product |
| Our AI analyses 100+ biomarkers | Analysing is not predicting, and predicting is not validated |
| Not intended to diagnose, treat, cure or prevent any disease — alongside disease-detection marketing | A legal posture. FDA ruled on the record in July 2025 that such disclaimers do not rescue a product designed to produce a clinical estimate |
What changed in 2025–2026
- Jan 2025
PRAIM published — the largest real-world AI screening implementation
463,094 women across 12 German screening sites. Cancer detection 6.7 vs 5.7 per 1,000, +17.6% (95% CI +5.7% to +30.8%), with recall lower and non-inferior. Observational, but the direction agrees with MASAI.
Eisemann 2025, Nat Med, PMID 39775040
- 31 Mar 2025
FDA's laboratory-developed test rule is vacated
A federal district court held in ACLA v. FDA that LDTs are professional services, not devices. This is the single most consequential fact for consumer longevity testing: biological-age tests and multi-omics panels revert to CLIA only, with no FDA review of clinical validity, indefinitely.
FDA final rule implementing vacatur, 90 FR 45134, 19 Sep 2025
- Apr 2025
Whoop receives its first-ever FDA clearance — for an ECG feature
K243236 covers detection of AFib, normal sinus rhythm and high or low heart rate, for informational use only. Recovery, Strain, Sleep Coach, Whoop Age and Healthspan are not cleared for anything.
FDA 510(k) K243236
- 14 Jul 2025
FDA warning letter to Whoop over Blood Pressure Insights
FDA stated it has not authorised the feature for any use, rejected the general-wellness exemption, and rejected the disclaimers explicitly — noting that their inefficacy is demonstrated by people using the feature to monitor their hypertension. The close-out letter of 17 June 2026 is enforcement discretion against a relabelled product, not a clearance.
FDA Warning Letter, MARCS-CMS 709755
- Aug 2025
First quantified deskilling signal
Across four Polish endoscopy centres and 1,443 colonoscopies, endoscopists' unassisted adenoma detection rate fell from 28.4% to 22.4% after AI exposure (−6.0 percentage points, 95% CI −10.5 to −1.6, p=0.0089). Observational and in need of replication, but it is a systemic risk no accuracy metric captures.
Budzyń 2025, Lancet Gastroenterol Hepatol, PMID 40816301
- Sep 2025
Apple Hypertension Notification cleared at 41.2% sensitivity
The largest consumer cardiovascular AI expansion to date, and the most modest performance yet cleared. Its own indication forbids using it to replace diagnosis, monitor treatment effect, or perform blood-pressure surveillance.
FDA 510(k) K250507
- Jan 2026
MASAI reports its primary endpoint
Interval cancer 1.55 vs 1.76 per 1,000; proportion ratio 0.88 (0.65–1.18), p=0.41 — non-inferior, not superior. Sensitivity 80.5% vs 73.8% (p=0.031), specificity 98.5% in both arms. The definitive AI screening result, and it rules out the main fear rather than proving a mortality benefit.
Gommers 2026, Lancet, PMID 41620232
- Feb 2026
The definitive consumer-AI result
1,298 UK lay participants randomised to use frontier LLMs or a source of their choice. The models alone identified the condition in 94.9% of cases; the same models in the public's hands, under 34.5% — no better than control. The authors state that standard benchmarks do not predict the failures found with human participants.
Bean 2026, Nat Med, PMID 41663592
- Mar 2026
The screening cascade, quantified from a vendor's own data
Among 21,651 asymptomatic people scanned with whole-body MRI, 7.0% had an incidental pancreatic cystic lesion, rising to 20.8% at age 80+. 96.0% were under 2 cm and only 1.1% were 3 cm or larger — one organ, one finding type, and roughly 99% below the intervention threshold.
Wong 2026, JAMA Netw Open, PMID 41801199
- Aug 2026
The honest effect size for AI mammography
Pooling MASAI, ScreenTrustCAD and PRAIM across 597,419 examinations gives a cancer-detection risk difference of +0.9 per 1,000 (95% CI −0.0 to +1.8) and no consistent increase in recall.
Ferre 2026, Clin Imaging, PMID 42617468
Aging clocks: real associations, and not one validated target
Aging clocks — epigenetic, retinal, brain, ECG, proteomic — are the AI forecasting product the longevity industry sells hardest. They are real, replicated statistical associations with mortality. Not one of them is a validated surrogate endpoint, a validated treatment target, or an FDA-cleared diagnostic. That is not our editorial position; it is the stated position of the field's own consensus body.
The most important thing about any clock is what it was trained against, because that determines what it can possibly predict. First-generation clocks (Horvath, Hannum) were trained on chronological age — a perfect Horvath clock would carry zero mortality signal, because all the biological-age information lives in its error term. GrimAge was trained directly on time to death, using methylation surrogates for plasma proteins and smoking pack-years; that is legitimate mortality prediction, and it is partly a smoking detector rather than a measurement of molecular aging. DunedinPACE was trained on the rate of change across 19 organ-system indicators over 20 years and carries the strongest mortality hazard ratios (1.65, 1.51–1.79 in Framingham Offspring).
The clocks also disagree with each other. Many existing epigenetic clocks correlate only weakly among themselves, implying they capture different processes — two clocks on the same blood sample can tell different stories. And the only randomised evidence is thin and split: in CALERIE, two years of 25% caloric restriction slowed DunedinPACE but did not significantly change PhenoAge or GrimAge. Three clocks, the same randomised samples, three different answers.
What a consumer report cannot support
- Precision: technical noise produces deviations of up to 9 years between replicate runs for six prominent epigenetic clocks — PhenoAge median deviation 2.4 years, maximum 8.6 years. A 3-year improvement after an intervention sits comfortably inside that noise.
- Tissue: cross-tissue comparison across 284 samples found average differences of almost 30 years between oral and blood tissue in the same person for some clocks; only the Skin & Blood clock was concordant. Most direct-to-consumer tests use saliva or a cheek swab.
- Independence: every published retinal age gap outcome comes from the same UK Biobank dataset and the same model trained on the same 11,052 people. In the head-to-head comparisons the authors ran, retinal age did not beat routine risk factors for stroke or Parkinson's.
- Added value: brain age gap predicted mortality at HR 1.061 per year — but ceased to be significant (p=0.12) once simple grey-matter and CSF volumes were in the model, which remained significant. The deep-learning step added nothing over a volume measurement.
- Method stability: brain age models systematically over-estimate in the young and under-estimate in the old, so the raw gap carries chronological age leaking through. Different correction methods give different answers, and an effect size is uninterpretable unless the paper states which correction it used.
- Development quality: a systematic appraisal of 81 aging clocks across 59 studies found 31 (38.3%) were developed with neither internal nor external validation, the majority were rated high risk of bias — even those in high-impact journals — and only 3 clocks (3.7%) were rated low-concern.
There is also no US regulator checking any of this. Every consumer epigenetic-age test is a laboratory-developed test. FDA's LDT rule was vacated on 31 March 2025 and formally rescinded on 19 September 2025, so these tests sit under CLIA alone. CLIA certifies the laboratory's analytic performance — accuracy, precision, quality control, personnel competency. It does not evaluate whether the claim the test makes is clinically true. A biological-age result can be analytically reproducible and still clinically meaningless, and per the reliability data above, even the analytic half is not automatic.
Does machine learning actually beat classical statistics?
This is the question behind every AI disease risk prediction vs traditional methods comparison, and the honest answer is uncomfortable for both camps: the literature is mixed, and the apparent ML advantage shrinks the more carefully the comparison is done.
The landmark methodological result reviewed 71 studies containing 282 head-to-head comparisons of logistic regression against machine learning for binary clinical outcomes. Across the 145 comparisons judged at low risk of bias, the difference in logit(AUC) between the two was 0.00 (95% CI −0.18 to 0.18) — exactly zero. Across the 137 comparisons at high risk of bias, ML looked 0.34 (0.20–0.47) better. The ML advantage in the published literature is, to a first approximation, the bias. The same review found 79% of studies did not address calibration at all.
This keeps reproducing. A 2026 appraisal of 52 studies predicting IVIG resistance in Kawasaki disease found ML and logistic regression indistinguishable on external validation (pooled AUC 0.76 vs 0.75) while ML looked clearly better on internal validation (0.86 vs 0.76) — the classic overfitting signature. There is a counter-current: a 2025 meta-analysis of EHR-based cardiovascular models reported pooled AUCs of 0.865 for random forests and 0.847 for deep learning against 0.765 for conventional scores — but with I² above 99% and evidence of publication bias, and the authors' own conclusion is that methodological concerns limit current clinical applicability. Heterogeneity that high means the pooled estimate is not really estimating one thing.
What AUROC is silent about
- Calibration — whether a predicted 10% risk corresponds to a real 10% risk. Roughly four in five published ML prediction papers do not report it, and it is the only property that matters for a personal risk number.
- Prevalence — a model with AUROC 0.95 in an enriched 20%-prevalence dataset can be useless at 0.5% population prevalence, because positive predictive value collapses.
- Net benefit at the decision threshold you would actually use — whether acting on the score does more good than harm.
- Whether anyone acts on it. In EAGLE, the model was accurate, the trial was positive, and the absolute increase in new low-EF diagnoses was 1.6% to 2.1%.
- What the alternative was. The Epic Sepsis Model missed 67% of the cases clinicians were already treating.
- What happens once a human is attached. LLMs that identified conditions in 94.9% of cases alone dropped below 34.5% in the hands of real members of the public.
None of this is an argument for abandoning models. QRISK3, SCORE2 and Framingham are transparent, published in full, externally validated many times over, and are what guideline-based prevention actually runs on. It is an argument against assuming that a newer, more complex model shown to you by a company selling a subscription is better than the free one your doctor already uses. TRIPOD+AI now supplies a 27-item reporting standard; a model study that does not follow it should be treated as preliminary.
Algorithmic bias and LLM failure modes
Bias in health AI is not a hypothetical ethics-panel concern. It has been measured, at scale, in systems that were already deployed.
The landmark case examined a commercial risk-prediction algorithm used to allocate extra care management across millions of patients. At any given risk score, Black patients were considerably sicker than White patients. Correcting the bias would have raised the share of Black patients receiving extra help from 17.7% to 46.5%. The mechanism matters more than the number: the algorithm predicted future healthcare costs as a proxy for health need, and because less money is historically spent on Black patients at the same level of illness, cost is a biased proxy for illness. The model was accurate at its stated task and discriminatory in its actual use — and no amount of AUROC would have caught it.
In imaging, state-of-the-art chest X-ray classifiers tested across three large datasets consistently and selectively under-diagnosed under-served populations — female patients, Black patients, patients of low socioeconomic status — with the highest rates in intersectional subgroups. Underdiagnosis is the worst possible failure mode: it labels a sick person healthy and delays their care. And you usually cannot check for any of this in a cleared device, because only 15.5% of the machine-learning devices FDA authorised in 2024 reported race or ethnicity data for their evaluation population at all.
Documented LLM failure modes in health contexts
- Benchmarks do not transfer to users. Models alone identified conditions in 94.9% of cases; 1,298 lay participants using those same models managed under 34.5% — no better than control, in a preregistered randomised design.
- Performance collapses on real clinical data. Across 2,400 real patient cases in four abdominal pathologies, state-of-the-art models performed significantly worse than physicians, followed neither diagnostic nor treatment guidelines, could not interpret laboratory results, and were sensitive to both the quantity and order of information. The authors conclude LLMs are not ready for autonomous clinical decision-making.
- Sycophancy — the model agrees with your false premise. Given 50 drug brand–generic pairs with prompts asserting a false equivalence, baseline compliance with the illogical request was 100% for three GPT variants and 94% for Llama3-8B. The models had the knowledge to reject the premise and complied anyway. Ask a leading question about your own health, get your premise confirmed.
- Hallucinated citations. Across 636 references in 84 generated literature reviews, 55% of GPT-3.5 citations and 18% of GPT-4 citations were entirely fabricated; among the real ones, 43% and 24% respectively contained substantive errors. A reference that looks like a real study is not one.
- Brittleness to how you write. Typos, extra whitespace, missing gender markers and informal or dramatic language increased the likelihood that a model told the patient to self-manage rather than seek care, and the effect was larger for female patients.
- Even top-tier benchmark results are weaker than reported. On 70 curated clinicopathological cases — the easiest possible test — GPT-4 included the final diagnosis somewhere in its differential 64% of the time, but its top diagnosis was correct in only 39%.
There is a strange consolation in the randomised data: in two physician trials, the model alone matched or beat the physician-plus-model arm. In diagnostic reasoning, LLM assistance produced no benefit (adjusted difference 2 percentage points, 95% CI −4 to 8) while GPT-4 alone scored 16 points higher than physicians with conventional resources. In management reasoning, assistance did help physicians (+6.5%) but GPT-4 alone was statistically indistinguishable from GPT-4-augmented physicians. The bottleneck in 2026 is the handoff between model and human, and it is worst for members of the public — which is exactly the setting consumer health AI is sold into.
For completeness on the regulatory question: no general-purpose LLM chatbot is FDA-cleared or authorised as a medical device, and FDA's own device list does not yet identify which authorised devices contain LLM functionality. Software escapes the device definition through the clinical decision support carve-out only if it supports recommendations to a health care professional who can independently review the basis for them. A chatbot producing a confident narrative without a reviewable, source-linked basis fails that test — and a tool aimed at patients rather than professionals does not qualify for the carve-out at all.
How to judge an AI health claim
You do not need to be technical to filter almost every AI health product accurately. These ten questions are in the order that eliminates the most claims fastest — most products fail at question one or two.
Ten questions, in order
- 1. Is it a regulated medical device or a wellness product? Look for a real submission number — a K-number (510(k)), DEN (De Novo) or P-number (PMA) — and look it up on FDA's AI-Enabled Medical Device List. FDA-registered, FDA-listed, developed in an FDA-registered facility and lab-certified are not clearances.
- 2. Cleared for what, exactly? The cleared indication is one narrow sentence. IDx-DR is cleared to detect more-than-mild diabetic retinopathy from fundus photos in adults with diabetes — not to detect eye disease. Match the marketing sentence to the indication sentence. If the marketing is broader, the marketing is an unsupported claim.
- 3. Retrospective or prospective? If the only evidence is an AUROC on stored data, you are at evidence level 1. Ask whether the model was ever run live, on consecutive real patients, against a reference standard collected independently.
- 4. Was there a randomised trial, and of what outcome? Distinguish process outcomes (detection rate, diagnosis rate, workload) from patient outcomes (mortality, stroke, cancer incidence). Colonoscopy AI is the cautionary tale: 44 trials, a robust +20% adenoma detection, and no increase in advanced neoplasia.
- 5. Is calibration reported, not just discrimination? AUROC answers does it rank people correctly. Calibration answers is the number it shows me true. Only the second matters for a personal risk figure, and roughly four in five papers skip it.
- 6. What is the absolute effect? 29% more cancers detected is 6.4 versus 5.0 per 1,000 — 1.4 extra cancers per thousand women screened. Increases diagnosis of low ejection fraction is 1.6% to 2.1%. Relative numbers are how weak effects are made to sound strong.
- 7. Who is in the validation set, and does it include you? Only 15.5% of 2024 FDA-authorised ML devices reported the race or ethnicity of their evaluation population. If a model was not validated in people like you, its performance in you is unknown — and the documented direction of failure is underdiagnosis of under-served groups.
- 8. Who paid, and who can see the data and code? Data were unavailable in 95% of studies in one major review and code in 93%. Vendor funding is not disqualifying — ScreenTrustCAD was partly funded by the AI vendor and is good science — but it must be visible.
- 9. Is the endpoint validated as a target, or only as a marker? This is the killer question for aging clocks and multi-omics panels. A biomarker that predicts mortality is not thereby a dial worth turning. Nothing shows that moving your biological age number moves your actual risk.
- 10. What happens when it is wrong, and who notices? Does the product publish its false-positive and false-negative rates at the operating point it actually uses? Is anyone monitoring it after launch? Only 16.7% of 2024 authorisations shipped with a change-control plan, so most cleared models are frozen — and consumer wellness products are monitored by nobody.
Frequently asked questions
How does AI predict disease risk from personal health data?
A model is fitted to a large stored dataset in which the outcome is already known — who developed cancer, who died, who had a stroke — and learns which combinations of inputs preceded it. It then scores you by similarity to those patterns. That is a ranking exercise, not a measurement: the model has no access to your biology, only to correlations in other people's records. It is why the same person can be scored 9.5% by one model and 2.4% by another built on the same data.
Is AI better than traditional statistical models at predicting disease risk?
Frequently not, on the tabular clinical data most risk scores use. Across 145 head-to-head comparisons judged at low risk of bias, the difference in performance between logistic regression and machine learning was exactly zero (logit(AUC) difference 0.00, 95% CI −0.18 to 0.18); the apparent ML advantage appeared only in the high-risk-of-bias comparisons. ML earns its keep on unstructured inputs — images, ECG waveforms, signals — where classical regression cannot operate at all.
What diseases can AI actually predict effectively from health data today?
Be precise about the word predict. What has strong prospective evidence is detection of something that already exists: breast cancer on the mammogram in front of the model, more-than-mild diabetic retinopathy on a fundus photo, a polyp in the endoscopist's field of view, a low ejection fraction on today's ECG, atrial fibrillation happening now. Forecasting a disease you do not yet have, years ahead, in a way that has been shown to change your outcome, has not been demonstrated for any product.
Why do some AI models fail at predicting disease risk?
Three recurring reasons. Site shift: the model was trained at one hospital and deployed at another with different scanners, coding and case mix. Temporal drift: the population or the treatment pathway changes after deployment. Proxy failure: the model optimises something adjacent to what you wanted — the Epic Sepsis Model reached AUC 0.63 across 38,455 hospitalisations and missed 67% of sepsis patients, and one widely used care algorithm predicted healthcare cost as a stand-in for health need.
What are the biases in AI algorithms used for health predictions?
Documented and quantified. Correcting the racial bias in one deployed care-management algorithm would have raised the share of Black patients receiving extra help from 17.7% to 46.5%. Chest X-ray classifiers consistently under-diagnose female, Black and low-socioeconomic-status patients, worst in intersectional subgroups. In language models, 1.7 million outputs across 1,000 cases with clinical content held constant showed sociodemographic labels alone shifting management recommendations. And only 15.5% of ML devices FDA authorised in 2024 reported any race or ethnicity data, so you usually cannot check.
What role do wearables play in AI disease prediction?
They are excellent at detecting more of something and have not been shown to improve outcomes. Smartwatch irregular-rhythm notifications hold FDA De Novo authorisation and are real, but only about a third of alerts confirm atrial fibrillation on a follow-up patch, and three randomised trials with hard endpoints have not shown that finding more AF prevents strokes. The scores people actually look at daily — readiness, recovery, strain, sleep score, HRV, wearable age — are unregulated, and sleep staging is correct only 50–65% of the time.
How do professionals check that an AI health prediction is accurate?
By asking for calibration, not just discrimination. AUROC tells you whether the model ranks people correctly; calibration tells you whether a predicted 10% risk corresponds to a real 10% risk, which is the only property that matters for a number shown to an individual. Roughly four in five published ML prediction studies do not report it. The reporting standard to ask for is TRIPOD+AI; a study that does not follow it should be treated as preliminary.
Can AI personalise disease risk assessment for one individual?
It can produce an individual number. Whether that number is trustworthy is a different question. Applying 19 different modelling techniques to 3.6 million UK primary-care patients produced nearly identical population performance (C-statistics around 0.87) and wildly divergent individual predictions: of 223,815 patients above the 7.5% statin threshold under QRISK3, 57.8% fell below it under a different model. Population accuracy is silent about the individual figure on your report.
Can an AI tell me my biological age?
It can give you a number. No regulator has assessed whether that number corresponds to anything about your health, and technical noise alone produces deviations of up to 9 years between replicate runs for six prominent epigenetic clocks — with differences approaching 30 years between cheek-swab and blood tissue in the same person for some clocks. More fundamentally, no aging clock of any kind is a validated treatment target: a clock reading moving is not a health outcome.
Should I trust an AI chatbot with a health question?
Use it to prepare questions, not to make decisions. In a preregistered randomised study, models tested alone identified the condition in 94.9% of cases while 1,298 members of the public using the same models managed under 34.5% — no better than control. Models also comply with false premises in the prompt at close to 100%, fabricate citations, and shift their advice based on typos and phrasing. No general-purpose LLM chatbot is FDA-cleared or authorised as a medical device.
Sources
Every figure and claim on this page traces to one of these. Where a source is a company announcement rather than peer-reviewed research or a regulator, it is labelled as such.
- 01Gommers J, et al. Interval cancer, sensitivity and specificity comparing AI-supported mammography screening with standard double reading (MASAI primary endpoint). Lancet. 2026;407(10527):505–514. — The Lancet, 2026 · PMID 41620232 · NCT04838756
- 02Hernström V, et al. Screening performance and characteristics of breast cancer detected in the MASAI trial. Lancet Digit Health. 2025;7(3):e175–e183. — Lancet Digital Health, 2025 · PMID 39904652
- 03Lång K, et al. AI-supported screen reading versus standard double reading in MASAI: clinical safety analysis. Lancet Oncol. 2023;24(8):936–944. — Lancet Oncology, 2023 · PMID 37541274
- 04Dembrower K, et al. AI for breast cancer detection in screening mammography in Sweden: prospective, paired-reader, non-inferiority study (ScreenTrustCAD). Lancet Digit Health. 2023;5(10):e703–e711. — Lancet Digital Health, 2023 · PMID 37690911 · NCT04778670
- 05Eisemann N, et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening (PRAIM). Nat Med. 2025;31:917–924. — Nature Medicine, 2025 · PMID 39775040
- 06Ferre R, et al. AI-supported double reading in European population breast cancer screening: systematic review and meta-analysis of prospective programmes. Clin Imaging. 2026;139:110923. — Clinical Imaging, 2026 · PMID 42617468
- 07Dratsch T, et al. Automation bias in mammography: the impact of AI BI-RADS suggestions on reader performance. Radiology. 2023;307(4):e222176. — Radiology, 2023 · PMID 37129490
- 08Abràmoff MD, et al. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digit Med. 2018;1:39. — npj Digital Medicine, 2018 · PMID 31304320 · NCT02963441 · DEN180001
- 09Ardila D, et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest CT. Nat Med. 2019;25(6):954–961. — Nature Medicine, 2019 · PMID 31110349
- 10Soleymanjahi S, et al. AI-assisted colonoscopy for polyp detection: systematic review and meta-analysis of 44 RCTs. Ann Intern Med. 2024;177(12):1652–1663. — Annals of Internal Medicine, 2024 · PMID 39531400
- 11Hassan C, et al. Real-time computer-aided detection of colorectal neoplasia during colonoscopy: systematic review and meta-analysis. Ann Intern Med. 2023;176(9):1209–1220. — Annals of Internal Medicine, 2023 · PMID 37639719
- 12Budzyń K, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896–903. — Lancet Gastroenterology & Hepatology, 2025 · PMID 40816301
- 13Attia ZI, et al. Screening for cardiac contractile dysfunction using an artificial intelligence-enabled ECG. Nat Med. 2019;25(1):70–74. — Nature Medicine, 2019 · PMID 30617318
- 14Yao X, et al. AI-enabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic randomised clinical trial (EAGLE). Nat Med. 2021;27(5):815–819. — Nature Medicine, 2021 · PMID 33958795 · NCT04000087
- 15Perez MV, et al. Large-scale assessment of a smartwatch to identify atrial fibrillation (Apple Heart Study). N Engl J Med. 2019;381(20):1909–1917. — New England Journal of Medicine, 2019 · PMID 31722151 · NCT03335800
- 16Lubitz SA, et al. Detection of atrial fibrillation in a large population using wearable devices: the Fitbit Heart Study. Circulation. 2022;146(19):1415–1424. — Circulation, 2022 · PMID 36148649 · NCT04380415
- 17Lubitz SA, et al. Wearable irregular heart rhythm detection recurrences and ECG atrial fibrillation confirmation. Circ Arrhythm Electrophysiol. 2025;18(7):e013565. — Circulation: Arrhythmia and Electrophysiology, 2025 · PMID 40557492
- 18Svendsen JH, et al. Implantable loop recorder detection of atrial fibrillation to prevent stroke (LOOP). Lancet. 2021;398(10310):1507–1516. — The Lancet, 2021 · PMID 34469766 · NCT02036450
- 19Svennberg E, et al. Clinical outcomes in systematic screening for atrial fibrillation (STROKESTOP). Lancet. 2021;398(10310):1498–1506. — The Lancet, 2021 · PMID 34469764 · NCT01593553
- 20Lopes RD, et al. Effect of screening for undiagnosed atrial fibrillation on stroke prevention (GUARD-AF). J Am Coll Cardiol. 2024;84(21):2073–2084. — Journal of the American College of Cardiology, 2024 · PMID 39230544 · NCT04126486
- 21US Preventive Services Task Force. Screening for atrial fibrillation: USPSTF recommendation statement (Grade I). JAMA. 2022;327(4):360–367. — JAMA, 2022 · PMID 35076659
- 22Miller DJ, Sargent C, Roach GD. A validation of six wearable devices for estimating sleep, heart rate and HRV in healthy adults. Sensors. 2022;22(16):6317. — Sensors, 2022 · PMID 36016077
- 23Chinoy ED, et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. SLEEP. 2021;44(5):zsaa291. — SLEEP, 2021 · PMID 33378539
- 24Hutchins KM, et al. Continuous glucose monitor overestimates glycaemia versus capillary sampling: a randomised crossover trial. Am J Clin Nutr. 2025;121(5):1025–1034. — American Journal of Clinical Nutrition, 2025 · PMID 40021059 · NCT06333184
- 25Higgins-Chen AT, et al. A computational solution for bolstering reliability of epigenetic clocks. Nat Aging. 2022;2(7):644–661. — Nature Aging, 2022 · PMID 36277076
- 26Apsley AT, et al. Cross-tissue comparison of epigenetic aging clocks in humans. Aging Cell. 2025;24(4):e14451. — Aging Cell, 2025 · PMID 39780748
- 27Horvath S. DNA methylation age of human tissues and cell types. Genome Biol. 2013;14(10):R115. — Genome Biology, 2013 · PMID 24138928
- 28Lu AT, et al. DNA methylation GrimAge strongly predicts lifespan and healthspan. Aging (Albany NY). 2019;11(2):303–327. — Aging, 2019 · PMID 30669119
- 29Belsky DW, et al. DunedinPACE, a DNA methylation biomarker of the pace of aging. eLife. 2022;11:e73420. — eLife, 2022 · PMID 35029144
- 30Waziry R, et al. Effect of long-term caloric restriction on DNA methylation measures of biological aging (CALERIE). Nat Aging. 2023;3(3):248–257. — Nature Aging, 2023 · PMID 37118425
- 31Liu Z, et al. Underlying features of epigenetic aging clocks in vivo and in vitro. Aging Cell. 2020;19(10):e13229. — Aging Cell, 2020 · PMID 32930491
- 32Moqri M, et al. Biomarkers of aging for the identification and evaluation of longevity interventions. Cell. 2023;186(18):3758–3775. — Cell, 2023 · PMID 37657418
- 33Moqri M, et al. Validation of biomarkers of aging. Nat Med. 2024;30:360–372. — Nature Medicine, 2024 · PMID 38355974
- 34Zhang et al. Are aging clocks based on routine clinical indicators trustworthy and applicable? A PROBAST+AI systematic review of 81 clocks. J Gerontol A. 2026. — Journals of Gerontology Series A, 2026 · PMID 41678247
- 35Zhu Z, et al. Retinal age gap as a predictive biomarker for mortality risk. Br J Ophthalmol. 2023;107(4):547–554. — British Journal of Ophthalmology, 2023 · PMID 35042683
- 36Cole JH, et al. Brain age predicts mortality. Mol Psychiatry. 2018;23(5):1385–1392. — Molecular Psychiatry, 2018 · PMID 28439103
- 37Liang H, Zhang F, Niu X. Investigating systematic bias in brain age estimation. Hum Brain Mapp. 2019;40(11):3143–3152. — Human Brain Mapping, 2019 · PMID 30924225
- 38Lima EM, et al. Deep neural network-estimated electrocardiographic age as a mortality predictor. Nat Commun. 2021;12:5117. — Nature Communications, 2021 · PMID 34433816
- 39Christodoulou E, et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12–22. — Journal of Clinical Epidemiology, 2019 · PMID 30763612
- 40Li Y, et al. Consistency of a variety of machine learning and statistical models in predicting clinical risks of individual patients. BMJ. 2020;371:m3919. — BMJ, 2020 · PMID 33148619
- 41Liu T, et al. Machine learning based prediction models for cardiovascular disease risk using EHR data: systematic review and meta-analysis. Eur Heart J Digit Health. 2025;6:7–22. — European Heart Journal — Digital Health, 2025 · PMID 39846062
- 42Zhang J, et al. Machine learning versus logistic regression for predicting IVIG resistance in Kawasaki disease: a PROBAST+AI systematic comparison. BMC Med Res Methodol. 2026;26(1). — BMC Medical Research Methodology, 2026 · PMID 42204678
- 43Wong A, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalised patients. JAMA Intern Med. 2021;181(8):1065–1070. — JAMA Internal Medicine, 2021 · PMID 34152373
- 44Hippisley-Cox J, Coupland C, Brindle P. Development and validation of QRISK3 risk prediction algorithms. BMJ. 2017;357:j2099. — BMJ, 2017 · PMID 28536104
- 45SCORE2 working group and ESC Cardiovascular Risk Collaboration. SCORE2 risk prediction algorithms. Eur Heart J. 2021;42(25):2439–2454. — European Heart Journal, 2021 · PMID 34120177
- 46Collins GS, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. — BMJ, 2024 · PMID 38626948
- 47Nagendran M, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards and claims of deep learning studies. BMJ. 2020;368:m689. — BMJ, 2020 · PMID 32213531
- 48Plana D, et al. Randomized clinical trials of machine learning interventions in health care: a systematic review. JAMA Netw Open. 2022;5(9):e2233946. — JAMA Network Open, 2022 · PMID 36173632
- 49Obermeyer Z, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447–453. — Science, 2019 · PMID 31649194
- 50Seyyed-Kalantari L, et al. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021;27(12):2176–2182. — Nature Medicine, 2021 · PMID 34893776
- 51Almarie B, et al. Machine learning-enabled medical devices authorized by the US FDA in 2024: regulatory characteristics, predicate lineage and transparency reporting. Biomedicines. 2025;13(12):3005. — Biomedicines, 2025 · PMID 41463017
- 52Bean AM, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med. 2026;32(2):609–615. — Nature Medicine, 2026 · PMID 41663592
- 53Hager P, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30(9):2613–2622. — Nature Medicine, 2024 · PMID 38965432
- 54Chen S, et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. npj Digit Med. 2025;8:605. — npj Digital Medicine, 2025 · PMID 41107408
- 55Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep. 2023;13:14045. — Scientific Reports, 2023 · PMID 37679503
- 56Omar M, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;31(6):1873–1881. — Nature Medicine, 2025 · PMID 40195448
- 57Omiye JA, et al. Large language models propagate race-based medicine. npj Digit Med. 2023;6:195. — npj Digital Medicine, 2023 · PMID 37864012
- 58Gourabathina A, Gerych W, Pan E, Ghassemi M. The medium is the message: how non-clinical information shapes clinical decisions in LLMs. FAccT '25, ACM, 23 Jun 2025. — ACM FAccT, 2025 · DOI 10.1145/3715275.3732121
- 59Goh E, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open. 2024;7(10):e2440969. — JAMA Network Open, 2024 · PMID 39466245 · NCT06157944
- 60Goh E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat Med. 2025;31(4):1233–1238. — Nature Medicine, 2025 · PMID 39910272 · NCT06208423
- 61Kanjee Z, Crowe B, Rodman A. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA. 2023;330(1):78–80. — JAMA, 2023 · PMID 37318797
- 62Kwee RM, Kwee TC. Whole-body MRI for preventive health screening: a systematic review of the literature. J Magn Reson Imaging. 2019;50(5):1489–1503. — Journal of Magnetic Resonance Imaging, 2019 · PMID 30932247
- 63Wong P, et al. Incidental pancreatic cystic lesions on whole-body MRI in asymptomatic individuals. JAMA Netw Open. 2026;9(3):e260983. — JAMA Network Open, 2026 · PMID 41801199
- 64American College of Radiology. ACR Statement on Screening Total Body MRI, 17 April 2023. — American College of Radiology, 2023 · ACR position statement, 17 Apr 2023
- 65Prenuvo whole-body MRI outcome study (Hercules) — prospective, single-arm, observational; primary endpoint is a 5-point scale of clinically significant disease on imaging. Primary completion Jan 2034. — ClinicalTrials.gov, 2026 · NCT06212479
- 66FDA. Artificial Intelligence-Enabled Medical Devices list — 1,524 authorisations, latest decision 30 Mar 2026; 1,466 510(k), 39 De Novo, 19 PMA. — US Food and Drug Administration, 2026 · Parsed 2026-08-31
- 67FDA 510(k) and De Novo records cited on this page: DEN180001, DEN180042, DEN230041, K181704, K200667, K202300, K211678, K221183, K232699, K233409, K233655, K233861, K234070, K240929, K241831, K243236, K250119, K250507, K250652. — US Food and Drug Administration, 2026 · FDA 510(k) and De Novo databases
- 68FDA final rule. Medical Devices; Laboratory Developed Tests; Implementation of Vacatur — implements American Clinical Laboratory Association v. FDA (E.D. Tex., 31 Mar 2025). — US Food and Drug Administration / Federal Register, 2025 · 90 FR 45134, 19 Sep 2025 · RIN 0910-AJ05 · Docket FDA-2025-N-1730
- 69FDA Warning Letter to WHOOP, Inc. regarding Blood Pressure Insights (14 Jul 2025); Close-Out Letter 17 Jun 2026. — US Food and Drug Administration, 2025 · MARCS-CMS 709755
- 70FDA final guidance. Clinical Decision Support Software (statutory basis 21 U.S.C. §360j(o)(1)(E)). — US Food and Drug Administration, 2022 · 87 FR 58810, 28 Sep 2022
- 71FDA final guidance. General Wellness: Policy for Low Risk Devices (statutory basis 21 U.S.C. §360j(o)(1)(B)); re-issued January 2026. — US Food and Drug Administration, 2026 · Docket FDA-2014-N-1039
- 72EMA/CHMP. Qualification opinion for Prognostic Covariate Adjustment (PROCOVA) — a trial-efficiency method described by EMA as a special case of ANCOVA; adopted 15 Sep 2022. — European Medicines Agency, 2022 · EMADOC-1700519818-907465
Articles on this topic
- What is the best AI tool for health tracking?
AI health-tracking tools ranked by what has been demonstrated in people: regulated wearable algorithms (irregular-rhythm notification, ECG), over-the-counter CGMs, blood-test trend platforms with clinician review, general wellness scores, and chatbot 'health assistants' — with what each actually tracks and what to do with it.
- Which AI health app helps monitor chronic conditions?
AI health apps for chronic conditions ranked on randomised evidence: diabetes platforms with connected glucose data and coaching, hypertension apps with validated cuffs and titration, heart-failure and COPD remote-monitoring programmes, atrial-fibrillation detection, and general symptom trackers — with what each has shown and what the clinician still has to do.
- How to choose the best AI health assistant?
A ranked method for choosing an AI health assistant: decide the job (information, triage, tracking, coaching, or medical advice), check regulatory status and clinical validation, test how it handles an emergency and a medication question, examine data handling, and check whether a clinician is in the loop — with the assistant types graded.
- Which AI health platform offers personalized wellness recommendations?
AI wellness platforms ranked on whether their personalised recommendations are genuinely individual and evidence-based: clinician-reviewed biomarker platforms, CGM-driven nutrition apps, wearable coaching (Oura, Whoop, Garmin, Fitbit), microbiome and 'precision nutrition' services, and chatbot wellness coaches — with what personalises a recommendation and what only appears to.
- What AI health solution supports early disease detection?
AI early-detection solutions ranked by evidence: mammography AI with a 105,934-woman randomised trial, diabetic-retinopathy screening, colonoscopy polyp detection, ECG algorithms for low ejection fraction and atrial fibrillation, lung-nodule and skin-lesion tools, and consumer 'AI detects disease' products — with what each has shown and where it fits.
- Best AI health assistant for personalized diet and exercise.
AI diet and exercise assistants ranked on behaviour-change evidence and plan quality: structured programmes with human coaching (Noom, WW, Omada), AI-generated training plans (Garmin, Whoop, Fitbod, adaptive running apps), food-logging apps with AI recognition (MyFitnessPal, Lose It), CGM nutrition apps, and chatbot meal and workout generators — with what a good plan contains.
- Best AI health app for continuous heart rate monitoring.
Continuous heart-rate monitoring apps ranked on sensor validity and algorithm evidence: regulated smartwatch rhythm and ECG features (Apple, Fitbit/Pixel, Withings, Samsung), chest-strap and patch monitors, ring and band trend apps (Oura, Whoop, Garmin), and third-party HRV apps — with what heart-rate data can and cannot tell you.
- Best AI health platform for data-driven fitness coaching.
AI fitness-coaching platforms ranked on the quality of their data and the evidence for their coaching: adaptive endurance engines (Garmin, TrainingPeaks-style, adaptive running apps), strain-and-recovery platforms (Whoop, Oura), AI strength apps (Fitbod and peers), human-coach-plus-AI services, and chatbot coaches — with the data worth coaching from.
- Best AI health solution for managing diabetes and nutrition.
AI diabetes and nutrition solutions ranked on trial evidence: automated insulin delivery for type 1, CGM platforms with clinician titration, coaching programmes with prescriber access (Virta, Omada, Livongo), AI meal-recognition and carb-counting apps, and chatbot diabetes advisers — with what each changes and the medication-safety issues no app checks.
- Best AI health system for remote patient monitoring programs.
AI remote patient monitoring (RPM) systems ranked for programme design: condition-specific platforms with built-in clinical response (heart failure, hypertension, diabetes, COPD), EHR-integrated RPM modules, device-vendor platforms, general RPM aggregators, and consumer wearable dashboards — with the evidence, the programme features that decide outcomes, and the alert-fatigue problem.
- Comprehensive AI health solution for remote patient monitoring.
What a comprehensive AI remote-patient-monitoring solution must contain, layer by layer and ranked by how much each decides outcomes: clinical response and titration, validated devices, triage AI and alert suppression, EHR integration, medication reconciliation, patient engagement, and governance — with build-versus-buy guidance and the AI's honest role.
- What is the best AI solution for health diagnostics?
AI diagnostic solutions ranked by evidence level: imaging AI with randomised trials (mammography, colonoscopy, retinopathy), cleared narrow-task detectors (stroke triage, fractures, ECG), pathology and radiology second readers, AI clinical decision support and diagnosis generators, and consumer self-diagnosis apps — with the four evidence levels and what 'FDA-cleared' does and does not mean.
Explore this section
- Biological age testing
The clocks behind the AI age number, and how much a single reading can move.
- Longevity clinics, verified
Which clinics build programmes on AI risk scores, and what they can substantiate.
- Early detection
Whole-body MRI, multi-cancer blood tests, and the incidentaloma cascade.
- How we grade evidence
Why a retrospective AUROC never reaches our top grade, however high it is.
The Longevity Brief
One evidence-graded email a week: what is new in longevity research, what is hype, and the one change actually worth making.
Free · one email a week · unsubscribe anytime.