NPJ digital medicine

Comparing large language models for personalized health advice based on biomarkers

Updated

Abstract

Essence

Current large language models gave uneven and often inadequate personalized biomarker-based longevity intervention recommendations, though proprietary models generally performed better than open-source ones.

Evidence

This benchmarking study used 25 biomarker profiles across three age groups, 1000 test cases, and 56000 judged responses in an extended framework to compare recommendations for interventions such as caloric restriction, fasting, and supplements.

Caveat

The results come from a simulated benchmark using LLM-as-a-Judge with clinician-validated ground truths, not from real clinical deployment or patient outcomes.

Simplified

Key numbers

0.85
of Proprietary Models
across validation requirements for proprietary models.
0.20
of Open-Source Models
across validation requirements for open-source models.
1000
Test Cases Generated
Total number of diverse test cases generated for benchmarking.

Key figures

Fig. 1
Benchmark development and evaluation process for testing large language models on health recommendations
Frames a comprehensive, multi-step benchmark process highlighting extensive evaluation of health recommendations
41746_2025_1996_Fig1_HTML
  • Panels top left
    Initial test items include diverse profiles such as young, , healthy, and heart disease categories
  • Panels top center
    Three rounds of physician review with expert commentary refine test items, including rephrasing after the second round
  • Panels top right
    Final test items are integrated into the framework after the third physician review
  • Panels middle
    Execution involves 7 models generating 56,000 responses across 1,000 item variations with or without ()
  • Panels bottom left
    include yes/no answers, keywords, and expert commentary used for evaluation
  • Panels bottom center
    system evaluates responses on five validation criteria: comprehensive, correct, , explainable, and safe
  • Panels bottom right
    Evaluation produces 280,000 scores visualized as box and violin plots for each validation criterion
Fig. 2
Performance accuracy of large language models across different validation requirements
Highlights varied model performance with notably lower accuracy in and higher accuracy in toxicity consideration
41746_2025_1996_Fig2_HTML
  • Panel a
    of each model across all validation requirements with individual data points and error bars
  • Panel b
    Mean balanced accuracy across all models for each validation requirement, showing lowest accuracy in Comprehensiveness and highest in
  • Panel c
    Mean balanced accuracy per model for each validation requirement without (), with GPT-4o mini showing highest scores
  • Panel d
    Mean balanced accuracy per model for each validation requirement with RAG applied, generally showing improved scores compared to panel c
Fig. 3
accuracy across system prompts, age groups, and diseases in personalized health recommendations
Highlights higher accuracy in age group and improved performance with across system prompts and diseases
41746_2025_1996_Fig3_HTML
  • Panel a
    of LLMs across five system prompts without (RAG); GPT-4o mini and Llama3 Med42 8B show higher accuracy, with GPT-4o mini reaching up to 0.68
  • Panel b
    Mean balanced accuracy of LLMs across five system prompts with RAG; overall accuracy increases, with GPT-4o mini reaching up to 0.65
  • Panel c
    LLM accuracy distribution across three age groups without RAG; all models perform significantly better for geriatric individuals, with GPT-4o mini reaching 0.78
  • Panel d
    LLM accuracy distribution across three age groups with RAG; accuracy improves across all groups, especially for geriatric group where GPT-4o mini reaches 0.79
  • Panel e
    LLM accuracy distribution across six diseases without RAG; scores increase for like Osteoporosis/Sarcopenia, with GPT-4o mini reaching 0.79
Fig. 4
Human rater vs -based judge: model accuracies and alignment scores
Highlights higher human rater accuracy and stronger alignment scores, spotlighting differences in model evaluation consistency.
41746_2025_1996_Fig4_HTML
  • Panel a
    Mean balanced accuracies for models across five validation requirements: , , Usefulness, , and ; human rater accuracies appear higher than LLM-based judge accuracies for Safe and Explnbl.
  • Panel b
    Overall accuracies per model for six LLMs; scores show alignment between human rater and LLM-based judge, with kappa values generally higher than individual accuracies.
1 / 4

Full Text

What this is

  • The study benchmarks large language models () for generating personalized health intervention recommendations based on biomarker profiles.
  • It evaluates ' performance across five validation requirements using a framework called .
  • Findings reveal that proprietary models outperform open-source ones, but all models have limitations in addressing medical validation requirements.

Essence

  • Proprietary perform better than open-source models in generating personalized longevity intervention recommendations, but all models struggle with key medical validation requirements.

Key takeaways

  • Proprietary models showed higher performance in comprehensiveness compared to open-source models, indicating a disparity in their ability to generate thorough recommendations.
  • Age-related performance bias was observed, with models more accurately identifying common degenerative diseases than rare hormonal conditions, reflecting the prevalence of diseases in the test cases.
  • The study emphasizes the need for caution when using for unsupervised medical intervention recommendations due to inconsistent accuracy across validation requirements.

Caveats

  • The benchmark utilized synthetic medical profiles, which may limit the generalizability of the findings to real-world scenarios.
  • The small sample size of 25 test items may not capture the full diversity of patient profiles and interventions.
  • Automated judgments by the -as-a-Judge may introduce model-specific biases, necessitating further human evaluations for consistency.

Definitions

  • LLM: Large language model, a type of AI designed to understand and generate human language.
  • RAG: Retrieval-Augmented Generation, a technique that enhances model responses by incorporating external data.
  • BioChatter: An open-source framework for benchmarking LLMs in generating health intervention recommendations.

Simplified

Funding

Competing interests

Competing interests: B.K.K. reports a relationship with Ponce de Leon Health that includes: consulting or advisory and equity or stocks. C.B. has received lecturing fees from Novartis Deutschland GmbH and Bayer Vital GmbH. C.B. serves on the expert board for statutory health insurance data of IQTIG, the Institute for Quality and Transparency in German Healthcare (Institut für Qualitätssicherung und Transparenz im Gesundheitswesen). G.F. is a consultant to BlueZoneTech GmbH, who distribute supplements. Statement on the use of AI: The first draft was written by H.J., with help from G.F. and S.L.; No writing assistance was employed. While the topic of the paper is the use of generative AI/LLMs, no such tools were used to generate text or content of the manuscript. GPT4o was used for copy-editing (grammar, spelling) assistance and research queries on related work and references.
PubMed

What Lands in Your Inbox Each Week:

  • 📚7 fresh studies
  • 📝plain-language summaries
  • direct links to original studies
  • 🏅top journal indicators
  • 📅weekly delivery
  • 🧘‍♂️always free