Pith. sign in

REVIEW 9 cited by

Towards a Personal Health Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06474 v1 pith:ALBS3363 submitted 2024-06-10 cs.AI cs.CL

classification cs.AIcs.CL
keywords sleephealthph-llmpersonaldomaindatafitnessknowledge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In health, most large language model (LLM) research has focused on clinical tasks. However, mobile and wearable devices, which are rarely integrated into such tasks, provide rich, longitudinal data for personal health monitoring. Here we present Personal Health Large Language Model (PH-LLM), fine-tuned from Gemini for understanding and reasoning over numerical time-series personal health data. We created and curated three datasets that test 1) production of personalized insights and recommendations from sleep patterns, physical activity, and physiological responses, 2) expert domain knowledge, and 3) prediction of self-reported sleep outcomes. For the first task we designed 857 case studies in collaboration with domain experts to assess real-world scenarios in sleep and fitness. Through comprehensive evaluation of domain-specific rubrics, we observed that Gemini Ultra 1.0 and PH-LLM are not statistically different from expert performance in fitness and, while experts remain superior for sleep, fine-tuning PH-LLM provided significant improvements in using relevant domain knowledge and personalizing information for sleep insights. We evaluated PH-LLM domain knowledge using multiple choice sleep medicine and fitness examinations. PH-LLM achieved 79% on sleep and 88% on fitness, exceeding average scores from a sample of human experts. Finally, we trained PH-LLM to predict self-reported sleep quality outcomes from textual and multimodal encoding representations of wearable data, and demonstrate that multimodal encoding is required to match performance of specialized discriminative models. Although further development and evaluation are necessary in the safety-critical personal health domain, these results demonstrate both the broad knowledge and capabilities of Gemini models and the benefit of contextualizing physiological data for personal health applications as done with PH-LLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. SleepLM: Natural-Language Intelligence for Human Sleep

    cs.AI 2026-02 conditional novelty 7.0 of 10

    A sleep-language foundation model trained with contrastive, captioning, and reconstruction objectives outperforms general LLMs and fine-tuned VLMs on zero-shot sleep staging, event localization, and cross-modal retrieval.

  2. OpenMHC: Accelerating the Science of Wearable Foundation Models

    cs.LG 2026-06 conditional novelty 6.0 of 10

    OpenMHC contributes the largest open-access consumer wearable dataset to date (67M hours, 11,894 participants), a standardized three-track benchmark, and the first open implementations of Apple WBM and Google LSM-2.

  3. HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    State-of-the-art LLMs are frequently inaccurate, and sometimes dangerous, when answering harm reduction questions about drug use, even when given retrieved source material.

  4. GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing

    cs.HC 2025-07 conditional novelty 6.0 of 10

    A group of LLM agents that collaboratively generate code for raw passive sensing data outperforms RAG on objective query accuracy, while remaining only moderately consistent across repeated runs.

  5. SensorLM: Learning the Language of Wearable Sensors

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.

  6. RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.

  7. ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data

    cs.AI 2025-05 conditional novelty 6.0 of 10

    ProMind-LLM integrates objective sensor data and subjective mental records via domain-specific training, self-refine formatting, and causal chain-of-thought prompting to improve LLM mental health risk classification.

  8. SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching

    cs.HC 2025-09 conditional novelty 5.0 of 10

    SePA combines personalized wearable-data risk prediction with a whitelisted web search pipeline to give cited, context-aware health coaching.

  9. Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring

    cs.LG 2025-06 reject novelty 5.0 of 10

    CAMA applies a three-step feature engineering procedure to LLM agents and reports 55 to 92 percent accuracy on ML monitoring report questions, outperforming six baselines.

Pith tools