REVIEW 9 cited by
Towards a Personal Health Large Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In health, most large language model (LLM) research has focused on clinical tasks. However, mobile and wearable devices, which are rarely integrated into such tasks, provide rich, longitudinal data for personal health monitoring. Here we present Personal Health Large Language Model (PH-LLM), fine-tuned from Gemini for understanding and reasoning over numerical time-series personal health data. We created and curated three datasets that test 1) production of personalized insights and recommendations from sleep patterns, physical activity, and physiological responses, 2) expert domain knowledge, and 3) prediction of self-reported sleep outcomes. For the first task we designed 857 case studies in collaboration with domain experts to assess real-world scenarios in sleep and fitness. Through comprehensive evaluation of domain-specific rubrics, we observed that Gemini Ultra 1.0 and PH-LLM are not statistically different from expert performance in fitness and, while experts remain superior for sleep, fine-tuning PH-LLM provided significant improvements in using relevant domain knowledge and personalizing information for sleep insights. We evaluated PH-LLM domain knowledge using multiple choice sleep medicine and fitness examinations. PH-LLM achieved 79% on sleep and 88% on fitness, exceeding average scores from a sample of human experts. Finally, we trained PH-LLM to predict self-reported sleep quality outcomes from textual and multimodal encoding representations of wearable data, and demonstrate that multimodal encoding is required to match performance of specialized discriminative models. Although further development and evaluation are necessary in the safety-critical personal health domain, these results demonstrate both the broad knowledge and capabilities of Gemini models and the benefit of contextualizing physiological data for personal health applications as done with PH-LLM.
Forward citations
Cited by 9 Pith papers
-
SleepLM: Natural-Language Intelligence for Human Sleep
A sleep-language foundation model trained with contrastive, captioning, and reconstruction objectives outperforms general LLMs and fine-tuned VLMs on zero-shot sleep staging, event localization, and cross-modal retrieval.
-
OpenMHC: Accelerating the Science of Wearable Foundation Models
OpenMHC contributes the largest open-access consumer wearable dataset to date (67M hours, 11,894 participants), a standardized three-track benchmark, and the first open implementations of Apple WBM and Google LSM-2.
-
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
State-of-the-art LLMs are frequently inaccurate, and sometimes dangerous, when answering harm reduction questions about drug use, even when given retrieved source material.
-
GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing
A group of LLM agents that collaboratively generate code for raw passive sensing data outperforms RAG on objective query accuracy, while remaining only moderately consistent across repeated runs.
-
SensorLM: Learning the Language of Wearable Sensors
SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.
-
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.
-
ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data
ProMind-LLM integrates objective sensor data and subjective mental records via domain-specific training, self-refine formatting, and causal chain-of-thought prompting to improve LLM mental health risk classification.
-
SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching
SePA combines personalized wearable-data risk prediction with a whitelisted web search pipeline to give cited, context-aware health coaching.
-
Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
CAMA applies a three-step feature engineering procedure to LLM agents and reports 55 to 92 percent accuracy on ML monitoring report questions, outperforming six baselines.
Discussion (0). Continue with ORCID to comment.