REVIEW 10 cited by
GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
There is an increasing interest in developing artificial intelligence (AI) systems to process and interpret electronic health records (EHRs). Natural language processing (NLP) powered by pretrained language models is the key technology for medical AI systems utilizing clinical narratives. However, there are few clinical language models, the largest of which trained in the clinical domain is comparatively small at 110 million parameters (compared with billions of parameters in the general domain). It is not clear how large clinical language models with billions of parameters can help medical AI systems utilize unstructured EHRs. In this study, we develop from scratch a large clinical language model - GatorTron - using >90 billion words of text (including >82 billion words of de-identified clinical text) and systematically evaluate it on 5 clinical NLP tasks including clinical concept extraction, medical relation extraction, semantic textual similarity, natural language inference (NLI), and medical question answering (MQA). We examine how (1) scaling up the number of parameters and (2) scaling up the size of the training data could benefit these NLP tasks. GatorTron models scale up the clinical language model from 110 million to 8.9 billion parameters and improve 5 clinical NLP tasks (e.g., 9.6% and 9.5% improvement in accuracy for NLI and MQA), which can be applied to medical AI systems to improve healthcare delivery. The GatorTron models are publicly available at: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/clara/models/gatortron_og.
Forward citations
Cited by 10 Pith papers
-
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
MedPath combines 513k+ expert-annotated biomedical mentions into a UMLS-normalized dataset with cross-vocabulary mappings and hierarchical paths for 11 vocabularies.
-
Towards Domain Specification of Embedding Models in Medicine
A contrastively fine-tuned GTE model (MedTE) trained on two million medical text pairs scores highest on a new 51-task medical embedding benchmark (MedTEB) built from the same data sources.
-
Abstract Meaning Representation for Hospital Discharge Summarization
An AMR-graph-based extractive pipeline generates traceable hospital discharge summaries, but its section classifier and human evaluation show poor coverage and organization.
-
Enhancing Clinical Models with Pseudo Data for De-identification
Continued pretraining on masked and pseudo-replaced MIMIC-III notes yields strong de-identification models, with masked RoBERTa large performing best and pseudo data helping XLM-RoBERTa more than RoBERTa.
-
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reinforcement learning on questions extracted from CRISPR expert forums improves LLM accuracy on a new benchmark (Genome-Bench) by over 15 percentage points.
-
From EMR Data to Clinical Insight: An LLM-Driven Framework for Automated Pre-Consultation Questionnaire Generation
A multi-stage LLM framework using atomic assertions and clustered causal networks generates pre-consultation questionnaires from EMRs, reporting 84.2% personal key-fact coverage versus 42.1% for direct LLM prompting.
-
Large Language Models as Unified Multimodal Learners for Clinical Prediction
Serializing all patient data — notes, vitals, labs — into one text sequence and fine-tuning an LLM matches or beats task-specific multimodal fusion baselines on mortality, graft-failure, and triage prediction.
-
Building Models of Neurological Language
A neurology NLP project moved from custom model training to RAG with small Gemma models, reporting QA results that are undermined by a contaminated benchmark and internally contradictory claims.
-
The Latent Space Hypothesis: Toward Universal Medical Representation Learning
The paper argues that all medical data modalities encode projections of a single latent physiological state, so a universal learned geometry could unify diagnosis, monitoring, and treatment.
-
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
MedOrch is a modular framework in which LLMs call medical tools to answer clinical questions; its headline results on Alzheimer's, chest X-ray, and VQA benchmarks are weakened by best-of-five scoring.
Discussion (0). Sign in to comment.