Pith. sign in

REVIEW 13 cited by

CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.09167 v3 pith:LYV7NO4E submitted 2020-04-20 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords annotationslabelingreportexpertmedicalbertchexbertlabeler
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The extraction of labels from radiology text reports enables large-scale training of medical imaging models. Existing approaches to report labeling typically rely either on sophisticated feature engineering based on medical domain knowledge or manual annotations by experts. In this work, we introduce a BERT-based approach to medical image report labeling that exploits both the scale of available rule-based systems and the quality of expert annotations. We demonstrate superior performance of a biomedically pretrained BERT model first trained on annotations of a rule-based labeler and then finetuned on a small set of expert annotations augmented with automated backtranslation. We find that our final model, CheXbert, is able to outperform the previous best rules-based labeler with statistical significance, setting a new SOTA for report labeling on one of the largest datasets of chest x-rays.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

    cs.CV 2026-01 conditional novelty 7.0 of 10

    AnatomiX, a two-stage anatomy-first multimodal LLM for chest X-ray interpretation, reports >25% relative gains on anatomy grounding and grounded captioning, but some aggregate benchmark numbers are internally inconsis...

  2. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  3. NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

    cs.NE 2026-08 conditional novelty 6.0 of 10

    NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...

  4. Scaling medical imaging report generation with multimodal reinforcement learning

    cs.CV 2026-01 conditional novelty 6.0 of 10

    UniRG-CXR, a Qwen3-VL-8B model trained with SFT plus GRPO reinforcement learning that directly optimizes the ReXrank metric components, reports state-of-the-art 1/RadCliQ-v1 results on all four ReXrank chest X-ray dat...

  5. Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Domain-adapted LLM encoders trained with masked token prediction and supervised contrastive learning improve chest X-ray image-text retrieval and external generalization, reaching GREEN scores of 0.308 on MIMIC-CXR an...

  6. Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis

    cs.CV 2025-07 reject novelty 6.0 of 10

    RadGazeIntent, a transformer model, predicts per-fixation diagnostic intention from radiologist gaze on chest X-rays, evaluated on three newly constructed intention-labeled datasets.

  7. Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation

    stat.ME 2025-07 conditional novelty 6.0 of 10

    REVTAF, a retrieval-augmented radiology report generator, reports average gains of 7.4 points on MIMIC-CXR and 2.9 points on IU X-Ray across nine metrics.

  8. Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 8-stage chest X-ray VQA benchmark and a context-aware model trained on it.

  9. Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Medical CLIP training with text, clinical, and graph soft labels plus negation hard negatives improves chest X-ray zero-shot and fine-tuned performance.

  10. RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores

    cs.CL 2025-08 conditional novelty 5.0 of 10

    RadReason trains a 7B language model with GRPO to output six radiology error sub-scores plus textual reasons, reporting Kendall tau 0.730 on ReXVal, best among offline metrics.

  11. RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

    cs.CV 2025-07 reject novelty 5.0 of 10

    A video-based eye-gaze prompt improved report generation and diagnosis for one general-purpose vision-language model, LLaVA-OneVision, but hurt or barely helped two others, and the main comparison to medical models re...

  12. Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.

  13. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

Pith tools