REVIEW 10 cited by
MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces MedMCQA, a new large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions. More than 194k high-quality AIIMS \& NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity. Each sample contains a question, correct answer(s), and other options which requires a deeper language understanding as it tests the 10+ reasoning abilities of a model across a wide range of medical subjects \& topics. A detailed explanation of the solution, along with the above information, is provided in this study.
Forward citations
Cited by 10 Pith papers
-
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.
-
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
An NLI-hypergraph audit with deterministic AND–OR search labels LLM reasoning segments as supported, unsupported, or orphaned, improving balanced F1 over LLM-as-judge on a new 40-case clinical benchmark.
-
Cura 1T: Specialized Model for Agentic Healthcare
A healthcare LLM trained with an agent-driven loop that converts benchmark failures into new training data outperforms frontier baselines on most healthcare tests while staying competitive on general reasoning.
-
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.
-
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
A new bilingual dental QA benchmark and corpus reveals large performance gaps in LLMs for dentistry, and shows that domain adaptation with the corpus improves accuracy.
-
HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models
On a new Portuguese healthcare benchmark, leading language models score high overall but drop sharply in several specialties, such as neurosurgery and social work.
-
MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
A multimodal offline RL framework that fuses LLM-encoded clinical notes with structured vitals and labs modestly improves sepsis policy scores on some datasets, with evaluation caveats.
-
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
AutoEvoEval applies 22 atomic perturbations and multi-round chains to MCQ benchmarks, causing average accuracy drops of 7.283% and up to 52.932% for long chains.
-
Enterprise Large Language Model Evaluation Benchmark
A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...
-
A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment
GraphMPA combines an embedding-similarity hierarchical graph with mode-seeking preference optimization to improve RAG question answering on six datasets.
Discussion (0). Sign in to comment.