REVIEW 11 cited by
MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces MedMCQA, a new large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions. More than 194k high-quality AIIMS \& NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity. Each sample contains a question, correct answer(s), and other options which requires a deeper language understanding as it tests the 10+ reasoning abilities of a model across a wide range of medical subjects \& topics. A detailed explanation of the solution, along with the above information, is provided in this study.
Forward citations
Cited by 11 Pith papers
-
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.
-
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
An NLI-hypergraph audit with deterministic AND–OR search labels LLM reasoning segments as supported, unsupported, or orphaned, improving balanced F1 over LLM-as-judge on a new 40-case clinical benchmark.
-
Cura 1T: Specialized Model for Agentic Healthcare
A healthcare LLM trained with an agent-driven loop that converts benchmark failures into new training data outperforms frontier baselines on most healthcare tests while staying competitive on general reasoning.
-
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.
-
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
A new bilingual dental QA benchmark and corpus reveals large performance gaps in LLMs for dentistry, and shows that domain adaptation with the corpus improves accuracy.
-
HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models
On a new Portuguese healthcare benchmark, leading language models score high overall but drop sharply in several specialties, such as neurosurgery and social work.
-
MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
A multimodal offline RL framework that fuses LLM-encoded clinical notes with structured vitals and labs modestly improves sepsis policy scores on some datasets, with evaluation caveats.
-
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
AutoEvoEval applies 22 atomic perturbations and multi-round chains to MCQ benchmarks, causing average accuracy drops of 7.283% and up to 52.932% for long chains.
-
Enterprise Large Language Model Evaluation Benchmark
A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...
-
A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment
GraphMPA combines an embedding-similarity hierarchical graph with mode-seeking preference optimization to improve RAG question answering on six datasets.
-
Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble
An ensemble of expert medical LLMs with triage and weighted consensus reports accuracy gains over single frontier models on medical QA benchmarks.
Discussion (0). Continue with ORCID to comment.