Pith. sign in

REVIEW 10 cited by

MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14371 v1 pith:7GCHZUZO submitted 2022-03-27 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords medicalquestionansweringdatasetentranceexamlarge-scalemedmcqa
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces MedMCQA, a new large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions. More than 194k high-quality AIIMS \& NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity. Each sample contains a question, correct answer(s), and other options which requires a deeper language understanding as it tests the 10+ reasoning abilities of a model across a wide range of medical subjects \& topics. A detailed explanation of the solution, along with the above information, is provided in this study.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 74 citations worldwide. Full citation record

  1. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

    cs.AI 2026-07 conditional novelty 6.5 of 10

    On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.

  2. Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    cs.CL 2026-07 conditional novelty 6.0 of 10

    An NLI-hypergraph audit with deterministic AND–OR search labels LLM reasoning segments as supported, unsupported, or orphaned, improving balanced F1 over LLM-as-judge on a new 40-case clinical benchmark.

  3. Cura 1T: Specialized Model for Agentic Healthcare

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A healthcare LLM trained with an agent-driven loop that converts benchmark failures into new training data outperforms frontier baselines on most healthcare tests while staying competitive on general reasoning.

  4. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.

  5. DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new bilingual dental QA benchmark and corpus reveals large performance gaps in LLMs for dentistry, and shows that domain adaptation with the corpus improves accuracy.

  6. HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    On a new Portuguese healthcare benchmark, leading language models score high overall but drop sharply in several specialties, such as neurosurgery and social work.

  7. MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A multimodal offline RL framework that fuses LLM-encoded clinical notes with structured vitals and labs modestly improves sepsis policy scores on some datasets, with evaluation caveats.

  8. AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    AutoEvoEval applies 22 atomic perturbations and multi-round chains to MCQ benchmarks, causing average accuracy drops of 7.283% and up to 52.932% for long chains.

  9. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  10. A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment

    cs.CL 2025-06 conditional novelty 4.0 of 10

    GraphMPA combines an embedding-similarity hierarchical graph with mode-seeking preference optimization to improve RAG question answering on six datasets.

Pith tools