Pith. sign in

Paper Citation Record · LEDGER

MedExQA: Medical Question Answering Benchmark with Multiple Explanations

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.06331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06331 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:22:52.006834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T06:43:10.331084Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 009a210e-faa4-4f71-b2de-e818d5809a10 · inbound

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models cites this paper.

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:22:52.006834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:22:52.006834Z digest=sha256:7e78301ddae6e493ce29137bc4bc814a784fb77951ed2e4b578f525c042e09cb

Observation 8331722b-b271-4f08-baaf-bc7ae2e6c368 · inbound

TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification cites this paper.

TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:46.910566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:39:46.910566Z digest=sha256:be3476159fa00c135c187fa96c260b9c1a242bef9b55245382075b37b2cd86df

Observation 840a9aee-0103-4034-9cc7-08e2094eff3c · inbound

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making cites this paper.

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:40.204720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:40.204720Z digest=sha256:db19da676fd4ebce974c3fcdb851d04ffd2b792b1f04641f08145451362b1055

Observation 9f3f7833-6bc2-49ff-96b8-04d12c84cac7 · inbound

BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain cites this paper.

BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:58.263719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:15:58.263719Z digest=sha256:5cd3eff5b133567f7bb2dce7f2f9b54d66aca04804a1e5f78cf98eaa4597d49a

Observation f4206f00-7e82-4c72-bd38-a53bb995c248 · inbound

AROMA: Augmented Reasoning Over a Multimodal Architecture for Virtual Cell Genetic Perturbation Modeling cites this paper.

AROMA: Augmented Reasoning Over a Multimodal Architecture for Virtual Cell Genetic Perturbation Modeling MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.744295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T23:07:51.076618Z digest=sha256:afb249f11ec0e5015213f8e0cb40e512a4071d56b34301d612ea3e1e06089196

Observation ff65055c-2306-4b75-8859-507df4433507 · inbound

How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking cites this paper.

How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.841217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T10:59:27.010954Z digest=sha256:aa26f50ce501131ade1634c43751fd9e44d3379c49704dd34dcb33086c853abc

Observation 665b1604-7557-4c46-a161-b421f6d7707c · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.332403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:07e035bc03240a21311196a3984fd9bb32b3e0c783101c5c6c4a3c7146aba7f3

Observation b1948ae0-5462-4dc6-ab3c-328015b895d2 · inbound

Between Knowledge and Care: A Mixed-Methods Evaluation of Generative AI for T2DM Self-Management from Patient and Physician Perspectives cites this paper.

Between Knowledge and Care: A Mixed-Methods Evaluation of Generative AI for T2DM Self-Management from Patient and Physician Perspectives MedExQA: Medical Question Answering Benchmark with Multiple Explanations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T00:24:45.348343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:24:45.348343Z digest=sha256:eb9b9ccc1b226bc64469f16136a6b441d278916337dd75fb4f4f806b569f6618