Pith. sign in

Paper Citation Record · LEDGER

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

As of 15 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 3 inbound Pith citation observations for arXiv:2505.24040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24040 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:53.123407Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T17:02:30.696295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:34.994254Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy46
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 896e37bd-4420-440b-b58a-955634184988 · outbound

This paper cites Can We Use Large Language Models to Fill Relevance Judgment Holes?.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Can We Use Large Language Models to Fill Relevance Judgment Holes?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.373137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.373137Z digest=sha256:f13b6127660715b5d237ff5eae50e1e20d3273c150cf9e413b9fc93c18410579

Observation 4ba8765b-c6d5-4427-a016-69206bdb2628 · outbound

This paper cites Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.540335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.454079Z digest=sha256:8bb1a5379f0cc7891875d5799278fc880741e7982bfe0de072e998ac794238ed

Observation edb08fed-c0a0-4444-a431-21e84c342fa6 · outbound

This paper cites Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.520925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.520925Z digest=sha256:7f11361d41dae6302ace970eb60e6610e84af54ac726fee7c2b62c263ff079df

Observation dcc6e987-7e02-414b-b536-0475620f76c7 · outbound

This paper cites LLM Stability: A detailed analysis with some surprises.CoRR, January 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering LLM Stability: A detailed analysis with some surprises.CoRR, January 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.374152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.573430Z digest=sha256:54be1cecc7cb704ff1d38ad311ace93c541ae845a49f955936a2476c757409fc

Observation a7088955-a11b-40fc-b0f9-5cbba743a2c1 · outbound

This paper cites Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.204581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.618744Z digest=sha256:6e0947231f3f307c254d912ea8d50d2a2004f10c5b61fcad8bd1fcbb5f8b411a

Observation 5c5091e5-d6b4-4eb1-9bbc-097a5aa693bf · outbound

This paper cites LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.026316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.677371Z digest=sha256:b4b154324c841c8942722bc4f02fcf4094c27fd1dc1f6964fbaa78da2b600724

Observation c962acac-362b-44e9-a52c-ab7da4651209 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.853146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.733566Z digest=sha256:bee8f2bd32411d33b2c173e379bb1dbe026e0704bbdfc4687092475bb7f8c7c2

Observation 2dab7938-f13c-4679-b569-71c0f1da83aa · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.766014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.825672Z digest=sha256:12fc3e6cb8ff00f813d863fffa9ac0eea8f2b5d74b14a708c4b1fb4df9469826

Observation 605ec00f-5428-4b24-8de1-0fb383a0b60b · outbound

This paper cites Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.641860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.859493Z digest=sha256:36ce85a2547559131ad6fd90b6d59f209063619065504845acc8e32d87e5f49f

Observation 0022eabe-eae4-4e43-a1d3-c13a3cc5b603 · outbound

This paper cites Barnhill, Mar Llamas-Velasco, Gabriela Poch, Sören Korsing, Wiebke Sondermann, Frank Friedrich Gellrich, Markus V.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Barnhill, Mar Llamas-Velasco, Gabriela Poch, Sören Korsing, Wiebke Sondermann, Frank Friedrich Gellrich, Markus V

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.520737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.899250Z digest=sha256:0bacd500dc95492c9f365c3953d044e38f012350d2101fc6484aa9801999a7bf

Observation 1aa9cc7f-e276-443d-bc5e-9ac402ac6630 · outbound

This paper cites Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.370899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:47.933942Z digest=sha256:f06c4cf0f3cef8dc2ac8323cabb76684c8bdeb908c26507a1928e7de5a6eb6aa

Observation 60895be4-10d9-4b8e-81bf-5fa59db03965 · outbound

This paper cites Reasoning Models Don’t Always Say What They Think.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Reasoning Models Don’t Always Say What They Think

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.262806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.003702Z digest=sha256:b6645e8b364ce36d023bd9ba5350344fb561cf5a3ab2a182556fc72e1fda010b

Observation a333b942-e8db-48f9-a0b5-f255f65bab64 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.169216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.037851Z digest=sha256:2a6a12cc98bc2353706315cb7ab3de56bbd15a250b0788bcd6de79dcc291e59a

Observation 3c51a8e8-fc08-426b-8379-074619e98521 · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.072916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.072916Z digest=sha256:170be99809b6b6b39f00f5cf3ed9a0b7dbc56ddcadf2320cd2c175c64e13800a

Observation 3bebac39-b22f-4dc8-b55c-b2c82906396f · outbound

This paper cites Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:42:53.527784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.135694Z digest=sha256:6321862e6843881713d3d444b1d8b8549b0fb6baa75d8526f698380f23ec94eb

Observation e558dbca-b791-4367-8d0d-00a90e38525e · outbound

This paper cites SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.196183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.196183Z digest=sha256:3cd37a192cc79482585adb8f6e831d1ab081e1cb8f1a4ee3fda62040791522d8

Observation 92323745-e460-41c4-859c-9603845f493c · outbound

This paper cites Learning to Attribute with Attention.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Learning to Attribute with Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.271464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.271464Z digest=sha256:1dd415b7214d011575b007ba22c61e541b6b3e5b421cf79b402e0f3d61fc163d

Observation dbbf625c-d841-4454-9cb7-1168e8422158 · outbound

This paper cites ContextCite: Attributing Model Generation to Context.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering ContextCite: Attributing Model Generation to Context

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.051715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.295228Z digest=sha256:afa6620ff25421e6fb31f1a76a0acdfab27fb9de2bad59dc96158482cd5aadbc

Observation 25c4cf0f-5e78-460e-a01f-5b4faf721827 · outbound

This paper cites Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–13, February 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–13, February 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.957124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.349220Z digest=sha256:56fed59a33c3de725584a2e3ac9cbecdf4ac3025f305a8106eaee24580753128

Observation 5cd13f48-69b2-4551-8d11-b8913902161d · outbound

This paper cites Skinner, Ariel Dora Stern, and David Wennberg.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Skinner, Ariel Dora Stern, and David Wennberg

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.790928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.427677Z digest=sha256:ae226564505b39d960cbc7d2733556894afd6ee60667b109fd929220336527a7

Observation 5f85e1c6-8d58-4b93-af8a-3d4ccc12c41c · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:59.646744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.515752Z digest=sha256:e782ab7db6dc62d99173ad95da7fca926bb0b2425ecf65833db39b502eb4d659

Observation 8646b9b0-66ce-46f4-a2ce-7eab95d78398 · outbound

This paper cites RAGAs: Automated Evaluation of Retrieval Augmented Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RAGAs: Automated Evaluation of Retrieval Augmented Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.489396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.596357Z digest=sha256:95224c48eec8396bf6ee95747bf3099ff7cdcc02ecf21a7ec63bdba4062f65ee

Observation ebedf70f-cde7-4367-bdd8-4953545d8098 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.358407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.762620Z digest=sha256:96ca4d8d947b852fb3f1579add861766a626bca134810177e2c721b4c8c3fa19

Observation 956e18ff-2a35-4c26-9997-0ba048f0648e · outbound

This paper cites CiteBench: A Benchmark for Scientific Citation Text Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering CiteBench: A Benchmark for Scientific Citation Text Generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.208232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.836182Z digest=sha256:3bc96678ffa3bbd97d781c6e3fd264db1241368b9ab958762b964e42d2c47660

Observation d2fafe29-dff1-4c5e-8cf7-779d624989c9 · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Enabling Large Language Models to Generate Text with Citations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.103626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:48.958146Z digest=sha256:e50f156330e7e912a2ef4e23bc54e139284ebb57065387ad8bd2408b802be9e8

Observation 7ac6e3ac-3e43-45eb-9f47-5d1ed96a87bc · outbound

This paper cites Koch, Matthias F.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Koch, Matthias F

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.021226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.059462Z digest=sha256:aaaddca7fe8008e4e44f6547e3f1948c7476d8e851d714f539dd1b3e1aeae8a4

Observation 933f896e-63d6-4172-a219-43e18b2ff4bb · outbound

This paper cites The Llama 3 Herd of Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.147633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.147633Z digest=sha256:e44ca3efd3d4b05bc53e7bb6eb1231fd3fe07ab65fd70c2635ad94d02ef1972d

Observation 3e4e0f94-6992-457e-849c-f11294198349 · outbound

This paper cites Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.926035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.223937Z digest=sha256:c938ab24a67263f8d87d0c90d100bfb057e4fa7c82dcdf29978033501be72ede

Observation 858e417c-5da7-4210-8a11-adbd5dc72bad · outbound

This paper cites A Survey on LLM-as-a-Judge.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering A Survey on LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.331240Z digest=sha256:0a15376ddf90c6c54e25a03a6b4b31fa65266825d0aeb1ec8150c66f40b65c7d

Observation 414ad4e3-3197-4a1c-9a74-43f0ede88883 · outbound

This paper cites McKone, Daniel K.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering McKone, Daniel K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.849701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.421050Z digest=sha256:3689baec523047408982da64b85e65e4ed56afbdc2f4c0228c6da726fa4efe4c

Observation 31db672f-c43e-4c3f-88fc-749c9f84ae2d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Measuring Massive Multitask Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.523183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.523183Z digest=sha256:88d875c5affe7da0ce6f01eb3698d8c073976bcacb37e4f369a2ef9a88ae5b31

Observation 49b3ef4f-48b4-4ec8-89a4-4895b76b0d5b · outbound

This paper cites Spurious.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Spurious

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.731232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.623543Z digest=sha256:83ce947e9ced314e859a3f406bdac2956f66858a9c32aeb446b8cfd517a6abd8

Observation 554300d1-91e4-4681-a9c6-7bc73458aaec · outbound

This paper cites RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.692217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.696190Z digest=sha256:5d9bec99cce9e5c1b146dca7629171f6a08e9b5cd661a6e163e31169c7f3816e

Observation 00d8c488-aa8c-4d0a-9c97-98bf7547a73a · outbound

This paper cites What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.634175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.775976Z digest=sha256:abd93971c1ebf69298debffff62367b8444463d7f7ec7161f9d063059abfeb0d

Observation e5109b91-6a33-4bdf-baf7-1efaf4f3b2da · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.548413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.864087Z digest=sha256:5d85b73303e410e51b578cc3b020c52ebc3a6394bdca4abd65d108ab12baa691

Observation 3e55d466-c684-44d0-8684-7aa9cee9bae2 · outbound

This paper cites Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.475273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:49.945421Z digest=sha256:6a3e165b43fe08fe659433832363166300afe81987f7af585fa88ede7cb7ca2f

Observation e2f9176e-86cf-4e67-948e-bbbf3a2d91a8 · outbound

This paper cites GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.398558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.051191Z digest=sha256:352c1299fd6f34f9a860ef2387e0376086c7944b3f7e747f28c65ad5fd05407e

Observation 1f555e5e-8ec6-48f4-a962-e52dba343d5a · outbound

This paper cites Baleen: robust multi-hop reasoning at scale via condensed retrieval.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Baleen: robust multi-hop reasoning at scale via condensed retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.326631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.133218Z digest=sha256:31d823d047ea9271b998aa373788c16a6ae28fc9971288c33ac5f620ffb5dce4

Observation f6434b9f-01aa-484f-a74c-930e59a325b9 · outbound

This paper cites Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.278636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.234128Z digest=sha256:d4ac466da865f870b58ba9a9744805b3c6451c326e9924af3afe5bf1d8e9370b

Observation 1f67cda6-ea80-4875-82ee-58408f191d64 · outbound

This paper cites AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.213188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.304300Z digest=sha256:43b0ca7aaf253e048b98ee701c0995f6d8c1bf9f8883e723380a96d80b3b3b6c

Observation 56fd67af-3220-4adc-af55-53276c8e0353 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.169801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.380832Z digest=sha256:310b60478810aad7e4736d5ad8b2e2876902b1af79a6003483504a1dfe9a1afa

Observation 64a4ff52-f181-4bd4-9cf0-d4020f078ffa · outbound

This paper cites Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.127920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.460951Z digest=sha256:492304a4773b7482e524a1720ffb09a761c94af368d79de5bbe7e43ed6ab3291

Observation 850b2af6-8012-430e-9571-07616fc036ad · outbound

This paper cites Context Example Selection for LLM Generated Relevance Assessments.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Context Example Selection for LLM Generated Relevance Assessments

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.985087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.551402Z digest=sha256:235e0434da820080028e0d6ffabf7b3298a588a304f8313c25ec550da176bdd4

Observation 2a0765b8-280c-4ef4-b802-90f2dee564db · outbound

This paper cites GPT-4o System Card.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering GPT-4o System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:50.653743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:50.653743Z digest=sha256:fc21527126c52e337e68f9cbab7b759cdac08bee529b501794b55ee6004f0b77

Observation 6203c3c4-7570-4e33-a864-ccb56c174205 · outbound

This paper cites MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.903870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.733849Z digest=sha256:8eef8b60956b491c86c27973fd7bad519910408b66f006fedb22fa2c98d25d1c

Observation ecd5867d-7b9f-47f0-9424-6e4c838f586b · outbound

This paper cites Bowman, and Shi Feng.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Bowman, and Shi Feng

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.665395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.789115Z digest=sha256:e9b542d41da758f050b4b021eff895e93e5dd5fadc2eec8adea8fb1b6e5aca08

Observation 41b30a8b-73d9-463b-89dd-a2fee267dff3 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:57.510099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.838304Z digest=sha256:036cf339b4d27836b5a2bfc8e98a9e1d1938a1e3f83cea42b7e4c48c0e7a50a5

Observation 572f7938-3816-440f-8d60-01db2dac20db · outbound

This paper cites Qwen2.5 Technical Report.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:50.889423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:50.889423Z digest=sha256:009dee632dd5a1b57a1fad653dbc6879d257d057cf038bdd8f3479fd6a2ef1b5

Observation 2a2b63cc-2784-4256-b394-ab564a89868e · outbound

This paper cites Benchmarking Prompt Sensitivity in Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Benchmarking Prompt Sensitivity in Large Language Models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.268717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:50.989585Z digest=sha256:3df3eb0a064a851d8753c6ae956acfa53be8897be62de687e49750cec01b0573

Observation 9ad4cffd-c04b-41da-8b46-f9574415308d · outbound

This paper cites Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.096053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.056410Z digest=sha256:e0701cc9938bf89613516dff827eeff3546f9cc36894e557aabfcfa163e3485c

Observation 03c0d856-2e1d-4243-ae7b-523cd8d01e8a · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:57.017659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.118620Z digest=sha256:98e218c050f713c81ed5b5fec94c159c702166fba575dfd3997458045303a0cb

Observation b6a919f8-20ca-478d-b579-b69def8345e1 · outbound

This paper cites Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.844972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.187616Z digest=sha256:389c7fa55e230293440c9428577afcb105577a0327563d0c290ef4f8661746b3

Observation 8e97d18c-f69c-4a7f-a6a0-6b9cd9d21e2a · outbound

This paper cites Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.272998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.272998Z digest=sha256:98ae8f6a92f67d4103ba307191a0949254319b5da64089be5a58c8b357a6bc38

Observation 618559c9-d48e-4adc-84e1-0c84343e0e0d · outbound

This paper cites Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.346660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.346660Z digest=sha256:3ba0d078ba5c55db1ad72eed7ad8f89412779dd7e3068f1af6c3b0efc28fbc7d

Observation 5f00423a-8a68-4d61-805b-db539a1b02dc · outbound

This paper cites Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.543815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.410573Z digest=sha256:0aadc05563fefcfe3d8a60e8130b3ffb5c9449d20cbf1725c1917e32683a30e1

Observation 00ce585b-331c-4895-a2de-ba017a47fc50 · outbound

This paper cites Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.273646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.483199Z digest=sha256:356a81530a4698717b2693914af6cbaeb4508c6fbe89f458146ad6472fbbc4e9

Observation 332a4770-99a4-4ef3-a481-e6fccb35aeae · outbound

This paper cites RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.005774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.499100Z digest=sha256:46d20c5f2ba4efb4f196effbb2cdab36baac15943b053e6a4a2bee1d5ff8454e

Observation 131b2e92-d6ca-4370-ae6e-8b32163ac763 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:55.813354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.519320Z digest=sha256:3b2c1ef39eb461f51513011a6ecebc7d268a4a1c7692c3cb04d816d615e59b71

Observation 5d75952b-8710-4d0a-be25-59fb2c34cfd8 · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.574225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.574225Z digest=sha256:ba5f0da45752df5ef1165711ca2d4b7fa51fc5682b0a65531e6ace49158ad6ad

Observation 4718ff6a-cf92-411f-8310-476c7ff3d371 · outbound

This paper cites Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.610609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.692771Z digest=sha256:97c34b89523e2c169fcdc64f16cdbd6a1dd3fb166e85de17e640235c25a3b17d

Observation 9773ce0e-f307-480e-abf7-c1e95ab90dc7 · outbound

This paper cites MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.752953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.752953Z digest=sha256:343d2e705b34316206633a02327e934c1aa2a5ef01521305928571d2b95de92c

Observation 012df2de-7e4e-421c-8f92-ba936bbb4b9f · outbound

This paper cites An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.489658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.855274Z digest=sha256:1753a9dd5b51f6c6784e3376d299cd588f8f7e733d2302489b1d95db92a64a82

Observation 1758cd8c-4c30-4e3c-9b32-823590261821 · outbound

This paper cites CARES: A Comprehensive Benchmark of Trust- worthiness in Medical Vision Language Models.Advances in Neural Information Processing Systems, 37:140334–140365, December 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering CARES: A Comprehensive Benchmark of Trust- worthiness in Medical Vision Language Models.Advances in Neural Information Processing Systems, 37:140334–140365, December 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.353717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:51.968438Z digest=sha256:7395b424f74d2812b4443c0bf5c8fdc9870b4752d59b7725ca06976ca92d2183

Observation fa70b999-29d9-4a59-a506-b046a9b0b448 · outbound

This paper cites Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.221327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.018572Z digest=sha256:3b1894bb7f9dca23d45fcc9603605396dde3a60443fb6b68c7b324a7d4cc0f4a

Observation edca9806-32b8-4907-b8a0-62b4c8b149a1 · outbound

This paper cites A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.092113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.106367Z digest=sha256:bc8431873c12320b64b15988b0897072d5e433d8a2531d38d527f414c1638ea1

Observation ed446d7e-86b8-46f0-bad2-1b772a81272b · outbound

This paper cites Meyer, and Steffen Eger.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Meyer, and Steffen Eger

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:54.974189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.206653Z digest=sha256:bc8ba7511f7371c8144a0e3164a43fde02e72ac489bc855e9384e3df2d9356b1

Observation 30d7ea68-3a86-48f9-8f68-d268b451553a · outbound

This paper cites Gonzalez, and Ion Stoica.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Gonzalez, and Ion Stoica

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:52.346225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:52.346225Z digest=sha256:64e7f73a835b4c4cefb88168d283e1e85fc8476c2e87fb1493ca6db3f25c2937

Observation 1ee80f7e-a2d4-4c6b-a799-220692457244 · outbound

This paper cites Melton, James Zou, and Rui Zhang.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Melton, James Zou, and Rui Zhang

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:54.796467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.455186Z digest=sha256:d85f3e361cce17040a298f190c9cfc5323e883bb70678a34555c6765be7cf215

Observation f548da07-6d9c-438d-bc51-f2528ac72cc9 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:52.565185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:52.565185Z digest=sha256:5c37fbfa5d870e4151c26f6cd75e683b038f5ec7936e5310b25fbcfd8db41398

Observation 3bf35d46-e107-48e7-bb1b-8f8c404da2b8 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.605443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.655992Z digest=sha256:f3139b6973a44d9914364c463259439419425457e7dcde808e9cc7dc55296799

Observation e67608e9-17db-4bdc-8e0d-51711bed76e4 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.420889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.759621Z digest=sha256:b153ae41f1bd3c3d37b073f72b61abe103474aeeb2df85a99c1581da9b61d9f1

Observation 753df1bd-8613-42a7-8c2a-9c5e9cbedbed · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.180946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.868840Z digest=sha256:4b7816a0a4c8b81727006657290b28332546eebfa85ecf8eb39124d843a69af4

Observation 76b99462-dda6-4b93-ac42-1f74f8952046 · outbound

This paper cites Her new job involves walking several miles daily across a large facility.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Her new job involves walking several miles daily across a large facility

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:42:54.061125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:52.971780Z digest=sha256:9426c1c7ab1476dc829d638f6ff281fb4b68379d0c19306b82ebd759879517c4

Observation 12000119-02f6-4bc8-b0e2-c5e33039fb90 · outbound

This paper cites Towards Digital Sustainability in Health Care: Developing Digital Health Products through Data-Driven User Insights.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Towards Digital Sustainability in Health Care: Developing Digital Health Products through Data-Driven User Insights

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:53.800495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:42:53.123407Z digest=sha256:7f1a79a934076070b52f1a37be164d75402d92efb26cd15b7af463bcdccbd2a2

Observation d72d57a3-c8b8-4355-b7c5-2d08ac677448 · outbound

This paper cites Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.761883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.761883Z digest=sha256:3dd38458ff40ddf3a79b9bea8d12cdb54ad992cdccc7e71aaa1dbb99065cf90d

Observation e2972f1e-6fff-451e-b275-05a5a6bfbc06 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.690834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.690834Z digest=sha256:6199835de9e124ec65e2c9be594ed9960017248c87158dc1b94ee3156504057a

Pith citing papers

Observation c2b5d90d-a592-4cfb-a612-656ac73f0cb7 · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.892747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:7a721a6fc114446702ce06ff6b7c98cb0d2d3b3a7e4a66da18de01c4c105e1eb

Observation 754b740e-9581-46b2-9fe4-fbb513d5b8b2 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.604352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:a323ce3a9da55eb03d0e13b1db616a26c7f6995826294db37ebb2c8d55175ea1

Observation 3911e0e9-f2f5-4471-a71b-d8a61893a70b · inbound

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining cites this paper.

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.996501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T17:02:30.696295Z digest=sha256:c1a500c605b20cca47880f4333ad8125daefaea2fa41ca39ca6715b9ee940917