Pith. sign in

Paper Citation Record · LEDGER

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 3 inbound Pith citation observations for arXiv:2507.07988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07988 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:03.576676Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:11:20.481528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T10:38:36.176606Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fea62828-9d94-41fa-8edb-b65d3f7ba34e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.602141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.538335Z digest=sha256:a38b1311c89e63a2166f142a6f1df38f06d620d90acf1874c93dad40a08501e0

Observation 82c86c53-cff5-4b11-9e22-becdd6e7bb11 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.584943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.605075Z digest=sha256:f70c8ac977a70026f0e515c4fcadb14345067fd6857c5f65ae331a0fbea83a14

Observation d024cad4-2a9f-4d7e-8650-c479bdefc270 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.569216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.699546Z digest=sha256:915501f706a8307df5a330e5a991ec9f71ae4d2ba36e16913c8b29f07150705a

Observation b1fa0974-30b9-4dc5-8a72-2a64f4e7e327 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.550385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.775519Z digest=sha256:18fbfbf0198fdb337a94063106ff7d8f5e8bfe4db91953c637e3f12ac41869e1

Observation 9938ec45-ea59-4fd7-9f6a-7656b396dd7d · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.532446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.849681Z digest=sha256:57404d9b999661c622be0391f2b900d59940fbefa0361662a60961e66abdcee3

Observation 2a1ea739-43f7-4404-acc5-266adaefde6a · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.510289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:57.926277Z digest=sha256:56c58f95ad0feb049be135db87a5b42c58e297152fb12f49f8b4c28b57befcbd

Observation cf93f74a-3642-4262-a184-22013c439aa7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.490531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.030842Z digest=sha256:6b8c0267eda8b16ef9aba24242387692d200485104c4977f0dc487dd67b41606

Observation b8ea57cd-4e6e-42c7-a217-dee22dfa13c9 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.469292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.126704Z digest=sha256:8ce7cbe49d283b79bd51829d62de26b54d5def15d1ebe28bcba15e1496fb6715

Observation 33371637-d575-458a-aaa0-b0ad6ed691c1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.454065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.206001Z digest=sha256:6aa7d75dbb25113863d1d873f2445037099dee5a37ded285212b8e398fbb41b7

Observation 78f6a361-d8f3-4917-b1aa-a90e6b6535bb · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.435785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.279514Z digest=sha256:7876f6ab5e0c8ccdc88f38186ebdeab424ac37cf8210894b814071d62947c205

Observation 8201bc19-11ec-40d2-987a-27f7010a8402 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.410700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.351768Z digest=sha256:572e9ed5c7d6310eed471ac4ad8364444f4fc50ccc4703b2132a1135968e9121

Observation 94d06134-d812-4eea-8618-3082cdef62c0 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.427520Z digest=sha256:4440cca5f2a5feec7de1d76fd570955a73afa9e71ae2487899935a9f752a5c2e

Observation 68a926b1-31da-48cb-ae54-ad162f7f71c5 · outbound

This paper cites & Ranisch, R.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Ranisch, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.387708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.512040Z digest=sha256:72d39af85abe434674b5a9c1c0e4f383a498cb77279842dca64c6ec09434d8c3

Observation 0ee57d6b-f59d-4a04-ae2f-9d27b508b54c · outbound

This paper cites & Chen, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Chen, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.360873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.585270Z digest=sha256:64143733f686c0be672f3f8e0e6231bb43ce24bf680083833a98be0c092612f5

Observation c410b272-f65c-4388-928f-d10e1bfbef9c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.343527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.645309Z digest=sha256:fd0e2664cae5cf528cab0de30b79ca58e9c971601c2330be8c42c99e7e250579

Observation df2e0a39-cbd9-401b-9a47-813ca40e9eb5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.315728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.703906Z digest=sha256:f22135666a47f1d58788f7d30b06e172c8a9266f51fa66e19d7514a7802e5c01

Observation 2b2ebbac-7dc1-4ffb-989e-a8cd2c496ec7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.296147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.787928Z digest=sha256:65dcab0bdeed6f1366a4dbd2e005eccf094803edf4c645637661a70c51cbe8af

Observation b6563a44-3ff1-455f-99b7-d1138c6e2e84 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.275469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.867942Z digest=sha256:959138b63bf9303b4cbc5869f75c11c93a65b379aa69ee1d2d7da2bed2404e18

Observation 9780ca1c-fb59-4ce2-a6fd-3f9d7771df50 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:31:03.894469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:58.912490Z digest=sha256:54b1395d1e2d17f107971a0eda50946724828779118f77b688aee760b5fd4e44

Observation 96bcb4f4-9993-4925-8232-c60a00c68819 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.963784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.963784Z digest=sha256:229b4b4fa4d7d3f0858d25fc93b2aa70e9457bc304f61b03f1512dfbae0e63fe

Observation 0edfd9dc-f811-4d8b-b2cb-163399547ac4 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.239650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.028935Z digest=sha256:6385fffd4ef9b7fbad5b8241291eb8f73539ef93e058cb5337e9f06b2b71fb10

Observation ce73f86f-a492-49a9-9699-29eccc6f86d7 · outbound

This paper cites E., Motzfeldt, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models E., Motzfeldt, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.223653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.088944Z digest=sha256:77c2bf1370984637e498b955acc50d927b14ae341e7ead1a6d09551183afae5e

Observation 4badec55-5a92-445a-beef-5119c443f365 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.199441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.135672Z digest=sha256:c8a7ff0e3c902308bdcbc07df887153151f4d27343614e4b8ff1f5957a7e44ab

Observation 7920dcd3-97d4-46be-82ea-5322ae45aee7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.155223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.176330Z digest=sha256:19200641194808b4a2bc2a4a1106ccc2ccf3ee2c6b41773877c8c46d1cecb9ec

Observation 9715e056-8792-4869-8056-d83e11810222 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.135472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.223086Z digest=sha256:9fec1ff8551092dac0f5331eaa950f01c24dd4fb8de5c85e6db636868877a4a7

Observation 15b7a50e-c4b9-4033-9ae9-2ee3170bf2e5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.270249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.270249Z digest=sha256:85b60b21cba0c72e91caa5b1aa881b3ca01534ccaa8f7fb3abe99a05171abed2

Observation 654eef0e-e4ea-44f0-818e-9ab4b5b51d0a · outbound

This paper cites & Zhang, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhang, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.095502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.323111Z digest=sha256:eccebbe658caba7d1defd70b2bf2d6e0cf18490fc421d060b2331616b53c4a48

Observation 595d3d64-676c-47f4-a650-f243b5b26c23 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.077073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.398355Z digest=sha256:cbded73f4315464d9cbb65435f9f0eb91ab2f132d5c5ad385cec0d03238ed30c

Observation 513cfb26-1ebe-446f-8783-d749c16559e5 · outbound

This paper cites & Zhu, W.-J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhu, W.-J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.448696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.448696Z digest=sha256:f15050898df76a0f99fb3b02d54ec36cd7e3c0694a784e23c0271e7239f7094d

Observation 32742a83-d522-476b-aada-9cafc088dc44 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models ROUGE: A package for automatic evaluation of summaries

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.056931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.526033Z digest=sha256:6b14096fd26d338a420e33f114cd01526d60c8d25140effebce182ca6707dcde

Observation cad86f9f-c74a-42b3-8a7d-996eca6af49e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.032771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.591080Z digest=sha256:1b09dd20d3e6bb5c7acfa599b04435f354062a24c966c7268032c8d939d88de7

Observation f41df4a2-84cd-4fbf-9d46-856873202ddc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.011481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.647833Z digest=sha256:dfcad8034639a602e0a443c69bd87a91ac85fba83d23a0c02cc06475df47abce

Observation 5325ec03-3582-46f6-96ef-30aba848691c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.992397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.731276Z digest=sha256:b2a1d76dcf9f78e0c0465c2d10fa80239c369bad05a3228ec1a7638c69b1a9d7

Observation 2b2bfa33-9464-4c8b-9c7e-74a7161d602b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.970046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.816765Z digest=sha256:04565c843121037b95f3d3ee57fa5666ee4c06581ec7cbe76efccd49648f3b5f

Observation 0924487e-7e62-4d87-bc51-a69a859c37ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.948488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.927540Z digest=sha256:b8cf0fac6821f7c4e2f6fea51c4dabedae6ade4d81a8d32ca39b1400f2c481d1

Observation 7d75a361-2d51-4ec0-b04e-42f68e294ab7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.933069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:30:59.986127Z digest=sha256:5e192067f4a099cfa01950d44086843829503f1ca082e93035262d584e5b4227

Observation 35697907-0a56-492d-9a62-502def21f83f · outbound

This paper cites & Yuksel, D.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Yuksel, D

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.063157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.063157Z digest=sha256:0dbcfe3df9fd784ad0c8530437d5ce049af43e095bcc6e85aa97d13d05e1d441

Observation 5104ce6d-cde6-4ee8-8a4e-d50de8cbda29 · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.901236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.143843Z digest=sha256:789487952d4c9b4dd64022339ff453e8d22bf23a98e66bbb5131f338c55e1a8a

Observation 9c211c11-1538-4043-be55-3e18d6957f65 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.711102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.227387Z digest=sha256:39751b8b11e90679c829223e45a86f2ed78dacc3b02bb455187b135bba90e968

Observation d77a8f62-3d2a-4da2-8746-98299c88fbc3 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.297754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.297754Z digest=sha256:f4adf0b8b82141e8cc8acf0874cba19c29f95b953e1883e666d506473315228e

Observation 7a873045-b05b-4cae-b63d-989c30b9dfd5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.436810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.380994Z digest=sha256:b3f5e910d4179846634767407a1d1f31340f780392e1c573eec7a94082475b17

Observation fc03a2dc-81fd-4c6e-8802-3e8c43caeb5c · outbound

This paper cites F., Goel, R., Wen, Z., Martel, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models F., Goel, R., Wen, Z., Martel, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.300854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.463207Z digest=sha256:f0984388a464b5cd21fb3fb7e642a89afc59fd73b3480c3f73953f4359eaf85a

Observation dba215f0-da18-488e-a5ef-b838c43c7cee · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:07.990827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.540262Z digest=sha256:cf225559917cf0e779b4db70d2369bede96f3ef68b8e8b2c179ac69c2adc1c68

Observation e012493f-0dd8-46c4-ba8b-5a9496c6af2f · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.630372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.630372Z digest=sha256:b86af3c6b1c4196f754710ee08b793540aecef5d1096d4af2be6f64f3d140425

Observation 62b7f041-cf32-42c2-aa67-16cd97084df2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.848943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.760960Z digest=sha256:93f409a859d84ec4ff86b689b9a17d8ff1d1c2c7800cad465599009e573ab168

Observation 7d9c4fd0-ba39-41b4-85de-4610ba30e2ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.530259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:00.867715Z digest=sha256:b0c44f4e514b36c786670db0f403bbea608d33c9db1e30a647862b764a407e3c

Observation 8c4aa8a8-8ab3-47c2-9ff8-51943695fa21 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.962024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.962024Z digest=sha256:33334450e4c4e45cdc0bc82242205cd050f2493160f78596336480c065f3ac91

Observation 3e22655b-e182-4afb-b696-df244c612cce · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.176482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.052752Z digest=sha256:24fbbe9cad55be60e0fe7d1380eb1219d1dbdb1f4251830dd6eb3726c33ee955

Observation a0da6c83-e3ed-4dec-bf93-7866be91e26f · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.997332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.146944Z digest=sha256:d1f506eb7c655219076e75734e77c2db6f1e90cdc1a379fbc8d099aca488d00d

Observation fc27fef7-a5bf-4b47-9e17-25612dec8a60 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:01.239886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:01.239886Z digest=sha256:7b35d47ef39af2a4521f9e504e10d7c082e4e1d4f2da6285c4cc4847162a7eff

Observation d8abe900-b029-4332-ac3d-dda2921d18d6 · outbound

This paper cites & Gómez-Rodríguez, C.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Gómez-Rodríguez, C

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.835297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.341500Z digest=sha256:0b69fe9695acf7cdcacd690b962293bdfc936c09f2094e203482b5744100dc5f

Observation 1cbece59-209a-468e-95be-d351c0166c44 · outbound

This paper cites & Lavie, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Lavie, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.693457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.473941Z digest=sha256:4d224e5eb6b30dd225e636d5c46674e5867d04513a2734fb35ff438a7b92c26e

Observation bfef7063-a6f8-4e3f-9838-a6b87e8c2f54 · outbound

This paper cites & Parikh, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Parikh, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.529950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.622469Z digest=sha256:629639f259010ea8e54df1eadc4e6c6fe6c9b07254f39eada98e4064e3998168

Observation cd95447a-adce-4700-a42b-c96546f6f4d2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.332867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.723726Z digest=sha256:0414fb19fc52d84a59b795d48f4a6075d0f8066eb60c977da9fdfffd0d9f6f36

Observation e13e57a9-178d-4d8b-943d-20462a45e7a8 · outbound

This paper cites GPT-4 Technical Report.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.153060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.826623Z digest=sha256:baedbe4db0308326306fba9cf521bc19b77475dae4cfccf46a0862fcbd74df80

Observation 3b6e1b6c-72f4-4171-b9eb-9aaf9216b8d0 · outbound

This paper cites GPT-4o System Card.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4o System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.021261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:01.969972Z digest=sha256:00972f3766c2bdd08de2697ce43ca99367aebdf939982bdd804d855c5a187872

Observation b93086a9-ac3b-448c-83d1-74cb23b0facc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.868662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.061257Z digest=sha256:e05fd987c3e1498eb5801e0c2d2a43b824cbba3b5a4ea2048dce8aeec74c9cf6

Observation a0709f7d-6805-4dda-bd90-2cd96bbb9273 · outbound

This paper cites claude-3.5-sonnet.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models claude-3.5-sonnet

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.736069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.213379Z digest=sha256:7570487718b1b2894ca75a86e65076a8f7f674216562eb38aa574d314c545d67

Observation ea5f7796-2527-4459-be53-fdab36e7d703 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.599003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.376098Z digest=sha256:cd332a385015c5292946bb4d59e0f6d0300c887f51bc427c705e1fde783f89fe

Observation ba711ad0-87ce-41e4-b5d4-656e21e35187 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.461146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.462287Z digest=sha256:dfd38c4650e025b0d2ae0137d152f4fa5e77c167d888d75a63221bc3c5cce666

Observation d7d83259-76cc-46e2-9e34-e33de1a424d1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.317895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.601521Z digest=sha256:6172648a544cc501d1f34170f4f773e85f2f6ec0522a617513b70fbe52766df9

Observation 35b2db4c-e65b-448d-8c84-ab6ea47ee2fd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.192439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.743876Z digest=sha256:f5e75c1f40384c7509b6624124307785093ed3147a3cfa2f78cd22796a2aab68

Observation 78e6ae41-eea0-4b7d-87de-8c495109db28 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.019319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:02.896069Z digest=sha256:1a2364a7a52d56adba7ab73aa48524f83cd5e9aa7227e5b56db515d6f10c2c6f

Observation 9026ac5f-3925-46e5-ab3c-acc744688791 · outbound

This paper cites K., Raha, T., Khan, S.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models K., Raha, T., Khan, S

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:04.840016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:03.071493Z digest=sha256:3cc8616302bab4336008788d227c502c23ae95e7ff1d1a014073c474f7b86bd2

Observation cb3674c9-db41-44de-9bfa-6e784ff8c288 · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:03.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:03.237211Z digest=sha256:24cf3843249e25aa1fd97618e8af2f1ab20328154f0a85866a40fe6d544e4e81

Observation c02f3d50-24f3-44ac-9322-865f2fb5b65b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.545657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:03.443895Z digest=sha256:b2bafcd4799c7c5eb288e22da90626b6f8d282f837a0957112cb8ceabcf3ddb3

Observation 8245cac9-65da-4272-af6c-45b915b618dd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.317615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:31:03.576676Z digest=sha256:1f0fdcdfe06366cbf43caecfbb20544984d16eff7aeba47134e1403d4140a3f3

Pith citing papers

Observation 89338951-d8e3-47ff-a305-7f5e7bfd8d5d · inbound

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models cites this paper.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.481528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.481528Z digest=sha256:9c0aeb703dffd4fb99065b995fab6959c30e26dc114ea6911b6afb77d0321b0d

Observation 3cfe754e-cd08-4ea8-9c04-f57d9a221b4f · inbound

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation cites this paper.

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:38:36.213313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:38:31.601683Z digest=sha256:8ab48f1bb2044a497bc055177fb4ea41f2a23177505bc84edd182aba67d7aee4

Observation e3612936-29d2-4b45-b3ec-42657fc5a18b · inbound

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering cites this paper.

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:28.955923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:04:28.955923Z digest=sha256:44e9f7185c23b9721a0764bc59f5fffee7904a262d9346d20095d476d3e35763