Pith. sign in

Paper Citation Record · LEDGER

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 3 inbound Pith citation observations for arXiv:2507.07988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07988 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:03.576676Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:11:20.481528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T10:38:36.176606Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fea62828-9d94-41fa-8edb-b65d3f7ba34e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.602141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.538335Z digest=sha256:1713a859a3a3af4633992479e58cb2f699771df963a4dc6476198d6607c6ef0d

Observation 82c86c53-cff5-4b11-9e22-becdd6e7bb11 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.584943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.605075Z digest=sha256:4163aa10bc3f3d6495fc540d3b598fbe880b065d5414bb07be43ac90b1eeeb84

Observation d024cad4-2a9f-4d7e-8650-c479bdefc270 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.569216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.699546Z digest=sha256:98ec601ef69f8c93b04fd6b99abc45e88ad224fa8f094d251f6a832cfcbf315c

Observation b1fa0974-30b9-4dc5-8a72-2a64f4e7e327 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.550385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.775519Z digest=sha256:06ebf4d55407800ad8c96d05d8eff7312e35f5570914f7306fed569e9518227c

Observation 9938ec45-ea59-4fd7-9f6a-7656b396dd7d · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.532446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.849681Z digest=sha256:84a55eec35b53c5c63d191f43bc1689e5e144cb2e8a150ff08e1fb549e634107

Observation 2a1ea739-43f7-4404-acc5-266adaefde6a · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.510289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:57.926277Z digest=sha256:7253d5769057be8bd0a06587cb91bc838af1be5ae375b737f5df2af8f57f2da5

Observation cf93f74a-3642-4262-a184-22013c439aa7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.490531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.030842Z digest=sha256:09eaf4b45fe35f291633c9f3b987b75446e0fd8169f2907e2dcf3bce66737832

Observation b8ea57cd-4e6e-42c7-a217-dee22dfa13c9 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.469292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.126704Z digest=sha256:8b50eb3d44d7024869e803ada64d938db5c6c37ad46648d532c6145c5790b2e4

Observation 33371637-d575-458a-aaa0-b0ad6ed691c1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.454065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.206001Z digest=sha256:a281b0d6c5e32f87552a2825b8542d7abd190bc66795ee374bb1fafe1696449e

Observation 78f6a361-d8f3-4917-b1aa-a90e6b6535bb · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.435785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.279514Z digest=sha256:de0011eca9f54be606a5145ea49811d2ab2cbcaaa22fef673a2cfb857c4f0e10

Observation 8201bc19-11ec-40d2-987a-27f7010a8402 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.410700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.351768Z digest=sha256:ec487325a10441a211e149844980bd511d1f1e8cb85fa15c05e10d61a96cc71d

Observation 94d06134-d812-4eea-8618-3082cdef62c0 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.427520Z digest=sha256:4440cca5f2a5feec7de1d76fd570955a73afa9e71ae2487899935a9f752a5c2e

Observation 68a926b1-31da-48cb-ae54-ad162f7f71c5 · outbound

This paper cites & Ranisch, R.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Ranisch, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.387708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.512040Z digest=sha256:5ea514e65d8f26c290dd26a0a2a94c7d26dad14697e727a2364533cbfc3d7076

Observation 0ee57d6b-f59d-4a04-ae2f-9d27b508b54c · outbound

This paper cites & Chen, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Chen, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.360873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.585270Z digest=sha256:7737938cee3063c99149a7c164d99cd60dd4b640e1c4824af37c6fb9e3a27e5e

Observation c410b272-f65c-4388-928f-d10e1bfbef9c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.343527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.645309Z digest=sha256:f0e96e4ac295c58ef42d2487eab29a33fc05c535b839999c93766e59208cdbf1

Observation df2e0a39-cbd9-401b-9a47-813ca40e9eb5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.315728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.703906Z digest=sha256:f98bead15e793511548ac749d461db95c5d9de54faa84b758818f33123305eec

Observation 2b2ebbac-7dc1-4ffb-989e-a8cd2c496ec7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.296147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.787928Z digest=sha256:6bfab82f5dd0973af9e0affdb4731454c15decdd41b21c3564d29a06529ac48e

Observation b6563a44-3ff1-455f-99b7-d1138c6e2e84 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.275469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.867942Z digest=sha256:1f4ed7680e05c049bb4a6769ed42ca40f3daecf8f9c47f24ac75197faaf98d04

Observation 9780ca1c-fb59-4ce2-a6fd-3f9d7771df50 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:31:03.894469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:58.912490Z digest=sha256:4e6e00f99cd6e71ea1a43f4ef01f61d7ff959c774ebbf52210fb4d18be11d134

Observation 96bcb4f4-9993-4925-8232-c60a00c68819 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.963784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.963784Z digest=sha256:229b4b4fa4d7d3f0858d25fc93b2aa70e9457bc304f61b03f1512dfbae0e63fe

Observation 0edfd9dc-f811-4d8b-b2cb-163399547ac4 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.239650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.028935Z digest=sha256:a0e10ac2a632dcd2fa2847c46f427942796d21ce12548d604871e35157775171

Observation ce73f86f-a492-49a9-9699-29eccc6f86d7 · outbound

This paper cites E., Motzfeldt, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models E., Motzfeldt, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.223653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.088944Z digest=sha256:5066125ed53f9fdda259bf08f6f1af0b72578bbff4c61ba5ccc3a26978398337

Observation 4badec55-5a92-445a-beef-5119c443f365 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.199441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.135672Z digest=sha256:814d1adc7349f63a2644d44b1ec4e978b27ed74c2de3ea6949645304b70471a8

Observation 7920dcd3-97d4-46be-82ea-5322ae45aee7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.155223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.176330Z digest=sha256:a7223537f5ff544a33191c80acaa2ce8e8bf832e3d95effe1e147a0c47af85d1

Observation 9715e056-8792-4869-8056-d83e11810222 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.135472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.223086Z digest=sha256:9bd2274702261c87853a90a9cae48d0908628c63dde7b73d2407ccab60112598

Observation 15b7a50e-c4b9-4033-9ae9-2ee3170bf2e5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.270249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.270249Z digest=sha256:85b60b21cba0c72e91caa5b1aa881b3ca01534ccaa8f7fb3abe99a05171abed2

Observation 654eef0e-e4ea-44f0-818e-9ab4b5b51d0a · outbound

This paper cites & Zhang, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhang, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.095502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.323111Z digest=sha256:eb30ab0547d22754bf3d5b6c3c773ea6b2570926d6abccaa9305d6e9c08dc1c4

Observation 595d3d64-676c-47f4-a650-f243b5b26c23 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.077073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.398355Z digest=sha256:d9634b9d645482affe320360d85a8a165d5c9eee9921b2ac2ecbab78775381e6

Observation 513cfb26-1ebe-446f-8783-d749c16559e5 · outbound

This paper cites & Zhu, W.-J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhu, W.-J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.448696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.448696Z digest=sha256:f15050898df76a0f99fb3b02d54ec36cd7e3c0694a784e23c0271e7239f7094d

Observation 32742a83-d522-476b-aada-9cafc088dc44 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models ROUGE: A package for automatic evaluation of summaries

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.056931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.526033Z digest=sha256:259bd5f5bda8fc05bef698592a512ec0982f53ab3bbcc4f4a6f82a5af6547234

Observation cad86f9f-c74a-42b3-8a7d-996eca6af49e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.032771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.591080Z digest=sha256:60784f293577416c78c08a53bc3baac62a29e43d5f3fa4b20658c3d0b2055a1a

Observation f41df4a2-84cd-4fbf-9d46-856873202ddc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.011481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.647833Z digest=sha256:6c31c52941ff668460844775bdefb47946e50fa2cbb725c415ad1a4759fff931

Observation 5325ec03-3582-46f6-96ef-30aba848691c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.992397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.731276Z digest=sha256:e91a4f5273ca3bf5b252b4431e6fe46920d191462a38fb9ea114db445779daf9

Observation 2b2bfa33-9464-4c8b-9c7e-74a7161d602b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.970046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.816765Z digest=sha256:a4515547cd862514fbf1ca25e4c71847569eecf65069e0d8a49bb789c02431e1

Observation 0924487e-7e62-4d87-bc51-a69a859c37ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.948488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.927540Z digest=sha256:a0a228374552a7811d66046d80adb82ec3a1ca70e6825071f23f0cc5d37ddec4

Observation 7d75a361-2d51-4ec0-b04e-42f68e294ab7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.933069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:30:59.986127Z digest=sha256:69efa7fbfa65c3506951961b45fcb33ace9b190ad161982cbe39aa53ce562cd9

Observation 35697907-0a56-492d-9a62-502def21f83f · outbound

This paper cites & Yuksel, D.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Yuksel, D

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.063157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.063157Z digest=sha256:0dbcfe3df9fd784ad0c8530437d5ce049af43e095bcc6e85aa97d13d05e1d441

Observation 5104ce6d-cde6-4ee8-8a4e-d50de8cbda29 · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.901236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.143843Z digest=sha256:1b47fcf7d784ffc390e753d043e864e371ff63a8e13620fea09c401e9b137747

Observation 9c211c11-1538-4043-be55-3e18d6957f65 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.711102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.227387Z digest=sha256:10e150e6bc39635edb09885823a408a9718088a654942e2cd7339919acf6d6d8

Observation d77a8f62-3d2a-4da2-8746-98299c88fbc3 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.297754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.297754Z digest=sha256:42e5a5a23618b84a1cfda9a17575ed2243305e07460889fa513ec645054a86b2

Observation 7a873045-b05b-4cae-b63d-989c30b9dfd5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.436810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.380994Z digest=sha256:d76d960750aa911b81cb4adc8b70af8f4b28d3357153f87d65d4a471b243893b

Observation fc03a2dc-81fd-4c6e-8802-3e8c43caeb5c · outbound

This paper cites F., Goel, R., Wen, Z., Martel, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models F., Goel, R., Wen, Z., Martel, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.300854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.463207Z digest=sha256:8dbe823910fb43eaeddac09fa81f8eb1f1d45b80d3c3c9f9700d5a802d24ceb8

Observation dba215f0-da18-488e-a5ef-b838c43c7cee · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:07.990827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.540262Z digest=sha256:bfc704c1061155dd59801ef8bd8e12f23981c7905fe6511b84701ed3f4fdc232

Observation e012493f-0dd8-46c4-ba8b-5a9496c6af2f · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.630372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.630372Z digest=sha256:b86af3c6b1c4196f754710ee08b793540aecef5d1096d4af2be6f64f3d140425

Observation 62b7f041-cf32-42c2-aa67-16cd97084df2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.848943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.760960Z digest=sha256:543ca9e9b88ad6a3b1050743cc4e67ccd9bcd22b7c15d6bacbf9d027083f0857

Observation 7d9c4fd0-ba39-41b4-85de-4610ba30e2ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.530259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:00.867715Z digest=sha256:cf9e38f6810a705ce90ba4131a46db88e8dc6c5da40a8350e7da6f9caa5ffdf8

Observation 8c4aa8a8-8ab3-47c2-9ff8-51943695fa21 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.962024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.962024Z digest=sha256:33334450e4c4e45cdc0bc82242205cd050f2493160f78596336480c065f3ac91

Observation 3e22655b-e182-4afb-b696-df244c612cce · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.176482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.052752Z digest=sha256:fde9dbb2cc393777d34ae4d57ca5a6037af97f759191a534e5521fd28b1a1a71

Observation a0da6c83-e3ed-4dec-bf93-7866be91e26f · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.997332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.146944Z digest=sha256:e5d9dc994cdb2521b2fcb57041345d3671d4bc6161d78b6efab43c81fdccc779

Observation fc27fef7-a5bf-4b47-9e17-25612dec8a60 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:01.239886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:01.239886Z digest=sha256:7b35d47ef39af2a4521f9e504e10d7c082e4e1d4f2da6285c4cc4847162a7eff

Observation d8abe900-b029-4332-ac3d-dda2921d18d6 · outbound

This paper cites & Gómez-Rodríguez, C.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Gómez-Rodríguez, C

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.835297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.341500Z digest=sha256:76deb2f48de27526fe6bef2f4be96245e102d9fb89281476d77d163d0227ab37

Observation 1cbece59-209a-468e-95be-d351c0166c44 · outbound

This paper cites & Lavie, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Lavie, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.693457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.473941Z digest=sha256:77098d3db918d2cf78b6bbc4d5e03c4bd4db4e9b941a9303fb46a93c8897c83b

Observation bfef7063-a6f8-4e3f-9838-a6b87e8c2f54 · outbound

This paper cites & Parikh, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Parikh, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.529950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.622469Z digest=sha256:5113c2ce9431b3d73e89ee8b16e6400e33c4f7c01b5ffcd3695e48a1e1ae9402

Observation cd95447a-adce-4700-a42b-c96546f6f4d2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.332867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.723726Z digest=sha256:dac0b66fcf0e404f527689a275d7a196f9c4d02c6871635c58b162ec19756a41

Observation e13e57a9-178d-4d8b-943d-20462a45e7a8 · outbound

This paper cites GPT-4 Technical Report.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.153060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.826623Z digest=sha256:23eaf9e177c1b36c35fa1e0ee47f91c568f966e6a1f523ae2c2dc7cd5d262218

Observation 3b6e1b6c-72f4-4171-b9eb-9aaf9216b8d0 · outbound

This paper cites GPT-4o System Card.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4o System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.021261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:01.969972Z digest=sha256:2938f8ce5add5754b9084821f90aa4aeda81442bc826c85bbc01d8d4a72ed32f

Observation b93086a9-ac3b-448c-83d1-74cb23b0facc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.868662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.061257Z digest=sha256:de6e7350d126bcf8a46c811b7de10e887532e80ab83c7a155fafed31c7add297

Observation a0709f7d-6805-4dda-bd90-2cd96bbb9273 · outbound

This paper cites claude-3.5-sonnet.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models claude-3.5-sonnet

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.736069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.213379Z digest=sha256:576e7429a00fdcf4d3f4a62e8ecb39dfaabba043c990b2901d694f7ed110f854

Observation ea5f7796-2527-4459-be53-fdab36e7d703 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.599003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.376098Z digest=sha256:3a2df962a84e41bf017d6ab6836fbfd0cca0a2b6dbf73b338c145610c77dec4c

Observation ba711ad0-87ce-41e4-b5d4-656e21e35187 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.461146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.462287Z digest=sha256:8d5957bef6cc4c8029585f7c120bcb38410322f295cd02763b43645e401b6f64

Observation d7d83259-76cc-46e2-9e34-e33de1a424d1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.317895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.601521Z digest=sha256:e2a745199d04c27ae627e85123e9cc40465b37fb01941e9ef89fcaac20412dd5

Observation 35b2db4c-e65b-448d-8c84-ab6ea47ee2fd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.192439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.743876Z digest=sha256:fabf4f2f46c4ee1fa0c6e70c6198c0e73e8c3659c5e6e86e7a8af5eeeac6038e

Observation 78e6ae41-eea0-4b7d-87de-8c495109db28 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.019319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:02.896069Z digest=sha256:6bb30b25c2736124f8442a34c2d3c8b06523dc38220cbcd48ca93df357a34ce0

Observation 9026ac5f-3925-46e5-ab3c-acc744688791 · outbound

This paper cites K., Raha, T., Khan, S.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models K., Raha, T., Khan, S

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:04.840016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:03.071493Z digest=sha256:c8314a377c694da484eb16aa6cc927711c37372b8f5b0dac69f9a0acfb0ccc4d

Observation cb3674c9-db41-44de-9bfa-6e784ff8c288 · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:03.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:03.237211Z digest=sha256:24cf3843249e25aa1fd97618e8af2f1ab20328154f0a85866a40fe6d544e4e81

Observation c02f3d50-24f3-44ac-9322-865f2fb5b65b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.545657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:03.443895Z digest=sha256:a9950ba20169412efcdf8a9b113c99a70fb66d3e6ca111b8fd1a86bee2da8073

Observation 8245cac9-65da-4272-af6c-45b915b618dd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.317615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:31:03.576676Z digest=sha256:0c13fe60a447f6fdb71121533b0e04c4bbb4d56e9fa5ffead52a82cd719cea2e

Pith citing papers

Observation 89338951-d8e3-47ff-a305-7f5e7bfd8d5d · inbound

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models cites this paper.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.481528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.481528Z digest=sha256:9c0aeb703dffd4fb99065b995fab6959c30e26dc114ea6911b6afb77d0321b0d

Observation 3cfe754e-cd08-4ea8-9c04-f57d9a221b4f · inbound

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation cites this paper.

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:38:36.213313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T10:38:31.601683Z digest=sha256:8162066d1cb9d80b1e1f7c1d0c8f1cb9851ce2bb3ae2c0ad9f59f3215d455e90

Observation e3612936-29d2-4b45-b3ec-42657fc5a18b · inbound

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering cites this paper.

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:28.955923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:04:28.955923Z digest=sha256:44e9f7185c23b9721a0764bc59f5fffee7904a262d9346d20095d476d3e35763