Pith. sign in

Paper Citation Record · LEDGER

HIVMedQA: Benchmarking large language models for HIV medical decision support

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2507.18143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18143 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:42:18.697628Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:18:59.283834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:24:01.661042Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00093d7-01e1-48eb-bb8c-7a925d5f04f4 · outbound

This paper cites & Topol, E.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Topol, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.345475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.327944Z digest=sha256:0a3a7995a1f9b6e3308bf5e52f30dbfd37153dd7fa81ae177450f6aff0fb2c01

Observation cbade836-7e62-48fc-a0c6-56b58dc1a4dc · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.333246Z digest=sha256:28afffb10d3ef05be68fe26e19a80b7aca083f84524417c258b5b798e0f32519

Observation 8cb3275f-2cd7-4c1e-87aa-50b5aa0d57b9 · outbound

This paper cites & Taylor, R.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Taylor, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.301557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.338407Z digest=sha256:48f3630300ba80d2c970a5ff32bbfed4cb79e387b1e2c9ce9a89f07486e37cf8

Observation a07a9f0c-7f11-4a62-aac3-8816d3a708b3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.279241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.348108Z digest=sha256:5a2774e83f930ece87b6fddb82f9df22f5484ca6b2de734b7c7012041e48bda0

Observation 7ab591a6-d8c3-434a-a2e7-c59205dfc5b0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.258544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.353270Z digest=sha256:755780a03068b712f1e7d2b80a48ec6c2c2701636317e2bbc52f8e9192adff4d

Observation d1145283-b369-458e-8b44-981598025c41 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.238457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.358167Z digest=sha256:0f81a830a3b25e239ad4f69b4f8de8c0fdf10057993a4ca6e9e4c05c23170cca

Observation 3bfe12e4-5e2b-4491-a1f7-701e880d3e76 · outbound

This paper cites Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial.

HIVMedQA: Benchmarking large language models for HIV medical decision support Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.216975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.364769Z digest=sha256:1c701825158b3d28e1716a432595427d68a5ff58654ab4020fa6ac9ac0bd2f61

Observation 9850c560-7d20-4d05-8d5e-bf9bf6e0f933 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

HIVMedQA: Benchmarking large language models for HIV medical decision support Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.370147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.370147Z digest=sha256:8c7a1c336e2245fdad36767559206965b1a59efd4ebc99c4d547602c68b70796

Observation b520f252-391d-40a0-bada-e08f7a2c7906 · outbound

This paper cites & Petro, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Petro, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.195521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.376914Z digest=sha256:34db15df75fdbf1f21d6d465ba6a21d77a1ea8770aa40aecec9422bf4a7d2d08

Observation a43450ec-d55d-429d-a722-69591d5051e0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.169555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.382624Z digest=sha256:79833998a4a977d8c2e14b6e0e5d498682fbbd326368b695aaa16f1a98abd5e4

Observation b9cd8862-9e6b-4482-940d-86bda2612c4e · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.147090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.388422Z digest=sha256:0f8bf20c54f22e9de7b98aff111e9e36222cafa1574e9d981b04ec131e664e18

Observation 6d58057c-469f-4d10-a7ea-6dfd1cfea59f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.112418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.396040Z digest=sha256:bf03b47aa6038fc0650540a11567094d709756ab305973d00d8471c145824b08

Observation a0268c01-e7f3-4dd3-b561-f2d9bc5a82d7 · outbound

This paper cites & Chow, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Chow, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.078979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.401352Z digest=sha256:92e2185b94a1442e3ae219b08f08340bb7125ec001bcc27a160ae9bb0dd5965f

Observation 4258786c-ff17-458e-9d9a-3d0308c087c1 · outbound

This paper cites & Dussault, G.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Dussault, G

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.049365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.407255Z digest=sha256:9cc9906f48f5fdf825538c4ccd35bf4e0e252d7b6e0c65da9f2d87c7ee6dd813

Observation df319221-9d75-4762-b780-023922aa6af9 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.027482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.414911Z digest=sha256:0e9f54ff267ff19f804c6b14e27fb28d19374cbfb5cae8b629c93a5928e08e56

Observation 4144b980-0807-4326-83eb-902099f163dd · outbound

This paper cites S., Link, K.

HIVMedQA: Benchmarking large language models for HIV medical decision support S., Link, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.996200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.423293Z digest=sha256:7cc591ff1fa73c4f9f312d0c710dbd36b4d2d52eea6a5c566478307946e36a34

Observation 8b1b808f-b918-46e3-90b2-adf5a6666d7f · outbound

This paper cites A., Lester, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lester, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.975130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.428884Z digest=sha256:9614178e0bb37ee0bce2e10520dacff93f65e61900062a995519f5ce2fa68b1d

Observation e9b44823-ec67-431b-95af-2be9221ba3cb · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.951455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.433357Z digest=sha256:351542a4f0fff11315017d5dfda8e466c422b7d1aae6325b30afdb340a49d7e4

Observation 194980ef-8d66-47ce-885c-05672c89046a · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

HIVMedQA: Benchmarking large language models for HIV medical decision support MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.438834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.438834Z digest=sha256:0bbdd200701517a9f7d873717004c248c8ab7e501fefc05975ddd426d2499bbe

Observation 283e7acc-25df-4505-9085-9bf869211366 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.927167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.444401Z digest=sha256:0c5bf458664319365f8b3f97a216b86365984dc67867a17c980095e291954aeb

Observation b09057d3-9bd3-4014-bb7d-26f8b9c2eb2f · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

HIVMedQA: Benchmarking large language models for HIV medical decision support AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.449467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.449467Z digest=sha256:ca1881798c63b751592c212a52daf6062f5a92a52257cc6b7ab5cd212644b236

Observation bbbbf14a-1c89-4f74-8a0b-584165d72658 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.900924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.457258Z digest=sha256:9d18a0721d54802cab222e8c095892afb7684e86a7e42c709150a4466c164c6c

Observation 54eaaba1-7eaa-4eb1-97b4-c421d08e0230 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-06T14:42:18.753396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.466737Z digest=sha256:7ce3d41b353f515b1dcb893af269c80c130de4c296beec807eab035330f3f01e

Observation 7100ea01-68b9-4964-9fa2-cf8de4a11d86 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.881456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.472116Z digest=sha256:2487db4fddc2018d54597beba22422cbb3a2036a932bb99d41064642814d7abb

Observation 905063a4-eb7c-40f9-a37b-15af8cd2b7b2 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.478987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.478987Z digest=sha256:6674091b1cd2c1ac429652c5ea6fcc6a91ba15af4f8b9d1b55666f6896946d00

Observation 30d1fa25-0606-4c56-873c-e0a04c54c838 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.839199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.483722Z digest=sha256:16d0fa3cdccaf3cff0e0d468cf511f2092ade80364800f0cb65c5101007b9c71

Observation f52043c0-b81a-4f05-9069-e414a76359e9 · outbound

This paper cites Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data.

HIVMedQA: Benchmarking large language models for HIV medical decision support Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.489506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.489506Z digest=sha256:0e973e540d1766cc098ac434fa542c1dc1d656c735b3ba0ff217d30a39859f71

Observation 32ad7521-810e-4550-8e2f-19abbdb99079 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.817052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.497292Z digest=sha256:835cf485a91ad00ee51e2b3a18ad6177899c3f77b2bba752109605f9a654fc35

Observation 127a8375-058a-45e2-ba9d-17c2f4ffb014 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.790721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.502209Z digest=sha256:cadfba30c8687bbb95152df31f19e0dc656d0b3617d9fd92104626cbc3772d33

Observation 2c305b8a-1052-421a-8afb-cbb5ccd6ec79 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.769605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.507177Z digest=sha256:8584d1655b55f7bd032db46b2c51f08d414178650cc84b11507b35cf2ee86d18

Observation af8f18bc-6c5c-4b95-84b3-8c981eaebaba · outbound

This paper cites A., Lingohr-Smith, M., Rogers, R., Lin, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lingohr-Smith, M., Rogers, R., Lin, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.731208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.515205Z digest=sha256:c5199141ba8b1acca0efbd44ea186171d5295d45367bec8b3a0ba76a788282fa

Observation 6d6b0bb9-fb89-4bcf-85f7-e95a6da22a90 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.712854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.519604Z digest=sha256:12b98290596891ac77ac83843de9f86b2ddf5cedbe26038b207563c30c1c2c7c

Observation 9b5f2c17-8b58-47fa-9b71-d29d63ef46c0 · outbound

This paper cites The Llama 3 Herd of Models.

HIVMedQA: Benchmarking large language models for HIV medical decision support The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.525401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.525401Z digest=sha256:75e3b655596349788b4ae1228d39077f5839625e66b65a35369cf65bfc0600bf

Observation 1a0ddae2-c61f-4989-99e2-13ed60bfe4f6 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.693675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.531488Z digest=sha256:53f24cae99d7401f8700c22e837910145496663601bafc6d4eeb8211e714a473

Observation e67a81bc-4767-41e5-ae98-d45116476693 · outbound

This paper cites Med42-v2: A Suite of Clinical LLMs.

HIVMedQA: Benchmarking large language models for HIV medical decision support Med42-v2: A Suite of Clinical LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.536520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.536520Z digest=sha256:928c093a4464969e7dd18be8aa5a5c61ab1fcf2c3cac82aad0c7b1b3dc5c7fa5

Observation b116b51f-60ef-42bf-a9f9-b39101152c97 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.542389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.542389Z digest=sha256:0991b1823c105b7d5e54fdc3248c92d10088c70809749ea65bd306c425c7cea3

Observation 91abcdb6-f3b3-422d-8517-ae17d85f68f3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.663539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.547648Z digest=sha256:fd3e624e7a9934e4e323a019adfc56d2fcc30d46990d5f577ede713ba7a3a1b8

Observation 85188c7c-f502-4aad-828d-c21a2ebc916f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.553615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.553615Z digest=sha256:5242bbd9f482f193b515d49a157f9a1d1b04f046a333ffba6bbf7058afaa0d4c

Observation 238f3bf3-6cd0-4784-8f72-113d9b51bffa · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.631293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.559307Z digest=sha256:7f5b7fcb81719d2cc314c58caa20bb08086ae52e34c4752cb237ee9a99400b24

Observation 733c0ab9-b198-48e8-a7fd-b0d6895220af · outbound

This paper cites A Benchmark for Long-Form Medical Question Answering.

HIVMedQA: Benchmarking large language models for HIV medical decision support A Benchmark for Long-Form Medical Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.565269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.565269Z digest=sha256:e57ddd5490375e388a316baa5c92b7bc126ca4e065e889324071b968aa1e4523

Observation 73429537-e222-466a-a55f-98764968841c · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.601340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.571529Z digest=sha256:44ef9b7cbdab7e783890363cc1044e427e6df0d3fd0bebb892bc506aad823cd4

Observation b656ce3a-8f40-43c1-8439-db8ba960243d · outbound

This paper cites GPTScore: Evaluate as You Desire.

HIVMedQA: Benchmarking large language models for HIV medical decision support GPTScore: Evaluate as You Desire

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.577140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.577140Z digest=sha256:405f142ffc4cdab27b2b3cb1f93fd90d9d72148f94b7844d074c409a6bfab33b

Observation 7a870550-04c4-4e4a-a3c1-716001ae2c5e · outbound

This paper cites JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability.

HIVMedQA: Benchmarking large language models for HIV medical decision support JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.584585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.584585Z digest=sha256:9d2bbb69f67b387cecbee9442f753c0bf4a3d9ef91444193398f2baf2abbe031

Observation 850d7db2-28cd-4daa-97ec-fc87bfbdcac5 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.590156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.590156Z digest=sha256:607c751b7684bbca4e7b7c11462d0176bf1d2e14bbf3543d806ecfd15cdd11f7

Observation bb43128b-c32b-4dc2-85f7-612df07471ef · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

HIVMedQA: Benchmarking large language models for HIV medical decision support Rouge: A package for automatic evaluation of summaries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.566562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.596357Z digest=sha256:906e1f921442dcf9253031e5e9d16f3d05c5681821f48723ad524bd92f44e084

Observation 0cb365e8-567e-4e8f-b125-853006d854d7 · outbound

This paper cites & Zhu, W.-J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Zhu, W.-J

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.546357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.602428Z digest=sha256:707d704144411f24b6568d309eb4c2151436ce7be48d544995ffaaf88d59ba5c

Observation 2e5c27f2-ab9e-46c9-9892-4b01b10ffeb6 · outbound

This paper cites The unified medical language system (umls): integrating biomedical terminology.

HIVMedQA: Benchmarking large language models for HIV medical decision support The unified medical language system (umls): integrating biomedical terminology

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.528066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.608825Z digest=sha256:3f5149cfe40efbd180616f97dd617d5e4254447442dea1cc3a6dc0c43be8bd29

Observation 6133ece7-f6fe-4a77-9e9b-4196eab01b6e · outbound

This paper cites ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing.

HIVMedQA: Benchmarking large language models for HIV medical decision support ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.614111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.614111Z digest=sha256:8246943be59b05f43fcaf13a32ecc77172105a568a0564e6ed9399b830c3afd3

Observation de0629ec-f942-4841-afa1-579c39959101 · outbound

This paper cites & Duclos, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Duclos, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.509088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.619442Z digest=sha256:c939e8607a2881e19282b22dd7ca6a579d642cfcb1559c954f7bf654a3fafbc8

Observation 88e426b7-3f18-457c-a08b-de8d3d112b1b · outbound

This paper cites Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies.

HIVMedQA: Benchmarking large language models for HIV medical decision support Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.492646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.624809Z digest=sha256:51bec18a0840804f7ef1eae91ee3244f34ac07ac23fd682e93d0e928f979bb37

Observation b556e891-6910-45b7-9227-35b94625b3d3 · outbound

This paper cites Unified Medical Language System (UMLS): 2024AB Full Release Files (2024).

HIVMedQA: Benchmarking large language models for HIV medical decision support Unified Medical Language System (UMLS): 2024AB Full Release Files (2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.474078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.630257Z digest=sha256:6b0290b8b249902206bc1526833999b2289dd4052c3af671b19a40737bf8c7b2

Observation 39c37b99-6a8e-48bf-8e40-05f3efd272b5 · outbound

This paper cites Nltk: the natural language toolkit.

HIVMedQA: Benchmarking large language models for HIV medical decision support Nltk: the natural language toolkit

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.457002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.637022Z digest=sha256:f3bbcfb7f72d453363a084bfe9e5994e16c752ed3473d838aff7afd08786110f

Observation d7074ab6-9d76-42e6-ad9c-628e233c462f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.433360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.645060Z digest=sha256:6a18af3a0da11720475aed5cbb018572d88e05e5e6a9fee5dc4ce2acbd54f338

Observation 611bcb64-c907-412e-b97e-675a73de8fda · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.412058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.652634Z digest=sha256:72084f76a4918429b1d3e9a706e444d39b9ea92a917a1c1cf01b740368f111aa

Observation 21637a13-bd05-44d4-b07f-63821aae5d57 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.385838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.663408Z digest=sha256:b0a97dec3707b7227620f5f558f3810480e70baec4e269f45cadf508b4a18266

Observation 21272b7e-acc2-410c-8499-524d04da7d5d · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.365142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.670150Z digest=sha256:176e927bee7ee37e189a8541db326ba33ab7f441329c7bd1a946b620a40d1a27

Observation 25b0ab7b-26fb-4b0f-acf9-01c3e1c304f1 · outbound

This paper cites - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations.

HIVMedQA: Benchmarking large language models for HIV medical decision support - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.332237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.676520Z digest=sha256:b72591e050c347831dfc61b9ba5fbda821a3dc0cd1c800e1ae4b396b677ea4c9

Observation 3cb5bd30-414a-4147-af87-b03d56049fd2 · outbound

This paper cites - Score low if the reasoning lacks clarity or is inconsistent with medical principles.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Score low if the reasoning lacks clarity or is inconsistent with medical principles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.315183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.681719Z digest=sha256:573ba6eac338e3e68aaf87c2507685c79ea0d25c93ebfbb0d6f497e0f9a60cfe

Observation cceefff4-bf91-4f9d-9ca7-aebe4c019f74 · outbound

This paper cites - A lower score should reflect the severity and frequency of factual errors.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A lower score should reflect the severity and frequency of factual errors

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.296570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.686903Z digest=sha256:5e2faffb62bdb60ecb4242aaf7da394015440c0d9cc54c0e4647188ef050f82b

Observation 0f45edbf-6d13-44e2-907e-40830b13a740 · outbound

This paper cites - A perfect score requires complete neutrality and sensitivity.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A perfect score requires complete neutrality and sensitivity

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.277039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.692350Z digest=sha256:0f885b2774a475bd8d64faeb79c384afeaa9cb328ca7744c05bc4f44f831ee75

Observation 7b71bcc7-b87a-49af-b57e-7544a5405a40 · outbound

This paper cites - Perfect scores require clear evidence of safety-oriented thinking.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Perfect scores require clear evidence of safety-oriented thinking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.256427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:42:18.697628Z digest=sha256:36c2fe3a5baa457ecc9d8f557ca000b0aecd3397a20c44779522eef062959405

Pith citing papers

Observation 19b33552-b332-4037-a819-f9cd47f40acb · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment HIVMedQA: Benchmarking large language models for HIV medical decision support

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.662324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:adc4914c3d80075c1ccd4fe11bf3c5e038a4f20577e081fd88b3c0ed80515cce