Pith. sign in

Paper Citation Record · LEDGER

HIVMedQA: Benchmarking large language models for HIV medical decision support

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2507.18143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18143 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:42:18.697628Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:18:59.283834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:24:01.661042Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00093d7-01e1-48eb-bb8c-7a925d5f04f4 · outbound

This paper cites & Topol, E.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Topol, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.345475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.327944Z digest=sha256:9b465384034b7adeaaec9e4928be2f33dbbadc59f9f90a1047a3c4d690327f34

Observation cbade836-7e62-48fc-a0c6-56b58dc1a4dc · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.333246Z digest=sha256:7217e4636659a8aef6b7e4a2f0b8782804c80936347a6df0741f24c88bfa2d8b

Observation 8cb3275f-2cd7-4c1e-87aa-50b5aa0d57b9 · outbound

This paper cites & Taylor, R.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Taylor, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.301557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.338407Z digest=sha256:bc2fefdfca3b78d24408a244b681ace988a429799fa9a5431808b9f62d4f6414

Observation a07a9f0c-7f11-4a62-aac3-8816d3a708b3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.279241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.348108Z digest=sha256:c9c56a4862a87eb08a89fd38ce3199951f6fa6a25ba56617376654f4a1122ed7

Observation 7ab591a6-d8c3-434a-a2e7-c59205dfc5b0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.258544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.353270Z digest=sha256:29d6583bdbe1e0c67a531e80b232c1dec149beaecb648d883e06d87e644812cf

Observation d1145283-b369-458e-8b44-981598025c41 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.238457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.358167Z digest=sha256:34c1138703d357b2a0e11432ca0d4d7064bce919d5d59940be8c8bb328cd9ce5

Observation 3bfe12e4-5e2b-4491-a1f7-701e880d3e76 · outbound

This paper cites Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial.

HIVMedQA: Benchmarking large language models for HIV medical decision support Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.216975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.364769Z digest=sha256:e5048a9e0fe37cb3d7d59c3337605fc9a6906148c5f9a307d807339077523e63

Observation 9850c560-7d20-4d05-8d5e-bf9bf6e0f933 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

HIVMedQA: Benchmarking large language models for HIV medical decision support Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.370147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.370147Z digest=sha256:3eb26352b76bd15f8c431f45d485a272cb5837ffa435ed5ee13daca12faa7972

Observation b520f252-391d-40a0-bada-e08f7a2c7906 · outbound

This paper cites & Petro, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Petro, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.195521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.376914Z digest=sha256:e9ab5de0aae93d3572ae852f0f83b9c5823cac104e600b92a7d2e615cfbc3fc1

Observation a43450ec-d55d-429d-a722-69591d5051e0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.169555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.382624Z digest=sha256:7bf4fee50f3a42f635b1cc231601a6114d4d8559a952a32fe4f687ce6c4d7eb3

Observation b9cd8862-9e6b-4482-940d-86bda2612c4e · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.147090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.388422Z digest=sha256:4160e2632adcc9e3b27c050c114c2f8481c38b58bca8d4f46dadc04fe2426459

Observation 6d58057c-469f-4d10-a7ea-6dfd1cfea59f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.112418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.396040Z digest=sha256:664619f0b5609d857b5f832b0ce83dc7620d66042d487d9e3843e4b0a4ce10a6

Observation a0268c01-e7f3-4dd3-b561-f2d9bc5a82d7 · outbound

This paper cites & Chow, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Chow, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.078979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.401352Z digest=sha256:8ddacf950d01633ce3e0ccb11ef3dca7910a12b68562bc9aedfc35606fa56c28

Observation 4258786c-ff17-458e-9d9a-3d0308c087c1 · outbound

This paper cites & Dussault, G.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Dussault, G

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.049365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.407255Z digest=sha256:93a341c6d8ecafff403ad602d820ce9e60473effb4f90d1e35366cb9b04f6176

Observation df319221-9d75-4762-b780-023922aa6af9 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.027482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.414911Z digest=sha256:3905194377f827730acadbe975c34b8ee62498bd37bd88e6655e901e193e1135

Observation 4144b980-0807-4326-83eb-902099f163dd · outbound

This paper cites S., Link, K.

HIVMedQA: Benchmarking large language models for HIV medical decision support S., Link, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.996200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.423293Z digest=sha256:863a1d40ab158c91781b5c1a39063033830393b5dbbf2af9096e61ecde5b6939

Observation 8b1b808f-b918-46e3-90b2-adf5a6666d7f · outbound

This paper cites A., Lester, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lester, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.975130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.428884Z digest=sha256:1ece37c6c1bc6d6a77c79a0f7ae7c49aab95f0c6059ce4d4236215149cd77f16

Observation e9b44823-ec67-431b-95af-2be9221ba3cb · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.951455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.433357Z digest=sha256:751adc538b2e308aaf9541a55e2bf50ce9a0e04dfe24991e18270b2c69711b02

Observation 194980ef-8d66-47ce-885c-05672c89046a · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

HIVMedQA: Benchmarking large language models for HIV medical decision support MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.438834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.438834Z digest=sha256:2395774d35d453ff35082dae380c88c50b9c98f65d9d0fd0a2598c8dc52b30c4

Observation 283e7acc-25df-4505-9085-9bf869211366 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.927167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.444401Z digest=sha256:15fa34a0dff985a3411e3b247b883728889be1682a5e27805bf6e27f1f5815c9

Observation b09057d3-9bd3-4014-bb7d-26f8b9c2eb2f · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

HIVMedQA: Benchmarking large language models for HIV medical decision support AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.449467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.449467Z digest=sha256:f8db43bbb316833b75b6f3937b83467c4659a53c90239eb89a037d3893353087

Observation bbbbf14a-1c89-4f74-8a0b-584165d72658 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.900924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.457258Z digest=sha256:0abbbb331988510a082df2962ee34bd1ff82f239585d2fdb055bae01c4092fad

Observation 54eaaba1-7eaa-4eb1-97b4-c421d08e0230 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-06T14:42:18.753396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.466737Z digest=sha256:ce0a0aebb7864f515c48db4a53ef34e22ba0c2a6010cc096003fdb3e701c93fd

Observation 7100ea01-68b9-4964-9fa2-cf8de4a11d86 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.881456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.472116Z digest=sha256:966a8feb3a555e88086824365a7c86bbd82fb6bc9f74b93e69a24b24d4a2be43

Observation 905063a4-eb7c-40f9-a37b-15af8cd2b7b2 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.478987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.478987Z digest=sha256:de97681b62f666df1fee776ea5de3f3006d7167ef883ec13c1889cd20e3ea697

Observation 30d1fa25-0606-4c56-873c-e0a04c54c838 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.839199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.483722Z digest=sha256:6aff0be4941815090800c2fa89928a918b3d7ba993e6bc3bafb2c01b7e710cf3

Observation f52043c0-b81a-4f05-9069-e414a76359e9 · outbound

This paper cites Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data.

HIVMedQA: Benchmarking large language models for HIV medical decision support Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.489506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.489506Z digest=sha256:70650599c21ad9883ad3e006a2ec3e4445deca569e2bc8a7a811238dc417cba0

Observation 32ad7521-810e-4550-8e2f-19abbdb99079 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.817052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.497292Z digest=sha256:b6747fb5aeb5d391a4827bbbd44f0e080a0d2713fe804c3d9842852677fb4a66

Observation 127a8375-058a-45e2-ba9d-17c2f4ffb014 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.790721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.502209Z digest=sha256:f983c324fcac24b31bd9126789b863bfc7e2b24dff2925341bdcbf89949eeffc

Observation 2c305b8a-1052-421a-8afb-cbb5ccd6ec79 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.769605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.507177Z digest=sha256:8c9493d0b9bbddbb4da4e05fb4250d275ce4978a82e093d606df5d0904205491

Observation af8f18bc-6c5c-4b95-84b3-8c981eaebaba · outbound

This paper cites A., Lingohr-Smith, M., Rogers, R., Lin, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lingohr-Smith, M., Rogers, R., Lin, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.731208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.515205Z digest=sha256:4563212987ca9ff7c813114c1b7dee7320ace049e59536d4694162fb9969526c

Observation 6d6b0bb9-fb89-4bcf-85f7-e95a6da22a90 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.712854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.519604Z digest=sha256:44942c2e1d0f0cc7fa863c36f833c26601e707120e1935356da8aed08ad4af27

Observation 9b5f2c17-8b58-47fa-9b71-d29d63ef46c0 · outbound

This paper cites The Llama 3 Herd of Models.

HIVMedQA: Benchmarking large language models for HIV medical decision support The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.525401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.525401Z digest=sha256:86647afe89423b29e7b0d9579ed16b8607a7dab2d7ce6fe838737778475a5290

Observation 1a0ddae2-c61f-4989-99e2-13ed60bfe4f6 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.693675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.531488Z digest=sha256:ecd8f1cb0bd907e63eb813832a09556412543280d1fe7d2413836f8f9819e936

Observation e67a81bc-4767-41e5-ae98-d45116476693 · outbound

This paper cites Med42-v2: A Suite of Clinical LLMs.

HIVMedQA: Benchmarking large language models for HIV medical decision support Med42-v2: A Suite of Clinical LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.536520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.536520Z digest=sha256:31a7ac15acaee59204d8dfa810c9caa69c28c51f1c3b6ac2d9896055130d3b4d

Observation b116b51f-60ef-42bf-a9f9-b39101152c97 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.542389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.542389Z digest=sha256:0073f738e98ca117920fb5cb384ce6f949bfec975d524c562dfaf59092939848

Observation 91abcdb6-f3b3-422d-8517-ae17d85f68f3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.663539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.547648Z digest=sha256:1421aa3cdcdb7b98305698becdd1bce36016e7b02e1295d3bc2160db8ba9c77c

Observation 85188c7c-f502-4aad-828d-c21a2ebc916f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.553615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.553615Z digest=sha256:16790cf3b896117022874303ae4b9565a0e1a2f9268bd01e55f651c3610787a8

Observation 238f3bf3-6cd0-4784-8f72-113d9b51bffa · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.631293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.559307Z digest=sha256:6d6bef94961b8c3f17b0542801887b50f2f5316258a118332579ebd691a1a8d9

Observation 733c0ab9-b198-48e8-a7fd-b0d6895220af · outbound

This paper cites A Benchmark for Long-Form Medical Question Answering.

HIVMedQA: Benchmarking large language models for HIV medical decision support A Benchmark for Long-Form Medical Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.565269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.565269Z digest=sha256:5f66808c05bacc6af8a3bb264202a476bc313e63bc90f9862b193676ab7b9743

Observation 73429537-e222-466a-a55f-98764968841c · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.601340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.571529Z digest=sha256:b14f38f8e3897310a3c455cd5b7a67bcb845d29c688c2fb169a079364b06fa1f

Observation b656ce3a-8f40-43c1-8439-db8ba960243d · outbound

This paper cites GPTScore: Evaluate as You Desire.

HIVMedQA: Benchmarking large language models for HIV medical decision support GPTScore: Evaluate as You Desire

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.577140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.577140Z digest=sha256:a05a56f8603d33524508e6e4e311f2d03033212c0e05f4b7328f9e6ba2d2ecab

Observation 7a870550-04c4-4e4a-a3c1-716001ae2c5e · outbound

This paper cites JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability.

HIVMedQA: Benchmarking large language models for HIV medical decision support JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.584585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.584585Z digest=sha256:3298f8797fc1bf1e32f3c5c77302b457b89a29f14af3377456964c3440e11c88

Observation 850d7db2-28cd-4daa-97ec-fc87bfbdcac5 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.590156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.590156Z digest=sha256:5f40ade7e1b7aa53dca724e674b4d03d196d061a164387f98fc2af5493928197

Observation bb43128b-c32b-4dc2-85f7-612df07471ef · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

HIVMedQA: Benchmarking large language models for HIV medical decision support Rouge: A package for automatic evaluation of summaries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.566562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.596357Z digest=sha256:6fb5d42e23c6f0cdc11c5238c546e2f231f8aefef4838b0cc3fd89ff6936d7d3

Observation 0cb365e8-567e-4e8f-b125-853006d854d7 · outbound

This paper cites & Zhu, W.-J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Zhu, W.-J

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.546357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.602428Z digest=sha256:7a8c8b06636ba536295d363c15ab578bfee89de5703b0601b1b58239a4d458d1

Observation 2e5c27f2-ab9e-46c9-9892-4b01b10ffeb6 · outbound

This paper cites The unified medical language system (umls): integrating biomedical terminology.

HIVMedQA: Benchmarking large language models for HIV medical decision support The unified medical language system (umls): integrating biomedical terminology

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.528066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.608825Z digest=sha256:12b84b014f169755e6870674cc62eccd7f4c76daa6f564643a820916aa814ad3

Observation 6133ece7-f6fe-4a77-9e9b-4196eab01b6e · outbound

This paper cites ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing.

HIVMedQA: Benchmarking large language models for HIV medical decision support ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.614111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.614111Z digest=sha256:dee7efb6d739b836fe17d3840cf1ad86d1ac720ba7ca65bf0b327ce9f1c543f5

Observation de0629ec-f942-4841-afa1-579c39959101 · outbound

This paper cites & Duclos, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Duclos, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.509088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.619442Z digest=sha256:50f8204d32744ef3c8fea8109c07f6a046b2a0478bf51383a47be58122a9cefb

Observation 88e426b7-3f18-457c-a08b-de8d3d112b1b · outbound

This paper cites Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies.

HIVMedQA: Benchmarking large language models for HIV medical decision support Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.492646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.624809Z digest=sha256:56f248ad30e166ccf6b240d3ca4b899a6518511d0498db1669ec5bde6ccbb23c

Observation b556e891-6910-45b7-9227-35b94625b3d3 · outbound

This paper cites Unified Medical Language System (UMLS): 2024AB Full Release Files (2024).

HIVMedQA: Benchmarking large language models for HIV medical decision support Unified Medical Language System (UMLS): 2024AB Full Release Files (2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.474078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.630257Z digest=sha256:77377deacf943ba4e43c13fbda64c60d9b0b56ce6ede73260a43a101fb81fadf

Observation 39c37b99-6a8e-48bf-8e40-05f3efd272b5 · outbound

This paper cites Nltk: the natural language toolkit.

HIVMedQA: Benchmarking large language models for HIV medical decision support Nltk: the natural language toolkit

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.457002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.637022Z digest=sha256:d8613492a591122e52584110068a1ffbaf3a7a72d764aeebe1d7f1e000a79146

Observation d7074ab6-9d76-42e6-ad9c-628e233c462f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.433360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.645060Z digest=sha256:46ef59e412a32ebb44eef67150c1452bb4ce2fa1d894d8dfe6e4624cf2ecc04e

Observation 611bcb64-c907-412e-b97e-675a73de8fda · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.412058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.652634Z digest=sha256:0fa4c070aff08ac81e4359d7c88bb390087ab54032a6a389c3de4b9c5df73c59

Observation 21637a13-bd05-44d4-b07f-63821aae5d57 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.385838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.663408Z digest=sha256:fb2818f3a8e138e80241bc0240c21cac1771265bf6db4746dd8e1bcabec5048a

Observation 21272b7e-acc2-410c-8499-524d04da7d5d · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.365142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.670150Z digest=sha256:574ff1c876e9cef0f35261458d3c94ce08c2ea399e4b970feec65cb31a2f5437

Observation 25b0ab7b-26fb-4b0f-acf9-01c3e1c304f1 · outbound

This paper cites - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations.

HIVMedQA: Benchmarking large language models for HIV medical decision support - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.332237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.676520Z digest=sha256:5949667c23101fda7dc25d2a2316808aabc60cf64c485a8d2442b9f134c5a100

Observation 3cb5bd30-414a-4147-af87-b03d56049fd2 · outbound

This paper cites - Score low if the reasoning lacks clarity or is inconsistent with medical principles.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Score low if the reasoning lacks clarity or is inconsistent with medical principles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.315183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.681719Z digest=sha256:26205c74d794dcf43abb951f02356c7e4d79a762dd05e7db0713970bad54d9ed

Observation cceefff4-bf91-4f9d-9ca7-aebe4c019f74 · outbound

This paper cites - A lower score should reflect the severity and frequency of factual errors.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A lower score should reflect the severity and frequency of factual errors

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.296570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.686903Z digest=sha256:9541a37616a587d2fbdfd1534feff8848246f95884f0bee0a4930c929a0b4966

Observation 0f45edbf-6d13-44e2-907e-40830b13a740 · outbound

This paper cites - A perfect score requires complete neutrality and sensitivity.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A perfect score requires complete neutrality and sensitivity

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.277039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.692350Z digest=sha256:8f3008408bc8c2ecd96bb9fc1d024a6d3d18ff76c730749924e16f03898bd521

Observation 7b71bcc7-b87a-49af-b57e-7544a5405a40 · outbound

This paper cites - Perfect scores require clear evidence of safety-oriented thinking.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Perfect scores require clear evidence of safety-oriented thinking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.256427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:42:18.697628Z digest=sha256:b4497ff92670cfa3a3d053606474b52803963d0f12041ccc1cf53a5c643f8940

Pith citing papers

Observation 19b33552-b332-4037-a819-f9cd47f40acb · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment HIVMedQA: Benchmarking large language models for HIV medical decision support

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.662324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:0905efc1b400df8adcfc141e5242431735396895ab22aa45cb1b4f68fe3cd87c