Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Match the Conclusions of Systematic Reviews?

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2505.22787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22787 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:05:45.754584Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:03:59.798126Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.432897Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abfb9d70-afe7-4cb1-855e-741140657f6c · outbound

This paper cites Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases.

Can Large Language Models Match the Conclusions of Systematic Reviews? Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.999188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.627088Z digest=sha256:9ba86ef06fce40ae53b64a40233c6a8fc1286dde7ab3a7311a2af6fcf0b37a4a

Observation 8d6b0df0-f2ba-4cb7-817e-201fb4cac29a · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:51.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.702126Z digest=sha256:a056a68d0dcf92efd6790845c3d8d5e008e59b4b5ce7cf9f51062579632d26e9

Observation 8823c06f-d137-4b0f-991b-818d8f54718c · outbound

This paper cites The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:41.788855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:41.788855Z digest=sha256:e341bd2f8d6b8c443ece9172b820309299024505dc74e8af81d3772a71494639

Observation 22686ff8-3484-41d7-a983-d27f398c74c3 · outbound

This paper cites How to optimize the systematic review process using ai tools.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to optimize the systematic review process using ai tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.728714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.857108Z digest=sha256:dc570488978294cb2f9161d16bc2c553a37a9481bdf57ec34c13345e7139f698

Observation 3ef3b132-53c3-4c9f-b155-040c13c904d1 · outbound

This paper cites Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses.

Can Large Language Models Match the Conclusions of Systematic Reviews? Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.608818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.917310Z digest=sha256:83b44040b68e848006a32cdc12297c43277cd35d6a86a226802b595051befdfc

Observation 41609f5e-f087-4fa1-a9c7-11f8ba592707 · outbound

This paper cites Deep research system card, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deep research system card, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.433866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.001265Z digest=sha256:5d539783783dfc6467ba28f991043c7f7daadb0e717e7a4b89c5b1647e297847

Observation 94eaa987-14a8-4cf7-9b2b-938abb39df49 · outbound

This paper cites Gemini deep research – your personal research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gemini deep research – your personal research assistant, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.069304Z digest=sha256:53b65646562d735bebd1ab8620950bfca9a1fc0e193230411b1b983b5ede0585

Observation 0f59b63d-401a-4432-b88c-b2d27d28c6cc · outbound

This paper cites Elicit: The ai research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Elicit: The ai research assistant, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.172471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.163546Z digest=sha256:678e678392635a5a0e61938edf501418c173b5bfed01d25b220f5758beed2a02

Observation 00702845-aabc-4db0-8d57-5d07882d3dff · outbound

This paper cites Open evidence: Ai-powered medical information platform, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open evidence: Ai-powered medical information platform, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.025871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.237160Z digest=sha256:66c35f616b28f4f9d29c4f75268f52a39f8369458fa2633130a2f02e4e6b79bc

Observation 63bb4f11-83cd-4faa-8b6c-af94bbb4ac6d · outbound

This paper cites Food and Drug Administration.

Can Large Language Models Match the Conclusions of Systematic Reviews? Food and Drug Administration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.851003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.333145Z digest=sha256:4313e5747ba9d699d7dfe108ac346c5317e048c065ed0798b3fa68ed30ed4f37

Observation fd923c07-34dd-4cab-ae8e-3e618ca7096e · outbound

This paper cites Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report.

Can Large Language Models Match the Conclusions of Systematic Reviews? Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.369820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.369820Z digest=sha256:bd234649094138037323e54978b672d23b31b618c49155ba6ff794c8fe6d7e1d

Observation 35007af2-d8e3-4837-b6b3-21daba273a0b · outbound

This paper cites Can large language models reason about medical questions? Patterns , 5(3), 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can large language models reason about medical questions? Patterns , 5(3), 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.753052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.424003Z digest=sha256:0dc0d695a7172346e7e07abdf2d45fdfd0b91856cef9da5a73d1b556c47d628c

Observation 182459d2-29d9-417e-9a0e-fbbf805386af · outbound

This paper cites Medalign: A clinician-generated dataset for instruction following with electronic medical records.

Can Large Language Models Match the Conclusions of Systematic Reviews? Medalign: A clinician-generated dataset for instruction following with electronic medical records

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.618921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.533830Z digest=sha256:59ef99bde23375b9672d0724c7dd48feef76f46aa4dca1bdb87dc6e9207b0e6a

Observation 85ac5650-3983-4a41-90cf-dd91d884f93c · outbound

This paper cites Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.445363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.587478Z digest=sha256:f7d8b94af655e46b2c2b874997e40f11fd4334457726c4b67d627100693a5c03

Observation 2e81b9be-4c87-4a88-86df-6691597059e5 · outbound

This paper cites Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.285844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.700952Z digest=sha256:df82d5b253145e535e376b26355c51599ac68ab9ce6f4db4133d358d361c2c1d

Observation 9f2c0cdb-5233-4d15-9719-d9efeaca81e6 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:50.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.778063Z digest=sha256:dbb60b1911244fb03333f3219b2c836ac51ebb465ec7ed160d58862953b7d7b8

Observation 38605347-300d-4d31-a656-f16fa7949a8f · outbound

This paper cites Assessing the risk of bias in randomized clinical trials with large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessing the risk of bias in randomized clinical trials with large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.934019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.850157Z digest=sha256:b3b32055a74b0fe87accae31913f11b7c131e2e1034e59f8bb7df77044660565

Observation 98e86c26-42ec-4b84-b016-a820ada89bb2 · outbound

This paper cites BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature.

Can Large Language Models Match the Conclusions of Systematic Reviews? BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.893682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.893682Z digest=sha256:712aabb13453083fde3375abd7cd94769f26c29266a193944211fcf4a7ece435

Observation d67b7e0e-e96d-483a-86f3-58ec3d060384 · outbound

This paper cites o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \.

Can Large Language Models Match the Conclusions of Systematic Reviews? o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.738498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.975698Z digest=sha256:fb6918a4c887e01447881c23eddfc6143b0c9ebd9e007f686a4681412c71b5ce

Observation 25f5ae01-0e52-451e-8498-0e05348f9d19 · outbound

This paper cites Generative artificial intelligence use in evidence synthesis: A systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Generative artificial intelligence use in evidence synthesis: A systematic review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.615255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.044004Z digest=sha256:3b69ef73a5580d64d6f0a8709ddbe9f7307a656aa71f89a3b658cbd2e3bbaa4c

Observation ba71cfcd-0fdc-43bd-b69e-7113e17e0ee8 · outbound

This paper cites M ed REQAL : Examining medical knowledge recall of large language models via question answering.

Can Large Language Models Match the Conclusions of Systematic Reviews? M ed REQAL : Examining medical knowledge recall of large language models via question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.417678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.127980Z digest=sha256:7095558703c026a7aae0d32d7213b02ea8dd5962afb77fb115edc3fe0d86b300

Observation 4ee122c0-75c4-4ee7-9002-b1a475e4b530 · outbound

This paper cites H ealth FC : Verifying health claims with evidence-based medical fact-checking.

Can Large Language Models Match the Conclusions of Systematic Reviews? H ealth FC : Verifying health claims with evidence-based medical fact-checking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.287045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.193185Z digest=sha256:d71319481595fedfda1dec0737abc724d42e5a159f8222c5b623f51270dab7f9

Observation 116dc038-7c89-44ff-94f7-91293199787a · outbound

This paper cites What evidence do language models find convincing?, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? What evidence do language models find convincing?, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.071576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.238814Z digest=sha256:ec8a80a9e6afb44ef5cb006fa5a3ef353c1719841772849db472e40631ac4bfb

Observation 1d2fea88-ace9-4c10-86be-61cdf76911e7 · outbound

This paper cites Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.888788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.332339Z digest=sha256:090d8b135da91d00d9472820fa34a9edeb678a838119411a9120e08a436c5da7

Observation 2fe2347a-368f-481e-9843-adec69ddf050 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.713111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.408496Z digest=sha256:219b22ab0cb1749deb68bc50cff5e8e40dc101c1aa5f4026a94d10f97ddd53f6

Observation 153c3248-884e-4b22-b754-fe928d4bb7b1 · outbound

This paper cites Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.535178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.481306Z digest=sha256:2ad33442ba0a3d415c247d72dc21f390c8375ef51f99cec3f118f8988d3d7b33

Observation c70947c1-3b0c-4aa1-b33e-b3a69558d131 · outbound

This paper cites How to write a cochrane systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to write a cochrane systematic review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.332514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.544854Z digest=sha256:7e437f7b6b55c45eee585188b51391278cbcfdac8521a5d90eb4b6d1f8e2301c

Observation b0a7dd8d-16d1-4069-84d7-ff2ff483006e · outbound

This paper cites Quality of cochrane reviews.

Can Large Language Models Match the Conclusions of Systematic Reviews? Quality of cochrane reviews

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.191484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.584060Z digest=sha256:f9c11f7346a8af75b6aceae81e74a16cf93d83e5dc3e9e4291196f72084886d1

Observation 0ee4b7c3-0587-4990-a353-0b1af38a86f4 · outbound

This paper cites What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011.

Can Large Language Models Match the Conclusions of Systematic Reviews? What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.680202Z digest=sha256:cac84034556ef0f07076d8fd31c57fbf647c2d411de51d7e2ee6b854d1e0f4c2

Observation f5763b0a-a6c1-4835-8b0a-348f3ea98e09 · outbound

This paper cites Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.846344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.833178Z digest=sha256:64c4a465c72a937491993b578a528ad16bf3eb96d7cb8102ea5292e133d09db4

Observation 57ebdff6-cfed-4775-83c3-ef9ba273a830 · outbound

This paper cites Bethesda (MD): National Center for Biotechnology Information (US), 2010-.

Can Large Language Models Match the Conclusions of Systematic Reviews? Bethesda (MD): National Center for Biotechnology Information (US), 2010-

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.705602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.916594Z digest=sha256:b643aed85a55dcb8c2a3edfd25203f2eb2671e7b9c68a6cc66896d615a9c77a0

Observation ddc25f37-4587-4ab6-9032-9003f348532f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:47.583430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.012378Z digest=sha256:c55367bcb6585d07a5d4db16c77eb16274c45a84b0496cc5c2c423209ae416ba

Observation 2eebce0a-e7a3-4f98-b7aa-e7752544f1f9 · outbound

This paper cites Assessment of the strength of recommendation and quality of evidence: Grade checklist.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessment of the strength of recommendation and quality of evidence: Grade checklist

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.424267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.097237Z digest=sha256:acfcb6b5618f8f2ac2e0a3dbf7a64b911a6eee26b8108545fdd69da1159e9cad

Observation ff297756-cfd5-48f9-916d-c8f0ef492680 · outbound

This paper cites Openai o1 system card, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openai o1 system card, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.179454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.179454Z digest=sha256:cfb94de9716201b254caec1acf94f7165a161dae0bb60416e37d83081b0e162d

Observation 052921c6-0370-49bf-ad06-b2da74cec50a · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.270040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.270040Z digest=sha256:2ca6bcdbebeee200b1e24ff6cdde8c80658070e10f778eabf7b2863e1a8e1e56

Observation a2576a7f-030f-4e82-b9a4-b5681c7e2546 · outbound

This paper cites Open Thoughts.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open Thoughts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.358789Z digest=sha256:4a81be4b2cc29ddda231886c8a3c85c27f36b238b63f01138a9577a7349ca8c4

Observation 74080c23-941b-4d06-9368-74c6f7e90f55 · outbound

This paper cites Gpt-4 technical report, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gpt-4 technical report, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.428543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.428543Z digest=sha256:edae367685f5fb686e84ef0d75be8053bd73dce823e1f4adbaff406b77ded5c7

Observation 305a4698-4bc5-4ff5-8564-cad8cac40ee3 · outbound

This paper cites Qwen3, April 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen3, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.486772Z digest=sha256:98d64c8a024708ef1e22f08ae34e79cbe8d42aed4e9cf73080b3281c79672480

Observation 4f7c6598-9e4c-44fa-b82c-930f310679ac · outbound

This paper cites The llama 4 herd, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 4 herd, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.057232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.547522Z digest=sha256:e7a587a2e8a14891e8e332c2fe77fe00a1490e7d9fb0faf1c1f0fade4d37a4f3

Observation f7222a5f-1b41-44f6-817a-c64f0cd26817 · outbound

This paper cites Huatuogpt-o1, towards medical complex reasoning with llms, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Huatuogpt-o1, towards medical complex reasoning with llms, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.623175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.623175Z digest=sha256:ad183c4dd6395aa1916db996c7a80e9d616e13f4581ab1ab8f15d3c1719c0d2e

Observation 935206a9-ac73-4ecf-90e1-0bddbf237818 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openbiollms: Advancing open-source large language models for healthcare and life sciences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.888598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.724796Z digest=sha256:a879295ace13d249ebd5d42ab00bd8928df798d8410f2e52ee57af288fbffdd8

Observation 01a60bc0-9a29-433e-9a9d-b6db76a4c5b4 · outbound

This paper cites Refinedocumentschain.

Can Large Language Models Match the Conclusions of Systematic Reviews? Refinedocumentschain

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.742395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.817332Z digest=sha256:246786cea8cd49622f4efcdceac270828cbc2ddf818d05be0aed4c19d71d62e5

Observation 0b21fda6-023b-4ded-8cba-254e44ea342a · outbound

This paper cites An introduction to the bootstrap.

Can Large Language Models Match the Conclusions of Systematic Reviews? An introduction to the bootstrap

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.604185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.916032Z digest=sha256:7f0884a45f465701b283a15c15ec582e2b46f1663ec1d0cc64899ada84204299

Observation d6ba0e78-aec5-4990-8dad-eee64931b492 · outbound

This paper cites Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.983710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.983710Z digest=sha256:66e5fe82e60347c763a127e54a44ace924a3c6b7583230238dc82cafc12e4f82

Observation a9a25759-e103-4159-9809-9ffbe1bc9358 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long-context LLMs Struggle with Long In-context Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.055957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.055957Z digest=sha256:ae8976b7dca8438aec182d6255bf69e889c98841f23a6c83f478fbbadb15cf85

Observation c1bc1f40-dfa1-4665-8f19-8e8aece3d297 · outbound

This paper cites Large language models are overconfident and amplify human bias.

Can Large Language Models Match the Conclusions of Systematic Reviews? Large language models are overconfident and amplify human bias

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.149556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.149556Z digest=sha256:34d1646626ab8d193ed9816e1156a2adf678ebbc4f5a05debcd68ba575f0d35e

Observation f740f987-927e-461a-aba9-43f980797571 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.189971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.189971Z digest=sha256:f1e7a8cbb6c4ec63fae640da88c8502abc115e0e326710d7652f7b6a027c58ca

Observation 607a5fc8-c3fc-456a-bcae-c39cd7a68ee4 · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Can Large Language Models Match the Conclusions of Systematic Reviews? Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.303527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.303527Z digest=sha256:ea40b6bae47341f6a57fa51597d36d1e16858f491ab98cbce4d8e3920c1cdf18

Observation c88dedd6-9da0-4f3b-bbde-83e6c12e97fa · outbound

This paper cites Fine-tuning is fine, if calibrated.

Can Large Language Models Match the Conclusions of Systematic Reviews? Fine-tuning is fine, if calibrated

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.443913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:05:45.381013Z digest=sha256:eff65d070c32ff992d0499f2279bd86aac13a4533a4001c2ded925afde85ce61

Observation 8ec6626c-2a83-4b3f-978a-1cc08e17c7ec · outbound

This paper cites Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data.

Can Large Language Models Match the Conclusions of Systematic Reviews? Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.447639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.447639Z digest=sha256:70fe4ccc9ecba250be2bfa30d1f5136c4ccaba089011f0b4092caaec2434a0ed

Observation fd2586f8-4202-4c76-92c1-6f458ae2bf17 · outbound

This paper cites FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?.

Can Large Language Models Match the Conclusions of Systematic Reviews? FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.491178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.491178Z digest=sha256:87abf8f52a64ff30dec2ae7bff4395a62ddfb0fbafa405d4424c9d29b1fdc9c2

Observation d4fce9a9-c421-4326-add9-a0d887425872 · outbound

This paper cites Deepseek-v3 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-v3 technical report, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.578828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.578828Z digest=sha256:9f1009504bf3a25b60b6a737ce5fa64299c843a66a96f5ec2918a0b398537023

Observation 0fb222eb-0441-4a2d-a7ac-c47d6547eed8 · outbound

This paper cites The llama 3 herd of models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 3 herd of models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.620661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.620661Z digest=sha256:3655327f6916c2319d4a23b21fba97c40f8545dbc0271e56bee5a44092caa8d7

Observation adbe0757-7821-4c02-bd0d-a20069d1bb3f · outbound

This paper cites Qwen2.5 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen2.5 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.680263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.680263Z digest=sha256:cbd0b837e50e17bf64dce943eef432ad08c132ccc5a3449572a170909fc04c72

Observation 5c7b44d0-a5e8-48b7-b2d3-b21b530bc83d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.754584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.754584Z digest=sha256:dc356dcb87f9d77cb259db04277a9c4abc87c085cf6555914758de2204b9571e

Pith citing papers

Observation b9ba4ec6-9052-4ae7-a9b4-e667da51b05f · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.764062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:9058be0ae93a499ff6e3be28588ac4416f82c9ef3b7ef41d8a2c3ba21ad1083b

Observation 0aeb246b-c148-408e-aa59-a286729bfe92 · inbound

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison cites this paper.

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.026099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T06:03:59.798126Z digest=sha256:daaed680bfdb424b6bccc3ad4182d758d56119f507a4c75d2b6f10855b09b664

Observation 94e14c8b-1f7f-4c93-9106-9f4d9a7c0e67 · inbound

Can AI Agents Synthesize Scientific Conclusions? cites this paper.

Can AI Agents Synthesize Scientific Conclusions? Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.435176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:05:07.718882Z digest=sha256:45a64615215e096a9cbf0565d5db0549598ec832b0286092b48320b080666ff8