Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Match the Conclusions of Systematic Reviews?

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2505.22787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22787 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:05:45.754584Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:03:59.798126Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.432897Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abfb9d70-afe7-4cb1-855e-741140657f6c · outbound

This paper cites Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases.

Can Large Language Models Match the Conclusions of Systematic Reviews? Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.999188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.627088Z digest=sha256:d3113158778699306ba19fad0de7ac61e902422b84f45e2eb731c8af5c3e2c8d

Observation 8d6b0df0-f2ba-4cb7-817e-201fb4cac29a · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:51.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.702126Z digest=sha256:e1119fe53153c5b0c49b8cacf6a88d419c39f717f0f95b05ad49aee628fd1ff9

Observation 8823c06f-d137-4b0f-991b-818d8f54718c · outbound

This paper cites The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:41.788855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:41.788855Z digest=sha256:e88cd5aaf927b2a63a8544e8a05b99f04806b28f635c3db75246be82e7f7d6e5

Observation 22686ff8-3484-41d7-a983-d27f398c74c3 · outbound

This paper cites How to optimize the systematic review process using ai tools.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to optimize the systematic review process using ai tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.728714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.857108Z digest=sha256:3eabb501df1b728dd1fdcfd890f404f01a345044883aaa2fdbedfad35a3f1153

Observation 3ef3b132-53c3-4c9f-b155-040c13c904d1 · outbound

This paper cites Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses.

Can Large Language Models Match the Conclusions of Systematic Reviews? Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.608818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.917310Z digest=sha256:c724b56d813a78224b9c8174de2ed1b894ee8e07d34d826011361cd64053011e

Observation 41609f5e-f087-4fa1-a9c7-11f8ba592707 · outbound

This paper cites Deep research system card, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deep research system card, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.433866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.001265Z digest=sha256:21be110fe1487f74edd36b13236c1638f33b5c9ddc7e53761114d2824e0a1060

Observation 94eaa987-14a8-4cf7-9b2b-938abb39df49 · outbound

This paper cites Gemini deep research – your personal research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gemini deep research – your personal research assistant, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.069304Z digest=sha256:a5fd78aebf2484def9205f44a9ff73a01f6c9092c0b4b02d0e9a8288b9b76034

Observation 0f59b63d-401a-4432-b88c-b2d27d28c6cc · outbound

This paper cites Elicit: The ai research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Elicit: The ai research assistant, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.172471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.163546Z digest=sha256:33152f57812fabe4c1ea34cdc223a8fa8f98a8a42e87522e6f4f82d202971097

Observation 00702845-aabc-4db0-8d57-5d07882d3dff · outbound

This paper cites Open evidence: Ai-powered medical information platform, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open evidence: Ai-powered medical information platform, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.025871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.237160Z digest=sha256:d3ba1730b705a6808486c62ad25095ed6a49d6b26b0a15f2e77c373b17b3a32b

Observation 63bb4f11-83cd-4faa-8b6c-af94bbb4ac6d · outbound

This paper cites Food and Drug Administration.

Can Large Language Models Match the Conclusions of Systematic Reviews? Food and Drug Administration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.851003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.333145Z digest=sha256:26846a37ea1ff1da4ffbabd5eacb0fd5185149ef3454376390b95c56c307f4bf

Observation fd923c07-34dd-4cab-ae8e-3e618ca7096e · outbound

This paper cites Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report.

Can Large Language Models Match the Conclusions of Systematic Reviews? Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.369820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.369820Z digest=sha256:36efa67144a1db813eea37b9e56d2cb1c3fcd40b3862313bf1c5a3675b23deb8

Observation 35007af2-d8e3-4837-b6b3-21daba273a0b · outbound

This paper cites Can large language models reason about medical questions? Patterns , 5(3), 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can large language models reason about medical questions? Patterns , 5(3), 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.753052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.424003Z digest=sha256:70e57b81ec1ce9457caf58340fb1d4a0e0d09bdd06e953fcec73681b381f2b79

Observation 182459d2-29d9-417e-9a0e-fbbf805386af · outbound

This paper cites Medalign: A clinician-generated dataset for instruction following with electronic medical records.

Can Large Language Models Match the Conclusions of Systematic Reviews? Medalign: A clinician-generated dataset for instruction following with electronic medical records

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.618921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.533830Z digest=sha256:0f87098a9eb5b34ab5c393642baeaa9d9cdfefba8a80ff6b96e2abe47664a343

Observation 85ac5650-3983-4a41-90cf-dd91d884f93c · outbound

This paper cites Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.445363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.587478Z digest=sha256:0102a0023bbd6a3ea52c9665981f705a41b8cb1ad4b2e1179a83d0a9dbe18766

Observation 2e81b9be-4c87-4a88-86df-6691597059e5 · outbound

This paper cites Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.285844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.700952Z digest=sha256:65bfbe03d7a8c73c1fd488ab2b10306cb850b1e38ae7905971c1adc4f89c7bbc

Observation 9f2c0cdb-5233-4d15-9719-d9efeaca81e6 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:50.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.778063Z digest=sha256:8f875d42d05952dc896c405586eb8dba8a2aa9ff86756510e983902800fb7e04

Observation 38605347-300d-4d31-a656-f16fa7949a8f · outbound

This paper cites Assessing the risk of bias in randomized clinical trials with large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessing the risk of bias in randomized clinical trials with large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.934019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.850157Z digest=sha256:ca579c0927ec44728601c13e6c59f9426648b5c891c54b48d43e7ed7434870c4

Observation 98e86c26-42ec-4b84-b016-a820ada89bb2 · outbound

This paper cites BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature.

Can Large Language Models Match the Conclusions of Systematic Reviews? BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.893682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.893682Z digest=sha256:0ffd6e485c1f000d7b30300a51e98a2695e0128c5c4d4a81028c988b530e3748

Observation d67b7e0e-e96d-483a-86f3-58ec3d060384 · outbound

This paper cites o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \.

Can Large Language Models Match the Conclusions of Systematic Reviews? o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.738498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.975698Z digest=sha256:70a6f01b13e92dd12ed7db40d7883c061aa04d27f77657e5e564462530b69eb7

Observation 25f5ae01-0e52-451e-8498-0e05348f9d19 · outbound

This paper cites Generative artificial intelligence use in evidence synthesis: A systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Generative artificial intelligence use in evidence synthesis: A systematic review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.615255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.044004Z digest=sha256:c3eecde5713012e6aaf31c91d3de537658adf2e408490ffb443cb82e4ef52b3b

Observation ba71cfcd-0fdc-43bd-b69e-7113e17e0ee8 · outbound

This paper cites M ed REQAL : Examining medical knowledge recall of large language models via question answering.

Can Large Language Models Match the Conclusions of Systematic Reviews? M ed REQAL : Examining medical knowledge recall of large language models via question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.417678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.127980Z digest=sha256:a328be201b22cf8f4780303c392d7f23e70556eb60a2c991bc70aca63f478908

Observation 4ee122c0-75c4-4ee7-9002-b1a475e4b530 · outbound

This paper cites H ealth FC : Verifying health claims with evidence-based medical fact-checking.

Can Large Language Models Match the Conclusions of Systematic Reviews? H ealth FC : Verifying health claims with evidence-based medical fact-checking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.287045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.193185Z digest=sha256:90c4cc3b127d2845813d76b1f2eb635f0c7869a89f2545df476a1fcb78977bea

Observation 116dc038-7c89-44ff-94f7-91293199787a · outbound

This paper cites What evidence do language models find convincing?, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? What evidence do language models find convincing?, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.071576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.238814Z digest=sha256:799090d8b723eee1d41b11d6c7f94626fa8c33577cfcbee9b8d164c5f09de0ec

Observation 1d2fea88-ace9-4c10-86be-61cdf76911e7 · outbound

This paper cites Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.888788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.332339Z digest=sha256:e179369a45ef398585c28ac21e7d89b8c3f3c2f449d4fd7925bb390863935a68

Observation 2fe2347a-368f-481e-9843-adec69ddf050 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.713111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.408496Z digest=sha256:e154af8700064ce10accf65f6ffe01e23b70f00b544ffafb8bb4d608d3c41b58

Observation 153c3248-884e-4b22-b754-fe928d4bb7b1 · outbound

This paper cites Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.535178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.481306Z digest=sha256:b4746e04262694bac686e15e276574ae3a94a96f81da8511b0d8e1e9225bec0b

Observation c70947c1-3b0c-4aa1-b33e-b3a69558d131 · outbound

This paper cites How to write a cochrane systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to write a cochrane systematic review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.332514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.544854Z digest=sha256:cdc278e706870c1c79473d209f8767e64d11ab905e1b9b9cd8e2775971b6e7df

Observation b0a7dd8d-16d1-4069-84d7-ff2ff483006e · outbound

This paper cites Quality of cochrane reviews.

Can Large Language Models Match the Conclusions of Systematic Reviews? Quality of cochrane reviews

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.191484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.584060Z digest=sha256:7cc74790d7f6027eb366aa4e461ca2da79a7645de217e6bee34a9e14e171d1b9

Observation 0ee4b7c3-0587-4990-a353-0b1af38a86f4 · outbound

This paper cites What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011.

Can Large Language Models Match the Conclusions of Systematic Reviews? What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.680202Z digest=sha256:247302e1aad890db0cbba7fef16ded09265256f8ca8b58d97609a2bfa7e3c2eb

Observation f5763b0a-a6c1-4835-8b0a-348f3ea98e09 · outbound

This paper cites Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.846344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.833178Z digest=sha256:fbdfc4221b55fc01cc8c202ee6abd1fc2a3f9d95c6de747fc54e7a8f83f9cd6b

Observation 57ebdff6-cfed-4775-83c3-ef9ba273a830 · outbound

This paper cites Bethesda (MD): National Center for Biotechnology Information (US), 2010-.

Can Large Language Models Match the Conclusions of Systematic Reviews? Bethesda (MD): National Center for Biotechnology Information (US), 2010-

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.705602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.916594Z digest=sha256:53b5e335603122649af0a66d24abe0f625ef0b9623a4745fc91381569a005fdd

Observation ddc25f37-4587-4ab6-9032-9003f348532f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:47.583430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.012378Z digest=sha256:f949f48990851f23d7eeafd334a3349d2307a2c78f8ff544c673271f705b19e8

Observation 2eebce0a-e7a3-4f98-b7aa-e7752544f1f9 · outbound

This paper cites Assessment of the strength of recommendation and quality of evidence: Grade checklist.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessment of the strength of recommendation and quality of evidence: Grade checklist

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.424267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.097237Z digest=sha256:8e42e2087d8e7b4513f234796de75d66f2e84dcfbd35aca71c8874b0430a8188

Observation ff297756-cfd5-48f9-916d-c8f0ef492680 · outbound

This paper cites Openai o1 system card, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openai o1 system card, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.179454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.179454Z digest=sha256:190bdaf70db8dde9eec11db9e86bbb3f8afe63727bedf0c6aa1799550fa474ab

Observation 052921c6-0370-49bf-ad06-b2da74cec50a · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.270040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.270040Z digest=sha256:23d95ce723f4e41ff3682c536ddf078c694e47806e2022a286fdf7a121398136

Observation a2576a7f-030f-4e82-b9a4-b5681c7e2546 · outbound

This paper cites Open Thoughts.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open Thoughts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.358789Z digest=sha256:62980e39efe6decd969576fc1cac904ba7986b2540e239df431592373a126bf0

Observation 74080c23-941b-4d06-9368-74c6f7e90f55 · outbound

This paper cites Gpt-4 technical report, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gpt-4 technical report, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.428543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.428543Z digest=sha256:b07bfdf104296275636a3fea74730ee2bc31c6c06242955fc863c6b626207494

Observation 305a4698-4bc5-4ff5-8564-cad8cac40ee3 · outbound

This paper cites Qwen3, April 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen3, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.486772Z digest=sha256:1358e2a4c8cdb460ca5332e4964ddd70ef3fbeed4e17977e38e28b2d3811257d

Observation 4f7c6598-9e4c-44fa-b82c-930f310679ac · outbound

This paper cites The llama 4 herd, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 4 herd, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.057232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.547522Z digest=sha256:54a277f1609e1f4bbd8a3464d2009b8cdfb16b79158287925bc7cc8f91c534dd

Observation f7222a5f-1b41-44f6-817a-c64f0cd26817 · outbound

This paper cites Huatuogpt-o1, towards medical complex reasoning with llms, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Huatuogpt-o1, towards medical complex reasoning with llms, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.623175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.623175Z digest=sha256:708ab39deba810817f72992cc98bac78d99c2c4d54a67bc4e3fa35e8a4c6c27c

Observation 935206a9-ac73-4ecf-90e1-0bddbf237818 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openbiollms: Advancing open-source large language models for healthcare and life sciences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.888598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.724796Z digest=sha256:9aa99acbf76e88787cb9b036e2dfa686e9072854ac519075954a8135a93bc8a0

Observation 01a60bc0-9a29-433e-9a9d-b6db76a4c5b4 · outbound

This paper cites Refinedocumentschain.

Can Large Language Models Match the Conclusions of Systematic Reviews? Refinedocumentschain

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.742395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.817332Z digest=sha256:39710463292316c7e9f4c3921140bd0c7f4301d585184aed82ef8caded101fe7

Observation 0b21fda6-023b-4ded-8cba-254e44ea342a · outbound

This paper cites An introduction to the bootstrap.

Can Large Language Models Match the Conclusions of Systematic Reviews? An introduction to the bootstrap

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.604185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.916032Z digest=sha256:92b2d45c88dca28434b46ac8c9dce0b1ed722386a8143b0e4524eec43b78aa5a

Observation d6ba0e78-aec5-4990-8dad-eee64931b492 · outbound

This paper cites Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.983710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.983710Z digest=sha256:9b3a3a9cbf0d39a353059c467b5c103fc3239de697b7015f3c553a4047bb771e

Observation a9a25759-e103-4159-9809-9ffbe1bc9358 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long-context LLMs Struggle with Long In-context Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.055957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.055957Z digest=sha256:d8a18ea32b1d728d152c2148f3a485be5acbfc53db24f5b37be0eb2c47d3daf4

Observation c1bc1f40-dfa1-4665-8f19-8e8aece3d297 · outbound

This paper cites Large language models are overconfident and amplify human bias.

Can Large Language Models Match the Conclusions of Systematic Reviews? Large language models are overconfident and amplify human bias

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.149556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.149556Z digest=sha256:0e6cc9bae9b55391ce44809c15a5e57c2de48800f6012a1dbf9ef322f9bd465a

Observation f740f987-927e-461a-aba9-43f980797571 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.189971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.189971Z digest=sha256:df94791d0ad4dca6d96bd1f0291d92217bace182ee0090ad819d402ba3b9dc8c

Observation 607a5fc8-c3fc-456a-bcae-c39cd7a68ee4 · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Can Large Language Models Match the Conclusions of Systematic Reviews? Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.303527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.303527Z digest=sha256:e7018257d042e432f6a32fe4a5d2f02c83abcd8b979217fc6b4738bbfcff2d53

Observation c88dedd6-9da0-4f3b-bbde-83e6c12e97fa · outbound

This paper cites Fine-tuning is fine, if calibrated.

Can Large Language Models Match the Conclusions of Systematic Reviews? Fine-tuning is fine, if calibrated

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.443913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:05:45.381013Z digest=sha256:fa283271813182eced159f8934f856cd3733f50af81750cd561b45f056daa437

Observation 8ec6626c-2a83-4b3f-978a-1cc08e17c7ec · outbound

This paper cites Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data.

Can Large Language Models Match the Conclusions of Systematic Reviews? Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.447639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.447639Z digest=sha256:b54dfe02a54a7603c6af5fb8e2e7c4412684d0e5ae91f70192d390083504ac96

Observation fd2586f8-4202-4c76-92c1-6f458ae2bf17 · outbound

This paper cites FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?.

Can Large Language Models Match the Conclusions of Systematic Reviews? FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.491178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.491178Z digest=sha256:e3ede7eaf95247e5abb3c8138f7d86188a9b78a28a0aebdee537e592ef543a65

Observation d4fce9a9-c421-4326-add9-a0d887425872 · outbound

This paper cites Deepseek-v3 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-v3 technical report, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.578828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.578828Z digest=sha256:f217b22240bcfb297e6b181ba1ecf263672e43dcce8e43142393647eb67d5a3f

Observation 0fb222eb-0441-4a2d-a7ac-c47d6547eed8 · outbound

This paper cites The llama 3 herd of models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 3 herd of models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.620661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.620661Z digest=sha256:6153e9ab96d2c901f0195456066f79d728c03b73cf7aff1d4a2955806afcaecb

Observation adbe0757-7821-4c02-bd0d-a20069d1bb3f · outbound

This paper cites Qwen2.5 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen2.5 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.680263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.680263Z digest=sha256:c559db53ce25dc9666fc1339c53fe66309a1cadd6ac48591cd58bba09b8228ab

Observation 5c7b44d0-a5e8-48b7-b2d3-b21b530bc83d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.754584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.754584Z digest=sha256:54e26f883976e276c64257e8528b915d424d28ed0a20b54ea638957ca06eafb4

Pith citing papers

Observation b9ba4ec6-9052-4ae7-a9b4-e667da51b05f · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.764062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:2268c7031154a145aa21821d38b217344931100e1db2f077b1bb74c77cc1d400

Observation 0aeb246b-c148-408e-aa59-a286729bfe92 · inbound

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison cites this paper.

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.026099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T06:03:59.798126Z digest=sha256:3b38a683b48edc581fc1af801885e28688ddf8e2311a3a73fbb386e46a1a0664

Observation 94e14c8b-1f7f-4c93-9106-9f4d9a7c0e67 · inbound

Can AI Agents Synthesize Scientific Conclusions? cites this paper.

Can AI Agents Synthesize Scientific Conclusions? Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.435176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T13:05:07.718882Z digest=sha256:caf7ee7474175e0c634ce9974609bb4b1072a70dff5736b6cc3032c7bf1234f8