Pith. sign in

Paper Citation Record · LEDGER

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

As of 15 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2507.15152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15152 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:46:08.376749Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:53.040834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:13:54.647415Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact16
  • verified fuzzy38
  • unresolved37
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59abdfec-2b79-400f-abb4-6896d381a4af · outbound

This paper cites Research Synthesis and Meta-Analysis: A Step-by-Step Approach.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Research Synthesis and Meta-Analysis: A Step-by-Step Approach

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.653529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.653529Z digest=sha256:9523da5c8397e758d5c11ec24d4d0fd96f411029fd1d5a01ad34cbd0f0b4bddc

Observation 1f67cc0f-6e6e-4c9c-8dd2-dd9074bee257 · outbound

This paper cites Analysing data and undertaking meta-analyses, chapter 10, pages 241–284.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysing data and undertaking meta-analyses, chapter 10, pages 241–284

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.664179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.664179Z digest=sha256:7df3f4aec032481299045f32978d9fc194538919335d50478a41e459449099a1

Observation d0762771-aa93-47e0-aeb8-895ffc8ebf2c · outbound

This paper cites Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.671805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.671805Z digest=sha256:6318e86ff6b4e35ef420086e4f6dbd148f870f651c716e5203b5179ddb86823e

Observation efdddffb-8142-469b-909e-5f43339c72f5 · outbound

This paper cites Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.038582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.680841Z digest=sha256:363cc34b30b098ef95e54d291075980a7516f0a143bf0724a010eb92b5a049f3

Observation 70ea14fd-5453-4d30-8447-05c6b28e9f97 · outbound

This paper cites Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.020674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.685632Z digest=sha256:d8616651a64e4063732ffb7b41a1c8bdc40864f91c0d3c061992c3e18f0f28dc

Observation 975c4024-e2d0-4308-ada1-f84d80119446 · outbound

This paper cites Toward systematic review automation: a practical guide to using machine learning tools in research synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward systematic review automation: a practical guide to using machine learning tools in research synthesis

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:07.690733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.690733Z digest=sha256:7abda77b656381341825d4a6bb3480a1767d86d4bf0c8a0c1551be4d5cdbf644

Observation b0629e68-bf64-4206-99b4-91252af8cde9 · outbound

This paper cites Exact: automatic extraction of clinical trial characteristics from journal publications.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exact: automatic extraction of clinical trial characteristics from journal publications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.697875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.697875Z digest=sha256:9af18ac884de50aa2b3bfd91e5e3b5cde4c5e4710644d6d23492d9e4991a23b8

Observation fdd4bafa-4184-4928-a2ff-b8fe42880703 · outbound

This paper cites Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz

Reference 8

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.968448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.746197Z digest=sha256:a700fb148d9570901ff74943dba3b92dcfdf7b58195611392c8ddfb6f4b7396e

Observation ca059850-7817-4bff-a971-e79644f87bd6 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.459526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.763071Z digest=sha256:8589abef006b44c76bb4b0cb5250b18e470f1caf615b552bf31dd09f6cfd67d0

Observation 1c86dbba-aac8-4c3c-ac6b-60cf0a5c8cc7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.784022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.784022Z digest=sha256:447d335815c871d1808cf8671bd5dab60d92fdf17dec0a90851c7f3bc40a7b00

Observation bc08e3c2-ce9c-41ab-ae05-793e5c754ed0 · outbound

This paper cites Automating meta-analyses of randomized clinical trials: a first look.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automating meta-analyses of randomized clinical trials: a first look

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.801622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.801622Z digest=sha256:561766ac3ded77a9301b61ea970f1666db1e0d9dff1a346f7cf705044b32c745

Observation c032b9f2-8819-4519-aed0-7709bda04041 · outbound

This paper cites Katz-Rogozhnikov, Kush R.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Katz-Rogozhnikov, Kush R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.812345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.812345Z digest=sha256:f7da1ebfbac1334fa39368ba3cfcd739c98bfa186bf36a2f3c2527576ee9a65e

Observation ec473968-5324-44d2-a235-b93418bdb282 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.817822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.817822Z digest=sha256:d9caebde0ebd531cdfd52b61e41a3c137879b5230383c97c16d99aa9452253f1

Observation 25b1e172-1af5-4a79-be85-76ea657b1997 · outbound

This paper cites Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved].

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved]

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.928184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.824462Z digest=sha256:a3c0e1620529259492f111b110cd9b9aa196cad6ed067cdc1c7de95d248e4310

Observation 094ae696-9487-4bcc-8bfd-ea6e8255c58e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.829830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.829830Z digest=sha256:5618321a6ed82042fa2c406ef0ecb21bbb82ae3d5c7b4a87a7db436713204136

Observation d8fbea24-066b-455d-b1ca-bda23239e5b1 · outbound

This paper cites The data is in: Deciding when to automate screening in your slr, November 2023.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The data is in: Deciding when to automate screening in your slr, November 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.834832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.834832Z digest=sha256:d16d11fa274a912fdd84e6615f4f329f87ebaa9440281f1c9d6eba5e5730be62

Observation f63edf05-4f58-41af-89ae-082feb262b6e · outbound

This paper cites Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.839801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.839801Z digest=sha256:3919567fab7202de042d2fa31b337bbf393400c0e3e7c202c81a5679c708b873

Observation 28e34ed0-8485-4fcb-8124-2dad653175dd · outbound

This paper cites MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.854237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.854237Z digest=sha256:82fc84a4a6fb2b8688e126d0c27b6f7e35a9debdbc193ae008b149e58ac12b17

Observation e7c78217-16e8-4a6f-8fd3-64c498887a46 · outbound

This paper cites Chatgpt: Large language model (mar 14 version).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Chatgpt: Large language model (mar 14 version)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.864218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.864218Z digest=sha256:4fcffc854d8350448365ea7bc577284dd3fcaea08e3cbe6e7c9ddfb4dbcad5d1

Observation cac6f8e3-00ef-4460-b6f3-4ea7f1d4f000 · outbound

This paper cites Claude 2 model announcement.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Claude 2 model announcement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.870017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.870017Z digest=sha256:fd88b56b15479595c7946336b3ff4292a46faca9238e8a96e3e57157a90d0c48

Observation 583f778e-b019-49a5-aa43-88fadb7779a7 · outbound

This paper cites Zero-shot infor- mation extraction for clinical meta-analysis using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zero-shot infor- mation extraction for clinical meta-analysis using large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.875116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.875116Z digest=sha256:57e4022004aaa46f3aa4349d1fe5c71fa90619a103ff0aa85f1ae6a6638b3cd8

Observation 75556778-c65e-4b3b-8f2d-8b86090c2e79 · outbound

This paper cites Performance of two large language models for data extraction in evidence synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Performance of two large language models for data extraction in evidence synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.880456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.880456Z digest=sha256:98753d1b014661e8ad03bf789337271e7022563990425c0d841a6e7a93d521fe

Observation 6705cf30-2390-4d27-9f7e-b02c0a6a2512 · outbound

This paper cites Automatically extracting numerical results from randomized controlled trials with large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automatically extracting numerical results from randomized controlled trials with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.885286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.885286Z digest=sha256:229afc3e515962604ffbc9557934e0f7b8d2ebc7e11d14c438cd7c19bde26d0e

Observation fe259602-0acf-4aa5-9e96-25afc38508d9 · outbound

This paper cites Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.891121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.891121Z digest=sha256:956a5428316b635753cd6a4e9aa8ccadafa8bf34dd766e2202ffb659748f8acb

Observation 94bbea6d-b5f3-4320-89d7-15055bdf1e24 · outbound

This paper cites Lee, Shigeki Yamada, and Tomohiro Mizuno.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Lee, Shigeki Yamada, and Tomohiro Mizuno

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.896068Z digest=sha256:13a45683ab231e4c7cd0aeeb6fa84437ea8c5ef1e7df2479f3768ad6df1933b4

Observation 1cfbaa80-85fd-4a5a-97d0-24a0a022a479 · outbound

This paper cites Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis

Reference 26

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.370112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.901443Z digest=sha256:6677f7d97265c33efb2e1bc0db7f9aa3d03bc1789297b8f3ec00245a683583cd

Observation 96dc95e7-b78e-47d4-bdca-e36962197bcf · outbound

This paper cites Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris

Reference 27

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.829269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.906139Z digest=sha256:984c3036118e54ebf23b4dc0d4cade6c111d5e419be396f6f9000dff18ae7fc4

Observation 64105b66-d969-41af-a918-8da6e2c4b62b · outbound

This paper cites Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.285708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.913293Z digest=sha256:58bbf3c59e0a575e7f2765b61b8a680d876a8872f57084c0fad0e7a06411574e

Observation 8b698dba-6c1e-4100-a69b-c5908f7b0a20 · outbound

This paper cites Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim

Reference 29

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.799935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.918186Z digest=sha256:adecba6ddc10853b438cce8dbb41de1c3e018c0b63d375e8487543d965283898

Observation 5c6db0f3-9237-4145-a5d2-9b5c829d2f1e · outbound

This paper cites V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.774773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.924380Z digest=sha256:2b501034c89356d7df13a51468485fdca6a70e6f63bc61aece328fa9895d574a

Observation 34192b22-a459-443b-8287-0cf5870d8db2 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gpt-4o mini: Advancing cost-efficient intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.934074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.934074Z digest=sha256:ce36807d415e924c841dbed22b8ecff04e95c6f1396862864c9883b5efb1d32a

Observation c32e8170-e0d4-4921-9966-1a46b81557ab · outbound

This paper cites Gemini 2.0 flash.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gemini 2.0 flash

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.939758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.939758Z digest=sha256:7f6d2dfd65e0011a3fe10c7cd10b8794c56bb23620fb3b3a5b7538e2dc728d48

Observation 93682bef-d44e-4c7c-8a57-315a4becb015 · outbound

This paper cites Grok-3 language model.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Grok-3 language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.945503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.945503Z digest=sha256:925c35affd1e011a9bc077fa6da43830bb6faeccb6d879eaab506f1e29fcb273

Observation 8b9cbaff-4f41-4132-857c-c60de2b330d0 · outbound

This paper cites The impact of temperature on extracting information from clinical trial publications using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The impact of temperature on extracting information from clinical trial publications using large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.951933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.951933Z digest=sha256:47f408318cf8dc79de0b7f79ecbb696fcf26a811b8604d9180a1961a62e3a560

Observation 4bbf6358-214b-43d1-afdb-4f0aea3b2d51 · outbound

This paper cites AI-Assisted Data Extraction for Systematic Reviews in Education.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction AI-Assisted Data Extraction for Systematic Reviews in Education

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.958889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.958889Z digest=sha256:4b6a45a258ade03713a86443cf2bdf3a33b5c671f2f1275ade5864a33057d60b

Observation 178abf96-a100-40cc-b4ea-fefa51c92b6a · outbound

This paper cites Use gemini 2.0 to speed up data processing.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Use gemini 2.0 to speed up data processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.965054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.965054Z digest=sha256:97a8628918327fcb4e01c872224e4dbdd0c60230ef640aa8fe4cfde85efb8a10

Observation f654e4b5-931a-4750-a1f4-2c9e8ea67a82 · outbound

This paper cites Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.748746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.971513Z digest=sha256:4e52d782ab42ee119ca3289bfedc13f19135ba017e11f11ba33018428550c727

Observation 759dce94-632b-4544-a756-d882cdaf18b0 · outbound

This paper cites Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.347631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.976676Z digest=sha256:7204fd3dfc52b2698414cb85156624b4a73481645bea1472ad3f1298de6d4828

Observation fc9056f3-42b0-4a30-89cd-34b59992be83 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Reflexion: language agents with verbal reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.318292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.989820Z digest=sha256:f9a0bd1c55a83d7f67f2dbb22f9019cbf88faf263d09b3730ade938902abc2ca

Observation 2acec01b-d982-4110-aa42-d3a6d238ea8e · outbound

This paper cites Towards mitigating LLM halluci- nation via self reflection.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Towards mitigating LLM halluci- nation via self reflection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.994474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.994474Z digest=sha256:9651c64732f5fdfae326fec7bfa3758c93d83c48550f6c7f2c33179dc6c04698

Observation 6c7a93d2-95b1-41d6-8356-4ae8b04c0c95 · outbound

This paper cites When hindsight is not 20/20: Testing limits on reflective thinking in large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction When hindsight is not 20/20: Testing limits on reflective thinking in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.295126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.000826Z digest=sha256:7e523e81b406e10c1dc25b0d3fd09dfaa507c57a96b22631def8f03bceca10b7

Observation ccc81064-a620-4159-a80e-ba3512ca49d4 · outbound

This paper cites Dietterich.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Dietterich

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.267940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.011272Z digest=sha256:c937d05d630954a30e025b5c747a511f54312df0e844cf4eabde09bd6f6806e9

Observation 23f4b757-ac28-47ef-8306-fc3ec919bcb0 · outbound

This paper cites A survey on ensemble learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction A survey on ensemble learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.018396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.018396Z digest=sha256:cbbb9a29dec99e4ecde42fc05cded7883e50365bb7505170efe1cf85ee758b2a

Observation e84df5e4-15d8-4caa-aca3-aba9d2a57a96 · outbound

This paper cites Ensemble pretrained language models to extract biomedical knowledge from litera- ture.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Ensemble pretrained language models to extract biomedical knowledge from litera- ture

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.645522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.023179Z digest=sha256:f9d947188ec89f417bc5d224f0f8f18f171f7662df3a96bbae431213725a1f55

Observation 8d7ef2bc-7931-463f-922d-5202489506dd · outbound

This paper cites Zhang and A.L.P.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zhang and A.L.P

Reference 45

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.618170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.030348Z digest=sha256:2e577329c195803bd27a0c79f13c32489694a5769373f6cf176572c133ba7422

Observation 31adf290-91b9-4d02-862c-b19873f71cee · outbound

This paper cites Comprehensive testing of large language models for extraction of structured data in pathology.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Comprehensive testing of large language models for extraction of structured data in pathology

Reference 46

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.587520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.036361Z digest=sha256:386756fe8b9d51af9f59c4e08e74b3cac5d8f4ec00fe16148d2859b341f45b92

Observation d28dc553-4477-4edf-b7d3-e393a2f59d3f · outbound

This paper cites Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.546789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.041977Z digest=sha256:446e4e6b5f721d91296c163c88801809768c0b0cfdfcd8a1cc228a09d6a7a2d3

Observation 18e1f353-74a7-4807-89c8-99ee1df24e4f · outbound

This paper cites Match, compare, or select? an investigation of large language models for entity matching.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Match, compare, or select? an investigation of large language models for entity matching

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.243358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.048626Z digest=sha256:24f745b745332a6cab1129a68d1345d8e7612c742f5f76453027628a767187fd

Observation cc3c1a78-35ff-4e69-a324-9c56813ca94d · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 49

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.516502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.053892Z digest=sha256:d27fc241a20299a5455f2263919a4dac29715365113ff243424bd02dcb06f8dd

Observation 84af77db-91b6-4c1a-b14e-2095284f718e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.061584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.061584Z digest=sha256:6e9ff934ef336cfb6a711765e9639db33afd0ddc559d3a5866adff1805b76961

Observation 3f397eca-7685-4cf9-8b04-69ba6bc821d3 · outbound

This paper cites Guyatt, Andrew D.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Guyatt, Andrew D

Reference 51

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.492238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.068084Z digest=sha256:0592c1480eb692a4fe964b0dc3773d1331c65349af2dc950beb10ca0f354ebd0

Observation 8614a940-0e89-4af5-add4-f99fb77eb40c · outbound

This paper cites Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.075310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.075310Z digest=sha256:cfa96a49ff51aa17f1458ded0db25a55e5a0fc539d8e21e7d364b3b1308393ca

Observation 449df189-eab1-4838-8f8f-8e5dc1abdc35 · outbound

This paper cites Agentic reasoning: Reasoning llms with tools for the deep research,.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic reasoning: Reasoning llms with tools for the deep research,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.198891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.083557Z digest=sha256:32f85fef25b4e1cd40bbaaaa650a508f601a32af4aa39b52f7bf9bcf9fabef05

Observation 2aa669b4-d09e-46ea-aa25-0ac7f3da646a · outbound

This paper cites TART: An open-source tool- augmented framework for explainable table-based reasoning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction TART: An open-source tool- augmented framework for explainable table-based reasoning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.171211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.099934Z digest=sha256:51744e3fe4dfecb47221e2feff50beea53d67a4b8a59e555d3807fe83fdfe35e

Observation b81958ca-a660-4689-bb8e-db65fdbe01a3 · outbound

This paper cites Medical hallucination in foundation models and their impact on healthcare.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Medical hallucination in foundation models and their impact on healthcare

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.105709Z digest=sha256:86d14108f7bde081d240ec5e28db0109eb4990525a89199d8cbbae5f01c655d1

Observation cabcab11-7bd8-46b3-ae71-b44625490a07 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.111858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.111858Z digest=sha256:802cb8f2215aa56e62f9af2affd978a6a26b0957a237eb1d56ce0c9c1e0343d8

Observation 2b1898d8-f3f5-489c-9aee-2d67e449992e · outbound

This paper cites Potential roles of large language models in the production of systematic reviews and meta-analyses.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Potential roles of large language models in the production of systematic reviews and meta-analyses

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:08.118777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.118777Z digest=sha256:f2c326be66bc7c214c61286b4d12ef70b3994847152639799858ef3150ab6de8

Observation 93db9ab6-c6b9-4861-98f0-f20db0315ff6 · outbound

This paper cites justification.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction justification

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.153899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.124679Z digest=sha256:e9b7487a9ba026029f5f10f52c1178b2fe971a08d7227bf5032ce0d62990074f

Observation e1d54069-2ad1-44a4-9a0c-74c7a7d12d35 · outbound

This paper cites other_time_points.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction other_time_points

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.136844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.130975Z digest=sha256:c1ce542d5825b2279b974e411b990bc2b3e8c87163e01cafe39d4238e7b59a25

Observation 69682989-e509-4562-9ede-0699148ce6d3 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.117550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.136861Z digest=sha256:ce0a13cec1cfe56a2a0b933218335c87331ce905612b3ca6fdf4cc6aca527734

Observation 5af7e017-6330-4c4b-9a15-b03cafc5d496 · outbound

This paper cites null"`, NOT `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"`, NOT `

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.100782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.142821Z digest=sha256:91931b4e5b6d534e5d11da124f73fd478c0c152d7e6bf6a8b4500c8bb3964f52

Observation 502eff7e-a05c-4706-a710-8504908387b9 · outbound

This paper cites more common in the intervention group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction more common in the intervention group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.078892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.149296Z digest=sha256:d817ca6d5228b97417519558f02603bdbaefead2bcb0478bb748cbc156f18d7d

Observation 03ae9405-e1de-40c5-a58f-b5f4d3b0af26 · outbound

This paper cites data_conflicts.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction data_conflicts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.058081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.156022Z digest=sha256:a6699f023c48fe71f75c06feaef2b6bec5e9cf1dfbc8ed053f16502199d346b2

Observation cbdb9462-4a71-409a-9055-a8c6b0f3411e · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:46:10.036874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.167425Z digest=sha256:3e8d25f84efecfa5e6b8bacb92fa6dbf879d469ebcc2a1110ec433e626c5d068

Observation be175a7f-d594-4d4a-b249-ba68e6a3e99c · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.018641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.175488Z digest=sha256:c349f18a1f4cef3fe72a4cc9e4093f416d40dd5e2a6dc6c9bb5b5313bb32afe7

Observation b871bc30-9e35-4490-80b3-93fad8c85ad0 · outbound

This paper cites - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.000536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.180143Z digest=sha256:f45bf75aedaa9cf0a9936760b718e75bb16f73a370d1dc9cada0218417e6af04

Observation 011d302c-57c1-4b6c-9e08-c0b603624a95 · outbound

This paper cites - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg)

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.981621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.185009Z digest=sha256:080f7f2b842e1e120b04fb84878b29ed6b1f533413d22a2c69943804d45241ec

Observation 7e25a773-219d-4351-b77e-f33ec5ef81cf · outbound

This paper cites - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.956642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.190287Z digest=sha256:1061af41a32397b4264c5fbb6c33dad8378a8d03a9b8637d0dbde8cc48bcd59a

Observation eeed05e4-b1e0-4c03-8474-33f38239b288 · outbound

This paper cites source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.929667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.194780Z digest=sha256:11e6d705dd2e7c6b76aeaf3d6ba93b94d3ed444ec09f334b9d60c1efc2b72ed4

Observation b82d1c6d-022f-4f65-b6be-1fa851dfde95 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.910425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.200814Z digest=sha256:55e8c12c690ed08a6e750c64ccd18428663a1cc60c035922347b91e955b1f238

Observation 3c9eca83-c241-4135-8854-e64f65e33fec · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.890938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.207241Z digest=sha256:daac732bf481bfe0976bee41be8728338be876c54f7db99bce2fc8e6d88f86b9

Observation 8fdb2b6c-3a11-49f8-9877-ccc8cae1ee5a · outbound

This paper cites null"` or `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"` or `

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.871709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.212807Z digest=sha256:68b77610f6519dbd3e2dc61e78138d331ed142df7e522a1e45fd81f0ad54426e

Observation 92559793-e3f6-4e01-bbf9-3f5fdee6ad67 · outbound

This paper cites revised_value.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction revised_value

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.855205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.217778Z digest=sha256:28e5c8bb83629d3eca43adb377eb27ab87f2504a8e2370b831a833f9e516c87d

Observation 9fc1478b-13be-463e-83ff-8a368299de13 · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.835083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.222572Z digest=sha256:1d3247e2ab17d11c2a3f477f7e05617d933161a4e032abb904f3aed35b79e7ef

Observation 90943ae5-6f81-415a-8a34-8356cbc00d3f · outbound

This paper cites confidence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction confidence

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.813283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.234818Z digest=sha256:0148f9eeb576f54a360912aa40d41404924e0f4405d4f94bd3cf70c129bae249

Observation aa8d47f0-2708-41cd-a7a1-9c15b0df1e89 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.793979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.240837Z digest=sha256:0bdfe5cd76c94922128b8c054d0ef12c37a45ebd68d7a51943e8882062c7f379

Observation d2d90d35-360b-48b9-b65e-437e63df213b · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.770068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.246699Z digest=sha256:e99a0baca8ba7973aa51d1eec677f0e6c565a637f9a9373d5708cfac073c0d68

Observation 41fcd1e8-5048-4941-909e-7a53cd38dea4 · outbound

This paper cites Just return the final merged JSON object.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Just return the final merged JSON object

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.252902Z digest=sha256:d598ec7a348483187a4a10b32421764b7f4ab44d7344fc1e0c516a1b8fec8577

Observation 9385cae5-652e-45cb-8c22-60b0402bf079 · outbound

This paper cites LGL_group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction LGL_group

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.713347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.270446Z digest=sha256:46217f68a15b077b77697071ec9719ad5f28a4fe5af0a57343bca6f1ba2ac732

Observation bd83a4e9-a0df-4780-ada0-9ec931fa8fcd · outbound

This paper cites **Note:** EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** EXT fields may be nested

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.694694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.277351Z digest=sha256:96a79e48282d1b0c79ec36edaa1e1fca3a2890820e2de216b5e9236d4053992d

Observation 6c0500bc-867b-4458-806d-71ba0e2a6d13 · outbound

This paper cites kg/m²"` and `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction kg/m²"` and `

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.679005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.284067Z digest=sha256:9104f6ddd6cf9176bc68eebaecc4742af6e647c14508fdb6f64c57085b3b82b3

Observation 7176815e-0209-41f6-8c8e-5775018994e1 · outbound

This paper cites low glycemic load diet.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction low glycemic load diet

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.664051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.291370Z digest=sha256:6463d60e7df512bbc9ff263eb3a8b8bed2e82fe47db7ecfc7fb5a60997dd0710

Observation 17f12667-e5e3-4b97-af9c-4857f228a268 · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.649218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.296906Z digest=sha256:83a4ac0e4a1afa66e67b4d383cd9d8336538767012a4e6d2e183dde367fd7eb6

Observation 1be37d85-a7de-4cf8-8428-1f2016d914b7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.635116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.305674Z digest=sha256:3ef7f8e2e67834b805622b1e571546bda574de4c5f28a2022afdad0cc6b8a6fd

Observation 2edfcde8-35c9-4f4d-9765-3301503073e9 · outbound

This paper cites randomised controlled trial.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction randomised controlled trial

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.620028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.316558Z digest=sha256:f6d2b988927cf9aeff9860e7e1ec7c7f6ab9fbd4537234553e57d84db7bc63a4

Observation fc808152-11b7-4421-add0-9536fa7c8b8c · outbound

This paper cites You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.605521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.322330Z digest=sha256:acca58f754d10214d7faf44abcf9e83543e3257ba0038408715d5c906874a7b4

Observation cdf87145-1676-4d00-b375-3100fed7d3fa · outbound

This paper cites not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction not reported

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.589766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.328561Z digest=sha256:e327e6c50ff8ad1e0bf3c9f8606af73fd1b098512b4c11752f09e24c04feed45

Observation 685b9bca-5db9-43cd-9dad-2c4663eac3aa · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.575016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.335133Z digest=sha256:393976efc0be4d6d1c1e38124d7b68b58bb40f1ee8c43e90080f295795f7dc2e

Observation 8c836ecb-5ab3-4107-a7c8-4d359f7eae8a · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.340397Z digest=sha256:70850e7f91b4b48c262fee3a8599b355cce8535ac89a41b8e5ec5f84bdb3b3e1

Observation 606b3b47-a88b-4036-b68c-99e978ad0083 · outbound

This paper cites This is the preferred method.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction This is the preferred method

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.731507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.345245Z digest=sha256:1749f9fa9791d50ac6032d32b83140ecd341c48e29f843401faf121e4aff918c

Observation eb4f5bfe-8964-4cb9-82f6-44340197761e · outbound

This paper cites study_characteristics.PC.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction study_characteristics.PC

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.543680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.350379Z digest=sha256:628820e286c247e81798039a166371f42fb42e8cf8f7aafd9d31f7a1632de2aa

Observation 6f81d274-411b-4662-8a29-b2a4b800eb0c · outbound

This paper cites **Note:** GT and EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** GT and EXT fields may be nested

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.528298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.357558Z digest=sha256:3bfb9bad8f2405e64887331fbb6a26e0f89be0923932c2f332267aa765a277de

Observation d9c6a707-ca25-4048-ab83-9962fac63c46 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.513048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.363013Z digest=sha256:7dd41a6bd172ec08e00bfa241878010aa5e4127a0f9ceb1ca1b3889e69870c73

Observation f85a4249-4b59-4dd7-8a96-a6d65085bd5f · outbound

This paper cites Hallucinated.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Hallucinated

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.495300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.370459Z digest=sha256:cc6bfb55f02b57f98bd21e2ce1d22a8f4225d338890ec958255e3d51f28b698e

Observation 630a5e2b-c286-48be-b87e-e5bf735924fd · outbound

This paper cites Not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Not reported

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.480859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.376749Z digest=sha256:6827e0e801ee643d8c4d5bf464b8e09f005dfd0e4c9ed9ae3b1f3e6996d7a0a0

Observation 6eb66810-3256-4c2e-ac1c-081f1c2e57ba · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.993029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.723776Z digest=sha256:7d3ce4bc7292c2448ecff29af7b255f7697678772fcac8fc33d966239715d613

Observation f2e4c449-39dc-4619-bcfe-c1fbc710d46c · outbound

This paper cites doi:10.2196/33124.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.2196/33124

Reference 2021

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:07.844989Z digest=sha256:6880053fce1cc4857a641663d899b0a986c207672866527f532bcb6b90fd6d7b

Observation d98a3bf5-63c6-4afe-962e-99dc9b911760 · outbound

This paper cites doi:10.18653/v1/2024.findings-naacl.237.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.18653/v1/2024.findings-naacl.237

Reference 2024

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:46:08.006169Z digest=sha256:558c1000f12a377a72e03d2eb9c6e941e40fe591dea2ba8b183cebc88a825c67

Observation eb66b4e2-c9ff-40f7-bcf8-a057212be15a · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.090594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.090594Z digest=sha256:666bbe11aae9958ea48bdeafb46e633eb46661cf1d4cef06e222dbab00b8b4e5

Pith citing papers

Observation 10d2e2ae-f6ec-48f1-a8fa-6aa090174121 · inbound

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis cites this paper.

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:54.709479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T17:13:53.040834Z digest=sha256:a2769e007b9ed5763b3b20833cacf0bd9ac190ad03dd8b28910fc66ef9b56ef4