Pith. sign in

Paper Citation Record · LEDGER

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

As of 9 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2507.15152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15152 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:46:08.376749Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:53.040834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:13:54.647415Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact16
  • verified fuzzy38
  • unresolved37
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59abdfec-2b79-400f-abb4-6896d381a4af · outbound

This paper cites Research Synthesis and Meta-Analysis: A Step-by-Step Approach.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Research Synthesis and Meta-Analysis: A Step-by-Step Approach

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.653529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.653529Z digest=sha256:41a6bc0df2a41c25649c88a9810284000bff974b001da0e453be5d17f571cafd

Observation 1f67cc0f-6e6e-4c9c-8dd2-dd9074bee257 · outbound

This paper cites Analysing data and undertaking meta-analyses, chapter 10, pages 241–284.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysing data and undertaking meta-analyses, chapter 10, pages 241–284

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.664179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.664179Z digest=sha256:24b42a6b6629a25a550007f15f1274125329d014ed06fdbd86fd4b023c78be1b

Observation d0762771-aa93-47e0-aeb8-895ffc8ebf2c · outbound

This paper cites Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.671805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.671805Z digest=sha256:5c63e617a486d564655d821b738f00267599594ffdad900267b2580d893988e4

Observation efdddffb-8142-469b-909e-5f43339c72f5 · outbound

This paper cites Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.038582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.680841Z digest=sha256:db4a0bbcbbeea81881ae69c15e1b9218355eb2d6574fa74e74e646fa5f28281d

Observation 70ea14fd-5453-4d30-8447-05c6b28e9f97 · outbound

This paper cites Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.020674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.685632Z digest=sha256:0c315b808bf00b82f8b9e0f8b08b808adfea3167f1cb8cd0fa3296e567637277

Observation 975c4024-e2d0-4308-ada1-f84d80119446 · outbound

This paper cites Toward systematic review automation: a practical guide to using machine learning tools in research synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward systematic review automation: a practical guide to using machine learning tools in research synthesis

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:07.690733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.690733Z digest=sha256:f6abb4776d520762156ca63739b0f7709d96019cc1d1889406419f8eae14a209

Observation b0629e68-bf64-4206-99b4-91252af8cde9 · outbound

This paper cites Exact: automatic extraction of clinical trial characteristics from journal publications.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exact: automatic extraction of clinical trial characteristics from journal publications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.697875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.697875Z digest=sha256:a80fabeb0eb1d089fd449efdeb159b224b6a571889455e34a5a13a15b6b09098

Observation fdd4bafa-4184-4928-a2ff-b8fe42880703 · outbound

This paper cites Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz

Reference 8

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.968448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.746197Z digest=sha256:7251250270974e16a59f6ffeef765e3a70e448c866098183cbef9dfcc0f3f1a4

Observation ca059850-7817-4bff-a971-e79644f87bd6 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.459526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.763071Z digest=sha256:9600803e10d43c95df2c80216320f8d9dd56ef6bb0fc1fa10627073941c3970a

Observation 1c86dbba-aac8-4c3c-ac6b-60cf0a5c8cc7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.784022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.784022Z digest=sha256:1b8129ece0b1453698ddc2273396d47c137fb333e7403ddad5272c2d6634ea75

Observation bc08e3c2-ce9c-41ab-ae05-793e5c754ed0 · outbound

This paper cites Automating meta-analyses of randomized clinical trials: a first look.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automating meta-analyses of randomized clinical trials: a first look

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.801622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.801622Z digest=sha256:bb79346af5806facb7334ed562c2c404b68f391bd77f164ae224b84d02881c7b

Observation c032b9f2-8819-4519-aed0-7709bda04041 · outbound

This paper cites Katz-Rogozhnikov, Kush R.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Katz-Rogozhnikov, Kush R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.812345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.812345Z digest=sha256:ce39452c012052598acdd282de6fc44b195fbdc6208d086f12f824242d39fe01

Observation ec473968-5324-44d2-a235-b93418bdb282 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.817822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.817822Z digest=sha256:5681fcba4f589ca2a98ea3a1190f1227db09d0074bb15c2a7545fa0754672381

Observation 25b1e172-1af5-4a79-be85-76ea657b1997 · outbound

This paper cites Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved].

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved]

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.928184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.824462Z digest=sha256:7db5a9184a3dd5eebd9348abc25acb9162beb40511e32680f9e47126823b6230

Observation 094ae696-9487-4bcc-8bfd-ea6e8255c58e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.829830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.829830Z digest=sha256:2b3e4b3032b11e555f7bb3503e28feb6e23bd1129e6a6e9bd21738a5deb4b889

Observation d8fbea24-066b-455d-b1ca-bda23239e5b1 · outbound

This paper cites The data is in: Deciding when to automate screening in your slr, November 2023.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The data is in: Deciding when to automate screening in your slr, November 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.834832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.834832Z digest=sha256:b0e72f3aa85f9512b3de7b3f98ff24dea4d9fb2da6aba633c888cccbb3d5fb45

Observation f63edf05-4f58-41af-89ae-082feb262b6e · outbound

This paper cites Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.839801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.839801Z digest=sha256:27814425c805980ee1b2e70111dc8356bda71d4f32fee668b80f1461722f7f01

Observation 28e34ed0-8485-4fcb-8124-2dad653175dd · outbound

This paper cites MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.854237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.854237Z digest=sha256:d64401f630cf70f9b35d494ced4334fa6ffeda740415ccee866b0c8b938b1d53

Observation e7c78217-16e8-4a6f-8fd3-64c498887a46 · outbound

This paper cites Chatgpt: Large language model (mar 14 version).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Chatgpt: Large language model (mar 14 version)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.864218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.864218Z digest=sha256:d02a11912c9915b274978539566f49f6d3748e57dd403461ba72860815862ce3

Observation cac6f8e3-00ef-4460-b6f3-4ea7f1d4f000 · outbound

This paper cites Claude 2 model announcement.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Claude 2 model announcement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.870017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.870017Z digest=sha256:e434148fe6f0317767e43d69226c9fa87cccda2897cae96dc9b6c85ad2ee12c6

Observation 583f778e-b019-49a5-aa43-88fadb7779a7 · outbound

This paper cites Zero-shot infor- mation extraction for clinical meta-analysis using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zero-shot infor- mation extraction for clinical meta-analysis using large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.875116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.875116Z digest=sha256:9831d4558a036c9ab9303d2c62491dfb95849b331eac75418a21237afb892b7d

Observation 75556778-c65e-4b3b-8f2d-8b86090c2e79 · outbound

This paper cites Performance of two large language models for data extraction in evidence synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Performance of two large language models for data extraction in evidence synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.880456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.880456Z digest=sha256:0eb4364c7b88d483bb59aa4f0319e1bf33ebbde27c7066320b2f20a9d86a93fc

Observation 6705cf30-2390-4d27-9f7e-b02c0a6a2512 · outbound

This paper cites Automatically extracting numerical results from randomized controlled trials with large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automatically extracting numerical results from randomized controlled trials with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.885286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.885286Z digest=sha256:3d226fb901b3f37189ccb63c592eef0b30d9231d0bbb1fe6e9fc667aac5b630d

Observation fe259602-0acf-4aa5-9e96-25afc38508d9 · outbound

This paper cites Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.891121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.891121Z digest=sha256:c75f85c8144aa8638207df835c07fbbb64d322e4f178bba321807e7a186e5440

Observation 94bbea6d-b5f3-4320-89d7-15055bdf1e24 · outbound

This paper cites Lee, Shigeki Yamada, and Tomohiro Mizuno.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Lee, Shigeki Yamada, and Tomohiro Mizuno

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.896068Z digest=sha256:393379f31d7426c52fee2f1e068fdad8aec9c2d66f4a4844c2490bddc78b73fb

Observation 1cfbaa80-85fd-4a5a-97d0-24a0a022a479 · outbound

This paper cites Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis

Reference 26

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.370112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.901443Z digest=sha256:9bd43cde77d2ef13853bcbe7ed755ef9a2c0061f8909ba6170ad2e7c62fa3237

Observation 96dc95e7-b78e-47d4-bdca-e36962197bcf · outbound

This paper cites Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris

Reference 27

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.829269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.906139Z digest=sha256:ea010f3a9d6915898f2401d0b92531e50435bd7aaab1db9d78c263cbbb890d23

Observation 64105b66-d969-41af-a918-8da6e2c4b62b · outbound

This paper cites Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.285708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.913293Z digest=sha256:8bbe6afb7d0333e57cf94858a05293a23eab9f7bceecfd181f4f0b9d5389b44d

Observation 8b698dba-6c1e-4100-a69b-c5908f7b0a20 · outbound

This paper cites Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim

Reference 29

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.799935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.918186Z digest=sha256:f65bd38b0bb38b9c4eaad420347fda337b0e3ae42c5e8ae1c7dd988fb1dbca33

Observation 5c6db0f3-9237-4145-a5d2-9b5c829d2f1e · outbound

This paper cites V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.774773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.924380Z digest=sha256:b161bd33c3a9c9fbef75aa8334758c7fa53df4f39207ed6e905aa4877ca875d2

Observation 34192b22-a459-443b-8287-0cf5870d8db2 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gpt-4o mini: Advancing cost-efficient intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.934074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.934074Z digest=sha256:aa6d70c6c8cccbec1205aad89bef1f80b163cde2442e660c197ab59f4f51e0b1

Observation c32e8170-e0d4-4921-9966-1a46b81557ab · outbound

This paper cites Gemini 2.0 flash.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gemini 2.0 flash

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.939758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.939758Z digest=sha256:423d288313fbdd607406289f195eb416a4886260a38dfb937425058c57af8f12

Observation 93682bef-d44e-4c7c-8a57-315a4becb015 · outbound

This paper cites Grok-3 language model.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Grok-3 language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.945503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.945503Z digest=sha256:cca5493567838af6a4d4c94a5fb594c69f7372a3e77467b71b02280e809a4c10

Observation 8b9cbaff-4f41-4132-857c-c60de2b330d0 · outbound

This paper cites The impact of temperature on extracting information from clinical trial publications using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The impact of temperature on extracting information from clinical trial publications using large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.951933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.951933Z digest=sha256:a91a8a601de5fbdf11e97453f47899f269ebd98e50a8ca42f62ca961c10dd083

Observation 4bbf6358-214b-43d1-afdb-4f0aea3b2d51 · outbound

This paper cites AI-Assisted Data Extraction for Systematic Reviews in Education.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction AI-Assisted Data Extraction for Systematic Reviews in Education

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.958889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.958889Z digest=sha256:6014ba26d53cf3b9696d960eb2254abecdf5497b5783dca17d4b67c961b5444f

Observation 178abf96-a100-40cc-b4ea-fefa51c92b6a · outbound

This paper cites Use gemini 2.0 to speed up data processing.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Use gemini 2.0 to speed up data processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.965054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.965054Z digest=sha256:8362c26792ccd1cbeda1b6e92feafa24372a9ccbc136167abc8e8084d1a9eb40

Observation f654e4b5-931a-4750-a1f4-2c9e8ea67a82 · outbound

This paper cites Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.748746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.971513Z digest=sha256:19db38ab911c66aa14440d8b93539d3782f578b8cc34313da7885b811427d8a3

Observation 759dce94-632b-4544-a756-d882cdaf18b0 · outbound

This paper cites Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.347631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.976676Z digest=sha256:1573f2d7e00e52559940b5408a49f7f4fcd9160a9712ff8c9c6095fb1c757e47

Observation fc9056f3-42b0-4a30-89cd-34b59992be83 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Reflexion: language agents with verbal reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.318292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.989820Z digest=sha256:1bffe59af63826054a54e404c3027f6bdb1e588b13cf2efcdd814f5f2219cf0d

Observation 2acec01b-d982-4110-aa42-d3a6d238ea8e · outbound

This paper cites Towards mitigating LLM halluci- nation via self reflection.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Towards mitigating LLM halluci- nation via self reflection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.994474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.994474Z digest=sha256:82d939e101babde9e44738041994c7cac09fcac0306bd2d5996e21b7b3167d47

Observation 6c7a93d2-95b1-41d6-8356-4ae8b04c0c95 · outbound

This paper cites When hindsight is not 20/20: Testing limits on reflective thinking in large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction When hindsight is not 20/20: Testing limits on reflective thinking in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.295126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.000826Z digest=sha256:17cda2d6c9068b2f0b4bd38d373d9c7e6b914986341e5152ca3ec5f65dc326b9

Observation ccc81064-a620-4159-a80e-ba3512ca49d4 · outbound

This paper cites Dietterich.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Dietterich

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.267940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.011272Z digest=sha256:2ade5fbc1b383ef3ba8eb445d3673c862d69120440ba59b80307b03037672feb

Observation 23f4b757-ac28-47ef-8306-fc3ec919bcb0 · outbound

This paper cites A survey on ensemble learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction A survey on ensemble learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.018396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.018396Z digest=sha256:27a986ec4c48416fb06f0ee26e2e081ffb6a07334b1dc36940ef13eb9f6339a5

Observation e84df5e4-15d8-4caa-aca3-aba9d2a57a96 · outbound

This paper cites Ensemble pretrained language models to extract biomedical knowledge from litera- ture.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Ensemble pretrained language models to extract biomedical knowledge from litera- ture

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.645522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.023179Z digest=sha256:dea946a53bda04a59fa096bff757f2e1000a8b1814d51b87e7184a630bf7a133

Observation 8d7ef2bc-7931-463f-922d-5202489506dd · outbound

This paper cites Zhang and A.L.P.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zhang and A.L.P

Reference 45

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.618170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.030348Z digest=sha256:7facbf675865881ee562debcc88295bac9b4b9d33c7c2c78c7f335bf2d799e2a

Observation 31adf290-91b9-4d02-862c-b19873f71cee · outbound

This paper cites Comprehensive testing of large language models for extraction of structured data in pathology.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Comprehensive testing of large language models for extraction of structured data in pathology

Reference 46

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.587520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.036361Z digest=sha256:ad96cfd5bd4c9eac395c9900f6a1af174a8d63df30ecd78ee5e8d51c1d4fbfd9

Observation d28dc553-4477-4edf-b7d3-e393a2f59d3f · outbound

This paper cites Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.546789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.041977Z digest=sha256:100da44338ebfcea8a55c1c5fec8abb9ab995173ea73e01b03068dba0b8223db

Observation 18e1f353-74a7-4807-89c8-99ee1df24e4f · outbound

This paper cites Match, compare, or select? an investigation of large language models for entity matching.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Match, compare, or select? an investigation of large language models for entity matching

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.243358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.048626Z digest=sha256:321726161a80e146584a7c4bd8b7899d7461910676a99c1e346b840543d08e88

Observation cc3c1a78-35ff-4e69-a324-9c56813ca94d · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 49

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.516502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.053892Z digest=sha256:9f3818b8849f7f6d534630e6328463a0b6166084fac8cfcd802a44c4f0addc4c

Observation 84af77db-91b6-4c1a-b14e-2095284f718e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.061584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.061584Z digest=sha256:4f838b3a814a0b696cad191fece0cbad7c4b3d58bc0a5c9ad5723a4db5835d8d

Observation 3f397eca-7685-4cf9-8b04-69ba6bc821d3 · outbound

This paper cites Guyatt, Andrew D.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Guyatt, Andrew D

Reference 51

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.492238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.068084Z digest=sha256:08da9e2110268aab3ba2254434bd5b78155d891f47806e5b147b6b6af7e5f4a6

Observation 8614a940-0e89-4af5-add4-f99fb77eb40c · outbound

This paper cites Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.075310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.075310Z digest=sha256:0f5bbbd95825f83eeaf12d95cd20bf1377485252c3e6d74d0cd51970b0068e1a

Observation 449df189-eab1-4838-8f8f-8e5dc1abdc35 · outbound

This paper cites Agentic reasoning: Reasoning llms with tools for the deep research,.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic reasoning: Reasoning llms with tools for the deep research,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.198891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.083557Z digest=sha256:ad0b531b63030001643ad0416304da969ebb43093da0956c9d2c7e0e5cc22fca

Observation 2aa669b4-d09e-46ea-aa25-0ac7f3da646a · outbound

This paper cites TART: An open-source tool- augmented framework for explainable table-based reasoning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction TART: An open-source tool- augmented framework for explainable table-based reasoning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.171211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.099934Z digest=sha256:e3a70a1c6caf44b04a17d49bfb759507f53ef1bb8758d2aa2f13a2068c2c65cc

Observation b81958ca-a660-4689-bb8e-db65fdbe01a3 · outbound

This paper cites Medical hallucination in foundation models and their impact on healthcare.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Medical hallucination in foundation models and their impact on healthcare

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.105709Z digest=sha256:58b8859fdedea96731881083da6ea91927d810b3fc136f236de1866d267c6ddd

Observation cabcab11-7bd8-46b3-ae71-b44625490a07 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.111858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.111858Z digest=sha256:c231a4cfc03ddf9afa7c0cca850618bd5e1076aa33f21c34fa13b10dee472528

Observation 2b1898d8-f3f5-489c-9aee-2d67e449992e · outbound

This paper cites Potential roles of large language models in the production of systematic reviews and meta-analyses.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Potential roles of large language models in the production of systematic reviews and meta-analyses

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:08.118777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.118777Z digest=sha256:a285559707cb0255e566ba113bb6b69047c085527f3b70ec6a5cef904a62bce0

Observation 93db9ab6-c6b9-4861-98f0-f20db0315ff6 · outbound

This paper cites justification.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction justification

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.153899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.124679Z digest=sha256:94e1ed15d81de1df654ad59c0e6933bf35b14218102edee9bcdf5a0628b4906a

Observation e1d54069-2ad1-44a4-9a0c-74c7a7d12d35 · outbound

This paper cites other_time_points.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction other_time_points

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.136844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.130975Z digest=sha256:b5ec45bd4b5058d38447d118338d43eb07a41f42fd2e35efaecfb304351d7d9d

Observation 69682989-e509-4562-9ede-0699148ce6d3 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.117550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.136861Z digest=sha256:9fd09f3e8c3fa3659231377fe0280d564d81f519c8c91955beb7acb634f894fb

Observation 5af7e017-6330-4c4b-9a15-b03cafc5d496 · outbound

This paper cites null"`, NOT `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"`, NOT `

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.100782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.142821Z digest=sha256:ece90dcb9bcf210fde565b7a08d485b02ff90dc0e15c9312e56aa336a3bc95cd

Observation 502eff7e-a05c-4706-a710-8504908387b9 · outbound

This paper cites more common in the intervention group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction more common in the intervention group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.078892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.149296Z digest=sha256:52b85945a193cf96d2c6e4c7bc44724b5c7f95157e8a8f0fba24cacb9f3cb132

Observation 03ae9405-e1de-40c5-a58f-b5f4d3b0af26 · outbound

This paper cites data_conflicts.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction data_conflicts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.058081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.156022Z digest=sha256:274f4bebea6eef3255ff69bcc78cb351529bb1b05960816b8f8f78f80e59efef

Observation cbdb9462-4a71-409a-9055-a8c6b0f3411e · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:46:10.036874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.167425Z digest=sha256:43978fc2dac51edc09171e48e122450541c7fa94207fa7e601721f28e100cbc3

Observation be175a7f-d594-4d4a-b249-ba68e6a3e99c · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.018641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.175488Z digest=sha256:7b6b4d8ea592862bc5c85fb5e284fd698e701dd30ac58967260758ccf72507ee

Observation b871bc30-9e35-4490-80b3-93fad8c85ad0 · outbound

This paper cites - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.000536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.180143Z digest=sha256:3f249bec3f426d212c306f1389bbcaf48acffae67f61013a326e94baf2ac5b4f

Observation 011d302c-57c1-4b6c-9e08-c0b603624a95 · outbound

This paper cites - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg)

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.981621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.185009Z digest=sha256:44b40c9630452143fcc3cac3734c5ea358631af946a653a873484da5fa764495

Observation 7e25a773-219d-4351-b77e-f33ec5ef81cf · outbound

This paper cites - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.956642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.190287Z digest=sha256:bb8c85e309278bf5f1a6c0c5c7c6733ffa9a71733bfcd7407e243d6d4e2c30c2

Observation eeed05e4-b1e0-4c03-8474-33f38239b288 · outbound

This paper cites source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.929667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.194780Z digest=sha256:fec7d52caa38230ee0e8d2ac2f7c6967c4f0bdc4032714cecc4b82d80e7157dd

Observation b82d1c6d-022f-4f65-b6be-1fa851dfde95 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.910425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.200814Z digest=sha256:ef5eafea093729799ff660fd43601084af3a2fb02cf4b14e4ee27dc023feb493

Observation 3c9eca83-c241-4135-8854-e64f65e33fec · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.890938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.207241Z digest=sha256:419e98c2ae6de5137a8e8219a3a24faad36d653057724b66f4ac6b8e96c9eeef

Observation 8fdb2b6c-3a11-49f8-9877-ccc8cae1ee5a · outbound

This paper cites null"` or `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"` or `

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.871709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.212807Z digest=sha256:25a527a64aac6b91124b8266368eb206499c5575c46c3c740776b8b6e3091b35

Observation 92559793-e3f6-4e01-bbf9-3f5fdee6ad67 · outbound

This paper cites revised_value.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction revised_value

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.855205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.217778Z digest=sha256:fa12b8ca121f18dfdd1cbdb280066b3304f964d6ffd1ca11e780d95390f24499

Observation 9fc1478b-13be-463e-83ff-8a368299de13 · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.835083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.222572Z digest=sha256:38943bdca31c07405b989ce1fd139ab87af44f5f9e9ebcf68101914b1dd5bbf5

Observation 90943ae5-6f81-415a-8a34-8356cbc00d3f · outbound

This paper cites confidence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction confidence

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.813283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.234818Z digest=sha256:c75599e43a7a8ae773fc1cd86cf193d19ee8510176742f2b8dcd3466f4f67da4

Observation aa8d47f0-2708-41cd-a7a1-9c15b0df1e89 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.793979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.240837Z digest=sha256:04adba9cb6d1fc8432476627222c3ff1c477f9d6c0c3bf3c9985a372d260a11f

Observation d2d90d35-360b-48b9-b65e-437e63df213b · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.770068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.246699Z digest=sha256:9bce9a31a7ebb059352c6830667eef7ee7f1477eea58fa94b5e80c9daa5439e9

Observation 41fcd1e8-5048-4941-909e-7a53cd38dea4 · outbound

This paper cites Just return the final merged JSON object.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Just return the final merged JSON object

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.252902Z digest=sha256:9e25dd4cc28217e7dc98c2bb4bb09935e70560b370007616952e09f9ab2663b1

Observation 9385cae5-652e-45cb-8c22-60b0402bf079 · outbound

This paper cites LGL_group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction LGL_group

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.713347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.270446Z digest=sha256:4cd2b9629411f4770d8be80749eb6ca31ab8fe8be5d17a45f43d129bb84884d7

Observation bd83a4e9-a0df-4780-ada0-9ec931fa8fcd · outbound

This paper cites **Note:** EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** EXT fields may be nested

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.694694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.277351Z digest=sha256:70b4599b85eef93da62ec4b44da575cb1b8a0d5596c4d91022555c9138f43aae

Observation 6c0500bc-867b-4458-806d-71ba0e2a6d13 · outbound

This paper cites kg/m²"` and `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction kg/m²"` and `

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.679005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.284067Z digest=sha256:f87d8fd64384a5e4e65212efc1e9e254c0d5bc0e4fd5294f7833ed0f7efd7c4e

Observation 7176815e-0209-41f6-8c8e-5775018994e1 · outbound

This paper cites low glycemic load diet.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction low glycemic load diet

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.664051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.291370Z digest=sha256:ceba566d6eb7fa7aedd69fca9aae88ec1dbae4915024502ac27d8e89802cd7bc

Observation 17f12667-e5e3-4b97-af9c-4857f228a268 · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.649218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.296906Z digest=sha256:322b5742f4b7fcdbd584f8a070263b648ab6ec941d18276b79b0102208524c18

Observation 1be37d85-a7de-4cf8-8428-1f2016d914b7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.635116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.305674Z digest=sha256:d3a7088d8b1ee71b1ad2cacc3c2c00691888715712d8cbab0d2b4c172cabbd8c

Observation 2edfcde8-35c9-4f4d-9765-3301503073e9 · outbound

This paper cites randomised controlled trial.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction randomised controlled trial

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.620028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.316558Z digest=sha256:11911da7a37414cbf52606cb49dc0a6dec7662e9fcd5a4ffc5370651abf288c1

Observation fc808152-11b7-4421-add0-9536fa7c8b8c · outbound

This paper cites You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.605521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.322330Z digest=sha256:c00c8507547f0b7e2dff090521749183a9953e7ed7247ece55bacf930b87d4a6

Observation cdf87145-1676-4d00-b375-3100fed7d3fa · outbound

This paper cites not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction not reported

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.589766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.328561Z digest=sha256:f816d8471ac5a5253fe5479d489dcaaeeb8b2bcb1ce8ebd59de0301088d4f1e5

Observation 685b9bca-5db9-43cd-9dad-2c4663eac3aa · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.575016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.335133Z digest=sha256:714005f4ceba7fc25ff65ddd80eff1e81fd6a059bfe6ce92039a5d3d004a3263

Observation 8c836ecb-5ab3-4107-a7c8-4d359f7eae8a · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.340397Z digest=sha256:c04772276d73c67cd78362d44c032fae7b2d4b148ef1d65d9072824e1796877d

Observation 606b3b47-a88b-4036-b68c-99e978ad0083 · outbound

This paper cites This is the preferred method.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction This is the preferred method

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.731507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.345245Z digest=sha256:6b2eee38878ce56a10ba46fdf7cb603d163e5e4d872980af52db14f02c0d0b04

Observation eb4f5bfe-8964-4cb9-82f6-44340197761e · outbound

This paper cites study_characteristics.PC.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction study_characteristics.PC

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.543680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.350379Z digest=sha256:871e2eb40f9c84cb36b01a08d4929bb542bc7615ecbc36277ac02067769f49a1

Observation 6f81d274-411b-4662-8a29-b2a4b800eb0c · outbound

This paper cites **Note:** GT and EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** GT and EXT fields may be nested

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.528298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.357558Z digest=sha256:32d975965f0ef11ae269933789e8d7c0fba91af13a254f76a4d721e26f1892e9

Observation d9c6a707-ca25-4048-ab83-9962fac63c46 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.513048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.363013Z digest=sha256:69b4bcb91493aca97a6fc53a9be04b213befb38e6c9e30a6b52e3185dcc33bd1

Observation f85a4249-4b59-4dd7-8a96-a6d65085bd5f · outbound

This paper cites Hallucinated.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Hallucinated

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.495300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.370459Z digest=sha256:a644ee3729bf9a25ab9490b8c86f16cd8f18f6055bf5488b4835c8dfe64e3f1e

Observation 630a5e2b-c286-48be-b87e-e5bf735924fd · outbound

This paper cites Not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Not reported

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.480859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.376749Z digest=sha256:8f2d83696d6052c856d4cc2c773dc76c88f2d818b5d724e1380eb93f8a9b11cd

Observation 6eb66810-3256-4c2e-ac1c-081f1c2e57ba · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.993029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.723776Z digest=sha256:7b9f162a23eba20e1638c7d54280b1a9027837fe3bc218512a8eb77b481baa96

Observation f2e4c449-39dc-4619-bcfe-c1fbc710d46c · outbound

This paper cites doi:10.2196/33124.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.2196/33124

Reference 2021

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:07.844989Z digest=sha256:c0d49f48b30453a7f9946d88a9bef1c63b3443577fba6b88e65c79e5b6ca411f

Observation d98a3bf5-63c6-4afe-962e-99dc9b911760 · outbound

This paper cites doi:10.18653/v1/2024.findings-naacl.237.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.18653/v1/2024.findings-naacl.237

Reference 2024

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:46:08.006169Z digest=sha256:377bf8337863c1956c1750403c1f93a19e5969080076f02c5c1001166eba34be

Observation eb66b4e2-c9ff-40f7-bcf8-a057212be15a · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.090594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.090594Z digest=sha256:d0c617588f076681ae182aa6c23c10e189c51c38ff50460d22f7fda96253783b

Pith citing papers

Observation 10d2e2ae-f6ec-48f1-a8fa-6aa090174121 · inbound

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis cites this paper.

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:54.709479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:53.040834Z digest=sha256:44cbeae5f32bf3d402aa5cb2c43459273e42e89ac05a56f8ed03de51a2edb6e3