Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:42.348441Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.15715.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:42.348441Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 27c96f69-cd38-4040-a144-6b5fa93f0f7c · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Kurtz, Edwin A
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e857067f-c455-44c1-8017-c896aefbafe8 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Henneken, Carolyn S
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16e9845e-ce20-4707-ab47-19d22826bbe0 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3b0184-0ec3-48f0-bbe2-5efbd72a176d · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Search B ots: User engagement with ChatBots during collaborative search
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b256564-1ae3-4dc9-ac4e-53ef4e34de2a · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Embedding search into a conversational platform to support collaborative search
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857cfe90-fb18-485b-8def-5517ff44a169 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs The effects of system initiative during conversational collaborative search
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f7e1caf3-44f0-4c07-86a0-d19a3b528224 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874cf6a4-dc96-4f38-9d90-68df6f0ff5e0 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Bowman and George Dahl
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e8f034-f8d8-4ee5-8e89-cc9f404679d3 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Toward algorithmic accountability in public services: A qualitative study of affected community perspectives on algorithmic decision-making in child welfare services
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b9716d-7221-49e5-aa41-abd3c7533141 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 043b933d-e42e-4251-ad1e-b29e00ad287c · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Fabbri, Wojciech Kry \'s ci \'n ski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6570ef50-b6b5-4676-b30c-279824ba7822 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 266e21d3-c56b-44e6-9fb6-0e17a90ecfb5 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Enabling large language models to generate text with citations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456fe5da-355b-4c50-8d98-531913e588bb · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs What can large language models do in chemistry? A comprehensive benchmark on eight tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29d230c2-2941-40ff-8f67-5c12f3144f92 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Measuring mathematical problem solving with the math dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf012bce-e408-49c0-86dd-64eb54026ca1 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Iyer, Mikaeel Yunus, Charles O’Neill, Christine Ye, Alina Hyk, Kiera McCormick, Ioana Ciucă, John F
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f31780-8ae0-4420-8746-a8aeabd65f56 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb8b2d2-134b-47ff-aa79-48e18b31b379 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs On scientific understanding with artificial intelligence
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f701a313-32ec-485f-a132-4e78c62403b5 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Wikibench: Community-driven data curation for AI evaluation on W ikipedia
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690c2d37-96c5-4c25-842a-3e23044ae119 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Chapter 8 - interviews and focus groups
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0218d1d6-2b04-4562-bc11-17f6776dba69 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs H alu E val: A large-scale hallucination evaluation benchmark for large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ef06a9-e284-49af-b016-443301616f0d · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64e7b96d-dc3b-424a-bdb2-ee5f5deb5e4b · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92f3370-7547-4c92-97ad-5a0957c5af95 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Vera Liao, Daniel Gruen, and Sarah Miller
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b993fe-9673-47a8-aa2b-6b9d3d39ce79 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs S elf C heck GPT : Zero-resource black-box hallucination detection for generative large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96673d3-1578-4deb-aa67-acf7f07d0710 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Mathewson, Jaylen Pittman, and Richard Evans
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dff6a3-e837-4a45-ab51-6548962615c2 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Wolf, Josh Andres, Michael Desmond, Narendra Nath Joshi, Zahra Ashktorab, Aabhas Sharma, Kristina Brimijoin, Qian Pan, Evelyn Duesterwald, and Casey Dugan
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc0bdb0-4283-4bf8-a8fe-4d6e40d4b2e3 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a529e512-12a9-4686-aae0-4702e446ffb3 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Rodriguez Mendez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill Naiman, Jesse Cranney, Kevin Schawinski, and Roberta Raileanu
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed222ae-9815-4590-82e8-c0b06816356b · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs emr QA : A large corpus for question answering on electronic medical records
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340bbaff-2ae5-4607-9b65-9f4e20bfe3f5 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs B io R ead: A new dataset for biomedical reading comprehension
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e74bfa0e-789f-4503-9ab5-483edc04a799 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs B io MRC : A dataset for biomedical machine reading comprehension
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ff9d22-4084-4f89-bdd1-f30853b47950 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Exploring temperature effects on large language models across various clinical tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3314d704-20c7-4a91-8284-519ee72dda91 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Smith, Huiling Liu, Kevin Schawinski, Kartheik Iyer, Ioana Ciucă, and UniverseTBD
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fca715d-5d5d-4316-a7e7-c0c8c5af6b18 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Bender, Alex Hanna, and Amandalynne Paullada
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db18018f-e43b-4cf7-8f0a-1e5d630aa9bd · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs The effect of sampling temperature on problem solving in large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a2d16e-6a6f-4b4e-b5e5-73306be032ec · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Improving evidence retrieval for automated explainable fact-checking
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 187fe219-bbef-4d96-8112-24a864213317 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71985320-c1e8-4ecf-a5dc-8468c119c5d6 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aa554d2-c2d9-4573-ac49-0400e7b08b79 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Scieval: a multi-level large language model evaluation benchmark for scientific research
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd25c18-98f3-4b28-91af-f17b7e0950cd · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Nestor, Ali Soroush, Pierre A
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d21ebef-9c75-4bf2-8a6d-fe3d9e583d95 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs FEVER : a large-scale dataset for fact extraction and VER ification
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21746c8-56d1-4ecc-b42f-5a61afa09ddd · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Group chat ecology in enterprise instant messaging: How employees collaborate through multi-user chat channels on slack
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 643ed9c8-a501-4895-bf7a-81bef5ad4cce · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Evaluating large language models on academic literature understanding and review: An empirical study among early-stage scholars
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194c1acc-3bff-422b-b52c-c2e22d164b3c · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs as an ai language model, i cannot
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cda476-4e8d-475d-a1c2-9704ddabbf1e · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Designing an Evaluation Framework for Large Language Models in Astronomy Research
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 347b5cc6-f90b-465a-87da-5b86b931a88a · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs C-pack: Packed resources for general chinese embeddings
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844b0afc-4d00-46e9-a8b5-eb4d4912d025 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs AI as an active writer: Interaction strategies with generated text in human- AI collaborative fiction writing
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 01bbfc41-c32a-45ea-bb2a-0b9b7c493064 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs User-Controlled Knowledge Fusion in Large Language Models: Balancing Creativity and Hallucination
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bde327-f37d-4162-9d84-bf16ed863ac6 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Weinberger, and Yoav Artzi
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40a5499e-16cf-4445-8e8f-cd3dde3432f5 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs WildChat : 1 M C hat GPT interaction logs in the wild
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 67552ce0-5813-426e-a34d-c33fa787ac2b · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98031c2-bf97-47b0-a7f9-a425c93788b5 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Navigating the grey area: How expressions of uncertainty and overconfidence affect language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe456e7f-3a03-4119-ac9b-93293c239cc1 · outbound
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Relying on the unreliable: The impact of language models ' reluctance to express uncertainty
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.