Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:58:03.190450Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2508.07827.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:58:03.190450Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T21:36:22.389047Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T21:43:59.674035Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b936667-020d-4e5e-a448-86eb1033055c · outbound
Evaluating Large Language Models as Expert Annotators write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c63d06-2427-4046-be84-810b6c61a8cf · outbound
Evaluating Large Language Models as Expert Annotators GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d58986c-7e14-4e38-8c82-ee71a1672364 · outbound
Evaluating Large Language Models as Expert Annotators Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c71ad9-7577-4825-b493-a4d9c14aa089 · outbound
Evaluating Large Language Models as Expert Annotators The claude 3 model family: Opus, sonnet, haiku
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529cf28d-af15-4939-80ac-bf7d8d69c3d7 · outbound
Evaluating Large Language Models as Expert Annotators Claude 3.7 sonnet and claude code
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb27085f-bb3b-4086-a198-57b7f2157896 · outbound
Evaluating Large Language Models as Expert Annotators Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f06ed9c3-bca1-4c83-b6eb-db6b4708b7c8 · outbound
Evaluating Large Language Models as Expert Annotators Must read: A systematic survey of computational persuasion
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae1414c-9156-4b55-b678-7dd155b038f4 · outbound
Evaluating Large Language Models as Expert Annotators Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a55427f-961a-4075-a876-0c55903314d1 · outbound
Evaluating Large Language Models as Expert Annotators ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae76b19-00d9-4629-9bc3-94eaa1c256df · outbound
Evaluating Large Language Models as Expert Annotators Evaluating Large Language Models Trained on Code
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517ec177-3791-4aa5-8cf0-ef478a09746b · outbound
Evaluating Large Language Models as Expert Annotators Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b4fce70c-4c5d-4064-8b83-7f70e29dbcd3 · outbound
Evaluating Large Language Models as Expert Annotators Chatgpt goes to law school
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6872d044-9693-4b34-a660-6621fcd6be4b · outbound
Evaluating Large Language Models as Expert Annotators GPTs Are Multilingual Annotators for Sequence Generation Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02c00c8-f313-47e0-b984-c949688251ec · outbound
Evaluating Large Language Models as Expert Annotators Is gpt-3 a good data annotator? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 11173--11195, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7c5c73a-8d72-48e5-9517-8f6c6df3b268 · outbound
Evaluating Large Language Models as Expert Annotators Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17132b9e-4e9c-4631-bf10-a58a9008b6d5 · outbound
Evaluating Large Language Models as Expert Annotators Measuring the persuasiveness of language models, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7121dc-276f-4728-ab52-9717e23e1123 · outbound
Evaluating Large Language Models as Expert Annotators Measuring nominal scale agreement among many raters
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e9c6e715-aa56-4a68-95f0-99378478ba55 · outbound
Evaluating Large Language Models as Expert Annotators Chatgpt outperforms crowd workers for text-annotation tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56a7c1a1-538c-4428-95bd-b23114d5e49b · outbound
Evaluating Large Language Models as Expert Annotators Introducing gemini 2.0: our new ai model for the agentic era
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 986d6739-07a7-47d7-b125-516761e6fd95 · outbound
Evaluating Large Language Models as Expert Annotators Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b75d865b-3672-4cb4-b829-b76d31677352 · outbound
Evaluating Large Language Models as Expert Annotators Large Language Model based Multi-Agents: A Survey of Progress and Challenges
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b214e203-5069-48d5-a24d-1bd72469d499 · outbound
Evaluating Large Language Models as Expert Annotators AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5df38a-bde8-4855-9660-b973c203d59e · outbound
Evaluating Large Language Models as Expert Annotators Measuring massive multitask language understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4a376f-c68b-4e41-8805-cc10a7c880d8 · outbound
Evaluating Large Language Models as Expert Annotators CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f533cea-5aac-4b70-9d42-58f50de4ed87 · outbound
Evaluating Large Language Models as Expert Annotators Coda-19: Using a non-expert crowd to annotate research aspects on 10,000+ abstracts in the covid-19 open research dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fad2b1ea-507f-4d2e-a830-d3b9e44fa91f · outbound
Evaluating Large Language Models as Expert Annotators Pubmedqa: A dataset for biomedical research question answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b92f1980-6828-457c-bc87-e8a9b0266c36 · outbound
Evaluating Large Language Models as Expert Annotators Gpt-4 passes the bar exam
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f50825cc-d97b-4807-baac-a61adcee6c86 · outbound
Evaluating Large Language Models as Expert Annotators Refind: Relation extraction financial dataset
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48a1eac5-5a42-4160-a50b-fd833586a3fb · outbound
Evaluating Large Language Models as Expert Annotators Large language models are zero-shot reasoners
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c933e7-86b6-4513-8c8e-6d0ebec54055 · outbound
Evaluating Large Language Models as Expert Annotators Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e66b29-847b-4193-bdea-3249e44aa8c9 · outbound
Evaluating Large Language Models as Expert Annotators Self-refine: Iterative refinement with self-feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac125cd0-dd4a-4c59-9341-46d9049aed29 · outbound
Evaluating Large Language Models as Expert Annotators Note on the sampling error of the difference between correlated proportions or percentages
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015ce650-1ee1-475e-ab60-f723fe4f1b30 · outbound
Evaluating Large Language Models as Expert Annotators Hello gpt4-o
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6380658b-753d-4b7e-a4a2-483069c51e97 · outbound
Evaluating Large Language Models as Expert Annotators Openai o3-mini
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bd27a646-139c-455a-a90f-9e1b16a8886b · outbound
Evaluating Large Language Models as Expert Annotators Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39af5a28-0b05-4d68-a5d0-524cab284565 · outbound
Evaluating Large Language Models as Expert Annotators GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4f6826-db1a-4fe3-bf6e-c0ffabb20674 · outbound
Evaluating Large Language Models as Expert Annotators Trillion dollar words: A new financial dataset, task & market analysis
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d3760b4-75fb-4393-b1ec-f8a4ad59da4e · outbound
Evaluating Large Language Models as Expert Annotators Large language models encode clinical knowledge
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6f28250-193d-4583-8c0e-52a1c449e99c · outbound
Evaluating Large Language Models as Expert Annotators Towards Expert-Level Medical Question Answering with Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23d2c8a-924e-4b8e-b24b-ea4a1e70e351 · outbound
Evaluating Large Language Models as Expert Annotators Large Language Models for Data Annotation and Synthesis: A Survey
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8488ec-0c4b-4283-b3ae-f7adc6aeb2bd · outbound
Evaluating Large Language Models as Expert Annotators Are Expert-Level Language Models Expert-Level Annotators?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf80663-a3d0-4615-ab02-2226a7e484fe · outbound
Evaluating Large Language Models as Expert Annotators Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666ab889-8119-4572-9afa-ea7a8261fe26 · outbound
Evaluating Large Language Models as Expert Annotators Foundational autoraters: Taming large language models for better automatic evaluation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dab46878-2ba9-4c4a-a4cc-256cdbdafad0 · outbound
Evaluating Large Language Models as Expert Annotators Cord-19: The covid-19 open research dataset
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d12a2d4-d2c7-4b0b-ba50-5b488092f484 · outbound
Evaluating Large Language Models as Expert Annotators Self-consistency improves chain of thought reasoning in language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c0cfab-0624-4c8e-896e-26cd8cd03869 · outbound
Evaluating Large Language Models as Expert Annotators Chain-of-thought prompting elicits reasoning in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80120f2-9e6a-4fc6-a9ad-4e0b6f294264 · outbound
Evaluating Large Language Models as Expert Annotators The rise and potential of large language model based agents: A survey
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f22bc1e0-d7fa-4bd8-9ba3-941f36e36b57 · outbound
Evaluating Large Language Models as Expert Annotators LLMaAA: Making Large Language Models as Active Annotators
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 687ed8e3-6ff3-4836-bc6f-f9702943e27c · outbound
Evaluating Large Language Models as Expert Annotators Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5ebd11-3889-42b7-a92d-ea7c3b9c517f · outbound
Evaluating Large Language Models as Expert Annotators @esa (Ref
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97540c71-885f-451f-a6d0-8d9139f62244 · outbound
Evaluating Large Language Models as Expert Annotators Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e82cdd5-a4b0-4c2c-8998-871c749e6f46 · outbound
Evaluating Large Language Models as Expert Annotators By leveraging additional inference-time compute, we explore whether individual LLMs can serve as a direct alternative to expert data annotators
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a8e6be-c643-4a74-8401-59d26806e229 · inbound
Agentic-imodels: Evolving agentic interpretability tools via autoresearch Evaluating Large Language Models as Expert Annotators
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7de7f1b9-5adf-4e10-b065-2fa6ed303ab6 · inbound
Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation Evaluating Large Language Models as Expert Annotators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.