Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models as Expert Annotators

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2508.07827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07827 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:58:03.190450Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T21:36:22.389047Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T21:43:59.674035Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b936667-020d-4e5e-a448-86eb1033055c · outbound

This paper cites write newline.

Evaluating Large Language Models as Expert Annotators write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:58.492704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:58.492704Z digest=sha256:c53c2fc7ee606a9305e705dea08783f3b7c72045a33f48350658d20343745732

Observation b4c63d06-2427-4046-be84-810b6c61a8cf · outbound

This paper cites GPT-4 Technical Report.

Evaluating Large Language Models as Expert Annotators GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:58.634448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:58.634448Z digest=sha256:8d41aba0a408ef76d529e399aa7990c74f6e7bee3c3532132eac0acd943693d5

Observation 1d58986c-7e14-4e38-8c82-ee71a1672364 · outbound

This paper cites Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning.

Evaluating Large Language Models as Expert Annotators Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:58.829185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:58.829185Z digest=sha256:162b991707f1878919f4775a4258709695be7bb25449c50f1e59af3d519d0d59

Observation b0c71ad9-7577-4825-b493-a4d9c14aa089 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Evaluating Large Language Models as Expert Annotators The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:58.988358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:58.988358Z digest=sha256:7249691921e330d1671eb3e030b18256b8d9d37e09a973eb3086fd780deb43ac

Observation 529cf28d-af15-4939-80ac-bf7d8d69c3d7 · outbound

This paper cites Claude 3.7 sonnet and claude code.

Evaluating Large Language Models as Expert Annotators Claude 3.7 sonnet and claude code

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.436139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:57:59.152154Z digest=sha256:e83384ca09e8fbae583ec8dee9bad04cde9a73d545a9568eb1e0c3dcdcd2a9f9

Observation fb27085f-bb3b-4086-a198-57b7f2157896 · outbound

This paper cites Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost.

Evaluating Large Language Models as Expert Annotators Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:59.306767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:59.306767Z digest=sha256:5fe1fdbe93da805c6268822bad7549618480d6334216181190e83560056b6587

Observation f06ed9c3-bca1-4c83-b6eb-db6b4708b7c8 · outbound

This paper cites Must read: A systematic survey of computational persuasion.

Evaluating Large Language Models as Expert Annotators Must read: A systematic survey of computational persuasion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:59.405114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:59.405114Z digest=sha256:d9de173e622b4f93f094f2de1ae3e8ebbdb2fc6c7dfe55997a26c69dab19207f

Observation bae1414c-9156-4b55-b678-7dd155b038f4 · outbound

This paper cites Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams.

Evaluating Large Language Models as Expert Annotators Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:59.514351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:59.514351Z digest=sha256:66f21323cb8348278c2eb8cd42c7088a95128c4323a3ad2ee9aafcfc0766ec6b

Observation 1a55427f-961a-4075-a876-0c55903314d1 · outbound

This paper cites ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs.

Evaluating Large Language Models as Expert Annotators ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:59.640281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:59.640281Z digest=sha256:dba2409734c84e80152a857b62f6a95982e7f5d2cfbf50268fb0d17ceb9b2d05

Observation 5ae76b19-00d9-4629-9bc3-94eaa1c256df · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Large Language Models as Expert Annotators Evaluating Large Language Models Trained on Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:59.744177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:57:59.744177Z digest=sha256:d57657861daa17ae09c0d41fb17c0852a86b53061b7b4241487542971c59a244

Observation 517ec177-3791-4aa5-8cf0-ef478a09746b · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models as Expert Annotators Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:58:05.427175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:57:59.875292Z digest=sha256:66997b2cf09f76c83ef1bad1a172f9bfbfd344c6f4f05de9ddab13caca3812e5

Observation b4fce70c-4c5d-4064-8b83-7f70e29dbcd3 · outbound

This paper cites Chatgpt goes to law school.

Evaluating Large Language Models as Expert Annotators Chatgpt goes to law school

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.417901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:57:59.945208Z digest=sha256:2801031e420703c855c4b7378835190e3e2626d80260fc7056c8e3d90e872fe1

Observation 6872d044-9693-4b34-a660-6621fcd6be4b · outbound

This paper cites GPTs Are Multilingual Annotators for Sequence Generation Tasks.

Evaluating Large Language Models as Expert Annotators GPTs Are Multilingual Annotators for Sequence Generation Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.004961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.004961Z digest=sha256:698f50404165f1c1483f3b032a9131df3b253c9b781733cad2142cef932d2f99

Observation f02c00c8-f313-47e0-b984-c949688251ec · outbound

This paper cites Is gpt-3 a good data annotator? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 11173--11195, 2023.

Evaluating Large Language Models as Expert Annotators Is gpt-3 a good data annotator? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 11173--11195, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.408117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.061528Z digest=sha256:2f6bfaa3e12b650e8a60176599426f29250764d02faa9fc7bd9cee4d67b2ac65

Observation a7c5c73a-8d72-48e5-9517-8f6c6df3b268 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Evaluating Large Language Models as Expert Annotators Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.167703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.167703Z digest=sha256:075d56fcca25459a87bada410c8b57f2ac9960caf1940ace8e8c331b0a6e2be7

Observation 17132b9e-4e9c-4631-bf10-a58a9008b6d5 · outbound

This paper cites Measuring the persuasiveness of language models, 2024.

Evaluating Large Language Models as Expert Annotators Measuring the persuasiveness of language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.256478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.256478Z digest=sha256:e3fab751670dc04a839ee6010f8beaaac24782aad3be774ef857ff97ae138ec4

Observation de7121dc-276f-4728-ab52-9717e23e1123 · outbound

This paper cites Measuring nominal scale agreement among many raters.

Evaluating Large Language Models as Expert Annotators Measuring nominal scale agreement among many raters

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.392774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.323738Z digest=sha256:9ee3b9f90abfd7c347d162a88b0f34ab1e806179a07dd0ee0dd9cbfa1c56c35a

Observation e9c6e715-aa56-4a68-95f0-99378478ba55 · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.

Evaluating Large Language Models as Expert Annotators Chatgpt outperforms crowd workers for text-annotation tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.383309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.395214Z digest=sha256:c17592cf0b8df837739343dc110d5e9cf082a2c6909012aee4fd50b18175fea3

Observation 56a7c1a1-538c-4428-95bd-b23114d5e49b · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Evaluating Large Language Models as Expert Annotators Introducing gemini 2.0: our new ai model for the agentic era

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.373126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.441674Z digest=sha256:ccfdf961320b47099e672fca5ef8277843e7b98417f1381bdac0617342cce010

Observation 986d6739-07a7-47d7-b125-516761e6fd95 · outbound

This paper cites Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models.

Evaluating Large Language Models as Expert Annotators Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.363839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.556464Z digest=sha256:84297be45474322a5fae4595be686ec2d47e375dea515f2ab7f4b94c82b92271

Observation b75d865b-3672-4cb4-b829-b76d31677352 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Evaluating Large Language Models as Expert Annotators Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.622649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.622649Z digest=sha256:3d5ec97df447cef96763762dd362daa976c0f1be38024c979fff270cb970584a

Observation b214e203-5069-48d5-a24d-1bd72469d499 · outbound

This paper cites AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators.

Evaluating Large Language Models as Expert Annotators AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.717175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.717175Z digest=sha256:0acaaebb085fd2fb1495b7ed02abbe98de916a40a6b0b98ac8e58a35229b9fa5

Observation de5df38a-bde8-4855-9660-b973c203d59e · outbound

This paper cites Measuring massive multitask language understanding.

Evaluating Large Language Models as Expert Annotators Measuring massive multitask language understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.790500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.790500Z digest=sha256:e3e1649f4bb82cab5ef7f778e6a3a914436086c5b540a038ff5998f08bef3a74

Observation 4c4a376f-c68b-4e41-8805-cc10a7c880d8 · outbound

This paper cites CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review.

Evaluating Large Language Models as Expert Annotators CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:00.863174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:00.863174Z digest=sha256:8411971ba1df0592eb5b62eb7a9aee9452648fcc7271f79080149e236d1af3bb

Observation 4f533cea-5aac-4b70-9d42-58f50de4ed87 · outbound

This paper cites Coda-19: Using a non-expert crowd to annotate research aspects on 10,000+ abstracts in the covid-19 open research dataset.

Evaluating Large Language Models as Expert Annotators Coda-19: Using a non-expert crowd to annotate research aspects on 10,000+ abstracts in the covid-19 open research dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.347992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:00.938390Z digest=sha256:d095b0d4dd21e05eb517a4cc68fdbb371b7fd5a4afb1563c25c682589f46190d

Observation fad2b1ea-507f-4d2e-a830-d3b9e44fa91f · outbound

This paper cites Pubmedqa: A dataset for biomedical research question answering.

Evaluating Large Language Models as Expert Annotators Pubmedqa: A dataset for biomedical research question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.234433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.028600Z digest=sha256:1b3a1610d815482db46199f0a45bef7a59d432b47335e402bb71688836a40440

Observation b92f1980-6828-457c-bc87-e8a9b0266c36 · outbound

This paper cites Gpt-4 passes the bar exam.

Evaluating Large Language Models as Expert Annotators Gpt-4 passes the bar exam

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:05.085182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.127299Z digest=sha256:c34d7512cabd76ade13d203cf0496e40571a8ec78a0f6f8974247bc7286260b4

Observation f50825cc-d97b-4807-baac-a61adcee6c86 · outbound

This paper cites Refind: Relation extraction financial dataset.

Evaluating Large Language Models as Expert Annotators Refind: Relation extraction financial dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.966746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.208891Z digest=sha256:86fae6b6f5354f435d47327f15869c3763c1b0eda26d4eb587bae0fdd248efd7

Observation 48a1eac5-5a42-4160-a50b-fd833586a3fb · outbound

This paper cites Large language models are zero-shot reasoners.

Evaluating Large Language Models as Expert Annotators Large language models are zero-shot reasoners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.287975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.287975Z digest=sha256:95d911a7b5bcf2982602f85b8cfd4ff3248bb0ae27db1f04fcf2580907043e17

Observation 75c933e7-86b6-4513-8c8e-6d0ebec54055 · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Evaluating Large Language Models as Expert Annotators Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.346989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.346989Z digest=sha256:853f2159a285621035b3f7887e31a6be46a4be6bd3f7241cdf9e1dd1d1f7c13b

Observation c6e66b29-847b-4193-bdea-3249e44aa8c9 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Evaluating Large Language Models as Expert Annotators Self-refine: Iterative refinement with self-feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.425203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.425203Z digest=sha256:2251c4cd7fe8263a9638cbc14f4bc8f4fdd2e311ff46e4189c52eca440aa696a

Observation ac125cd0-dd4a-4c59-9341-46d9049aed29 · outbound

This paper cites Note on the sampling error of the difference between correlated proportions or percentages.

Evaluating Large Language Models as Expert Annotators Note on the sampling error of the difference between correlated proportions or percentages

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.502147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.502147Z digest=sha256:165e2e4dc23715bb15d498134476f9ba2ee54be78d8afc0a5c75d845d5b7ec84

Observation 015ce650-1ee1-475e-ab60-f723fe4f1b30 · outbound

This paper cites Hello gpt4-o.

Evaluating Large Language Models as Expert Annotators Hello gpt4-o

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.825842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.606341Z digest=sha256:e4c0749e91e240fc906b79025680eb6066c671ef2928d7d30f74d36f83689fe6

Observation 6380658b-753d-4b7e-a4a2-483069c51e97 · outbound

This paper cites Openai o3-mini.

Evaluating Large Language Models as Expert Annotators Openai o3-mini

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.688878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.663633Z digest=sha256:be2a7349d9970b9e0b4f2c5ecc916ac76fb2dec8579362675eb574ee9f40b39d

Observation bd27a646-139c-455a-a90f-9e1b16a8886b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Large Language Models as Expert Annotators Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.771395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.771395Z digest=sha256:1ab3e7fd42d35c509f98d37ea3a420c5498966ffd9ac73ff3954bf065b460b78

Observation 39af5a28-0b05-4d68-a5d0-524cab284565 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Evaluating Large Language Models as Expert Annotators GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:01.848446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:01.848446Z digest=sha256:6bed5879953424ad5c944be104f6d7e3a66e66dd1783c9ad35e8b61e938b1a16

Observation 5f4f6826-db1a-4fe3-bf6e-c0ffabb20674 · outbound

This paper cites Trillion dollar words: A new financial dataset, task & market analysis.

Evaluating Large Language Models as Expert Annotators Trillion dollar words: A new financial dataset, task & market analysis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.560419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:01.919674Z digest=sha256:85932c888981c840510bea5bd5de1ac79511b5fe179d4772358b37415ad2d916

Observation 8d3760b4-75fb-4393-b1ec-f8a4ad59da4e · outbound

This paper cites Large language models encode clinical knowledge.

Evaluating Large Language Models as Expert Annotators Large language models encode clinical knowledge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.399509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:02.006829Z digest=sha256:e8324a54013f2189d7bfb83778cdfd670dc1ee9f9675a9d042953fc2cba71617

Observation e6f28250-193d-4583-8c0e-52a1c449e99c · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Evaluating Large Language Models as Expert Annotators Towards Expert-Level Medical Question Answering with Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.067890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.067890Z digest=sha256:fc37949d88cd994a6a6dd95249cda66e14d8d12171c75036550b64a8b826a987

Observation b23d2c8a-924e-4b8e-b24b-ea4a1e70e351 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

Evaluating Large Language Models as Expert Annotators Large Language Models for Data Annotation and Synthesis: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.162770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.162770Z digest=sha256:d26380079ac2d587c913783fce01d945be9ec36c209e5f73e152a079792a5131

Observation ae8488ec-0c4b-4283-b3ae-f7adc6aeb2bd · outbound

This paper cites Are Expert-Level Language Models Expert-Level Annotators?.

Evaluating Large Language Models as Expert Annotators Are Expert-Level Language Models Expert-Level Annotators?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.244277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.244277Z digest=sha256:b13f85ea2df92918760dcedd40f94ed976d3bee7e13001e3a043fe76c5abeb35

Observation 6bf80663-a3d0-4615-ab02-2226a7e484fe · outbound

This paper cites Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization.

Evaluating Large Language Models as Expert Annotators Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.325509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.325509Z digest=sha256:1925e567fdfa7db46ed72e00d23b45b5a62d3799e150e877b380383b6ed47318

Observation 666ab889-8119-4572-9afa-ea7a8261fe26 · outbound

This paper cites Foundational autoraters: Taming large language models for better automatic evaluation.

Evaluating Large Language Models as Expert Annotators Foundational autoraters: Taming large language models for better automatic evaluation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:04.187402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:02.380263Z digest=sha256:08d611288aa73447fe5b36dd6cdd88fbc71f19cd0f85228efcda93f351f543b4

Observation dab46878-2ba9-4c4a-a4cc-256cdbdafad0 · outbound

This paper cites Cord-19: The covid-19 open research dataset.

Evaluating Large Language Models as Expert Annotators Cord-19: The covid-19 open research dataset

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:03.971639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:02.438362Z digest=sha256:930bc0885a6c0addf822d2edd4fbfe3689aecf6007765e4d946f3db5961fc356

Observation 1d12a2d4-d2c7-4b0b-ba50-5b488092f484 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

Evaluating Large Language Models as Expert Annotators Self-consistency improves chain of thought reasoning in language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.534703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.534703Z digest=sha256:2a26f5049d00fbc93c396d8c637570f2f04696a47633cbf05792e1a0f4688716

Observation 70c0cfab-0624-4c8e-896e-26cd8cd03869 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Evaluating Large Language Models as Expert Annotators Chain-of-thought prompting elicits reasoning in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.595024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.595024Z digest=sha256:63db9d8ff66aed91b404afd16484cae8b2185f9f59d78e140404ed630fb6da44

Observation e80120f2-9e6a-4fc6-a9ad-4e0b6f294264 · outbound

This paper cites The rise and potential of large language model based agents: A survey.

Evaluating Large Language Models as Expert Annotators The rise and potential of large language model based agents: A survey

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:58:03.823825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T21:58:02.612999Z digest=sha256:3b7f380dabfed26ddcd31805d9ff6fc1945c713f49f9d9a9b6a74651cdae35ae

Observation f22bc1e0-d7fa-4bd8-9ba3-941f36e36b57 · outbound

This paper cites LLMaAA: Making Large Language Models as Active Annotators.

Evaluating Large Language Models as Expert Annotators LLMaAA: Making Large Language Models as Active Annotators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.659226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.659226Z digest=sha256:b6fa20e0c5bfdf7992ca6cc7d7fd4d836d3facab7ccfb00e1b54407807727551

Observation 687ed8e3-6ff3-4836-bc6f-f9702943e27c · outbound

This paper cites Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks.

Evaluating Large Language Models as Expert Annotators Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.738693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.738693Z digest=sha256:514d5487aae34b4b0706f9ac1c82a80acdc3a204af295ee486de0675757355db

Observation cc5ebd11-3889-42b7-a92d-ea7c3b9c517f · outbound

This paper cites @esa (Ref.

Evaluating Large Language Models as Expert Annotators @esa (Ref

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:02.878965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:02.878965Z digest=sha256:0f77e503ae1f47d7f0caeed53c204d45c1aa35f41d01b120a9cf2ea3faa705f0

Observation 97540c71-885f-451f-a6d0-8d9139f62244 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models as Expert Annotators Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:03.021696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:03.021696Z digest=sha256:fb17acd1672513638e066ecf7ebe97a12fec12b85ec0ddadf44f075b06ea6ec4

Observation 5e82cdd5-a4b0-4c2c-8998-871c749e6f46 · outbound

This paper cites By leveraging additional inference-time compute, we explore whether individual LLMs can serve as a direct alternative to expert data annotators.

Evaluating Large Language Models as Expert Annotators By leveraging additional inference-time compute, we explore whether individual LLMs can serve as a direct alternative to expert data annotators

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T21:58:03.190450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:58:03.190450Z digest=sha256:3b9670a78126af00e582fe0042dc54444b629c241a6695c9f30574a1ea075872

Pith citing papers

Observation c8a8e6be-c643-4a74-8401-59d26806e229 · inbound

Agentic-imodels: Evolving agentic interpretability tools via autoresearch cites this paper.

Agentic-imodels: Evolving agentic interpretability tools via autoresearch Evaluating Large Language Models as Expert Annotators

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.282145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:37:43.371592Z digest=sha256:f282c11d028fdf1b1d79c3112bb895ba9014e4e250bb35cbca96ed23eebe3b6c

Observation 7de7f1b9-5adf-4e10-b065-2fa6ed303ab6 · inbound

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation cites this paper.

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation Evaluating Large Language Models as Expert Annotators

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:43:59.676021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T21:36:22.389047Z digest=sha256:159eaceb93da8175934b6e76732be8bdf08fd7958637a66e90879939bc313893