Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals

As of 10 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.08809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08809 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:11:41.723677Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1155b71-e107-4bbe-aeae-b27a94747926 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.593705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.593705Z digest=sha256:3d817edde67311db79fa380e338f2029663e0d4417030b390f3205787d850684

Observation 780f3e7d-73e0-4609-81b2-53a1f4c6ece5 · outbound

This paper cites Self-teaching prompting for multi-intent learning with limited super- vision.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Self-teaching prompting for multi-intent learning with limited super- vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.606910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.606910Z digest=sha256:3933b6aa934f91c244f64eb8a267458a0c999c63d91745c0181e27ff3576bffc

Observation fbd2193e-118d-4d9c-bd11-fc22c1d5cc00 · outbound

This paper cites An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.630126Z digest=sha256:8a9693f1ae84e6798e7ee77f021950ae3a63b95335e4bbaf0e0aa0f0473c0b7c

Observation ec452bde-1b91-4f30-835a-199a972eb0e2 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Model Guided Tree-of-Thought

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.647592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.647592Z digest=sha256:8757f7191b55726291067f64957ccd113eaebbcb4afdb455d8ebc10e753c1bfd

Observation 92527c59-d016-460c-89fa-24f1b07f8594 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models for Data Annotation and Synthesis: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.653118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.653118Z digest=sha256:f1a248c3b118e43bc63f59fcc599992ac9ed1e4ef69d0a6e11401f0b60242aab

Observation 8b125de1-3aa7-49cb-8222-c159da1a5b8f · outbound

This paper cites CodecLM: Aligning Language Models with Tailored Synthetic Data.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals CodecLM: Aligning Language Models with Tailored Synthetic Data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.660463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.660463Z digest=sha256:f3f43fab3ef66e497655ef0aedc271a6ed1b4d8d2d580ff3b7904b256af50349

Observation 1f8d1606-cbe6-497b-9bbe-dbc658520440 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Finetuned Language Models Are Zero-Shot Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.668846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.668846Z digest=sha256:323e2dbc5a51ba3229a7d2ac7a7f0e087315cb9ace467ac353600a3a137a447b

Observation eaf3db79-79cd-464d-983f-d6e95b491a94 · outbound

This paper cites Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Unigen: A unified framework for textual dataset generation using large language models.arXiv preprint arXiv:2406.18966,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.676658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.676658Z digest=sha256:6a30069c6030edd39750b48d31439c25ff288ab58798d16ba079b892a94ab931

Observation 2e045291-eb2a-472f-b5bc-c7a70f28a214 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.683557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.683557Z digest=sha256:84302fbec74c1e066b32266b6aebe0a84f90f6ee354831bd6b6dba4e358e2d51

Observation 338a3504-808b-4418-8baf-c4d8a489ba21 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.691968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.691968Z digest=sha256:13727f2b3bbd44a8ebed863137b9268173c6a4e3501de0a58a391cabf2c28491

Observation 30e50ffc-9452-48bb-ac5d-c14f7fac065b · outbound

This paper cites ZeroGen: Efficient Zero-shot Learning via Dataset Generation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals ZeroGen: Efficient Zero-shot Learning via Dataset Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.698533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.698533Z digest=sha256:594126a118aea9661c2c1a839f5f05be98fdfe75a4510d1b370ebbdc808bd76a

Observation 07457aea-99d2-4fe8-a7a1-b60ed49e0470 · outbound

This paper cites Clusterllm: Large language models as a guide for text clustering.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Clusterllm: Large language models as a guide for text clustering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.711658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.711658Z digest=sha256:47a5fdb9242185a2867c08ea3ddd064fbee263495241b4536293c58f59fef8d9

Observation e8ead15a-d86e-4033-a881-23c70738845c · outbound

This paper cites Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.717627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.717627Z digest=sha256:e4fb97f39a18141bededa5069466134881971bde20fa2dec6cfb759006e24064

Observation 7c32b9ed-be89-420c-b367-3de21fab166e · outbound

This paper cites yi denotes the LLM annotation accuracies.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals yi denotes the LLM annotation accuracies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.723677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.723677Z digest=sha256:d32a323a405f88eb695f1c6b9e7377cb8eeeae719b86a6d1515eac12b3286a31

Observation 70bfafd6-52d9-4e8e-a797-b91e80301d73 · outbound

This paper cites Large Language Models Can Self-Improve.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Large Language Models Can Self-Improve

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.623904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.623904Z digest=sha256:2178b932b6ab3e25e66e453c0d1f517414cb4a0719281b93f85aac170537bf0e

Observation 8ddb5ba5-2b9b-4000-8a18-fbe2b55a60cb · outbound

This paper cites MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.635830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.635830Z digest=sha256:195fc7a96259bd1a1c98e4e8a6420da835881b4edece7a80971af161535ef4f5

Observation 7c8b7f60-0c63-4701-888c-dcdcc7e57049 · outbound

This paper cites Efficient Intent Detection with Dual Sentence Encoders.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Efficient Intent Detection with Dual Sentence Encoders

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.601157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.601157Z digest=sha256:2af30997133279451e38ec35a56960c9d657a4e2d6bfa9a3e1acbaee91811e67

Observation 4ace878f-ff85-441e-ae45-d2f847d49ad1 · outbound

This paper cites New Intent Discovery with Pre-training and Contrastive Learning.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals New Intent Discovery with Pre-training and Contrastive Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.704280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.704280Z digest=sha256:2d9616b0e9ea782f615df6b2ea1fc11a6e0c3c54d34ec0e1c9438db426a70127

Observation 71e14911-340e-4b2c-bdec-063d2428ef9e · outbound

This paper cites TWEAC: Transformer with Extendable QA Agent Classifiers.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals TWEAC: Transformer with Extendable QA Agent Classifiers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.612448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.612448Z digest=sha256:1fc0277b0ef7fa1ce40ecd255b6fddcfc6d616fb8aa82b954221e91256cd5014

Observation 8d39323f-e64d-429f-abf4-1048cb2d4f7d · outbound

This paper cites FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.618338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.618338Z digest=sha256:65562645dd842e284cf73043ec9188664ba2dda27392a6d2a6c57c387efee6b2

Observation b97cea59-8b75-4c4d-b098-1dd48215103d · outbound

This paper cites Graphprompt: Unifying pre-training and downstream tasks for graph neural networks.

Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals Graphprompt: Unifying pre-training and downstream tasks for graph neural networks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:11:41.641752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:11:41.641752Z digest=sha256:9229be223181a5a5f358bd268ab1f238b2ca657bf467123d70ff0d45406c40ea

Pith citing papers

No inbound Pith citation observations are available.