Pith. sign in

Paper Citation Record · LEDGER

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2501.10970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10970 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:04.898211Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba4a2884-781c-47f3-be0c-4ac0a44f9252 · inbound

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective cites this paper.

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:15:21.289225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T03:14:00.526112Z digest=sha256:eaa55e4773ca1a0e05d82c1dc3a5e47ec1376ffb8fe1b9530dc706ffbf171b1b

Observation 23413ca2-6ffe-41c4-ab1f-82c3100367c4 · inbound

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs cites this paper.

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:04.898211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:04:04.898211Z digest=sha256:941c372ff3ebba0f6d414aad33ccaa80cd9dd5e46ee223a0be6942278eba5964

Observation f76b4449-d78a-4ec0-baf0-65a16d2ba282 · inbound

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods cites this paper.

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.541169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.541169Z digest=sha256:f0f33aaf4f6617e754c05c3b14bbdb4cff7f003b649ef91a500ea58ad3593c1e

Observation 88be4fb8-bd12-4cb5-88ef-726c42c7928a · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:06.332229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:06.332229Z digest=sha256:81f474cf004680b3879fce767d75741ccb4a734c462a6ca9042f0dfcfa0e3bb2

Observation bb2a8f9b-2df4-4727-9cb3-89df77f03b3c · inbound

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering cites this paper.

ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.542384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.542384Z digest=sha256:f526da01a84ef6011b56a019d672ce92ee6b2e720c6e98c3221d48a5b55f8afb

Observation b958903d-046f-4147-8229-74b87bd1188b · inbound

EduCoder: An Open-Source Annotation System for Education Transcript Data cites this paper.

EduCoder: An Open-Source Annotation System for Education Transcript Data The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:37:05.540440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:35:52.219876Z digest=sha256:bc6975d4fa45fe0db2d08c3ad0c65c510e2cf5839aefc8a13c2a884ae18f60a8

Observation cbd87949-a508-4ff1-9728-fbf3ffe6c530 · inbound

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications cites this paper.

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:40.826384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:49:40.826384Z digest=sha256:4c70852c61bac9634351ff6b96b0b31bc02e5f43ca69946c59390745aae0d118

Observation 0aeb6484-a2ca-438d-bea8-95f5b98646f7 · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.683550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.683550Z digest=sha256:e8b94a3a8cbe981b0248a30280cf9e3691b175310d8e0b8e8ccb0dfd4397c3ac

Observation 05fbd995-8ae2-4e05-9f4d-6350a5a8bb39 · inbound

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models cites this paper.

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:33:07.548115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T15:33:00.984344Z digest=sha256:fecacec05ccde73eb1f9a6553350ed6555274a6024e0da3aedf4f60a31de9195

Observation f17c4560-d09f-4ffa-9ef6-05541150e15c · inbound

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations cites this paper.

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.726078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T18:27:50.159580Z digest=sha256:22f137b344f4f954bdaa64c4ad20bbd9e8f0ff6c902757b07bd4e3a38b6dacca

Observation 04c52b5d-503c-4852-b7e2-411c3c0f2c35 · inbound

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe cites this paper.

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:58.149738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:58.149738Z digest=sha256:db0c0d8e7e4761a03b3540fa25f2de00a7599fa514eed27fe87cf67d6c7d8df9