Pith. sign in

Paper Citation Record · LEDGER

Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2407.10817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.10817 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:26:06.997800Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:27.459119Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2976c262-f0a4-454f-b32e-3bf84bdadcd8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.575421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:3752237108f0808b717b7b065e8c16d6128688196d3924350eab9d14beea05ef

Observation af6ecb3f-2016-40f0-95d1-db9311c2b413 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.997800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.997800Z digest=sha256:25108c3c6fcc0b2fc53625962930c56b20df9108a00d474a478a73ac7cb3ff62

Observation 795ba532-f5e8-431e-a83c-0bb54cc36c61 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.469756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:9a1944ea884aa19054d0de2123471635a1a7f4303f6c9753eb1eafd2d4f5c069

Observation 23974def-50e9-40c3-b46c-5dc512e23a69 · inbound

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety cites this paper.

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:13.271655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:13.271655Z digest=sha256:38b09f1206138e6a1db6310ca86dd2d6c20b722ce556ccc597ae1b63ae3f9cc2

Observation 240afa34-a1fe-46c9-a726-ea6ec02e62c6 · inbound

Do Biased Models Have Biased Thoughts? cites this paper.

Do Biased Models Have Biased Thoughts? Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T22:39:28.084796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:39:28.084796Z digest=sha256:cab2698e1d1b10a018628e0852bb996033b42cf447ce51ef2c1331c694f69bc7

Observation 9da63934-2330-48fc-a9c6-0ad8e789d7aa · inbound

CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments cites this paper.

CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:36:50.269769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:35:29.338034Z digest=sha256:db77360dbbb8d962355890b165f715354f97606f3c3ae83707a05227f2dd2ad0

Observation 9fa42a41-edc6-4d17-ac23-407057dcd7dc · inbound

EvoSkill: Automated Skill Discovery for Multi-Agent Systems cites this paper.

EvoSkill: Automated Skill Discovery for Multi-Agent Systems Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:25:05.453288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:25:05.418500Z digest=sha256:777daed3f57574108ec2c8db37a8cb7e7cb4a5101a7d0d1a6ce2d19e5d904217

Observation bb20f579-9ca9-4c12-84b4-8ce6d72d0dbc · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.460931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:ca00786a876cfd7c40eacdd8c765d5797317608e9f65774319cd412e675de93d