Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2401.16788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.16788 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:07.304009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:25:53.828114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd5030dc-b7ad-4f50-91ec-f625b2d85ee4 · inbound

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data cites this paper.

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:07.304009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:07.304009Z digest=sha256:a002a1c1bddee4e088710e1dc958b93e38854cd6e80389ef895beb5818bd47fa

Observation cd6a15b7-77aa-44f1-8b53-a9df7295ca61 · inbound

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate cites this paper.

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:10.537701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:11:10.537701Z digest=sha256:97fd5fe0c66c2a5377a8492cc7ff370c63739bb78c5c7190b16fe058a9eb1722

Observation e6ee4100-cff8-4116-ad67-63afce0e9c73 · inbound

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials cites this paper.

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T22:56:20.064068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:56:20.064068Z digest=sha256:b353201d6899a0bda777a86ddaa32b26a019f0bf09f1424817277e0949831a63

Observation 93a1ffad-dabd-437f-8aa7-8d98a91b72ca · inbound

Learning to Interrupt in Language-based Multi-agent Communication cites this paper.

Learning to Interrupt in Language-based Multi-agent Communication Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.830634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T19:08:45.818851Z digest=sha256:055d76ea1d0d9a15e445e63d0314b1a75349b10c85bdc839d890a02d53142270