Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2506.09443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09443 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:25:46.744451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T14:17:10.144588Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 37ec9b1c-1c4c-444f-9ef0-5ce82bf4230c · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:93e25c7169b5eef4c1e2e86503d7095faf45e26b8efa09c9487d543613df8623

Observation 1a800c6c-9f97-4cb9-bd4c-7c1e5c164ddb · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:02701614f1b2a297561e7c16e6e5f1846beebd661bd984a0758c13bd54406d41

Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · inbound

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs cites this paper.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:730f4a1e45c18d8a0f0e86013392542836c218f1942cbe421a9f00e138640456

Observation 5758604b-1fca-4aaf-a88d-965dea1cfe97 · inbound

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following cites this paper.

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:24:33.375282Z digest=sha256:d7a948227fa104328dcd65d1555942e052c53b4ad43c196b73af04841faa3977

Observation 0a58087e-4ad6-4a36-ad0a-4bcf37cd46a6 · inbound

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges cites this paper.

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T10:23:58.198464Z digest=sha256:357bd23fa18e95a67b046999e14e15da81336ae0663704b7f613c9f229a6489f

Observation cd2a0700-e44f-4903-9643-9dede2d6fcab · inbound

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator cites this paper.

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T08:50:44.523468Z digest=sha256:72f9964faedd89661ab754c1d690b480d3dc0f03f95d15d0c1cee977db28173e

Observation 42e23d4e-3b66-4f5d-b430-9de3132b0757 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:315132895c5bad3ddfb8b6a5b24d73c12c0042a163380215c97c02d827f6094b

Observation f252d99f-fc69-4c7f-b2f0-429b1dfa1656 · inbound

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions cites this paper.

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T06:37:45.175776Z digest=sha256:0217e946904b0fb6485db43a1526ad007198b916c666c36ef740569e18e91862

Observation 04277a2a-a34a-40df-ac8d-b9258216665a · inbound

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems cites this paper.

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:49.493310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-10T14:09:11.328304Z digest=sha256:a6c3ce4a87f1ddfa01c6f6b745ee4e04b9457971e38bde5a6932e7421aa0eecb

Observation 32b44645-bdcc-4df4-89d6-af2a75c8fa76 · inbound

Policy-as-logic for robust reasoning over rules cites this paper.

Policy-as-logic for robust reasoning over rules LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:25:46.744451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:25:46.744451Z digest=sha256:e3ca710735a239aa00dd752ceaf720fa762f058ccf99f4f6006e5d083f5ece4c