Pith. sign in

Paper Citation Record · LEDGER

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2503.05061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05061 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:48:47.164830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.653154Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07198add-f6ef-4dbe-a58c-f159f398f067 · inbound

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation cites this paper.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.164830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.164830Z digest=sha256:099d710b2cd1b5e15140b2fe9539e144beb41de08ee71f7f82f61e23266e0d9f

Observation e7697d9a-a7b0-4a90-af30-9147f7f8c8f3 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:458656f9513edcc592a3ae4750beb359777012b422692401992543f96543ebd9

Observation e90a2fbc-d89e-4962-a6a7-3d9496f2359f · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:7d9e1e8c75cc4b86abbc1a102231736117a88b61a07fee85023599f87d50c70e

Observation d3be9a4e-a668-452a-acfc-b1b922cab7d6 · inbound

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models cites this paper.

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T21:57:16.279587Z digest=sha256:ed2a43adca81ae2f850c5c060cd9de4b5f78bfe5cbf5ac152b18b85e885ecb31

Observation fcfa7d4a-a0b7-4b81-92a0-0671ba477b79 · inbound

When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation cites this paper.

When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:57:21.981236Z digest=sha256:b1da00472f125998a8527958c44b67789ea15190b6f1e84505bb73f550e51009

Observation 141e6aa8-1ef3-4c27-b759-9604657c2c0c · inbound

Medical Reasoning with Large Language Models: A Survey and MR-Bench cites this paper.

Medical Reasoning with Large Language Models: A Survey and MR-Bench No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T10:21:39.892271Z digest=sha256:78bdf331f784a6a1f475e4bffeb675b8dbf2c7b5ec96fb2a371868eebc00e065

Observation 66cff144-3f95-4d56-8869-a36f86835261 · inbound

Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance cites this paper.

Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T05:27:50.049798Z digest=sha256:5fd5f1179228b40805d5471b4fee737540104a9e45f7ce13a0078ca5220fcade

Observation e1ced17e-5ab4-446b-9e9f-30a7ac1159fa · inbound

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety cites this paper.

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T04:17:43.880661Z digest=sha256:61cc385ebecc836ffd27466e590303d96f220c0b9c3db4bcb108dd64daf9a0c3

Observation 4a2f157e-b045-4c09-ae2c-2b4d601ff26b · inbound

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents cites this paper.

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:27:16.974169Z digest=sha256:fd18913bc01edf06d56a1bd59fbe17a609164da234c3cd6c95e4a7db426ed427

Observation 8ef0b962-b5b7-4453-b862-3e0aecdd86e5 · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:2eaf51d02915dd131720d62de66185f3dd531b49270d217a193ca7590de1872e

Observation a84fdf79-e3e6-48f3-aede-20d1233afc3a · inbound

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring cites this paper.

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:29:01.722453Z digest=sha256:0f72a55443e165e6bb54131262b246d43017e42e88d4857d6e4aa98df7581506

Observation 9cef6ea9-6279-44ea-9322-f04cdf25958e · inbound

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark cites this paper.

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T19:17:17.373463Z digest=sha256:d49ad35549c6c5508144f8cbc40efdccd3a51a9618fa17c91e654714b3bf325a

Observation 935e1686-125a-4747-8f98-7cda50e740e7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:26:03.833426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:580b0f3a726a542a0c3228257798e6719537537779541c83c6eab390aae197f9