Pith. sign in

Paper Citation Record · LEDGER

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2608.07813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07813 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:11.112009Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfb53ac5-7f8d-4fd9-bf39-5e8b2690071c · outbound

This paper cites Let's Verify Step by Step.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Let's Verify Step by Step

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.048811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.048811Z digest=sha256:6f637fd26cc18b134586090407eddd788dc8c3cddddc05e9811c06b978e02337

Observation 32bcac5d-3cb4-4e06-86ba-4e576eae9dcb · outbound

This paper cites Factscore: Fine-grained atomic evaluation of factual precision in long form text generation.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Factscore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:18:11.385320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:18:11.059931Z digest=sha256:460e89c6cfd6b0ee839037ed03b267256f977fd9faab9c1c84c20f8fa6a6522f

Observation fdf26fe2-76bc-4506-95a7-fb8b3621f7d1 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.741.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines doi: 10.18653/v1/2023.emnlp-main.741

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.065092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.065092Z digest=sha256:84397a3afed40a9dc0dc2cda6e7f0244388517f5f07c2a9b1d97701eb649ec9c

Observation 5ead05c0-3feb-4beb-8912-c0bcb3e264be · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines WebGPT: Browser-assisted question-answering with human feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.070300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.070300Z digest=sha256:9cff300482c38dc614bfa10b141f1b47c9a73089adef4e4b942889ac0f9286a7

Observation 928b9da8-5f40-40aa-971a-76b986046821 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines LLM Evaluators Recognize and Favor Their Own Generations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.076461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.076461Z digest=sha256:7317202a5a0121028c261c268a6922035314a72fb94400ebb672a949f0d2d817

Observation a61de5b6-5a26-4e55-84ce-97b33d266fbe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.081607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.081607Z digest=sha256:2439fde6a70bc4b4a0d2f9b73351ab184eb4d3598000672570b6ad8127ad88c6

Observation f5e30d92-2f1f-40ae-8c06-6fc1ad964346 · outbound

This paper cites Large Language Models are not Fair Evaluators.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Large Language Models are not Fair Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.086784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.086784Z digest=sha256:bcb423467a85be23132cc0ac67d71a2299f12111c3a218908aab9d1b7c9407b3

Observation 3c0b3d1d-a3da-4811-8d8c-9bf99f87963d · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.091647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.091647Z digest=sha256:1d47b365003099c099e6a212ccf73c75b7b4a60e6fa9ab1a87a18b695fea9833

Observation 15392555-41bc-488d-9c08-e7e00bd8c9fb · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:18:11.368928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:18:11.097139Z digest=sha256:4ec2dd4ec2d5827dbcb259a18bf515997050de17d6cbb9cf62bd9fd1c75e290a

Observation 12235552-9bb1-4cdf-aa37-c5782f4aaa01 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines ReAct: Synergizing Reasoning and Acting in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.107248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.107248Z digest=sha256:33d633364c2284314743f5dba720ad143ea9ee540283a821e42fadbd87c12049

Observation 4b695ca7-2953-4cd4-8267-83efb4937583 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.112009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.112009Z digest=sha256:7cb32fbe9197440f0da85056e939b4f786f3ee0c79ad62d91b2ba9f0a49899bd

Observation 741af49e-e7b8-48f2-a4c7-fa7288e9f527 · outbound

This paper cites doi: 10.18653/v1/D18-1259.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines doi: 10.18653/v1/D18-1259

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.102124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.102124Z digest=sha256:9ec2c777b7adcfbd3f35023f61e53ee1c6dcaed59b0e107c3720366f8f201f31

Observation c6f116b2-f598-4d8d-aa97-87f9135693f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.043369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.043369Z digest=sha256:a5dc8088b397077ed095a7ea92fff082074977a19889cec8b6ea5bfd81ac78b3

Observation 50a2fe24-2a3a-43d7-b97e-4cd1651ea685 · outbound

This paper cites FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.038075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.038075Z digest=sha256:66dbb389d199bfc913601551e74bcb0a9aea6ae4f8ed857674046a4fdeb5ee7c

Observation 705fdad2-3aba-444a-9f4c-cd2385e72fc3 · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.032070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.032070Z digest=sha256:5dbde09516b4de97f9acf041b8b54fc6b614b21840cb1b284d1430c3daa23469

Observation d7be8218-0e2a-4828-bc27-b9e00949bb21 · outbound

This paper cites SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.054622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.054622Z digest=sha256:4d606162819311db2d45efa9946a0d9b204641f18050665b0f821f70805208fb

Pith citing papers

No inbound Pith citation observations are available.