Pith. sign in

Paper Citation Record · LEDGER

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2504.09946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09946 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:26.281619Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:46:58.798974Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f023a88-ad66-4856-bf90-666a5247da9f · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:26.281619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:26.281619Z digest=sha256:4ed2d8c42c9a11b3f14eaf164bf9b0f8d97d478552a64df8cf16d699831f67ab

Observation 4602b950-ac16-45f8-9cdb-edb20ef4be25 · inbound

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading cites this paper.

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:11.379436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:57:11.379436Z digest=sha256:2a8156c0f689802b8760a5fc417dc454705dd4fdafe11ff374c8dc5c0cc7144b

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · inbound

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning cites this paper.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:0fa7f569ae585d5f7adfa1f7f75c60f1a29c0791b57bfc8d7cbd754d422e0ea1

Observation 5c18f31d-decc-49ed-9033-a5fc6ef24b86 · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:39.513701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:39.513701Z digest=sha256:cff17d9e066a172287487bbabafe8b064dae20ddafadfcc8216cd4c7ff702ed3

Observation 8bdcb376-98e8-4552-9ed3-189f805dad09 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.098978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.098978Z digest=sha256:57c90190557b83ee335193718257d02ac815e4f72bea1159560612db7622176e

Observation 38f9930e-ce24-4312-8ac3-6a5e0fb3a48c · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:09.271466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:09.271466Z digest=sha256:feffdc542d27dc02d2fcf12aab7b6ee91fdd81c1f66f288827707314328c77db

Observation 079480e8-7c06-46c0-a94e-dc321e49ed16 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.228860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:eb3e2765e835583a8f10ab6c4c9c2a4402f311020b28af82020509a438239520

Observation 8a2caa14-9a8c-4dca-b0f2-00ba14ba491e · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.880554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:9975d7d86b86b2fcb155e328921cf120103adf04534a860ca4d3aaf2e7a73fc6

Observation c37bc6d8-22b6-47f0-9ab2-6df4471574f2 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.715805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:a7f566dd653e803e5898b156078f392ce287119f7dc19333bbfef77804de7db1

Observation 1fdf6cc4-d633-4575-a4fd-0dbcdededa35 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:28.071739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:72cb1ec152cc09abd496e65dd9d9a1178f84c95fe0969605f078131fed8b8f7e

Observation 04ef2b08-c858-4f54-97a1-7758616c055c · inbound

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents cites this paper.

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T09:53:00.677605Z digest=sha256:1aae46b66c21ed4c67fcf0099a2c0569fce859b777cb766401bbcd91f024f218

Observation de7ab540-249c-4c37-88ba-c09b73a8695a · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.979242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:8b0dac3da3e0d719f478aa59d851f8ec7a25d1cfce1149adef355e795a7b38f0

Observation 5c98fb3c-8520-4226-b3ec-997ec28c0470 · inbound

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents cites this paper.

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.117102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T09:27:48.400371Z digest=sha256:2fdaeb136859f25bd73d1449166221ae35309badfca3f6b92ecbc0037f1eb153

Observation 2e2cdb51-dc98-4b39-ac21-03865044b9a3 · inbound

A Mechanistic View of Authority Hierarchy in LLM Sycophancy cites this paper.

A Mechanistic View of Authority Hierarchy in LLM Sycophancy Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.800444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-02T13:41:18.279864Z digest=sha256:b4fbe56cebff34148dc123220330b181d91b01f3c87f35292e41e0b5fb20eee3