Pith. sign in

Paper Citation Record · LEDGER

OffsetBias: Leveraging Debiased Data for Tuning Evaluators

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2407.06551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06551 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:00.091046Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T17:35:43.819146Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b8114625-c371-4b9f-ba32-9776f8bf6193 · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.653875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:53bc6b4205cb6c43f88be8e77b6bcce1881f526a20619dd0c7d56569a0f46101

Observation 6e23e892-f64e-402d-9d1a-76d508738a51 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.821753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:d2fdafaed7749b1006c5d05658935d060506b447431fffb8d2cb8e6473c87a67

Observation eef2f67f-fc87-40da-a019-8cb8cd92a048 · inbound

Interpreting Language Reward Models via Contrastive Explanations cites this paper.

Interpreting Language Reward Models via Contrastive Explanations OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.883420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.883420Z digest=sha256:7590e0790add0aa90efd812a003d1fe6dba37fc82eaa3e377df25a074e07719a

Observation b809d533-f053-4a48-bf02-b40e061d3cc8 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.188993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.188993Z digest=sha256:e064a1749bb1776f8dea7585018f47b4752d4b9474e6ebdb8c5ffa09b96d39e9

Observation d10d3d48-12cd-4bb2-a8f7-0f6ab458808a · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.149064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:0aa36136d01b1f287f6b1600c2daa996771b26c1454f8b56af2dbfe459b2e147

Observation 54600f9e-9665-49c2-8dac-a1a8bb733340 · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.926599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.926599Z digest=sha256:e2a199ad99cf5ff07853b0bf77f4b78bd8e84c42e6992092308721b56444f4de

Observation 31398427-a194-46fc-9c48-0a2b2dd6c92b · inbound

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer cites this paper.

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:00.091046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:34:00.091046Z digest=sha256:a7d6dbc382e9ee68086158606e3eb684944dcc002ed474477987baa02dc2df26

Observation c8014eb0-9e0e-4c9d-b102-143e4b4e20af · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.491812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.491812Z digest=sha256:ffa3aa1fc3ba6fac7e7a6e598c865dff2fdd41326ca267171d26fad639cc8a1a

Observation ee68f132-4b5e-4f95-aeee-ddce04755a16 · inbound

Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking cites this paper.

Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:40:34.455462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:40:34.455462Z digest=sha256:d4eac12ed995ebb9da42330f42f19a2acd59aa83f854358f486d8dcec5c9c853

Observation c0978775-496a-4fb0-932b-a78846273445 · inbound

WorldPM: Scaling Human Preference Modeling cites this paper.

WorldPM: Scaling Human Preference Modeling OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:52.701060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:12:52.701060Z digest=sha256:b8d55842b31fd1cf1a3f85566c5e7725e2be888d5f24341a164a93ad771e8d25

Observation 9032f033-466d-4408-bd5f-3777967ffaeb · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.626884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.626884Z digest=sha256:c2d300f813d80f89f2650bea81ebf7648e372d25673ba3c6f2e0b8343f3a5aad

Observation 07db7b63-fc3d-4db5-8548-104e0b6028c1 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.601136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:6de5115dc6278718f6cbd3feef75c78295d529c2d175c692b4abc4814bc863ff

Observation 56114032-d189-4cae-8ef5-06c0cdd39060 · inbound

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows cites this paper.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.605360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.605360Z digest=sha256:5b1aaf8712d80f7f05b3f383beb3ad31fab666ee3a37ea2c901fe99ebd7ae77c

Observation a27bbd21-8242-4092-b6a8-dd026c14e7e7 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.819184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.819184Z digest=sha256:b014dfb90e39aae37afa461d5e735d337b42c4497cc4667a0dfd087346eab04b

Observation a6e2eded-492c-4857-b642-92f634971cc9 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 287

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.252751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.252751Z digest=sha256:5984ecb763c9ebfefcf9fbf711519a4e816b385c65a8409cb01822a8b47c8d60

Observation c125dec4-740c-4de0-b143-a7f6f052dd4f · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.029108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.029108Z digest=sha256:3718e53266fa6c73f27ce562e5d87998f8c84bd158b2edb94ec5fbd459648285

Observation 0a394f64-376d-46e7-b9ed-4fdaf1a71a9c · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.937026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:8ec094ee336f0f0381e6343cedfa74a1bd1d63b1d6092f524bf99f12eeadc712

Observation 0448b508-fee8-48e2-aa94-56507906638f · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.550549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:97400baa3a3ef911d5b9088054c0a736987f7a07a5b29d18aa38dd83a974a1bc

Observation 5ef0d219-294c-4580-92d3-33a2f1d989a2 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.046679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.046679Z digest=sha256:db37247bbad0bf85ee0b92399e2cadf10312cb420132e8aef2600124122054f3

Observation d2012570-b487-402f-af39-a05be86e5a6b · inbound

SafeScreen: A Safety-First Screening Framework for Personalized Video Retrieval for Vulnerable Users cites this paper.

SafeScreen: A Safety-First Screening Framework for Personalized Video Retrieval for Vulnerable Users OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:25:30.934598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:23:43.022505Z digest=sha256:4efac63c5196e63a40154ebab01b4dc50005af36443f31cd249311e263fef8ac

Observation f60d5783-df58-4dea-be63-361f7b5252de · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.032476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:d1939f09be5c542cff7a69ae6040ae583abce99230c0e63fe9d9eb437f473c53

Observation e7a42db8-9698-428b-85c8-67fa49688360 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.061387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:60ad88990f3fe57868e397a5c83346d0d92bcc54a5f1f831483f835ac22939cc