Pith. sign in

Paper Citation Record · LEDGER

ODIN: Disentangled Reward Mitigates Hacking in RLHF

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.07319.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07319 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:34:56.907788Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:46.669571Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2d964af-fd96-4424-8667-bbc72aba3a2b · inbound

On Teacher Hacking in Language Model Distillation cites this paper.

On Teacher Hacking in Language Model Distillation ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-09T11:34:56.907788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:34:56.907788Z digest=sha256:4f9b92438c19939f0f768c232a1bc4ac3e4d4a1d4db925c97d67b4007fd6a22f

Observation f595ec0a-4750-44c9-afa0-fc88a87ae2a9 · inbound

Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving cites this paper.

Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-08T12:07:04.976745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:07:04.976745Z digest=sha256:a4031af0e2d79a7bf152ae9736daad65bbf0dc8e498a6e9643fc244e32e1f20e

Observation c7e10994-40b2-4706-903f-2421b967a51a · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:17.249774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:17.249774Z digest=sha256:453da80940c8bf6a55533d5981b76cd5cf43e9a75cb0f0064d1952c5d1bc3b93

Observation 45abfca5-0ae6-49c3-943e-6ac160b9dfdd · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.979266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:791d9fdebbe78b842bd8b2c5d5c397dd4b519ee8a8dbd6087ab178b3dc28b90e

Observation 2e00a555-b139-4013-8ce2-8655891135f1 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.862282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:323d3d4a511344fbd09f589e06f68048894897b052590c41ad1edbb997dd48da

Observation 163fec65-6cc6-462f-b789-08baac871e75 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:25:45.371973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:bf8936659f65461f207c290ee28c91994c7b6981da16179da47b98d37abeaf4e

Observation 674e5b0e-29ed-44e1-8a23-8126f4dbeea9 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.711581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:a32f1a4a3b2ec1146955361b88f44e92a5b52415a51fd9929b2ef23db5284aac

Observation 29b958c3-ab11-43ba-bbbc-8b7b9dceb1b5 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.555672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:2a42e3c2659c2c00fc1ca4bf091fdcee1f1facded09cb82e715568ae1888e0c6

Observation 0ae14f58-bc00-4810-96d7-48ceae5e544f · inbound

RVPO: Risk-Sensitive Alignment via Variance Regularization cites this paper.

RVPO: Risk-Sensitive Alignment via Variance Regularization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:08.473849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T15:00:11.237293Z digest=sha256:7596a7641f9d13a4ebd1631114f81ca01bf47b285bbd7e7ab48d87cce5a08db6

Observation e6bc8d08-6f0e-4077-aa87-a0c5207aa80b · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.073790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:1f697d13b56b8459649be94c3a5ff5c8e216dfe0e94b44f1d28565e34e2c7d7b

Observation ec71211d-e745-4e88-9e62-9c33004311fc · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.527008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:db3e494b2248ec26b002d69ef0cdb8be5c56ec4827303c0b810cdb240b04e901

Observation 5eca7ad6-eb1a-4f45-9ac3-6ff0ab240b1b · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.735216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:4905b5ee98063e96b9a9b0d950b452efccee86932e3f68d0e087359c60199135

Observation e9f9d502-d305-4086-9927-bedb017f52d8 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.859083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:e5fe0f4039d9b69af658fbd672dff2408baf3187709f60c4da40ad5f1d50e78c

Observation 184967ed-9231-46dc-8553-08bb083bfc90 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.746637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:aa2ca6dcf1d94a3cfb07f9538210566bbade7be25f275f4ec1784faae8383f0b

Observation 86b8fada-495c-4ee4-8766-514451a80c4e · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.505006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:580d1e4c6d1264d4a27c1cacd718da9bfca4db9588fb265ca6851dcefd8cd52f

Observation 16d00a7b-976c-4d46-909b-80e40d11902d · inbound

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling cites this paper.

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:46.671372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T17:56:48.024298Z digest=sha256:52cd749e39b5ab0c8b669f68a88684fba46ff236d934b34e34ec0ff9ab902d9b

Observation e09f281b-e0e9-4d96-8bb5-3d7d060122fa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:4c5e57d2086efc49e8bb9e8e8d2d258eed55a58cc285aa0b90ec032f688a77c4

Observation 768ec640-82d1-4148-9f71-238efeaac97c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.315887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.315887Z digest=sha256:2b14ac478b827d8f88f936ca11af40232b39c7497120197bf21abef555e5f736

Observation de695394-6551-43ea-8b6d-e119fa33956d · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:57.056053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:57.056053Z digest=sha256:967cec17d22687f7a834355085835bcffbae3a25c272701dfdc3cb1e6ddcee9c