Pith. sign in

Paper Citation Record · LEDGER

Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2204.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.06601 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:06.259098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.237337Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc961c3e-c66e-430a-b4e3-ddb6d8c0a42c · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:46:56.819389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:d2aa1d7602a920ea08f4aea8d5487c2a8ca5caf7e135f5e173914b7c82d42777

Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.259098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.259098Z digest=sha256:c0a6c88ea75db224a42eb4e183fb8a8cfb1439db7762b2705d0575b048da62b9

Observation 3a528559-9b07-4aa3-8710-8ad63f060415 · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:17.017683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:17.017683Z digest=sha256:286629993114a21ec25ee9006df8e2ffe03d4554559f53340a7297e49cdcfd6c

Observation 8873c3f5-2282-4a05-aae8-6661ff942edf · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.046448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.046448Z digest=sha256:28cc6b9b01e6e2240ff6bb0c20eb229ab7f8b3f0ca6d28d56d7950965463655a

Observation c39d7299-8903-4d26-857d-5f897ba91aaa · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:de10a06f7d998ea9713633fb076930257dd995a7bb5cfc50f724987037c70216

Observation d30fb4d1-e33d-41f5-81fb-75d86b60e1bd · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:24.068249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:04792b6c7f5c764b2827317c1dc09a755bf840ff233730d5c96ca5c65ddf1193

Observation eba63cda-02e1-4808-a11d-e0a19d880f55 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.193120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:0979944942095ac2a6dd80886256df5de7243e601302a890f63a4a377315ae5f

Observation 09df6ec3-a0e2-473b-add8-634482ab72bb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 173

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:ebfec2c53c9dfc5c855176f35f9ccb48435ee78134c563684c88739144e0a409

Observation a339ab8c-88e3-443c-9121-9a77aef70368 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.349555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.349555Z digest=sha256:451a67a73ec5e1be8436ee6bf38199479ea1a5b96fba2b0ce85a966b05b6f8a0