Pith. sign in

Paper Citation Record · LEDGER

Explaining and Preventing Alignment Collapse in Iterative RLHF

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2605.04266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.04266 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:37:18.229897Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:26:02.889788Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T11:26:03.433332Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8981857-b564-41f4-a54d-d462b092b225 · outbound

This paper cites P´ asztor et al.

Explaining and Preventing Alignment Collapse in Iterative RLHF P´ asztor et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.559912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:9d2f320df0c453311f2fb260729ce273b390915de0ac41e472962a8a80e7ea06

Observation 7a66c6d1-6f9e-4504-8c9a-b6038792b215 · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.571097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:c5e0b63322f77f5eb37e865dbdcc824d3d0749efd5815727a4d088084a3301f2

Observation 6255ff2c-3e68-433d-b930-72492ce0f157 · outbound

This paper cites Figure 3 demonstrates that standard RLHF policies increasingly exploit noise dimensions, achieving high proxy rewards but abandoning true utility.

Explaining and Preventing Alignment Collapse in Iterative RLHF Figure 3 demonstrates that standard RLHF policies increasingly exploit noise dimensions, achieving high proxy rewards but abandoning true utility

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.553283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:926f2fc0c01332af6658696c2a4ac36414650b956c10a2130d842f1fa83f21be

Observation 1508fd0d-ec0a-478c-b3b7-2e3297ea823e · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.556353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:e2d4d87f8ff56b2fbeba2a973a0dc3c9d4c4480dc2f895b70d4aff9fa2a9ca01

Observation 5f0b6385-4c40-4b6c-9866-bc701ba4625a · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.598294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:bc096f74abcd37fb55108d2a1b1b23f2311d56c57295659782e66484cfb7ec47

Observation 5e054673-8b96-4f47-871e-45dc355466fa · outbound

This paper cites The relaxed penalty multiplies this inner product by the overconfidence proxy σ(rϕ(x, yi) −r ϕ(x, y′)) −σ (U(x, yi) −U (x, y′)), where σ is the logistic sigmoid (cf.

Explaining and Preventing Alignment Collapse in Iterative RLHF The relaxed penalty multiplies this inner product by the overconfidence proxy σ(rϕ(x, yi) −r ϕ(x, y′)) −σ (U(x, yi) −U (x, y′)), where σ is the logistic sigmoid (cf

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.584703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:6dd59104468c40b989f08b629b476bf7ed9adacb44389db1f8a34709d1d8867a

Observation 15032004-61a2-4ff3-be2b-5aa50e2b3eca · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.594988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:d8fc754ae59f128a4eac75fcb6d1dec864fb92a6cf1159bbd2590becb825c428

Observation 2aa457e6-936f-40b4-b569-bf4263836e0e · outbound

This paper cites Both updates use the 8-bit paged AdamW optimizer [Dettmers et al., 2023, Loshchilov and Hutter, 2019] with gradients accumulated over four steps before each optimizer step.

Explaining and Preventing Alignment Collapse in Iterative RLHF Both updates use the 8-bit paged AdamW optimizer [Dettmers et al., 2023, Loshchilov and Hutter, 2019] with gradients accumulated over four steps before each optimizer step

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.548930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:2c68ccdbe71ef1c1320ce7bb3258aee1e89fa6f4943e35bb0901b4991f5faae5

Observation 47ae9d58-1425-470e-85a4-d99848e67f82 · outbound

This paper cites Punish models that confidently state false info.

Explaining and Preventing Alignment Collapse in Iterative RLHF Punish models that confidently state false info

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.588269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:0d74254e3c7d3b7470c88a73114fb5ec6c4ffd1a4f7f61e5c8ab466133d10d0f

Observation 5ccb1f38-95c8-4839-b8b2-4926e19552f9 · outbound

This paper cites reasoning.

Explaining and Preventing Alignment Collapse in Iterative RLHF reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.577637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:16d7da26e45100cfd0b9aef77b2f149ddf4720dfa82a2ca497bb88058a42f257

Observation 492eaf3c-e480-44c5-9dc8-2d46961e4c97 · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.591621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:70b039117d84e570121df8a8ad8ae700b7506aa3fdcd1f2182221cf76206ba1c

Observation fa317903-27eb-4b84-9eb6-adaac08e533d · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.574459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:e2298529501aec61cdbcca5ce178ceb4114bef7ecd3a010992b5177db57b9511

Observation f8aa3efc-dc27-4243-8757-7331765b4e8e · outbound

This paper cites spheri- cal.

Explaining and Preventing Alignment Collapse in Iterative RLHF spheri- cal

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.581112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:c3fc1ba32105fc0713f29bcc8ec373b548ba5b948ac704f410a1ca633b063461

Observation db703653-8228-47b4-8050-be4497ff7228 · outbound

This paper cites Its training objective is the pointwise prediction error:ℓ train(z;ϕ) =ℓ(z;ϕ).

Explaining and Preventing Alignment Collapse in Iterative RLHF Its training objective is the pointwise prediction error:ℓ train(z;ϕ) =ℓ(z;ϕ)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.567753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:1b439ad184264935fd29f950b08235a1be9be17a73f0b1d2909cde542e675f00

Observation 83ed3f0e-95ac-4523-bc44-85c0db65b96e · outbound

This paper cites To frame reward maximization as a test loss to be minimized, we define the Leader’s objective as ℓtest(z; ϕ) = −rϕ(z).

Explaining and Preventing Alignment Collapse in Iterative RLHF To frame reward maximization as a test loss to be minimized, we define the Leader’s objective as ℓtest(z; ϕ) = −rϕ(z)

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-05-26T04:27:23.563431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:7ab405abdc46e583dfd76d55c3ca2004fa2139c93cce51884a19b02d9064ffaf

Pith citing papers

Observation 81d5aaa8-a6b5-4379-9f87-0dd63c65de62 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:579cc4be391527283bdcefa42b8734aeeeb7d189de883fef876e14b6c7f86c63

Observation 444421fd-0641-4f09-ab05-109e196003cc · inbound

Efficient Hypergradient Descent for Inverse Reinforcement Learning cites this paper.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:26:03.437865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.889788Z digest=sha256:7e12d6696cb818343da5e8ca226d9366caa44d3f0b9bf0f4e9468b46ceeee8f6