Pith. sign in

Paper Citation Record · LEDGER

Transforming and Combining Rewards for Aligning Large Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.00742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00742 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:52.776300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:07:47.876920Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3fffca85-55af-4cda-a00a-bed68e84db18 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.399932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:8a103069375e6b9690000a2afb1cedad86c9e891e325efc49bbd4d5f87cbf5d5

Observation 7f4a72c2-4579-406e-9786-47ad96cedb5e · inbound

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning cites this paper.

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning Transforming and Combining Rewards for Aligning Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:43.697884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:43.697884Z digest=sha256:c9bf09b811f8a061313f6321f9c39dc5f936343d94657d255636e22d98b0ae8b

Observation a80e9000-c40a-4121-964b-61f1b9dc0c67 · inbound

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models cites this paper.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.712536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.712536Z digest=sha256:a428176712fe08002348db8b8bb69d30718aef1eb549b32683dbda02793b4841

Observation ab6a48ed-00b9-4874-a99c-bc7856e6a796 · inbound

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation cites this paper.

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation Transforming and Combining Rewards for Aligning Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:04:57.363613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:04:57.363613Z digest=sha256:d0fde6a35d4e7e9d96ae8f5752a84c2b31a82cf748379c9b31463f92e48b1b55

Observation 4909cd43-a8e3-4144-8f5e-775542ef432e · inbound

Language Model Networks: Supervision-Efficient Learning through Dense Communication cites this paper.

Language Model Networks: Supervision-Efficient Learning through Dense Communication Transforming and Combining Rewards for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:51:41.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T14:50:45.735917Z digest=sha256:1bbd063d57d72f2f74aa55ec00a78ec309586b5f74c3a05b5fe7d2ef71c1eacc

Observation 984a176d-4cc5-4aad-99dc-7fb317dde5bb · inbound

Language Model Networks: Supervision-Efficient Learning through Dense Communication cites this paper.

Language Model Networks: Supervision-Efficient Learning through Dense Communication Transforming and Combining Rewards for Aligning Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:52.776300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:52.776300Z digest=sha256:3df06d8f0d79b15f25f7e6a6178a5dfcb1fc0fabd8912333ca0fd85d7cb2d47e

Observation c7c2a000-22cd-47a4-8c5e-76e767e818af · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.839256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.839256Z digest=sha256:8e31c8060107084f4b5817ac63d1a11b8cb7a2cd5c1cf611223ed5b3c7b02825

Observation 1378a620-b05f-49a2-b875-d8cf129227b0 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Transforming and Combining Rewards for Aligning Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.003533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.003533Z digest=sha256:d304a43449279dd7c8d928856b04a98531eed18c4da81c9e20ff9efaa056582f

Observation cf7406b8-da57-48b7-a5b5-fb213d6d1f82 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Transforming and Combining Rewards for Aligning Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:12:14.083693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:b93a4b9099875e469a071dbad9080cb3fa2a4d34547bc1c54cf530b94e46e334

Observation aec05c2e-5af9-45a5-9870-bee120dd3755 · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Transforming and Combining Rewards for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.878366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:df9e68e64800a2c65c542c88b542e87fd3f212e20cd706365670b219e48b3bd0