Pith. sign in

Paper Citation Record · LEDGER

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

As of 17 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2501.06248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06248 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:45.734231Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63b0bb0d-2444-4081-ac61-41292b01466b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.586146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.586146Z digest=sha256:da35cd14dc0d222a1f0905116cc1d1a18f49dc9756cc61c1dcf4596af2246050

Observation 69553556-f3f6-4317-9aeb-0f9ea8d6161c · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.591475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.591475Z digest=sha256:69a6cf9612a9b5ec533a2761766d52f2e5e02e33c70f7aed735c4f24f33472d9

Observation 056526bc-d7af-4d73-a5af-30a5c2f08331 · outbound

This paper cites Language Models are Few-Shot Learners.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.596208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.596208Z digest=sha256:1554721bc490447d7ae960a9fc65236e180f4eafde5b78ed12d8834966a7f100

Observation e3b3b00e-561a-4824-ad5b-31f4003490f7 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.602059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.602059Z digest=sha256:cb7e62ff08a4ea7693c65f1229880ef6d7b218b37343d54937442e7e6e1a7050

Observation 817eda52-d0f5-431c-bd61-25f2726e6f93 · outbound

This paper cites Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.607085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.607085Z digest=sha256:e3f74a4941e5bd0887d9712a718e7daf2784cf850ec2ed38318439f0f6bc09d6

Observation 3abef2ae-4105-4fe2-bd5a-372bbf6d1b1c · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.612012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.612012Z digest=sha256:68e3595ad930f28762db591ff4530375e5899b990344e05cef1b8a89000ee5c4

Observation 0ef1db20-cc05-4cec-9123-ba7a076ec3f8 · outbound

This paper cites Attention Flows are Shapley Value Explanations.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Attention Flows are Shapley Value Explanations

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:46.036410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.617402Z digest=sha256:a59cd1ba7cf6cdd6bd15d222c00e09df8937efd319c4a33a677a9f6bc1f37a07

Observation d43c383f-e8ab-48e3-9731-411802b1b3a6 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Axioms for AI Alignment from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.621875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.621875Z digest=sha256:1aaed3e0bf8f0120c23bba95d18645c2bcfe6c5e845fa5419d68eaf046a93c1b

Observation 1c32782a-3e91-42a1-9c00-e763f65a879d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.626337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.626337Z digest=sha256:5a53a3f6acdac6a2ff9e4b98ed69470c7256a9ab704ab6797ca8ce0b0bd9c698

Observation 12272422-9d9d-4b21-a3ac-c2ff4ce161fb · outbound

This paper cites Steering Language Models with Game-Theoretic Solvers.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Steering Language Models with Game-Theoretic Solvers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.630739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.630739Z digest=sha256:e8ebc9caa49dcc361a58c2953b7bbc4adecfbef07b533f652678b0fc332dd47f

Observation 24db5438-2529-4b7c-8f0c-ccc38c6f2f56 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Improving alignment of dialogue agents via targeted human judgements

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.635195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.635195Z digest=sha256:73ae7659286696e24d0a8bda0d23677021fa02a529dfdc8bffaba6caa237f2c3

Observation 6080f8c3-ea81-4161-ad10-30f11ebf0a5e · outbound

This paper cites The Consensus Game: Language Model Generation via Equilibrium Search.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models The Consensus Game: Language Model Generation via Equilibrium Search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.639950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.639950Z digest=sha256:61cb4612644e60b0c3e4e74c94d832a48bf3550c6bcb663b581d063860e46924

Observation c3d976a8-3a21-4061-9a45-7ad1402a7f05 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.242306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.644489Z digest=sha256:062f162cbcb8849c8bb1c5810d163a5385db8705988b062cbfa4b429840de843

Observation 23e0c1d8-1691-480e-baa4-8a0aa3aa6c1a · outbound

This paper cites Confronting Reward Model Overoptimization with Constrained RLHF.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.648887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.648887Z digest=sha256:1957af49af1e98f8e5afe3b3d8ea9b6cedc6aae627bb9d595001f5a973f473a2

Observation 4cfdc89b-b94b-4f18-8c90-3004031e06b8 · outbound

This paper cites Nash Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Nash Learning from Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.654182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.654182Z digest=sha256:c34e62030d363727ed87f257fa944d141074a4c0ef2c514da7812ec90eb4ddbb

Observation 1b4998d3-5525-4a81-abfd-67e7a4ba4e99 · outbound

This paper cites Learning Social Welfare Functions.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Learning Social Welfare Functions

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:45.906234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.658803Z digest=sha256:33994a4e7d4d94d656d8d22af6bae5087c150188372fe77c6472cba6ee760aa1

Observation 847d409e-a360-4a39-83a3-f8fdf8999ddb · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.227446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.663300Z digest=sha256:4dbfd3e91c7bbe735fadb6e1c69a292ca7b0742f6eda69a4a708b85a2e60806c

Observation 771fc87a-285d-4986-bfda-c994e177d059 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.212247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.667488Z digest=sha256:290e65ae5373ac1add0626052239e18d59d545877b6fcd2f1ff72aa9f4ef4545

Observation b1bd732e-37d6-44c9-aa74-bade1ffe93d1 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.197705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.672008Z digest=sha256:7d032ab565b7e1fc19372f635e06206df4b82b890b935bb486930dcbda246b6c

Observation ddb1ab5c-24e8-42e3-bc45-a5ee8f733e34 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.183274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.676572Z digest=sha256:262d31d1d0526e0e8da2634838b0ca7a7cedbd9d6128c33aa92518dc540d14a5

Observation 65e807e5-2b58-4273-b5a1-8043d6d25af1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.681223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.681223Z digest=sha256:c90e2ff68f6e2a5057a60849b08c76df2928aad9b64b36923750e925067b330c

Observation 571b9802-8a55-474e-81c0-c2e5a4ec0f20 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.685491Z digest=sha256:66bcbd12bc22d4df15a00ee3c5f2b39582cf570c586c829573f9962f78029698

Observation b86165bb-7e5f-458a-a718-31a4d8330e34 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.690199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.690199Z digest=sha256:a84b7b3a478f9e6a2265a98a72b01279c5f9d046e39858b7e54b2dd274a630d0

Observation e699ef17-cf16-4f60-a1e9-a2a258580c43 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.694507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.694507Z digest=sha256:efa8179ff184fcd5f7c4ba06f17598bac9470784e9b44bae6eb4cc631ede53aa

Observation aab78833-519c-4606-8f87-c64c2f219b4b · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.698903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.698903Z digest=sha256:21b4116c30e9e5ac874fb36f47313817688a1a58e4fe9287b3b223443a53a7a8

Observation e46d02ab-af9a-4a41-91e5-acf244c16eeb · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.158280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.703389Z digest=sha256:4592219086cd7025875ae2ec103209e1c2f8d3ecd6746d5bb5f4440f122c080c

Observation 97accfb3-d797-48ae-9bd5-3cb04c4287d0 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.707647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.707647Z digest=sha256:0755d260f3c8348d6afe5b4f3450dee0f3eafdec5e181a3cdbf8a08fb600462c

Observation a80e9000-c40a-4121-964b-61f1b9dc0c67 · outbound

This paper cites Transforming and Combining Rewards for Aligning Large Language Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.712536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.712536Z digest=sha256:a428176712fe08002348db8b8bb69d30718aef1eb549b32683dbda02793b4841

Observation a26d0694-f2fd-4af0-b501-231962d711fe · outbound

This paper cites Ethical and social risks of harm from Language Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Ethical and social risks of harm from Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.716752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.716752Z digest=sha256:62ce90a83878aedbad285c1241caea4dad727006b94d62fb28c8ba512f020980

Observation 9c7c0d23-7c6c-47ca-9aef-7828554af0ff · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.721017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.721017Z digest=sha256:14c9099e6c2639886e49903662a459cf3485b0591088935e982fdebac812dfca

Observation 146e3566-1b63-4247-bd29-bb6993028f04 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Fine-Grained Human Feedback Gives Better Rewards for Language Model Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.725087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.725087Z digest=sha256:81cd3f4944aefd32e5461a709038691300c1d576badc2cd788aca627a6a17bd9

Observation 7a496f50-7965-4c1a-b2dc-cddd7aa76bc5 · outbound

This paper cites online" 'onlinestring :=.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.729401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.729401Z digest=sha256:837d5a50c3a76a741749a77d774cc20c700ffe739ce03b625eae7e51c3b38a3d

Observation 6b515aae-a9cf-4cbc-919a-138da253c004 · outbound

This paper cites write newline.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.734231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.734231Z digest=sha256:221602f19eb19ecfc031938d2929d4ad3c5c50d7a4a844a6548b4cfaebce9408

Pith citing papers

No inbound Pith citation observations are available.