Pith. sign in

Paper Citation Record · LEDGER

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

As of 17 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 5 inbound Pith citation observations for arXiv:2501.02790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02790 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:07:54.567013Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.189439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.511368Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52ba7fba-dc28-4045-8b69-1abab7b9054e · outbound

This paper cites Adding hops early in the boil (usually within the first 15 minutes) primarily contributes to the beer’s bitterness.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Adding hops early in the boil (usually within the first 15 minutes) primarily contributes to the beer’s bitterness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.866830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.489403Z digest=sha256:a1f291e7b1fac0c07992447b5d8c3447843e6df3659a61fd82f042cb64a7b1e6

Observation a4808866-602a-4262-885f-8183dba9a357 · outbound

This paper cites The bitterness level is moderate, and the hop flavors and some aromatic compounds are preserved better than in the early boil, thanks to the shorter exposure time.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model The bitterness level is moderate, and the hop flavors and some aromatic compounds are preserved better than in the early boil, thanks to the shorter exposure time

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.849886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.495412Z digest=sha256:786d7ecc398fcaa02ac4d112514b9311524a8be925c8de2ddd3ecb1a57b8050e

Observation f8a7153d-d537-44ad-80cd-1803f0f73d12 · outbound

This paper cites This is because the shorter boiling time allows the volatile aromatic compounds to remain intact, while the alpha acids responsible for bitterness are less extracted.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model This is because the shorter boiling time allows the volatile aromatic compounds to remain intact, while the alpha acids responsible for bitterness are less extracted

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.832782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.502124Z digest=sha256:497bbbf6f248886bf02007c5dc40e7d4343ec9916a318d5e63e5ddcb113859c6

Observation c53aa61d-2c2f-467d-aabd-d7ecc2e0ebb0 · outbound

This paper cites This is done after the primary fermentation has completed, and the beer is transferred to a secondary fermenter or directly to the bottle/keg.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model This is done after the primary fermentation has completed, and the beer is transferred to a secondary fermenter or directly to the bottle/keg

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.815952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.507540Z digest=sha256:d53b1415387fe525fd72cbb7186d4e3f1f4b59fda5860abe09103df4f36b5c98

Observation c11a862a-8a5f-4fc8-b73e-49575bc40e86 · outbound

This paper cites Spinach and Feta Egg Muffins.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Spinach and Feta Egg Muffins

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.798773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.513187Z digest=sha256:d7e697f63bdb0c1d7717b6bb9cabd850378860fc53a5ad4960a1e5a5fa5edd08

Observation e29cbc92-e1e1-44a7-bc90-a3cf0df62939 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.696961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.546752Z digest=sha256:8dc2813eb0fb0214dda6daf7b0b9e86bd0f89845e9ab59c3b468619c4c8e01d0

Observation ae4b4585-9915-4c02-88a2-e5cdf1012f28 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.782680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.518774Z digest=sha256:beb1072deda9e637f9128ca2d879c6dee1e999467f11c74bc5736084162b80f2

Observation c5d1c61f-8afa-473e-a930-88a2e9e7b443 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.764859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.524240Z digest=sha256:46b67d1321333c4ea3e39dd7892ed623a5f7765180cd0024471491dea4da9644

Observation d217ee92-7bcc-44e6-909b-8d5c96c465e2 · outbound

This paper cites Stir until all the ingredients are evenly distributed.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Stir until all the ingredients are evenly distributed

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.747472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.529974Z digest=sha256:19e2d192cb52f75d69622c40192444276df44e3bfdee61e99e18fdf74e13adaa

Observation a7003008-5cf3-4e7e-8df5-2692f73807c1 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.730212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.535681Z digest=sha256:f505d53c72cb7265dc99bbf8f146df2dac088015c0bfee02b9e09fdeac3baf68

Observation d86e7d58-6d46-4edc-a883-1ca0aef06910 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.713758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.541559Z digest=sha256:29f733489ab108a02707860fc2b2d4e1fa6de33f22dc760822dce4b69426262d

Observation 82ee1389-56bd-40b2-aa7b-54c17ac78e0c · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.680867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.551598Z digest=sha256:aa9857b627a1e17a2fd87087a7b0991a3d49553cc3d11b4edb0857e4e9db426c

Observation 351e48cd-4793-4374-84c4-bd526a8923d9 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.663422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.556760Z digest=sha256:53e31db16975d00c76a6ad11d179c527492d0b8c8c9a8f559d652f7b99d2bab9

Observation 71260a24-a3bf-4002-af9f-1297cf6609bd · outbound

This paper cites The roots of the equation are: {roots[0]} and {roots[1]}.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model The roots of the equation are: {roots[0]} and {roots[1]}

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.646193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.561505Z digest=sha256:d6be4844cada3d672c0c47068310d6ddcb42f0f31eab62312f4cc339d7c2cede

Observation e8176557-d097-4005-b745-46e67f6e0387 · outbound

This paper cites Global Statistics of All.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Global Statistics of All

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.628064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:07:54.567013Z digest=sha256:628e28753719bc6769d88dcc3b8ed7b5663e3e7e592c02cb6b72d422b379810d

Observation 2eee8d5b-c2e1-402d-8d55-a6ea8292b7b7 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:07:54.481652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:07:54.481652Z digest=sha256:83c4164578900700ccfd4953473607d2045517db85c32a9ba4939ee83478bded

Pith citing papers

Observation c7e0e76c-dfb0-42dd-a0ac-40850cb5cca2 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:31.014205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:31.014205Z digest=sha256:e5b496c0d5f3f6425fd50b798c801a483acd5aa596cdcec46928befb9b52f004

Observation 87939929-964a-4e5a-88ea-1058304cd6c7 · inbound

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization cites this paper.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.189439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.189439Z digest=sha256:911104a60addbcef3db931769f990962259eec4171ef2a6e64f8a8698320ac23

Observation b5923323-4fcb-4b42-9458-093c65a4a7df · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.401522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c655c0414f448ae890039165ddb586a75b6ce5bc69eb2a8ea20069a01282598e

Observation 35917a05-ad3b-432d-881f-e8f9db0b75b6 · inbound

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment cites this paper.

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:18.030277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T23:00:14.031709Z digest=sha256:5d42a10077c64120c38676cda8864b5b1031465e3f5662d1394710b418256aa9

Observation 26de573c-73ff-40fa-a7b8-64a60a72c502 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.512754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:c211dc5f057209f6e9fd605b886de25bafccbecec4a6a4a9e319bf2c93e21aea