Pith. sign in

Paper Citation Record · LEDGER

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

As of 17 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 5 inbound Pith citation observations for arXiv:2501.02790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02790 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:07:54.567013Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:20.189439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.511368Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52ba7fba-dc28-4045-8b69-1abab7b9054e · outbound

This paper cites Adding hops early in the boil (usually within the first 15 minutes) primarily contributes to the beer’s bitterness.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Adding hops early in the boil (usually within the first 15 minutes) primarily contributes to the beer’s bitterness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.866830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.489403Z digest=sha256:7f33bde774dc2d3d0961c5f7dcf113bf5d6d2e31c26df2a0949d34a09e07b2ae

Observation a4808866-602a-4262-885f-8183dba9a357 · outbound

This paper cites The bitterness level is moderate, and the hop flavors and some aromatic compounds are preserved better than in the early boil, thanks to the shorter exposure time.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model The bitterness level is moderate, and the hop flavors and some aromatic compounds are preserved better than in the early boil, thanks to the shorter exposure time

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.849886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.495412Z digest=sha256:b77dda2e0c96eca72f354c02d8ca9c8169cb35a98427a33cf8a3fee415c350d4

Observation f8a7153d-d537-44ad-80cd-1803f0f73d12 · outbound

This paper cites This is because the shorter boiling time allows the volatile aromatic compounds to remain intact, while the alpha acids responsible for bitterness are less extracted.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model This is because the shorter boiling time allows the volatile aromatic compounds to remain intact, while the alpha acids responsible for bitterness are less extracted

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.832782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.502124Z digest=sha256:238912ddfb5402f9b2667dde04d1bd2fcd3e48b0833acde1f8c973da0ebaa88e

Observation c53aa61d-2c2f-467d-aabd-d7ecc2e0ebb0 · outbound

This paper cites This is done after the primary fermentation has completed, and the beer is transferred to a secondary fermenter or directly to the bottle/keg.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model This is done after the primary fermentation has completed, and the beer is transferred to a secondary fermenter or directly to the bottle/keg

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.815952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.507540Z digest=sha256:54c03aa0771ff1f216f4aa4fe518c661ebc01ccfe12f5384310bd570329445c2

Observation c11a862a-8a5f-4fc8-b73e-49575bc40e86 · outbound

This paper cites Spinach and Feta Egg Muffins.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Spinach and Feta Egg Muffins

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.798773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.513187Z digest=sha256:fc750e383f2eb01c0d101b4ae4efad7724eb0ac7d5eebcbea22d95a7ab4af955

Observation e29cbc92-e1e1-44a7-bc90-a3cf0df62939 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.696961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.546752Z digest=sha256:131e3b84b949489cb9566f3c1a85f82c7bebb99db6c7e3835549631632ca187a

Observation ae4b4585-9915-4c02-88a2-e5cdf1012f28 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.782680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.518774Z digest=sha256:815ec2a9f1ef799cbae250116ac76de259689943a50aa70f5822f2d89c85164f

Observation c5d1c61f-8afa-473e-a930-88a2e9e7b443 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.764859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.524240Z digest=sha256:477922ce6c016be763cd8281881b5afcd63998a211397d5d788b311d0426e064

Observation d217ee92-7bcc-44e6-909b-8d5c96c465e2 · outbound

This paper cites Stir until all the ingredients are evenly distributed.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Stir until all the ingredients are evenly distributed

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.747472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.529974Z digest=sha256:8b535ea711e74e2828344547ef31a613292662e896c490b3642f366e2b8e5143

Observation a7003008-5cf3-4e7e-8df5-2692f73807c1 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.730212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.535681Z digest=sha256:ab96111d9e715c89170ba39455a015483ca3cea880fdbcd25f380a358ebefdf1

Observation d86e7d58-6d46-4edc-a883-1ca0aef06910 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.713758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.541559Z digest=sha256:d3f27001d1f7c6bcad2df23ddcaab530aa0178e9fef8d02a0dd8d9ba2d7d3147

Observation 82ee1389-56bd-40b2-aa7b-54c17ac78e0c · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.680867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.551598Z digest=sha256:5baa77b350eaeeb9f8009ceb2e149c02eb20eac117b929ddc6f6e42b6a99ebf2

Observation 351e48cd-4793-4374-84c4-bd526a8923d9 · outbound

This paper cites an unresolved cited work.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:07:54.663422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.556760Z digest=sha256:3fa6c0b28da208d487a0b4cc8da449bca5ac511a38b88f6e0ef26cdfea5f593c

Observation 71260a24-a3bf-4002-af9f-1297cf6609bd · outbound

This paper cites The roots of the equation are: {roots[0]} and {roots[1]}.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model The roots of the equation are: {roots[0]} and {roots[1]}

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.646193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.561505Z digest=sha256:0db6127cf8b53c18fc686691adeda1264428e252e1f2e0d16546df83a300ae8c

Observation e8176557-d097-4005-b745-46e67f6e0387 · outbound

This paper cites Global Statistics of All.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Global Statistics of All

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:07:54.628064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:07:54.567013Z digest=sha256:46b334e5632d8db7340cb2bb444d0fd6a388fde8dc5cea36c23544a3640a6a49

Observation 2eee8d5b-c2e1-402d-8d55-a6ea8292b7b7 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:07:54.481652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:07:54.481652Z digest=sha256:83c4164578900700ccfd4953473607d2045517db85c32a9ba4939ee83478bded

Pith citing papers

Observation c7e0e76c-dfb0-42dd-a0ac-40850cb5cca2 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:31.014205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:31.014205Z digest=sha256:e5b496c0d5f3f6425fd50b798c801a483acd5aa596cdcec46928befb9b52f004

Observation 87939929-964a-4e5a-88ea-1058304cd6c7 · inbound

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization cites this paper.

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:20.189439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:20.189439Z digest=sha256:8a8a9c58bb56ee28c3ab065a6b5568b5d8cdca430835e2979e9c708929db1bf4

Observation b5923323-4fcb-4b42-9458-093c65a4a7df · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.401522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3a726c38b8824255f956b551d68537ac3f103c8f51e0e3e370493b3dfad3149e

Observation 35917a05-ad3b-432d-881f-e8f9db0b75b6 · inbound

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment cites this paper.

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:18.030277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T23:00:14.031709Z digest=sha256:ec75761d423440722de982eaa67e041de497ea3b47bc6ad14ca13b4a2185bbc8

Observation 26de573c-73ff-40fa-a7b8-64a60a72c502 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.512754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:16ed4e4781d97b038f2ab40fe91a54cc347cd2a0a2b54d0e20ad1f6ff334707f