Pith. sign in

Paper Citation Record · LEDGER

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.02951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02951 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:02:45.665708Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee422279-496a-4840-9e7d-bfb307e0ec00 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.558196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.558196Z digest=sha256:b2288cb3f5cfca215e098337ac68298f4154cdb510cbc3d1856472682f8e253c

Observation d290f816-01cf-4562-8b67-462bbce1b352 · outbound

This paper cites Contrastive Preference Learning: Learning from Human Feedback without RL.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.577850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.577850Z digest=sha256:84b935beef02fa7b56cd6e04f810d7c48902216dd045c00a2ecc282d6e0b0db4

Observation 98b5c4aa-31eb-42e3-a4ba-65e39627fcad · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.582408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.582408Z digest=sha256:ebb23b5b01ef73e2c3199c76223733c7f26aef2a7825c53a50c52be889cef39a

Observation 3a480b73-c493-4684-92e3-dccca44d5ce0 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.609451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.609451Z digest=sha256:d8e4f17a4c8eb64d25ddb352f26bc6688fda0d775841ce759fd3d6bb5f0aec99

Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · outbound

This paper cites Reinforcement Learning with Segment Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:02:45.752291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.617996Z digest=sha256:9ebfad3692c9cb963b572ebf26f8c5f0d5fd944054643408bfa2141f115562c6

Observation 605b043e-97f6-4fc4-89d8-ad25e4b7c7df · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.622442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.622442Z digest=sha256:0bd1b4747cd489e8d7017426b750431395da5287d16f7266c11a38bde9d2f5a8

Observation 08c4f306-f481-4970-b9c9-d8c977c836e5 · outbound

This paper cites Decoupled Weight Decay Regularization.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.631715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.631715Z digest=sha256:49958c836bbfc6d82d875650f605d3c82d580b010febc9c5f7fcc311120e3e3a

Observation e28fc222-6012-45fa-b84a-03663fb3c60f · outbound

This paper cites Approximating kl divergence, 2020.URL http://joschu.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Approximating kl divergence, 2020.URL http://joschu

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.087144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.636173Z digest=sha256:d40a44851d0488ed187c8bdb2c3e3d8e4cc24f30a57f38bbab316c0dcb09c2aa

Observation 4fb484bb-1eb7-4c64-a6d9-305513f72586 · outbound

This paper cites 15 A.2 Comparison of Preference Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling 15 A.2 Comparison of Preference Models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.071148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.641390Z digest=sha256:ec16a41cbe89adf90269561354a12f6144732fc9ffdba5b8dfeca331f3015bbe

Observation ba2a8584-11b2-422a-9c6f-84a761d8d5d0 · outbound

This paper cites This makes them inapplicable to many modern RL problems.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling This makes them inapplicable to many modern RL problems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.055188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.646471Z digest=sha256:cdce40987aeec8c41cecfa948fc23453d825170ec4c80b0d3d40144c9fe3ec4a

Observation 13ac6ae2-d619-4207-9ba0-d10702023097 · outbound

This paper cites Specifically, even if the partial returns of two segments are the same, 15 humans could still prefer one trajectory segment based on the value of the last state and action.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Specifically, even if the partial returns of two segments are the same, 15 humans could still prefer one trajectory segment based on the value of the last state and action

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.040674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.651001Z digest=sha256:5b5ebc2be221138ae59091643e8a1b52c9cb3d2e861e23b9c53c26ba12376a2d

Observation 156a5eb1-7f77-42e1-b4bb-c7055b8afd38 · outbound

This paper cites During training, policies were Gaussian with a parameter controlling the log standard deviation for each dimension in the action space.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling During training, policies were Gaussian with a parameter controlling the log standard deviation for each dimension in the action space

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.010880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.660952Z digest=sha256:544482b33c44911e1667140a180227d9c6b09b91a8b80a039b3c5c687afb96b4

Observation ff72a96f-415a-4d8e-9695-9a1ab3b7e032 · outbound

This paper cites The global gradient norm is clipped to1.0.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling The global gradient norm is clipped to1.0

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:45.996416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.665708Z digest=sha256:47feb602269389438c21ef79a9b9442018de3951c4690d6b24fed0d16a30d75c

Observation fe301495-77e8-4b26-9e33-1856d6f1db27 · outbound

This paper cites We observe that in Ant-v5 usingπθt´1 as the second policy out performs setting the second policy toπθref.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling We observe that in Ant-v5 usingπθt´1 as the second policy out performs setting the second policy toπθref

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:02:46.025703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:02:45.655771Z digest=sha256:41871af7428d5f3ccf1cd578bb8ba28e8caa27059bf1132baa3a2ed6ce28a1e4

Observation 34a6412f-ff93-4201-886c-6617e19fa7e0 · outbound

This paper cites Models of human preference for learning reward functions.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Models of human preference for learning reward functions

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.573168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.573168Z digest=sha256:ce90ec2434b02f47adfceb5c0a56c56fcaed4bb8db570fc1005620e4dacb9f44

Observation 14ed6ea2-811c-45d0-84b0-51b5553374de · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.591812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.591812Z digest=sha256:1da7fad89cb73d8ec699bc654c4d4c1ba9edb499b1f3e9fd9f5deacd7a299667

Observation 1ff964b8-3fa5-467a-8dfa-88ee91dd45af · outbound

This paper cites Proximal Policy Optimization Algorithms.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.548076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.548076Z digest=sha256:c000929c6d44fec8b6589c821e2ce2c9e71c65314aff70b4137926c325e4c588

Observation 52bf253f-7a60-4fdd-acb2-db84ff3810f2 · outbound

This paper cites Making Reinforcement Learning Work on Swimmer.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Making Reinforcement Learning Work on Swimmer

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.626921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.626921Z digest=sha256:c29a54c983983f706a4522648f70fdee0c708bdf6f3ecf1c1cc49591530d92e4

Observation 526c189c-caec-415d-a723-96ba1e20aceb · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.601173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.601173Z digest=sha256:589f9cf821e4bfe6654a6c3369c35572fc18fc46feeef1e09a6f5f44f562b489

Observation 179e76b8-8da8-4c1f-8fe0-de3fa777a8c2 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Aligning Text-to-Image Models using Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.567940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.567940Z digest=sha256:9e266abbf9443b904b0cb01f17460f026b696f4fab4ffb4897c717ef01555ad4

Observation b0d7fa27-dac7-474f-9532-0b7e02d9daa7 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Playing Atari with Deep Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.542769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.542769Z digest=sha256:5318bb2459540478125836ea135c04b169ae1d2b3474e99061eafb2ba8ed8bb6

Observation 38abb7ee-237b-4218-993a-661c3d4087de · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.613860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.613860Z digest=sha256:6fe958d824549c6dcce994ba4e1cf655166b081274c7d770fe5dbb32ca581508

Observation b85b8801-d6b7-4c14-a2b0-d595bc8b65a0 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.605223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.605223Z digest=sha256:62026dc2485933fe74844bb71626734b9ddb4b7b5a0d19397824121ddac7091e

Observation ea2ce2c8-9538-4056-8515-e1ea757041a9 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.552796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.552796Z digest=sha256:48b2fa2effa45e46f2c56c278e80ea45143262a3e7d0b69c0d9fef2e92c4d27d

Observation 227b8e35-4610-4640-bc74-c6d145f81fd4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.563041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.563041Z digest=sha256:6486fbb919f9200b4b3bda68b6d9a7911ae4b6462fc3c46215bda19ea867ef6c

Observation f7f7b12a-67c5-4df6-b868-14f03998d776 · outbound

This paper cites Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.587036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.587036Z digest=sha256:3b22f29cd807405629798294e39012c74c2ac46b8357ebc52433fb0ed409765d

Observation 5d978a87-4205-43fd-bd9a-0b116899fbf5 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.596461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.596461Z digest=sha256:5f11a49b66ec41bf9b8fb6ec20e06244b2e315724d8afcae1d75d746ade70b31

Pith citing papers

No inbound Pith citation observations are available.