Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:58:44.792769Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.18258.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:58:44.792769Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a1d01a30-cc13-41bc-87df-4428c04ff6a0 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e997443e-0234-45a4-966f-0222cca11a2f · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fine-Tuning Language Models from Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a641f918-90d0-4565-ac0b-88e965e5dc29 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d72c03d-8f2a-4a5f-a77a-4149ff0e054b · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Rlhf: A comprehensive survey for cultural, multimodal and low latency alignment methods.arXiv preprint arXiv:2511.03939, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a6dedc-8cd2-4cb2-8b7c-7e324c731514 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6d387b-8abf-4ccb-9373-7382e95a2bdf · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Confronting reward model overoptimization with constrained RLHF
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7161b2-cfe7-4a4f-aa95-d690a487ce79 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Regularizing hidden states enables learning generalizable reward model for llms.Advances in Neural Information Processing Systems, 37:62279–62309, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811c9ace-3bff-4de1-b814-c5edaaea8256 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f1947b5-3dd1-4def-8064-62d01c363751 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DPO meets PPO: Reinforced token optimization for RLHF
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c691c864-bdc1-4f50-a799-165888626e13 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 466a9e3d-1a39-45d3-9d58-f200606e3a30 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Capo: Towards enhancing llm reasoning through verifiable generative credit assignment.arXiv e-prints, pages arXiv–2508, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bc4ba3-52a9-44f0-8c10-e91d630d0ece · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b3e7637-bf81-47cb-905e-67f9f28ff4d7 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68823f1-e86b-47a1-976e-7c63d2205a7d · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Finetuned Language Models Are Zero-Shot Learners
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2e5afc-4658-4ede-8538-ee45c4e8f27b · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to summarize with human feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f21a7eb-3ce9-4467-86ea-e1bea4692410 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26ac867-d237-4dfe-8199-f00c18ea6625 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF The Alignment Problem from a Deep Learning Perspective
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3930e48f-e5d7-44e7-930b-044532b77b35 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Stepwise alignment for constrained language model policy optimization.Advances in Neural Information Processing Systems, 37:104471–104520, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1f8bcc-148c-46c9-a36d-ed815553aac2 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Panacea: Pareto alignment via preference adaptation for llms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bf01e07-aabb-4491-b581-22710502770a · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Preference learning for AI alignment: a causal perspective
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0e8ae0d-f299-40f2-8fd8-e7a528581d5c · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Direct preference optimization: Your language model is secretly a reward model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382bfbfe-7516-4062-ae85-35d2c6448ef3 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e82d91-7de6-41a1-b1bf-5bfc59681804 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ORPO: Monolithic Preference Optimization without Reference Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d6b48d-d7e1-4024-be27-3c0143647f9f · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 935ed82c-84b8-4118-b846-e028770ca025 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Scaling laws for reward model overoptimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9812646-da51-41eb-a4e6-fa5dfe955922 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Learning to Understand Goal Specifications by Modelling Reward
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5a709f-7554-4495-9096-3cfa72d7bed2 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc5c481-6322-44df-8ee8-1de5b8809906 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3806584-e0f4-46d1-af92-90f3e81efe65 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Privately Aligning Language Models with Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960fd674-aabb-4829-8aad-7e28efe7dab5 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Dense reward for free in reinforcement learning from human feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91756c88-c717-4cf3-b916-2d149021361d · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033c1848-ef05-4e5d-80ec-ad830e59491a · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF SCAR: Shapley Credit Assignment for More Efficient RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b31347-93a8-45a4-91e0-514061b8e5b6 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498542d1-fddb-41dd-8ee0-ee271e8119cc · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396a5801-d84e-4043-b36b-9754894356f4 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Judging llm-as-a-judgee with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ecb033-fee4-43e5-8fe7-f7a3200cdc60 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4o System Card
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3221b3-a337-4f26-97bb-7572eddb06bc · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF GPT-4 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e76b477-3e0a-403b-91ae-c82d6ca1d08c · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Raft: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3cf37e1-420e-40a5-9ee9-a40c56a6fdbf · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Cambridge University Press, 1999
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c71261-88d4-4fa8-b19b-c396ef0f6bd3 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Multi-Task Learning as a Bargaining Game
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e27c840-2e87-4f65-a49d-bdf9865982a5 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Fairness-aware meta-learning via nash bargaining.Advances in Neural Information Processing Systems, 37:83235–83267, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c329ffb8-72cb-4586-a3e9-6ff234c50431 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Hindsight-DICE: Stable Credit Assignment for Deep Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a37414f-9070-4741-b658-b8279522a888 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 230a840b-9267-4b04-b53d-8fdb971f7559 · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832150d9-4421-4da7-935a-8d3d61499e3a · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c598da-52c3-485c-99dc-b42c5ac3be2d · outbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.