Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:25:13.230661Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2501.15453.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:25:13.230661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:27.438057Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:19:34.990786Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2580469f-61db-4e46-bfda-2c09e04d1bf5 · outbound
Data-adaptive Safety Rules for Training Reward Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27bce344-63e0-43a5-b1e8-cf48a2dcc71c · outbound
Data-adaptive Safety Rules for Training Reward Models Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0094cbfa-69e6-4ce1-b644-a9f29f1420e6 · outbound
Data-adaptive Safety Rules for Training Reward Models {severe level} harm question:
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f470421d-ab73-4e1c-b06d-bb39dce0e019 · outbound
Data-adaptive Safety Rules for Training Reward Models SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafb5bb8-e090-46e9-96e6-715b3ae5428a · outbound
Data-adaptive Safety Rules for Training Reward Models Quantile Regression for Distributional Reward Models in RLHF
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baeecbd7-6222-423c-bd54-a7e288f10e8e · outbound
Data-adaptive Safety Rules for Training Reward Models The Capacity for Moral Self-Correction in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ddd42a-ac36-44ab-bb67-176891b13f68 · outbound
Data-adaptive Safety Rules for Training Reward Models Improving alignment of dialogue agents via targeted human judgements
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fce955e-fa59-4530-816f-4f37da03cc29 · outbound
Data-adaptive Safety Rules for Training Reward Models Collective Constitutional AI: Aligning a language model with public input
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b12fa28a-d167-4fcf-aeb5-fa8ad8b9afc4 · outbound
Data-adaptive Safety Rules for Training Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a941f6f1-daac-4ccd-922d-6c52dd840f81 · outbound
Data-adaptive Safety Rules for Training Reward Models Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778afdcc-4115-4c0f-a689-d5149d6a5281 · outbound
Data-adaptive Safety Rules for Training Reward Models Mixtral of Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b18964-317a-47b4-aff0-a183c9e5f0b2 · outbound
Data-adaptive Safety Rules for Training Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52271593-4ae9-48d8-9291-b2b50125b6ad · outbound
Data-adaptive Safety Rules for Training Reward Models Rule-based data selection for large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2119d49-2202-4e31-b9a1-0c82e0ae652b · outbound
Data-adaptive Safety Rules for Training Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60177fe9-4da5-4bc5-9d79-4b58bed8ef16 · outbound
Data-adaptive Safety Rules for Training Reward Models Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ece55a-59bb-4ec9-ac52-06f13615a23a · outbound
Data-adaptive Safety Rules for Training Reward Models Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c4542a-fc8f-46fe-a415-90a40a27d33f · outbound
Data-adaptive Safety Rules for Training Reward Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a34b5fb-7a42-4c67-b238-2a08df5dd680 · outbound
Data-adaptive Safety Rules for Training Reward Models LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589cc8e1-0d44-499e-99c3-9300818f6f27 · outbound
Data-adaptive Safety Rules for Training Reward Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15725ff4-32ee-4da2-80d6-910a317613ff · outbound
Data-adaptive Safety Rules for Training Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44391cd-317e-447f-a0d8-44b23ac6a679 · outbound
Data-adaptive Safety Rules for Training Reward Models SafetyBench: Evaluating the Safety of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c4d6ac4-34cb-4be4-85e9-df16ad2ca076 · outbound
Data-adaptive Safety Rules for Training Reward Models − X t P+(t) logP+(t) # + 1 2
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e82345c2-ad06-45c2-b429-4143a1f3a7b1 · outbound
Data-adaptive Safety Rules for Training Reward Models 16 Proof of Theorem 3.4
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9badda7-99c6-439c-8cfb-18042465e3b4 · outbound
Data-adaptive Safety Rules for Training Reward Models Results are averaged over 2 trained models with different random seeds for optimal hyperparameter selection
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d4d4fd1-2249-4b0a-836a-59f9ed0a4bb3 · outbound
Data-adaptive Safety Rules for Training Reward Models Specific versus General Principles for Constitutional AI
Reference 1951
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8552f2a-d285-4119-a6c9-ee9beb47689c · outbound
Data-adaptive Safety Rules for Training Reward Models Language Models are Few-Shot Learners
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0689ef8-c045-44ef-a253-98bd8de39d5f · outbound
Data-adaptive Safety Rules for Training Reward Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 1975
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba69f6bc-1fa7-4417-837c-fa5160970cdd · outbound
Data-adaptive Safety Rules for Training Reward Models The Llama 3 Herd of Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c204dd-5cdf-441e-b506-db0b9521d109 · outbound
Data-adaptive Safety Rules for Training Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ed5287-008a-404d-ba01-f8b749518143 · outbound
Data-adaptive Safety Rules for Training Reward Models Qwen Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433e024a-c09c-44cd-898b-c5318d5171dc · outbound
Data-adaptive Safety Rules for Training Reward Models Llama 3.2: Advancing ai on edge and mobile devices
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69708228-e735-41ca-9311-dc586d66fc1f · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Data-adaptive Safety Rules for Training Reward Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26cf5ff7-62cf-458f-8a51-58f0d698fee8 · inbound
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining Data-adaptive Safety Rules for Training Reward Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.