Pith. sign in

Paper Citation Record · LEDGER

Data-adaptive Safety Rules for Training Reward Models

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 3 inbound Pith citation observations for arXiv:2501.15453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15453 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:25:13.230661Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:56:15.621439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:34.990786Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2580469f-61db-4e46-bfda-2c09e04d1bf5 · outbound

This paper cites GPT-4 Technical Report.

Data-adaptive Safety Rules for Training Reward Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.088401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.088401Z digest=sha256:0bb29d1dc7302a4e21799196be16c9353ce3171bf4b4710a1d7f1929cce97acd

Observation 27bce344-63e0-43a5-b1e8-cf48a2dcc71c · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Data-adaptive Safety Rules for Training Reward Models Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.142485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.142485Z digest=sha256:f88fc4331e311bc0f54f39b8bdfd3d0370ccf8dc74e9ddd9077c3f0caa1f321c

Observation 0094cbfa-69e6-4ce1-b644-a9f29f1420e6 · outbound

This paper cites {severe level} harm question:.

Data-adaptive Safety Rules for Training Reward Models {severe level} harm question:

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.756734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.226125Z digest=sha256:e66e22dbf8ec3eef77dd6276e68e90776fdced51a5da3f3104165ebf38e9ba36

Observation f470421d-ab73-4e1c-b06d-bb39dce0e019 · outbound

This paper cites SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF.

Data-adaptive Safety Rules for Training Reward Models SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.113800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.113800Z digest=sha256:e219efebb402f362073c03436b79c05f21263534d603804fb16e86216b187398

Observation cafb5bb8-e090-46e9-96e6-715b3ae5428a · outbound

This paper cites Quantile Regression for Distributional Reward Models in RLHF.

Data-adaptive Safety Rules for Training Reward Models Quantile Regression for Distributional Reward Models in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.118759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.118759Z digest=sha256:895b7acb0edd3468c9f6a7c0f2222791023a9decc80366e3484c385a666338e9

Observation baeecbd7-6222-423c-bd54-a7e288f10e8e · outbound

This paper cites The Capacity for Moral Self-Correction in Large Language Models.

Data-adaptive Safety Rules for Training Reward Models The Capacity for Moral Self-Correction in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.128436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.128436Z digest=sha256:6326635cc72867e4c896f5c9e2a5ee7f3b3ac3ddefce77beb37d362561a60baf

Observation 61ddd42a-ac36-44ab-bb67-176891b13f68 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Data-adaptive Safety Rules for Training Reward Models Improving alignment of dialogue agents via targeted human judgements

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.132837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.132837Z digest=sha256:8678b7ea8c6e50438649738d25ac46d3f61cf88d8e55adda8b3810e1b913b6d7

Observation 7fce955e-fa59-4530-816f-4f37da03cc29 · outbound

This paper cites Collective Constitutional AI: Aligning a language model with public input.

Data-adaptive Safety Rules for Training Reward Models Collective Constitutional AI: Aligning a language model with public input

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.815217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.137934Z digest=sha256:0bf33520ea3a5f0973d839777e47fa257be7c754fb94319540337fb45934f1d0

Observation b12fa28a-d167-4fcf-aeb5-fa8ad8b9afc4 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Data-adaptive Safety Rules for Training Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.146914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.146914Z digest=sha256:6a27544726acac1db5387e29be17d610195bbbc94a61b508df4a9bd92ea593cc

Observation a941f6f1-daac-4ccd-922d-6c52dd840f81 · outbound

This paper cites Mistral 7B.

Data-adaptive Safety Rules for Training Reward Models Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.151591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.151591Z digest=sha256:dfc90a7bb186a31c65ab34c4327381b422229473a336d5c313c7ff5df830ea55

Observation 778afdcc-4115-4c0f-a689-d5149d6a5281 · outbound

This paper cites Mixtral of Experts.

Data-adaptive Safety Rules for Training Reward Models Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.155955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.155955Z digest=sha256:69a11f06d3914ba955e72115e26343267224514cbd12a63d9a7d1c49a9a8b7e2

Observation 95b18964-317a-47b4-aff0-a183c9e5f0b2 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Data-adaptive Safety Rules for Training Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.165585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.165585Z digest=sha256:f24833ff9bca24f0dba3b3a14ab572e8200d7a34864c649e53ec73df76389093

Observation 52271593-4ae9-48d8-9291-b2b50125b6ad · outbound

This paper cites Rule-based data selection for large language models.

Data-adaptive Safety Rules for Training Reward Models Rule-based data selection for large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.170124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.170124Z digest=sha256:dee4edf25074dafa4e0ac132ce8d6abb2770a7b073e360ba9cc6f59fd231ae38

Observation b2119d49-2202-4e31-b9a1-0c82e0ae652b · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Data-adaptive Safety Rules for Training Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.174468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.174468Z digest=sha256:a926412dcadca76b910a33467e4deae8598d396ee96f173d865e616b24e4cf6a

Observation 60177fe9-4da5-4bc5-9d79-4b58bed8ef16 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Data-adaptive Safety Rules for Training Reward Models Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.179261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.179261Z digest=sha256:1d87be4a1f5396a01508e88a2e01d3c9ba0f17b24516ff533b876f896edcdc2a

Observation 80ece55a-59bb-4ec9-ac52-06f13615a23a · outbound

This paper cites Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization.

Data-adaptive Safety Rules for Training Reward Models Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.188226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.188226Z digest=sha256:975094c2ad339d485b9cbcdc7a677d7469fc96e7e6d714c822ddc05d504c6ac7

Observation 98c4542a-fc8f-46fe-a415-90a40a27d33f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Data-adaptive Safety Rules for Training Reward Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.193055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.193055Z digest=sha256:7edfea7e71d0065b61105b9fbaf661e63cb2339a7430bffe187252242595cdcd

Observation 7a34b5fb-7a42-4c67-b238-2a08df5dd680 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Data-adaptive Safety Rules for Training Reward Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.197665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.197665Z digest=sha256:8ff88724922f95e0b663a76f517eca3fd723d37c558eb07da60bf43761fac69b

Observation 589cc8e1-0d44-499e-99c3-9300818f6f27 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Data-adaptive Safety Rules for Training Reward Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.202305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.202305Z digest=sha256:84ce2107ad440fd39d923dde4f63147317db65db78a141befd834b42fe5638bb

Observation 15725ff4-32ee-4da2-80d6-910a317613ff · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Data-adaptive Safety Rules for Training Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.206994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.206994Z digest=sha256:9ba08e2d4989dd0e3731ea49dcf07db6b284bcbaeb541e0a1125359dd45e4d3d

Observation a44391cd-317e-447f-a0d8-44b23ac6a679 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

Data-adaptive Safety Rules for Training Reward Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.211691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.211691Z digest=sha256:536f472749cf9c8b0b31579fcfc04a7d7a376267c334242c678478491a18d69f

Observation 9c4d6ac4-34cb-4be4-85e9-df16ad2ca076 · outbound

This paper cites − X t P+(t) logP+(t) # + 1 2.

Data-adaptive Safety Rules for Training Reward Models − X t P+(t) logP+(t) # + 1 2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.788799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.216632Z digest=sha256:d09b685e6ee4a78835ac760ba891916f4ec77c7e146ce39ae90d5d0498d554da

Observation e82345c2-ad06-45c2-b429-4143a1f3a7b1 · outbound

This paper cites 16 Proof of Theorem 3.4.

Data-adaptive Safety Rules for Training Reward Models 16 Proof of Theorem 3.4

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.771723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.221460Z digest=sha256:40234a9060781997a5616ab2ec9917ef8b92305e4dcc0affff255cf7786c016e

Observation f9badda7-99c6-439c-8cfb-18042465e3b4 · outbound

This paper cites Results are averaged over 2 trained models with different random seeds for optimal hyperparameter selection.

Data-adaptive Safety Rules for Training Reward Models Results are averaged over 2 trained models with different random seeds for optimal hyperparameter selection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.741166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.230661Z digest=sha256:57a7faf70ed159a8d4cfc70347deec1113262868b2583ef16206dd32bdacda33

Observation 2d4d4fd1-2249-4b0a-836a-59f9ed0a4bb3 · outbound

This paper cites Specific versus General Principles for Constitutional AI.

Data-adaptive Safety Rules for Training Reward Models Specific versus General Principles for Constitutional AI

Reference 1951

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.160754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.160754Z digest=sha256:d850e7ffb1e287c8f19be5ce6bce09564f71a5d345c020127c2915049d05ba3b

Observation c8552f2a-d285-4119-a6c9-ee9beb47689c · outbound

This paper cites Language Models are Few-Shot Learners.

Data-adaptive Safety Rules for Training Reward Models Language Models are Few-Shot Learners

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.108657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.108657Z digest=sha256:08ca5c976ba399f2388705446250a6aec0357e466e6aad33a71beddfbc7bfc0a

Observation c0689ef8-c045-44ef-a253-98bd8de39d5f · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Data-adaptive Safety Rules for Training Reward Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 1975

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.183662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.183662Z digest=sha256:c3299123170e4ff71e871eeba6160a64811c71adbba60a90b868d2910ad67f49

Observation ba69f6bc-1fa7-4417-837c-fa5160970cdd · outbound

This paper cites The Llama 3 Herd of Models.

Data-adaptive Safety Rules for Training Reward Models The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.123550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.123550Z digest=sha256:1d28c16bca0b9960f61ed77a8e4204a5142ed88e1ed607952dd46c973471daff

Observation 88c204dd-5cdf-441e-b506-db0b9521d109 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Data-adaptive Safety Rules for Training Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.103534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.103534Z digest=sha256:12fbd1818b97b9f097bff2f6b38d7f12e051020f4a55a11563e43fdd98c92284

Observation c5ed5287-008a-404d-ba01-f8b749518143 · outbound

This paper cites Qwen Technical Report.

Data-adaptive Safety Rules for Training Reward Models Qwen Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.098361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.098361Z digest=sha256:e2ac58f23caaec571e15cf3b67523a5bc6ee46add938f2d422ad20e295d70e0d

Observation 433e024a-c09c-44cd-898b-c5318d5171dc · outbound

This paper cites Llama 3.2: Advancing ai on edge and mobile devices.

Data-adaptive Safety Rules for Training Reward Models Llama 3.2: Advancing ai on edge and mobile devices

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:25:13.832460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T14:25:13.093692Z digest=sha256:3e11f3abb8c33c4d43d795dd847c91b9af2bed05af0987472af44e4cdd3c1d71

Pith citing papers

Observation 5d6ef36b-588b-4687-8898-ad4c1ce30ca2 · inbound

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models cites this paper.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Data-adaptive Safety Rules for Training Reward Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.621439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.621439Z digest=sha256:a1b8705f75426d923c88ca9187d7888278d9b462554f9306dc72610ba242f079

Observation 69708228-e735-41ca-9311-dc586d66fc1f · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Data-adaptive Safety Rules for Training Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.438057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.438057Z digest=sha256:85b9930229330e036eb49c27a5fc4b96bff5ef3fad5832c956734bb577e548d2

Observation 26cf5ff7-62cf-458f-8a51-58f0d698fee8 · inbound

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining cites this paper.

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining Data-adaptive Safety Rules for Training Reward Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:19:34.992897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T17:02:30.696295Z digest=sha256:c2cf74e8a30fa0682ea737b7a6cb80f1586dcf7ca64364c8e31f9697e805012c