Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

As of 15 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2506.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08266 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.815824Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:19.102106Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T08:31:16.863245Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56377b31-6649-45fe-88dc-a8000252c936 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.599591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.599591Z digest=sha256:6b3920cfdd0f617e3c2f2b10a342c4553ea9b54fcc1e9df59616b1a239db8efc

Observation 1d856b60-19ac-4048-827b-1bafa03f8893 · outbound

This paper cites Constrained Markov decision processes.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constrained Markov decision processes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.604175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.604175Z digest=sha256:731d366a6a2bf3af9973fa5d269c2383c5bba9cb337d928acb892090b44a22c5

Observation d6aa3d0b-ca63-4b80-8d0d-c16bd0df323a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.608048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.608048Z digest=sha256:2cca460e0b6274bb79aa4a87bc6d60ad4f836456fd0a5083112926366039a6d5

Observation fbe55287-cb2e-4b07-84f9-2299ce6ab95d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.612077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.612077Z digest=sha256:afcb9514f6d0513211d1eefa4f0e9ee9da46819d8a4b12c5f9d580f98990c8e3

Observation c12c71ee-0016-4ee0-9047-9f03ee6fe616 · outbound

This paper cites Convex optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.615943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.615943Z digest=sha256:76093194395452d0427aca4095a653f21d8c55405ba8c974870be5f4ecdd5a26

Observation d59d16e5-e0a7-46dc-a0cb-a7bed72ca393 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rank analysis of incomplete block designs: I

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.520895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.619931Z digest=sha256:1826b5d89f135cf6c7c81c5635a3754437c4e9288ac6519d862072f99aed40d9

Observation f2e01823-14fb-4cb4-b533-7bcffd74649e · outbound

This paper cites Deep reinforcement learning from human preferences.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.624040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.624040Z digest=sha256:aacb452246593f4d9ef29fe4434d12dd2e58787a211995ce4f76b170c06481d1

Observation 43af2ee8-fb62-4003-a3d9-36fbd6627858 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.628042Z digest=sha256:b942025a51bf1bf5118b38746f103e3edf940af50731dce161a996b602ed5324

Observation d0462d37-6aa5-49ff-b43f-4347a7448210 · outbound

This paper cites Policy Gradients with Variance Related Risk Criteria.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Policy Gradients with Variance Related Risk Criteria

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.225698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.631826Z digest=sha256:a25c2c700b86ce7ce6e1f9cbafeca01ea072dccec21b602d7480e11ce101027b

Observation e182abdc-9e22-47b8-9852-ac3d6f0eb28d · outbound

This paper cites Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.636063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.636063Z digest=sha256:ff70fd12bc05923bd6dd65348d291ccb896859bc5417fa8a744ada3ab60b58c4

Observation e8c042f1-7573-493f-aac2-c9a443c337f4 · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.639987Z digest=sha256:0c80ba34847294943f4408f342c7f45c8bbddd36f3d3dbd8f36c09ee24e0b819

Observation 2f07684d-8a34-45e2-9aaa-acf16bc1838d · outbound

This paper cites Fundamentals of optimization theory with applications to machine learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fundamentals of optimization theory with applications to machine learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.509176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.643947Z digest=sha256:c906840c75e595f67d2602bdb6d5c5ff97a10afb8fd3d2b8b02ec08808c78580

Observation 16df702a-b512-4412-854f-3cd7bae3fe74 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.647177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.647177Z digest=sha256:542df35631980b64ec1771b470f1a8a818fbdb20b2b87eedff15a96ac7816998

Observation 3268aad4-55be-4c92-9fa7-318e0e010cf5 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.497961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.650410Z digest=sha256:a4225148c683304565745b63987f27356e6e1eacdd006e38b5e7af36d8c00691

Observation 94183662-19cc-44e8-b867-3065a1951642 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:21:59.485906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.653546Z digest=sha256:120ff646dfed28692a3cf18cce845cc1ed728ad79dbe3b21e0c96c069cd08690

Observation adafa1f0-2f03-4377-afbc-cc0c6935a6eb · outbound

This paper cites Fairness guarantees under demographic shift.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fairness guarantees under demographic shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.475161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.659981Z digest=sha256:a889a835fb06f24c9c74aca31319e3f21b3d3e977055de2a25363cf071c15fef

Observation fb60a32b-e206-4001-9b95-87797be7d16a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.668478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.668478Z digest=sha256:3b847bc74a0df2da808dc333203ae6500debe89e287756c7a79352175f955a76

Observation b4fca4c5-6a67-46c0-8bf4-05c628320eea · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.671965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.671965Z digest=sha256:16fac7af8181b7d7e2ce5aad73429063437ddd7eb2732d1fbef586d847805806

Observation df3cfe86-df38-4464-b643-b1f6920c81ed · outbound

This paper cites Ethical Challenges in Data-Driven Dialogue Systems.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical Challenges in Data-Driven Dialogue Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.675367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.675367Z digest=sha256:82ab3265dfb9f786f3469cf6a408732b52c5ee0cffa8ccea41dd3963c742f039

Observation d3b9665f-e183-41e6-9990-aaf7b1ae7336 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Probability inequalities for sums of bounded random variables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.679396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.679396Z digest=sha256:fd69418ad8f54420d68792b264dd974f74ed063924408dbdf0ddcc7a7a5ebd41

Observation 5bf73396-d333-484c-8515-8a597fe4ae00 · outbound

This paper cites One-shot safety alignment for large language models via optimal dualization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints One-shot safety alignment for large language models via optimal dualization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.456542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.682720Z digest=sha256:ec099bf535f29205f5df71ce6f7da7170728e955986f3d808ec81bb6ce50d9f0

Observation 085d504c-bea5-441e-8287-5688ca3b64fa · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.685785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.685785Z digest=sha256:52f6077e6cd3ecc6e981d74a358e58fa3c5e774f2e80f141b3e73527534f82b2

Observation 4b85f2f4-dc88-4f8d-a4d0-1933183454bd · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.689205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.689205Z digest=sha256:ba2ce242fdd6b0f4276c6eddd7105d31d9fdc1e5344ce1adfa2bff324f3a2049

Observation 3a4a982a-e92c-4f25-8552-7c6e80c6392b · outbound

This paper cites ChatGPT for good? O n opportunities and challenges of large language models for education.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints ChatGPT for good? O n opportunities and challenges of large language models for education

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.444817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.692927Z digest=sha256:d2ae6af3efea350008980bc2a2095208afe8bcecc06e4f9fb70b55452416a8c4

Observation 78efc2c2-f73b-463f-bdd9-560b19c3c745 · outbound

This paper cites GPT -4 passes the bar exam.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints GPT -4 passes the bar exam

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.432588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.696131Z digest=sha256:7334bf33cb4c9a494a87cacb70f03a860beab4c770108c9135ebcd0d33eb80b5

Observation a9a7b0f8-a552-4b74-bd8c-f14dc96d2620 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.699651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.699651Z digest=sha256:15656ac8ed9bb0bded1fd09ef6ddd36e2d3c7720f00766666920d1386f61b869

Observation 31cb7100-2d54-4f7c-bc55-069b11578069 · outbound

This paper cites Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.413594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.703396Z digest=sha256:1ec3dfcf2fd5b51cd57d9c8e6fc2846b3bf61642c95cb1df0fe2e546a9e1e18e

Observation 094a5a8d-af98-4b91-83a0-bad6ed9f6a95 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.706600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.706600Z digest=sha256:a5f524becaf6971a6c919c7929347bca1b79ff1f8abd0e6d34b1b1b812296fab

Observation 00e2e0a5-26c9-4697-956f-4b4fb406a4e0 · outbound

This paper cites Offline contextual bandits with high probability fairness guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Offline contextual bandits with high probability fairness guarantees

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.401561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.710273Z digest=sha256:fc6f27275522e374ae1d22f9205902dbbffed643fbfe22856829b72f1653ae56

Observation fc679d97-1a53-4ad8-b491-8692792fbc38 · outbound

This paper cites Krumholz, Jure Leskovec, Eric J.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Krumholz, Jure Leskovec, Eric J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.389891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.713452Z digest=sha256:0a55d6df258e94bd3334db91abec56fb8665bc367b85817808d699009b908fc8

Observation bce4a920-bc4e-44b7-b7fd-fb5af4712eee · outbound

This paper cites Rule based rewards for language model safety.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rule based rewards for language model safety

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.379311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.716656Z digest=sha256:7c49f15e8de8c51cc3a7ea9948df382ab294ecbfb3a1f46be4348ca903145cc3

Observation ce0cb1d9-f734-47cb-b6ed-741e6ad2addd · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.720364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.720364Z digest=sha256:c33bb077b177d6b94f38ed866d137393dd207d06311c722d225c92ad867b6d60

Observation 08ef3aa0-5ad6-4b8b-a146-bd0f8c20961f · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.026775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.724147Z digest=sha256:6f9a085255541ec678a4752a24ae09c8d4a1a57166066d790cc458acd3cf9dc6

Observation f5776140-9477-4cb4-93ed-fd1f3ff2abba · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.727951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.727951Z digest=sha256:c05e7c97aaced9bbdc8fc457334fe8ad8d7b4a0bf5416f3be88b52df179ff0da

Observation 6dba9141-8a58-4692-a9b2-e1d3b097c395 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.731769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.731769Z digest=sha256:05a8566dfb0df8edb131e1978f3ab7de6b372598282be76aff585325ff116f35

Observation 14082912-204e-42f1-be50-db845eab2700 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.735489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.735489Z digest=sha256:03872632886c7089fddca91879c59623e9c2addf2ab78de81abe1bc6853704e9

Observation 681def66-995e-4f7e-b389-928d39c912c2 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.367354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.739826Z digest=sha256:4387dd810b37674ffc6fae5f772dc7af56c1679370b88895fc09bb931cfd3a90

Observation feb3e995-1e75-40f7-870f-c1a9fbca64fb · outbound

This paper cites Simultaneous statistical inference.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simultaneous statistical inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.356121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.743188Z digest=sha256:0734b69cdc40f90167820410911729e50bace87a57291e051b4f825aa2d08422

Observation 5fcea960-3005-4e63-bfb7-1b218532842a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.746310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.746310Z digest=sha256:0827a6a5cd4f5bb23989264fd9bdb835ceeea2d60991614193f43e1e1b4e8c14

Observation 0425a0d8-0177-4019-85af-d75be4267884 · outbound

This paper cites Learning to summarize from human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Learning to summarize from human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.749826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.749826Z digest=sha256:dad9251a6dcea37bedee98d672b25b4c749ecbcf042fdb460a98ac5f25c65833

Observation 449cef09-448b-4918-8f7b-1c2c972aa1c0 · outbound

This paper cites The probable error of a mean.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The probable error of a mean

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.345555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.753669Z digest=sha256:24f9355853c4e2b878e6e9c0994bfb20a5b4466e1167006652972c5d534352d3

Observation 7de895b4-008b-4581-abd2-f98f8bea42bc · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.756875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.756875Z digest=sha256:128a82329793fe2072c83904c051ff6fd4207864fc9e4d30bd433aec1f846cac

Observation 5fc9254e-270a-4133-9587-da6866bd2b74 · outbound

This paper cites Hashimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Hashimoto

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.760319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.760319Z digest=sha256:35b3730fad05a09ca920f97509ad11d18a9e00b1bde74537944c495cc8567d33

Observation d515259b-92f9-4ecb-b751-de47abce81f2 · outbound

This paper cites Preventing undesirable behavior of intelligent machines.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Preventing undesirable behavior of intelligent machines

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.327648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.763641Z digest=sha256:7551464076749b5881a39fef5e1a1924c4d0965a0833a54aa11def6db3bd4f57

Observation 64dc4375-c341-49e2-8d1f-aa55230916cd · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.767276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.767276Z digest=sha256:b1e036e72d25b171972f7f3049772277194350e49d6172cd4c57923b35c4c578

Observation c6401466-52e9-49e9-8e0b-1907d3b73769 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.770901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.770901Z digest=sha256:644d6b6c1a3e3b5d1bd1e8778017458dc96c6d3ab97b7a7e6d98f5a9a6e5d66d

Observation ce9dc7c0-7761-480a-a53a-8b110137fd31 · outbound

This paper cites Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.316321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.774310Z digest=sha256:701cce6ec88f78d02a234aaa85904fb312730927da7cffca2c8c089efced1483

Observation 46904a93-d126-448b-ac6c-959027d47ce8 · outbound

This paper cites Enforcing Delayed-Impact Fairness Guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enforcing Delayed-Impact Fairness Guarantees

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.932492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.778283Z digest=sha256:fa9700bbf3ff394b3fafcb7b5cdae558fb12102138bf9db8a7f4a5e296def258

Observation 5b6357af-c971-41de-9d6a-10a85f23d23e · outbound

This paper cites Ethical and social risks of harm from Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical and social risks of harm from Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.781959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.781959Z digest=sha256:1885067a3afa4f020999bf809b74f396cf9ac7fddae1270e1aa186347df1486e

Observation 367a0044-733d-4030-86de-d0ac856dc876 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.785508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.785508Z digest=sha256:538bf2c3405eb6ca539b9cd7b7f01c2a3ed1f339fcb1fdc36464463212c98bde

Observation 17688df0-e3b8-42e0-91ee-ffcc0610464f · outbound

This paper cites Recipes for Safety in Open-domain Chatbots.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Recipes for Safety in Open-domain Chatbots

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.788856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.788856Z digest=sha256:e8801a3b9b28cfe721be3398d17a7d8557826c7389d020e10e60aa7e7964764b

Observation 215d5810-a9fe-434f-b1cf-63c3a55c3b00 · outbound

This paper cites Qwen2 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.792512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.792512Z digest=sha256:b244223a1e8009ecda8a9a20289e7c4d51682b4c390544d9b82d469b262818be

Observation 3e7ecda3-d383-418a-9fd8-e16ffa389a83 · outbound

This paper cites A large language model for electronic health records.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints A large language model for electronic health records

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.298310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.796231Z digest=sha256:a06d5d0fe7daf219c1736187ee40e8470629f279467615bc66ad378309fcd9e4

Observation f8037334-eadc-4a1f-8d46-ec3f0acf2601 · outbound

This paper cites Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.886951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.799911Z digest=sha256:580bf44096fbfc9efdfa1e9d2126e7b1c7c0375c2974b7dbdfad129f7756a191

Observation 373be2dd-3351-4c21-85a9-93d010d415bf · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.803923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.803923Z digest=sha256:ac46e751946860fe23ff55119f3df72f805e785d19bb687d0523bd4ba98e7cef

Observation 4e532ecc-f56f-4b18-b185-9ca210e2b8c5 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Secrets of RLHF in Large Language Models Part I: PPO

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.808116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.808116Z digest=sha256:f5aa96ab73b1382a0866d58e46f6293c0282f263aa0a67479571df7cf779d489

Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.812008Z digest=sha256:6d402da9a11c4e6387f1305c99f0d08f9f8c05d054b80fc5b0b39a94138a1907

Observation f98a0e96-415d-4e86-9bf3-5e8cacb70470 · outbound

This paper cites write newline.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.815824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.815824Z digest=sha256:9aa9b8ff9634c16db27896b82b6a3f4aa8f5deba1c2051b613730edd5af5f4ec

Pith citing papers

Observation 32637181-00ce-4796-a516-eaa778829233 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:19.102106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:19.102106Z digest=sha256:bd90b86b02ee5fe5ecad2d35e3e5ca808582960362e7e62929d2b0e6ebe24ffe

Observation 1437125e-c3ed-4c8c-85f0-a701b5c4ff4e · inbound

Reinforcement Learning from Human Feedback: A Statistical Perspective cites this paper.

Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.609624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T20:10:43.578904Z digest=sha256:3dc3e4bac1fd0c3c3d94940cad9b6980169ce1643ab558a7d0e9acd4ab83393d

Observation 8345e638-7f56-407a-8140-7d5adc8af28e · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.865653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:4258f07531ec0f0af82b6a5cf276fb675ecc2db024624bb5597404215143702f