Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2506.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08266 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.815824Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:19.102106Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T08:31:16.863245Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56377b31-6649-45fe-88dc-a8000252c936 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.599591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.599591Z digest=sha256:6b3920cfdd0f617e3c2f2b10a342c4553ea9b54fcc1e9df59616b1a239db8efc

Observation 1d856b60-19ac-4048-827b-1bafa03f8893 · outbound

This paper cites Constrained Markov decision processes.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constrained Markov decision processes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.604175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.604175Z digest=sha256:731d366a6a2bf3af9973fa5d269c2383c5bba9cb337d928acb892090b44a22c5

Observation d6aa3d0b-ca63-4b80-8d0d-c16bd0df323a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.608048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.608048Z digest=sha256:a47967c9011fb10bcc4f5e982b158fca323f38af05235f158b1221c77f1ae20d

Observation fbe55287-cb2e-4b07-84f9-2299ce6ab95d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.612077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.612077Z digest=sha256:afcb9514f6d0513211d1eefa4f0e9ee9da46819d8a4b12c5f9d580f98990c8e3

Observation c12c71ee-0016-4ee0-9047-9f03ee6fe616 · outbound

This paper cites Convex optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.615943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.615943Z digest=sha256:76093194395452d0427aca4095a653f21d8c55405ba8c974870be5f4ecdd5a26

Observation d59d16e5-e0a7-46dc-a0cb-a7bed72ca393 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rank analysis of incomplete block designs: I

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.520895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.619931Z digest=sha256:6a48a7e0c1641acbd6e06e010db015dccf8547f695f73290d8f13bd75f84ebc7

Observation f2e01823-14fb-4cb4-b533-7bcffd74649e · outbound

This paper cites Deep reinforcement learning from human preferences.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.624040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.624040Z digest=sha256:aacb452246593f4d9ef29fe4434d12dd2e58787a211995ce4f76b170c06481d1

Observation 43af2ee8-fb62-4003-a3d9-36fbd6627858 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.628042Z digest=sha256:b942025a51bf1bf5118b38746f103e3edf940af50731dce161a996b602ed5324

Observation d0462d37-6aa5-49ff-b43f-4347a7448210 · outbound

This paper cites Policy Gradients with Variance Related Risk Criteria.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Policy Gradients with Variance Related Risk Criteria

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.225698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.631826Z digest=sha256:1c070709371cdbb4d21ed04b20aecabf2032cffc8f02c991abcf1b5b1a992522

Observation e182abdc-9e22-47b8-9852-ac3d6f0eb28d · outbound

This paper cites Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.636063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.636063Z digest=sha256:4233d2182bb171169ad58a14f99fc3e4175976d027154e05c5fb9a7ac7fbd856

Observation e8c042f1-7573-493f-aac2-c9a443c337f4 · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.639987Z digest=sha256:e93432999aa65955ed864a98e806a9ee1e456631b5990ae3c29cc63cba38f379

Observation 2f07684d-8a34-45e2-9aaa-acf16bc1838d · outbound

This paper cites Fundamentals of optimization theory with applications to machine learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fundamentals of optimization theory with applications to machine learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.509176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.643947Z digest=sha256:d3dc7bbe6cad153912c05a36a5556231496892cfc9f61bc1d7b4e3d04cfe4e84

Observation 16df702a-b512-4412-854f-3cd7bae3fe74 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.647177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.647177Z digest=sha256:542df35631980b64ec1771b470f1a8a818fbdb20b2b87eedff15a96ac7816998

Observation 3268aad4-55be-4c92-9fa7-318e0e010cf5 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.497961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.650410Z digest=sha256:e2a33b412ed636a6db1f2031705181719a4ab3ab7328ef5b9a858cb70067a051

Observation 94183662-19cc-44e8-b867-3065a1951642 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:21:59.485906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.653546Z digest=sha256:d5b66d14bba4500f2eb262f03a0d065a7b9163e34a1bd90ccbb4a9fea23a679a

Observation adafa1f0-2f03-4377-afbc-cc0c6935a6eb · outbound

This paper cites Fairness guarantees under demographic shift.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fairness guarantees under demographic shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.475161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.659981Z digest=sha256:4246f29a32bbbece380452e68e03ca02e5248d2e4e94065eaeeada7a6afa8784

Observation fb60a32b-e206-4001-9b95-87797be7d16a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.668478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.668478Z digest=sha256:b0e8d7ec0d0a35356a5f1b895570977662531173ef499cd807b38f83f7aad6e0

Observation b4fca4c5-6a67-46c0-8bf4-05c628320eea · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.671965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.671965Z digest=sha256:16fac7af8181b7d7e2ce5aad73429063437ddd7eb2732d1fbef586d847805806

Observation df3cfe86-df38-4464-b643-b1f6920c81ed · outbound

This paper cites Ethical Challenges in Data-Driven Dialogue Systems.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical Challenges in Data-Driven Dialogue Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.675367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.675367Z digest=sha256:82ab3265dfb9f786f3469cf6a408732b52c5ee0cffa8ccea41dd3963c742f039

Observation d3b9665f-e183-41e6-9990-aaf7b1ae7336 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Probability inequalities for sums of bounded random variables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.679396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.679396Z digest=sha256:fd69418ad8f54420d68792b264dd974f74ed063924408dbdf0ddcc7a7a5ebd41

Observation 5bf73396-d333-484c-8515-8a597fe4ae00 · outbound

This paper cites One-shot safety alignment for large language models via optimal dualization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints One-shot safety alignment for large language models via optimal dualization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.456542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.682720Z digest=sha256:9a790eba3bce1d5138b66ae6630994820d521108afdd8e7daed916e9ee7b6224

Observation 085d504c-bea5-441e-8287-5688ca3b64fa · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.685785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.685785Z digest=sha256:52f6077e6cd3ecc6e981d74a358e58fa3c5e774f2e80f141b3e73527534f82b2

Observation 4b85f2f4-dc88-4f8d-a4d0-1933183454bd · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.689205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.689205Z digest=sha256:ba2ce242fdd6b0f4276c6eddd7105d31d9fdc1e5344ce1adfa2bff324f3a2049

Observation 3a4a982a-e92c-4f25-8552-7c6e80c6392b · outbound

This paper cites ChatGPT for good? O n opportunities and challenges of large language models for education.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints ChatGPT for good? O n opportunities and challenges of large language models for education

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.444817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.692927Z digest=sha256:21047cd02bc45f0c576456610e432c556149e7bea31dd1911e00bfbc24ba58e5

Observation 78efc2c2-f73b-463f-bdd9-560b19c3c745 · outbound

This paper cites GPT -4 passes the bar exam.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints GPT -4 passes the bar exam

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.432588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.696131Z digest=sha256:63e72f7a9d5aebdeadaf838b93025e4a9ea7d836851e9a87389b3ae9e2dc803e

Observation a9a7b0f8-a552-4b74-bd8c-f14dc96d2620 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.699651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.699651Z digest=sha256:15656ac8ed9bb0bded1fd09ef6ddd36e2d3c7720f00766666920d1386f61b869

Observation 31cb7100-2d54-4f7c-bc55-069b11578069 · outbound

This paper cites Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.413594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.703396Z digest=sha256:bc2771a2f758026a9feac06a947d883a569c65a10e22cc5c306e75f2af43d433

Observation 094a5a8d-af98-4b91-83a0-bad6ed9f6a95 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.706600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.706600Z digest=sha256:c800e9c1ecc4960b559faa54452fc91342749d3afffbdf3a5a637ee351e26d9e

Observation 00e2e0a5-26c9-4697-956f-4b4fb406a4e0 · outbound

This paper cites Offline contextual bandits with high probability fairness guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Offline contextual bandits with high probability fairness guarantees

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.401561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.710273Z digest=sha256:edf2c224044886af5fc70322d2b09c135ae125a7c35942d8b9cbec8df95f5226

Observation fc679d97-1a53-4ad8-b491-8692792fbc38 · outbound

This paper cites Krumholz, Jure Leskovec, Eric J.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Krumholz, Jure Leskovec, Eric J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.389891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.713452Z digest=sha256:536823f7db03af7b683a78b95256f4b870e1d165f0fbb3da7f7d50ab8b29f32b

Observation bce4a920-bc4e-44b7-b7fd-fb5af4712eee · outbound

This paper cites Rule based rewards for language model safety.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rule based rewards for language model safety

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.379311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.716656Z digest=sha256:491abef1b58bafd36dfa3863a798902680f8cb81b36d076e27cea3d2b31f205c

Observation ce0cb1d9-f734-47cb-b6ed-741e6ad2addd · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.720364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.720364Z digest=sha256:c33bb077b177d6b94f38ed866d137393dd207d06311c722d225c92ad867b6d60

Observation 08ef3aa0-5ad6-4b8b-a146-bd0f8c20961f · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.026775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.724147Z digest=sha256:9697220865f9a096f99d1cd728ab3af7a897861e4968aa398b338e27d000e894

Observation f5776140-9477-4cb4-93ed-fd1f3ff2abba · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.727951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.727951Z digest=sha256:c05e7c97aaced9bbdc8fc457334fe8ad8d7b4a0bf5416f3be88b52df179ff0da

Observation 6dba9141-8a58-4692-a9b2-e1d3b097c395 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.731769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.731769Z digest=sha256:05a8566dfb0df8edb131e1978f3ab7de6b372598282be76aff585325ff116f35

Observation 14082912-204e-42f1-be50-db845eab2700 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.735489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.735489Z digest=sha256:907d9f6f385387321ef4f506701f0d8d62341544480e1abbb537938bc860eee2

Observation 681def66-995e-4f7e-b389-928d39c912c2 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.367354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.739826Z digest=sha256:9bf8a0b17e96143cc5a58c1527c229db4543f20d59acf1d19a801b6f1543edc5

Observation feb3e995-1e75-40f7-870f-c1a9fbca64fb · outbound

This paper cites Simultaneous statistical inference.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simultaneous statistical inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.356121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.743188Z digest=sha256:50f575a8ddb5dc77a288cb27fd1809a1685cf8facf93d010d3f57a49e817f3fe

Observation 5fcea960-3005-4e63-bfb7-1b218532842a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.746310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.746310Z digest=sha256:0827a6a5cd4f5bb23989264fd9bdb835ceeea2d60991614193f43e1e1b4e8c14

Observation 0425a0d8-0177-4019-85af-d75be4267884 · outbound

This paper cites Learning to summarize from human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Learning to summarize from human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.749826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.749826Z digest=sha256:dad9251a6dcea37bedee98d672b25b4c749ecbcf042fdb460a98ac5f25c65833

Observation 449cef09-448b-4918-8f7b-1c2c972aa1c0 · outbound

This paper cites The probable error of a mean.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The probable error of a mean

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.345555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.753669Z digest=sha256:95b755dd50bfc8c20d4d6db115ef19eb23e5bbd8aa5339966d7d116c3eadc498

Observation 7de895b4-008b-4581-abd2-f98f8bea42bc · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.756875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.756875Z digest=sha256:da2cfe92842900d4b3fb192fdf5aa3e8d37de56f193690730cc696d1e503a63e

Observation 5fc9254e-270a-4133-9587-da6866bd2b74 · outbound

This paper cites Hashimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Hashimoto

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.760319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.760319Z digest=sha256:35b3730fad05a09ca920f97509ad11d18a9e00b1bde74537944c495cc8567d33

Observation d515259b-92f9-4ecb-b751-de47abce81f2 · outbound

This paper cites Preventing undesirable behavior of intelligent machines.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Preventing undesirable behavior of intelligent machines

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.327648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.763641Z digest=sha256:a76e076826c3c350c4eba7f3b7224e0a5fdc8602dbeb27048a3b24e210fa7fe4

Observation 64dc4375-c341-49e2-8d1f-aa55230916cd · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.767276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.767276Z digest=sha256:b1e036e72d25b171972f7f3049772277194350e49d6172cd4c57923b35c4c578

Observation c6401466-52e9-49e9-8e0b-1907d3b73769 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.770901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.770901Z digest=sha256:644d6b6c1a3e3b5d1bd1e8778017458dc96c6d3ab97b7a7e6d98f5a9a6e5d66d

Observation ce9dc7c0-7761-480a-a53a-8b110137fd31 · outbound

This paper cites Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.316321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.774310Z digest=sha256:affa4b46562ead129c4885b0c4fdf71164b7ec41a77785879f980be3bbec0d9e

Observation 46904a93-d126-448b-ac6c-959027d47ce8 · outbound

This paper cites Enforcing Delayed-Impact Fairness Guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enforcing Delayed-Impact Fairness Guarantees

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.932492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.778283Z digest=sha256:d74d3a6043a419ba331b5824c2199343a7aea650ed0391f42082007cc5aa3db4

Observation 5b6357af-c971-41de-9d6a-10a85f23d23e · outbound

This paper cites Ethical and social risks of harm from Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical and social risks of harm from Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.781959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.781959Z digest=sha256:1885067a3afa4f020999bf809b74f396cf9ac7fddae1270e1aa186347df1486e

Observation 367a0044-733d-4030-86de-d0ac856dc876 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.785508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.785508Z digest=sha256:538bf2c3405eb6ca539b9cd7b7f01c2a3ed1f339fcb1fdc36464463212c98bde

Observation 17688df0-e3b8-42e0-91ee-ffcc0610464f · outbound

This paper cites Recipes for Safety in Open-domain Chatbots.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Recipes for Safety in Open-domain Chatbots

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.788856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.788856Z digest=sha256:e8801a3b9b28cfe721be3398d17a7d8557826c7389d020e10e60aa7e7964764b

Observation 215d5810-a9fe-434f-b1cf-63c3a55c3b00 · outbound

This paper cites Qwen2 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.792512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.792512Z digest=sha256:b244223a1e8009ecda8a9a20289e7c4d51682b4c390544d9b82d469b262818be

Observation 3e7ecda3-d383-418a-9fd8-e16ffa389a83 · outbound

This paper cites A large language model for electronic health records.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints A large language model for electronic health records

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.298310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.796231Z digest=sha256:03f8aa1cddfc4949d6836b29a2d7162411166972a70cbbfa5dab1fe9c8b07d39

Observation f8037334-eadc-4a1f-8d46-ec3f0acf2601 · outbound

This paper cites Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.886951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.799911Z digest=sha256:0edb5aa23707d4af633b90c6e8a009a4ae55b739f8bf062e0085df39a06bcfc6

Observation 373be2dd-3351-4c21-85a9-93d010d415bf · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.803923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.803923Z digest=sha256:ac46e751946860fe23ff55119f3df72f805e785d19bb687d0523bd4ba98e7cef

Observation 4e532ecc-f56f-4b18-b185-9ca210e2b8c5 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Secrets of RLHF in Large Language Models Part I: PPO

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.808116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.808116Z digest=sha256:ba1908a9b346c596da674d4980064a9ccbfcb705b610e0ca2d1417e4ed93d184

Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.812008Z digest=sha256:a1e37ede50a167927623d39d2062d0a7763df68967636a136273459d501fe9b7

Observation f98a0e96-415d-4e86-9bf3-5e8cacb70470 · outbound

This paper cites write newline.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.815824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.815824Z digest=sha256:9aa9b8ff9634c16db27896b82b6a3f4aa8f5deba1c2051b613730edd5af5f4ec

Pith citing papers

Observation 32637181-00ce-4796-a516-eaa778829233 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:19.102106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:19.102106Z digest=sha256:bd90b86b02ee5fe5ecad2d35e3e5ca808582960362e7e62929d2b0e6ebe24ffe

Observation 1437125e-c3ed-4c8c-85f0-a701b5c4ff4e · inbound

Reinforcement Learning from Human Feedback: A Statistical Perspective cites this paper.

Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.609624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T20:10:43.578904Z digest=sha256:2f9fd680fcd795c0797b1646795cdaf87c81062ca095aa9292b6a8ea3c144bb6

Observation 8345e638-7f56-407a-8140-7d5adc8af28e · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.865653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:7d20edcecb94aed98f1245bbcd7db4e97991ebabaa3e306d60c851d759ab5aa7