Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:54:05.242057Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2501.06911.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:54:05.242057Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3cfd475d-d010-4627-b1c5-cf309afa68a1 · outbound
Risk-Averse Finetuning of Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c731d022-e0b5-4ab2-af41-1e260a7b6081 · outbound
Risk-Averse Finetuning of Large Language Models Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a00f026-83bf-4318-82e2-df7b8510debd · outbound
Risk-Averse Finetuning of Large Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee09e131-6f1e-4b0b-93e2-907218c1c4d8 · outbound
Risk-Averse Finetuning of Large Language Models Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd73537-6261-4a51-8ceb-ce1b2d892da4 · outbound
Risk-Averse Finetuning of Large Language Models WHEN I AM UNBLOCKED I SWEAR I WILL GO F**K YOUR M C
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45be7a25-cdac-471f-b276-17624c861a8d · outbound
Risk-Averse Finetuning of Large Language Models DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdee585-fe79-4fe5-9c7c-71cf1f75512e · outbound
Risk-Averse Finetuning of Large Language Models EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381e7469-1002-47d5-870c-8417caabbeba · outbound
Risk-Averse Finetuning of Large Language Models The Woman Worked as a Babysitter: On Biases in Language Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f092efe-31a3-4711-b7da-334899d6cfbe · outbound
Risk-Averse Finetuning of Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ecb8305-c134-450f-a5f8-5a22cb0f3d93 · outbound
Risk-Averse Finetuning of Large Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274f79a1-7817-4abc-971e-aded10f64a4d · outbound
Risk-Averse Finetuning of Large Language Models Ethical and social risks of harm from Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b72b7d-80c9-4cd4-bb81-ddc0e655768b · outbound
Risk-Averse Finetuning of Large Language Models Fine-Tuning Language Models from Human Preferences
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3b5f31-4961-4eea-be33-3b00704744f7 · outbound
Risk-Averse Finetuning of Large Language Models In our work, we primarily focussed on generative tasks, and not the Question-Answer (Q&A) format
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5df3b17b-9423-4bcb-b6db-115326511e04 · outbound
Risk-Averse Finetuning of Large Language Models Furthermore, even aligned versions of LLMs are not immune to exploitation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74a7bf16-41b5-47ee-98a0-f4b04e0894e9 · outbound
Risk-Averse Finetuning of Large Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66a36157-066c-451a-9918-dd0d8af762c8 · outbound
Risk-Averse Finetuning of Large Language Models Furthermore, even aligned versions of LLMs are not immune to exploitation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1c3e45d6-d485-47d1-be7a-2b5b69d76344 · outbound
Risk-Averse Finetuning of Large Language Models Risk Averseness in RL
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94cd4264-4635-402e-bb16-775317b7b7be · outbound
Risk-Averse Finetuning of Large Language Models The dataset utilized in this task is introduced by Gehman et al
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a81c51f-ad58-4a6c-bccd-ce0535e95d7e · outbound
Risk-Averse Finetuning of Large Language Models We choose to work with regularized reward for two reasons: I
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10bec7ae-4ad5-4f6c-ae8e-0e0cc07abb20 · outbound
Risk-Averse Finetuning of Large Language Models [PAD]", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 100: AddedToken(
Reference 768
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f6894a9-0bec-4174-b2b5-448f8278d076 · outbound
Risk-Averse Finetuning of Large Language Models Language models are few-shot learners
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799074a3-0bed-4808-93aa-00648c9f1c55 · outbound
Risk-Averse Finetuning of Large Language Models Proximal Policy Optimization Algorithms
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c71568-bd16-4288-abaf-efb32a7a799a · outbound
Risk-Averse Finetuning of Large Language Models WebGPT: Browser-assisted question-answering with human feedback
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e66f78b-29f4-4884-9793-12598b607891 · outbound
Risk-Averse Finetuning of Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4b1f60-ffc2-4b51-b397-a44e9bbb9d2d · outbound
Risk-Averse Finetuning of Large Language Models Worst Cases Policy Gradients
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8abc197-a505-4958-bbce-d00cd6bf48bd · outbound
Risk-Averse Finetuning of Large Language Models Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa4f4c4-9db1-4bf7-a623-f4a8a538f8ad · outbound
Risk-Averse Finetuning of Large Language Models GeDi: Generative Discriminator Guided Sequence Generation
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de18ffcd-46cf-4c53-8b3a-d60c18d3f827 · outbound
Risk-Averse Finetuning of Large Language Models Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b19e35a-fec8-443e-acd2-7b43bfcb77db · outbound
Risk-Averse Finetuning of Large Language Models Universal Language Model Fine-tuning for Text Classification
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b67231ad-3244-4c56-be41-254761be35a1 · outbound
Risk-Averse Finetuning of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3607c869-9333-4d05-a8d4-8411918d39c6 · outbound
Risk-Averse Finetuning of Large Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b296f1-c8bb-4395-9709-51788a0ce629 · outbound
Risk-Averse Finetuning of Large Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d11885-206f-4c57-956a-741d2a3ce73f · outbound
Risk-Averse Finetuning of Large Language Models Systematic Rectification of Language Models via Dead-end Analysis
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.