Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:38:43.203572Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.13998.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:38:43.203572Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.399277Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:34:26.163405Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43689ceb-c541-4db3-b210-86441f99eeb6 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250908ab-3120-4f44-8c85-56326a1db72a · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1e4c42-4b68-4d93-ade1-8f6705b801c9 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26b12ce-55ee-4cc0-ac18-8e93e5019c61 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Pareto-Optimal Learning from Preferences with Hidden Context
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c9dcdbbf-3471-4d72-9556-16611bf5b0a1 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c046b64a-76df-4ac3-becd-caba3cc8f1f9 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes V., and Syrgkanis, V
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a85868-da9c-4cbf-b000-666680618db7 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9966b605-cc5a-40fd-a080-678a3f0bdd6d · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes W., Rezende, D., and Eslami, S
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation deb0e917-eefd-4d4f-bc5a-1d50bafec7a7 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Neural Processes
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528954e4-c12f-4004-b160-bd3384c4d079 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba85029f-7f41-4750-bb58-e62d39f7e4e9 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes LoRA: Low-Rank Adaptation of Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb0deea-a249-478e-8cc3-65d4c56f5484 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9259effb-283c-4e17-a72e-1a043393b3f2 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Preference Transformer : Modeling Human Preferences using Transformers for RL
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 649ec6cd-7114-4f2c-893e-5699022ca650 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Adam: A Method for Stochastic Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e15f9e-bd09-49f0-92b1-4f4d3adfcc70 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Training language models to follow instructions with human feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39d24ee-4f7c-4713-96aa-8e4d92794d4b · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Automatic differentiation in pytorch
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 640a6259-da4a-4a4d-ad6f-e33fabf67bcf · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes FiLM: Visual Reasoning with a General Conditioning Layer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 025fb091-50e1-4470-a2b2-e333053699b1 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aca3f66-cddd-4ac6-95f2-b1d7bc3ba8f0 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381a1218-56d0-47fc-bfa7-658bb07a7b81 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162becfb-7cf4-4aae-9f36-4588d7d45ddf · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes DISTRIBUTIONAL PREFERENCE LEARNING : UNDERSTANDING AND ACCOUNTING FOR HIDDEN CONTEXT IN RLHF
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 07f9abf7-96ec-4007-a446-4ed10da6cfdc · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes A Roadmap to Pluralistic Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27b4d01-c81a-416c-b870-d071856d62b4 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation da1bb913-e86b-4ec3-a51c-8e5b0c49a762 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7225a9a4-7606-4b12-81ed-c638cc3613ac · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286c9f84-ce67-4444-8e5b-41ba393b1719 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7dc0ca9-b69e-47a1-b330-91bd90b61821 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Deep Sets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd68763f-dc41-4f9f-aaed-8695034a5913 · outbound
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Panacea: Pareto Alignment via Preference Adaptation for LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acdb000-0686-4932-9d12-8726a2b6d43c · inbound
Activation Reward Models for Few-Shot Model Alignment Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e12ffca-9b85-4c25-ba3e-61cd0bbcb210 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.