Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:59:38.022401Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2602.15206.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:59:38.022401Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c23f11f1-2fd9-4976-8cd6-1aa4647a3a10 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1ee73f-9612-41ad-ac88-2d0b75328dbb · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Decoupled Weight Decay Regularization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f846c9cd-60c5-40d8-9c43-cb51d9fb91f2 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference URL https://doi.org/10.1145/3623384
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268e2b95-fc1e-44cd-a3ed-ac98b11670dc · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ecc515b6-c2dd-428d-85e2-c4b2f8ccd695 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Vivek Myers, Erdem Bıyık, Nima Anari, and Dorsa Sadigh
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35cae80d-7fd8-4431-8600-5ccbf5b974b7 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Training language models to follow instructions with human feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01800d95-c7cd-43c8-8e6c-b74caafb3306 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2246660-9ab4-48c7-9e67-7fdeef9e3261 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference 13 A Implementation Details A.1 Model Architecture The reward encoder qθ and Q-value estimator Qϕ were implemented as two-layer MLPs with Leaky ReLU activations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c228a7f-f9e2-4010-9580-9aea33351a72 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Shaunak A
Reference 1980
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06f97f8-0610-4967-b1b9-6aab1ab2e25b · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference ISBN 978-1-58113-838-2
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8244124f-9067-429d-b58d-6bfdd43df72a · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference ISBN 978-1-60558-658-8
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57aefb9-b22b-45d8-a9d1-09c2763386ab · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Auto-Encoding Variational Bayes
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40ae02c-c435-419f-a524-3fa31e9516e3 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Unsolved Problems in ML Safety
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f29625-6a8f-4974-87b8-5875edba4c1e · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Reward-rational (implicit) choice: A unifying formalism for reward learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1e099b-a8b8-4ce1-870c-a28b782a31ad · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Learning Multimodal Rewards from Rankings
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c9328504-8c06-47f4-9c62-3db756a303e8 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference doi: 10.1177/02783649211041652
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb5badb-d0db-496c-8459-160db2383dc8 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference doi: 10.1609/aaai.v37i5.25740
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f93f15a9-5612-410e-a186-8e7972475d65 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6b0a61-a1ed-4f96-ac42-0b337c3049d6 · outbound
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference Scalable Bayesian inverse reinforcement learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.