Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:06:27.935915Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2411.16502.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:06:27.935915Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:05:09.068412Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T05:17:05.898300Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4260fdf5-defc-4136-ba5e-e23247df9345 · outbound
Interpreting Language Reward Models via Contrastive Explanations GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8407de-de5f-4ac3-b3d0-59503440c3a6 · outbound
Interpreting Language Reward Models via Contrastive Explanations Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf93350a-3f2f-4db3-a1c9-46396bea7ff0 · outbound
Interpreting Language Reward Models via Contrastive Explanations LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7ab228-acbd-4c88-948b-60baf42d76f6 · outbound
Interpreting Language Reward Models via Contrastive Explanations UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ef8185-27d1-4a00-8b97-20b8f0e31e14 · outbound
Interpreting Language Reward Models via Contrastive Explanations Research agenda for sociotechnical approaches to ai safety
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78c0272d-2e56-4c3c-b869-616027a62d37 · outbound
Interpreting Language Reward Models via Contrastive Explanations RewardBench: Evaluating Reward Models for Language Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806e2465-5eb5-476e-bc01-c1b1365b123a · outbound
Interpreting Language Reward Models via Contrastive Explanations Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e709d8-f3bf-411e-bd0d-f11e9e3d3622 · outbound
Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 799ea82e-acc6-4756-a6cd-24210f785702 · outbound
Interpreting Language Reward Models via Contrastive Explanations A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda6611a-462d-4a5e-a4f9-bfdf8590a683 · outbound
Interpreting Language Reward Models via Contrastive Explanations A Long Way to Go: Investigating Length Correlations in RLHF
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce12127a-aba2-41c0-aedc-27fc53cc946c · outbound
Interpreting Language Reward Models via Contrastive Explanations HelpSteer2: Open-source dataset for training top-performing reward models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b60194-1e5c-4833-91d6-83c0220a4464 · outbound
Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 66cf38e7-971b-45ae-98a9-ca71a2643add · outbound
Interpreting Language Reward Models via Contrastive Explanations As the former is not the focus of the datasets we experiment with, we only additionally include honesty, relabelled to avoid-to-answer for better relevance
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bad868db-86aa-44ce-b783-ccd53aae4468 · outbound
Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8ba8052-758f-4831-9ab8-a62d6d24fecc · outbound
Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eef2f67f-fc87-40da-a019-8cb8cd92a048 · outbound
Interpreting Language Reward Models via Contrastive Explanations OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db6ba71-ca46-4a62-bae2-e4c9d57fd807 · outbound
Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 179382e3-09a0-4fee-93bb-aa3b7174946c · outbound
Interpreting Language Reward Models via Contrastive Explanations Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976b738c-f2fe-4636-bcb8-7c61adfba7ac · outbound
Interpreting Language Reward Models via Contrastive Explanations LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5cada5c-4f92-4315-88a9-e02035c7ed41 · outbound
Interpreting Language Reward Models via Contrastive Explanations Sentence-bert: Sentence embeddings using siamese bert- networks
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd32ea03-4936-41ca-8d3c-19452fd37c60 · outbound
Interpreting Language Reward Models via Contrastive Explanations Vera Liao, Rania Abdelghani, and Pierre-Yves Oudeyer
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f241e812-535e-4449-89cb-9b246075f74b · outbound
Interpreting Language Reward Models via Contrastive Explanations Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd0785c-10d6-4195-afdc-11c6e3de6449 · outbound
Interpreting Language Reward Models via Contrastive Explanations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1581f68e-0b0f-4f49-9b4a-763df987c986 · outbound
Interpreting Language Reward Models via Contrastive Explanations RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb42c8bb-e6f5-46c5-bd0c-475a44b6d185 · inbound
Multi-Domain Explainability of Preferences Interpreting Language Reward Models via Contrastive Explanations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a250044e-9fae-47e7-a728-f8814d8e2bf9 · inbound
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Interpreting Language Reward Models via Contrastive Explanations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.