Pith. sign in

Paper Citation Record · LEDGER

Interpreting Language Reward Models via Contrastive Explanations

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2411.16502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16502 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:06:27.935915Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:05:09.068412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T05:17:05.898300Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4260fdf5-defc-4136-ba5e-e23247df9345 · outbound

This paper cites GPT-4 Technical Report.

Interpreting Language Reward Models via Contrastive Explanations GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.823032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.823032Z digest=sha256:1d04fef3e468cb489166fee271122eb90caeff6cad9842cae895924381765e9a

Observation 2a8407de-de5f-4ac3-b3d0-59503440c3a6 · outbound

This paper cites Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis.

Interpreting Language Reward Models via Contrastive Explanations Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.844021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.844021Z digest=sha256:f08f4852e0946794bda651cd6a9337b5844efe9a37544e7217577a5c210f2919

Observation bf93350a-3f2f-4db3-a1c9-46396bea7ff0 · outbound

This paper cites LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding.

Interpreting Language Reward Models via Contrastive Explanations LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.849086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.849086Z digest=sha256:ad69981f937385d07a37186b606754fdca59ba7d4c57f465392b1f63436a210c

Observation 1f7ab228-acbd-4c88-948b-60baf42d76f6 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Interpreting Language Reward Models via Contrastive Explanations UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.854568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.854568Z digest=sha256:4672039f70d3f0ae86a53c5c82292cf71eed30592d03dfae32bf87b8c3d8ac81

Observation e0ef8185-27d1-4a00-8b97-20b8f0e31e14 · outbound

This paper cites Research agenda for sociotechnical approaches to ai safety.

Interpreting Language Reward Models via Contrastive Explanations Research agenda for sociotechnical approaches to ai safety

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:06:28.329790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.859612Z digest=sha256:93dbf45e1c43e7814d539c3cbb2fd5ce421f90b0f4781b25062b2be58e19c01b

Observation 78c0272d-2e56-4c3c-b869-616027a62d37 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Interpreting Language Reward Models via Contrastive Explanations RewardBench: Evaluating Reward Models for Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.864403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.864403Z digest=sha256:101cb76ed2f467f810b8a0a1c6bfab6bb0f9f151c115cde98ccd78f2208bd97a

Observation 806e2465-5eb5-476e-bc01-c1b1365b123a · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Interpreting Language Reward Models via Contrastive Explanations Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.869259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.869259Z digest=sha256:96910bdef3c0e25b1010f6ffe4043d38a31b9bc75a65fa229fb9ab2d1be4e55b

Observation 54e709d8-f3bf-411e-bd0d-f11e9e3d3622 · outbound

This paper cites an unresolved cited work.

Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:06:28.315477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.878777Z digest=sha256:0f2c44e9e34e72967719fb0ef46b4d216a4e5474524cbafe7f0eb83cf17e0f70

Observation 799ea82e-acc6-4756-a6cd-24210f785702 · outbound

This paper cites A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift.

Interpreting Language Reward Models via Contrastive Explanations A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.888214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.888214Z digest=sha256:03761357ca5c89e734d741e04105c04eef6b1871b24125e434ef647f602c9e57

Observation bda6611a-462d-4a5e-a4f9-bfdf8590a683 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Interpreting Language Reward Models via Contrastive Explanations A Long Way to Go: Investigating Length Correlations in RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.901889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.901889Z digest=sha256:442d06ec3231cf65c0e814ecfcdd351a9b7e5e10479c83631fd7eeec1fe947ef

Observation ce12127a-aba2-41c0-aedc-27fc53cc946c · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Interpreting Language Reward Models via Contrastive Explanations HelpSteer2: Open-source dataset for training top-performing reward models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.910817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.910817Z digest=sha256:a62e7cd8b5dd7b85649e03b80bc3d3f8592ae99d47dbe96d923a709b01041405

Observation e6b60194-1e5c-4833-91d6-83c0220a4464 · outbound

This paper cites an unresolved cited work.

Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:06:28.254424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.920396Z digest=sha256:fa1001d635d551e6e722c89a9cce614f6c088e4a85c002b91ad8919bb3fca54d

Observation 66cf38e7-971b-45ae-98a9-ca71a2643add · outbound

This paper cites As the former is not the focus of the datasets we experiment with, we only additionally include honesty, relabelled to avoid-to-answer for better relevance.

Interpreting Language Reward Models via Contrastive Explanations As the former is not the focus of the datasets we experiment with, we only additionally include honesty, relabelled to avoid-to-answer for better relevance

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:06:28.239664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.925624Z digest=sha256:4ca8a79896a7df35c06d91de1242657400efcf5075cc187558ca031ea39021d9

Observation bad868db-86aa-44ce-b783-ccd53aae4468 · outbound

This paper cites an unresolved cited work.

Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:06:28.224411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.931064Z digest=sha256:df6f0076410d645b8acdac9ea4020df024e02721a8d0b34c7b9b92d5622645a6

Observation f8ba8052-758f-4831-9ab8-a62d6d24fecc · outbound

This paper cites an unresolved cited work.

Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:06:28.209311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.935915Z digest=sha256:9b46b87800d8dc1b33d957525434702a14875f5dfe6ed0465b53f00a4511b3e0

Observation eef2f67f-fc87-40da-a019-8cb8cd92a048 · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

Interpreting Language Reward Models via Contrastive Explanations OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.883420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.883420Z digest=sha256:e5796339422c46f2d23ea3a5f9cedc0668bbb02783b0ea7f941e98fb26fb900a

Observation 0db6ba71-ca46-4a62-bae2-e4c9d57fd807 · outbound

This paper cites an unresolved cited work.

Interpreting Language Reward Models via Contrastive Explanations Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:06:28.285887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.897374Z digest=sha256:be83a0c660c8791f01d2625a35c9d010200638ff44426241ec21920505f998d2

Observation 179382e3-09a0-4fee-93bb-aa3b7174946c · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Interpreting Language Reward Models via Contrastive Explanations Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.906346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.906346Z digest=sha256:5191c47c784e327d744150bf2ace7c83b08a6d2e26d74d6da7fa11a2903f391a

Observation 976b738c-f2fe-4636-bcb8-7c61adfba7ac · outbound

This paper cites LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study.

Interpreting Language Reward Models via Contrastive Explanations LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.873949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.873949Z digest=sha256:083787410e8df902e28d2be376c72eab7426ca78e3ec285f4bef118538726716

Observation c5cada5c-4f92-4315-88a9-e02035c7ed41 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

Interpreting Language Reward Models via Contrastive Explanations Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:06:28.301157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.892865Z digest=sha256:f23785959822891d373b3002f89acd95f1da5a686b88cddac94c6e1f0afd1bce

Observation dd32ea03-4936-41ca-8d3c-19452fd37c60 · outbound

This paper cites Vera Liao, Rania Abdelghani, and Pierre-Yves Oudeyer.

Interpreting Language Reward Models via Contrastive Explanations Vera Liao, Rania Abdelghani, and Pierre-Yves Oudeyer

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:06:28.269477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:06:27.915793Z digest=sha256:5e89dfd532729a7151229912231d464c77d2c33f317fe1364337a472ae1d8d29

Observation f241e812-535e-4449-89cb-9b246075f74b · outbound

This paper cites Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation.

Interpreting Language Reward Models via Contrastive Explanations Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.833762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.833762Z digest=sha256:5ec50d9d05a9cc7977bcadb5aa54709ac12b767d79156c44670ef582b6829713

Observation cfd0785c-10d6-4195-afdc-11c6e3de6449 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Interpreting Language Reward Models via Contrastive Explanations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.828613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.828613Z digest=sha256:ca99ac74f64608511f7a440f662d17d2419216fa29b257f2f4b46057b31df772

Observation 1581f68e-0b0f-4f49-9b4a-763df987c986 · outbound

This paper cites RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs.

Interpreting Language Reward Models via Contrastive Explanations RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.839076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.839076Z digest=sha256:f5ec04a4264d453663751e82890f41b9ecf825374563ea06cfc31e798419e7f4

Pith citing papers

Observation eb42c8bb-e6f5-46c5-bd0c-475a44b6d185 · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences Interpreting Language Reward Models via Contrastive Explanations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.068412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:09.068412Z digest=sha256:2349a63f16d8d8d01f5354c2e3cee1b6444eaf4fe2fb53fee09b3afea6c1e071

Observation a250044e-9fae-47e7-a728-f8814d8e2bf9 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Interpreting Language Reward Models via Contrastive Explanations

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.901461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:8ce00b438ba5f1395266015b136492aad215f26ed617aa6b60549a37e96deb09