Pith. sign in

Paper Citation Record · LEDGER

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.19567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19567 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:44:39.682811Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cf30fcf-e54a-45fa-9ab9-c9fea9ad904a · outbound

This paper cites Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T15:44:40.886960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:38.297770Z digest=sha256:ec611c5c3f7e7c7d74ce111267bb17d02be53a0d87930d74de14575bfd2e0fb9

Observation 0054fdc0-6a7b-4d2a-954e-9f7e225cbcba · outbound

This paper cites Bubbling of K\"ahler-Einstein metrics.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Bubbling of K\"ahler-Einstein metrics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.462427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.462427Z digest=sha256:9126b18f4ae431a61104410041d331de5bb6a08930181f981e2ff3434bf6c1d2

Observation 6687d88d-2e56-49f3-ac17-78c763673605 · outbound

This paper cites Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:44:40.151591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:38.520437Z digest=sha256:98353950f4652f07e9e4d00d48130b90105f731f26112533260f880826b58f72

Observation 7c1c29ba-f599-4abb-8e1f-7192e8ac808e · outbound

This paper cites On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.622610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.622610Z digest=sha256:03035bfddd721b521cff02bb3ed38a68ec235b6566a5439e4a9e610c7010aefc

Observation 93716715-9978-42be-a103-bf34678de6da · outbound

This paper cites ACL Long Paper (2025).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning ACL Long Paper (2025)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:44:41.374021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:38.744210Z digest=sha256:05c8c06e801b3e8cd61f0b41e43596c30c2229a81712d3b9d54f1a638777f811

Observation 7c5af0fe-90f3-4854-93b8-aefbc3283826 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Axioms for AI Alignment from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.879132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.879132Z digest=sha256:7a6e98bba186a0dea4489a62f1155039cd3e533f14d4ac998558469f18a5a293

Observation 4b4df1d4-e9d8-435f-ae1b-2cbf8d1d07bf · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.975171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.975171Z digest=sha256:bb5487978c0151c87a4d7133d1110919c3bc5b3546d9eb5f0cd70fade1e543aa

Observation 923a96db-b7f5-4e3f-8c12-29fc44fb86ba · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.099633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.099633Z digest=sha256:48577e84bdaba4ebea11b2b64b10060a78d9072dd028858a1541df261071acf2

Observation 79ff5b0a-2487-4a5b-988f-5aaa386772ad · outbound

This paper cites IJCAI (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning IJCAI (2024)

Reference 9

Resolution
verified exact
doi, observed 2026-08-05T15:44:39.890905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:39.196330Z digest=sha256:087fc7bc70bd2b0e9e86b185a0f5613d72b92fb092af6fc3ef9849ace3872386

Observation a68d5bed-fb1d-4a56-9f91-be3f92577fb8 · outbound

This paper cites F AccT (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning F AccT (2024)

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T15:44:40.665521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:39.300014Z digest=sha256:cb7beeac8a7e097797b6b58fa188d30f45a111ad59ce208b7644591983ba168a

Observation a4af2564-5cdd-4f01-a626-dd9f741fb5ef · outbound

This paper cites NeurIPS Workshops (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning NeurIPS Workshops (2024)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.360105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.360105Z digest=sha256:85983cf71a82c3dc52206cba7d231654debbf23bc2ac4fbafaea65557ef1dde2

Observation 3d59fb5e-1a28-4c03-9309-037afe1233a1 · outbound

This paper cites Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T15:44:40.354785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:39.466931Z digest=sha256:6a1b616aef1d7a98069953267b186a84704744c03ff3a0ef00387b2a818cafe1

Observation f32b68ef-3f2b-4d9b-89dd-26e0656748e7 · outbound

This paper cites arXiv preprint (2023).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning arXiv preprint (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:44:41.127320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T15:44:39.567689Z digest=sha256:e078393316f950c4daa9dbc453d2fb1703bbfc9be552b3056d36e2a61a99c7cd

Observation e1b5eae3-ff19-4adf-a37f-730502636d27 · outbound

This paper cites The doubly librating Plutinos.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning The doubly librating Plutinos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.682811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.682811Z digest=sha256:d9c24f139f578ea638120c8f5f3793b29d5491649135685067861f8cfb697638

Pith citing papers

No inbound Pith citation observations are available.