Pith. sign in

Paper Citation Record · LEDGER

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.19567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19567 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:44:39.682811Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cf30fcf-e54a-45fa-9ab9-c9fea9ad904a · outbound

This paper cites Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T15:44:40.886960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:38.297770Z digest=sha256:0fd4242db41aedb105a5a257f52aa43ac0cb39cc16964bc10585568373a882eb

Observation 0054fdc0-6a7b-4d2a-954e-9f7e225cbcba · outbound

This paper cites Bubbling of K\"ahler-Einstein metrics.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Bubbling of K\"ahler-Einstein metrics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.462427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.462427Z digest=sha256:892cca30bec4d944321c41f9336d2c38d083fbc5a9c2dc9244b59e8c57a1a3f7

Observation 6687d88d-2e56-49f3-ac17-78c763673605 · outbound

This paper cites Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:44:40.151591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:38.520437Z digest=sha256:c02c830e20afe8adc6cb8865fc83af78ff52dfa8115cd456c9eeef10ca7ab903

Observation 7c1c29ba-f599-4abb-8e1f-7192e8ac808e · outbound

This paper cites On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.622610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.622610Z digest=sha256:b024a6e6c55ecc348ff54a6fccf12ca3b9ae7da302f08a1cd29f8f9e8fd8d341

Observation 93716715-9978-42be-a103-bf34678de6da · outbound

This paper cites ACL Long Paper (2025).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning ACL Long Paper (2025)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:44:41.374021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:38.744210Z digest=sha256:c0db2315867c3f36ea5735e428eece192f26da464eddfe134ea692742083af54

Observation 7c5af0fe-90f3-4854-93b8-aefbc3283826 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Axioms for AI Alignment from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.879132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.879132Z digest=sha256:d42674b0f80c5ff7312393d4e33527dd9c867fcbeed747489cc4fcdcc5ff99ea

Observation 4b4df1d4-e9d8-435f-ae1b-2cbf8d1d07bf · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:38.975171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:38.975171Z digest=sha256:0a7cfb09fd644380e095db8d09e1ba881c5a35045454f0ae0a1a8d5106e9ada9

Observation 923a96db-b7f5-4e3f-8c12-29fc44fb86ba · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.099633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.099633Z digest=sha256:5344ca0fbc13f9cbe2e644ffe0e2eb3c19abcbb3cb449b29a771cddff43b5239

Observation 79ff5b0a-2487-4a5b-988f-5aaa386772ad · outbound

This paper cites IJCAI (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning IJCAI (2024)

Reference 9

Resolution
verified exact
doi, observed 2026-08-05T15:44:39.890905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:39.196330Z digest=sha256:c221ddf58dafd5c7b57ce3bcb1c49e6b385650c4bede0458e562b61cdf00ec4f

Observation a68d5bed-fb1d-4a56-9f91-be3f92577fb8 · outbound

This paper cites F AccT (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning F AccT (2024)

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T15:44:40.665521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:39.300014Z digest=sha256:7b8518d9d20cab9d2ff35b6ac34ae1f6202b3f76f7b438c34258956b1f299d19

Observation a4af2564-5cdd-4f01-a626-dd9f741fb5ef · outbound

This paper cites NeurIPS Workshops (2024).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning NeurIPS Workshops (2024)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.360105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.360105Z digest=sha256:1e571884003115794a2d338bd97adbd11695e52e1624b04a7954adc8ce96b79c

Observation 3d59fb5e-1a28-4c03-9309-037afe1233a1 · outbound

This paper cites Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T15:44:40.354785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:39.466931Z digest=sha256:5a6303150f7260235e12e682df5536aef4fa8cffb7c39f4f41e8466157354bf8

Observation f32b68ef-3f2b-4d9b-89dd-26e0656748e7 · outbound

This paper cites arXiv preprint (2023).

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning arXiv preprint (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:44:41.127320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:44:39.567689Z digest=sha256:f6208bc035b400ef694b74e8ac0fb7ec64eccbc61f91f0c18f046127a1adb6ee

Observation e1b5eae3-ff19-4adf-a37f-730502636d27 · outbound

This paper cites The doubly librating Plutinos.

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning The doubly librating Plutinos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:39.682811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:39.682811Z digest=sha256:d5425243ece0a872a88a641d484cd54d68e7ccb1e911999237383f91dff31c75

Pith citing papers

No inbound Pith citation observations are available.