Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:44:39.682811Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.19567.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:44:39.682811Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1cf30fcf-e54a-45fa-9ab9-c9fea9ad904a · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0054fdc0-6a7b-4d2a-954e-9f7e225cbcba · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Bubbling of K\"ahler-Einstein metrics
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6687d88d-2e56-49f3-ac17-78c763673605 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c1c29ba-f599-4abb-8e1f-7192e8ac808e · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93716715-9978-42be-a103-bf34678de6da · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning ACL Long Paper (2025)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c5af0fe-90f3-4854-93b8-aefbc3283826 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Axioms for AI Alignment from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4df1d4-e9d8-435f-ae1b-2cbf8d1d07bf · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923a96db-b7f5-4e3f-8c12-29fc44fb86ba · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ff5b0a-2487-4a5b-988f-5aaa386772ad · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning IJCAI (2024)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a68d5bed-fb1d-4a56-9f91-be3f92577fb8 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning F AccT (2024)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4af2564-5cdd-4f01-a626-dd9f741fb5ef · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning NeurIPS Workshops (2024)
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d59fb5e-1a28-4c03-9309-037afe1233a1 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f32b68ef-3f2b-4d9b-89dd-26e0656748e7 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning arXiv preprint (2023)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1b5eae3-ff19-4adf-a37f-730502636d27 · outbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning The doubly librating Plutinos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.