Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2210.01241.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.175303Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T00:46:56.861091Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 60163b01-834e-4577-85aa-4a58338630d1 · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30583c24-e39b-4dfe-aa0d-406303ddb552 · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ed6cbbe-99a6-41cd-ae6a-9d427b1592b3 · inbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4315a4d-6138-4a5e-a77a-56c8f4f8748a · inbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5f0ec8-367f-4403-92b0-338e21d2aa72 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 217
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58995126-1e7c-4486-8e9b-dd9328fd994d · inbound
Red-Teaming Vision-Language-Action Models via Quality Diversity Prompt Generation for Robust Robot Policies Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62096ec7-e79e-4dbc-a449-9330c3fc8199 · inbound
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6935e888-670c-40ab-bb41-51594d2b5b4a · inbound
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0a86a7a-ce88-4642-b809-653e5a94be19 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a17cbd27-48fa-4f76-94d0-784c84a0ef1f · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 289
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa3f6ec-309c-4f64-95e9-78ba1ad164c0 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 290
Source-reported events for the cited work
Unavailable: canonical work link unavailable.