Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.10848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 68551c4e-dc65-4aed-bdb1-f151fabf8187 · outbound
Predictive Divergence Masks for LLM RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ac8a75-1c16-4bfe-a1ca-1fa17084e47c · outbound
Predictive Divergence Masks for LLM RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906cac7a-da9c-4681-b459-2d42a2f87604 · outbound
Predictive Divergence Masks for LLM RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6deee50-f3d8-42a3-8e35-22a2bb1d132e · outbound
Predictive Divergence Masks for LLM RL https://thinkingmachines.ai/blog/defeating-nondeterminism-in- llm-inference/
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c6f435-b655-48f1-9fec-eb2b48ee8b53 · outbound
Predictive Divergence Masks for LLM RL Understanding R1-Zero-Like Training: A Critical Perspective
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb8cef3-6fd1-42ae-b224-2feaf23564c4 · outbound
Predictive Divergence Masks for LLM RL Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c6814d-7b83-431a-8c00-a43aeddf6ae4 · outbound
Predictive Divergence Masks for LLM RL Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0215ada-3e65-4c29-8cd0-763b780cc13d · outbound
Predictive Divergence Masks for LLM RL Rethinking the Trust Region in LLM Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31a58da-a056-41e9-8ed3-7d959e170055 · outbound
Predictive Divergence Masks for LLM RL Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b62590e-aeb5-47dc-8912-7f00e983fdd1 · outbound
Predictive Divergence Masks for LLM RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5423e8eb-5d03-4520-bb39-c284be433c82 · outbound
Predictive Divergence Masks for LLM RL HybridFlow: A Flexible and Efficient RLHF Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa69b9c-a5e3-4df8-9ea9-5222a7380d00 · outbound
Predictive Divergence Masks for LLM RL Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28266e3-e8ba-4885-91e6-b672dc761455 · outbound
Predictive Divergence Masks for LLM RL Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff94ef2b-28cd-49bc-9af9-e63f2b4be31b · outbound
Predictive Divergence Masks for LLM RL Simple Policy Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e72984-97cd-420a-8b54-9113b30ba513 · outbound
Predictive Divergence Masks for LLM RL Qwen3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5861436c-563c-4e29-9d3e-c224264a6331 · outbound
Predictive Divergence Masks for LLM RL Rethinking the Divergence Regularization in LLM RL
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335fd98d-8447-453d-adf1-5664ee146f3d · outbound
Predictive Divergence Masks for LLM RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80c74fc-c51f-4766-a16c-9c8c9ee4f54c · outbound
Predictive Divergence Masks for LLM RL Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · outbound
Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d701038e-e666-4992-8d4a-dd92dcb8800c · outbound
Predictive Divergence Masks for LLM RL Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 722c673d-87a4-41cb-9887-8a4da861ba9a · outbound
Predictive Divergence Masks for LLM RL GRPO (Shao et al.,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfebd66-a1dd-4f1a-b6ab-4a7422b3d5ea · outbound
Predictive Divergence Masks for LLM RL Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd4f3d8-062a-4429-92f4-e9cca86232e2 · outbound
Predictive Divergence Masks for LLM RL These methods modify the objective or the clipping rule, but they still decide the update direction from the sampled importance ratio
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb1750d-f617-4116-8345-fc7dadec84cf · outbound
Predictive Divergence Masks for LLM RL (5) that we build on
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38072ebf-d0b6-44c5-bc05-9d4b3b38beb3 · outbound
Predictive Divergence Masks for LLM RL Aggregated-tail estimator.Under the top-K aggregated-tail construction, the support contains the retained tokensK and one tail bucket
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee75953-0c84-48f1-b679-d164f04cf47b · outbound
Predictive Divergence Masks for LLM RL By default, both training and rollout use BF16
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6836379-670c-4122-b8d1-2462534f9de2 · outbound
Predictive Divergence Masks for LLM RL Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.