Pith. sign in

Paper Citation Record · LEDGER

Predictive Divergence Masks for LLM RL

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.10848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10848 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68551c4e-dc65-4aed-bdb1-f151fabf8187 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Predictive Divergence Masks for LLM RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:928e727524c763c662bfe73a37dec981120e67696f9eff1712b7cc3c6e86fcd3

Observation 52ac8a75-1c16-4bfe-a1ca-1fa17084e47c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Predictive Divergence Masks for LLM RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:ecd8d0758c01e7fbdb21e34628a11d757a7508eea1e341b406a5964cd82ac7b4

Observation 906cac7a-da9c-4681-b459-2d42a2f87604 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Predictive Divergence Masks for LLM RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:40102a1f05d90f92b75d0fecb8b9d975cc199049f3e92e76924c2485400d3094

Observation a6deee50-f3d8-42a3-8e35-22a2bb1d132e · outbound

This paper cites https://thinkingmachines.ai/blog/defeating-nondeterminism-in- llm-inference/.

Predictive Divergence Masks for LLM RL https://thinkingmachines.ai/blog/defeating-nondeterminism-in- llm-inference/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b3bbeacff82402926b2920247e01066e417c379ab0883c716a8ecd04b483d7f2

Observation f6c6f435-b655-48f1-9fec-eb2b48ee8b53 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Predictive Divergence Masks for LLM RL Understanding R1-Zero-Like Training: A Critical Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:10d24a17ed73bdf80cf66ea39d1a94b103c5d98e65b85fa4fa0315de274f1989

Observation 0eb8cef3-6fd1-42ae-b224-2feaf23564c4 · outbound

This paper cites Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning.

Predictive Divergence Masks for LLM RL Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:79b47dd60009043b5519ae2a265727567ba273b6dabd0c8a943975be7f249e57

Observation 70c6814d-7b83-431a-8c00-a43aeddf6ae4 · outbound

This paper cites Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

Predictive Divergence Masks for LLM RL Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:ee8480884d19e80b136d567b6551e7273ca8466acd80b00582b1a204b4d40a11

Observation f0215ada-3e65-4c29-8cd0-763b780cc13d · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

Predictive Divergence Masks for LLM RL Rethinking the Trust Region in LLM Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:1956ae8f528e1cd479ba1a07ef69f3e7c2314ec1b16bb2355151a422a7104d24

Observation a31a58da-a056-41e9-8ed3-7d959e170055 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Predictive Divergence Masks for LLM RL Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:5bba2c8bda09d8ad5f985d336d67097321254889588a794fdbc24e93d3d4fbda

Observation 9b62590e-aeb5-47dc-8912-7f00e983fdd1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Predictive Divergence Masks for LLM RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:799c7304a141c7283c98e936137350b74d514efa9abf4790d06099f799f424d1

Observation 5423e8eb-5d03-4520-bb39-c284be433c82 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Predictive Divergence Masks for LLM RL HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:27605ef3fc2c4e3774607d5c6da2a60db903d39aac45fed4a82bffe76b0d7b64

Observation 7fa69b9c-a5e3-4df8-9ea9-5222a7380d00 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Predictive Divergence Masks for LLM RL Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:33116387f940f58f2c9a63d1cf87f48655c3331bfed521406931fc88433493c7

Observation d28266e3-e8ba-4885-91e6-b672dc761455 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Predictive Divergence Masks for LLM RL Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:2a18ed29276a714d342f80653641842098fe0c91effc6d4ecc109fcdf5b86dd5

Observation ff94ef2b-28cd-49bc-9af9-e63f2b4be31b · outbound

This paper cites Simple Policy Optimization.

Predictive Divergence Masks for LLM RL Simple Policy Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:a9d13ff1808c92a893a2cc841d7239bf11760b324b4da10ca0a06401ce84b3af

Observation 45e72984-97cd-420a-8b54-9113b30ba513 · outbound

This paper cites Qwen3 Technical Report.

Predictive Divergence Masks for LLM RL Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:bdc827ffd20dda98bc722c5c9df3b226f5989976a32c84b279def82bf3f6069f

Observation 5861436c-563c-4e29-9d3e-c224264a6331 · outbound

This paper cites Rethinking the Divergence Regularization in LLM RL.

Predictive Divergence Masks for LLM RL Rethinking the Divergence Regularization in LLM RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:3ed1c0da6884f53e12e0156e5e7d00407fc6e2c1ccd4b792775c59b767e9a368

Observation 335fd98d-8447-453d-adf1-5664ee146f3d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Predictive Divergence Masks for LLM RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b934e43241bbe368e773c4cbcf0f01d766cf1d389bfb074b8572d756b81f5471

Observation a80c74fc-c51f-4766-a16c-9c8c9ee4f54c · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374,.

Predictive Divergence Masks for LLM RL Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:e5132f794a67ebfe2e557b227ddd56a6b67b40fd1afca281b3b429ce49f5b0de

Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:488df3dfec0f4c9683d06c67f418ced5774f6db53c0231d982f84d934fc8ed10

Observation d701038e-e666-4992-8d4a-dd92dcb8800c · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:59834675d5b016292a5f6750a9445b44eeca617a3319f3f9c9e98beb5e09872c

Observation 722c673d-87a4-41cb-9887-8a4da861ba9a · outbound

This paper cites GRPO (Shao et al.,.

Predictive Divergence Masks for LLM RL GRPO (Shao et al.,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:67a66afe16e69c45097bd9cf8ec5726584221614a2708cf9cba5aec1000a0152

Observation 6cfebd66-a1dd-4f1a-b6ab-4a7422b3d5ea · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:a64edb073160121da9c796e2cb7a8f97d123ac3526f42ee295eb58aa9491819f

Observation 7dd4f3d8-062a-4429-92f4-e9cca86232e2 · outbound

This paper cites These methods modify the objective or the clipping rule, but they still decide the update direction from the sampled importance ratio.

Predictive Divergence Masks for LLM RL These methods modify the objective or the clipping rule, but they still decide the update direction from the sampled importance ratio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:ba9f629b794ddf44b8b9025ccd7454c50655f767bcaf106f5f45c44e805aa05b

Observation 0bb1750d-f617-4116-8345-fc7dadec84cf · outbound

This paper cites (5) that we build on.

Predictive Divergence Masks for LLM RL (5) that we build on

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:b39467695df54f016bd3f28b09c0bb99d60311f49f55d0384573dfd55f0fd713

Observation 38072ebf-d0b6-44c5-bc05-9d4b3b38beb3 · outbound

This paper cites Aggregated-tail estimator.Under the top-K aggregated-tail construction, the support contains the retained tokensK and one tail bucket.

Predictive Divergence Masks for LLM RL Aggregated-tail estimator.Under the top-K aggregated-tail construction, the support contains the retained tokensK and one tail bucket

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:9373fa57462f89f0db5f3f7d2f514c8ab5b44bad04d06712019128e538d52f2f

Observation 7ee75953-0c84-48f1-b679-d164f04cf47b · outbound

This paper cites By default, both training and rollout use BF16.

Predictive Divergence Masks for LLM RL By default, both training and rollout use BF16

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:a57f5266e7a40e70c83d7c158fac2460631a9238dd9eaafd7be961f11c5653d1

Observation a6836379-670c-4122-b8d1-2462534f9de2 · outbound

This paper cites an unresolved cited work.

Predictive Divergence Masks for LLM RL Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T08:49:37.615924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:49:37.615924Z digest=sha256:42606cc83862c8a33d0ed99be43a0a120b721ebe54d21b6332274f239f6850f1

Pith citing papers

No inbound Pith citation observations are available.