Pith. sign in

Paper Citation Record · LEDGER

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.10481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10481 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.344213Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e4e57c1-b755-49ca-a60a-fa13eeb3b440 · outbound

This paper cites OpenAI o1 System Card.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:23.928177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:23.928177Z digest=sha256:1e5d92ff05946ca78dbe5078bccfc77a6defc0154d124f8c9d6e718d6e9b85f7

Observation 450c0916-6b95-4f93-9157-9fdf49aee130 · outbound

This paper cites Nature , volume =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Nature , volume =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.006590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.006590Z digest=sha256:e8c853cc9f4d92f77670c9cc2383ee32e3931b2422a181a3b07cbd2561b75db4

Observation e10c48c1-b611-4d29-88b3-ecf63eb507da · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.088146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.088146Z digest=sha256:c3283bc3d3fe9f1dc317e244b41f0af9132fe9c264162c2ddb49e00411eb1736

Observation b2340d31-664f-444e-90bb-5f6515277157 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.173789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.173789Z digest=sha256:62f150fae7b367e81ddafd69dd437f5a45705d5673ddf5ef960600d509bde891

Observation 64a91aa6-9641-4b7c-8eed-7f521c90c960 · outbound

This paper cites Qwen3 Technical Report.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Qwen3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.240938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.240938Z digest=sha256:0acbc209e94400115456741d5070ca52551d78668d6e855c7ef55558ca94cbe3

Observation 360b6a8f-6196-442f-b298-6d016242c44c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.296105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.296105Z digest=sha256:eb51393cfb96c24c8ed9d134f5da7e2403aa9efd425308c06c78951c9a87dd29

Observation e7e73dca-5877-4ff9-be43-1fb21fe231ba · outbound

This paper cites 2024 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2024 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.367087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.367087Z digest=sha256:dcd927919119c6a825459b13e8d91d36724f75e591c88d621213ae6dda0933fb

Observation 104fd3d2-3d38-4978-896d-7c756d512d2f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.456293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.456293Z digest=sha256:82bd193083348710118573e045723040d812e7a46f461df4eb6fc67efe3f613d

Observation ac960a03-5beb-4333-ae8a-a92078c66deb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.544737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.544737Z digest=sha256:9925bb42eedebcd1582c527b23d9ff2f946e1f29273dbe04ab471a9d8a5cf953

Observation cd2cf552-6097-47ec-87ef-e3049d793cd4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.608784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.608784Z digest=sha256:2c63330fa731298616dac63c51877a7863c45cedb090b5ff5a6ac6a33214ce96

Observation 986de566-4ebb-4d99-add4-7fd0c747c883 · outbound

This paper cites International conference on machine learning , pages=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International conference on machine learning , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.718817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.718817Z digest=sha256:4e6c14594e737c744c757472d9c6c3907c2b82eedc8f60565419ebba8d771d16

Observation 1def997a-e4da-48fa-b1cd-a2048e2d81e7 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.802584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.802584Z digest=sha256:b672ed8ccbf6527e811cec51d9a16b78411540c83cef5e955676b73e889b5faf

Observation bf3897f7-3cf1-4dd2-8641-827a86680b12 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.906088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.906088Z digest=sha256:2e38a04fa9f0105c5f1529e95d3bcefe886751085cd01c0a9995e1bbd571ecab

Observation 9dca93f1-90c3-4970-bf38-5bce292e677c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:24.982944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:24.982944Z digest=sha256:04f85eaaf03d0e3ae0ebe347852a60450f4332661af5e4c444e811ef999f27da

Observation d398004d-dca2-4410-b88d-32b70387c35b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.052567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.052567Z digest=sha256:9c1f06dddda5effc2d643fb65e5c80ac261697ff93fa4220129737686a3ca599

Observation 56514245-048e-4c88-a373-e180751db1be · outbound

This paper cites arXiv preprint arXiv:2507.20673 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2507.20673 , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.055939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.055939Z digest=sha256:1ea63b94a581bcb696865e5a11551efb0f5972cce15644cf5c25ad5f984c3b52

Observation 8522327a-21c8-42ba-924a-983938f2330c · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.058589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.058589Z digest=sha256:8f05547fe335be2063c61a589edf2896375381bf88d1f3b9533a66d4b415ddde

Observation 786d679e-0833-4e47-a0a1-ed845b8e4855 · outbound

This paper cites Group Sequence Policy Optimization.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.134121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.134121Z digest=sha256:9a99628b6c08d69ab2cc6d0b00b10e6b159c1eaf42a3deba42015b2d2da97414

Observation 1de4cc21-8ae0-4c05-b957-4778999137c0 · outbound

This paper cites Your Efficient RL Framework Secretly Brings You Off-Policy RL Training , url =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Your Efficient RL Framework Secretly Brings You Off-Policy RL Training , url =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.298585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.298585Z digest=sha256:0c912bb5739af6f06b01b74995065a49490ed2d3092a57640e96e1fe36a3ead1

Observation e4c92e99-8e6f-4e18-a5d5-b5e39c27a706 · outbound

This paper cites When Speed Kills Stability: Demystifying.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples When Speed Kills Stability: Demystifying

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.446955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.446955Z digest=sha256:a25247bb80b253053c31643263e43c531f5555decb0848a6adece58afe8a83b0

Observation 8aaa4bc6-f4aa-421c-9d84-a588d0c4a3fb · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.588238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.588238Z digest=sha256:691a96f2fbd7820bc48464e516ecb047f9cedd9c3e4a41f1b2a72c82c59bdd52

Observation c7960eeb-f82d-4099-8439-dedf72049e49 · outbound

This paper cites 2025 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.770620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.770620Z digest=sha256:301734f99c7c55a3b84e7baff4f1afec844c5d06863bb7e6f106fa8c0593b7a3

Observation 4d271953-69c3-468e-ae90-1b7a0bc3e4af · outbound

This paper cites Small Leak Can Sink a Great Ship--Boost RL Training on MoE with IcePop! , url =.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Small Leak Can Sink a Great Ship--Boost RL Training on MoE with IcePop! , url =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:25.894447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:25.894447Z digest=sha256:18a82f8cf6994b14fe6194757c0fead19ba788d7d07b3cf547c231d92bb8f20e

Observation 2bbb3002-cd38-4609-962a-455a163933af · outbound

This paper cites Reward Hacking in Reinforcement Learning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Reward Hacking in Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.043996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.043996Z digest=sha256:51b3982bff3c7327751d45d12ab38ddb0fadf124485624041c3355fa1107539d

Observation ce7b66e2-fe00-47f2-9e65-f8d5409ac127 · outbound

This paper cites International Conference on Machine Learning , pages=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples International Conference on Machine Learning , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.169322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.169322Z digest=sha256:606627bcd9687364de8cd917f66c139d94814e6f2119aef671a32ea69356c91b

Observation 5181479d-b2d4-4764-b6a3-63147c4bdc3e · outbound

This paper cites an unresolved cited work.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.324565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.324565Z digest=sha256:b1c73149fcff7bb11f020a12bfbcf513587e55ac56d45736d5151f650d3e9ef8

Observation 75e1e575-ceae-4910-b4be-8f1b294b5501 · outbound

This paper cites arXiv preprint arXiv:2510.22543 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.22543 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.485479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.485479Z digest=sha256:55774ee77b480696e1818e603b75afff938b87bd0c92b99665f5cbd99a7fcd6e

Observation f6ee98a5-88fd-4b88-afd1-bd1fff59d1b6 · outbound

This paper cites Why Language Models Hallucinate.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Why Language Models Hallucinate

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.610302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.610302Z digest=sha256:6adf3887a23bf762845191c9cc9970177985c675ce4b2a68c331fe7a57fe3043

Observation 644ec384-f034-42a5-b532-3af159802735 · outbound

This paper cites arXiv preprint arXiv:2509.09177 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.09177 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.753190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.753190Z digest=sha256:e87ac176ceb3d9e5af0b27aea88656c9fc503e68d88b0940c0a06b64e5a9394d

Observation a1f20914-edae-452f-bf14-017970f5e5bf · outbound

This paper cites arXiv preprint arXiv:2508.17850 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2508.17850 , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:26.879211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:26.879211Z digest=sha256:340d605b488d82f23b9959bddf6b47fd6cfa2b01098b4f83454e3e50f1a45171

Observation c05e314c-dd90-4047-aaac-b1b841762423 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems (NeurIPS) , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.022813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.022813Z digest=sha256:d3afb79c8ef00179289c51eaeca081146f41e8c07e5690a20474ef3053e5041f

Observation 0a741d53-da2c-4f08-9894-e3d11f428091 · outbound

This paper cites Advances in neural information processing systems , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in neural information processing systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.135881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.135881Z digest=sha256:3c27dc3be33e1c44ffff0beb5a0ee55ee23fe130c6d3764f09ce4559eed99dae

Observation 52d4cd12-08db-4a46-9185-101a8416eb73 · outbound

This paper cites On a few pitfalls in KL divergence gradient estimation for RL.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.338911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.338911Z digest=sha256:088f6443f6e3dd3e514b339c5674203d17b76849a742641ca73e756f7fc8ffdd

Observation 8268c740-42ef-4ac1-8371-7bfda21e2a1e · outbound

This paper cites arXiv preprint arXiv:2512.21852 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2512.21852 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.495012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.495012Z digest=sha256:814ca8396b01f838e3e8fb2e55d9b1ee9e1f1e878243805b4d8d57ae58d9518b

Observation 2250deb4-8107-49ac-9686-a8c2265794d8 · outbound

This paper cites arXiv preprint arXiv:2510.20817 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.20817 , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.584045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.584045Z digest=sha256:33d679199e97adad22f98a5a4f02cf9d667f7a8114a955d644aefd6463c4d9b2

Observation 791ae013-e99d-4e5c-9fe0-14095100ee33 · outbound

This paper cites Beyond Reverse.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond Reverse

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.635901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.635901Z digest=sha256:f9fd9410c1307d610734f41853dade160cff10d5f447d05f20769cd0e030d188

Observation 2a94f418-4dc3-47e1-900f-7c6878ac2504 · outbound

This paper cites First Conference on Language Modeling , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples First Conference on Language Modeling , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.740294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.740294Z digest=sha256:2c62ad7786f7166a05761030a152953dcfd174d41837ebdc9109d2ecb9fc3320

Observation 4e73d123-5477-4d87-b305-e231911add6d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Advances in Neural Information Processing Systems , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.805767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.805767Z digest=sha256:67c222fb35c90753ab8e58475a789d2094bf81d2221518cd0556bde550155282

Observation bf03e8f2-4be6-42ff-8570-0a3750bf6877 · outbound

This paper cites 2021 , eprint=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2021 , eprint=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.878462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.878462Z digest=sha256:89aefb2ed42ea1c5209aee56359d0840e9e697c07b362f44c035ef66be0abb0c

Observation 0af955a8-16ca-4abe-8d0d-10a7dafbf55a · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.968153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.968153Z digest=sha256:a5b4eef1c5c831801c56a48021ad9470efd2151cdbdfbd7bdbf5998ab6229673

Observation 7d932ad8-5103-4abf-8d8e-3a68b7e1d25b · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.043198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.043198Z digest=sha256:9e6d35432ed4d0d56abb9c9f4662afe5da5d7a1cc8c0fdf59b01f594e54bb864

Observation a6f435ce-a6d7-4e87-9064-dc2720359b4f · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.120890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.120890Z digest=sha256:56b0a8803739532d79353dc725ceeae920905d67f97b5c289e2ce17846e099f4

Observation 12a94e09-3547-486a-90a6-3ad997509fb6 · outbound

This paper cites 2023 , cdate=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples 2023 , cdate=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.195301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.195301Z digest=sha256:fe66db96c1280ce2d2e5d9c9a25d201549de18a2fc7e9344e0dbb286fc9a8eaf

Observation cfdfc1eb-6ad9-477e-8379-337302f35f1e · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.266130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.266130Z digest=sha256:e1040a21da04657c670a51ffdfcc7894f61606f539af338ed21d65221d31eea6

Observation 53f87756-ad26-448a-93e4-ff1191eadeb9 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Learning to Reason under Off-Policy Guidance

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.362049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.362049Z digest=sha256:78ed15f23c75368eb040c1960dbd19ee5dc7b3e7ca554de127117c95b3625851

Observation 7de04c6c-5efd-4ba3-ab2a-02f49bb48455 · outbound

This paper cites arXiv preprint arXiv:2506.07527 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2506.07527 , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.452592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.452592Z digest=sha256:7774d64a818a50934f46a011f8c29f730c9a9807272cc7eeb2d6e2a7e867898e

Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.545828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.545828Z digest=sha256:93174a58fb529959ea00d55185efe153faa0ac7ded49f035269cc192fad76662

Observation ebad1479-c453-443e-9776-30c06c72f2f6 · outbound

This paper cites arXiv preprint arXiv:2509.04419 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.04419 , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.622105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.622105Z digest=sha256:0d2c3fd9a6391a7425959eabd5feb4d2547ba9e83405e2cc136af5ea74335194

Observation 35b8c770-1527-4917-8fb3-60869d0916a8 · outbound

This paper cites RePO: Replay-Enhanced Policy Optimization.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RePO: Replay-Enhanced Policy Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.717603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.717603Z digest=sha256:bb1642040840ebd5a7486868a5cf283fccec2d52bf30d952bc0690b719574165

Observation a6ba5694-badf-4fad-8793-20ff8e40f96b · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.818300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.818300Z digest=sha256:784d8a578a590964d790e054ed5a924029715f401477b89c7747c35e6641e96f

Observation eb6a0f3c-58e8-49d2-94f5-7c88251d8adf · outbound

This paper cites arXiv preprint arXiv:2510.02245 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.02245 , year=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.913789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.913789Z digest=sha256:e292b0378da2f6ee5045c23e9144cf49780f0fd6595a9db2a5637face10e7294

Observation e769e426-7214-4d05-bd84-4105242bdd2e · outbound

This paper cites RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.013207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.013207Z digest=sha256:f0e0665d6ac8a0f61fbc7599f90a4bd10a148c14838ff03f8ebe32ff4b3acf8f

Observation f7fe8cac-6531-4408-99c4-1bd10dd168f0 · outbound

This paper cites Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.108343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.108343Z digest=sha256:7503cfb6bd9917be4e23f06fa59784d45fe0a4dc2abe2403970781d351c9f48b

Observation 8577ebe5-249b-4800-bb18-c975f86e57a1 · outbound

This paper cites arXiv preprint arXiv:2510.03865 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2510.03865 , year=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.154734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.154734Z digest=sha256:f4cda9243b796de567f35986a8947e2d10818ae72dbb3f3d6aa7c79422afe37a

Observation 44e25a45-8a51-4105-a098-e97129606755 · outbound

This paper cites arXiv preprint arXiv:2509.07430 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2509.07430 , year=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.245333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.245333Z digest=sha256:d98e14dc9cb96b704ece214843439125f91884fd9a51d1e621cf3cec21a2e4f1

Observation dd617188-aaca-41c7-8cff-f6e816320646 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples The Twelfth International Conference on Learning Representations , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.404192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.404192Z digest=sha256:2678c466dfe470a1fff6f7d9ef7af11cc622f4a83c2b4511889283f1bd3df443

Observation 9b022685-e7d2-4fbb-aa66-842bfd501e22 · outbound

This paper cites MiMo-V2-Flash Technical Report.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples MiMo-V2-Flash Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.526424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.526424Z digest=sha256:40c4e11b08aa6e3a3608eb88d51ee505a5187470936ddf3196014e985a7f52f2

Observation 28cdf359-2843-4333-b14d-b21f1a92d30f · outbound

This paper cites arXiv preprint arXiv:2603.22117 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22117 , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.693770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.693770Z digest=sha256:9f7e9a88ad20dbb47d85d9b62bf328cea4f843dafcf2cf3ab45593930c601a49

Observation 682775fd-3a03-44a7-aa4c-dab4e66ddea5 · outbound

This paper cites arXiv preprint arXiv:2603.19835 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.19835 , year=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.817926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.817926Z digest=sha256:e4c6b6d35515db9d5cf8e48666b5c69b30833ef79bcfab7d059ba81f5bfde83e

Observation 06234128-6f8c-46a4-a1cf-4929522df424 · outbound

This paper cites arXiv preprint arXiv:2603.22446 , year=.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv preprint arXiv:2603.22446 , year=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.943795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.943795Z digest=sha256:9cc3d5527259877a776fcff0e02d93ad770af5a975dd5bd47854d7eb3f32c3b5

Observation 79e6168e-577f-4be5-ba8a-bb1cfab2dcca · outbound

This paper cites Experience Augmented Policy Optimization for LLM Reasoning.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Experience Augmented Policy Optimization for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.079906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.079906Z digest=sha256:ab3413692574b08732dfa08f9adf7f2e861c1f4a63510132aa9f91117084ce11

Observation 9b5031bc-8e80-4294-bedd-c73a5983a727 · outbound

This paper cites One-Way Policy Optimization for Self-Evolving LLMs.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples One-Way Policy Optimization for Self-Evolving LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.215870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.215870Z digest=sha256:fdccba5be4e44e4276cfd6e12ad3fa5d615b09058ee1b4cdadcca267f3717e0a

Observation 50e9359a-fb92-42cc-a5c2-a388db996a82 · outbound

This paper cites Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.344213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.344213Z digest=sha256:1ac42c72bd4e65b00f5cda5b596fb40a76a1e49978655f83e394ba8ca372ce22

Pith citing papers

No inbound Pith citation observations are available.