Pith. sign in

Paper Citation Record · LEDGER

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2608.01743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01743 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.368038Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df80e58d-a917-48ca-be6d-d8d867e8e48b · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.245730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.017940Z digest=sha256:3344861d5a9df84b929dead979dea7856982f7d14c041c6f9f25a76b989731de

Observation c8b89b9c-8b80-427b-bf58-57a1510d7f77 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.230777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.050129Z digest=sha256:f21c96a90eea4302e29134f2a8c20213f6bf61d957c5c883000ee7e25917a8b1

Observation f20e3a33-f3ed-47aa-af3c-01ea0516b4d9 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.217308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.060067Z digest=sha256:bf299d0bda1186e3d1af003fcd1cd3b930d6f218301b72eababc852a0169ba30

Observation d4b0c03d-f000-49f3-a49b-7cb98d1adb80 · outbound

This paper cites H.; Gonzalez, J.; Zhang, H.; and Stoica, I.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning H.; Gonzalez, J.; Zhang, H.; and Stoica, I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.202641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.069509Z digest=sha256:26d17948055e0ed4e9e652406948c995e51dc1e6c836861f952a030154aaebb7

Observation d9356796-39dd-404c-85a1-c7e61bc5c2d1 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.188389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.097267Z digest=sha256:4d0bd87c225a01202f180038d99795d4737e2ffd73e201a9b2b69d417bfd4cc2

Observation 65f5e0f5-bb4a-4284-a5a1-c16010cdfcc8 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.129231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.129231Z digest=sha256:c16f37839b7b030af3c63ceda9922b4839fc003431fda76ef934b35e01243f2d

Observation 74af9aad-919c-4603-a4f7-4284782058ba · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.137374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.137374Z digest=sha256:8ccea4a2ac0b6029d18d1c1f18d8200af9fe24d17db2677684579e3ae0805a44

Observation 7b0d8687-957a-413b-b072-9ba612e3c9c7 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.161193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.141974Z digest=sha256:7d639cbc01f5d0aa84c64aa90116c72cc79a9dc2d1fc3f89804b38219edc9947

Observation 959a227d-d93c-4771-a638-169caa91aef7 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.146636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.145654Z digest=sha256:e0490df54aa66b95434f362bef483aece27c8e2d5590b8925cc0529768027e7d

Observation 495d56fd-0403-4cdb-a2c7-d82029c5ba61 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.132103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.174384Z digest=sha256:f5f8cd2ab83ad53f081def2358d172cd01d7c981a5df65398abf53316f1ce52e

Observation 4c92224e-9ff6-4b52-b7dc-a0a876477ea1 · outbound

This paper cites 2024 , journal =.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning 2024 , journal =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.187684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.187684Z digest=sha256:dc7b838e2d91ec152f9bdd5748affa84b73493668962e43e3728d1290fd8f70a

Observation 09e37af4-78c6-4be7-863b-ec28a04e474f · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 29th symposium on operating systems principles , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.192272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.192272Z digest=sha256:80eade1997ddf723ae722fa6c56e3c1a15ee271519d0fd1556d197bb0fd5bfe5

Observation f05b61a8-05ed-4c07-954b-c723093d7e39 · outbound

This paper cites Qwen3 Technical Report.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.196474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.196474Z digest=sha256:c91c4d3f32121fbfc4f75efff8eb93ca560d6e64f73bd6c415de46309a5510ff

Observation 81afb144-f8cc-4f4a-be3e-bc8db4b9f015 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.200389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.200389Z digest=sha256:40adea6a8f86ed0ea624348ad2b27cd94581285dd7570b8622054dcafd2f7597

Observation f77e03be-e80f-4784-8276-969b8fc49b18 · outbound

This paper cites Advances in neural information processing systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in neural information processing systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.204431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.204431Z digest=sha256:a801301220a902e88f207c8954f0e3f1c9665ba62ca2f7219371724e86b611f2

Observation d4614ed2-ffee-4162-a9b8-5f7d2d541dc7 · outbound

This paper cites Advances in neural information processing systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in neural information processing systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.208940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.208940Z digest=sha256:18f54a5c9d0ea6d0452ff7b6266247c0c0056c99210609f688925c5079d0dbbe

Observation c5d7d783-687b-4adf-a1dd-938cddac8947 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.213111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.213111Z digest=sha256:b9c947de289bab87933e7e6270590c283946af1c352c59fbde915454803be1d4

Observation 2aa3844b-7052-4b9d-9a1e-1610e23bdac9 · outbound

This paper cites arXiv preprint arXiv:2509.07430 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.07430 , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.217118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.217118Z digest=sha256:210e077cfcca63c6e15912d7d51d2117d825585331f05f932e5e2573b4172b93

Observation a98d7c5b-c8f8-4eba-b5bd-07ff69b01084 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.221531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.221531Z digest=sha256:2fcd419efaab65b212acf7f13b7da90b5cdc2ac439ef9ecef4bdb69e667efb73

Observation 2a521534-ce38-4f85-9d5e-09031face8e8 · outbound

This paper cites Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.514401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.225531Z digest=sha256:3c0a5ee0b1776eb88050dbcec802c86fc975ecc01d31f25d4936ad1a88d1f9b3

Observation 4fea440c-0f4f-4aa6-9af2-c869ce602c86 · outbound

This paper cites arXiv preprint arXiv:2510.03865 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2510.03865 , year=

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:54:44.888734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.229256Z digest=sha256:c08b2dc0ef4d92321ece259cd6692a4f946302becba70ef6ef927f4103c5d804

Observation 51920178-91f2-4f57-aa29-e7275570cb94 · outbound

This paper cites SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.782159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.233162Z digest=sha256:af8dd1e6bb99d58f40eb29748ae14b837babd7de70eedd94e2e6b151e2e25638

Observation c5c3b979-95d4-4927-86f1-31f82e92fb49 · outbound

This paper cites arXiv preprint arXiv:2510.20817 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2510.20817 , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.237573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.237573Z digest=sha256:63ea07c8bf5e063626b333d6f6905a8414300f7f7180d30abc2bf083091498e3

Observation a819ad68-2a25-48bb-ba76-c9fe1bf79dfd · outbound

This paper cites expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.715269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.241343Z digest=sha256:77d2a998daca006b9e1b4c7ef765319ee99b5aa63f28b5e87fce3762de01d1f5

Observation 69b0009d-c43c-4ea1-8c06-e629fe386bee · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.245197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.245197Z digest=sha256:a95ff77dbc47ef3febe538f1c44a2c29db942aff87591b0d2af53b90963cae41

Observation 51f81fca-0199-4d8e-8479-56237d36f1c1 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.248752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.248752Z digest=sha256:0f4d032af7c1a19953cfe6fafb4aa44903c5037bbc4fe3330547928bfab5f8db

Observation 4b9db30a-cc20-4107-8274-997583ac6163 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.069692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.252587Z digest=sha256:1ce92276c3fa646bb1123dae619a2d4180400a9e332d5c4893e0d5af0e96d18a

Observation cb170a3c-c32c-4233-8b8a-7deb11a1ecc8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.256399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.256399Z digest=sha256:002464c53e6d0423360cbcfe56f02433966fd7f3dcb12dae34df6c782a02785a

Observation 0008add1-0006-4660-b26c-238db2c742fd · outbound

This paper cites Group Sequence Policy Optimization.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.260001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.260001Z digest=sha256:d236b6b7ddfbe8cf28e40966dcc6317adaab6935100027e87925ae2364fd2b1b

Observation 0b978fe9-44cc-42d7-996e-074b439cce26 · outbound

This paper cites SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.263730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.263730Z digest=sha256:56ffcaf005aa6ae3adaa407daf1ce3bc5c3eb15c75098bcafba0a3a6abbc6239

Observation 61c442a8-efe1-4f94-811d-0f65aa07c352 · outbound

This paper cites arXiv preprint arXiv:2509.21826 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.21826 , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.267496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.267496Z digest=sha256:79746640004afee4f6cf114f6da765c4a945a55a607be5ad201184787efa3f9b

Observation beb7572e-2aad-4fb6-9fb2-415e538b6b4d · outbound

This paper cites ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.693998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.271066Z digest=sha256:3d893096aee8cb46bc2f1a6d70d0d835de52b21fd386334f77dfdb86202582ea

Observation 1760ee1a-2a76-4fac-b3b8-d602362fe35f · outbound

This paper cites Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.275053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.275053Z digest=sha256:74c503a304f551b7535c47b3691af155d96a20d41eea8392b4239c4d9742b23b

Observation 11a28c1b-8692-4df2-bc0f-b48642f39907 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.055465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.279008Z digest=sha256:a9d9e293dbe6386be783c0981a4f87571de5e49c556719d4340bd4afd2fe7b21

Observation 513eb1db-d952-4e0c-abc8-c1e864933d55 · outbound

This paper cites arXiv preprint arXiv:2507.14783 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2507.14783 , year=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.282831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.282831Z digest=sha256:76a46dc1aa41c2f6ab4c0922dd2d4ee3a18b98713a4fc09cb159e700839dd5bc

Observation b0b85819-9a32-450d-990c-f7710a6cf889 · outbound

This paper cites Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.286875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.286875Z digest=sha256:680e3a7cecd6037990565c86ffec8e8c662d206515151ff53e69d9110619d1ce

Observation e48110fa-5932-4e42-a981-6716065cab6a · outbound

This paper cites arXiv preprint arXiv:2602.12566 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.12566 , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.290891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.290891Z digest=sha256:5105c2306ae511db6b3f5ebcb9124c5ee2251689b430cf81541cf54f6c3efbc5

Observation 78f8ee3c-a707-4ca5-881b-389d9ef6b591 · outbound

This paper cites arXiv preprint arXiv:2602.02301 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.02301 , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.294669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.294669Z digest=sha256:7f276d0e37f30ae6fe3f0a03c743a739bf321d7a0cc6615f2e48417f1dda05f4

Observation 2de762b3-6690-4c75-a29c-ff5df8769829 · outbound

This paper cites arXiv preprint arXiv:2505.17508 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2505.17508 , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.298515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.298515Z digest=sha256:d1dadf8a3aa583b3f25df3b456ae9c371d854e354987a805a4be31721e5db299

Observation f0ab0c6d-c983-4837-ade3-88399ac72e09 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.041362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.302583Z digest=sha256:823aa789d5b2de4f49563155d8134b3ee5773366740c1d3226473bb29cbded01

Observation a25a14df-98a3-4de8-be31-241cdd07debb · outbound

This paper cites MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.306632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.306632Z digest=sha256:ffe897a92aa6862d31c57745d6a07bef7fd8daa2080f989663581a83a18246d9

Observation e66e8458-af75-4c5b-8783-51bc09ccd8b2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.025570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.310926Z digest=sha256:472f2e1318d2c655aaa0431784236848f895e8db4ac5caf6f0cb474e98e2e805

Observation a5b07d83-b46e-4da2-af82-47e44e99f983 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.315385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.315385Z digest=sha256:aa057b7e1d15b60552c4668be7c672584fe2927afded8fc62e5c31b8f21c64aa

Observation fefc93de-90e7-4a81-a7e9-de149a2a8cef · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.319080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.319080Z digest=sha256:5015ca305dfaf81b1cc008700b3c5edc5cb46b7458ffcb171fe4a5bb7322b877

Observation 440a5857-0cef-47f0-9e65-01ef9da28d69 · outbound

This paper cites WildChat: 1M Chat.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning WildChat: 1M Chat

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.322742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.322742Z digest=sha256:1293e929e06c9e0b6c4f68e4c2e7628e27687b3c348631bf871bfa9dc58e382f

Observation d85754cd-3a0f-47a7-96bc-de263b245193 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.326437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.326437Z digest=sha256:2fecbcb4d2e9440d5d62eb2b38a1027c7e78d924123253a0ab4623ef963a6de3

Observation 6d3c600e-de81-43fc-b9b7-4ad1eebf5f11 · outbound

This paper cites arXiv preprint arXiv:2512.15489 , year =.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2512.15489 , year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.330165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.330165Z digest=sha256:0633cf4d9f3a3dbe600ffae80f63e9c75f2d2f868acb4f09b209748f34b808b8

Observation 5b129d2e-ab42-4199-8d96-22397aa477a2 · outbound

This paper cites arXiv preprint arXiv:2509.20357 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.20357 , year=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.333909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.333909Z digest=sha256:466cd6661463550a93e7fb4a45e5f4daebe0edb3320edd43ac0d919df668710c

Observation 9f41ce78-a012-42b9-ab64-9f1f61dd0740 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.337692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.337692Z digest=sha256:66015bbdc47b28d6cff2dd7e9e41df8be8400b15f70295656a29943fcfc20211

Observation bbd38b4d-117c-43b4-81ca-2cbdc185adab · outbound

This paper cites arXiv preprint arXiv:2512.05962 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2512.05962 , year=

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:54:44.806705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.341274Z digest=sha256:4e4dadd7b25270a2635af633c4b22b6e39988fa6815c04df0aabbad792347146

Observation 490492f5-7b91-4f81-b8db-d1e281bbd6c6 · outbound

This paper cites Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.626662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.345264Z digest=sha256:80181f8ab38771cccb7e6fc9f67f8afb738320343a23e76d17b65c3fbae33f7b

Observation 14284321-4222-44e3-864a-f7944fb9fa81 · outbound

This paper cites arXiv preprint arXiv:2602.19895 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.19895 , year=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.348703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.348703Z digest=sha256:e2b6d0d968f497b0cdb7b3eda32791929052ea5bc394f309318b36b6a8465b73

Observation d40d7bb1-73a0-4e5b-804e-e4a11b18e0cf · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.352386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.352386Z digest=sha256:e0768691d002028c598aceff31046292b7cb285c1f990c095a8fdad4ecdf0ff9

Observation 0309255f-307c-41e5-8890-66f275465eba · outbound

This paper cites International Conference on Learning Representations , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning International Conference on Learning Representations , volume=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:44.966856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.356061Z digest=sha256:d05ef18aeee4ba50e5ddf2105640bf75bd80dce99946dc691fd15d6685996782

Observation ee7e51f3-7ccf-4a53-b5e3-504e3936633f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.359793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.359793Z digest=sha256:d83f984083dc3f38bc5e2ac048f1aa7165d6f0144628ae91231130c6f2f53ca2

Observation 567c4997-dfc7-44bc-90ee-c1b8dd195bdc · outbound

This paper cites arXiv preprint arXiv:2507.14843 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2507.14843 , year=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.363782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.363782Z digest=sha256:fda6ada66fb31bc79a6053eebf267fe67db4b2a1e134205ab50f43b8b6b59d3f

Observation 4a43e7f0-9217-49ef-97ad-685026069a89 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.368038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.368038Z digest=sha256:afddff79ccb6df5966a617bbfbbed56b14e9e32e551651a4295ca91fa93c552d

Pith citing papers

No inbound Pith citation observations are available.