Pith. sign in

Paper Citation Record · LEDGER

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 10 inbound Pith citation observations for arXiv:2502.06533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06533 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:11:08.217318Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:01.278862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T11:41:29.623621Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4335a7de-95b4-4289-beb7-31d9e2830a4a · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.598573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.108360Z digest=sha256:09e96bee80b2043683eb46e3717c6a51f322fa6b47eea9c0d96f5f8940af8f01

Observation 2be403d8-dfcc-4cbe-acd2-a0877eae64d9 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.590593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.111751Z digest=sha256:cb192664cb3b15421ad44975e3e605b3e332816e1a891dc3ca8336e0466b349c

Observation cb02443b-5e7c-454d-80a1-cdaf2f15c3f6 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.583473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.114846Z digest=sha256:11b8a3f9d3edd462e25d595278fb58d60d66588cfb1617c494ab0bafeb85639b

Observation f5e03d1f-0a1a-4099-98da-8e6dfe7a5b76 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.119207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.119207Z digest=sha256:98162c0c0032f8f73c8bef06127f2e35f11dfdc1f83291bfd453a0d480a84ca5

Observation abae6b32-0469-4597-84ce-2779d1d152b1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.574958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.122210Z digest=sha256:1ee66b737795372758f4ac28ab4ac0005c57f65a7cd0012a6db7716564c6bd66

Observation 918f833d-6546-4426-aabd-16754a3d7bdd · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.125025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.125025Z digest=sha256:2fb0e05ed626e07c5bff9bda93b26cb64fce246e2d44a4fd2404d21d08a4187e

Observation a65dbc7e-62ff-433b-8364-619b03875d86 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.567131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.128531Z digest=sha256:c81fec0139926aba33c5be0e27cd76151ce8faf93296a72f642b698b8b4e1af2

Observation 3db7b570-de64-4666-901b-8d936ac93fe0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.131846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.131846Z digest=sha256:7d8be31661594b34aa995fe3c25d22088639b382c85c194ed0c571598a605a3d

Observation c368a59c-9816-4431-9826-845747dd1018 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.134977Z digest=sha256:6d70fc37232fbc4c42993730c9322115e6cab16630fa2ab78649dcd2db37d19d

Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.137792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.137792Z digest=sha256:82ff676c8f1e813cf1784bbe3e4687bdd74c428939c2c4240338c79e76ce0d73

Observation ab25b46b-525a-4587-ac65-4057203baa0b · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.140623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.140623Z digest=sha256:9f3819c6579faf28e51b6da24fda384fcc7f336234873f7312d28211909e118b

Observation fabb473f-eb47-46c3-a79e-e01d1d4da8f1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.555008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.143465Z digest=sha256:b50ce288af64881951d2438051e0d04136e8bbb28c879e83c443bcc7678b34be

Observation ef6d4f98-1704-467a-a0a5-21484918da71 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.547831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.145539Z digest=sha256:1bc63cdfaced00656d40944dae64664ca7024079455480649911929a136bae69

Observation f1b2ed8f-87c8-44a9-ade0-4021d5e97acd · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.147904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.147904Z digest=sha256:2623b87633d0b53d6202d270832d7dd8fe08642b9a632e8fc234b75e402e07d1

Observation 6086b97c-1a84-413e-a8a6-fa899a26aca7 · outbound

This paper cites Lee, Kangwook Lee, and Dimitris Papailiopoulos.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Lee, Kangwook Lee, and Dimitris Papailiopoulos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.541335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.150117Z digest=sha256:1355b8c64481aca45de03af67fd4b7644a696905d478de9aebb411554b71c9a4

Observation 028bfb75-43f9-46a5-9400-58b7aba0a563 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.152169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.152169Z digest=sha256:fe0f645eb97e6da85d514c484127c021a72ab54f1c8aabb01e63fa63390716be

Observation 4389c21f-a67e-4a21-8939-300e35901201 · outbound

This paper cites Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.154302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.154302Z digest=sha256:cf123aff5fd88804a4cc6add012b0925c834bd675fdc97aa54aafe9ded1d61a1

Observation 92dfa1ca-53fb-45a5-93cf-af728f0d6dee · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.534555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.157948Z digest=sha256:a64dafc416132988978c62da8111c2a3f2e117ae196871cbbbef4ab65425ca2b

Observation a0a11701-511c-4602-a25e-807baca51cc0 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.526289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.160447Z digest=sha256:7a3fc23dc1c857334748313afc553e2539294f9cd598be677f030b3be3cbd355

Observation 020d0dcb-a310-4071-aa29-30b33186f5fa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.518595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.163134Z digest=sha256:cde6ae43cffdc0107be7fffd136f200556d572b3d880bfc25e41599973a1d5cc

Observation 24008427-3b92-4768-b228-d42bd07c16e4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.165824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.165824Z digest=sha256:00e76774ca55d6a134b46fb0ad4a424e263951f920f39a55710ed1bd7b84721f

Observation 757c0598-cfca-444c-9323-03d517276ec4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.506127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.168861Z digest=sha256:6d1d64ce6d64bca142b967a520b14ffc7eac34f2b44cee44d8ae0ac26f8c9c73

Observation c9a8232f-e00c-44de-8223-7c30eedb9ef1 · outbound

This paper cites Positional Description Matters for Transformers Arithmetic.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Positional Description Matters for Transformers Arithmetic

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.172446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.172446Z digest=sha256:8a83d41da074879f238b8e96057bc30f0c1823c9f877274ce600994b8682fb1f

Observation cac9a057-a2aa-416f-9395-58ec1ee3f4d2 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.498250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.176269Z digest=sha256:31492ff24c2a3f00354b5bf8f62c6176cf35e439244aa34ead044bb0c2368f7f

Observation 55a321ac-facf-48ee-b2ac-08c9b019dc82 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.179605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.179605Z digest=sha256:3a9d086a0040025c6e3d719138d04d75bd217401ea9e4b4e19ee91bdca7d6cd7

Observation 24c77f65-7e3f-4ac4-acfc-a7da8bfd6b90 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.182556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.182556Z digest=sha256:cef1cf675726ab2956c64edd4900fc1ff2b2b2fff486984d8229c43a53b68349

Observation 636c5415-2dd8-4bb2-a643-7afc60724c4b · outbound

This paper cites Chi, Quoc V.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chi, Quoc V

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.185354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.185354Z digest=sha256:3a33b3d680a33e567acbba8b802a4df12dea94486b9830781a32fba46c56aac3

Observation 665bf2e8-fd54-4a58-9104-00a2bd0f3aaa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.486457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.188614Z digest=sha256:80eb33a96634b6e8aebd948c92aee6872d6bf78a07e9b34013a3674581a08740

Observation 0515158b-359e-4933-9066-0b29407c7474 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.191550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.191550Z digest=sha256:86453afe7ab3b8de5a1c6f07b4b9b7cae3e4093f7b8c6c5740fc25e1a709821b

Observation cde243c6-e19a-47f4-9383-71f9af953329 · outbound

This paper cites Conditions for Length Generalization in Learning Reasoning Skills.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Conditions for Length Generalization in Learning Reasoning Skills

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.194619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.194619Z digest=sha256:d2313219c02aa6eb7aafbe4bb1bcb55970bebcf8ae07e397d7a5ad94023ca514

Observation d5e1c47b-f1a6-4b0f-9def-eda2b3f42635 · outbound

This paper cites Hausknecht, and Karthik Narasimhan.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Hausknecht, and Karthik Narasimhan

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.197539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.197539Z digest=sha256:a5070c06de56c74d01575db1da0cce4c99d7c3413404bd38b1f526d8dca39449

Observation cdf1b3b7-0ef0-4a28-98f7-d25e3ee8bb00 · outbound

This paper cites How well do Large Language Models perform in Arithmetic tasks?.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning How well do Large Language Models perform in Arithmetic tasks?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.200325Z digest=sha256:fcbcb7ca28240f62d79f967ed81bcea3cea03282b64c67227e45c621dff3ab10

Observation 6c344090-0431-4f4e-add7-abf0ab2e5722 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.474833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.203740Z digest=sha256:4108aff45cab6e28d5aa45039f06487639342faafd2883511e31eddc9cbefd08

Observation 18beb684-bb45-4677-b7c7-b82c595574a7 · outbound

This paper cites Chain-of-Thought Reasoning is a Policy Improvement Operator.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chain-of-Thought Reasoning is a Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.206499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.206499Z digest=sha256:a05d6d73d5141f570d521ec1bea6ac76bbc7f94f6c95065b3ff4bff4a87be55c

Observation c16742db-1a7f-40b6-9c3a-95ea4da776e1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.466799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.209904Z digest=sha256:cf20461e7c1ae590a2df8f7bf0c7bd9488ee6abbe302ed99cde47e1c0a90cc49

Observation 9c541d36-6aa4-4a9a-a160-5566293ec96d · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Fine-Tuning Language Models from Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.212092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.212092Z digest=sha256:87ab43bf4184b0d0bddc5a6fd94cecbf71baf5f9c3dfb2fc8df4970c464f5ef8

Observation dfd56579-ae4d-4fdf-b3ae-cb16ce856c98 · outbound

This paper cites online" 'onlinestring :=.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning online" 'onlinestring :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.214627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.214627Z digest=sha256:2d0e4457b9b601f164285a4bc867a6ea955bd089e3975d06b9b8f54987c40b35

Observation 453d04b3-385b-43cc-b44c-fb4d7f650d30 · outbound

This paper cites write newline.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.217318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.217318Z digest=sha256:38aa0683b5784e805d1c96ca1845de9efee1f632214622d752034d06bdd66299

Pith citing papers

Observation 7674cb9a-22f7-4bfd-b688-55d4dd1f049a · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:08.960644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:0299afed35e4001d592c9f52c86648db87b70236444349bc118b707872e2c08b

Observation 44b0d37e-a9a3-4184-9cb0-6a9915ea55e8 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.278862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.278862Z digest=sha256:e9c7d60e52b91b9d85dd0ecd49fa87021def7729d0d3882267cc631d81f26230

Observation 6a1c392c-5e2c-4a08-a44c-b3dcb06b6e52 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:38.761833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:38.761833Z digest=sha256:51314340d4d3b8853f3f7d1dde2ecb27cb8a7c3860ea0030085992f3bee8a351

Observation 1eab7712-b1b2-44ff-98fa-4910f7e8c24c · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.241888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:5a34d4c72d0d79c515a9be6f38928df6b762181d900ed7b7fd91e0005f5c7949

Observation 376af9d9-ebde-4618-9197-5700c26f2387 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:43.057152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:43.057152Z digest=sha256:126fb908206a3d1e58bdba5e63e8ed8f440e113ddf3f29a3bc7c668db2384b50

Observation 0b57fbfe-e87d-4eb8-bf7b-53889a8b50a3 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.263325Z digest=sha256:b0b51152c3b4a6a7393517a466c4e4483a59f7aa5c9df5b36113173bad5b29eb

Observation 68e48558-ca96-4c19-b535-78a32c01ac09 · inbound

Training-Trajectory-Aware Token Selection cites this paper.

Training-Trajectory-Aware Token Selection Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:29.626509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:41:21.275802Z digest=sha256:fd40912ceebecf932f4db62901d06abc2adc6b4281162ba419e4e86a3ecd9f69

Observation 713e4f03-97af-4015-a2a7-e47620c85e4b · inbound

Embarrassingly Simple Self-Distillation Improves Code Generation cites this paper.

Embarrassingly Simple Self-Distillation Improves Code Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T14:33:35.834383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:33:35.834383Z digest=sha256:d4a2ec4923e7c9ff16ac86bf53c2c5bec49a63c2f2e93d32acba318f36369edf

Observation 9ccb437f-a756-4f65-9150-38507645632c · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.709068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:68fc078e280b1009f8392b1b54abc509e0737fc66e625e7562e1299234655c87

Observation 3164a8ee-072f-4e86-80d7-3c6c4139884a · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:14.645501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:54775d6b41334acf02ad2993a57614f11c303789e473f3c9ff0880181598ace7