Pith. sign in

Paper Citation Record · LEDGER

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 10 inbound Pith citation observations for arXiv:2502.06533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06533 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:11:08.217318Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:01.278862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T11:41:29.623621Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4335a7de-95b4-4289-beb7-31d9e2830a4a · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.598573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.108360Z digest=sha256:63cf0081c05a0c0770b291637f991708fb7ca5192c599105ffd019264d45be5e

Observation 2be403d8-dfcc-4cbe-acd2-a0877eae64d9 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.590593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.111751Z digest=sha256:6b768e8bfc9497f8451f66368496783509b37f74ec99d64df5451327f3ac28d8

Observation cb02443b-5e7c-454d-80a1-cdaf2f15c3f6 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.583473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.114846Z digest=sha256:c00c78d2620e74288022a5fc82ce0a65ed60bd19cea111b3905a86c287394424

Observation f5e03d1f-0a1a-4099-98da-8e6dfe7a5b76 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.119207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.119207Z digest=sha256:d2b6eff80028377f33f76498883db54cab65e42fe982f2cb3978e415298e2d58

Observation abae6b32-0469-4597-84ce-2779d1d152b1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.574958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.122210Z digest=sha256:37149a1f16600d41c620c499429f418215ff8092552995959185173960d06a84

Observation 918f833d-6546-4426-aabd-16754a3d7bdd · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.125025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.125025Z digest=sha256:58a0e9264179a138ce836a89e4b544934df569938316ad67bb92080399f6db2e

Observation a65dbc7e-62ff-433b-8364-619b03875d86 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.567131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.128531Z digest=sha256:25d2c4a046c2ddad79a9069108e65a9852485122bedbdbb6b2a783057d3e7703

Observation 3db7b570-de64-4666-901b-8d936ac93fe0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.131846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.131846Z digest=sha256:7e83e7a282a7dc6996d9f8b30f6138b29f3e36d340ca1236a1a0ad0aaa2dcd7d

Observation c368a59c-9816-4431-9826-845747dd1018 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.134977Z digest=sha256:13c5d2232038dc89099d8cd5d7a2b65f87f29d2d9f08c9ad01a9423534100ee8

Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.137792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.137792Z digest=sha256:bf4934e2de356a2f60c83414c5973242114959d06eb1bf5593b372dc527fd538

Observation ab25b46b-525a-4587-ac65-4057203baa0b · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.140623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.140623Z digest=sha256:06b1c2e538b466f3beb045561bb387aacb3c13d16c1a4250e78f7fe917ddac7c

Observation fabb473f-eb47-46c3-a79e-e01d1d4da8f1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.555008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.143465Z digest=sha256:b6eb23d5ab6c1f45f1c81cc6de75de162b69f43f393c7d2a85f2d24caa72f124

Observation ef6d4f98-1704-467a-a0a5-21484918da71 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.547831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.145539Z digest=sha256:d93a5e857657b67209d35a647d448db06c9090ad6f640b4017d19395be542f6e

Observation f1b2ed8f-87c8-44a9-ade0-4021d5e97acd · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.147904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.147904Z digest=sha256:91fbadbc6806f163be37b8ede246cd6aad65a93507c0af7aeee4a9a997b9d5bf

Observation 6086b97c-1a84-413e-a8a6-fa899a26aca7 · outbound

This paper cites Lee, Kangwook Lee, and Dimitris Papailiopoulos.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Lee, Kangwook Lee, and Dimitris Papailiopoulos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.541335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.150117Z digest=sha256:a31e2066fe40f1f8d09049bca6db19b73895e668368bb0712691199a8f89394b

Observation 028bfb75-43f9-46a5-9400-58b7aba0a563 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.152169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.152169Z digest=sha256:0635cc8fc34edd1f1e4e98fdd8b66535ebfe34c548fb84cb2b1ff8e40328ce2d

Observation 4389c21f-a67e-4a21-8939-300e35901201 · outbound

This paper cites Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.154302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.154302Z digest=sha256:68fd697366116e439b9899b429b8d4c92b02d7152c616c98d35cbb6177944e80

Observation 92dfa1ca-53fb-45a5-93cf-af728f0d6dee · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, and Tom Goldstein

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:11:08.534555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.157948Z digest=sha256:6dfa875b5855b84b1a0ab6173e9364cac3b5a7db2f5fb68808f91a29269f252b

Observation a0a11701-511c-4602-a25e-807baca51cc0 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.526289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.160447Z digest=sha256:f2aaedcd052df0ee8dba20c7ba0d67111392d0543d94eb7246bb41ee475f639f

Observation 020d0dcb-a310-4071-aa29-30b33186f5fa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.518595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.163134Z digest=sha256:a717b705819706a81b461e75a223d449e94831ed2a035f90503f6be6fb42dbab

Observation 24008427-3b92-4768-b228-d42bd07c16e4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.165824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.165824Z digest=sha256:1decc3aed3349209d63167e82edde0615f8a5dc9b9fbe692cec1fc7f4b313d12

Observation 757c0598-cfca-444c-9323-03d517276ec4 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.506127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.168861Z digest=sha256:1fa35685a7c8c6295d5f4bd1169494735839d3b3a7a432d0ca6745aa0153f384

Observation c9a8232f-e00c-44de-8223-7c30eedb9ef1 · outbound

This paper cites Positional Description Matters for Transformers Arithmetic.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Positional Description Matters for Transformers Arithmetic

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.172446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.172446Z digest=sha256:7399192682d33476e1163816cdf6ba44bc7bfc0ee01817590d005e49951a49fa

Observation cac9a057-a2aa-416f-9395-58ec1ee3f4d2 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.498250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.176269Z digest=sha256:1968d840533b7aec8e32c34a792e494720b4a8f6b22938b15f63e863372ce4d4

Observation 55a321ac-facf-48ee-b2ac-08c9b019dc82 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.179605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.179605Z digest=sha256:0c610bda50d615be8fc848536c7df1b2e2907d8584fe8e88e0dceb32ae53b949

Observation 24c77f65-7e3f-4ac4-acfc-a7da8bfd6b90 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.182556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.182556Z digest=sha256:5a89bda80c11f8970a4b28a5b68a5c1c4f0e3020a1d68bcf5472f437e3141cb3

Observation 636c5415-2dd8-4bb2-a643-7afc60724c4b · outbound

This paper cites Chi, Quoc V.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chi, Quoc V

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.185354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.185354Z digest=sha256:12499be6b9bdd3196266d2d49e48a969951c8478f4c0f41fe4fc87b31a4e226a

Observation 665bf2e8-fd54-4a58-9104-00a2bd0f3aaa · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.486457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.188614Z digest=sha256:d0ae0b0b685c708c029b9ad891767abadcd2243c1a7f8f750b9e58b5959b7ff5

Observation 0515158b-359e-4933-9066-0b29407c7474 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.191550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.191550Z digest=sha256:85fe957b5b0988b4986c1a0cf3950d725faa71bc87e27adff42ade0cf269ae1f

Observation cde243c6-e19a-47f4-9383-71f9af953329 · outbound

This paper cites Conditions for Length Generalization in Learning Reasoning Skills.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Conditions for Length Generalization in Learning Reasoning Skills

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.194619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.194619Z digest=sha256:db00aa42d1f80cf9d7a0cb4ba4cafbba0ffecb16633491a02820b5309f103710

Observation d5e1c47b-f1a6-4b0f-9def-eda2b3f42635 · outbound

This paper cites Hausknecht, and Karthik Narasimhan.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Hausknecht, and Karthik Narasimhan

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.197539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.197539Z digest=sha256:d1c664d4996a401b16618e6c54600799085e8ed0a373047e826fb1a8ec092bdc

Observation cdf1b3b7-0ef0-4a28-98f7-d25e3ee8bb00 · outbound

This paper cites How well do Large Language Models perform in Arithmetic tasks?.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning How well do Large Language Models perform in Arithmetic tasks?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.200325Z digest=sha256:14f0e1ae65132f3ec517b0dfdb2dca33d7c59f6654b426e57570cdd2b89cdd1a

Observation 6c344090-0431-4f4e-add7-abf0ab2e5722 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.474833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.203740Z digest=sha256:cfbe97c107b726557701b1f8a6d58f79353b5d813a9adc96a8267d7de2e24747

Observation 18beb684-bb45-4677-b7c7-b82c595574a7 · outbound

This paper cites Chain-of-Thought Reasoning is a Policy Improvement Operator.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Chain-of-Thought Reasoning is a Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.206499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.206499Z digest=sha256:7421bb935cefda47a159e4653a67e8fe4bec1d233bb180ea608caa8650c33a4b

Observation c16742db-1a7f-40b6-9c3a-95ea4da776e1 · outbound

This paper cites an unresolved cited work.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:11:08.466799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T15:11:08.209904Z digest=sha256:536abcd8bc5b5184f11ba8e1a89fae65336097f2ee6895f30018badbcac55264

Observation 9c541d36-6aa4-4a9a-a160-5566293ec96d · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Fine-Tuning Language Models from Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.212092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.212092Z digest=sha256:190d6ed05af056ab87e0ca09aa6f0bae826ec6ab2d51b97d4b9eede65a39ba3b

Observation dfd56579-ae4d-4fdf-b3ae-cb16ce856c98 · outbound

This paper cites online" 'onlinestring :=.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning online" 'onlinestring :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.214627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.214627Z digest=sha256:388baa59cadff53610905bad4a0506be27a71722284485eef9eea4a93708146e

Observation 453d04b3-385b-43cc-b44c-fb4d7f650d30 · outbound

This paper cites write newline.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.217318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.217318Z digest=sha256:d91cc817ce4778ea9100054727eb12e8985bb39819cf9ac27941b695ccb9f916

Pith citing papers

Observation 7674cb9a-22f7-4bfd-b688-55d4dd1f049a · inbound

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning cites this paper.

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:08.960644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T12:12:08.724844Z digest=sha256:76434e210f2bb808ffd43993a3ff9c125311d9128510d2f659840d247a88b0a6

Observation 44b0d37e-a9a3-4184-9cb0-6a9915ea55e8 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.278862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.278862Z digest=sha256:abfb80774f9da287f634b7089e11b9c892f571ae27fb889d6fbfdbb53f38725b

Observation 6a1c392c-5e2c-4a08-a44c-b3dcb06b6e52 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:38.761833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:38.761833Z digest=sha256:0c24cd7b0a03d76bc8878526617cc6a88399fcd96dc0bebbb633c1db65cbf55c

Observation 1eab7712-b1b2-44ff-98fa-4910f7e8c24c · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:24:26.241888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:3e7ae3c4d16a52d09c6295b358fe33d30bdc2b25b7d6564f9da02fc5cb361ca3

Observation 376af9d9-ebde-4618-9197-5700c26f2387 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:43.057152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:43.057152Z digest=sha256:5d2f83cac67a4315c1bdc69d0b05d51e4dbdfeeb5c51c05bb04b18454143f125

Observation 0b57fbfe-e87d-4eb8-bf7b-53889a8b50a3 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.263325Z digest=sha256:03a6651b52defd7f209d293ca7ace6c7b616a0063314e58ce96638fc40021661

Observation 68e48558-ca96-4c19-b535-78a32c01ac09 · inbound

Training-Trajectory-Aware Token Selection cites this paper.

Training-Trajectory-Aware Token Selection Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:29.626509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:41:21.275802Z digest=sha256:fca52fd0818ca45048c9ed562ee4db206e75f4b7c905bc0c33e57353e4a41714

Observation 713e4f03-97af-4015-a2a7-e47620c85e4b · inbound

Embarrassingly Simple Self-Distillation Improves Code Generation cites this paper.

Embarrassingly Simple Self-Distillation Improves Code Generation Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T14:33:35.834383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:33:35.834383Z digest=sha256:ff98931e6e1f8b9bb3c6de793a0d6ddd7137ff6bf9c2b2387dc68137bfc1707e

Observation 9ccb437f-a756-4f65-9150-38507645632c · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.709068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:c8fa0c0fb34927574c811d09a823a17aace1e36a1765e053f9e208cfde4e7840

Observation 3164a8ee-072f-4e86-80d7-3c6c4139884a · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:14.645501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:92aad188d71f0db8b86bd6b053be0929fde6fc5f8e4cd4ebf438d465ba5c0e07