Pith. sign in

Paper Citation Record · LEDGER

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.07976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07976 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy14
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f939aa2-c97f-4ece-a945-caa962951124 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.513959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:633302d4e553984e1952b2c62a8470006d82accc260a04df4d1c4cbfc9c84339

Observation 62f0072d-42cc-49d5-8032-ab5669f3fbfa · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 3

Resolution
verified exact
doi, observed 2026-07-10T14:27:07.150173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:bc669b6b2494ed79512a55bc431f68267a87f74854c8cd2d842bfd9bd6a92a5f

Observation 889c4786-37d5-48d0-b766-416aac88aafa · outbound

This paper cites The curious case of neural text degeneration.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning The curious case of neural text degeneration

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.522253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:f837839a4706bab3922b5882c31376535e35430b773a151b3b8554d4b6926a9e

Observation 371fe2a7-ef04-4efb-918f-889151363546 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Large language models cannot self-correct reasoning yet

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.510160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:5c8e9c512567f3ab96bdb9781ee652bf3beb5856fb9fcef82ae97d029271a5ac

Observation 5ca50a82-9f01-4344-85cb-904541018532 · outbound

This paper cites URLhttps://openreview.net/forum?id=IkmD3fKBPQ.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning URLhttps://openreview.net/forum?id=IkmD3fKBPQ

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.519445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:b871b5a4f414c8fe9b5cd60b6d67c7a8481c4f3d464bfccdf2481395e4ffe2a2

Observation d9cacb44-7e90-4b73-813b-888370759ec5 · outbound

This paper cites Execution-grounded credit assignment for grpo in code generation.arXiv preprint arXiv:2603.16158, 2026.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Execution-grounded credit assignment for grpo in code generation.arXiv preprint arXiv:2603.16158, 2026

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.362244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:c10fb9601815d6ea1bc8c31094c29b98049749024308860f83c430a4cb785352

Observation 8f841fa7-bbd9-4c3c-80e9-94cd255008cc · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.520180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:964d7e82050cd8aa291fc7218df1cc137dfa4debb445d2278e5ec3fe86cf9553

Observation c6b3cc15-1af5-43fc-a7ed-70a27c62fdc7 · outbound

This paper cites Let's Verify Step by Step.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Let's Verify Step by Step

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.380941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:17e2709b03b0a60a660000e00f4c90f918858df0a730e67113a537b80b1f1f2c

Observation 3c72e548-1c77-44d7-a5c6-9aee68ec78be · outbound

This paper cites Heterogeneous adaptive policy optimization: Tailoring optimization to every token’s nature,.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous adaptive policy optimization: Tailoring optimization to every token’s nature,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.514418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:14d0de983d8a0bee7cf43e51c2ce1f5a0468053a83fef020e314f4cee3c494ca

Observation d853adc5-ed7e-40fa-8fcc-0e2ba6282484 · outbound

This paper cites Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.371938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:c3c01f2b22621be8ac3ba4e0b4711ac067a0d1883579183ffe2c9e75b92b060d

Observation d5004d3a-3942-49af-8c16-a1ef2d2c52fc · outbound

This paper cites AMC23: American mathematics competitions 2023 test set.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning AMC23: American mathematics competitions 2023 test set

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.512049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:70984b978c5cb0273f5c96279ec6fea56549caa4097937b8cdcdb6e347107cb0

Observation bf9449d8-30be-46c4-91ce-3ad63950abf8 · outbound

This paper cites Grpo- λ: Credit assignment improves llm reasoning, 2025.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Grpo- λ: Credit assignment improves llm reasoning, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.518110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:666e465abcbc03f4d66ee068bd6aa1f1ed5ef7b10f4391b05dd1a099b5f1cd5d

Observation 43e0cafc-c7ab-40f5-b264-040dd8c5f203 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.384161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:b7891a63b37278ac0483eee67f5675a461e673b86a8985908dff86f03cb3c8c9

Observation 6ab7f1fc-f10e-4e0c-a54e-7d8c23c246df · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.366684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:195df2336a5453d3187397524cfea3cc714c84d41a41a5b177db33ad92955783

Observation e66dfbd5-cc9b-494f-9305-edaa05cb4812 · outbound

This paper cites Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.360085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:1fa2fee2ed89643eaac17768493d519d4a62cdc4b17570ef0425e9cd9adfbe0e

Observation c4b2a6d7-bbc5-4613-9c20-d6e97f25b4ac · outbound

This paper cites verl: V olcano engine reinforcement learning for llms.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning verl: V olcano engine reinforcement learning for llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.517642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:6e7f7f483b74ccb1df5141ef1bbb39d6c8b61910db5567cd127df610fd8a9e6b

Observation a476cc92-eedb-4256-9b16-207b25f08939 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.499701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:ce409c81ccfe3619ebcd726a1a20c367327c840f3951a11564dd2d52eb05f1de

Observation dd27a7bb-7cf1-493e-a2b6-7b234622c20e · outbound

This paper cites Neural text generation with unlikelihood training.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Neural text generation with unlikelihood training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.497765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:932e17d554467632357940e000f5bff8f880a28ce942a98a01eacb64eae2fee4

Observation 8a490e27-cb6f-4625-a5fb-a8c44d8a8a0d · outbound

This paper cites Self-ensemble: Mitigating confidence distortion for large language models.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Self-ensemble: Mitigating confidence distortion for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.501760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:97ed2335668dcce9264c6d40dc69939bd96f6571a12764fa9cc97e1e05f9752a

Observation 226c4206-eba7-4215-ba25-34ba9a82457d · outbound

This paper cites an unresolved cited work.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-07-10T14:27:07.516170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:633efd44fbd10d62ea98780a38a03fdfdab0c3286d0ac972ae678710cc342b87

Observation e54299af-ac5a-4999-9d10-86526f99c937 · outbound

This paper cites Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.377916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:27aefda9e0aa6a3f497bea571184796bda11bf63c7a4be491a3d8a7fef851505

Observation be331791-00b3-4a40-975e-9490a234aa7a · outbound

This paper cites Qwen3 Technical Report.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Qwen3 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.365296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:16fb4b32cdc0f25eb2b21b923b4f179321ae55a7ace3654d384d0b706b3ba76d

Observation 103c5305-bbf0-425a-91ae-708c593246de · outbound

This paper cites Int: Self-proposed interventions enable credit assignment in llm reasoning.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Int: Self-proposed interventions enable credit assignment in llm reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.369395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:362a6bc3d3fd796ab1f16144ca4e30e0c0fbfd5aa7f5a985e54169cca028c2f4

Observation de480b0e-a72d-4343-9e30-f44cbd760865 · outbound

This paper cites Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.374977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:9cc4bff3688183219359637f8ea078fd9616530ca1d88d97955926f187aa2f50

Observation 552159d3-a99d-4695-b361-2ce62ababa40 · outbound

This paper cites American invitational mathematics examination (AIME) 2024.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.524036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:299ea18b096a01f68d4051d2a59209c1af775e2b7a4d179242ba3417f9992715

Observation f0ae9229-216b-4b9c-8ece-02462de4d7f8 · outbound

This paper cites American invitational mathematics examination (AIME) 2025.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.512495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:f6b34aec0d74caace45fb365fd967be287d69fd63533b9dcd1b0322cff028528

Observation d144154f-0194-4292-88c3-3a36c45b58a6 · outbound

This paper cites Group Sequence Policy Optimization.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 42

Resolution
malformed identifier
local_arxiv, observed 2026-07-10T14:27:07.148343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:15f6bfc493db2983d344938abcd9320a8b247209e589a6701a874c566c425cb6

Pith citing papers

No inbound Pith citation observations are available.