Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.07976.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6f939aa2-c97f-4ece-a945-caa962951124 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 62f0072d-42cc-49d5-8032-ab5669f3fbfa · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 889c4786-37d5-48d0-b766-416aac88aafa · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning The curious case of neural text degeneration
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 371fe2a7-ef04-4efb-918f-889151363546 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Large language models cannot self-correct reasoning yet
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ca50a82-9f01-4344-85cb-904541018532 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning URLhttps://openreview.net/forum?id=IkmD3fKBPQ
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d9cacb44-7e90-4b73-813b-888370759ec5 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Execution-grounded credit assignment for grpo in code generation.arXiv preprint arXiv:2603.16158, 2026
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f841fa7-bbd9-4c3c-80e9-94cd255008cc · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6b3cc15-1af5-43fc-a7ed-70a27c62fdc7 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Let's Verify Step by Step
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3c72e548-1c77-44d7-a5c6-9aee68ec78be · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous adaptive policy optimization: Tailoring optimization to every token’s nature,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d853adc5-ed7e-40fa-8fcc-0e2ba6282484 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5004d3a-3942-49af-8c16-a1ef2d2c52fc · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning AMC23: American mathematics competitions 2023 test set
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf9449d8-30be-46c4-91ce-3ad63950abf8 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Grpo- λ: Credit assignment improves llm reasoning, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 43e0cafc-c7ab-40f5-b264-040dd8c5f203 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6ab7f1fc-f10e-4e0c-a54e-7d8c23c246df · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e66dfbd5-cc9b-494f-9305-edaa05cb4812 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c4b2a6d7-bbc5-4613-9c20-d6e97f25b4ac · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning verl: V olcano engine reinforcement learning for llms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a476cc92-eedb-4256-9b16-207b25f08939 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd27a7bb-7cf1-493e-a2b6-7b234622c20e · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Neural text generation with unlikelihood training
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8a490e27-cb6f-4625-a5fb-a8c44d8a8a0d · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Self-ensemble: Mitigating confidence distortion for large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 226c4206-eba7-4215-ba25-34ba9a82457d · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e54299af-ac5a-4999-9d10-86526f99c937 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation be331791-00b3-4a40-975e-9490a234aa7a · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Qwen3 Technical Report
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 103c5305-bbf0-425a-91ae-708c593246de · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Int: Self-proposed interventions enable credit assignment in llm reasoning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation de480b0e-a72d-4343-9e30-f44cbd760865 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 552159d3-a99d-4695-b361-2ce62ababa40 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f0ae9229-216b-4b9c-8ece-02462de4d7f8 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d144154f-0194-4292-88c3-3a36c45b58a6 · outbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Group Sequence Policy Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.