Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:24.786925Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 26 inbound Pith citation observations for arXiv:2506.06395.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:24.786925Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:08.381567Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:19:34.020533Z
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0385cc3-5f01-4916-aeb6-4aca1a275831 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bce929-3162-4235-bd94-72cf6baab1ad · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60935ee-dde6-4400-8c29-27d6730082b3 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a63697-9523-438e-9935-b50034c49f05 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e948f6-3226-4c58-8c9a-f9dac71cdc30 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training verifiers to solve math word problems, 2021
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34376507-921d-4aa4-84dd-2e75bb0394ab · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Evalchemy: Automatic evals for llms, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec6c734a-ce12-4135-972f-d7448f132af8 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd03621-def7-4ce8-8627-b0353544e97a · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2602154c-cc1e-43d2-9bf9-6bbd3b08a05d · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3d1a66-1da3-42e3-afa3-a46e03d8a014 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2afed7-b944-4fab-87b0-569a5b564642 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b3a446-3282-4565-b297-577d8e3a4712 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring mathematical problem solving with the math dataset.NeurIPS, 2021
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 744166e6-7752-41a8-9cfa-c805b0bcef76 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebfdc093-9a0a-4435-9381-159df18db4e3 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39af67f3-817a-4648-8860-cff01d02a93c · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-V3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc871d76-c04e-4477-881b-39ea09382062 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models ReFT: Reasoning with Reinforced Fine-Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8693aa70-2e9b-46b7-b29c-0809c2fc14ce · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f09791-8594-43aa-ac55-612b59c5fcf4 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9e2248-f6ed-4d03-98a9-0fae86185402 · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1635733-65a9-48b1-a985-b3287d33c85f · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c1b0b8-9e28-4160-ae5f-bdb20ac72d4b · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8665494-fd5b-4915-b6a6-ca21438b272f · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb293c4-1fa1-4ded-af98-3ac56e326fbc · outbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models TTRL: Test-Time Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b0ad1c-dfe7-4e4a-adae-ca65f8b91856 · inbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc426df-37f3-4055-877d-698b692486d3 · inbound
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b606fae2-3e58-4654-a66b-1e6442e05a64 · inbound
Revisiting LLM Reasoning via Information Bottleneck Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f726fd13-53e2-436d-a482-265d307afd9b · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 280
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b366e46-e0ea-4246-8b46-8beaf67c8df4 · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0657594b-20a4-42e4-84e5-f1d2760997d2 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c31a09d-140a-4ba0-b9d4-c89ed43e3847 · inbound
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb0f7f68-99bc-46f6-9fd1-a97d1f67acd1 · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80755571-d9f4-4a77-9222-4be56575688e · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a48785a-3065-4a6b-a1a4-f3a9b0ac43cd · inbound
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7a5f517b-1ed4-42c9-ab8b-c8c24149ceb9 · inbound
Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8e06bf-4516-44e3-9dd5-59aa413d6e3f · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1b5a1f0a-dbe7-4c35-8427-3e4795804469 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 136bdb10-f5df-4391-bbbf-5716b98d57f1 · inbound
Hallucinations Undermine Trust; Metacognition is a Way Forward Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2225a65a-24bb-45d7-918a-647dc569ecf7 · inbound
Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 91676cf1-799a-4ac2-8233-5761a453ac46 · inbound
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 65dc5c60-d756-42d2-8f0d-482551526225 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0adf9594-fec1-4beb-906f-ffcf2f486f1f · inbound
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0fb3c622-79d3-4020-b0d5-d46f3c71b144 · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4b05f7d-a7a9-483f-9ea8-8cd9c4ece2ea · inbound
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c1d205cb-439a-471a-a099-f5e14cd33be0 · inbound
Trust Region On-Policy Distillation Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3fcabd54-4272-4f59-805a-86b1790d9ca2 · inbound
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 83ccfe0e-b653-4d2c-9dac-b0f2cfa1cb3c · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a1e595f4-6dc3-404d-8ca8-e0eea46b1270 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bab38903-170f-4d40-b597-a461b8816476 · inbound
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · inbound
On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.