Pith. sign in

Paper Citation Record · LEDGER

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 26 inbound Pith citation observations for arXiv:2506.06395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06395 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:24.786925Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:08.381567Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:34.020533Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0385cc3-5f01-4916-aeb6-4aca1a275831 · outbound

This paper cites GPT-4 Technical Report.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.377877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.377877Z digest=sha256:930a66586f9b241fc41bedca765a067164b7078255a265864d72b4853679cedf

Observation 25bce929-3162-4235-bd94-72cf6baab1ad · outbound

This paper cites Qwen Technical Report.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.420167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.420167Z digest=sha256:794d15a29688b5d09e13583af9e113ae12ebac3ab58b9982d61e40be3588966e

Observation a60935ee-dde6-4400-8c29-27d6730082b3 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.520322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.520322Z digest=sha256:bc326f79c49434aff2c4dbb067186e0b56baba54e4bf2a1b664e6c2c0793d528

Observation a5a63697-9523-438e-9935-b50034c49f05 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.548791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.548791Z digest=sha256:8aa3d4fd3b598f385d85d12460fca6da5ad23d2224594138d877f3238a1737f9

Observation f3e948f6-3226-4c58-8c9a-f9dac71cdc30 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training verifiers to solve math word problems, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:25.145528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:24:24.573746Z digest=sha256:ef3a717306c2c39eee897326d1a325ab2ad3c2af55f6b7af662424475616d068

Observation 34376507-921d-4aa4-84dd-2e75bb0394ab · outbound

This paper cites Evalchemy: Automatic evals for llms, 2024.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Evalchemy: Automatic evals for llms, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:25.128092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:24:24.589221Z digest=sha256:b7716e12b4bded1c02313c1b9d122baa54a39be8dad873a39b23f2a45a04c6a0

Observation ec6c734a-ce12-4135-972f-d7448f132af8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.604801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.604801Z digest=sha256:45d49daaefa5c74f37822fc953285e7c05d33bf6a95394d5608fa9ea59944f51

Observation acd03621-def7-4ce8-8627-b0353544e97a · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.614813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.614813Z digest=sha256:561e99263cc9683c342186f19ad94d1b42c153c404df9aa2631c934c7593de85

Observation 2602154c-cc1e-43d2-9bf9-6bbd3b08a05d · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.624793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.624793Z digest=sha256:14c94f5d12c0152914de0e8cf8f87ba8c20c839f8ed367d417b1b5ed71f78abd

Observation bf3d1a66-1da3-42e3-afa3-a46e03d8a014 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.633027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.633027Z digest=sha256:de5ac60f97c806244682270a0327429d8e658dd31bbd17d331ec8034b29f38f3

Observation ca2afed7-b944-4fab-87b0-569a5b564642 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.643359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.643359Z digest=sha256:fb4f51665c22d38fab72539ef307d81ca760578413c48b35f34fa9a661a73583

Observation d2b3a446-3282-4565-b297-577d8e3a4712 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.NeurIPS, 2021.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Measuring mathematical problem solving with the math dataset.NeurIPS, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.652806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.652806Z digest=sha256:183bd89c02e84c98c520038bb0c2ad89403f80cb9866edd13e8d7e1bd62a7b82

Observation 744166e6-7752-41a8-9cfa-c805b0bcef76 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.663552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.663552Z digest=sha256:51b62eccaf710b40550e91a1d56794072e5a1dc941a3eed44317366ab18cca1e

Observation ebfdc093-9a0a-4435-9381-159df18db4e3 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.685376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.685376Z digest=sha256:39321f4a01952aed67516100368b8a79876a5412e459aa7a4dff802f3e41322c

Observation 39af67f3-817a-4648-8860-cff01d02a93c · outbound

This paper cites DeepSeek-V3 Technical Report.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.692747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.692747Z digest=sha256:b237d38585d44d9ee186a8549fe1ba953726777151fa79a1c821ecf1df009776

Observation bc871d76-c04e-4477-881b-39ea09382062 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models ReFT: Reasoning with Reinforced Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.701431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.701431Z digest=sha256:49e2fa06ed58847d7e745ff554d916a4ac2b29b0c77762fae9b31cc05e09803b

Observation 8693aa70-2e9b-46b7-b29c-0809c2fc14ce · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.710826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.710826Z digest=sha256:ca0bcacf0a3844f0c3dcc9d76ceccafb26752e921bf76d7eaae409cd28d09574

Observation c3f09791-8594-43aa-ac55-612b59c5fcf4 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.721564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.721564Z digest=sha256:d205485310b420098231fe957667c5f99a252a474cf8978e63f00f1a51c2ad0c

Observation 6d9e2248-f6ed-4d03-98a9-0fae86185402 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.731345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.731345Z digest=sha256:ed17357c0b51ad5c609c5517382466289b5ec33a39c5da31584abb400edc64af

Observation a1635733-65a9-48b1-a985-b3287d33c85f · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.744921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.744921Z digest=sha256:6f7c232a7767187fc0fc762c36ab24f4d721e457f1d6a7db4607336151b6c0a2

Observation 54c1b0b8-9e28-4160-ae5f-bdb20ac72d4b · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.758614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.758614Z digest=sha256:d0d3f26ec26ee166bb89e03033c9933b55e7d9605e165177c38af5e11fe62c5e

Observation e8665494-fd5b-4915-b6a6-ca21438b272f · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.773742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.773742Z digest=sha256:0b392e842f1c4909d5ae31f137a1608672ba146a9a5387cee97c3865dfa66deb

Observation 0cb293c4-1fa1-4ded-af98-3ac56e326fbc · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models TTRL: Test-Time Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.786925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.786925Z digest=sha256:4939f9b69767ba0ace09d7a5d92a2e97759492ce7dc988487d48c9b8afa75797

Pith citing papers

Observation a7b0ad1c-dfe7-4e4a-adae-ca65f8b91856 · inbound

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning cites this paper.

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:24.315457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:24.315457Z digest=sha256:fa5f7c734ecab27ddcc32d40b3c64320c8558b6b6139a06dfabc625319246f7a

Observation bfc426df-37f3-4055-877d-698b692486d3 · inbound

Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning cites this paper.

Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:52.966250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:52.966250Z digest=sha256:f29a3b2ca2775ee11b08eda76579cb802cb289904c274e15fdfd5d5e8838f508

Observation b606fae2-3e58-4654-a66b-1e6442e05a64 · inbound

Revisiting LLM Reasoning via Information Bottleneck cites this paper.

Revisiting LLM Reasoning via Information Bottleneck Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.381567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.381567Z digest=sha256:838d6c2d44be6c7647105cd8a09ce7f7e78d9c3cd94c440c81a3a04e5d217537

Observation f726fd13-53e2-436d-a482-265d307afd9b · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 280

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.731798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:90df3c1ff17bad3f194cd24b44e0f929557a02a80b081e4980a74670f4165a46

Observation 0b366e46-e0ea-4246-8b46-8beaf67c8df4 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:46:34.049081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:b2a1abdd4c4f106ac8c10f21b3e886438a3f8fe645cde2b1592f2d907ece4034

Observation 0657594b-20a4-42e4-84e5-f1d2760997d2 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.243783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.243783Z digest=sha256:e5ce43fea12a80abce658ee6ef95edc2998e80413e87e39636b38a246aff1bf6

Observation 5c31a09d-140a-4ba0-b9d4-c89ed43e3847 · inbound

Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking cites this paper.

Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T13:42:31.638150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:42:31.638150Z digest=sha256:47cde49792a5516cec6794a666cd6c46472c70c216bd2a2ffebb526f41335f3d

Observation eb0f7f68-99bc-46f6-9fd1-a97d1f67acd1 · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:30.709604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:30.709604Z digest=sha256:862e5f35d43f7febafecba3a1e22c4ed760dbf8cde3377a17abf1794c44a4ef6

Observation 80755571-d9f4-4a77-9222-4be56575688e · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:19.263434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:19.263434Z digest=sha256:e039e7cc51ee9e59ece0744911e985a87ea22ae3b205fc9ae0ed32216ef99566

Observation 0a48785a-3065-4a6b-a1a4-f3a9b0ac43cd · inbound

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards cites this paper.

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:50:02.385698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T13:49:11.758336Z digest=sha256:f6c5856e2273951e3804a61a892309db254af9a5f36b09c1a199f9343a5d2738

Observation 7a5f517b-1ed4-42c9-ab8b-c8c24149ceb9 · inbound

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care cites this paper.

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T20:22:12.729190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:22:12.729190Z digest=sha256:e6b684245476b5aaff368bb304c705c2b16706777f977deb63ecfe5aa238a3ec

Observation 5b8e06bf-4516-44e3-9dd5-59aa413d6e3f · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:08:01.328397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:fb84d88464c5702380ca2b1685ca2a5ff30141c5e05eea30f9b01b77221762fd

Observation 1b5a1f0a-dbe7-4c35-8427-3e4795804469 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.520524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:bab7e06328a9a02336bca670fe7a6dd650332137cdc361584057a985d6df592e

Observation 136bdb10-f5df-4391-bbbf-5716b98d57f1 · inbound

Hallucinations Undermine Trust; Metacognition is a Way Forward cites this paper.

Hallucinations Undermine Trust; Metacognition is a Way Forward Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.150385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T14:29:54.924293Z digest=sha256:1e1de5dd64b6825024c0cc03b2c4b02c8b78ad5eba5e779c8331fb70245bcdba

Observation 2225a65a-24bb-45d7-918a-647dc569ecf7 · inbound

Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models cites this paper.

Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:21:08.878934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T17:20:19.586214Z digest=sha256:90480cc11ac4ff360b7fd5083eba492795b6d430c65c05b9787163a7944ea6ee

Observation 91676cf1-799a-4ac2-8233-5761a453ac46 · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:45:59.928030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:b5951e511779a5939a1f3357c8f046136d6d4d0d8d8abfd4fe77bb38072b229f

Observation 65dc5c60-d756-42d2-8f0d-482551526225 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:00:54.981604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:093807b427b20b54ea7c22cf285624d4cb65cc84f055dfc68b647d42dc4fb02b

Observation 0adf9594-fec1-4beb-906f-ffcf2f486f1f · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:22:50.601594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:c07614152fa20c18d050a1992c4d09f7defbf265f6d61e57d7fa3ad8cbe94993

Observation 0fb3c622-79d3-4020-b0d5-d46f3c71b144 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.052901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:d410220395e2c19aee292ae067ef589716fb98dddeed895b2e22ce187ab6b817

Observation f4b05f7d-a7a9-483f-9ea8-8cd9c4ece2ea · inbound

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting cites this paper.

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:06.927905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T07:20:50.835826Z digest=sha256:8413bf63c5aef6a8bcf9d43f44f36bbab2fc5dd397e192fa8f4eae55655ee552

Observation c1d205cb-439a-471a-a099-f5e14cd33be0 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.714700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:6f2837523c4feac1874fb8d534938ba9462d3a0a54cf7090d69f7a33d847718e

Observation 3fcabd54-4272-4f59-805a-86b1790d9ca2 · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.294521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:c7a70aa9085f7039061a2efb12f087973bb3d498b9f26a27c331b64d6331df10

Observation 83ccfe0e-b653-4d2c-9dac-b0f2cfa1cb3c · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:40.756454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:7cc2892af3e308bdc32594d2e59e8687cd8760816c0c5a931cbd2fa2be88feb8

Observation a1e595f4-6dc3-404d-8ca8-e0eea46b1270 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.871201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:e07f97cbbe17ff482e0f08d24d288cacba54aa2d8ed0ce4d7b8274e9d53b5951

Observation bab38903-170f-4d40-b597-a461b8816476 · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.022356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:102d21c3ad997bbdcd9d2bda4587e4adefa3107350d2f8afe601356003722c6e

Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · inbound

On-Policy Self-Distillation without Any Supervision cites this paper.

On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.827630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.827630Z digest=sha256:f561e8d157792895b9518e9a3c23bad6847e1e0e456a4d7d61acedea709db01f