Pith. sign in

Paper Citation Record · LEDGER

The Hallucination Tax of Reinforcement Finetuning

As of 21 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 9 inbound Pith citation observations for arXiv:2505.13988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13988 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:36.534742Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:52:50.299350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.598004Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee8a17e6-6a68-4f1f-a790-a77f4d267511 · outbound

This paper cites online" 'onlinestring :=.

The Hallucination Tax of Reinforcement Finetuning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:29.905006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:29.905006Z digest=sha256:439a363f5aa73f0dffda22c78b951a8da2b77c448cbd48efcc11ae1679d5b373

Observation 1bd65997-4777-4fe7-9e09-495e03bcfbc3 · outbound

This paper cites write newline.

The Hallucination Tax of Reinforcement Finetuning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.081158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.081158Z digest=sha256:b84a3ee9fbe2b61aa48ea791c04523ad7439bdb52dc751fbd614ac029a09ad83

Observation 995a819f-5fbc-4bb5-90c7-4cf4c8a0ac33 · outbound

This paper cites Phi-4-reasoning Technical Report.

The Hallucination Tax of Reinforcement Finetuning Phi-4-reasoning Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.194934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.194934Z digest=sha256:ce3d5fd9b53057efd5595fb621971c6d3c8ceb8b94e6dfca7b5ce396b4620668

Observation 6c6d2e74-fa0b-4a8b-bb16-c419cb0eefa0 · outbound

This paper cites HalluLens: LLM Hallucination Benchmark.

The Hallucination Tax of Reinforcement Finetuning HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.304843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.304843Z digest=sha256:c0890a43d3214e4798f4f83e81f6b8e85e189a38c47ac9c78c11538b8dd1a1d8

Observation 24326b6b-640a-4b9a-abe6-962ebb645b7e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.416483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.416483Z digest=sha256:b8d233cfe86ad388db9918268349fd4ecbf5ab4c5f73f6ac1b277d8737d7983f

Observation 5c1703b3-86d1-46ea-9da9-93516346b420 · outbound

This paper cites Mitigating Open-Vocabulary Caption Hallucinations.

The Hallucination Tax of Reinforcement Finetuning Mitigating Open-Vocabulary Caption Hallucinations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.575619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.575619Z digest=sha256:e8adbdc961536abddcee180b68481a88c31179ad6ece98129ceb0a9b30b5ea9d

Observation e0c669cd-5eb0-4592-82c1-95bfa02b08bf · outbound

This paper cites Evaluating Hallucinations in Chinese Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Evaluating Hallucinations in Chinese Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.675099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.675099Z digest=sha256:cf9dabb36cb0c483231bdbc0455856f5f15b34052b1a980b4548b0bed6f58300

Observation 41939ca4-a260-4759-91f5-d40a2cfe6b20 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Hallucination Tax of Reinforcement Finetuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.756069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.756069Z digest=sha256:7f1b2eaf72c0b05451267f2cdbdf64ea0e3f10684121ca0b13405071f26e0f73

Observation 3a56fac1-4f09-493f-ad1f-d88b1ad02e51 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.864841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.864841Z digest=sha256:638ccf2215350678fdfe7e9473e71fded966c4d229171892738735e2b3082030

Observation 71483c38-9e7a-49ef-8260-ff5294e1f711 · outbound

This paper cites Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?.

The Hallucination Tax of Reinforcement Finetuning Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.995190Z digest=sha256:84d5d39f1e5efdd380bd9c8e6086c705840a86707183ff383d02220938f30567

Observation 854c0351-2a20-4b30-92cb-d939b26418a0 · outbound

This paper cites The Llama 3 Herd of Models.

The Hallucination Tax of Reinforcement Finetuning The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.106531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.106531Z digest=sha256:03c5513b3dff3819e3af99117865220ceb7c5ba98a7b5af7bb9f7dbb4f8a94a7

Observation ce49eb88-0866-4122-8767-c1fdbd6c4c21 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Hallucination Tax of Reinforcement Finetuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.222493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.222493Z digest=sha256:dae70c85a7ed117c324027116984240dc8d11c0afc3b6f13b9c11db752cded91

Observation 819102f4-c5fd-4cb0-b98d-60b315192ed4 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

The Hallucination Tax of Reinforcement Finetuning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.355068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.355068Z digest=sha256:6561ccf6f50cf4d3b726f5f23b5eaef4e7524cf36b7347280535c27dd409930d

Observation f5f2b607-eb5a-428d-ac1d-fb36a23b78d3 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.447643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.447643Z digest=sha256:957198d4fdc2bc6d9bdeca55180f96b5878f5776792b1b4f45fffbc95c243bde

Observation 8f7481b6-dca4-4af9-aaed-449f5323eaa3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

The Hallucination Tax of Reinforcement Finetuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.518305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.518305Z digest=sha256:a9d2e6684227a1202d89dcb1e6b318d8fc6c65104abf6b5f8b5024f3bd3fa265

Observation 84052823-48f1-4594-8298-fe0688dfcfe8 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

The Hallucination Tax of Reinforcement Finetuning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.690376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.690376Z digest=sha256:77a30af6131e8ed07244393808037d30e5c5b6992d88d2fe8476043ca7869f82

Observation 2b37d6ff-c0f1-4c99-a109-81665178032f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:43.373694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:31.804578Z digest=sha256:2349b8a01e68b178476ccf8662fd32d3b31dbfd8ad63ede7a3fbb101942aa65e

Observation 25972b8f-464b-49e8-94e8-f8f5994bcaff · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

The Hallucination Tax of Reinforcement Finetuning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.005515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.005515Z digest=sha256:f776dc747c5a60f46f4d5cb097738503c1a543433246450eb21d29d2638295e2

Observation 5f1ffac1-a515-450a-8696-fd8c930e9a5e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.124833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.124833Z digest=sha256:8a8a9978a9d5ef766aab4c8fdb6a4c2bbd638d822aff0f92043607eab99ae752

Observation 911ec246-e3ca-4ef7-a54b-253350311b84 · outbound

This paper cites The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.201925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.201925Z digest=sha256:138abe73a118097d19a15c70fa6d455c953c90860893fcb9594ae96d159b0da4

Observation 27b411fc-cbc0-4d0d-9643-2eeee438d1af · outbound

This paper cites Treble Counterfactual VLMs: A Causal Approach to Hallucination.

The Hallucination Tax of Reinforcement Finetuning Treble Counterfactual VLMs: A Causal Approach to Hallucination

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:44:37.888022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.260799Z digest=sha256:912b80a8e3e3de02c74a5b8fc3fff762427072c0890ec31f1a3052e3653901f4

Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · outbound

This paper cites LIMR: Less is More for RL Scaling.

The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.385679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.385679Z digest=sha256:f1cbd40c7f2c626b4d2b16c618dfae7cc07e82b75ea569efadeb66c11135ae5d

Observation bc76dd61-de76-4e95-b56a-b6224cd025a2 · outbound

This paper cites Let's Verify Step by Step.

The Hallucination Tax of Reinforcement Finetuning Let's Verify Step by Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.534874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.534874Z digest=sha256:33faaff4cfc0ba7a06306e33fee0acc58b84f37df96b4947179ca0475f094df8

Observation 711946dc-cdd6-42db-a1c6-d3ed63c0ed2f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.681966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.681966Z digest=sha256:ebd76b72b326fe8021ddd1c3ef714ced87aec6fbc9bc33b1cc03d70ab1baac13

Observation 152c4b32-31b7-4b31-b8a7-ce1516fce7a9 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

The Hallucination Tax of Reinforcement Finetuning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:44:43.042890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.794896Z digest=sha256:83c8edd7a91536eb033d32ade56e627f77cd4094c0e51b07e117094a2d6ee1f1

Observation 655dcc39-5a30-44d3-ab7d-6c4fd97b7cc5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.954751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.954751Z digest=sha256:cabacff6cd6aa739452367c19b793e802f1636d378c169cd8a96e6fdfdd35fcf

Observation 10c551ec-e18d-4b78-987f-de8e7b0ceff8 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.634345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.126414Z digest=sha256:bb8657ef7b80121f8bd3c574998e4c8e551bf2c169dac3ae66146b0412e48c88

Observation 2ee8666a-b3cf-42e7-b1d0-e4cff21c07ea · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.366454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.324958Z digest=sha256:907818c51ac5f353e5bfa9d69d0b3f0156d4b2e6098d7057556b1bcf7785db6d

Observation e825bd42-c4fd-4479-9e05-a2f9c9a63638 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.545500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.545500Z digest=sha256:3d1cf8d547de16b8fd9c14db6440c03bd9f1b65913ce1a5787ccfa6e1ddb30c8

Observation 4e638485-7ee6-4076-8934-97bbd59cd17a · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.964745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.628322Z digest=sha256:6e41d2e913d33852a5d70acee516e09625b900d2c887b640c9858f39d053d44d

Observation be302502-a318-4dc2-87ee-44644080a929 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Hallucination Tax of Reinforcement Finetuning Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.782472Z digest=sha256:e6928a51e1389cea9537d9a8891b9cd73783d521d5b1b328cf2e8e6275e15ed7

Observation 7c226e53-6050-4b90-b6c2-edf442f7a5d1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Hallucination Tax of Reinforcement Finetuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.894983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.894983Z digest=sha256:2facbd9673b3721d75ef681c7369c839ce7074c971a4dbeaa1d1c6e263df5f71

Observation 16abeb0e-4769-4dee-bee0-62ab852816d5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.007922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.007922Z digest=sha256:c55998e3fdf276e9434ca92715bf8f264d682012a1b1e3ea83ff73db3d5e6c85

Observation b6932d08-7faf-4c4a-9d61-9ad113926933 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

The Hallucination Tax of Reinforcement Finetuning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.170709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.170709Z digest=sha256:1a7f7af89e666754ac36cb9a15774c895b84bdb4f93631876e48b98094b318b1

Observation d8527608-d5de-4952-b276-8553a0ab47df · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.272336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.272336Z digest=sha256:eebc84a354275b3a68648f3189b693d19c6e8ad4dba319d9d074cb79bb1cc2a1

Observation ed215d06-63d0-4e1a-96ed-001d416d4174 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.595034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.448645Z digest=sha256:814056e557a4da30a203e8582bf9eace9f39b1f958eccecc4549dc3ce88a9460

Observation 1e62e25e-fcd5-4aa5-9bbe-597e3bacba13 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

The Hallucination Tax of Reinforcement Finetuning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.610778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.610778Z digest=sha256:f0c25e5f2d24638ca30d4142fa88ac87204c343d164084e4d03aab054a561544

Observation 1e57963e-b949-4603-86b2-f160291ca989 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.284869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.775008Z digest=sha256:4635e46e7b3658cd23eda800cf7df6fa5a356ee8740dd62752f7c9496ae9708c

Observation fdd7360a-d90f-4de7-905a-48d9044a6cfb · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.912984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.912984Z digest=sha256:d3b532dd2efa4c9faffd1006970cb4c1196aa35a358872cde5f54b5b8b4928ba

Observation 1e86e968-b904-4aa8-af69-2222fd6b1be5 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Hallucination Tax of Reinforcement Finetuning AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.037495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.037495Z digest=sha256:7c5b77888f1d6d4817e61482589881c812d31977824eccb336b6a44ed4ad57d3

Observation 9e7e68fb-7fad-46bd-a097-8001bb4f84c5 · outbound

This paper cites Tina: Tiny Reasoning Models via LoRA.

The Hallucination Tax of Reinforcement Finetuning Tina: Tiny Reasoning Models via LoRA

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.167445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.167445Z digest=sha256:f38e03f3195a392def215983c95bca96f3eae28c9853fd788c8d681028dc9956

Observation ec9d1ffe-a12c-470d-9379-96f40be417bc · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

The Hallucination Tax of Reinforcement Finetuning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.310902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.310902Z digest=sha256:3528e724e1da70ff7c29fdf4a8272978990fefd9916d3e9969c87d02c51b4ab0

Observation 176734b8-7600-48fd-a725-72d93dc4c722 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

The Hallucination Tax of Reinforcement Finetuning Simple synthetic data reduces sycophancy in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.430765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.430765Z digest=sha256:087114daf2890a8aa5b739b6717c2a1a690f033c98509702efe32d463701bd50

Observation 9aebf6e4-c24a-47e1-9f0a-9503c9f6beb1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.994636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.515063Z digest=sha256:9c75862bf4cefdeab7ee858db9f578f1a37535854a147deb20513e114a144982

Observation 007b1964-85c9-4886-b1c9-f9788c4ad213 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

The Hallucination Tax of Reinforcement Finetuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.621729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.621729Z digest=sha256:57ffa0efe618d649b4034ff4d974c05079e6733b92845090a511ad7489b373bb

Observation 7be9aab1-c81d-42c3-9abd-b00308dcd272 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.628475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.719317Z digest=sha256:fbc95e9306d1383dc5cba68df3fcdaff641c018d6ca0ec2fedd0ae388f73274d

Observation 16848e90-3df9-41b2-bccf-5254ad9f571f · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

The Hallucination Tax of Reinforcement Finetuning Do Large Language Models Know What They Don't Know?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.850344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.850344Z digest=sha256:85c32033aa63426ed752f4b1592e4aa4e72e8f4c655d835ee1830d39a4485915

Observation b778e018-1636-4f23-891f-dbaaaa36b509 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Hallucination Tax of Reinforcement Finetuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.961914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.961914Z digest=sha256:76dcb45c69477a8ab9bb7287af11e5930f64a6b40fba27223958d6f034bd0575

Observation 9a8df657-c924-449c-b6ab-22dd12a69bd1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.383837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T15:44:36.094973Z digest=sha256:a1673cdffcefde90f4d7f8362c27bff941214839ed876b7f8990fbafaf629540

Observation 9268f459-e4da-4b2a-901c-79f32205026e · outbound

This paper cites How Language Model Hallucinations Can Snowball.

The Hallucination Tax of Reinforcement Finetuning How Language Model Hallucinations Can Snowball

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.213918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.213918Z digest=sha256:ecb44c5a8b5cdfb76d7b7064b893dc00bfdc0fe163a701919f10246de760582c

Observation f32edaa0-bd51-4cb7-81a7-533956c446c1 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

The Hallucination Tax of Reinforcement Finetuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.354899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.354899Z digest=sha256:192f30218e079cfff195c195db380139c80f8fe538f15692b02ab129521fe52a

Observation 13000c1c-0526-4a67-bb6d-f74ffe031b72 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

The Hallucination Tax of Reinforcement Finetuning Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.534742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.534742Z digest=sha256:51fec8a5b41b3de8954bbf9a5eb608b3d0a534c9bfb759e0cf6b56f91d5be803

Pith citing papers

Observation f96d3bbd-0892-49ad-9c12-ce012dde8f0d · inbound

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration cites this paper.

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration The Hallucination Tax of Reinforcement Finetuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:52:50.299350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:52:50.299350Z digest=sha256:809fae0598e947f98a6b72d7e58f44b9220ac5369ff1f46e896de63c576b0392

Observation d1ed60c3-31ad-4d33-8b55-7c3267e55276 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.113610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.113610Z digest=sha256:6f09be1eba562d08568ecaf189f8c1621039bc70c61ede59e8e3a90c066d7697

Observation 1d18fcaa-1364-4c4f-9724-4d6f3d7debac · inbound

MoCo: A One-Stop Shop for Model Collaboration Research cites this paper.

MoCo: A One-Stop Shop for Model Collaboration Research The Hallucination Tax of Reinforcement Finetuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.795988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T10:17:37.129753Z digest=sha256:3951884e5f4f7c23ec1bcd0680cd4ed49aec4a29c37bea3b98b0c2d830622c21

Observation 7a4542a9-ef81-404b-97c1-25a1b7813634 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.704183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:2b28157baecea94c4b0e51cd6b1f478844226d0f3fbdf41112f15724b5a20e0e

Observation 92e46cd5-8b6e-4d71-83a9-2633117e1cc6 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.560735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:7cd34395cd0e7f068d842dc1ae4780c183b20bd7ef65a86914a8b5efb5747fd9

Observation 1686a16f-1bbe-407e-b0bd-10c6fe6c2dc3 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.737194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:a25689c13c1506050720582f6845550041226f1402505febbe9ab94254c44ddd

Observation d9620be2-8cc6-45fd-88ce-ff858571b6e4 · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information The Hallucination Tax of Reinforcement Finetuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.821929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:d2203aca024705e77b5bdbee89c5e212e55f5c0ae158eb4fcd7663f5c5960f62

Observation 74983e86-3310-4efc-aab7-7357ede6e45f · inbound

Scaling Participation in Modular AI Systems cites this paper.

Scaling Participation in Modular AI Systems The Hallucination Tax of Reinforcement Finetuning

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:16.599301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:49:27.042616Z digest=sha256:460aef3b20763ad855087122b9b19a22178f8fbf59215261d2fde355e3bf364d

Observation a0a77b3d-b3f9-4a9d-8d75-9ebe41e5c81a · inbound

Mechanistic Attention Guidance for Agent Memory Refinement cites this paper.

Mechanistic Attention Guidance for Agent Memory Refinement The Hallucination Tax of Reinforcement Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:30:40.711089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:30:40.711089Z digest=sha256:295085a6a9f425fb6bda0dafd0efa075517931c454a9f78920cad1e54f2c2ba8