Pith. sign in

Paper Citation Record · LEDGER

The Hallucination Tax of Reinforcement Finetuning

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 9 inbound Pith citation observations for arXiv:2505.13988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13988 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:36.534742Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:52:50.299350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.598004Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee8a17e6-6a68-4f1f-a790-a77f4d267511 · outbound

This paper cites online" 'onlinestring :=.

The Hallucination Tax of Reinforcement Finetuning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:29.905006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:29.905006Z digest=sha256:def7828936c3b6c31268d95d410306c8994a0f79665f2f56d18d9f30a0807a45

Observation 1bd65997-4777-4fe7-9e09-495e03bcfbc3 · outbound

This paper cites write newline.

The Hallucination Tax of Reinforcement Finetuning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.081158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.081158Z digest=sha256:b94f1913c37e217bed9ab32a8dd2ff104abc99f6d06fb55ab9d26d12c479874e

Observation 995a819f-5fbc-4bb5-90c7-4cf4c8a0ac33 · outbound

This paper cites Phi-4-reasoning Technical Report.

The Hallucination Tax of Reinforcement Finetuning Phi-4-reasoning Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.194934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.194934Z digest=sha256:54831c1a0e81e93be68d104b6ec5c04b586af3a63f74cbf3fcf90d61fbb0c934

Observation 6c6d2e74-fa0b-4a8b-bb16-c419cb0eefa0 · outbound

This paper cites HalluLens: LLM Hallucination Benchmark.

The Hallucination Tax of Reinforcement Finetuning HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.304843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.304843Z digest=sha256:aa1872c82c5b9ea430fd1cfc1a487d5a2faec90d2cdfe36fa86d49ba30a90e38

Observation 24326b6b-640a-4b9a-abe6-962ebb645b7e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.416483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.416483Z digest=sha256:d6a35b9fea086b375677167f2df11444426562fa159247c36c482f75fa6ed6ff

Observation 5c1703b3-86d1-46ea-9da9-93516346b420 · outbound

This paper cites Mitigating Open-Vocabulary Caption Hallucinations.

The Hallucination Tax of Reinforcement Finetuning Mitigating Open-Vocabulary Caption Hallucinations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.575619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.575619Z digest=sha256:82c89633f122e6538afd15751f25ded202bfc45a0bf0448fef08e2897a3cc370

Observation e0c669cd-5eb0-4592-82c1-95bfa02b08bf · outbound

This paper cites Evaluating Hallucinations in Chinese Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Evaluating Hallucinations in Chinese Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.675099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.675099Z digest=sha256:bbaf049c37b39824b2e506ed19c36c215ba69b6c8e551dc3b38e5253dd54cc03

Observation 41939ca4-a260-4759-91f5-d40a2cfe6b20 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Hallucination Tax of Reinforcement Finetuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.756069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.756069Z digest=sha256:5419831d0e4282a8a0c3a9dc3782099d667bbe0525d3c5ce112923b9c82b42df

Observation 3a56fac1-4f09-493f-ad1f-d88b1ad02e51 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.864841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.864841Z digest=sha256:c4fd19033bb84928aec7b0434c076c80b56f4962cdf2b5b6dcb9eaec8622ce3f

Observation 71483c38-9e7a-49ef-8260-ff5294e1f711 · outbound

This paper cites Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?.

The Hallucination Tax of Reinforcement Finetuning Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.995190Z digest=sha256:5c0254325404b76a6039a7feb49592477efcb84df9f9bef95ae08a65f9acacc0

Observation 854c0351-2a20-4b30-92cb-d939b26418a0 · outbound

This paper cites The Llama 3 Herd of Models.

The Hallucination Tax of Reinforcement Finetuning The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.106531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.106531Z digest=sha256:e24f0d90de6191329b84ecbaf4072311643c03d1137c42fe310e8cdd4cdad264

Observation ce49eb88-0866-4122-8767-c1fdbd6c4c21 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Hallucination Tax of Reinforcement Finetuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.222493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.222493Z digest=sha256:6a52efc36704b45d018c194bbb459cba71fad32c156e29523115d3672cc6ae82

Observation 819102f4-c5fd-4cb0-b98d-60b315192ed4 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

The Hallucination Tax of Reinforcement Finetuning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.355068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.355068Z digest=sha256:f6f8bd71c45445aa29696da002b7689016f574a45824cc78b75d3350d35a4140

Observation f5f2b607-eb5a-428d-ac1d-fb36a23b78d3 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.447643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.447643Z digest=sha256:02d07a4c253ccfe9ccb3938e52f690b2fc6afbdbb28909ff9ea6d74d47842e61

Observation 8f7481b6-dca4-4af9-aaed-449f5323eaa3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

The Hallucination Tax of Reinforcement Finetuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.518305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.518305Z digest=sha256:58a49a7f6272f8eff0845d99859deb84d18dc84875dfefc393abd7aafac5f52b

Observation 84052823-48f1-4594-8298-fe0688dfcfe8 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

The Hallucination Tax of Reinforcement Finetuning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.690376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.690376Z digest=sha256:3ef99c43e61dd8aa79fcf1dd8e34aa2beab69c163246cb7e406ce3d1c28faeb2

Observation 2b37d6ff-c0f1-4c99-a109-81665178032f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:43.373694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:31.804578Z digest=sha256:2bd63af7764ee2d551093dc7c9f4e08582d3bfa66d7a15726a51057611f67a36

Observation 25972b8f-464b-49e8-94e8-f8f5994bcaff · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

The Hallucination Tax of Reinforcement Finetuning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.005515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.005515Z digest=sha256:57b441dfa13e78ec6ad17d9b2e89f3d856d7012c5f6d6a1e4716cd35b8c738b2

Observation 5f1ffac1-a515-450a-8696-fd8c930e9a5e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.124833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.124833Z digest=sha256:33ef8c1ac878e47597c72256d364431d06cd53cd405657d2c7d7a715d01df862

Observation 911ec246-e3ca-4ef7-a54b-253350311b84 · outbound

This paper cites The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.201925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.201925Z digest=sha256:74b950adfc4d2153d64e28822b23f4ff662209df2127a42c85582d45b8a7c9ce

Observation 27b411fc-cbc0-4d0d-9643-2eeee438d1af · outbound

This paper cites Treble Counterfactual VLMs: A Causal Approach to Hallucination.

The Hallucination Tax of Reinforcement Finetuning Treble Counterfactual VLMs: A Causal Approach to Hallucination

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:44:37.888022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.260799Z digest=sha256:c0a1887bc823ebd6219be097bab9ad43e96f1639fbf31bdf7f14541d0db06abd

Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · outbound

This paper cites LIMR: Less is More for RL Scaling.

The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.385679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.385679Z digest=sha256:1211678bebb8bc286b7e5ddc734a93ad28dd4a072bd607e8b698cd6575207dfe

Observation bc76dd61-de76-4e95-b56a-b6224cd025a2 · outbound

This paper cites Let's Verify Step by Step.

The Hallucination Tax of Reinforcement Finetuning Let's Verify Step by Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.534874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.534874Z digest=sha256:77ba96a6619a4ef46f8e71aa8c0add836d1222e14c05d3b2f7451b97185da166

Observation 711946dc-cdd6-42db-a1c6-d3ed63c0ed2f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.681966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.681966Z digest=sha256:897fd465326d548e270d93a86969235c6f5a6be712fea2be115fac3aff4abdc6

Observation 152c4b32-31b7-4b31-b8a7-ce1516fce7a9 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

The Hallucination Tax of Reinforcement Finetuning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:44:43.042890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.794896Z digest=sha256:6e1d579e9663d56f5e3a542c4cdb243001f1db5088e01a5882c3dc488b02fefd

Observation 655dcc39-5a30-44d3-ab7d-6c4fd97b7cc5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.954751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.954751Z digest=sha256:6277c68bb77c1d7e4f80371428dcfbb1db11af1d997e58676d1d4f6d6f5dd1d6

Observation 10c551ec-e18d-4b78-987f-de8e7b0ceff8 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.634345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.126414Z digest=sha256:e5a900fbcc42c60e0d79f7b759c00d459c08a04ac19e7b8e0b6e2d820f4ed907

Observation 2ee8666a-b3cf-42e7-b1d0-e4cff21c07ea · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.366454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.324958Z digest=sha256:2b59af390eb522110c282639de8affa7018e9116124d478e99e612c9da041798

Observation e825bd42-c4fd-4479-9e05-a2f9c9a63638 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.545500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.545500Z digest=sha256:416cdc7a6a20ddfa1935bdeba517cb2c9f83baa5d5cfbd3d1e73937c650dbdf3

Observation 4e638485-7ee6-4076-8934-97bbd59cd17a · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.964745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.628322Z digest=sha256:38e5a3f4e9c12c29e3eecea785b60936332251de589e5267a2621b30a1d54a0e

Observation be302502-a318-4dc2-87ee-44644080a929 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Hallucination Tax of Reinforcement Finetuning Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.782472Z digest=sha256:db98b82a50dc94075d5567ba85005c23159c0c3080f24daba35bd7929e2d5bb5

Observation 7c226e53-6050-4b90-b6c2-edf442f7a5d1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Hallucination Tax of Reinforcement Finetuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.894983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.894983Z digest=sha256:ae5908f3e39e86bd1a2755aec3fc802d37c4a10a9d2b766e945b31d9bc781d33

Observation 16abeb0e-4769-4dee-bee0-62ab852816d5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.007922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.007922Z digest=sha256:d8814268b35b358cb926b2ce77d836b82cca5c3b0c76d2131380081b368394fc

Observation b6932d08-7faf-4c4a-9d61-9ad113926933 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

The Hallucination Tax of Reinforcement Finetuning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.170709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.170709Z digest=sha256:4cdcca222db07b123dcddd5d19f5e86054a5d8c22d34454d398a2afecfc151f4

Observation d8527608-d5de-4952-b276-8553a0ab47df · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.272336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.272336Z digest=sha256:b5bde167d3ee132abb72268c70733b64aaae94d50b36210f4c67e4b8d0d7bb5a

Observation ed215d06-63d0-4e1a-96ed-001d416d4174 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.595034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.448645Z digest=sha256:9c9622a006ebc1ce59bbd652859600b168f26d241bf2865b54200e99589e135f

Observation 1e62e25e-fcd5-4aa5-9bbe-597e3bacba13 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

The Hallucination Tax of Reinforcement Finetuning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.610778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.610778Z digest=sha256:5743e463de881888bf0d229da54fc4e86c40b13e791f22c43fe4710d9cee9038

Observation 1e57963e-b949-4603-86b2-f160291ca989 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.284869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.775008Z digest=sha256:4efb6f93fb3208b86b1a5e265cc63eaa06076bc91e5e23cc3b54ae566466f9aa

Observation fdd7360a-d90f-4de7-905a-48d9044a6cfb · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.912984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.912984Z digest=sha256:0d81de7b048a5dea030a387801429f5a765a8d0bcaee9c30865ef6989893b0d9

Observation 1e86e968-b904-4aa8-af69-2222fd6b1be5 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Hallucination Tax of Reinforcement Finetuning AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.037495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.037495Z digest=sha256:4d1cfd2a7aa8061823786f125bd731bf51208ee25900aeed870ce9f2989830da

Observation 9e7e68fb-7fad-46bd-a097-8001bb4f84c5 · outbound

This paper cites Tina: Tiny Reasoning Models via LoRA.

The Hallucination Tax of Reinforcement Finetuning Tina: Tiny Reasoning Models via LoRA

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.167445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.167445Z digest=sha256:ed753f1de0feb29397736210e25327609ba7965ed8afac92cde05dccaf95e89b

Observation ec9d1ffe-a12c-470d-9379-96f40be417bc · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

The Hallucination Tax of Reinforcement Finetuning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.310902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.310902Z digest=sha256:a4247c07512392d79085d635bc917cf6eba563c7ddc09ba0c49cfb809625cb1a

Observation 176734b8-7600-48fd-a725-72d93dc4c722 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

The Hallucination Tax of Reinforcement Finetuning Simple synthetic data reduces sycophancy in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.430765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.430765Z digest=sha256:9d3a45e81b388d650fb8f823a2f6b59397659d5fb01a11b3e87513cc39a4e95d

Observation 9aebf6e4-c24a-47e1-9f0a-9503c9f6beb1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.994636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.515063Z digest=sha256:7672969911c081e8f1f02f65eccbff44f9e34aee78d8bc7ddef55e95f8564b7c

Observation 007b1964-85c9-4886-b1c9-f9788c4ad213 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

The Hallucination Tax of Reinforcement Finetuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.621729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.621729Z digest=sha256:6a0de58fbba90a150b25f69fef421acee398c6ceaa09ff02bae02190f56e44a2

Observation 7be9aab1-c81d-42c3-9abd-b00308dcd272 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.628475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.719317Z digest=sha256:26e1d7b27f7192bb966c9d41e628407c7112a2943541d07005370dc62d654aa3

Observation 16848e90-3df9-41b2-bccf-5254ad9f571f · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

The Hallucination Tax of Reinforcement Finetuning Do Large Language Models Know What They Don't Know?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.850344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.850344Z digest=sha256:f926838a42066354a82807833d1953bf1490c9ac169026cb13332934cdea5a13

Observation b778e018-1636-4f23-891f-dbaaaa36b509 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Hallucination Tax of Reinforcement Finetuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.961914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.961914Z digest=sha256:18c5f32b973cb0923cf03b7efd7f27c878488a5e37adb95e1e217b58d11394fd

Observation 9a8df657-c924-449c-b6ab-22dd12a69bd1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.383837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:44:36.094973Z digest=sha256:bc157ed05445c0883083788c402b32be176a49bd233dc67bb65c08fa7232e217

Observation 9268f459-e4da-4b2a-901c-79f32205026e · outbound

This paper cites How Language Model Hallucinations Can Snowball.

The Hallucination Tax of Reinforcement Finetuning How Language Model Hallucinations Can Snowball

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.213918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.213918Z digest=sha256:87e1125a6046a591154fda08e4f48b7478d6cbc07a5b5f255a7e1a384ff7f952

Observation f32edaa0-bd51-4cb7-81a7-533956c446c1 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

The Hallucination Tax of Reinforcement Finetuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.354899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.354899Z digest=sha256:3cdc52e0e6665c11f881cca4701004672971665c9468ea6e08e769dc54627217

Observation 13000c1c-0526-4a67-bb6d-f74ffe031b72 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

The Hallucination Tax of Reinforcement Finetuning Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.534742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.534742Z digest=sha256:158d2b843fa6fecf8e4b5b18bad664719d947e60f9c940c55fab6bd0629d0389

Pith citing papers

Observation f96d3bbd-0892-49ad-9c12-ce012dde8f0d · inbound

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration cites this paper.

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration The Hallucination Tax of Reinforcement Finetuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:52:50.299350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:52:50.299350Z digest=sha256:f1fbf84c431c77a003f56a8f75760a13feee3adfa58a31596a2a2cd2457cdad2

Observation d1ed60c3-31ad-4d33-8b55-7c3267e55276 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.113610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.113610Z digest=sha256:ff53d26ef65bebdae9f2024d9e7610586b8e1e8c08abb2cc0dfab3220f0eefe6

Observation 1d18fcaa-1364-4c4f-9724-4d6f3d7debac · inbound

MoCo: A One-Stop Shop for Model Collaboration Research cites this paper.

MoCo: A One-Stop Shop for Model Collaboration Research The Hallucination Tax of Reinforcement Finetuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.795988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:17:37.129753Z digest=sha256:31affad6bb738f750308265cc7e7561c7bcb32a7a448b13557b8fa803b055bb9

Observation 7a4542a9-ef81-404b-97c1-25a1b7813634 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.704183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:db866793ebb77a144dacd3f7fc822171cca5232a0cfc4862a4d444d51cd49440

Observation 92e46cd5-8b6e-4d71-83a9-2633117e1cc6 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.560735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:f51d7213e0716653f34d4758f6719049fca8a9078cf7b002ec52f9d9f32cedab

Observation 1686a16f-1bbe-407e-b0bd-10c6fe6c2dc3 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.737194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:b54268d835a0c3b4e72ca6f135eabc8586743d0838f45b27201911824b3a2be6

Observation d9620be2-8cc6-45fd-88ce-ff858571b6e4 · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information The Hallucination Tax of Reinforcement Finetuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.821929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:777273474876003ef1d321254af9b9fdad464c8bf2474b09cd944b8a89a20d8b

Observation 74983e86-3310-4efc-aab7-7357ede6e45f · inbound

Scaling Participation in Modular AI Systems cites this paper.

Scaling Participation in Modular AI Systems The Hallucination Tax of Reinforcement Finetuning

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:16.599301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:49:27.042616Z digest=sha256:693ab4dcc541fa3969e8333d383c17e30227bcd6cc81907cb82af1d8398a432b

Observation a0a77b3d-b3f9-4a9d-8d75-9ebe41e5c81a · inbound

Mechanistic Attention Guidance for Agent Memory Refinement cites this paper.

Mechanistic Attention Guidance for Agent Memory Refinement The Hallucination Tax of Reinforcement Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:30:40.711089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:30:40.711089Z digest=sha256:4c3083b785688ebf09216c89bb5da9377140331d150317e112f1eb06fc98af09