Pith. sign in

Paper Citation Record · LEDGER

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 1 inbound Pith citation observation for arXiv:2602.05897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05897 v2

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:09:49.243797Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:52:42.813309Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91a99dbd-85d5-472e-b9bf-4716629a7e93 · outbound

This paper cites an unresolved cited work.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.916643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.916643Z digest=sha256:aff98d5d20d3c6ede5bbde6252bcc9355ed05cf7e77d23882a17d8965b15f6c9

Observation feb71d4c-4332-4e4e-94ea-3b92ad96a4f4 · outbound

This paper cites an unresolved cited work.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.990491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.990491Z digest=sha256:5548007d88e6d576d530d504ce50918a4969af94268541d186695272a104a2ad

Observation 287c4a17-b522-4b7a-8127-d4e9a8460bf3 · outbound

This paper cites an unresolved cited work.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:49.135992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:49.135992Z digest=sha256:20b8f3e4944a6df038df67663d66e2bb63e3ba79d715436572632704f3899c84

Observation 3d98b1e1-bc17-4a8e-acae-a8fc774ebfab · outbound

This paper cites new_question_1.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models new_question_1

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:49.243797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:49.243797Z digest=sha256:300ed670b18c580a0e991d1fa3646506ed64d73879dac131d5d54eaaac069735

Observation d5b6b225-781e-4335-b577-abe861a02ca9 · outbound

This paper cites The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.472543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.472543Z digest=sha256:5d3c74c513cc2446a65a8245f3afd13e09893af7749cec8b7acaeb49a71dff4d

Observation afd2a6cd-2ad7-4aac-ac82-e01d76411e67 · outbound

This paper cites an unresolved cited work.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.763037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.763037Z digest=sha256:dd12669f1b686665008c30a7d06b286fe86c54f3bdcf69679d313f84dd29ff8c

Observation 8924ffb2-89f7-4ce6-994b-ae8d3aea172b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.842765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.842765Z digest=sha256:a84d371ff701cdb205b1bd3da9f03df09cd2ddfc7db3c378e3b63703b50047d6

Observation e2a3ce71-11dd-4a11-9d7d-e8459ca4b964 · outbound

This paper cites Why Language Models Hallucinate.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Why Language Models Hallucinate

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.609050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.609050Z digest=sha256:c4bf38aa19c4fab459a1a85bffee72a3e8e7d6583f05a3dc87bc04cfe4fef834

Observation 62fab9f8-67f8-495f-ae62-a0028f908a72 · outbound

This paper cites Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.708609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.708609Z digest=sha256:1695170a80469c56a243f5a2c16b34e827fc3330cb730be0e84cf03b2e0e1b97

Observation 1832476d-24f2-466d-8b07-a60996cc1d98 · outbound

This paper cites Haque, M.

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Haque, M

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T04:09:48.340314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:09:48.340314Z digest=sha256:e6deb601ba3a528f9451b693054b7f62d3dbd49a00a67e8a8385f17563dd3e84

Pith citing papers

Observation 2befc4e1-7fe5-4a7a-8ba7-22bc161c998d · inbound

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series cites this paper.

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T14:52:42.813309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:52:42.813309Z digest=sha256:fa45d0ba527bba4c6fa908ed4767fe2f6c94d76d0140914fdddae64ea65cd2cf