Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 3 inbound Pith citation observations for arXiv:2508.04848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04848 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:49:05.708180Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:23:52.536677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:01:22.618150Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cd5b0c6-2b61-4a8b-802e-9cd4f7868183 · outbound

This paper cites Evaluating LLMs and Prompting Strategies for Automated Hardware Diagnosis from Textual User-Reports.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Evaluating LLMs and Prompting Strategies for Automated Hardware Diagnosis from Textual User-Reports

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:07.511703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T23:49:02.285130Z digest=sha256:407c727f16ebfbf4697ab75252cdbbfdd9c58ec4bcbdb64addab9684238b69bd

Observation 26e2ce09-01a4-4690-ac40-280bd2d77a25 · outbound

This paper cites an unresolved cited work.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:49:07.820130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T23:49:02.550966Z digest=sha256:e1807b67eec1681c0aa055aac60ba27f1a7cd6b394adaff980ab6747b0d9cebc

Observation 3a6443b6-1c35-4fb7-a39b-93fcaba8ab0b · outbound

This paper cites arXiv preprint arXiv:2503.12434.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2503.12434

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.715209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.715209Z digest=sha256:24c73d318f30c85024558a0fdba8799617e28688aea4e718000e1147cdbb360a

Observation cc53431a-943a-4af7-afb6-ffb268577a7a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.929694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.929694Z digest=sha256:6f38c778a820309cf83ac19c4c452076ec7fef05949ff77cd9b136cd62c188fb

Observation eed7fb36-b81f-4e42-afde-45576bac2ccc · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.216906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.216906Z digest=sha256:2404476c98fd5a548e3e51b11349230f3d2fdd392793a76093d69fbc91e9c86b

Observation c089ee1b-f9ea-4421-8d1d-ce90170f4779 · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.474747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.474747Z digest=sha256:ac3b2b372764ab52254385f331349d9e9a777655fdc0d8c2e7451e2adea9e38c

Observation 5069f2cf-b853-4e39-a927-395baf2b1c31 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:04.041085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:04.041085Z digest=sha256:a4c6defdc470fcecd63ee55c2fd363f27263b5d576a8cb3cc1510d1cb49b100b

Observation 2a099c1a-ee77-4545-bebb-118df17d7438 · outbound

This paper cites Using Causality for Enhanced Prediction of Web Traffic Time Series.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Using Causality for Enhanced Prediction of Web Traffic Time Series

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:06.723730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T23:49:04.374494Z digest=sha256:d3024909f2e29bf6b61c9507c221dd3ff986e9e7ae72f1f6e3388526c861d2c4

Observation 9453a80a-4f89-45ac-adb3-85b33cfdf8d6 · outbound

This paper cites What is the Alignment Objective of GRPO?.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning What is the Alignment Objective of GRPO?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:04.813342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:04.813342Z digest=sha256:ea86ef4d9e55623e6fe96ced250473b8f2c3c00faee35f67452c384c7be3f506

Observation da8eb2be-757c-408d-9572-55982b3900b1 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.025670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.025670Z digest=sha256:d79f792733a3c9626f466ae152ad8e557101f2115d28d7e4a7e5d25ed241b1df

Observation 4dd15b6d-ed00-430b-b9bf-739dc07bb09e · outbound

This paper cites An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.199743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.199743Z digest=sha256:74b04947b77585b46099a5ce206d69a3ea6d5618a6578c7a3c9ed5a5436276c2

Observation b51dbaea-cabf-458e-bd33-21a885642e6d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.395955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.395955Z digest=sha256:6f032691047b50af630ca5595ca09be01499533a9b50b4f7a1484ed0cac75c40

Observation 8197227e-d83e-46e8-b5bb-5485649366a8 · outbound

This paper cites arXiv preprint arXiv:2505.17508.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning arXiv preprint arXiv:2505.17508

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.541952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.541952Z digest=sha256:35d57618861171ce3a11a078de81d7972181a9750ff2d37bcd8cf091b29aa29c

Observation 8a413b4d-c7e9-4a91-b607-8b6cf0a66171 · outbound

This paper cites Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:05.708180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:05.708180Z digest=sha256:3e0722b21e61a8d2140589ef298e9aee14b694b6553ba6f48fc011119b7f7b4c

Observation c5a07c9f-a069-4938-9dbe-c3d61df449e2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.820612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.820612Z digest=sha256:638aaf558e1ec753686fdb3cc4e6a555bad5e4b84667532234577f3a71538f74

Observation e9fbee50-7381-4d9e-9880-7e86a8620486 · outbound

This paper cites A Generic Method for Fine-grained Category Discovery in Natural Language Texts.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning A Generic Method for Fine-grained Category Discovery in Natural Language Texts

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:49:07.013350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T23:49:04.213823Z digest=sha256:bae3d6982d51ce543d81c44710ec337b9d7479f533fcda368174f198f2524def

Observation 10b799cd-fa19-409b-8e5e-13ff2688b4a9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.062950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.062950Z digest=sha256:0bff75231e339498b0b83289956573f4aea8acc45460d6c2981bab51a9da37f0

Observation 9fdfde4a-26c7-4548-9b83-561e8375ee4e · outbound

This paper cites Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.658162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.658162Z digest=sha256:f57f5d156ef1d7e57aadb2afbeba1139d6cd4a925b19b6ef72bf055d26786b6e

Observation f8585f12-a48c-4d6d-b420-834437982951 · outbound

This paper cites Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:49:06.307243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T23:49:04.536418Z digest=sha256:44ebc85cf556ab1f0cab2c0559bf9a92b5a6863a8da1527f0505a32edb001bd2

Observation 22c06452-16ca-4c32-b0ab-d55ec42d5471 · outbound

This paper cites Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.409134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.409134Z digest=sha256:231c0a8e18f86b42d8ac6c86f5962ad20e92a1499f839f5bc0efff6ba3ea02ca

Observation 6c9329a9-57f5-468f-aeec-0c66a4b91d01 · outbound

This paper cites Qwen2.5-VL Technical Report.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:02.173108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:02.173108Z digest=sha256:96e4aff59c0d3118e41cfb5b35494581f9cff04f2d3e5d953ca319e71d0b0cf0

Pith citing papers

Observation f869d388-280c-4d69-8949-c4af9c297b9e · inbound

Using Causality for Enhanced Prediction of Web Traffic Time Series cites this paper.

Using Causality for Enhanced Prediction of Web Traffic Time Series Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:52.536677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:52.536677Z digest=sha256:9d38f465a0ce06ad15e0fcef039bf4326787916588946e41d78cdbcbf9b1e134

Observation b426df3d-129f-4e23-a706-9cb26d330ce0 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.621931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:edf4e0da2d7960a23ae9ff8b338cc6aae28140a8a1472be75a5a8db4586ef62b

Observation b60805a6-808a-4981-b738-8bea31069308 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.662264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:8e75adf15c256d7b13be198d871eca15ac00d8ee3557541a5b5d203568b13e9e