Pith. sign in

Paper Citation Record · LEDGER

Reinforcement learning fine-tuning of language model for instruction following and math reasoning

As of 23 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2506.21560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21560 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:30.520795Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:40.924883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:39:43.095193Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b22cb6-de3c-4a24-b323-b1a84269d16e · outbound

This paper cites arXiv preprint arXiv:2412.15287 (2024).

Reinforcement learning fine-tuning of language model for instruction following and math reasoning arXiv preprint arXiv:2412.15287 (2024)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.512978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.512978Z digest=sha256:bf97ab1ebbd4df4e298c0cd1c836f55f319e650da1c1b9c4d57060123df1575c

Observation b4301b55-41a7-4964-a3b9-40574ff41443 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Stream of Search (SoS): Learning to Search in Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.746418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.746418Z digest=sha256:a159428955f5b15d95e54279c79963d7fdf54778861a929e31df5c933a64b3bc

Observation 3b9a162a-60ac-446a-8ddc-30c8f0d508da · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.805270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.805270Z digest=sha256:3841999759e0e327184562ddccb8c72b02eeac3a545a792a961bfe33ca7dc1c9

Observation 46b454a4-95ff-474f-bee4-d78c48d97283 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Qwen2.5-Coder Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.960575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.960575Z digest=sha256:f56c1af4fb808b5e31073f75679b3f402e32eca0b444aab967703d0430bdc143

Observation 59233e2a-8955-4a6b-a21e-f05e28ca516c · outbound

This paper cites GPT-4o System Card.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.042682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.042682Z digest=sha256:664c5d7b0b71add5f181e56037b81f0c84167602aa0a7ced4150e21c2e212706

Observation 9e941f99-4fb5-4202-af35-0900f6a5caa0 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.178554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.178554Z digest=sha256:adca8a8b04536819b186bea78e3090451a309c1327a6335344a955089268f5be

Observation 259ad9ec-5d56-40a1-b865-db09449d6c89 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.234365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.234365Z digest=sha256:0fdfad23907bdfa0573ff02bb4992a983bd5f0f541622da94b1ecfcee7a18e91

Observation 1ac134b5-0f06-4d92-bcf6-248a6fe79e99 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.328965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.328965Z digest=sha256:10edd39dc846e1608dcec8c8fdc8e1a3fcb643e0a6e5bb047f182b1aef7b22d2

Observation a2596d1b-9eaa-4455-a10c-3bd2053ee55f · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2023), 68539–68551.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Advances in Neural Information Processing Systems 36 (2023), 68539–68551

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:30.992778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:37:30.416102Z digest=sha256:0c603253c1e23cc10b2b608797062f7ca4175836640fc8d43f736d1ae70135c0

Observation 640b320e-fa86-4451-90c8-29f6488264bd · outbound

This paper cites ALJP: An Arabic Legal Judgment Prediction in Personal Status Cases Using Machine Learning Models.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning ALJP: An Arabic Legal Judgment Prediction in Personal Status Cases Using Machine Learning Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:37:30.687681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:37:30.462054Z digest=sha256:ef6d2c6ba5b006f6f4d3637ff820fc50163eb8a0e1959ce1d03c484301da5537

Observation 366625d7-fecf-4096-be1e-d31814a019e7 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.520795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.520795Z digest=sha256:b2dd2e2bef252c6a1d1298cc089fb0cf7d586f45471b792dad1c12dd4883cb68

Observation 02b4070b-e6d0-4dc9-9a42-fcc2995941ff · outbound

This paper cites In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers).

Reinforcement learning fine-tuning of language model for instruction following and math reasoning In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:31.172059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:37:29.669550Z digest=sha256:07be1dfe6af7de0dbcdea01e628ad8b2d513c5621a58f751e142a593b26596d7

Observation 5a2ca701-2b9b-4b81-b3e4-abec734aa4bc · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.907111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.907111Z digest=sha256:9ba9394587c0453d7b9367860707782cb43bbd609d39772a882048bc089af748

Observation 787ecde3-dc8e-4df0-b6c3-4c3626bb138b · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.608840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.608840Z digest=sha256:21ece01b4f29ac079a7c247c10b66279ad2f7f504ed45a77cbec8d02a33e2008

Observation 5a339a03-76b0-496e-9880-10cae60cd3d1 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.438985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.438985Z digest=sha256:0cb6b59c1f54f39c750435ad64e0ea9bdd86e003d58a6b756b3fe5270b9b10b1

Observation 072ae991-3bc8-4471-9ed3-6a86b39a3cf6 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.121543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.121543Z digest=sha256:e3046f4037fc53d8ae9e954b47c5584b0a7046c4a562e1fdc32e46b715453430

Pith citing papers

Observation 8059ffd7-6d45-4600-a3c4-3192ececd1f8 · inbound

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models cites this paper.

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models Reinforcement learning fine-tuning of language model for instruction following and math reasoning

Reference 142

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:39:43.100701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T00:39:40.924883Z digest=sha256:9061b80a68c8814d6722065b05832e3e997354647299c0e792afb8bccd578769