Pith. sign in

Paper Citation Record · LEDGER

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2505.17988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17988 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:42.861182Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:02:41.773793Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T19:03:11.044453Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac8f7d77-a107-4077-89fc-af3e6c7306dc · outbound

This paper cites On exact computation with an infinitely wide neural net.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning On exact computation with an infinitely wide neural net

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.977890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:38.977890Z digest=sha256:b8bc7a7a426e49f96956136a39df7a86aa86eda91a55b594a0bd1fdde9ab5357

Observation 297d458b-dbcc-4f72-ab49-ff4512919862 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning A general theoretical paradigm to understand learning from human preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.971192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:39.029884Z digest=sha256:68d82c44fac3f0e4c5f93387c01c1ce680d6911ae2301cb5c51acc940984eae9

Observation 31fed710-f59b-407c-a173-eda5d9b6604a · outbound

This paper cites Quantifying memorization across neural language models.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Quantifying memorization across neural language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.730875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:39.164577Z digest=sha256:4998750f5008d7a2eb66bd1e6bbc2c570b9604c2ebd943c4e6edb88727e558d2

Observation cf4a11ce-1b22-4ae5-9bd5-12572d62f4d8 · outbound

This paper cites The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.281110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:39.281110Z digest=sha256:ab1541771dee4373ac3b8e5be0222ea265aa9e111f977155ebe422d7ee6856ed

Observation 23683e5a-5c50-443f-9523-3803e6c70954 · outbound

This paper cites Generative AI for Math : Abel , 2023.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Generative AI for Math : Abel , 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.581085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:39.412297Z digest=sha256:00976f41e20b6fb6b5e807667bb5a68f0c676aca885a03631a51e02675c60666

Observation 52973908-ca4a-4c27-81c0-1b03f73d6470 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.559892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:39.559892Z digest=sha256:57b251298b2ce960ad7c0df137d7b34c422b1969578f59d8b00a26db056600c3

Observation 7956933a-7f09-42e8-9f54-64b27e12fbf0 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning KTO: Model Alignment as Prospect Theoretic Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.679644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:39.679644Z digest=sha256:edcc2bbbfaa336d35fd1c5a5a8448a6e6b7a61ec899781734bb20596598fe206

Observation 2ad3645d-df46-470d-85af-b60d0ad396a4 · outbound

This paper cites Open R1 : A fully open reproduction of DeepSeek - R1 , January 2025.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Open R1 : A fully open reproduction of DeepSeek - R1 , January 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.410444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:39.817035Z digest=sha256:8c403b6c670ba33088befd1e5c9e3f4bfa96123c65e75985ef0d25d9bc3cf0cf

Observation 790e28f7-f863-4365-9a47-094c6fe7d020 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.932382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:39.932382Z digest=sha256:a3bda1de452f621680489d569f12bf3da889a1b6c3fd35ead4e3d4f4cf5e0254

Observation 2ec1ae1b-3119-44a3-a895-9b83d959e8e2 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.069115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.069115Z digest=sha256:e987a403ab744bfd2a2e60ffd1a40431209684889ac04ee55156eb97d0f425b0

Observation fc27d130-dbf0-48e8-8537-4ea25f18070a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.220148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.220148Z digest=sha256:01ce12d378c843349e21d75daa24a28f8192844bffe31ce097e209084d666c68

Observation e2e60055-69b5-49b5-91f2-8f32525c57dd · outbound

This paper cites SoK: Memorization in General-Purpose Large Language Models.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SoK: Memorization in General-Purpose Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.334702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.334702Z digest=sha256:2ecf637937b4cb731ab693925f76eda832ab46e4214b02ee3675d86cd15d369c

Observation ff073668-a698-4c69-968f-456875c472b6 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.442788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.442788Z digest=sha256:26adb2c3a3065b370097092f9af7c490c5d0c4c9b2c1fa168040982dda2f3e08

Observation 9805a4cc-3efa-4317-96f9-3e0f371be348 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.517731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.517731Z digest=sha256:d8b7b88383f5cef68d4edb9870a00bef29f5f5dd72d7fa0119d8abaa4bba750c

Observation 1544743b-b843-42da-ba7f-096799465de8 · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Neural tangent kernel: Convergence and generalization in neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.283754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:40.654506Z digest=sha256:b89d63d1bf08cd29198242aa627e2edcf4a0bd8dc498480d1c3bcd5a3a8aa906

Observation d205c502-cb5c-4978-aa94-648277d293ff · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Towards Efficient Exact Optimization of Language Model Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.748666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.748666Z digest=sha256:a6d1aaa5503f8043b21d75147f5089daeed2e77679825a97df62e13c57e028db

Observation 7e5bb8d4-9530-415e-92f3-10efb6609422 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.825875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.825875Z digest=sha256:9695a1ba5c7a045eb7f1632030a1cf80aa9c1bbec78973b42f39c0cebfcd901f

Observation 73bd9f18-90c4-457d-aaab-815e50d0a226 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.896626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:40.896626Z digest=sha256:155a2b1d1408f878b1825a415a83f27af42729d45a8c81b16526418ef6557c59

Observation dc35b232-f6e9-40ef-8eb2-488e2bb593b7 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Efficient memory management for large language model serving with pagedattention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.003001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.003001Z digest=sha256:23743174e3e7b826ed24336b8445b9708032cc698b99320bffd6a12470ed66ef

Observation e183c9cf-5934-4b17-8d45-09e7a17989e7 · outbound

This paper cites NuminaMath , 2024.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning NuminaMath , 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:44.100819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:41.097977Z digest=sha256:307f7b0561cbe37c4f2cbcc9be9e848521598a8590eb630d2f8d9d1fdc6e1d7e

Observation 6f9b73f4-d090-4ab2-93a1-d8eea78a0da2 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.222191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.222191Z digest=sha256:64af35359d4654df160d2772f2ea6481374beb6ebac2ca6a9f3801cd06fd6d7a

Observation 307b8061-d6b0-40f8-adf4-7db7c44fbc69 · outbound

This paper cites s1: Simple test-time scaling.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning s1: Simple test-time scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.319715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.319715Z digest=sha256:0b5913d46f1b83115ced240f6c5436eeb846696ce7b92c92ac3edf8c6934e9cc

Observation f6015ee9-ea2d-4add-bf2a-9479d800535f · outbound

This paper cites OpenAI o1 System Card.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.399866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.399866Z digest=sha256:a620441f83699f731d3fe5464ee87f1908be1764151aa5c65b955746df190618

Observation 0878badd-2b52-442b-baf4-111d63fbec82 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.482480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.482480Z digest=sha256:b1081b4ee5cfa2b0083e3be128a1fd24200f41ef6937b65e5e92b163e7c8064c

Observation e4859605-fb3c-40c5-ab4b-5fd7867bf32f · outbound

This paper cites QwQ - 32B : Embracing the Power of Reinforcement Learning , March 2025.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning QwQ - 32B : Embracing the Power of Reinforcement Learning , March 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.993712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:41.561535Z digest=sha256:f768c18bd13e5676889cd84a0083eddde623c68afb6a9ac184a297b1bc739e00

Observation 158a710d-d6a8-4309-8aa9-5b8ac4361a2e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.655714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.655714Z digest=sha256:865b1a9b2d785808486b72b7445ed31735d9e5951a7ed733150dba44547ed0ee

Observation cd2a7a68-530a-49df-9b75-a23b2a061db7 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Learning Dynamics of LLM Finetuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.745767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.745767Z digest=sha256:7957434c2b21362cf432c8b55bf757dca19d164f0d214a69a51f2e70d07f3426

Observation 4aeecf9d-8acd-4dc7-afc4-42c11985a68d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.846490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.846490Z digest=sha256:0e721dde3ec5289dc0402d4dc28b6a755dfe22accfb366f5897da3cfc1a3cc2b

Observation 42448170-c7f8-432e-835a-fada56558a22 · outbound

This paper cites Policy Gradient Methods for Reinforcement Learning with Function Approximation.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Policy Gradient Methods for Reinforcement Learning with Function Approximation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.846736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:41.946416Z digest=sha256:51fa928128602cb745223df1a4d86197b64f2890350372c322ef0d4664c6243e

Observation 3e902f7e-4ce1-41cb-b93e-cbbfac471a4d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.669174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T14:41:42.046397Z digest=sha256:f0cee4c9fd1b085c635d188011c8250ec01c2215d7330678e77bd30524553735

Observation 18575e88-56ee-489f-9617-fde83ef0a83a · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning On Memorization of Large Language Models in Logical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.174397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.174397Z digest=sha256:fc9272e1d2382a1788eb0310583ce0f74e20608679d450a3a4b27899000fa097

Observation ed7b0955-9280-4627-9ee8-773c114f2c72 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.276120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.276120Z digest=sha256:86263a01a68197d50881b4df9543ab11893bca32b44c84cf829749ad750ed6b2

Observation 4d295d10-5b4a-4dd3-bec7-ac4081772af4 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Qwen2.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.361609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.361609Z digest=sha256:c01921dd2cdab7036b29fedbb25adc4ac8a9a95ad964e25aa5fad3a6c219e080

Observation e130b08a-299f-443b-a932-1891b1f24973 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.429652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.429652Z digest=sha256:4bc4fdb7be6af841ad0e63aa70cd42552ccf9e0da6f05c08618b9efddeea1810

Observation 0e849a46-9f46-4624-a898-8e79d93d41b8 · outbound

This paper cites LIMO: Less is More for Reasoning.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMO: Less is More for Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.502205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.502205Z digest=sha256:fc1878aeb09f91c255633951b94f612eb12f1e8af5d88c9eaafafe778dc56557

Observation 876f7086-f4b9-43fe-933f-97658e4fe52c · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.610566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.610566Z digest=sha256:869e1af63726715cd5904ed253bbf6466f74ca22352530d5f108e15b6ddb34e7

Observation 63cbc4b7-953c-499d-986d-d3fc73d92757 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.685105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.685105Z digest=sha256:3f10f3c256c1efd44922659232a2741355150f511abe302a5ae322369d70ba19

Observation 8016a355-3219-458e-862c-f60374c936ef · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.792713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.792713Z digest=sha256:05c68382fa17f3ce05d776bbd7d66df872539f34779d53fc94394617e549704b

Observation b1265abc-8c31-453d-8e78-17fdf1016739 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.861182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.861182Z digest=sha256:d98ba4741297cac20d2d41f6677a80b1b414c68063bbe78bfc73f26719750d70

Pith citing papers

Observation d1d75131-f636-457c-aea4-1ff0062c884a · inbound

Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance cites this paper.

Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:03:11.133072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T19:02:41.773793Z digest=sha256:60c928ccbb21a17a667979ea052195b5ffc1644690f79b338bfc98f840c39fbd