Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:42.861182Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2505.17988.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:42.861182Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:02:41.773793Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T19:03:11.044453Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ac8f7d77-a107-4077-89fc-af3e6c7306dc · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning On exact computation with an infinitely wide neural net
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297d458b-dbcc-4f72-ab49-ff4512919862 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning A general theoretical paradigm to understand learning from human preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31fed710-f59b-407c-a173-eda5d9b6604a · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Quantifying memorization across neural language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf4a11ce-1b22-4ae5-9bd5-12572d62f4d8 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23683e5a-5c50-443f-9523-3803e6c70954 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Generative AI for Math : Abel , 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 52973908-ca4a-4c27-81c0-1b03f73d6470 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7956933a-7f09-42e8-9f54-64b27e12fbf0 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning KTO: Model Alignment as Prospect Theoretic Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad3645d-df46-470d-85af-b60d0ad396a4 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Open R1 : A fully open reproduction of DeepSeek - R1 , January 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 790e28f7-f863-4365-9a47-094c6fe7d020 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec1ae1b-3119-44a3-a895-9b83d959e8e2 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc27d130-dbf0-48e8-8537-4ea25f18070a · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e60055-69b5-49b5-91f2-8f32525c57dd · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SoK: Memorization in General-Purpose Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff073668-a698-4c69-968f-456875c472b6 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9805a4cc-3efa-4317-96f9-3e0f371be348 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1544743b-b843-42da-ba7f-096799465de8 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Neural tangent kernel: Convergence and generalization in neural networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d205c502-cb5c-4978-aa94-648277d293ff · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Towards Efficient Exact Optimization of Language Model Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5bb8d4-9530-415e-92f3-10efb6609422 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bd9f18-90c4-457d-aaab-815e50d0a226 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc35b232-f6e9-40ef-8eb2-488e2bb593b7 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Efficient memory management for large language model serving with pagedattention
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e183c9cf-5934-4b17-8d45-09e7a17989e7 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning NuminaMath , 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6f9b73f4-d090-4ab2-93a1-d8eea78a0da2 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307b8061-d6b0-40f8-adf4-7db7c44fbc69 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning s1: Simple test-time scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6015ee9-ea2d-4add-bf2a-9479d800535f · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning OpenAI o1 System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0878badd-2b52-442b-baf4-111d63fbec82 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Competitive Programming with Large Reasoning Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4859605-fb3c-40c5-ab4b-5fd7867bf32f · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning QwQ - 32B : Embracing the Power of Reinforcement Learning , March 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 158a710d-d6a8-4309-8aa9-5b8ac4361a2e · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2a7a68-530a-49df-9b75-a23b2a061db7 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Learning Dynamics of LLM Finetuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aeecf9d-8acd-4dc7-afc4-42c11985a68d · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42448170-c7f8-432e-835a-fada56558a22 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Policy Gradient Methods for Reinforcement Learning with Function Approximation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3e902f7e-4ce1-41cb-b93e-cbbfac471a4d · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18575e88-56ee-489f-9617-fde83ef0a83a · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning On Memorization of Large Language Models in Logical Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed7b0955-9280-4627-9ee8-773c114f2c72 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d295d10-5b4a-4dd3-bec7-ac4081772af4 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e130b08a-299f-443b-a932-1891b1f24973 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e849a46-9f46-4624-a898-8e79d93d41b8 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMO: Less is More for Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876f7086-f4b9-43fe-933f-97658e4fe52c · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cbc4b7-953c-499d-986d-d3fc73d92757 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8016a355-3219-458e-862c-f60374c936ef · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1265abc-8c31-453d-8e78-17fdf1016739 · outbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d75131-f636-457c-aea4-1ff0062c884a · inbound
Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.