Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.408738Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 11 inbound Pith citation observations for arXiv:2505.00551.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.408738Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:56:46.353798Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T10:27:02.431242Z
100 of 164 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 539d1072-9f97-459b-94c5-7a17770fecb0 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b45ede3-5bee-4a45-8609-33f85f139e16 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737fb8a7-6034-40b5-ba2f-2d8f6dc4fb6c · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a3da34-7ee6-46b7-a0f3-3e436693dc6f · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e654e9a2-1a7a-4f1e-b88a-6aedfa1803be · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Concrete Problems in AI Safety
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc8145a-006b-49c0-9d7f-b6a959fe05b1 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Aops wiki:competition ratings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94de2831-4712-4f84-b463-4ee6d5318713 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Training language models to reason efficiently, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70909e96-10be-484f-b6c4-ec5351f37f12 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models o3-mini vs DeepSeek-R1: Which One is Safer?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af693a3e-e438-4504-8094-8afd781e87a6 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Bespoke-stratos: The unreasonable effectiveness of reasoning distillation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9cdd5d-fd91-406a-aa1d-a25151abec72 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Seed-thinking-v1.5: Advancing superb reasoning models with reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a341fbb0-cda3-4f41-842b-17b079732d41 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fb42a0-2d4d-4977-a683-767cd202b69b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5977f55c-6888-45e0-971d-2879b09060ea · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Extending Context Window of Large Language Models via Positional Interpolation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235a3bdc-f2ac-490e-b2bd-e46f64d65e1f · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bda6d7d-b413-499a-aa0a-b04a3300dce2 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9950968d-2503-4b3d-951f-3c365170bf79 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9445de0-8023-4fc4-9509-4ad79f9fb677 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Gpg: A simple and strong reinforcement learning baseline for model reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68977e9-67bf-4d12-aec5-ae17b98c33c2 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwen2-Audio Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4846c4d0-69cd-442b-88f6-4765013aaede · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9fa991f-8f09-4fb8-87b8-1f122e2e85fe · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Process Reinforcement through Implicit Rewards
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0d52865-b231-429b-b59c-60801388833e · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Graph-Based Multimodal Contrastive Learning for Chart Question Answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd0dc57-facd-48dd-8928-45bc8c6ef586 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Ai as algorithm designer: Teaching llms to improve sorting through trial and error in grpo
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc382601-9c73-468a-91ee-18e603a66191 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Teaching language models to invent or optimize efficient sudoku algorithms through reinforcement learning, 3 2025 b
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3f1e1f-ee3d-482b-b7de-ae5f61ab6369 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Teaching language models to solve sudoku through reinforcement learning, 3 2025 c
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a90ae2-7eda-4475-bbf7-f4748bd22c03 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Security and privacy challenges of large language models: A survey
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 842c4169-4948-4a16-b1ef-fcfba4180718 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepSeek-V3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a398ea-39f4-471e-9915-14d4d3093035 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ed7328-d0ea-4d23-b07b-6259ad79cc68 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6325cf-c9c2-4a59-bb4a-59bedf5c1323 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Rl, reasoning & writing - grpo on base model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c16ef4-b66c-4c1c-92f3-71d812140c20 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e165a717-3b67-4b88-b811-db7d64cf7964 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ad7230-6a31-4da7-be4c-ae366e42e6d4 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reinforcement Learning with a Corrupted Reward Channel
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9291da73-4fc6-4dca-b42d-1a1de474dd68 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06504b25-192d-4195-bbf4-a58ab1f83f72 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reasoning Does Not Necessarily Improve Role-Playing Ability
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3680ba6f-d40b-4cd3-a2c6-0a43099e1473 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d52c3e7-5fb4-44d5-9ecc-e8f3e5dd051e · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8791a3b-9453-4a93-81df-3f1645c89627 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75ba7e0-d583-433a-b684-a9e26f4c86c6 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Llama 3 Herd of Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e9d880-62e1-42dd-9f0c-435166cbec4c · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ca0d26-8934-405b-a674-102b5789d6e1 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6279c5-b26a-45cd-873b-c5e51377c159 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab3e8d5-0853-44ff-8370-0c71dc4c0d84 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Skywork open reaonser series
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b44c6e-d536-4be2-a50f-6b218cab3c47 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0e7856-73f4-4ed7-8bd3-d410e6b5035b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6368f4-43fa-4ac0-900c-1c5cef74b7dd · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264223ce-1e6f-4a4f-a21c-bc173a677681 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Curatedthoughts: Data curation for rl training datasets, 2025 b
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e704c7ad-f7a5-4fa3-aa49-08195c4722c9 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7919bc4a-540c-4556-8b10-0fa91a12a07f · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8a8724-a5cc-4de1-932a-8be56b37bd1d · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2742fabe-428d-4cfc-9ca8-4c6dfdeafd59 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b758c15-4145-4132-b9c1-afb7fa7159b5 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81325677-905f-4287-9ffb-5f781852b2b4 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Ii-thought : A large-scale, high-quality reasoning dataset
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94593af2-4786-4bbf-966a-26c4180612fa · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models OpenAI o1 System Card
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8911b6-556d-4591-be50-69962d61e3cc · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Rlsf: Reinforcement learning via symbolic feedback
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e1ff9e-115f-4005-a912-78a432dafbfd · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15495c8-f115-4aa3-bb20-e30f3f6da3e2 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Challenges and Applications of Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55d0dec-73e3-4d7c-918b-40de3708833d · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd814f1-5d9f-420c-8655-7239f1518449 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Overthink: Slowdown attacks on reasoning llms
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7479e89-db88-42e0-be82-2cc3d1119f42 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5991164e-8c26-405b-b9c5-f9f3c0aa6e6b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Math-verify: Math verification library, 2024
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332fcf98-a569-44f1-a0ea-be3d802202b5 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60b414b-4b68-4c0c-8972-bab8c3b42498 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Introducing superalignment
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e73aa25-0587-4aac-a69a-6e59d6163a0a · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Exaone deep: Reasoning enhanced language models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60de9cde-9c75-4160-9e52-0c8fde14d64b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d2baa89-6623-4f8f-82aa-17b6953abd41 · outbound
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb18a27-11ac-4e4d-9b38-88f173ee4f50 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models TACO: Topics in Algorithmic COde generation dataset
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a587b3-fe87-4db8-9ebb-748a1039dd17 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Large Language Models Can Self-Improve in Long-context Reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f16efdf-be7c-48aa-bbfc-13374a57dd7c · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5e8f2f-9fcf-4fb2-a0e4-46bbdeb4f310 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models LIMR: Less is More for RL Scaling
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c52a279-373d-4c57-aa91-d593b5b4e3b8 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59986456-26a4-49de-8c40-e5c2807722e3 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Small models struggle to learn from strong reasoners
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9424f790-a9de-4953-8fb5-eb7d4cdf526a · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models A survey of multimodel large language models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf94511-93a2-499a-9dc3-7b262485f806 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb84d433-49da-4ddb-8d63-5f4ec874af0b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025 b
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0d6079-6d76-4cfb-9aae-20c604b4ff1e · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Code-r1: Reproducing r1 for code with reliable rewards
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3496bbc7-00de-4828-a669-bd813145e4e5 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Diving into Self-Evolving Training for Multimodal Reasoning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca07818-0809-4e4c-bdfc-d31e0c5db510 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Guardreasoner: Towards reasoning-based llm safeguards
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94eeefe6-ae10-40f3-abdc-2618ca81a267 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models There may not be aha moment in r1-zero-like training — a pilot study
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2a3a0b-18ad-4406-8525-a2b1c50b5294 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb61198-0b13-4456-b6e9-17cd2c61a76c · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498f85de-3130-4e88-8950-cb711b149688 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba918591-2bd3-4234-b5e5-24366a42275f · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f50789-98d9-4c5f-ba51-f3255275429c · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52293daa-1209-4ad3-ad7a-f8bc2b0277cf · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203b37a9-0ece-4088-8dba-772aae9fc564 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1, 2025
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bd7f95-618f-4bea-8cde-bd8ea93d63b9 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models s1: Simple test-time scaling
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8786328-0db2-4dc0-8151-e9e2c3f9fbb5 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a9eb36-f8c6-4273-ab7d-16aa9933bfde · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Spirit-lm: Interleaved spoken and written language model
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e08302b-f578-4164-bca1-4c07b65cc56a · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Gpt-4o system card, August 2024
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a7770e-8221-46bd-97e8-37b142f98e7b · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Introducing openai o3 and o4-mini, 2025
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb88ca0-4883-4445-a83e-6e7374252f7f · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open Thoughts
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7bc8b6-6ca1-4e87-b5bf-d7b80afbf320 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Training language models to follow instructions with human feedback
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef7f58a-5a65-41e8-a0cd-20f4f5801c74 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc5f446-4463-48de-a024-fa8be02cf558 · outbound
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5372249b-0709-40fe-880c-5eec9d852c48 · outbound
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06719fe8-cfc3-4917-b2c2-8699f22cad61 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c7723f-ef14-454f-afd2-358cdb54fec0 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f967b2d8-ea1a-4767-84a9-440d24fe44b7 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwen3: Think deeper, act faster, 2025 a
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9f0f9f-edce-4c41-9927-3e9ab03d65a9 · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwq-32b: Embracing the power of reinforcement learning, 2025 b
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bff36c-2ddb-4015-990d-ffda5b89905e · outbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Manning, and Chelsea Finn
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e17fa188-ac60-485a-8b92-9650eba75225 · inbound
Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd93588-3794-49d2-b589-10b7e5da2a54 · inbound
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b48aa1-d17b-4230-8db2-f10c5ecc337c · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426d76a5-1199-4889-a3fb-0e9c2a936848 · inbound
WebDancer: Towards Autonomous Information Seeking Agency 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · inbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0404d396-0421-4503-b026-64648534fb39 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 238
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959f1a18-0010-458e-a947-05e298a39acd · inbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3712cee-75c6-4c3d-ab03-29e56857fac0 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4c8a8ab-0d59-4890-a6fe-6f8b65bba7fe · inbound
Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7cfe54ec-dac8-4315-9e43-a6338af90cf7 · inbound
Trust Region On-Policy Distillation 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc303744-dedc-4b76-b87f-a00cd8383882 · inbound
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.