Pith. sign in

Paper Citation Record · LEDGER

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

As of 20 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 11 inbound Pith citation observations for arXiv:2505.00551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00551 v3

Coverage vector

measured 100 of 164 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.408738Z

measured 111 of 111 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:56:46.353798Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:27:02.431242Z

Reference resolution

100 of 164 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 539d1072-9f97-459b-94c5-7a17770fecb0 · outbound

This paper cites write newline.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.653066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.653066Z digest=sha256:31c1f1c80eefc0758c5bc1a9fdbf3826b8318b05410a14d29f3c40f7f8af1336

Observation 7b45ede3-5bee-4a45-8609-33f85f139e16 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.660637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.660637Z digest=sha256:341810136ec212e1a4c75194da7024d1f5d02d766f696ba951d329e6781e2fcb

Observation 737fb8a7-6034-40b5-ba2f-2d8f6dc4fb6c · outbound

This paper cites Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.665427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.665427Z digest=sha256:41c1854909fb06f52bba8acaaa5c03ce2961d409e3a5a408e47a9f6b1ac2f178

Observation 33a3da34-7ee6-46b7-a0f3-3e436693dc6f · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.670016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.670016Z digest=sha256:69a229c6e10dbb743b397f63b3651ef552f1159eba554cf21a1da4bb540baeab

Observation e654e9a2-1a7a-4f1e-b88a-6aedfa1803be · outbound

This paper cites Concrete Problems in AI Safety.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Concrete Problems in AI Safety

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.675591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.675591Z digest=sha256:2a18438b3cdd232c1632df0b86aab801f963bece9bf422394e2c844b86f43ca1

Observation fcc8145a-006b-49c0-9d7f-b6a959fe05b1 · outbound

This paper cites Aops wiki:competition ratings.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Aops wiki:competition ratings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.681164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.681164Z digest=sha256:174b9217b3efe89e03cb16f8a25679b9c826aa4a4e9d3f4e1fd246777575406f

Observation 94de2831-4712-4f84-b463-4ee6d5318713 · outbound

This paper cites Training language models to reason efficiently, 2025.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Training language models to reason efficiently, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.685923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.685923Z digest=sha256:128c7fd31d5a03ea8f64c39f57f16a746944214dee8df980c48d8616f62dc174

Observation 70909e96-10be-484f-b6c4-ec5351f37f12 · outbound

This paper cites o3-mini vs DeepSeek-R1: Which One is Safer?.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models o3-mini vs DeepSeek-R1: Which One is Safer?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.692600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.692600Z digest=sha256:87ef50faea949cc2a6d432ae7fe14a71597107e2afc92d5a8c35f859c8c03cf3

Observation af693a3e-e438-4504-8094-8afd781e87a6 · outbound

This paper cites Bespoke-stratos: The unreasonable effectiveness of reasoning distillation.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Bespoke-stratos: The unreasonable effectiveness of reasoning distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.698742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.698742Z digest=sha256:7b802ae27be32f5490be4fa67436c3f15571d4fe5ba0e64f392e18358e5f17d9

Observation aa9cdd5d-fd91-406a-aa1d-a25151abec72 · outbound

This paper cites Seed-thinking-v1.5: Advancing superb reasoning models with reinforcement learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Seed-thinking-v1.5: Advancing superb reasoning models with reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.705095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.705095Z digest=sha256:705a67e33461c0a873e0f8644b20bff0a316ca990843c15ac41da3b9b7d25210

Observation a341fbb0-cda3-4f41-842b-17b079732d41 · outbound

This paper cites xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.709923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.709923Z digest=sha256:64326c17f22a00b3362a01a47779c8578f03b05d3b51e1f23c6175105eaa6ef2

Observation e8fb42a0-2d4d-4977-a683-767cd202b69b · outbound

This paper cites FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.716688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.716688Z digest=sha256:78f38b40badcd21e0c91f62c4a68dd62e1fb4b3be13f3eaf55e0bb7691375d07

Observation 5977f55c-6888-45e0-971d-2879b09060ea · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Extending Context Window of Large Language Models via Positional Interpolation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.726938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.726938Z digest=sha256:356ad0a15b3e5e972b41dcb7c1d638f3d29720cbd9c5c8963f233849f0602e41

Observation 235a3bdc-f2ac-490e-b2bd-e46f64d65e1f · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.734066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.734066Z digest=sha256:33eefb90e1f79c6d1a917e13d1b176c3f181a9995296a4e0d4dbd76dfdbe4338

Observation 9bda6d7d-b413-499a-aa0a-b04a3300dce2 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.740471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.740471Z digest=sha256:a7a111bec7e198dde2f340d062baa0385653e79a8ddb338de5b35eba60784f48

Observation 9950968d-2503-4b3d-951f-3c365170bf79 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.745341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.745341Z digest=sha256:f9810a905c41bb7c5a0502cb0aa673524dfd91cccd75e3bd324b75cae0b4b1a0

Observation f9445de0-8023-4fc4-9509-4ad79f9fb677 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Gpg: A simple and strong reinforcement learning baseline for model reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.753574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.753574Z digest=sha256:85e02e50dac237f1e611bb30d901c3e517958eb9a61ec653b198a90c288cfb67

Observation f68977e9-67bf-4d12-aec5-ae17b98c33c2 · outbound

This paper cites Qwen2-Audio Technical Report.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwen2-Audio Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.758427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.758427Z digest=sha256:3b04d860571364c87decbaca9f32d81c0b6060ceff2b4a8283377b39c60e5dc9

Observation 4846c4d0-69cd-442b-88f6-4765013aaede · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.765953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.765953Z digest=sha256:5d7f520703f3238e3c6ad5c36407b61ce6b5d10a2cb557eca8886a6a07135836

Observation d9fa991f-8f09-4fb8-87b8-1f122e2e85fe · outbound

This paper cites Process Reinforcement through Implicit Rewards.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Process Reinforcement through Implicit Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.770838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.770838Z digest=sha256:e25dabb10cc2506d6af9d21ff4228f7fc69d14d700222d4e448b9ddb1e4042ce

Observation c0d52865-b231-429b-b59c-60801388833e · outbound

This paper cites Graph-Based Multimodal Contrastive Learning for Chart Question Answering.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.953062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.953062Z digest=sha256:a04ec061782df134a1d5b6f063de669750fbf11c1d3009b2b3359fa84008213d

Observation 9bd0dc57-facd-48dd-8928-45bc8c6ef586 · outbound

This paper cites Ai as algorithm designer: Teaching llms to improve sorting through trial and error in grpo.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Ai as algorithm designer: Teaching llms to improve sorting through trial and error in grpo

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.958215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.958215Z digest=sha256:9dca16387cc514e64be7b9a219dc08b169941c4f551a0d41b003a48018486913

Observation fc382601-9c73-468a-91ee-18e603a66191 · outbound

This paper cites Teaching language models to invent or optimize efficient sudoku algorithms through reinforcement learning, 3 2025 b.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Teaching language models to invent or optimize efficient sudoku algorithms through reinforcement learning, 3 2025 b

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.963622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.963622Z digest=sha256:833396f57d63f77db7b45019259fbd7ec7eb8883433f6c994311abad0a5b1027

Observation fa3f1e1f-ee3d-482b-b7de-ae5f61ab6369 · outbound

This paper cites Teaching language models to solve sudoku through reinforcement learning, 3 2025 c.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Teaching language models to solve sudoku through reinforcement learning, 3 2025 c

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.967613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.967613Z digest=sha256:c61083da662feb231f6556d0ba5a1b83d8e0d8925d6a4fd2ddd12368a1605f73

Observation 68a90ae2-7eda-4475-bbf7-f4748bd22c03 · outbound

This paper cites Security and privacy challenges of large language models: A survey.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Security and privacy challenges of large language models: A survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.970955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.970955Z digest=sha256:67b3f2015bb776c2fbb98fb2cb592292ccb0f2bc6eb76699dc9af98beba4d4a0

Observation 842c4169-4948-4a16-b1ef-fcfba4180718 · outbound

This paper cites DeepSeek-V3 Technical Report.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepSeek-V3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.976613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.976613Z digest=sha256:9b15465ce3641c6f05e3c651c60d058350a5b71b5bc6440c56f4544414fbbf20

Observation 91a398ea-39f4-471e-9915-14d4d3093035 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.981170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.981170Z digest=sha256:8232ba5679ef77588489c3f8e07343a128101454c433c0068028cba5e96fd827

Observation 15ed7328-d0ea-4d23-b07b-6259ad79cc68 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.986239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.986239Z digest=sha256:028752b2adc8863d0c31a7263492175f0d780cae24ad51bb38d691735c66a890

Observation 7d6325cf-c9c2-4a59-bb4a-59bedf5c1323 · outbound

This paper cites Rl, reasoning & writing - grpo on base model.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Rl, reasoning & writing - grpo on base model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:43.993212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:43.993212Z digest=sha256:f3ebd2474fac529f16a46bc874aaac46e3b7e96480cd86c1cd72c643da2fae66

Observation 10c16ef4-b66c-4c1c-92f3-71d812140c20 · outbound

This paper cites Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.001435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.001435Z digest=sha256:9464b8c701dbcb85a8c9b67a807d55d1f45603c39f432c386e6244eb3cabfa00

Observation e165a717-3b67-4b88-b811-db7d64cf7964 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.006174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.006174Z digest=sha256:f4aca700e3e04df24bf0bcbd03f66fbb072db2c9b494076c47f14b334b04bc94

Observation a2ad7230-6a31-4da7-be4c-ae366e42e6d4 · outbound

This paper cites Reinforcement Learning with a Corrupted Reward Channel.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reinforcement Learning with a Corrupted Reward Channel

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.019777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.019777Z digest=sha256:d833927af912b9605643c87346e3c351277cf8928e455c8c694cbc4662d5812e

Observation 9291da73-4fc6-4dca-b42d-1a1de474dd68 · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.025443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.025443Z digest=sha256:874231ca9a3809313f6d402ba9899801081fdbade0cfff19fa4dc2cf3087bbf8

Observation 06504b25-192d-4195-bbf4-a58ab1f83f72 · outbound

This paper cites Reasoning Does Not Necessarily Improve Role-Playing Ability.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reasoning Does Not Necessarily Improve Role-Playing Ability

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.032419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.032419Z digest=sha256:951bbb598f7a6e36468f5ab7f352ce9962125d1257e27b4e90363a9d87d9508d

Observation 3680ba6f-d40b-4cd3-a2c6-0a43099e1473 · outbound

This paper cites MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.037426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.037426Z digest=sha256:61a2e573dd7d4798876428e9c590d36510fa990c9e52e4de964648601a984232

Observation 8d52c3e7-5fb4-44d5-9ecc-e8f3e5dd051e · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.045196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.045196Z digest=sha256:81c35ee2b29bd943485742392e7431ca510bec2a402dee598c88bfe29529c27a

Observation c8791a3b-9453-4a93-81df-3f1645c89627 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.050673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.050673Z digest=sha256:062351f17c57c3e1ef24316aa758ee71b5cfcc1ad73288a074cdb5e43a532906

Observation b75ba7e0-d583-433a-b684-a9e26f4c86c6 · outbound

This paper cites The Llama 3 Herd of Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.056956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.056956Z digest=sha256:eae5c75f9aa9469526f514ac52e3f107d3daf23a475357e2aff9ce1045c3a132

Observation b8e9d880-62e1-42dd-9f0c-435166cbec4c · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.061399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.061399Z digest=sha256:bc57fff452a2fe8d49a6113ef71933fb2e01a5c94e53181377ed8681849c8519

Observation a5ca0d26-8934-405b-a674-102b5789d6e1 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.071707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.071707Z digest=sha256:374fe8a491406c2ecc805a356cd1d672f091dd33c8f4331cb49a6c20b3246621

Observation 1d6279c5-b26a-45cd-873b-c5e51377c159 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.082711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.082711Z digest=sha256:9f630f34ec94113e2a9960962d9b17a026529683c058b8aeb456d24a32f91596

Observation dab3e8d5-0853-44ff-8370-0c71dc4c0d84 · outbound

This paper cites Skywork open reaonser series.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Skywork open reaonser series

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.087659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.087659Z digest=sha256:488c5620341508c67f13e13db2f451a50c69bd5df7eacdfe86c0b7394cdc10fc

Observation 36b44c6e-d536-4be2-a50f-6b218cab3c47 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.092433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.092433Z digest=sha256:52753884773ae49d5e9891a2fac32ce02e2918a817f29241ad33d780bc716eb6

Observation 9d0e7856-73f4-4ed7-8bd3-d410e6b5035b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.099187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.099187Z digest=sha256:18e0d36cd20d9142e4c37beb628b94070a1816a7786b1f758f2578ad36f3f814

Observation 3d6368f4-43fa-4ac0-900c-1c5cef74b7dd · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.105789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.105789Z digest=sha256:c3b86369dbb8dcacf329fdc52679ab905b6fa496df8c0a82c86235c4440aaa90

Observation 264223ce-1e6f-4a4f-a21c-bc173a677681 · outbound

This paper cites Curatedthoughts: Data curation for rl training datasets, 2025 b.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Curatedthoughts: Data curation for rl training datasets, 2025 b

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.110512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.110512Z digest=sha256:4111bd7086258db30474dccf20dae4f1cfa58b5ebf4db97dad163b9732e9b1e9

Observation e704c7ad-f7a5-4fa3-aa49-08195c4722c9 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.115501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.115501Z digest=sha256:9a8af98e3928ddf1007f66b4f051505a8c0b5982975875f2f686d735a2244430

Observation 7919bc4a-540c-4556-8b10-0fa91a12a07f · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.120856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.120856Z digest=sha256:d3c03e6f7f4fcf6da12a61dd31ebc13806ef793552b4c0cf6a81d2653ad417e2

Observation dc8a8724-a5cc-4de1-932a-8be56b37bd1d · outbound

This paper cites RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.124530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.124530Z digest=sha256:6f03f77a5bd78f40960be2f8b91e7c78a5553cb5cfd1e71b3bb65e4ce7bb9067

Observation 2742fabe-428d-4cfc-9ca8-4c6dfdeafd59 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.129405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.129405Z digest=sha256:253602e76050f00af76ab125600ddb309d30f298ee1b0032da073a6d3a8a7ab1

Observation 9b758c15-4145-4132-b9c1-afb7fa7159b5 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.133728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.133728Z digest=sha256:9bb2a8daa3e678b74bb6d63d9f5109c081360a89ef366ccbbe56bfb6245f0b05

Observation 81325677-905f-4287-9ffb-5f781852b2b4 · outbound

This paper cites Ii-thought : A large-scale, high-quality reasoning dataset.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Ii-thought : A large-scale, high-quality reasoning dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.138731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.138731Z digest=sha256:e431bb1dbf2cad4ba755dd5c3c786120a8724dc9840fd6ddfa72690caa63afe1

Observation 94593af2-4786-4bbf-966a-26c4180612fa · outbound

This paper cites OpenAI o1 System Card.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models OpenAI o1 System Card

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.144299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.144299Z digest=sha256:495edf80d66bf0600f151ecc32c9f549583e7ff06dae0ef289e937e7e3e821b6

Observation 2b8911b6-556d-4591-be50-69962d61e3cc · outbound

This paper cites Rlsf: Reinforcement learning via symbolic feedback.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Rlsf: Reinforcement learning via symbolic feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.149253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.149253Z digest=sha256:a6e0db242bde66bc210292dace7199db72cf971d256a16238f8cad81d5523537

Observation e4e1ff9e-115f-4005-a912-78a432dafbfd · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.154696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.154696Z digest=sha256:2ed4407c1eed923a66790548a99457bfa57785a9f6c72bae09f40ec98be79768

Observation c15495c8-f115-4aa3-bb20-e30f3f6da3e2 · outbound

This paper cites Challenges and Applications of Large Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Challenges and Applications of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.159814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.159814Z digest=sha256:018c23a0a620e83873917db4ec5aa19bfbd4176826156ce5e1d6d30b871e1118

Observation b55d0dec-73e3-4d7c-918b-40de3708833d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.165181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.165181Z digest=sha256:34df085bfe475ae46febc6ed9c15bd170c616ea5874f73c2ae1d5dd1b81b998c

Observation 8fd814f1-5d9f-420c-8655-7239f1518449 · outbound

This paper cites Overthink: Slowdown attacks on reasoning llms.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Overthink: Slowdown attacks on reasoning llms

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.170960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.170960Z digest=sha256:e3d95724409b270a56568c7d26c43cdd3d448230139e32e456d7bf322bd064cd

Observation d7479e89-db88-42e0-be82-2cc3d1119f42 · outbound

This paper cites H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.183385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.183385Z digest=sha256:dbfa386a08c870531dbeda085ee32946e424825539d50cfc75c0efc59a8d7c96

Observation 5991164e-8c26-405b-b9c5-f9f3c0aa6e6b · outbound

This paper cites Math-verify: Math verification library, 2024.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Math-verify: Math verification library, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.190597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.190597Z digest=sha256:2cd0419927777dee44dfacb60a101744549801c2bfdded8135dee8b21888cfff

Observation 332fcf98-a569-44f1-a0ea-be3d802202b5 · outbound

This paper cites mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.195659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.195659Z digest=sha256:b98b711018df07d15bce4f627163332e53c0f05ac3fb958c1f3a19b088c6fde0

Observation e60b414b-4b68-4c0c-8972-bab8c3b42498 · outbound

This paper cites Introducing superalignment.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Introducing superalignment

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.202387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.202387Z digest=sha256:03f33a6c2795139579cd73113f18bff3101fc68163e95c51211bf0e8f91f0d3a

Observation 5e73aa25-0587-4aac-a69a-6e59d6163a0a · outbound

This paper cites Exaone deep: Reasoning enhanced language models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Exaone deep: Reasoning enhanced language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.206947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.206947Z digest=sha256:4f959fac7050cf5664f39deb2c1a62d03c6f678293c918c3e5992e3eea194d0c

Observation 60de9cde-9c75-4160-9e52-0c8fde14d64b · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.212452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.212452Z digest=sha256:c097b67845a2a984fd1ac1e698be720e716b148aa4199203c167c3b5fdefd729

Observation 4d2baa89-6623-4f8f-82aa-17b6953abd41 · outbound

This paper cites Numinamath.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Numinamath

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.217133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.217133Z digest=sha256:9edb166b328c50f2e812355bf0f08abbd98f144df8b02e8c0a8a202c6061cc47

Observation ccb18a27-11ac-4e4d-9b38-88f173ee4f50 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models TACO: Topics in Algorithmic COde generation dataset

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.223637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.223637Z digest=sha256:b2a870a4868dce782ea462c100b1c9733fdb1df0fec8ead6abed2c34ae3b281c

Observation 14a587b3-fe87-4db8-9ebb-748a1039dd17 · outbound

This paper cites Large Language Models Can Self-Improve in Long-context Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Large Language Models Can Self-Improve in Long-context Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.229927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.229927Z digest=sha256:ba9044f91e4cb228e5ae41a7b3913c1eb2f1752651351da69b9b0ff32226416d

Observation 3f16efdf-be7c-48aa-bbfc-13374a57dd7c · outbound

This paper cites Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.234950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.234950Z digest=sha256:9f0f6a11f50c4b0e497d76f9e70a358fd42460f0905e847b2c8848c452d4d91c

Observation bb5e8f2f-9fcf-4fb2-a0e4-46bbdeb4f310 · outbound

This paper cites LIMR: Less is More for RL Scaling.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models LIMR: Less is More for RL Scaling

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.239798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.239798Z digest=sha256:e683b2e309b71cf9a5b708c8d482bb3e0bdf2c9a47b5b586267d0354a4cca589

Observation 8c52a279-373d-4c57-aa91-d593b5b4e3b8 · outbound

This paper cites Output Length Effect on DeepSeek-R1's Safety in Forced Thinking.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Output Length Effect on DeepSeek-R1's Safety in Forced Thinking

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.243755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.243755Z digest=sha256:03d529a5b0e413987c542809de203ddc61a661fa84cdf94c77ee100ae4057dc0

Observation 59986456-26a4-49de-8c40-e5c2807722e3 · outbound

This paper cites Small models struggle to learn from strong reasoners.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Small models struggle to learn from strong reasoners

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.247422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.247422Z digest=sha256:b68aa02a582dbd0e398fd4cae804dfde86e3724a2497387de640cb9f98444015

Observation 9424f790-a9de-4953-8fb5-eb7d4cdf526a · outbound

This paper cites A survey of multimodel large language models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models A survey of multimodel large language models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.251211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.251211Z digest=sha256:4693f7f18b06c674d4150359c7811f2ef79cbd6d6593d57160f28ee5be07c0a4

Observation 1bf94511-93a2-499a-9dc3-7b262485f806 · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.255840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.255840Z digest=sha256:f5efc5fd97e8b845a51d55f05bedee8e165a119df57531384c41ab0b539e16af

Observation cb84d433-49da-4ddb-8d63-5f4ec874af0b · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025 b.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025 b

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.260455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.260455Z digest=sha256:06fa76960ad40f634a77b8085b89d813abe78d5398b9f6f695f0d641b1f24a3f

Observation bc0d6079-6d76-4cfb-9aae-20c604b4ff1e · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Code-r1: Reproducing r1 for code with reliable rewards

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.265382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.265382Z digest=sha256:0612a84d2350e5695286086273a8057c96cb4d6cea904532674a0545045dfdd6

Observation 3496bbc7-00de-4828-a669-bd813145e4e5 · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Diving into Self-Evolving Training for Multimodal Reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.269509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.269509Z digest=sha256:999a5dcd10e81db2ad26a16fa59d696da3af00b0ceb93744ad703e7b109e3e64

Observation 9ca07818-0809-4e4c-bdfc-d31e0c5db510 · outbound

This paper cites Guardreasoner: Towards reasoning-based llm safeguards.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Guardreasoner: Towards reasoning-based llm safeguards

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.274565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.274565Z digest=sha256:b3c4defc4dacdb27baafeff37c369bc167ae561fe8ab9209818f097420291a42

Observation 94eeefe6-ae10-40f3-abdc-2618ca81a267 · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models There may not be aha moment in r1-zero-like training — a pilot study

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.279831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.279831Z digest=sha256:e090950b5fa959965454745aec9551e34d0ce8cc65f6c26e012114481145e8b8

Observation 4b2a3a0b-18ad-4406-8525-a2b1c50b5294 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.286554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.286554Z digest=sha256:cac90f3c1b93a81e3cf7341d01f86077c505445fd48822bcac2532988a311cf8

Observation 4cb61198-0b13-4456-b6e9-17cd2c61a76c · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.293430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.293430Z digest=sha256:2a6d38faeaf70681286189a6331df106c3b98ecca326255a52048d647caf188e

Observation 498f85de-3130-4e88-8950-cb711b149688 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.297852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.297852Z digest=sha256:a40702b3acae50d04cf765b89a13fb89e6298ac484000eac117f7fb41f2c6319

Observation ba918591-2bd3-4234-b5e5-24366a42275f · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.302502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.302502Z digest=sha256:7273ecf745b26a0036a83716792f8cfac8bc2d71d0201d80424ff3dcbc627e37

Observation 21f50789-98d9-4c5f-ba51-f3255275429c · outbound

This paper cites S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.306712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.306712Z digest=sha256:264d89c8183a0abcd5f12502ed61b8bd199d11f50a7f7733f2c0dbecff86de15

Observation 52293daa-1209-4ad3-ad7a-f8bc2b0277cf · outbound

This paper cites Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.310863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.310863Z digest=sha256:3f6a3acb8d19731314f0d56e4eb0696bc63862a32fc6972b5f47f8ecc2b3930c

Observation 203b37a9-0ece-4088-8dba-772aae9fc564 · outbound

This paper cites Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1, 2025.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1, 2025

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.316099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.316099Z digest=sha256:810c82379cf93ed9ccaf6043921e68dcd31a0b8f731ad75450572065f25a121c

Observation 52bd7f95-618f-4bea-8cde-bd8ea93d63b9 · outbound

This paper cites s1: Simple test-time scaling.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models s1: Simple test-time scaling

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.320176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.320176Z digest=sha256:f1fb2ce143dbe85d5dcfeb7b66841c743cba74db390f0fdb26e3568b9696c1bf

Observation f8786328-0db2-4dc0-8151-e9e2c3f9fbb5 · outbound

This paper cites CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.325608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.325608Z digest=sha256:2df378abc20bbef18da2e8799a9467188fa1c0e78ed8e0229ffece01795f7d61

Observation 90a9eb36-f8c6-4273-ab7d-16aa9933bfde · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Spirit-lm: Interleaved spoken and written language model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.329887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.329887Z digest=sha256:3507f23a827d04f731769bac6cad5f1f9ca7f3e72f6fc8665b58509a46bd29ab

Observation 5e08302b-f578-4164-bca1-4c07b65cc56a · outbound

This paper cites Gpt-4o system card, August 2024.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Gpt-4o system card, August 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.333747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.333747Z digest=sha256:3b91e6f672e01da9b153261f4b01b638e470b7f25654111db1c2ab42cfcf292e

Observation 94a7770e-8221-46bd-97e8-37b142f98e7b · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Introducing openai o3 and o4-mini, 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.338042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.338042Z digest=sha256:68139eeadf76e0c15067c72ea3064f5c02cda1b8b33111efe24272ec8f8433a8

Observation 8bb88ca0-4883-4445-a83e-6e7374252f7f · outbound

This paper cites Open Thoughts.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Open Thoughts

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.344468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.344468Z digest=sha256:ff1f1f9fa6211e95d52ee80a9caa197c3860f89df59494535e182a1d92bd5b75

Observation 0c7bc8b6-6ca1-4e87-b5bf-d7b80afbf320 · outbound

This paper cites Training language models to follow instructions with human feedback.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Training language models to follow instructions with human feedback

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.351074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.351074Z digest=sha256:abf88f3e7c49e867ed8262e02c4d648051137dd325080e690ef2f6de22e07ac4

Observation 2ef7f58a-5a65-41e8-a0cd-20f4f5801c74 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.362469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.362469Z digest=sha256:97b2c5abb26c073bf0b3b619864d4f58bd161297b3856e3f725412bac90563f2

Observation 0bc5f446-4463-48de-a024-fa8be02cf558 · outbound

This paper cites Tinyzero.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Tinyzero

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.370194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.370194Z digest=sha256:e57a8e3def0ff85a04e21cac87ea34fe233dd3d6653ee2b40c482cd26f8bdafa

Observation 5372249b-0709-40fe-880c-5eec9d852c48 · outbound

This paper cites Codeforces.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Codeforces

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.379471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.379471Z digest=sha256:bd63fb2301a59c8a5d59d683827d80a3be37d7ed98b90c534b805bb492d27130

Observation 06719fe8-cfc3-4917-b2c2-8699f22cad61 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.385068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.385068Z digest=sha256:a84dea800e504fc48bfd627d05a78bbc40c6c4d6dcecee16de8299cd9d29d1b7

Observation c8c7723f-ef14-454f-afd2-358cdb54fec0 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.393377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.393377Z digest=sha256:43e732d8e25ab6914a2e6d22a84bffded1e246fc7a50c1782904ff33d2956160

Observation f967b2d8-ea1a-4767-84a9-440d24fe44b7 · outbound

This paper cites Qwen3: Think deeper, act faster, 2025 a.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwen3: Think deeper, act faster, 2025 a

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.397390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.397390Z digest=sha256:62a0626724a0da6330409e692125350455881852d82629e9e93d4718644b2313

Observation 0e9f0f9f-edce-4c41-9927-3e9ab03d65a9 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2025 b.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Qwq-32b: Embracing the power of reinforcement learning, 2025 b

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.402354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.402354Z digest=sha256:5abb6d975f341cf3d8d15cb16c0762a451b70f7891035863aa068395cfea38df

Observation c4bff36c-2ddb-4015-990d-ffda5b89905e · outbound

This paper cites Manning, and Chelsea Finn.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Manning, and Chelsea Finn

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.408738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.408738Z digest=sha256:538be4db9f382e1cae13ef6a777904011e6be2970f3014575a8c9d72cb56962e

Pith citing papers

Observation e17fa188-ac60-485a-8b92-9650eba75225 · inbound

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models cites this paper.

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:46.353798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:56:46.353798Z digest=sha256:58e5caa630399788d6a93b0cbad83a7ccc4f566f9a96e610d6e2cbaee7f2fee1

Observation fbd93588-3794-49d2-b589-10b7e5da2a54 · inbound

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning cites this paper.

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:19.843039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:19.843039Z digest=sha256:4ad6ce789e5799a4d523fd4c4e5ebf9ba5557382ee9cca10699d79eb80a12e75

Observation 41b48aa1-d17b-4230-8db2-f10c5ecc337c · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.242194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.242194Z digest=sha256:e7d401644e0682089006054292f5daf4e5cf0938a9deb28692469e33db64f86e

Observation 426d76a5-1199-4889-a3fb-0e9c2a936848 · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:59.831868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:59.831868Z digest=sha256:33372b196e5b8048dfbf3fff57a49791d57f8f15fd73b93b53732e67addcb1a1

Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · inbound

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy cites this paper.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.860058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.860058Z digest=sha256:748092cd35f5bb79a696741d70bd719c1ac868207325d2826d8620333ddf0a14

Observation 0404d396-0421-4503-b026-64648534fb39 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 238

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:18.056863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:18.056863Z digest=sha256:7341b3aed453b0ed0e85137ef243949bbf1d7ddbb5f90b67ad0bd7f441c59587

Observation 959f1a18-0010-458e-a947-05e298a39acd · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.618381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.618381Z digest=sha256:bafcbafc26e1e3cb22bfad037000d49972e00cb773f25f659e7c4666820b8989

Observation b3712cee-75c6-4c3d-ab03-29e56857fac0 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.945828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:dbcc7c7961c4a7be2d7d8610d84f74a5631bdc9b68be9abda0d072521993de8f

Observation a4c8a8ab-0d59-4890-a6fe-6f8b65bba7fe · inbound

Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction cites this paper.

Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:09.438194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:10:53.620510Z digest=sha256:dddb1277a0b11790b540194ee87923302261c5dc5bd9a99a1c0e61455a61ef7d

Observation 7cfe54ec-dac8-4315-9e43-a6338af90cf7 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.444066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:ee9a3e965b72cb288013e0e3060689377c5ab5f3f8260a5de58646e7e1a204ef

Observation dc303744-dedc-4b76-b87f-a00cd8383882 · inbound

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment cites this paper.

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T10:27:02.432515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T10:20:46.103409Z digest=sha256:5fd189db2f20ec10c4b49d41c12fc9d0c101a543fbcf90c76142d9cd3533f0af