Pith. sign in

Paper Citation Record · LEDGER

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2501.12948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12948 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 2561 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:51.467877Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9d11db15-e491-4f63-8eca-9369f2313f81 · inbound

Scaling and renormalization in high-dimensional regression cites this paper.

Scaling and renormalization in high-dimensional regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T01:55:55.085863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-24T01:54:48.781227Z digest=sha256:c91f10e5381d1d0563ef49973864c7fe4292740d7b2a47e53ec16180ee87d100

Observation 92948f84-05c2-4d2e-9fcd-845885754947 · inbound

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework cites this paper.

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:28:57.146429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T03:28:57.008431Z digest=sha256:94a793f13ae4ff3adf5225e788ebd7af305d24ea82a13d4e27c0a71786d7a902

Observation 762eac57-ccd7-48ff-bcfe-646e64a5986e · inbound

Retrieval-Augmented Generation for Natural Language Processing: A Survey cites this paper.

Retrieval-Augmented Generation for Natural Language Processing: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:08:35.723552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T23:06:41.081461Z digest=sha256:757434f4ab35667d6a387fb1251f2e66de9fcdf3f17e48ef01263f5486aa8e2e

Observation 27a7bd9b-26cd-48c0-bf15-ddcedfb63719 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:16:04.675720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:b6e92bc98d77b3431e81663caf6411d6e5fa3b107018b623d629cc451c4084bb

Observation 74586e5a-1166-476f-ab26-1e2418675231 · inbound

Enhancing Clinical Trial Patient Matching through Knowledge Augmentation and Reasoning with Multi-Agent cites this paper.

Enhancing Clinical Trial Patient Matching through Knowledge Augmentation and Reasoning with Multi-Agent DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T17:03:12.477184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T17:01:30.638046Z digest=sha256:459103374b31cdd7ad65c33673bb7f7d653d792d35d6078f32aa49c8a1596218

Observation 6a332db2-c5c0-4cf7-943e-6ebdfb87ea26 · inbound

Training Large Language Models to Reason in a Continuous Latent Space cites this paper.

Training Large Language Models to Reason in a Continuous Latent Space DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:29:05.677264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T10:29:05.384381Z digest=sha256:78f523a0ab1f2b4584634389c4c8fdc5107a3cb5cd565efaeefe5dd1c044c623

Observation 206a5d9f-1c6e-4102-880a-1b846521b70b · inbound

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation cites this paper.

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:22:43.149402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T07:20:57.953573Z digest=sha256:969497d39ca37982e7ee6f849154c65a68428f9e87e36975c53a29cc4227372c

Observation 1776ca05-fe61-4453-90f9-f98e4894e39e · inbound

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation cites this paper.

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:33.975137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-25T08:12:15.133695Z digest=sha256:0038122f163147c3159a5cb7348627be652bae1610650641296eddd0b58f0720

Observation cfe1a6af-393e-4e2e-b8f2-dee6f9727076 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T15:51:29.231425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:eff26139c0d72dd1750d3bc4488a16c90dad48938e2674a2d77d2bca11708b37

Observation b707e1b1-845e-4c85-8edb-415edac1f881 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.448157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:e0533e0bc24fff1c7a3a65a8f0d2087669cff734634c4aa3694fac786e23c5a4

Observation 8767a48a-0425-4028-a7f3-803aa0e0d932 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:23:30.917683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:8c3085bb25000bcef2a761f9e61a07a81b4145883960e80d85f49c3d28496e54

Observation efd852f4-5519-4f3b-86e5-9a19c3a810d1 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.228380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:e24b222e036ead108744c91bf7214efb25b8c7123332554417de4c0fda3bcdfe

Observation 32bd76dd-b6f7-4929-bc4c-c2f67ca14f06 · inbound

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models cites this paper.

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:34.073158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T04:27:59.317818Z digest=sha256:b3d5b57b6387b88d462e25524a2af35f71b4c11bbb2d774f500db6566329d16e

Observation 3cf708f7-52c2-48eb-b279-124bee20a0d3 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 130

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:11:37.552410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:f2c95f855dcfacdf506c55e89f0aae9f8631d02518e545fe51c0a30ab6cec0b0

Observation 9a1ab93c-6bbf-442c-80fa-e03deec7a1a1 · inbound

Large Language Models for Multi-Robot Systems: A Survey cites this paper.

Large Language Models for Multi-Robot Systems: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:32.071591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T04:32:05.138744Z digest=sha256:e67975c306736e2bb637327ba43558716eb0f4744ab651013586d3fc062f638c

Observation 38cfc498-da59-42ea-9fa1-cd41c6c5a4cc · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:39:41.085481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:3baa338b30a38ff350f95f24dd0bf1653959c9d0282dcf717245ab610cde31bb

Observation cbcc6279-61f2-4c9a-bc04-55a63117a991 · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:42:54.763954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:3203917eee7da2e669dbd2437076e6a43f6fa735f99c2c9bf3edf34d99558b02

Observation f6b0c620-dc27-4df5-b667-8f25f95a4d3e · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.544061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d9d518d7ed7eb20e5f781e5cd9878d73de90be46ecc907258944c2a4ba1dff82

Observation 4c454568-6b64-4975-afca-cc83155628b1 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:46:30.038088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:b06e60bfdf92a5e15190098c2380cab05d0d014c59128073ede7b71c85183fec

Observation 7ca610c1-2904-4f2a-b3e5-93500c54d47b · inbound

A-MEM: Agentic Memory for LLM Agents cites this paper.

A-MEM: Agentic Memory for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:47:28.994970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T00:47:28.898112Z digest=sha256:53b7d37898e092334dde47dff057168e5f5b4f67cc6613476106459f9be8707b

Observation 11497c5e-639e-4a54-b725-c48e6ac33320 · inbound

Hallucinations are inevitable but can be made statistically negligible cites this paper.

Hallucinations are inevitable but can be made statistically negligible DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:25.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T02:41:59.640979Z digest=sha256:d3e75d45c57875ad15cf01fc12dfe2878c188d660c6ae19519e787b6c07b677f

Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.069190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8599042ecf3c4272973699d8f2479ae650af1f7b621962f5e591f14d0501c629

Observation 44924d88-cb1c-4f33-938c-4cb66f2ec26c · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.162426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:6eaf2addfe390cbc35222e435c4f6a7b2ce1179adbfc4f0d1f597bba0b4763d0

Observation 7f780ee7-29de-460e-9c07-66b590fa0526 · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.441515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:6432a0c7e730e157c9b3fb3e9775b8016b1f244e30d9d0305491a988f52a5466

Observation 84b52af0-bc9d-4b88-8ac3-3f9fd90962fa · inbound

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning cites this paper.

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:52:22.898183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T01:51:12.410547Z digest=sha256:e34d21c4d8eb75e8bfef5f15aef3acfae509f948fecffa8b894ae36b4b988ca4

Observation 8555a31e-9eef-4501-81e5-0bff8f3eb4fe · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:36:24.570662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:4bc6163f7c64b5d320d6c17bf903fe7bcfb87237690d845412429c5af461c96e

Observation 8d8df19b-750f-4734-9816-e2ac6d4dd1b3 · inbound

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification cites this paper.

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:57:22.756536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T01:57:19.398245Z digest=sha256:313571e9792832c3fd89d0034d752ad3b93f7a0ffdb3ea400a40dc17f23b9ce4

Observation aea0d5fa-1a47-408f-b966-5346630614bc · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:27:56.335474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:da7c585f6a6c02139d17fa82e5930560077d42a076d1b2a69e5c9d77c92da97c

Observation 36fc85fc-539e-4463-9562-b0c8006de306 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 242

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:45.420842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:7dce058e99979b323360b4b64f8c906b7236b3d1394b8432e299f39f5bdccc04

Observation 8a9b378f-7981-4cff-ba94-67a27c73d481 · inbound

Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models cites this paper.

Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:47:26.380429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T02:46:20.578680Z digest=sha256:b4c47cbcc219d62f7e0305e4af1968a8393e8a5d6e46e2838a88afb27f9f6b29

Observation 603a0f70-7dc0-45ef-acb9-4082d8366ef4 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.002945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:1de232bcba38c0e1ad1eaadbe68dad65369aad03a740dd10d3430b886c7c982e

Observation 116978b6-fa7d-444b-9bb9-b75b9f716a96 · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:16:16.557271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:b0cb5338dbd1f6d0aac11115eff329931dd52fe1d75584feb92f02e758552b6a

Observation 90502821-c995-4e39-828a-1d268bdfcf66 · inbound

LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion cites this paper.

LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T01:07:19.886029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T01:05:38.969311Z digest=sha256:0c24e865e49d2d7ab86472a032670101804e4ebe6870e208af18a4e45de9e0c3

Observation 1dd3c76e-bd63-468f-9143-dfa452f4a92c · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.201688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:22dc6f3fd000ee9f5fad35da059b2a0070b8a64cab354b9ac8fbdf85ff8c9ccf

Observation d149f89e-0910-4480-bdbb-6328b9b2f829 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:44:30.828729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:6462e3b9b7dc466fe34e44e74970dcea2693464beeae76fbb38e238d226edfe9

Observation 54605935-9cbe-4788-9153-4689e2ae971b · inbound

R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning cites this paper.

R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:37:17.756376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T18:37:17.725984Z digest=sha256:89cf4d648aec67ce2bd0942f89ccc069666c32f2af9a696c1df2997bb1d9798b

Observation ad47c6a7-71a2-478d-b321-b7802d3d672c · inbound

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement cites this paper.

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:31:43.568223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T12:31:43.494099Z digest=sha256:3abd83fc803c6090fd1a6e387b061601160e2033f1fd79cdce6d7c16a0c7e6ea

Observation 67bd6635-cb2d-48ec-9332-2d3ca94b3429 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:15:46.399734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:47207bd4d2b07083913c6b10cdc142fbd7b1d9da817a28d8e607bdfb73370e32

Observation 7d7f2f93-19a2-46c9-af1a-2229716c3e92 · inbound

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning cites this paper.

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:06:27.200708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T20:06:27.136345Z digest=sha256:5f15bf22ff0cc0d630d9a687950d4bb73374466701a051dd3b9d9ff4c028e489

Observation 35165ca8-d2ec-4554-95ea-87583becc981 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:22:18.727457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:e67dc80ccb6fc4bc2c8143657c9e2f5aa64d3da6919e00fccef10b580afd1943

Observation 74d9fc96-deb9-4561-b7fd-adf08d6711fb · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 231

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:40:41.437202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:02421591465f37dfb6a5cef933c8db874aa0fdde0779140e6dae9fb17eca2dab

Observation 0732cb83-ecc1-484e-a2c3-88b974073d01 · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:32:18.650809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:8b19f4664c49a317b73ffc3ee36875041617f963f9b64725773f6f9210d9d104

Observation 971aa50b-a162-4170-82f0-7306414ca55a · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:19:20.514762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:3d357889f0d4ad3e48c82c20243b4d686335c9c0b079cf602d3267c8684fb37a

Observation 85e0a5c0-80ae-4228-a69f-fe56867a6191 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:24:12.926621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:81a5035e7618cc1764f8c35113dd06869750e9461d331f83bdea7e2876fe2cf2

Observation f962c689-0269-4631-9644-d409a64b5968 · inbound

Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios cites this paper.

Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:52:19.444673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T00:50:05.344148Z digest=sha256:5dbddb9b2c1b2dd4897bcc01fd4e8381479189d746198bb2713dd088bea3a5e5

Observation d96d1e8e-4f65-4208-ae60-924eb7d998fb · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.213897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:29c7ab97fb4c56727bd8329727d9866cfc713b4095c6c41c67912bbc5ef66edf

Observation bd13eaa7-9b6a-4cac-849b-0abf7dd9a7b5 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:04:22.832542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:a788e9a256cbc4403902f925b945535ef8bcf3d34c23626d034406857020fb60

Observation 096c1a33-4843-4c20-92da-e0ce2affb49c · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:40:06.573742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:17cc2c7b0a05d299066b619ab8064cefb30b0982be002d47a01a49336c8989b2

Observation ea8125e8-dd16-432a-a319-8b2672e690c2 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.859435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:43e29050d40b8369611f76bf7444878c12103ac04430bf26f190a55d571b0265

Observation 82c25e79-80e3-4b0f-b478-56ec6b11807a · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.475103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:ea0a620c08f0824d66db2773f1617a6fa1d993f261ad4bc7b481deee815df90f

Observation 6802f323-9631-45de-ac66-59b5ea8423a7 · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.211498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:09d0f66f7d2a71727b2688d2cab6f74a6422c634f2def1d6adae1d8ecbacebd4

Observation 92cc2cb5-b885-4f17-893d-2f4b9dcb799a · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:29:57.336032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:d9fa0c049679c657e74a50942eb2a5f4640a595127666105a3032ff9a12befda

Observation 11f8cdf8-21c4-44ef-a526-4bca0cf0b6f6 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:57:13.221779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:988ca5e49381b30b310eb8253e6cdd8d6346e78671d06403c9b1275ef736fe26

Observation 1e93d0ff-364c-4c1c-8112-1ce67243dcc3 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.392088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:ebd118726b3ca43e68c256baa5ccadc5175c4d426bd856f9a46293dc899123b9

Observation 68e98038-2889-4a71-9910-6bc93cbc006c · inbound

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark cites this paper.

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:42:16.458629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T23:40:25.111680Z digest=sha256:267e805358c3bd2df27104ad24ffde37c0911b2c44c785ea33b684b71771b4e8

Observation 99ae216b-12e2-47a8-b664-f60cedaed612 · inbound

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning cites this paper.

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:48:34.759536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T15:48:34.705881Z digest=sha256:f37db9f069db3019e96db4bc0a8483656b2c37adad125597df14cd8b6f011a96

Observation 35bee7cb-f4ae-47c3-9f8c-c53efb43866f · inbound

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models cites this paper.

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:15:11.202017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T22:14:53.522206Z digest=sha256:1298afd65c1c2ef5d5b7ab16d13cee20f3bca6d3f3da25b69b7961c76d662d24

Observation 738d9a85-dec2-4ef6-b95b-21871b661bb4 · inbound

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning cites this paper.

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:02:41.372022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T11:02:41.335059Z digest=sha256:ddea63f713224167ac8470fad4f158d455650cfcb2d7adc6b3aeaf5db57ba222

Observation be9eb4e3-65b2-4cf1-b05b-22bdbf1e52b1 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:43:00.371525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:5b0291219d244672ae45465818c208719cfac015628e64f272c8055e64120b5a

Observation 3b310225-95b6-406f-b598-7f7f51011791 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:42:13.491102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:1c27e8d14a4e9584732f148ab70c90329fbaaee48251710489df431494659140

Observation bfa048df-635d-44b4-8afe-d09eac9af652 · inbound

Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model cites this paper.

Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T21:59:02.206491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T21:59:02.173820Z digest=sha256:7ea7bd62003f50e8b7db0eacedae5028e375cf9cba55844355316182cee192f1

Observation cfaa2480-7902-44c2-ab06-62f7dda100ca · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.778898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:b37a14d3acf981ae148a8c8a956a813030580240d66d5b263c10bafec19cbf48

Observation a2e71d8e-7606-43d6-80e6-042a3c4bd402 · inbound

OpenCodeReasoning: Advancing Data Distillation for Competitive Coding cites this paper.

OpenCodeReasoning: Advancing Data Distillation for Competitive Coding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T19:21:42.194865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T19:21:42.081762Z digest=sha256:a92f52773299927869b62daf274022aa72b23a3bf59cb9eb96a87cec80f87f95

Observation 2ce22385-4116-418f-bb7e-be2b3ffc90df · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 119

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:42:10.935652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:1b586b1f1a66f912356356de74a9fa69014d5b4e1d4b9abf5241de91d6aa1487

Observation bd97f8b2-89f7-4e8c-93fe-d37092050f0b · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:22:09.055779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:2d7934b1dcb408f995286080baf1ba3ce040d2515d7c5bfb11120d726223f462

Observation 4f456f52-de83-49c3-8b8e-8e5b282cc5e4 · inbound

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks cites this paper.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.755391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:6932fb3820e5c5a69256c8e605a6ddfc88ee80f8cd4810586d336d204648745f

Observation 8dfef9c8-9d04-43a7-9dbc-8400afbf2c66 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:23:51.684873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:ec18d00d49d27f611c3f62979abb379e4c0e097c363a03a4b2700a26c6f7d253

Observation 4280e525-0f88-4f93-b04a-b1fa86727d1d · inbound

ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs cites this paper.

ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:12:08.328703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T21:11:46.405944Z digest=sha256:35a7625a2a25450114ff9516feecbf6017fb706f14773aa303d318992874d77b

Observation 0949eb1c-3dcc-4606-8923-aed3099d416c · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.786208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:2c0d146571ca1002fff6e424e6934c8cc79c4c45812a955d4edc8149d214c8bd

Observation d65c834a-2372-447d-9e05-dbe0880a3c18 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.420952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:037dbab6c372e649326310697bc4d9d8c1fdbd42a2b3226e15f440e3bf3a7d1a

Observation ee895e6a-0b4b-4093-a63d-638f1b070568 · inbound

Exploring the System 1 Thinking Capability of Large Reasoning Models cites this paper.

Exploring the System 1 Thinking Capability of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T20:07:02.597495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T20:05:13.849797Z digest=sha256:7d1524babd4ef3fb1ea89d6741d329a5410225f534b9628eb094ddae5fbc00ec

Observation 20418ee8-95e6-451f-97c9-cb463c22b704 · inbound

GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents cites this paper.

GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:10:58.123955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T02:10:57.976448Z digest=sha256:cc06d7a72e947913b964696c85bce1a2c05365ae1cc3b8adf8db6fe34c9a8822

Observation 995d756a-08af-4938-8ce5-992406dc104c · inbound

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning cites this paper.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.799231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:f25135a0aa509f18c5d8fe6c07855a508ba4858cdb4a15bb95996ee8bdc15028

Observation 63eae499-6855-45f0-9437-3404fb437802 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:42:39.053511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:aa4acf252bf9a98b0aa28b106e3798850f1a4ace57387c495a036128f1d41271

Observation d8cef190-a21f-45ea-8c0f-5ddaead9f219 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.362542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:7439ad3dd4a047662af56a8999eeda4dc790685a6f957b69ff08f0278c2e4942

Observation 3f634a63-c95a-4e6b-a10d-c6a773c21517 · inbound

Design Topological Materials by Reinforcement Fine-Tuned Generative Model cites this paper.

Design Topological Materials by Reinforcement Fine-Tuned Generative Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:16:58.685680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T19:15:26.493687Z digest=sha256:82670d6afd49e977e66c3a4fb2f10b91d886eb97d6ebe4e534053d2b2ecbd829

Observation 59837342-2064-481c-b8cb-4758d66da240 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:46:57.178986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:c41b23ae81b56ebbd5ae0a796fdff2b476d439d5bdac8e75fb3ad7ea7ac8d465

Observation 0c311622-675b-42de-9781-55f79ba65698 · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:15:09.400280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:5d6b89fd12e431f6d7989f140183c790061ff900f3a8dc18ef5d61584ab45be2

Observation ba59a3e8-a1ce-413b-a6f7-379e4d7c8f52 · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T00:26:48.396196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:db5c94b276b66cb59e9d1f0f446dc0bd0d28652caf8ffe1d80476ca22fdf808f

Observation d4e41a9c-3bc4-432c-85fd-f29cb7206a30 · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:17:02.754382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:716c8ebb15bd6f30e20ea685c466ec75b5e3257466220a94eefd7c4dd3c5d6ed

Observation ddef81f6-addb-460a-8034-11e9c34cc2ee · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:36:58.742184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:a122f105af9f091903eb78f3b53587aa910fc9b6a5a35d7a47eba38e14af820d

Observation c061af56-75aa-4fb8-a8f1-a3c5ab10911a · inbound

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese cites this paper.

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T22:04:49.942408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T22:04:49.915916Z digest=sha256:3c87735a1ab5059b34ee657f016b1d4a99c86d29793a5a30a7700620bf2c9ab4

Observation 009483ec-02f4-4879-afa0-17775769c5d3 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:37.963947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:0a107cbf995dfd6eb4bce3d0983e47838bb54b08edc7b7cfc1be46f34a848961

Observation 843fc5fc-219d-41b6-8d90-e2ff72eedd0a · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.844671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:4cb9690a8ed5e42a041d82553e1b601c000cb163a88d032ab5ef6e247a2e165f

Observation caaa20f3-36fb-4418-b8d3-d12c69e68d78 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.767220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:5528236576692e554e5341bd9f71e8c5966524772a481b2b09f3532f4535bcab

Observation 6479e86a-1f1e-4d7a-b287-167f95917dc3 · inbound

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition cites this paper.

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:32:05.672242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T09:32:05.642611Z digest=sha256:48cbea6448b4bf3aba30340bb98d34b00009305d7b021568d47380395d826dd7

Observation dd04b8e8-a459-453c-b41b-44d68b16a649 · inbound

Always Tell Me The Odds: Fine-grained Conditional Probability Estimation cites this paper.

Always Tell Me The Odds: Fine-grained Conditional Probability Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T16:31:46.933633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:9dc1d5cd09c8fe8663763ae57e452a9ef3b3422c182e633bf38ac2842b27edd8

Observation 0322336e-a359-49ae-b64a-48297f7bbb6f · inbound

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference cites this paper.

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:35:25.153201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T15:59:04.724780Z digest=sha256:d32e92688938ab64deda0578fdd378fa2585a038584d4707dd104ec330deea58

Observation dacaf138-2ff8-4cdf-96be-7db464020e06 · inbound

SocialLM: Social Signal Processing of Patient-Provider Communication using LLMs and Contextual Aggregation cites this paper.

SocialLM: Social Signal Processing of Patient-Provider Communication using LLMs and Contextual Aggregation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T17:05:00.355095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T17:04:00.508789Z digest=sha256:a087979e80b8f6a2a9823adebb0c3eaac7d4d7cb35274f60a49cb017fac1aec5

Observation ce75d320-2c9b-4c7a-8e03-01b105ceaa31 · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:45:17.030802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:20c341dd144bd6e47ceaf76edacf167a871a2df3a26f76fa90a9b2de484d1d1d

Observation 2488832d-2931-470f-9507-c33112e66903 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:11:09.174697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:cf09cf7c0dcf77c98b8c6b0647cef811c0434f79194d499b371bb29bd9880fdc

Observation a7dd1a63-4a2e-42ab-bd40-8e5d44fd080b · inbound

A Survey on Foundation Models for Personalized Federated Intelligence cites this paper.

A Survey on Foundation Models for Personalized Federated Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:34:57.704097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T15:32:15.293888Z digest=sha256:fe0d3429b95369c54f059aacb4ab9d4ffb0287166a9130b7fb92eb35627d94ae

Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.042133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2a5804396296d5a8587ecd66aaf439a327fb112b8194d6982586757ab4bd6b53

Observation 5421bf75-b042-4d23-904e-9e889de9bd6d · inbound

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning cites this paper.

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:51:45.690186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T15:49:44.263123Z digest=sha256:900e4085834883caa6f1c9a4726ed3e3c1be57112c1e191928c8157d072b8567

Observation 11809140-14c3-4f97-ae15-67bf556b4edb · inbound

DanceGRPO: Unleashing GRPO on Visual Generation cites this paper.

DanceGRPO: Unleashing GRPO on Visual Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:28:26.684350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:28:24.929046Z digest=sha256:b5ce89c090b01e7360eeff0175b296047dd74f312270da27ec071011965cbf5c

Observation 76b95b6f-5f57-413a-b0db-a92597925aed · inbound

Not that Groove: Zero-Shot Symbolic Music Editing cites this paper.

Not that Groove: Zero-Shot Symbolic Music Editing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:31:47.462937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T16:27:44.903349Z digest=sha256:6439df5559a0d3df209bc58b89e26f857d6e2e11dc58e6efc2ebc729517fd5a7

Observation de132211-36e5-4928-943e-0d0bfd0276ff · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:35:28.514089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:c27475dacb0d4bfef0c3ebc580bcfd3bd2b99608d4a85b2e1fa60594e36be602

Observation 757208c1-62b3-4bbd-b38f-eb9bd58d792f · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:44:58.108944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:89b447fd8a4337c76aaa385fb49eff26657847919653a358cf2a84cf7abac8a5

Observation 73db4015-62f4-4f9e-8870-f6bc33d7eda8 · inbound

Superposition Yields Robust Neural Scaling cites this paper.

Superposition Yields Robust Neural Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 63

Resolution
malformed identifier
local_arxiv, observed 2026-05-09T06:36:25.553937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T14:38:44.789822Z digest=sha256:89b89d926fcc9438c7b8885d9d7c0876a8e8abfb4df4521babb47ea5b73f0f6d

Observation d3df8d7b-dcd4-451e-8682-aedd51add6ac · inbound

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction cites this paper.

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:21:45.070163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T15:18:15.475294Z digest=sha256:d49dde71ecb73c405d220b775b98bdb38a272b05894dc8910de0dcb14d2cfc34