Pith. sign in

Paper Citation Record · LEDGER

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 17 inbound Pith citation observations for arXiv:2505.17667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17667 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:48:00.129155Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:13.872622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.208570Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac731f42-ec22-4d6c-9ee9-351aa5787b09 · outbound

This paper cites Claude 3.7 sonnet system card, Feburary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Claude 3.7 sonnet system card, Feburary 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:04.276318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.108255Z digest=sha256:5b75ede8a8619239e2b2ea0e0d0aec752f81e878a36d88cdecfd5da75078561a

Observation 428b97af-075e-4a64-b2da-e51e870947cb · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LongBench: A bilingual, multitask benchmark for long context understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.174129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.174129Z digest=sha256:c8bb2b2c019e62b3e6c0308f1b4ceffc7b2ca268e68c38dd9a2b620b41855854

Observation dccea908-1f63-4ee0-a623-c2e2f93b4de8 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.301320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.301320Z digest=sha256:83a0c5680504d4e5af357e3f37e9c5250972a0ec779afc4f97bac433fe2ea928

Observation 5bcfd899-cebd-4adf-8532-464381b043e8 · outbound

This paper cites Thinking, fast and slow.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Thinking, fast and slow

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.367277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.367277Z digest=sha256:ac0cd228e7e563880793589b33130194f4bb4918818bc241d88084bfca791c43

Observation 2c80ca3c-6bcd-4ddb-804f-82c9bc6d6d7d · outbound

This paper cites A dataset of information-seeking questions and answers anchored in research papers.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A dataset of information-seeking questions and answers anchored in research papers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:04.052534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.465404Z digest=sha256:9eb581be02054086b61aa887e524179192361eb2441647d8b9e356d55bb4d76e

Observation 0ffeb1d1-5335-4866-8c4a-d046b8ec48f1 · outbound

This paper cites Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power!, November 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power!, November 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.542819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.542819Z digest=sha256:6662909c959a178270f53a9725e68c0a4db19fca0727047941232aa4eec63f06

Observation 76e19ce2-d0ac-4641-90bf-8a32e2b248e1 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.628265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.628265Z digest=sha256:0fbcc2cc826a2b547eac6b89d088465d6c0f62f0789a9ef586db29eeb7b6bf36

Observation 0555fc97-8e07-4f1f-a9d8-d1aadd3e6470 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.849948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.701599Z digest=sha256:fda3677e57ca9fb78bc4961b078bcff6096fef4972c1293ebae1fbaa93310c8f

Observation c4e7b238-38c0-4755-b60d-e34b0766870d · outbound

This paper cites Data engineering for scaling language models to 128k context.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Data engineering for scaling language models to 128k context

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.614383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.798913Z digest=sha256:b497569d4de9f74738a69a734a4e203716efa2c259fd72295646952e71896abe

Observation 50bb3bbf-3af0-44b5-b063-9be49de896b1 · outbound

This paper cites How to train long-context language models (effectively).

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning How to train long-context language models (effectively)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.906238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.906238Z digest=sha256:9e187020c9495bf644d0860715bc92f6ed90b42226729e9f98f887fb682bee33

Observation c031e961-f33c-4414-901d-bee4467a7c34 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.988796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.988796Z digest=sha256:d49d53f7aba34f714ebc3afb929ec33904f5c0dc4ded61ea00fc55c9904fff4a

Observation bbe9bc61-4b9f-4c90-9c46-294a5ab5a65e · outbound

This paper cites Retrieval augmented language model pre-training.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Retrieval augmented language model pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.096120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.096120Z digest=sha256:9a0145da809b4713eeca8bf36d765eb69caaeb4970c273c01e00f374b6e8ba9f

Observation efa39379-c508-4126-a5be-ad8c50dec571 · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.175370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.175370Z digest=sha256:1bdf82167ccada4dcfc847c9d2cf30365e234e4c8c4dbcfa8bed25bf8cef04f5

Observation eea31d13-f54a-4447-b04b-bdea4d8d786e · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model, 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.241433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.241433Z digest=sha256:26416badd535b564d2e36346cf4e034821e26882a5b4c3ac8323fb579a818f32

Observation 921fe487-7fff-4a7b-9e4d-dd87a0991bb0 · outbound

This paper cites OpenAI o1 System Card.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.301210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.301210Z digest=sha256:bcd96f0ebde7ec691f58e7f9405a426e51cd8a00c8a8ee3924ca247b331092e5

Observation ad90fe84-00f5-444b-b5b8-99ee3c97fe0c · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.375344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.375344Z digest=sha256:ed0c24a556765ef667c1df5c3671f4862f9182a7c580e20e51bfd215e5da04e7

Observation 1947ad9c-5449-4368-a48d-3c3da885c057 · outbound

This paper cites The narrativeqa reading comprehension challenge.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The narrativeqa reading comprehension challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.450413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.450413Z digest=sha256:854336a0dd8849169d89fc4bc76786b66097ce552cca4df1716bc71c84675fc3

Observation ba0e791c-8210-4b33-b1f6-9fe334c8d8d3 · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.502010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.502010Z digest=sha256:17e11c149405f90f7b4085e47f093751ad85f3850451a7e35319afa26c672fe8

Observation 0c3a635d-f495-4a84-93f9-e0578ea83ab5 · outbound

This paper cites From quantity to quality: Boosting llm performance with self- guided data selection for instruction tuning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning From quantity to quality: Boosting llm performance with self- guided data selection for instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.394518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.546278Z digest=sha256:0f4eef8d3f4e94a27f40365457933aaa61e70d321523c6a69a4aad53a6899205

Observation 8a195b50-4509-4525-9d5e-244ecf2c38d2 · outbound

This paper cites The unlocking spell on base llms: Rethinking alignment via in-context learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The unlocking spell on base llms: Rethinking alignment via in-context learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.197931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.637448Z digest=sha256:cc972594433958c98df9ab7ad3ee90c730d55b870ff219007c69d24577bc65ca

Observation d3c2ef07-9e66-4961-9b0b-8b5bf28a5099 · outbound

This paper cites DeepSeek-V3 Technical Report.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.693830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.693830Z digest=sha256:5e86fc1f03bb46978bdf347401de1dbe8c12e544c63a2cfd49836b14ec827bd2

Observation bee67213-7a99-47aa-8682-d01960789dbf · outbound

This paper cites A comprehensive survey on long context language modeling.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A comprehensive survey on long context language modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.774087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.774087Z digest=sha256:acc78e991b20f7503fd0ab6c73201076487e1823567b0a37ab21b2b993d80bc0

Observation 35645a35-7a41-449a-bae3-8555332e4569 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.813222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.813222Z digest=sha256:5b5e54ac1855e9f03e95c21e4e5ea1e786b0865b5332295690fbb4c97a94725f

Observation d4991f24-b25a-4657-a3fc-dc21deb99386 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.003415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.885862Z digest=sha256:ed5f851a300f36a474b06339b5b2d1c73df8cb2bd0c5758b59c9f7d4d935f38c

Observation c693e4ac-101f-4d0b-9120-f05c5ed32ef5 · outbound

This paper cites s1: Simple test-time scaling.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.954302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.954302Z digest=sha256:718df35324684c2fdbc2eec809a71adbe2dcc8bfc9b32abebdd68aba116433ac

Observation 9bfba96c-05b1-46ab-9419-19da37f515f1 · outbound

This paper cites Learning to reason with llms, September 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Learning to reason with llms, September 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.850468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.988945Z digest=sha256:e963fb41e45323b06f6089223997e0bf7967587908f4fe2a8a6faca1caa5f526

Observation 3ccb6168-1768-4eed-97fd-f2548393fea0 · outbound

This paper cites Introducing deep research, February 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing deep research, February 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.652281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.027224Z digest=sha256:6bbc865db351b3996793001d1a8f90698d18546e5124fe42051a8c7c8967bedb

Observation 9f07ac99-fddc-4d2a-8db4-0fa1168e53ee · outbound

This paper cites Openai o3-mini system card, January 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Openai o3-mini system card, January 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.469980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.084078Z digest=sha256:0c449c495915f0528698c7b277add9c13c8c83fc47438abdc5f96f3d7eb9cb19

Observation 3927f977-7b70-432c-8e69-c6e25b9015d2 · outbound

This paper cites Tinyzero: Clean, minimal, accessible reproduction of deepseek r1-zero, Janurary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tinyzero: Clean, minimal, accessible reproduction of deepseek r1-zero, Janurary 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.269767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.128326Z digest=sha256:3e3688ef5124659cc0abc74a955ba1deefe07fce709033d17a2ae1e15d819d31

Observation fd97e77d-61a8-4e78-9536-d4a81f9f505c · outbound

This paper cites In-context retrieval-augmented language models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning In-context retrieval-augmented language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.194579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.194579Z digest=sha256:0d65597e5c79f9a23f5371a7c1e56f82012a945d60f12d4d3290db67aaacfec5

Observation 8747f35d-13d6-4bb5-acc9-a7750adfb6cf · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.270125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.270125Z digest=sha256:58936601097edf511f7d015a1772fbd86654f22fdc53c1eaf65a24682dcd760e

Observation 86e6ef85-53f6-4a39-b12b-e47400314f6b · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.330384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.330384Z digest=sha256:dba979986e0af9991edc0d7acc80200fbc9cdee915b133cb9acb1a58ca8e2296

Observation c1cceeef-07a9-4641-b23c-6f4a9cfadceb · outbound

This paper cites Proximal Policy Optimization Algorithms.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.395201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.395201Z digest=sha256:b23edef978595a1d03b816e6c32e2d0ac4a552d0ace7b86e4718daa3dd47b64f

Observation 7dc88da3-2d69-4ee0-ba0b-79fc4edee777 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.433924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.433924Z digest=sha256:842d1ddf87f6e310542a27d8dbfcd0d86ed1d0d6f139c1fd8b044cef79af5806

Observation 6f5925a4-045a-4a55-8e01-e68cd626353a · outbound

This paper cites Defining and characterizing reward gaming.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Defining and characterizing reward gaming

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.480607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.480607Z digest=sha256:bedfd92a43719cf2d3eb7a8f09ccf2298c765727ce4a7421bef2d09d244f7c43

Observation 6e84b73d-ea36-4c11-ad42-f9bbcf2091de · outbound

This paper cites Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.108745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.488150Z digest=sha256:99287c9f4bb3c325c63cd19d4c302fdf095c75b48d039907acc0c49e373241dc

Observation 9818ed07-8b4e-4ff3-b356-dd225e445b1c · outbound

This paper cites Gemini 2.0 flash thinking, December 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Gemini 2.0 flash thinking, December 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.935285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.582817Z digest=sha256:b3959dc7ec239bfbdefedfc5c01a7e785bdd97dabc7c7e16b973837a85fa4e32

Observation 968e510e-3a68-492e-8913-8f473198698e · outbound

This paper cites Try deep research and our new experimental model in gemini, your ai assistant, December 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Try deep research and our new experimental model in gemini, your ai assistant, December 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.794465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.804958Z digest=sha256:62ab3fb8c34030dd84e0573ce19e380e6a16870dfdb7c293c6c19a44cbd605ef

Observation 18de4cb0-00e7-4f38-acc4-d36a3dd35d13 · outbound

This paper cites Unlocking the potential of reinforcement learning in improving reasoning models, Feburary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Unlocking the potential of reinforcement learning in improving reasoning models, Feburary 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.640372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.977349Z digest=sha256:0ff1abde7569df9e049c8fd06dbac8965d86a30ab05519ee0abe81697f571650

Observation 5af499fc-4208-4031-a3fc-e6890f0434cf · outbound

This paper cites Introducing perplexity deep research, February 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing perplexity deep research, February 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.519489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:57.194679Z digest=sha256:05d9924f466d8418614da7137f96fbfbd75984391aa3411d50be31b1fc85734c

Observation f0836f6c-480a-4a5d-84a6-8255c3886505 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.355705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.355705Z digest=sha256:8751acae110e0f281b54e0f00ccaef543b7284ca9b7cd30d6bf9abf366dca60a

Observation affe8d92-c044-42e4-b242-76e2a097a1c1 · outbound

This paper cites Qwen3: Think deeper, act faster, April 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen3: Think deeper, act faster, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.604942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.604942Z digest=sha256:9b0807a8a8bf35501cad4ab4ed5008e36716edb1a3be48996fabfbacabb3e587

Observation 6fe7b7b6-69f2-4bc4-9acd-954aedd08ca0 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.786971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.786971Z digest=sha256:b1d9face0d6ec1780ceaf7c6d882e8176a2c8a2fd69acc36e61b325333b4466c

Observation 4c15f17e-4a72-4d80-9006-0108a4394cd8 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Musique: Multihop questions via single-hop question composition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.926741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.926741Z digest=sha256:27bec024e361dd92a300ecfe151ca9c4d50926ee6167894adcd0ddb695e0fe31

Observation 9b3a869c-772d-4d03-b820-8926f46311ae · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.110674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.110674Z digest=sha256:dffc2a51e4f9cd188876f84059b273f9d03f91365456b5f470ffbc4a4cbc90a4

Observation 413f453b-55d9-49b4-9e26-7dcb2f9329b8 · outbound

This paper cites A Comparative Study on Reasoning Patterns of OpenAI's o1 Model.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A Comparative Study on Reasoning Patterns of OpenAI's o1 Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.227150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.227150Z digest=sha256:32ee19bb648ee633a650467b1749e1218a4a175df8ca3f89167e75d9e8d6b663

Observation ff648891-44b4-41c7-9ca2-2d54eee6e314 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.340088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.340088Z digest=sha256:3df2290f34283b22a1319c530dcc0fa9d1f3630c26ee6dd80f5b37063db7bf67

Observation 3f3f1ab7-7205-4d9f-b143-57ad1e783f38 · outbound

This paper cites Effective long-context scaling of foundation models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Effective long-context scaling of foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.336932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:58.419410Z digest=sha256:64aa5e7064934f4242ca5f1516d3f173773edcfd6b7e2d47a094e1812a326386

Observation 4191795d-4f7c-4b52-9837-1803d4e0eff9 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.564817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.564817Z digest=sha256:3c0db7d1676d034d8d1bead367027fc0f8754612e21127c149b803f84dff5e32

Observation 89ed0a4a-dba0-441f-9389-97d418ad9cd5 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.698401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.698401Z digest=sha256:e0babbfeb07f147cacabbd936f2cbee821c4d77dd19304416feb919a18382bbf

Observation b11d0c1e-9563-403e-b370-905692e76ce2 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.837526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.837526Z digest=sha256:0b339758ee604f51e151aaebf276ad5823658aa1877b8ba90fe0170df2aecee6

Observation d6873614-d83d-4412-896a-aac6382cb4e6 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.029063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.029063Z digest=sha256:13f25897526f903d22c292f1a0c598259d464e1b63ee8284b692a02e51251e22

Observation 46851b85-18b4-449a-a820-f04a16a9e52e · outbound

This paper cites LIMO: Less is More for Reasoning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LIMO: Less is More for Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.250891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.250891Z digest=sha256:ed2f8e339514413d05074dafe0aabdb1a54a07dd76655f502c44265df36afc08

Observation fd620298-cab5-4451-a2aa-84df904a8cb4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.379000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.379000Z digest=sha256:4ae4f8efb283a22c1c361537dd06508c4c045fbbe5c987a533527b0124fbac4a

Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.582127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.582127Z digest=sha256:50671477fb7818cf980b31bfdc0f3c0ec7a7c6e3b2da7be88d1d10b415b8e7dd

Observation 9ae0d5af-1f74-45e2-a223-e20895d4876f · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.741379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.741379Z digest=sha256:4d951fc577035d0022961435b31b9b33f6017f04ccbacb5f2278de81c771d4ac

Observation 6bfde056-f405-44df-87e7-d97d9a19bdc3 · outbound

This paper cites Docmath-eval: Evaluating math reasoning capabilities of llms in understanding long and specialized documents.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Docmath-eval: Evaluating math reasoning capabilities of llms in understanding long and specialized documents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.159664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:59.889395Z digest=sha256:027ae451e8227ea427d8f70272f09cf2d68d708d58e5d8b94d6eddc05f1f52c1

Observation f1041c25-8c63-4478-a8a1-19c0df1792a5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:00.020817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:00.020817Z digest=sha256:85e19de9cc5ffead27fad7247a02f49d557bfe8645f709fc87d66341b09b7562

Observation 42577c6c-3348-41c1-818b-8500d531da8f · outbound

This paper cites Lima: Less is more for alignment.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Lima: Less is more for alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.030781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:48:00.129155Z digest=sha256:b0826af11fcb0931edb134b55fd45aa8c31ff78b21bbb529be56e0d2cd0e0b48

Pith citing papers

Observation 3e123b75-b515-4470-9423-af84373323fd · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.509718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:117e45b8743708f9889249e9f88788c24e0050f15eee0880db3ebffa04f46f5c

Observation a10e8dbd-42af-4c9b-a662-fe5340c17dce · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:13.872622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:13.872622Z digest=sha256:754c07476091a61a9c683150ddb19414a72505b8a41c563e9ca7a50bc375e680

Observation e837cc81-3032-4a7f-a807-9b9901e3ce50 · inbound

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration cites this paper.

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:55:15.281023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:55:15.281023Z digest=sha256:3527f629d6516477557fee8fe18f331b41a119641302b5ba5eadee10695b978c

Observation cf7f53f6-838c-425f-82d3-2ef30ef5c63e · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:16.164641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:16.164641Z digest=sha256:f7c16dc9b5c8933ac65d51beead2cdd478fa70f9beebd4422afbdfdd433c46bf

Observation 8b38a9a7-7814-4433-8295-e3566f2c684b · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.524006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:76e818badd801a79c2dc1d6999425fdd15be92fa6996170f9196a1fc2f97f52e

Observation aa3c5603-7403-45e7-86eb-b38a7ab0e77a · inbound

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety cites this paper.

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:05.968531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:05.968531Z digest=sha256:659b7ca69f30bca063a8d46fd2d29b643aa3b8b3c4b9d6b9d189132bb9ccda50

Observation 2cbaf1f8-7acd-4ada-b28d-537dea77f30a · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:53:28.051553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:4f43fa1e814f8804f37794074e3ad455b773b8786d0b972d9a350f22c1f04b47

Observation 7416320e-45b1-46f1-b594-5636eb518597 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:cd643dfdee904142f0cea1e4173185c89765fe465ff80f79cee97c5ff04a97f8

Observation 12109965-067d-4e1f-af18-3eafcfa9dc17 · inbound

A Decomposition Perspective to Long-context Reasoning for LLMs cites this paper.

A Decomposition Perspective to Long-context Reasoning for LLMs QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:31:00.039292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:05:34.666937Z digest=sha256:58071068e7eaf63d74508e57b57ebdbe3b0d492abd667f28cd4d7f00eb1509ed

Observation 36cfaa94-33ae-4009-9417-ef4af0b4d499 · inbound

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning cites this paper.

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.482511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T11:17:43.769244Z digest=sha256:78cf1f59aefdff3f0b4ac772da950a7e66e7594a12bbe7b2ebe4e09d8255e2a7

Observation a4c46b9d-afc2-44c0-8e74-c1e1e34d5cff · inbound

OPSDL: On-Policy Self-Distillation for Long-Context Language Models cites this paper.

OPSDL: On-Policy Self-Distillation for Long-Context Language Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:11:20.402410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T06:07:36.830550Z digest=sha256:8cbedbcc2cb2ca45ad375e1bbb1fac92ef9d29898943aacb3ec02bb5b167663a

Observation 6e5e2916-47b9-4935-9b3a-d04314fb4766 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.683022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:df91c9bf86baa9647bafbd220e3ab0e5baa166cda6edf99a47b64fe75f7b31e7

Observation e6d73681-e8af-4b44-bf0c-fa94aad81275 · inbound

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation cites this paper.

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:32:19.432348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T05:29:00.576006Z digest=sha256:e402a33cf4c16e07e07416288bf7b6ebb9ad620503ec64da3061cc75bc03913f

Observation c837ddf0-367f-4f48-9c06-6fa5362361eb · inbound

Evidence-State Rewards for Long-Context Reasoning cites this paper.

Evidence-State Rewards for Long-Context Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:28:18.210189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T13:25:14.844589Z digest=sha256:0ba4d93c48ed4ecddd5293e808d419fce1daf2b73e5cf6b3f10b020ab63a4fc4

Observation 4e9f7c26-e957-4b75-b019-4d3c7c6d1dbf · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:03.589144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:03.589144Z digest=sha256:829aa6e4cfd66f46f16f2563ac9015a0f95b5d79acce2c8c7e9324865109fb06

Observation 9f3e8f69-c2da-48ed-90f6-f9289cc9115d · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.449669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.449669Z digest=sha256:7b35a7cfe5d0821866da6a0db7577840c708496d96efb738f44edf008e6746d8

Observation ed144155-1032-4489-b5af-d358ca4d980e · inbound

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning cites this paper.

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T09:18:46.780870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:18:46.780870Z digest=sha256:8a6826d52a2b365fa81d6d080f59f1ffd8e0850145e360e2103d276fe0d0ccb8