Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:48:00.129155Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 17 inbound Pith citation observations for arXiv:2505.17667.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:48:00.129155Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:13.872622Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T13:28:18.208570Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ac731f42-ec22-4d6c-9ee9-351aa5787b09 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Claude 3.7 sonnet system card, Feburary 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 428b97af-075e-4a64-b2da-e51e870947cb · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LongBench: A bilingual, multitask benchmark for long context understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dccea908-1f63-4ee0-a623-c2e2f93b4de8 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcfd899-cebd-4adf-8532-464381b043e8 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Thinking, fast and slow
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c80ca3c-6bcd-4ddb-804f-82c9bc6d6d7d · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A dataset of information-seeking questions and answers anchored in research papers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ffeb1d1-5335-4866-8c4a-d046b8ec48f1 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power!, November 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e19ce2-d0ac-4641-90bf-8a32e2b248e1 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Competitive Programming with Large Reasoning Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0555fc97-8e07-4f1f-a9d8-d1aadd3e6470 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4e7b238-38c0-4755-b60d-e34b0766870d · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Data engineering for scaling language models to 128k context
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 50bb3bbf-3af0-44b5-b063-9be49de896b1 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning How to train long-context language models (effectively)
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c031e961-f33c-4414-901d-bee4467a7c34 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe9bc61-4b9f-4c90-9c46-294a5ab5a65e · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Retrieval augmented language model pre-training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa39379-c508-4126-a5be-ad8c50dec571 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea31d13-f54a-4447-b04b-bdea4d8d786e · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921fe487-7fff-4a7b-9e4d-dd87a0991bb0 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning OpenAI o1 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad90fe84-00f5-444b-b5b8-99ee3c97fe0c · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1947ad9c-5449-4368-a48d-3c3da885c057 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The narrativeqa reading comprehension challenge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0e791c-8210-4b33-b1f6-9fe334c8d8d3 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3a635d-f495-4a84-93f9-e0578ea83ab5 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning From quantity to quality: Boosting llm performance with self- guided data selection for instruction tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a195b50-4509-4525-9d5e-244ecf2c38d2 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The unlocking spell on base llms: Rethinking alignment via in-context learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d3c2ef07-9e66-4961-9b0b-8b5bf28a5099 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-V3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee67213-7a99-47aa-8682-d01960789dbf · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A comprehensive survey on long context language modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35645a35-7a41-449a-bae3-8555332e4569 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4991f24-b25a-4657-a3fc-dc21deb99386 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c693e4ac-101f-4d0b-9120-f05c5ed32ef5 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning s1: Simple test-time scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfba96c-05b1-46ab-9419-19da37f515f1 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Learning to reason with llms, September 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ccb6168-1768-4eed-97fd-f2548393fea0 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing deep research, February 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f07ac99-fddc-4d2a-8db4-0fa1168e53ee · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Openai o3-mini system card, January 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3927f977-7b70-432c-8e69-c6e25b9015d2 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tinyzero: Clean, minimal, accessible reproduction of deepseek r1-zero, Janurary 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd97e77d-61a8-4e78-9536-d4a81f9f505c · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning In-context retrieval-augmented language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8747f35d-13d6-4bb5-acc9-a7750adfb6cf · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e6ef85-53f6-4a39-b12b-e47400314f6b · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cceeef-07a9-4641-b23c-6f4a9cfadceb · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc88da3-2d69-4ee0-ba0b-79fc4edee777 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5925a4-045a-4a55-8e01-e68cd626353a · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Defining and characterizing reward gaming
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e84b73d-ea36-4c11-ad42-f9bbcf2091de · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9818ed07-8b4e-4ff3-b356-dd225e445b1c · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Gemini 2.0 flash thinking, December 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 968e510e-3a68-492e-8913-8f473198698e · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Try deep research and our new experimental model in gemini, your ai assistant, December 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 18de4cb0-00e7-4f38-acc4-d36a3dd35d13 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Unlocking the potential of reinforcement learning in improving reasoning models, Feburary 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5af499fc-4208-4031-a3fc-e6890f0434cf · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing perplexity deep research, February 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f0836f6c-480a-4a5d-84a6-8255c3886505 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affe8d92-c044-42e4-b242-76e2a097a1c1 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen3: Think deeper, act faster, April 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe7b7b6-69f2-4bc4-9acd-954aedd08ca0 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c15f17e-4a72-4d80-9006-0108a4394cd8 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Musique: Multihop questions via single-hop question composition
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3a869c-772d-4d03-b820-8926f46311ae · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 413f453b-55d9-49b4-9e26-7dcb2f9329b8 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff648891-44b4-41c7-9ca2-2d54eee6e314 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f3f1ab7-7205-4d9f-b143-57ad1e783f38 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Effective long-context scaling of foundation models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4191795d-4f7c-4b52-9837-1803d4e0eff9 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ed0a4a-dba0-441f-9389-97d418ad9cd5 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11d0c1e-9563-403e-b370-905692e76ce2 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6873614-d83d-4412-896a-aac6382cb4e6 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning React: Synergizing reasoning and acting in language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46851b85-18b4-449a-a820-f04a16a9e52e · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LIMO: Less is More for Reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd620298-cab5-4451-a2aa-84df904a8cb4 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae0d5af-1f74-45e2-a223-e20895d4876f · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfde056-f405-44df-87e7-d97d9a19bdc3 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Docmath-eval: Evaluating math reasoning capabilities of llms in understanding long and specialized documents
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1041c25-8c63-4478-a8a1-19c0df1792a5 · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42577c6c-3348-41c1-818b-8500d531da8f · outbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Lima: Less is more for alignment
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3e123b75-b515-4470-9423-af84373323fd · inbound
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a10e8dbd-42af-4c9b-a662-fe5340c17dce · inbound
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e837cc81-3032-4a7f-a807-9b9901e3ce50 · inbound
MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7f53f6-838c-425f-82d3-2ef30ef5c63e · inbound
Observation of momentum dependent charge density wave gap in EuTe4 QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b38a9a7-7814-4433-8295-e3566f2c684b · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa3c5603-7403-45e7-86eb-b38a7ab0e77a · inbound
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbaf1f8-7acd-4ada-b28d-537dea77f30a · inbound
Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7416320e-45b1-46f1-b594-5636eb518597 · inbound
Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12109965-067d-4e1f-af18-3eafcfa9dc17 · inbound
A Decomposition Perspective to Long-context Reasoning for LLMs QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 36cfaa94-33ae-4009-9417-ef4af0b4d499 · inbound
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a4c46b9d-afc2-44c0-8e74-c1e1e34d5cff · inbound
OPSDL: On-Policy Self-Distillation for Long-Context Language Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e5e2916-47b9-4935-9b3a-d04314fb4766 · inbound
StoryAlign: Evaluating and Training Reward Models for Story Generation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6d73681-e8af-4b44-bf0c-fa94aad81275 · inbound
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c837ddf0-367f-4f48-9c06-6fa5362361eb · inbound
Evidence-State Rewards for Long-Context Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e9f7c26-e957-4b75-b019-4d3c7c6d1dbf · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3e8f69-c2da-48ed-90f6-f9289cc9115d · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed144155-1032-4489-b5af-d358ca4d980e · inbound
REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.