Pith. sign in

Paper Citation Record · LEDGER

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 17 inbound Pith citation observations for arXiv:2505.17667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17667 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:48:00.129155Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:13.872622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.208570Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac731f42-ec22-4d6c-9ee9-351aa5787b09 · outbound

This paper cites Claude 3.7 sonnet system card, Feburary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Claude 3.7 sonnet system card, Feburary 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:04.276318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.108255Z digest=sha256:5dccca6b6af86db7a37476e3ec4e6261f47563c8e7cce344fdb3b0d664cd6fc9

Observation 428b97af-075e-4a64-b2da-e51e870947cb · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LongBench: A bilingual, multitask benchmark for long context understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.174129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.174129Z digest=sha256:8c2c2813fabeb61d8242ec4cf1a159569b6e7a263d6a2453ccb3e033cbc88f24

Observation dccea908-1f63-4ee0-a623-c2e2f93b4de8 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.301320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.301320Z digest=sha256:d3d53cf6122a635d3a61a6c0b7029cd133ab2db442180d05ffc75b79f295509b

Observation 5bcfd899-cebd-4adf-8532-464381b043e8 · outbound

This paper cites Thinking, fast and slow.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Thinking, fast and slow

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.367277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.367277Z digest=sha256:72a5cbcb8f4789b0c2b3857ff2320a732138a6b72cf4ec461b72a9fac61e814a

Observation 2c80ca3c-6bcd-4ddb-804f-82c9bc6d6d7d · outbound

This paper cites A dataset of information-seeking questions and answers anchored in research papers.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A dataset of information-seeking questions and answers anchored in research papers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:04.052534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.465404Z digest=sha256:5a220c9065c52879bb55377e09d05bd7c0c83ff0f54ef722495d5427dd18df05

Observation 0ffeb1d1-5335-4866-8c4a-d046b8ec48f1 · outbound

This paper cites Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power!, November 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power!, November 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.542819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.542819Z digest=sha256:5e1496ec5139b6426bbbfe28e200131c6b912f67f72e1a3564c60eccbebc1f75

Observation 76e19ce2-d0ac-4641-90bf-8a32e2b248e1 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.628265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.628265Z digest=sha256:7c5e6a7cc9e2e1d1b3920e70e1f44fb83688b3e5cdf2d433794271d68c8c3d72

Observation 0555fc97-8e07-4f1f-a9d8-d1aadd3e6470 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.849948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.701599Z digest=sha256:3af020efd3f00f4f3ee8a5eff9cb64f6a4b6f6ecde77bff303b347e3cd923733

Observation c4e7b238-38c0-4755-b60d-e34b0766870d · outbound

This paper cites Data engineering for scaling language models to 128k context.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Data engineering for scaling language models to 128k context

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.614383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:54.798913Z digest=sha256:f057b14f9a11dc7cc2336bd41b3db061f5de40ead2c9c7fe2f1750a108790aba

Observation 50bb3bbf-3af0-44b5-b063-9be49de896b1 · outbound

This paper cites How to train long-context language models (effectively).

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning How to train long-context language models (effectively)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.906238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.906238Z digest=sha256:309aefe722fe44df2b5a122686d6e623c9302f7ab411b2f4d30e87430828ba61

Observation c031e961-f33c-4414-901d-bee4467a7c34 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:54.988796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:54.988796Z digest=sha256:67cc83b762243e2bc4732176029a67b14118c3f53e80d5a383037fa6adf8052d

Observation bbe9bc61-4b9f-4c90-9c46-294a5ab5a65e · outbound

This paper cites Retrieval augmented language model pre-training.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Retrieval augmented language model pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.096120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.096120Z digest=sha256:7e4057a11556c5cf70e2f3ed26aa97f6fb1748a0c04a2fdf5e56003f27d9f38c

Observation efa39379-c508-4126-a5be-ad8c50dec571 · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.175370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.175370Z digest=sha256:46525db034ff20226ce4985e4e31a2b51516d5fd400d012a408e5fbdd7e6d75c

Observation eea31d13-f54a-4447-b04b-bdea4d8d786e · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model, 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.241433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.241433Z digest=sha256:14229da8b080af39f96486d9a09e6e5f744f5ab2b663f1a15dca025cf3d2e395

Observation 921fe487-7fff-4a7b-9e4d-dd87a0991bb0 · outbound

This paper cites OpenAI o1 System Card.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.301210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.301210Z digest=sha256:d70ffbc7cdec57a44711035f86433510dd4d96a488b0f7d65824c9d3b7b6d782

Observation ad90fe84-00f5-444b-b5b8-99ee3c97fe0c · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.375344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.375344Z digest=sha256:7f98e5408bdc17e3c2968743195dede7be91bb3dcd085083342e664526f06662

Observation 1947ad9c-5449-4368-a48d-3c3da885c057 · outbound

This paper cites The narrativeqa reading comprehension challenge.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The narrativeqa reading comprehension challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.450413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.450413Z digest=sha256:59788ffaf80961e870f611d87748fe0383459687ecbc0582c5e18cc5157aeb11

Observation ba0e791c-8210-4b33-b1f6-9fe334c8d8d3 · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.502010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.502010Z digest=sha256:271ccf8c9a94811c6ff026e4482a22b452d55596fda88252a09ebd947eb2217d

Observation 0c3a635d-f495-4a84-93f9-e0578ea83ab5 · outbound

This paper cites From quantity to quality: Boosting llm performance with self- guided data selection for instruction tuning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning From quantity to quality: Boosting llm performance with self- guided data selection for instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.394518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.546278Z digest=sha256:4a2fdd81ddac558b2e73a0f0f725e98b46661e7a96b969d5ab988dc5092bf638

Observation 8a195b50-4509-4525-9d5e-244ecf2c38d2 · outbound

This paper cites The unlocking spell on base llms: Rethinking alignment via in-context learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning The unlocking spell on base llms: Rethinking alignment via in-context learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.197931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.637448Z digest=sha256:27de79d23d7f7b4121127b34ccb33212acce93ecb873ff5275d2ec17671f0bc1

Observation d3c2ef07-9e66-4961-9b0b-8b5bf28a5099 · outbound

This paper cites DeepSeek-V3 Technical Report.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.693830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.693830Z digest=sha256:836b095123dc375c3dc68f7381d3a5b9898d873cf2d0fed8bcc84b6a369dd08c

Observation bee67213-7a99-47aa-8682-d01960789dbf · outbound

This paper cites A comprehensive survey on long context language modeling.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A comprehensive survey on long context language modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.774087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.774087Z digest=sha256:61e81b62b234118cee2c520aa685fef7cce5ec7992525aee72468e23bddd584b

Observation 35645a35-7a41-449a-bae3-8555332e4569 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.813222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.813222Z digest=sha256:ff66d973e2b20bac4df276aee5f479354659abd968d1b101623eb92cdc74ccf9

Observation d4991f24-b25a-4657-a3fc-dc21deb99386 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:03.003415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.885862Z digest=sha256:ef72909e6da71cbf906bbf1153765b0171164f70e9cad19e01aee9b029631f5f

Observation c693e4ac-101f-4d0b-9120-f05c5ed32ef5 · outbound

This paper cites s1: Simple test-time scaling.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.954302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.954302Z digest=sha256:ce8ef25702cc77246a4a9424c310a7a7e4bc09ca6f4ec9668a65bcb460fdc73f

Observation 9bfba96c-05b1-46ab-9419-19da37f515f1 · outbound

This paper cites Learning to reason with llms, September 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Learning to reason with llms, September 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.850468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:55.988945Z digest=sha256:f35358286f40803d653517dd185db4bc64c51755c6b2c21974bb39792c20af15

Observation 3ccb6168-1768-4eed-97fd-f2548393fea0 · outbound

This paper cites Introducing deep research, February 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing deep research, February 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.652281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.027224Z digest=sha256:d65c1ed34767340c50085910c1fdc9148773f646a720fb11395678ebe9c9c600

Observation 9f07ac99-fddc-4d2a-8db4-0fa1168e53ee · outbound

This paper cites Openai o3-mini system card, January 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Openai o3-mini system card, January 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.469980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.084078Z digest=sha256:d2bad23113eb72f123e96e13dca2f70438c7ce4a5e5e2dd9b5e96b5a0f46074e

Observation 3927f977-7b70-432c-8e69-c6e25b9015d2 · outbound

This paper cites Tinyzero: Clean, minimal, accessible reproduction of deepseek r1-zero, Janurary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Tinyzero: Clean, minimal, accessible reproduction of deepseek r1-zero, Janurary 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.269767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.128326Z digest=sha256:f496a42f58459208234fdc77c638349dd82d9bcaaaf3e0c29b8efa5b89a0ba2e

Observation fd97e77d-61a8-4e78-9536-d4a81f9f505c · outbound

This paper cites In-context retrieval-augmented language models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning In-context retrieval-augmented language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.194579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.194579Z digest=sha256:4f38bf037a116fb796b95fe625fdf79f33edef474783b9e08704bd0a8c5ed4d8

Observation 8747f35d-13d6-4bb5-acc9-a7750adfb6cf · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.270125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.270125Z digest=sha256:f1a20702a8634574f0e3fe01d7641e143cef1338bd88712ab8b2379e6c23a972

Observation 86e6ef85-53f6-4a39-b12b-e47400314f6b · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.330384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.330384Z digest=sha256:c35828824b9da79b19479ea7714d314f7fa90459eb64638e42e9440baa7f597a

Observation c1cceeef-07a9-4641-b23c-6f4a9cfadceb · outbound

This paper cites Proximal Policy Optimization Algorithms.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.395201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.395201Z digest=sha256:6e5f8dff86756445e37827aebca571cd103dd6ba8b862acd9902320a9e78a6e7

Observation 7dc88da3-2d69-4ee0-ba0b-79fc4edee777 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.433924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.433924Z digest=sha256:ebcdae8daa211dcf6edd51d8a470dfa4e93494f2f6507b8616738eb31000db8e

Observation 6f5925a4-045a-4a55-8e01-e68cd626353a · outbound

This paper cites Defining and characterizing reward gaming.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Defining and characterizing reward gaming

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.480607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.480607Z digest=sha256:e1f567bde4d28ccf25eafbd4c80157bd46c8b43bfa0cceb7be9ea2b2fa57fd73

Observation 6e84b73d-ea36-4c11-ad42-f9bbcf2091de · outbound

This paper cites Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:02.108745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.488150Z digest=sha256:23bf8d77f067e64691340a5d38a2a4f0fbce5d192a8ca6791e6058e6e297bd11

Observation 9818ed07-8b4e-4ff3-b356-dd225e445b1c · outbound

This paper cites Gemini 2.0 flash thinking, December 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Gemini 2.0 flash thinking, December 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.935285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.582817Z digest=sha256:0819cd79f800bb5f9ceba5e3fd10b4f10fde7fb5811e4002af23b9ef19da8861

Observation 968e510e-3a68-492e-8913-8f473198698e · outbound

This paper cites Try deep research and our new experimental model in gemini, your ai assistant, December 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Try deep research and our new experimental model in gemini, your ai assistant, December 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.794465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.804958Z digest=sha256:6943e60af30550f695d5597016865b3e907ab5885edfb0186396f743f9dcbf20

Observation 18de4cb0-00e7-4f38-acc4-d36a3dd35d13 · outbound

This paper cites Unlocking the potential of reinforcement learning in improving reasoning models, Feburary 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Unlocking the potential of reinforcement learning in improving reasoning models, Feburary 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.640372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:56.977349Z digest=sha256:0084230c8d9befa882c351fed1dbf46210e1e6bc7554270591ab2a2dd34b954e

Observation 5af499fc-4208-4031-a3fc-e6890f0434cf · outbound

This paper cites Introducing perplexity deep research, February 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Introducing perplexity deep research, February 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.519489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:57.194679Z digest=sha256:e17d5de25745d25b9d61be694ae507fc0a312f0fa4a8581cd4b37058197947a4

Observation f0836f6c-480a-4a5d-84a6-8255c3886505 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.355705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.355705Z digest=sha256:bfa9b2345798e11c5d201eb2713d52c8d2a6f85c147bd16279ca8d1221d80e63

Observation affe8d92-c044-42e4-b242-76e2a097a1c1 · outbound

This paper cites Qwen3: Think deeper, act faster, April 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen3: Think deeper, act faster, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.604942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.604942Z digest=sha256:ec8d2bfb577f606620ac5d4f3339770e6dd26a10cc61606ce954b5b5c2a01271

Observation 6fe7b7b6-69f2-4bc4-9acd-954aedd08ca0 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.786971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.786971Z digest=sha256:deae36abe2dca673c424d7fb45b788ba629a7fb0093a72817805a7e8ae5f923c

Observation 4c15f17e-4a72-4d80-9006-0108a4394cd8 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Musique: Multihop questions via single-hop question composition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:57.926741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:57.926741Z digest=sha256:2204f11703d96724d160c5c5cccd98ea9bfa24e32113bdb4824ee91b200054d5

Observation 9b3a869c-772d-4d03-b820-8926f46311ae · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.110674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.110674Z digest=sha256:125af1517fe8974f9e7eb4f530f21438f5535ebe796c988a04387371d509eed9

Observation 413f453b-55d9-49b4-9e26-7dcb2f9329b8 · outbound

This paper cites A Comparative Study on Reasoning Patterns of OpenAI's o1 Model.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning A Comparative Study on Reasoning Patterns of OpenAI's o1 Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.227150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.227150Z digest=sha256:ba09f301e99d1b44235fdf69e3c244ef76b3935c8270ebc51ab5a4975cb3fb9e

Observation ff648891-44b4-41c7-9ca2-2d54eee6e314 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.340088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.340088Z digest=sha256:eae1dc431579562a3065c59068b2f1c80cf2798721cfee7ceb37f762266153a5

Observation 3f3f1ab7-7205-4d9f-b143-57ad1e783f38 · outbound

This paper cites Effective long-context scaling of foundation models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Effective long-context scaling of foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.336932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:58.419410Z digest=sha256:6ee8e56ea57ed6976352ed09d2098be8720465248b30496a19e8dbd2b03bc0e9

Observation 4191795d-4f7c-4b52-9837-1803d4e0eff9 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.564817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.564817Z digest=sha256:9569d064df5a9c83654e4162922e5bd387dd5fd985b43edbf89fab89f798ded9

Observation 89ed0a4a-dba0-441f-9389-97d418ad9cd5 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.698401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.698401Z digest=sha256:52db4e0835c7b6c39f00f08a8a0c9cffd1cbaaf4b9783f96d70877f334f66826

Observation b11d0c1e-9563-403e-b370-905692e76ce2 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:58.837526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:58.837526Z digest=sha256:d2a03aa7cb31d3d7773fa2cbf4fc3ce1d4ecd2851d7a5f1bd0745f23198fc045

Observation d6873614-d83d-4412-896a-aac6382cb4e6 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.029063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.029063Z digest=sha256:f982adc17048f513a1a38175131f1aab492b17880273a4b84f1d64ab90865a89

Observation 46851b85-18b4-449a-a820-f04a16a9e52e · outbound

This paper cites LIMO: Less is More for Reasoning.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning LIMO: Less is More for Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.250891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.250891Z digest=sha256:cc821d143f99a234f1f305fabdeb45e87c5024a3d9f571bb36cb54692bbcdcbc

Observation fd620298-cab5-4451-a2aa-84df904a8cb4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.379000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.379000Z digest=sha256:55b8c14db204808afcaf45fa89b93a1996fbfbc45d4e7cc4f4f2b20df79dee71

Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.582127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.582127Z digest=sha256:9390131d774854ef076cecf09571a82a77ba545976d759bc671bab06c721bdb4

Observation 9ae0d5af-1f74-45e2-a223-e20895d4876f · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.741379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.741379Z digest=sha256:0cbc570b33f24ce2310cb81134f61255916cb64408110f8ef70a383a0d932aaa

Observation 6bfde056-f405-44df-87e7-d97d9a19bdc3 · outbound

This paper cites Docmath-eval: Evaluating math reasoning capabilities of llms in understanding long and specialized documents.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Docmath-eval: Evaluating math reasoning capabilities of llms in understanding long and specialized documents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.159664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:47:59.889395Z digest=sha256:e45cdde97a01eb0395a5982fbcbd6245045be945183a828c3450166a09b65035

Observation f1041c25-8c63-4478-a8a1-19c0df1792a5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:00.020817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:00.020817Z digest=sha256:850bb09d470fd802943b0b61da0a2ad7c58df127b20bb6093b0875b8b93b2458

Observation 42577c6c-3348-41c1-818b-8500d531da8f · outbound

This paper cites Lima: Less is more for alignment.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning Lima: Less is more for alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:48:01.030781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:48:00.129155Z digest=sha256:faf2b5882a39e4b01eb5df18a72236fe82b3f89d5bd582b0220ed73817f0f9a8

Pith citing papers

Observation 3e123b75-b515-4470-9423-af84373323fd · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.509718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:c5e44fbcce65362c196cdf0280e3e8bc8111a5c1ebdcdfdcf9fb9008d360f52e

Observation a10e8dbd-42af-4c9b-a662-fe5340c17dce · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:13.872622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:13.872622Z digest=sha256:60c5437034069adf52bf2773c5b0fad62d12ca8bcb7c4cc7e01ffa274b13189a

Observation e837cc81-3032-4a7f-a807-9b9901e3ce50 · inbound

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration cites this paper.

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:55:15.281023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:55:15.281023Z digest=sha256:0bf535d393c5245f5605ff2d1a5b0258df2959023f7622283d677937dca36419

Observation cf7f53f6-838c-425f-82d3-2ef30ef5c63e · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:16.164641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:16.164641Z digest=sha256:6b538e8b464c03f3659d7f6904d83ce563c2004e5b50808dcd87b468e0f31c8c

Observation 8b38a9a7-7814-4433-8295-e3566f2c684b · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.524006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:0d12b8e9cdeec6a8f5d3b7c91fc58ac9fead83f33c73cc1e211ff62ad2e8556e

Observation aa3c5603-7403-45e7-86eb-b38a7ab0e77a · inbound

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety cites this paper.

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:05.968531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:05.968531Z digest=sha256:22d2b7cd48d476fef95fee8748eb71c19398814677321730cbd2f4ea4f22214a

Observation 2cbaf1f8-7acd-4ada-b28d-537dea77f30a · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:53:28.051553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:b3aaf020f24caea15712cc9be6013f14c7e668c0b6dd897c25ddb30006a33c4f

Observation 7416320e-45b1-46f1-b594-5636eb518597 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:b11d3568f1fcc46a92196fc05198abd72819f556b936f70a6923879e7202f727

Observation 12109965-067d-4e1f-af18-3eafcfa9dc17 · inbound

A Decomposition Perspective to Long-context Reasoning for LLMs cites this paper.

A Decomposition Perspective to Long-context Reasoning for LLMs QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:31:00.039292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:05:34.666937Z digest=sha256:e342034575c328d4581662f67f148c45b8732e60c3e207f60c1cf6671fc52126

Observation 36cfaa94-33ae-4009-9417-ef4af0b4d499 · inbound

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning cites this paper.

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.482511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T11:17:43.769244Z digest=sha256:699f51f45c2757610a621466450ea940dc7d98aae9afe61de219f8526f5251f2

Observation a4c46b9d-afc2-44c0-8e74-c1e1e34d5cff · inbound

OPSDL: On-Policy Self-Distillation for Long-Context Language Models cites this paper.

OPSDL: On-Policy Self-Distillation for Long-Context Language Models QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:11:20.402410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T06:07:36.830550Z digest=sha256:8c3fe91572888a2163edf42358374eaa576bb4985d07a6f3cc8016af5471d697

Observation 6e5e2916-47b9-4935-9b3a-d04314fb4766 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.683022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:334674fc145579e6aeaeeaabc5207d512bd6d02308a236ea27c8b8f81648fd33

Observation e6d73681-e8af-4b44-bf0c-fa94aad81275 · inbound

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation cites this paper.

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:32:19.432348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T05:29:00.576006Z digest=sha256:7d882048a62c3a5ccdcc9936ad3c872b2b35ef67d0341af1913d03920fcacee3

Observation c837ddf0-367f-4f48-9c06-6fa5362361eb · inbound

Evidence-State Rewards for Long-Context Reasoning cites this paper.

Evidence-State Rewards for Long-Context Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:28:18.210189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T13:25:14.844589Z digest=sha256:de7daafce01c8592b272c2bbda36bf6cb0aa8e05abc87c59a21d3191b1557789

Observation 4e9f7c26-e957-4b75-b019-4d3c7c6d1dbf · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:03.589144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:03.589144Z digest=sha256:030eab8146dfd61cfcfae86532c054a15f52a6338384e908325c151406b01202

Observation 9f3e8f69-c2da-48ed-90f6-f9289cc9115d · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.449669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.449669Z digest=sha256:ba628ec056010dc96d1c28ef2157634fa9b62c4b3bdb58fce6a1a2ef578db80b

Observation ed144155-1032-4489-b5af-d358ca4d980e · inbound

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning cites this paper.

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T09:18:46.780870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:18:46.780870Z digest=sha256:2d569a62a2c1f95b2003460589cfcf3085ab9817ce89134d905bcb8bb5f95fcc