Pith. sign in

Paper Citation Record · LEDGER

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 17 inbound Pith citation observations for arXiv:2509.07980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07980 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:25.998787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:30.028956Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbdb3f4-4df4-4b74-b593-063bbb42ed12 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.417972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.417972Z digest=sha256:ed9a22bf57cb92de7800403d77d7bff9452f25cdc3fe9b79ebc7abd7e5e79a4a

Observation ed11c2c4-520a-4729-99d2-62093fe0dd0c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.842982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.842982Z digest=sha256:95b86f183cafe59e9951c984ec2635230412480f27cdb1869259f20641c7785c

Observation b198ef7d-3422-4d4b-8989-f7f26f6cbc2a · outbound

This paper cites Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:28:53.391615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:28:49.059722Z digest=sha256:7b3b06847f2e570a37143a0e16566c6be9f75cd4204784e3649f84a510cd8f5f

Observation d60036a1-5e84-4f5c-a3d6-b9a9c62b5162 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Efficient Test-Time Scaling via Self-Calibration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.143122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.143122Z digest=sha256:2c9b509097056f9b2948fd6ba1d08455d28be1f794c236627251ee5badd04b66

Observation af4ec024-5037-42b1-81e7-713b7e716b3d · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.329292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.329292Z digest=sha256:339e8feaeee63589e4d8a12d614dcc9bbdd73b2cd195f3c1c81a3734f20b8528

Observation 9547c6ec-a53a-4d31-bff0-449c73ccfa24 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.390681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.390681Z digest=sha256:7a39725cda967b44cd1acb07a7e72198901a44a66a14bd0a004b6afc0d89e142

Observation 80287147-8dee-4a21-9caa-501098559db8 · outbound

This paper cites Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:28:53.693150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:28:49.445211Z digest=sha256:302d3c43dbaf61057db2852fa094bdb93e46fb7e00a15bb771dd89d261aa3a89

Observation 6fc2d1c6-c9bc-49c0-ace7-5ab8591c63f1 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.523071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.523071Z digest=sha256:00d507e2642aa272d4842948ede51f802c05ce1a972546a9a52a37beeed8781e

Observation 713939a4-7a0f-458e-ad97-07f78449db3c · outbound

This paper cites Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.606573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.606573Z digest=sha256:8d520891dfb8fdcced68fee44531137de29f61d934c2b300d06b34158ae52a90

Observation 405ed9da-be08-441c-a7c5-f99b938dd518 · outbound

This paper cites Adversarial Reasoning at Jailbreaking Time.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Adversarial Reasoning at Jailbreaking Time

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.697568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.697568Z digest=sha256:8f072df47b872bd20364d97ef89595237f9c21b78412207fc28c6eb8e832ec6a

Observation e08612b1-25fd-4410-b5dd-12cf2946f872 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.771954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.771954Z digest=sha256:8dfdbb53fb9e8eb4db37fc61bfe0d0d356f3d5d046b3c9cf9cd19b07e6f4cef5

Observation 92990e46-971a-4b49-802d-bc35ea597896 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.856838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.856838Z digest=sha256:1a12c289628e7fdcf08daf345d73698babd529a00d8a550b545cbe0fe8cc72f8

Observation 3a8e54aa-775a-42e5-89c2-6ac17b513180 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.949465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.949465Z digest=sha256:d38d08fb3e41c61affab48010ff128122e8ca692b9fc4811286f909685baa3f4

Observation bf208120-e27a-4f5a-a3f3-df5416e19825 · outbound

This paper cites On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.022641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.022641Z digest=sha256:cfbe511c10e3dfd2ecd5b89b74874b25b63ad76679f5e621b72609a3d6400d64

Observation bf8f7d3c-d7c9-40f7-b544-a1fa7f6a0420 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.169747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.169747Z digest=sha256:656fcabeec10c782d175b6e73bd3dd20c2f2f46fbe866a3f61e8fcd014c1f063

Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.583503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.583503Z digest=sha256:7eb82660cbf53bd075bd09765f6972d735de280f868d465d4b30f45a1931f10d

Observation 7449f66d-9235-43f2-b462-499ca818af64 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:51.031284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:51.031284Z digest=sha256:4eddcf85c2883a7fbe65cafdb24997be66a50dbe763eb7497aca46705d9c006c

Observation 2128fd9f-9219-4868-92bb-21e14ce554b4 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.055915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.055915Z digest=sha256:7dd74adb6d848a6feb830d2e8ec412d29f40e6f3d6fdaa940376a4b642fb0c9a

Observation 42cfe4e3-fc94-478f-9ca9-40b95fdf9932 · outbound

This paper cites Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.354156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.354156Z digest=sha256:f5b6060a76598f342d7adf8f45b5fe3d0dff06c1f928362300aad892b3737394

Observation 2220d7ac-1e80-4b0f-bb7a-a0aa533bc019 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.581372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.581372Z digest=sha256:7d5be212776126f3f25650d4fb11906384db71263191c9edbc51bedfb1f4a698

Observation fdcdd31d-64c2-4f09-9a52-63b8c5b54c8a · outbound

This paper cites Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.635933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.635933Z digest=sha256:86d833a35681d9c631b6790ddf5e7ac2035f65b133524beb698b727641871bf5

Observation c59ba061-c2cb-49e9-8181-ebcf4f50a036 · outbound

This paper cites Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.250131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.250131Z digest=sha256:b09aab7551c228bd49daba673aba5f34cb8a596dfaf87247f7e71e1312963056

Observation d32bd04e-8be1-4b51-beb1-40f512cf0fcb · outbound

This paper cites Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.990999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.990999Z digest=sha256:8d1f40e48c091ea26210744585a2a67dedf8c08b1c6dfc9ded43e05c4bf66532

Observation 545b25e9-4311-4970-99e9-24fbce537175 · outbound

This paper cites Qwen3 Technical Report.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Qwen3 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.298882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.298882Z digest=sha256:e465cf43ac211807cccc19e257c6f325e6a0ec6158eb199189eb6506fab998ec

Observation 40ebbcef-d119-4401-8470-c8073730f54a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.365945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.365945Z digest=sha256:7505a68551870b5ad705c2fbdccf59ba8bf82e25e86978e36737fd70a7ace978

Observation bf4b7a5f-abee-4dad-b9af-69f47d76e78e · outbound

This paper cites ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.497263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.497263Z digest=sha256:aa0cedf3c23a96112f58f4a77db8c9439bab4ce6966bbd9b60d60eb0dde3fce0

Observation fe6b19d7-fe42-4231-b346-f3a6e7612a58 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.745667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.745667Z digest=sha256:91a6bc34ad0d3b18739cdca6e42b4b7f21b3e281b1b5dbd0eaf01a322acda285

Pith citing papers

Observation 4cb5c848-dc4a-4312-8813-dc690d41f70b · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.998787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.998787Z digest=sha256:9ecd8761e992ac61a1196ab56d9081f48e9c3bd1b90e5ac62f36e1d6fed028e9

Observation 35545380-d09f-42e0-aa19-7af9659e144f · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.999081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:6c1464e19920fb95cd7474c9798a76b2234544f976f49a1be237e7f4926cc205

Observation 101c8fcc-0766-4ad4-937a-0f5757a1c1a7 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.796120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:5558d268beb9fba1f84a03bf3613fc86639daddb27cfff8c46341beb0dbb0830

Observation 1e4f5c74-0e8f-4a1c-bbc1-d90185439540 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:51.382604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:f2773cd0d635f50dacfd8f3b665f4771f3dd4b9c3f4c11553cb4cb3dc059a922

Observation 0050ce8a-a01e-4356-bc8e-2e84d34e485d · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:079e9f5f22a95f8c50da562df2349b2c62242454a039027c7800804da40e6e88

Observation 1a452e8f-2252-4f76-a376-3f16b8a823dc · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.481619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:72e38ad2b16eccd8ff87441987cc622fa94b60299bdbfc878c26552324d919d1

Observation d33036cc-7cf4-4082-92ae-3eff1461b24f · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.243161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:ffbeb921ba7a0e011dca01ae52de973dfc305ef666532b657efd9200463a158a

Observation 87bd6cf4-27ac-4581-98b7-e984461203ab · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.152538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:3dcaaab35302249311de813028d3af496fa66f905046c7179ffcbffc83821110

Observation 27be7139-36e9-413a-bbc7-69b562b1b901 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.484141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:002f0d371d7e587d61555b6c21da62c9b3b153aa30043270246820e98901bb3e

Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.009207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:b4fcfbe357121659709428b8198a76ffefef53b19dceaa55806eb56dd69547aa

Observation cf37284b-4403-415a-868c-d46f9d7fc1e8 · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.362781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:56:53.056292Z digest=sha256:2f47147ba5173d4a32be2b6f8e7d153ebe2dd45437e56f41c63fd1c0c552e1ee

Observation fd7d90a0-e220-4ac0-9b16-9863a5e3d58e · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T14:54:16.625002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:54:16.625002Z digest=sha256:1bba92699cc6c511ec39820ff7f8c61193b85375ea25962fa00b3786c759d5b0

Observation 6dc354b1-2c7a-42b3-823a-96c3d3d4d9f8 · inbound

Regulating Branch Parallelism in LLM Serving cites this paper.

Regulating Branch Parallelism in LLM Serving Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.350558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:00:53.308946Z digest=sha256:b95b9fa5832e989c85f24cbe18ab92a1c5c5cd9d370c563b18602f78cbcb4482

Observation d9179099-d64b-4beb-9db5-5f92b6033424 · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:25.161203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:9ae9ec95e99412b17d3f9ed98f0e94f102d106c175b51e3634e26bb029f7998f

Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.337093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:8ec03dc57624652a109265544e981d9feb939ea53597fb3a79db82139593caeb

Observation 0aeec0c5-f4f8-4218-be30-738bdc2b06b9 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.030900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:a642b42c2dd1513cb98cfc1f07b54001a05684dc55faeff55a95d8f17b952bc9

Observation a49cfe8a-2894-4485-ab0b-ea0df3835026 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:39.525266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:39.525266Z digest=sha256:295c09f00b23d17cdf346f3b6490f313c881aef1d56f08a1eb61744e62398bdb