Pith. sign in

Paper Citation Record · LEDGER

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 17 inbound Pith citation observations for arXiv:2509.07980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07980 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:25.998787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:30.028956Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbdb3f4-4df4-4b74-b593-063bbb42ed12 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.417972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.417972Z digest=sha256:f308158036688080643fd42976d0154ae5323f2f5b1c4be47d1bcedc433356e2

Observation ed11c2c4-520a-4729-99d2-62093fe0dd0c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.842982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.842982Z digest=sha256:545d5fe53c14efb54007b69f8cbb8ef7d81bc334d3602b8aedd77a72fbddfea8

Observation b198ef7d-3422-4d4b-8989-f7f26f6cbc2a · outbound

This paper cites Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:28:53.391615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-04T21:28:49.059722Z digest=sha256:7252c2da473e50ff39036b114fa3df20b19b27c51e92782f75cd95cae2a80e74

Observation d60036a1-5e84-4f5c-a3d6-b9a9c62b5162 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Efficient Test-Time Scaling via Self-Calibration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.143122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.143122Z digest=sha256:dcf76330ad102b221269e63bb9251715ba7145fd5847e13111d53d9a403836d9

Observation af4ec024-5037-42b1-81e7-713b7e716b3d · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.329292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.329292Z digest=sha256:006e4d15d6434177da3f265990205ce0a0fcadec426018dcd42928cae22e19db

Observation 9547c6ec-a53a-4d31-bff0-449c73ccfa24 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.390681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.390681Z digest=sha256:42cd1736447eb3abfc07a77fb1fd5056c8722b4a78a30ec944591bb2acc1c380

Observation 80287147-8dee-4a21-9caa-501098559db8 · outbound

This paper cites Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:28:53.693150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-04T21:28:49.445211Z digest=sha256:d0ac308682e5b23dfbcd2f6a87a604a0afd493051e04a75d5b3b9cb9a7642f7f

Observation 6fc2d1c6-c9bc-49c0-ace7-5ab8591c63f1 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.523071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.523071Z digest=sha256:8baeb38e282fd089ffc643e595d5db3d171ad864dc2b2b15cfa90e849e0a77af

Observation 713939a4-7a0f-458e-ad97-07f78449db3c · outbound

This paper cites Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.606573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.606573Z digest=sha256:4539e0f1cbde4ec36e438c8586ba782c334d0e1456288dafd1829c3af839b0da

Observation 405ed9da-be08-441c-a7c5-f99b938dd518 · outbound

This paper cites Adversarial Reasoning at Jailbreaking Time.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Adversarial Reasoning at Jailbreaking Time

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.697568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.697568Z digest=sha256:72b6891e81aacf5ec8d0992d4f098eaf73439b313abe1874a709105d33a893ab

Observation e08612b1-25fd-4410-b5dd-12cf2946f872 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.771954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.771954Z digest=sha256:2c11939decfc909ce33ef12487b9a220784e372f2a61d765fffe9c2fd704b65b

Observation 92990e46-971a-4b49-802d-bc35ea597896 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.856838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.856838Z digest=sha256:85d603715e356976b30033c8c68d90e36dab318476f6ed48b0a2dd434b84e51c

Observation 3a8e54aa-775a-42e5-89c2-6ac17b513180 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.949465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.949465Z digest=sha256:d9b07e5ab9554d623e2e990e73677d65205a8ddacbdfe1238d1155413c287109

Observation bf208120-e27a-4f5a-a3f3-df5416e19825 · outbound

This paper cites On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.022641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.022641Z digest=sha256:8ec8fcac0cb34b242a413833f71ce39020e71b8451e7728d6b897ee9787002bd

Observation bf8f7d3c-d7c9-40f7-b544-a1fa7f6a0420 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.169747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.169747Z digest=sha256:9251844fc07873d55f1a1b113631e18efa16847de9676269ab67ebf001d238ef

Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.583503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.583503Z digest=sha256:d53e657088d5eb596a411f5308cf2f1e81d5d739845d691162517b547de45bb5

Observation 7449f66d-9235-43f2-b462-499ca818af64 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:51.031284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:51.031284Z digest=sha256:7824ea8045fb2618543633b8a1e0df2cd02a61dfb5b15371de9bd0b2f1fef96c

Observation 2128fd9f-9219-4868-92bb-21e14ce554b4 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.055915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.055915Z digest=sha256:e64c299e0768a1104906f0d08fe0d4eb76a4087023f041d4833e495649645cc6

Observation 42cfe4e3-fc94-478f-9ca9-40b95fdf9932 · outbound

This paper cites Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.354156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.354156Z digest=sha256:28b3ecd05ec0280b3b1d22636d6850dd997a4272919fea40c2bd1cd4b229c083

Observation 2220d7ac-1e80-4b0f-bb7a-a0aa533bc019 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.581372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.581372Z digest=sha256:a5ae8439089c4c0cf282d23c34ddf3fa740ca1d2c70f9b45eb64cdab8c7fcacd

Observation fdcdd31d-64c2-4f09-9a52-63b8c5b54c8a · outbound

This paper cites Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.635933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.635933Z digest=sha256:b78a4ce0a2105f30936e3abd7543bbd095cf58b4d97a441dce837ad9431d3645

Observation c59ba061-c2cb-49e9-8181-ebcf4f50a036 · outbound

This paper cites Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.250131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.250131Z digest=sha256:b2e9a61f4d009c04acad1e68aa11704ef5a54728e9e13fd3680b62d7ceb86b78

Observation d32bd04e-8be1-4b51-beb1-40f512cf0fcb · outbound

This paper cites Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.990999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.990999Z digest=sha256:49e733d0eb9f6b81bb051ebd9b98769ada0ed25ed58c060d1d75719613415d6e

Observation 545b25e9-4311-4970-99e9-24fbce537175 · outbound

This paper cites Qwen3 Technical Report.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Qwen3 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.298882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.298882Z digest=sha256:ed595968f37df60992db2fa3ebf4231674ef0989d57d29ebe2f1760a9f345dae

Observation 40ebbcef-d119-4401-8470-c8073730f54a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.365945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.365945Z digest=sha256:e38bdf444996bc3106cc8f53c9128e3b0389b67cefe950e2905a06490b3f231b

Observation bf4b7a5f-abee-4dad-b9af-69f47d76e78e · outbound

This paper cites ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.497263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.497263Z digest=sha256:4844c5fefbd257f5ce621e1c983b223bc7ceb8d6ed7bd74a69c3feb4d623cf00

Observation fe6b19d7-fe42-4231-b346-f3a6e7612a58 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.745667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.745667Z digest=sha256:6562bf10df66cc22ca560adaa4f68f3617282727ccf6070663647097fa9d2b83

Pith citing papers

Observation 4cb5c848-dc4a-4312-8813-dc690d41f70b · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.998787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.998787Z digest=sha256:b6560c58c15dae2d29d77b79c4ad6b5b1d094b3d43e0f82842f6ad583798c0f8

Observation 35545380-d09f-42e0-aa19-7af9659e144f · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.999081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:1102b588536440c6d8b0757cac8e16929d6ca2137d9dec93c7bdd78b3024e12d

Observation 101c8fcc-0766-4ad4-937a-0f5757a1c1a7 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.796120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:eccb10af806f46b6f46d292baf5a5ffeeadf35d671c68c88505a673430d8f5bc

Observation 1e4f5c74-0e8f-4a1c-bbc1-d90185439540 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:51.382604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:21c0a7592fbbc35113c4cf1ba1d9bdd34487a41d0fa31bf40139b71a994342fa

Observation 0050ce8a-a01e-4356-bc8e-2e84d34e485d · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:4322cbb3342c5da63af800cbc2d53b1482cf6f1eb408b90b8db6afd75d156192

Observation 1a452e8f-2252-4f76-a376-3f16b8a823dc · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.481619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:777d71683bbc8cfd1164d9281d0ab575f7cc017d02dc9ab68179468d7c7facb4

Observation d33036cc-7cf4-4082-92ae-3eff1461b24f · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.243161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:2515509702dafa370de6a5dc6f88b1825092fd60791531a1ddecf6c53be1abed

Observation 87bd6cf4-27ac-4581-98b7-e984461203ab · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.152538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:b6c9a8911694236485cb97e518a6f512d461e8e8034661a1000a02c053bb18f0

Observation 27be7139-36e9-413a-bbc7-69b562b1b901 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.484141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:d11de723b824632983c465bf8337ee046cd2eacf4de1de7b5bb09fd4c3063756

Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.009207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:9430894e072e1c556a4f7d0796197d00056e9265dcb488b0930bd9e15b254a6e

Observation cf37284b-4403-415a-868c-d46f9d7fc1e8 · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.362781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T16:56:53.056292Z digest=sha256:b3ab897973973dce8a2f34612ad8ad607ffb9411f8c3f5a7fb7e54d5ecfd0d79

Observation fd7d90a0-e220-4ac0-9b16-9863a5e3d58e · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T14:54:16.625002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:54:16.625002Z digest=sha256:a242e2fc98294724f384ff7f4799e0942869ee13606afd6ff23eea3be3d8de61

Observation 6dc354b1-2c7a-42b3-823a-96c3d3d4d9f8 · inbound

Regulating Branch Parallelism in LLM Serving cites this paper.

Regulating Branch Parallelism in LLM Serving Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.350558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:00:53.308946Z digest=sha256:975e46c5b8d7414b9b7d7b3ac2047645f882dd4df6b6d7be55a0cdf5b6520052

Observation d9179099-d64b-4beb-9db5-5f92b6033424 · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:25.161203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:94c0557e8274f7aa4b5a42b55e174448fa3dc22a767f6a1b92b80b5de1c6255e

Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.337093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:70230d482aeb107a82d045452a21c5ca7c81dc750c30fb92c2567862bcc1d5c6

Observation 0aeec0c5-f4f8-4218-be30-738bdc2b06b9 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.030900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:5cf1ea60faed18a98ba244cf69401452379cb3ab08bcc6bf4643101133adf884

Observation a49cfe8a-2894-4485-ab0b-ea0df3835026 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:39.525266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:39.525266Z digest=sha256:14e426b5e525bf8960c244c388c2a3f4ef02d4b527fe510bb917d46db6571e43