Pith. sign in

Paper Citation Record · LEDGER

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2508.09303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09303 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:12:12.793725Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:00:11.440671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T22:32:44.029593Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f254660-1248-4545-8767-f6d6303d4f1e · outbound

This paper cites GPT-4 Technical Report.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.149968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.149968Z digest=sha256:4a70930a042e73009ed433f9bcbad4b1fac06b7ce817ecebcc37a4a9f9e5b43c

Observation 61f23d39-ef57-499b-bc83-f6391a3f148f · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.666111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.265720Z digest=sha256:1d2b5ed657352efff0615ecb02c6d37334aeb924ed66a80a62d6b30fc2cf266c

Observation 60debba3-24aa-4666-9f1f-d21fd3254f03 · outbound

This paper cites A review of factors influencing user satisfaction in information retrieval.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A review of factors influencing user satisfaction in information retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.330163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.330610Z digest=sha256:b3798ee24c726e1f42d526d6f5ede67b65895166133a13215508e8bd4587d1a9

Observation 9cf214fc-bd03-4f87-b810-6759d0682650 · outbound

This paper cites The Llama 3 Herd of Models.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.409056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.409056Z digest=sha256:47c2a6b59aa580b393b4ded414a2054524f1126a7fa3ee34a7a4be969f9af3fb

Observation 926ad8e5-0928-40fc-98fd-ff6f831f0d3b · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.477639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.477639Z digest=sha256:9c311f7c8ba3a3c56059e09cc6220e36078354348ec8c37c56be2ce2c93ca8b2

Observation 2e51088c-9156-452a-8866-ac64a9521b25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.565175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.565175Z digest=sha256:d5911cc560a095ff62381e190d8b737828b7b834b5480f990cea5d9562f320b5

Observation 6ab708f3-dc0f-46a0-8c7d-86458547d365 · outbound

This paper cites Retrieval augmented language model pre-training.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval augmented language model pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.102872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.645837Z digest=sha256:abaf628b484b8f75305bcc07a14bd703c433e9c8980c1a7589d8eef21b5c9391

Observation 26e37146-a3b6-4175-ba2c-26680d9e7982 · outbound

This paper cites Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.861437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.710959Z digest=sha256:0bc403f2e1ea4351ffe773101b0734ae63492f0aff5ff957cb02bf1fa6ee8dfc

Observation c7e6ef23-f227-443f-aa1e-de78ad7f9904 · outbound

This paper cites ORPO: monolithic preference optimization without reference model.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ORPO: monolithic preference optimization without reference model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.593997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.741836Z digest=sha256:262e8a97891d56fe2d79169f61acb410a883ed667bb95d6c429800fc94d7bf0f

Observation 7a300d47-82ef-406e-89b7-5e75e66634e9 · outbound

This paper cites Survey of hallucination in natural language generation.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Survey of hallucination in natural language generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.354547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.773407Z digest=sha256:d8090724d86696fd673dd510c6875f5d7c4febc5f5e5e122ded10a7aca19bc52

Observation afe7e04b-abae-4a2b-96fd-ebc3a778285c · outbound

This paper cites an unresolved cited work.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:12:16.098728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.826331Z digest=sha256:3d823e92357a7a91d733c9f589392d56269dc9de6adad3af12f25e7667d0d594

Observation d6e171a5-39d2-41d1-aa9d-37d2404192fc · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.890698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.890698Z digest=sha256:9deb02ac797ca895d1d796d30675ff6b63da8f577bcbb0930e8f09b9e1dd4d68

Observation 2aa8d116-1426-4df6-b0c9-70faad697aa7 · outbound

This paper cites Weld, and Luke Zettlemoyer.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Weld, and Luke Zettlemoyer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.848643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.958980Z digest=sha256:8bee74550efd23365d451ab04e8eb64633dbdf655e427f349b048b399541cbaf

Observation 38f07e05-e737-479d-8b23-f5f8e54beeb0 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Dense passage retrieval for open-domain question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.722055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.008903Z digest=sha256:cccc15bca0cbeae55379e3c4cb59eaa5a6e0e3511e9f02de19803254a18a1d11

Observation 85cf4b4f-2a8f-42ac-ac19-09a0614378ef · outbound

This paper cites A survey of reinforcement learning from human feedback.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A survey of reinforcement learning from human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.064527Z digest=sha256:69192a170961e40dea5640830fddcb49ceacca7733d818b4c3b4a45d1ae1b44f

Observation 499862f2-3f98-46d6-9e1f-15841e55680a · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.594794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.160381Z digest=sha256:6128282c7677d12bb3f4fc192833cbd6fa966ba67a96409860bd08395bde08e3

Observation 3733b631-9114-4e2c-bd79-634a9bc85621 · outbound

This paper cites Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.477053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.216923Z digest=sha256:e3adb0c56ad71f0bdd5f9da04c965271e794a0204112022ddcbbc3fa24d4507c

Observation 76aff35a-0537-44c0-b9bd-34f2948e9609 · outbound

This paper cites Large language models in finance: A survey.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Large language models in finance: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.303679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.274785Z digest=sha256:af82ac7eedd7ac32a7d02ad20524e147f10588514ccb72298023bdd6fabb1359

Observation 40799862-6eab-41dd-9159-cb04c7eeef93 · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.155608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.328199Z digest=sha256:3c2365b2695a7dd02c9aeb51a7c7a16eaa1df1dbeabd08459082e69caa113d16

Observation db71740e-3dcd-4e2e-9a94-e4126188259d · outbound

This paper cites O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.380262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.380262Z digest=sha256:203afba6864cb3ccb5be010738907fd0ac96ceecc77a8298ef0bbfaa9a68f636

Observation dd37ca0f-46cc-4136-87e3-e5f3dc0f416d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.987611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.469102Z digest=sha256:a07b5c595c378699723dc606879059fe7603ace096dcb8acddaa1cbfffcacd6d

Observation d2c768b6-3469-4726-a7d1-178707bc66ec · outbound

This paper cites an unresolved cited work.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:12:14.854532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.531717Z digest=sha256:b19a199101e6d4f10c8d8aa7267b33f996c6f4f18831015ab0a7671c9ec73600

Observation c510312e-8ad1-4642-9804-b8d81dcab088 · outbound

This paper cites Iterative reasoning preference optimization.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Iterative reasoning preference optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.620173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.614989Z digest=sha256:4642b0aa38576d80069373196649440c2c1d6c9df821c3e3d1a320de176fa197

Observation 88087109-df13-4f68-9096-5b191e5b10e7 · outbound

This paper cites Smith, and Mike Lewis.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Smith, and Mike Lewis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.444756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.660534Z digest=sha256:63b8d091d2472bf587b6dc078f1559b1e1040252bae0e1c93a65afc9a9d1deda

Observation 0f20ea70-41b3-4e1d-ad52-84df737196e9 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.294887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.710877Z digest=sha256:12b10d088097701ee4715ac586e0d6e1372ae9862d93f61339f982c534d634cf

Observation 9fe1ab6b-48d0-449b-8029-bcf0917ca917 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Toolformer: Language models can teach themselves to use tools

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.082981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.771234Z digest=sha256:c5874493f0d02631e6a1b625e52006e3c324df5692d04c3d8592d6cb9bce54c5

Observation 81d2ab38-8b65-413e-8d7a-fe17069685e5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.811800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.811800Z digest=sha256:84324e0516fae2c87d6be8737ef6765c032e0c9b1485976bade9feaabaaff3e1

Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.863783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.863783Z digest=sha256:e2101eb9480122b2b7e81e646b3d66268db1020670cc05f033d5ccb6f7610a50

Observation 64186f76-6f6c-4130-a97c-02440d495a05 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.902466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.902466Z digest=sha256:215d6512a6952aa3ef70b4bf8d06104f8267db98bb4444160f41d78d3adca35d

Observation f3fca3a1-431d-40fe-a749-5676a5df1ae8 · outbound

This paper cites Sutton and Andrew G.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Sutton and Andrew G

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.926922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.938138Z digest=sha256:63c932be50ce80217dc40b375360333c5404981db9ea410bb399b0085899e70e

Observation 6d2d5a50-a4e4-4e37-bc26-1ccd91cc38aa · outbound

This paper cites Multihop- RAG : Benchmarking retrieval-augmented generation for multi-hop queries.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Multihop- RAG : Benchmarking retrieval-augmented generation for multi-hop queries

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.800825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.022501Z digest=sha256:9b39b7a784ba7da5968aa14801205c2528968404ea5377a8e571711bc6c53cd0

Observation 0b75f438-f758-478e-99dc-a32f4f18f262 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.111745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.111745Z digest=sha256:390be2e1b3690a637b3cbc882a2b533bf2b30aee57a555b002345425d2497d30

Observation 2a37d315-32e8-4778-82d9-7287715f7736 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Musique: Multihop questions via single-hop question composition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.171420Z digest=sha256:b630952cbd86825e7981fb7793c3689561f33357f6a31d3e0eab85912e4b36c8

Observation c3f928e5-172d-4496-8e0e-f6a3ea73ec76 · outbound

This paper cites Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.509777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.220676Z digest=sha256:a2a424236dfcc5e8780bcd29cb658dad80872d004ab31466e5a57658affe22a2

Observation d89a6b42-e2b5-4355-8559-36edbc6133c3 · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.308929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.308929Z digest=sha256:bace8d17bdf7aa072d4147db84056c8803247793f7165cbdce3bb614100c2025

Observation 8cdbb16b-0648-4682-840d-c6757cefeb33 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.389009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.389009Z digest=sha256:57e4043358b3bb87ee9b289438f6e72ca23dfdd9a2836b87634665cee6a853fb

Observation 82c3f128-cd6b-4ca1-b174-dfc382795218 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.451731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.451731Z digest=sha256:c244741dace1e49c09b25611dad9e91e45ae37deb459d306cfcbd91694adb69a

Observation 84dd0aec-2dac-43c6-b372-30f4b55a0925 · outbound

This paper cites Reasoning or memorization? unreliable results of reinforcement learning due to data contamination.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Reasoning or memorization? unreliable results of reinforcement learning due to data contamination

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.500175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.500175Z digest=sha256:a0e6f52e3a7fdfd5f4099100796e54873930fdb64ad27f9815192c7f86246ad2

Observation dc6a072b-c9c2-460c-8d00-ee5ef2d0a98b · outbound

This paper cites MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.567646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.567646Z digest=sha256:dd14ef49ad6a51b45e8a3ad8e19b02cbd77562f9306e8fe17dc8f41ae9d2d9ae

Observation 8b525432-1de0-4413-ac36-c9212bd25ad7 · outbound

This paper cites Qwen2.5 Technical Report.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.658091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.658091Z digest=sha256:99849b8e136970d78b13a11355471729f3a8a7d7f6205c30fc92d88f65860b67

Observation 64153c56-4968-433a-b4d7-a68af4fd5931 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.389477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.715033Z digest=sha256:ece7672819b98b672992d681e3264f63e0abe399a7bc313403418272365e8318

Observation 549c28d9-061a-40bb-a086-9bd47816549a · outbound

This paper cites Narasimhan, and Yuan Cao.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Narasimhan, and Yuan Cao

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.239715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.760593Z digest=sha256:eca257697a8ae240ec9c7cdbb0fd88a210b20b9fa60e9a422327400e0eba5d10

Observation a035ada7-ce11-4bbb-8836-820e9b7b309f · outbound

This paper cites R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.793725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.793725Z digest=sha256:91710e952f344323df64b78a5dc1f4641104517f8a4edb2a4efeb121413bd48f

Pith citing papers

Observation 66d5ae84-a72a-457b-a78b-f5679185acc6 · inbound

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning cites this paper.

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:11.440671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:00:11.440671Z digest=sha256:258aa99d021fa064e98c5b9f417b533ea5a2142c5c9f60b9e831ab620eeb8d1e

Observation 74f3c35a-f61f-4efb-a4ee-568e19fcdf03 · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:18.137585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:9e0cfebb9b01a3f7d2df9836a06bb0b17d39ae5b7d2366146cce5919462478ff

Observation 5d40a41f-8c7f-4ca1-bec9-e8442c3461bc · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.974701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:bc20a5eca79c65602f7cab7e77dd6ef25bda2264db07815e76f67acd654d9060

Observation e0200f79-ff96-474c-947e-e6565d15f8f8 · inbound

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG cites this paper.

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:12.210309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:27:00.257353Z digest=sha256:607be2b6fb6431321b104d95951c0ca5d6e6e916d94a0a50bc2da8198798b480

Observation 4ea68c04-6587-46e0-964f-045f83cba423 · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.310629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:28:36.266167Z digest=sha256:d39dd28b1d73d7132a24fe30e4f8028ab9903f642652545014547496df0778cf

Observation b5213207-8b55-42ad-bdcb-e74528680afa · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:19.032968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T03:18:01.006274Z digest=sha256:5f4f04cdc62bbd744a0cd9f91643d6d98647b665f017a43425e6753eef7508e4

Observation c1b3fcab-7c07-4700-a910-7f540d833870 · inbound

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward cites this paper.

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.031157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T22:30:00.735630Z digest=sha256:20ddfba3f67dfd33dad5a978374c1c7bb061361bcd527f82cb5f1a62bfa515f2