Pith. sign in

Paper Citation Record · LEDGER

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

As of 9 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 2 inbound Pith citation observations for arXiv:2507.08960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08960 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:13:00.348063Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T22:47:51.676289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.756820Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4260ea4-7abf-4edc-b541-3b2a19422f04 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Graph of thoughts: Solving elaborate problems with large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.005763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.005763Z digest=sha256:c5aade4813253ebe3f03176c1c9c2b7a1790add9a8a9133b9bc8933f65a4e634

Observation bde69cd1-60e0-4620-90ec-6d3d5dcda0ab · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.029371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.029371Z digest=sha256:096319f89ef0ca20c0015214e9a01297222e57e1bdb09cf8f2046b2140b1bd19

Observation 4a557c60-7ca1-4d22-b4a3-e3bae4c1f77a · outbound

This paper cites Language Models are Few-Shot Learners.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.051900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.051900Z digest=sha256:359440ff751ca34a5b96f3b286b2a69b1e53d60d551e841e537a05738f42cbab

Observation a67cc61c-e0fe-4db0-8a0b-f7eff4fda5f7 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.078990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.078990Z digest=sha256:c60315c11e2e7309da8d72a16807f3b1dd67bc02e8dab110528b6b7901090c13

Observation 12af781f-5054-43ed-9f3b-ab547c44287a · outbound

This paper cites SocraSynth: Multi-LLM Reasoning with Conditional Statistics.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SocraSynth: Multi-LLM Reasoning with Conditional Statistics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.107269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.107269Z digest=sha256:518e86c6e56313579b3cd5d6f2432a9089511c7d4fa95a05feb4a19ab0c09a1f

Observation a3fcc3d6-af83-4d07-9272-a7a0c28e18cb · outbound

This paper cites ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.136451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.136451Z digest=sha256:906f66c01ac3dad3596037cd8cb11ecc796dd8aefafc23bd6d64f8470d732e16

Observation f61326d0-7de9-4c33-887e-afeeaeab4906 · outbound

This paper cites Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.172268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.172268Z digest=sha256:cb2adaef221070cdbd30abc32345c4bddddd047afcef7bf21e4ce392f03c8397

Observation 60908b6d-180d-4b21-a7e5-e772302339b5 · outbound

This paper cites Universal Self-Consistency for Large Language Model Generation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Universal Self-Consistency for Large Language Model Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.200045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.200045Z digest=sha256:751a8f77e5967c6ea610d99a749c41917430709652416f69c1864047ee947d55

Observation 5ef96167-bebf-4553-b765-e9a53fc887f1 · outbound

This paper cites Cost-Effective Online Multi-LLM Selection with Versatile Reward Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Cost-Effective Online Multi-LLM Selection with Versatile Reward Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.230490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.230490Z digest=sha256:24e180bf219da15149647ade0db222ede3f914cdf57d1f84badedda9071fa5b7

Observation 4048bd57-f74b-4197-8b87-0f335d1a653b · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Improving factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.254960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.254960Z digest=sha256:8619d76acb7ea8297b25c0fe5f25e368568b314ff250f9b4cc33b9d7faf4b83c

Observation 1808ef03-445f-432a-b6b7-ec0674130b18 · outbound

This paper cites Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.282695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.282695Z digest=sha256:553df34d79805e029cf5bb3d0eaa15a17229ad3be529d7f92dee8686fc3b03f4

Observation 2add6b6c-d084-4bb7-91b5-4ef3434c6652 · outbound

This paper cites Multi-llm debate: Framework, principals, and interventions.Advancesin Neural Information Processing Systems, 37:28938–28964, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Multi-llm debate: Framework, principals, and interventions.Advancesin Neural Information Processing Systems, 37:28938–28964, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.384428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:58.316555Z digest=sha256:4732b1a9b23c2c4df6a0069f007d467dcef1aff738d5061a238f4c49755a1935

Observation 5cb5f25c-4895-4e1c-b7b0-8109771efe23 · outbound

This paper cites ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.333530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.333530Z digest=sha256:e13b3d7979c152561e76e1e4034f33027f600537fde5143c16cf5780bbfd58d1

Observation ceb1668a-586b-4453-b857-772546007899 · outbound

This paper cites Acc-collab: An actor-critic approach to multi-agent llm collaboration, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Acc-collab: An actor-critic approach to multi-agent llm collaboration, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.218235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:58.359750Z digest=sha256:dbce9dcf7b8a44d6751cb814e49fc402d97b0f21006746d9495d27dd0c4273b6

Observation 4c63ecf6-b762-4b69-a2a7-99d65e5c23fd · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.377961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.377961Z digest=sha256:aef230b4ccedd93ef0323d67d1d735f2c92393b40b2602f9ff6be2fad318c18b

Observation 62cfc9c8-6a77-45ce-9136-e555f92dca6e · outbound

This paper cites Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.398205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.398205Z digest=sha256:caeeccc4f708398a155215fcd19a62be5148a8855add1948a4f8037198f1118a

Observation 0ab1ff3e-dcc9-448b-a654-72d8a3ac5e3a · outbound

This paper cites When One LLM Drools, Multi-LLM Collaboration Rules.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs When One LLM Drools, Multi-LLM Collaboration Rules

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.412984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.412984Z digest=sha256:53918711d56c1bf45f606e956e65b3b54e83fb7ced6df5253c0c0282b6f34857

Observation 98037f10-d385-4b10-96be-593571730124 · outbound

This paper cites Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.440669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.440669Z digest=sha256:37a21a7f6dc550f7bdc21959690cce6889a706bcfa365a964c0cf4d1fac1b974

Observation 0c44c63f-9bf1-481e-9251-6cf1b33cc292 · outbound

This paper cites The Llama 3 Herd of Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.461508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.461508Z digest=sha256:d3b0178f3b2f7938264424eeb1c7629d8648e95682c483b08127a16376a07789

Observation de01ddc7-bbd2-4760-9bb2-e1a489f7016b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.485112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.485112Z digest=sha256:e18913cfc2ab7efaa1c5f1ef1ed1415fb3e5e1f5032caddc38041e261d1ae605

Observation 8bc2de8f-deff-42d0-a3d8-50001e383221 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.517427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.517427Z digest=sha256:2835fe07f637fa110e0cdff1a81160679bf52dc3cc21de79d4f3ee04613e9f4b

Observation d07b6f32-3df4-4d8e-80bf-56647a1cfb0d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.548499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.548499Z digest=sha256:6096cc7b399061b39bf54f3480445a41e4087b1b8f9261a35672bea3a765b631

Observation 644f457c-aea3-4539-905c-75dd4b6023b2 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.583192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.583192Z digest=sha256:128ad6d777388c42416dc87e0d5a6f07c4d4d9941dd7a3e61dc85f3a4179c262

Observation 597c834d-ee7b-4050-a9dc-87ef7e557153 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.605891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.605891Z digest=sha256:944f1ebe3c15d8dc13b501ecb3050511dbc22a9bffd506cb0377fad77771dd98

Observation 99bc94e4-dcf3-486f-a5bd-b69e6a12db7c · outbound

This paper cites Ensemble learning for heterogeneous large language models with deep parallel collaboration.Advancesin Neural Information Processing Systems, 37:119838–119860, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Ensemble learning for heterogeneous large language models with deep parallel collaboration.Advancesin Neural Information Processing Systems, 37:119838–119860, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:03.050076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:58.630989Z digest=sha256:dcb88653f4731cd6c310229ba5198d0ef75a725d2a05ba3cb93c1dd8cd3c27a5

Observation f2669eb8-0f89-4653-984a-3a5a2f1c3b18 · outbound

This paper cites OpenAI o1 System Card.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs OpenAI o1 System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.649279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.649279Z digest=sha256:c49af490e27a35f7894e139f6cf0612d46c92ed5d220c28533662bb755b9bc33

Observation 908b419e-cb3d-45e8-b949-634ed9bfb66c · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.673023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.673023Z digest=sha256:983a9f5bd3d1db0ed3b6fb8f8d98e1fae01e37821315b628a3aeb00d05224682

Observation f49e5ddd-dd47-4ff4-bbfa-4f50ecc3c14d · outbound

This paper cites Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.703046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.703046Z digest=sha256:fc324f056c1ff2fa9c79f6c8ab8a28b56fa84226ee2326c23d3b519916b5f63d

Observation 12f9681d-b8c2-4344-b6c2-8546819eb797 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.734477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.734477Z digest=sha256:3396e6659cfbffe30121bcfc8f17da2045dc2fc4fed226d4c8ccc01073cdb8fe

Observation af1bd348-84c9-48e1-bb11-b38be3e4ea4f · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Training Language Models to Self-Correct via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.743208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.743208Z digest=sha256:23b0256ff8bcf1c38829c1d5bbba013a6c06c9372f8cc6dfce55f753b4ed82c0

Observation 065edb2c-6337-4c27-b62e-6ffbc0d040f6 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.768620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.768620Z digest=sha256:7ab72720695e26c55aefe93fd416a82886c0557bbf222d78791a3235ca934292

Observation b06ebd05-8852-4536-be17-d922927703b6 · outbound

This paper cites Two heads are better than one: Dual-model verbal reflection at inference-time.arXiv preprint arXiv:2502.19230, 2025.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Two heads are better than one: Dual-model verbal reflection at inference-time.arXiv preprint arXiv:2502.19230, 2025

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:13:01.369146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:58.791281Z digest=sha256:4ee252ec696441898a5541c9a98814da695676d5898d711e06aaf2d10cdf19c5

Observation 7db4046e-e50e-45f8-82f3-bc13de926a5a · outbound

This paper cites PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.824124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.824124Z digest=sha256:398cb39fb2202307af0af3f0285556b0e86b643f15f9447b02f1457eb3e8450f

Observation 700d64f3-8551-4e27-b970-ba6321acecd4 · outbound

This paper cites From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.843340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.843340Z digest=sha256:1be0a58668587ab1b8e891ea1fbcc82d32f038f79296869f4f0651f5ee4ff32f

Observation 139d4ef8-f847-45a9-8eb3-94ccb2f62650 · outbound

This paper cites Improving Multi-Agent Debate with Sparse Communication Topology.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Improving Multi-Agent Debate with Sparse Communication Topology

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.866755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.866755Z digest=sha256:f4c51115d9da75e4d545f85a5dac784e46da16ec035a9f11e0b3880c2a141bda

Observation 51167679-277f-4271-a834-e45d56d82b17 · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.901843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.901843Z digest=sha256:fa4392e40c543f786147c57018b3edd11f054c1d31971dfeca07835d7bd4ba82

Observation a90916ee-2ddb-465e-bff1-f8c54a8c7fda · outbound

This paper cites MARFT: Multi-Agent Reinforcement Fine-Tuning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.937766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.937766Z digest=sha256:c83f40d084b87baa0826b7531b940edb76349a5ccdf4b5c98125913f6b016430

Observation ed24df84-fe46-4fae-8838-6e7361642413 · outbound

This paper cites Groupdebate: Enhancing the efficiency of multi-agent debate using group discussion.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Groupdebate: Enhancing the efficiency of multi-agent debate using group discussion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.973810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.973810Z digest=sha256:376353804bf4cac45d8eeed01a5e3e727f725622fd522d96e23de701d036b678

Observation 341a368d-dece-4147-a542-a14dfe331d4a · outbound

This paper cites Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Towards Hierarchical Multi-Agent Workflows for Zero-Shot Prompt Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.009997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.009997Z digest=sha256:6c5fa06c5201313b40d93c1bdf9672638ff295385e203fb9b2ac67b42e139505

Observation 5dc73893-7337-4db7-ae30-b87f74102874 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.047193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.047193Z digest=sha256:781c1056304f037e734f23fc95b5bd24472889658362efe5606555a77678b6f7

Observation 9b776a90-25db-4ee1-a943-321d814100c8 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.901335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:59.082303Z digest=sha256:7366d6ac1d08c569fb97d9b723738e423f9450c6fde280a5b6b3f43be457622a

Observation 88dd62b2-115b-4d69-99f7-df26b77abd66 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.115363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.115363Z digest=sha256:24cf464c7e237df7ec0cc5be43c5d20506b1782891912dd02d8578d992294f02

Observation 7c3f980c-8035-4390-b9e6-2d0ab0c2195e · outbound

This paper cites SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.145103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.145103Z digest=sha256:3dd8286eccae721e7a56e99f1491281fa6b7934785b1a17730fddb57d17ba8ea

Observation 97d93bbb-9545-4b74-b30d-9c455ceb48e4 · outbound

This paper cites Beyond accuracy: Evaluating the reasoning behavior of large language models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Beyond accuracy: Evaluating the reasoning behavior of large language models

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:13:01.062647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:59.181998Z digest=sha256:0b96ffdc84c1fd1f56f11d0be0de1a29709b9ff185b7f1b21b5d448ff7785a4f

Observation e7966f11-a285-4ac9-b61d-d8d1ac569a1b · outbound

This paper cites Motwani, Chandler Smith, Rocktim J.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Motwani, Chandler Smith, Rocktim J

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.216898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.216898Z digest=sha256:70870c31737473de9e9900e2e61cb4bde6b06403af2c4ee23d08e3829d7967b1

Observation 933c70b5-5273-4ec6-b898-39f5003ff99f · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.245532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.245532Z digest=sha256:f47950f4e2fe84081ed26d6a2c0a14590a6a917c5163f2b53e5cea0d7d1e2aeb

Observation fcab31e2-8508-4fe2-9861-2d58435f8910 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.282290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.282290Z digest=sha256:b7cd539c35474bc1a986c3bf5d6e3644b901fe9e0b8192e6a90fa379b43e2837

Observation 6f94231f-a55b-4841-ab74-3b74cc6f3f99 · outbound

This paper cites Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.318146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.318146Z digest=sha256:2181804186ecc0284c3d7a56ba70d4f52c430159c1625a902859e998acfc3095

Observation 290f5497-d606-43da-af7d-05ffa0e50a00 · outbound

This paper cites Qwen2.5 Technical Report.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.353123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.353123Z digest=sha256:806306975290370d135f49440ec5823e5f07fdba47705060d73dc7073982a887

Observation 35070d5d-222d-4fbd-9321-603e556d56d2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.389070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.389070Z digest=sha256:a2ef8d5646df69a4520b0895430d7f73e2985d496431b755145f38dbaf21505f

Observation 6f8fdf02-dab1-4ce9-8149-e33179bf2602 · outbound

This paper cites MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.444849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.444849Z digest=sha256:dc5a2d6e1aaf7297d34e428902bb785dbddef4142998a35c25a020a7c4a339aa

Observation afbc721e-c949-4498-8155-5eb6cdfc1040 · outbound

This paper cites Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.473520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.473520Z digest=sha256:8ccd16605631254979a91f059c6f7de7323cb73143077f3b4c445ea22a3e56a8

Observation b08e34d0-5cb3-47a4-a98d-741b4b28bd2a · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.492882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.492882Z digest=sha256:665786bff3c9c3a5598d1eeecfb49a25ad31da2b1de4ecb948da684ce186cc52

Observation 42978e99-da8c-48d2-b500-364fcb55df68 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.819843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:59.524423Z digest=sha256:97ff075d6b3619f4991b0eb3560e418c9382b07345595b2ce20a68aa78aad6dd

Observation ec833267-3708-47c0-93a1-ed34501fb33a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.592171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.592171Z digest=sha256:8a32cbc36db1403f24ff77a7b442338e278327065c08366676ffd7acfc26b46d

Observation bee4dd27-919b-47ab-a771-c8f20354d44b · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.617501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.617501Z digest=sha256:0fa73e0c6b31e1c3baa43e5886df791d23a17a54cc17b48fa34945c58a54c89d

Observation 69d857ca-a367-4744-b812-1b5a3b746474 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.646978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.646978Z digest=sha256:ef958d0319268a8d81a62d2ec0a13c28697a8e0540ed71d8edc6e83f0777a0e8

Observation 456c1972-61bd-4e5d-8cb0-60ac7f7abce5 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.678125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.678125Z digest=sha256:2501dd85bee7ebf95b3529b450303c303b10ec68d6a7bad7c863c508bc97c0d0

Observation c6f7f492-6fe2-4edf-b028-048728c7dd10 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.703873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.703873Z digest=sha256:5a4ec5069959afd08d8627c7c07c9af07fd95f17d8e588ec9b0d98961e65e77d

Observation 7ba63848-f0ec-402a-82db-1d4aac63c743 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.735760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.735760Z digest=sha256:a92e589e33f1b962dd12c9cbad038114d0eefb4d775fb627254b49828fb9ae97

Observation a60e3e36-e01e-43d2-b319-9f9fcb5d416a · outbound

This paper cites Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.763216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.763216Z digest=sha256:24a0daf50326cb7af02cb68787b547ab579809fb990333bf5a61330ef7a65430

Observation 8e1bdd83-af97-4041-b0e1-521d2b1b7820 · outbound

This paper cites Qwen3 Technical Report.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Qwen3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.793527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.793527Z digest=sha256:5247348bd52c7722d6b77e2d789a9f38bc970adabf35fcaaa02cf0402890429d

Observation 608823d2-2b24-4911-9c2b-a6368629f91a · outbound

This paper cites Multi-LLM Collaborative Search for Complex Problem Solving.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Multi-LLM Collaborative Search for Complex Problem Solving

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.822314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.822314Z digest=sha256:6c0857eff7cff6f19cfcd4f6a1b496ef57dd9ca1ba8c0b3afa5cd62ea243b66a

Observation eca49996-f6e0-415e-9532-a107d19292dd · outbound

This paper cites AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.855011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.855011Z digest=sha256:26c94c73e14fbd4fc0c89a664b575194039a279e3e7fc0a90ee9ff2948964553

Observation 2873cd1d-8021-4c09-a3b6-8f98b5d851d0 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.892203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.892203Z digest=sha256:df99443572db1bf5811fff315ad95f4b5069c95936029e9197b7e7bba582e9af

Observation d1dd632c-f6eb-455b-9ab8-cdd6c3d1d633 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.917749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.917749Z digest=sha256:3060604cd3b4b788d1c1723a0af8369849c7f1b67504f1c37dd4f14814e598e3

Observation 3a0deef6-b5a5-46c0-893e-2f70b92aef8b · outbound

This paper cites X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.939425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.939425Z digest=sha256:2822e0b0888aec277a3285cd50fd4b991356d6bafc54636ac2afcbe5f2bdaec2

Observation 98fead6a-7a30-415d-8930-bc0ea65e6463 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.958609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.958609Z digest=sha256:6bb4ffc6b456bf81eddd3c0058c5621397ce0f92cd6e04f722a6a988ff694882

Observation 3afbd129-6a71-4eba-ac6e-3003fa257a68 · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37: 132208–132237, 2024.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37: 132208–132237, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.727265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:12:59.992721Z digest=sha256:2131370b03e67b0e1183cd073eac38409b6f698f45500ab62ecd960718cc88de

Observation 633bb555-fc2c-4003-9188-8f313ed67ce6 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.021058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.021058Z digest=sha256:a7665ddbcf1ff63d995cbe042e02774aa66b0aa5e6b88a8c77fda0ce0f11f5bb

Observation 062c7dff-3ce0-47f1-a21d-290ef93ee002 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.055272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.055272Z digest=sha256:056c638724fb84c4eb5af71a5bb200f3e134e15e304424141cd25b62bfc8f555

Observation 0a111599-c2f5-4e98-98f3-4d3ed6bff955 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.606895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.086857Z digest=sha256:7903fb193efaa0de802ec6152d89f500bc6899bcf255b67b3d09b1e9acc93c9d

Observation 1cbd8334-b5ec-46dd-9f99-393fcc307ac6 · outbound

This paper cites Address each question raised where relevant.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Address each question raised where relevant

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.462769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.122225Z digest=sha256:3a840e13a6499a9eac34a3dbafd67452b9e7346a0b3e0eb7e579a2028d67e459

Observation 671f858a-d84e-4cca-bd6e-490bd5d1cb56 · outbound

This paper cites Regardless of the approach, always conclude with: 25 Therefore, the final answer is: $\boxed{[answer]}$.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Regardless of the approach, always conclude with: 25 Therefore, the final answer is: $\boxed{[answer]}$

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.402514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.149962Z digest=sha256:458d57f33aeb0b9bc13a78789090c0c8dd7b8d4c115d8f9762ed996a752394e5

Observation add2d73d-c7af-4113-8882-af81644c6a15 · outbound

This paper cites - End the answer with: Therefore, the final answer is: $\boxed{[answer]}$.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs - End the answer with: Therefore, the final answer is: $\boxed{[answer]}$

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.222163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.212428Z digest=sha256:394677199f32910204d13fd8d3c31da4075a60ab3556cc8535683cafb13de0b1

Observation 5d7ebf72-172c-4525-afe0-5910b2999433 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.140011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.238645Z digest=sha256:5923911aa4ebafb31d687b337e3964b96ae0296690b3b7bea42a1583288d3bd0

Observation 53a71cad-ead2-4e6d-871b-a79d06efe8c2 · outbound

This paper cites These should be carefully reviewed for mistakes.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs These should be carefully reviewed for mistakes

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:02.058682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.261331Z digest=sha256:c765053ad03e0b777643e1782affe05196228bf6daa0cf6b58c47dbb77252b0b

Observation 26db6557-6743-4503-9159-44a94183d0c3 · outbound

This paper cites Wait, that doesn’t seem right.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Wait, that doesn’t seem right

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:13:01.989643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.284581Z digest=sha256:04013d29113bcda674b2c9f65c8770933f9d3e3a8c92d6ebd60fc0a38937017d

Observation f1997a2b-2b83-42f9-9fda-666bbccea7ed · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:02.326585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.320125Z digest=sha256:afa09d7237f2cb33784bc9cf1242ed40128a22cab742eadd50510835fc9acbdb

Observation 3cb29c6f-d52a-47f0-903e-2ddab44a79c9 · outbound

This paper cites an unresolved cited work.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:13:01.896301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:13:00.348063Z digest=sha256:7efd8a82cbb0ce3d49ac03fadd5a780269a4a1b11ff9c1df6b004b2051719d28

Observation 20202767-143f-4d82-9975-74f515d63946 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs Gemma 2: Improving Open Language Models at a Practical Size

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:59.556457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:59.556457Z digest=sha256:6428ffb9752ca50210e6f110a4b84de0aaada58ff0f724c21c1311897799cdf3

Pith citing papers

Observation 6fddc6f9-9dac-44d6-9fe6-1bb5a442007d · inbound

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection cites this paper.

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.554324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:31:44.815249Z digest=sha256:c590098354f24c390142b01dd74db518d7e71ed6aa3e7e22b2cec4a9bd617e1d

Observation 4b4c0f79-3c6a-4168-bb90-8a1d61ec4fc3 · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.758040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0521f03d62c0ec889092a82c3697e36fd6b017e5676a97a6408729f2bdd8eb57