Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:57:38.586633Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 210 outbound references and 17 inbound Pith citation observations for arXiv:2508.13167.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:57:38.586633Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:37.907213Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
100 of 210 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 03a0143b-e32f-42b7-8b9a-3f79642a9777 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Towards Effective Code-Integrated Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec18274-6b2c-4ac8-8320-231110c0218b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Multi-agent reinforcement learning: A review of challenges and applications
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9241b2-eeb8-4ea6-b3a8-f5ced24ab1b6 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78de714f-8f74-4c57-b429-3147a784ba24 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Process Reinforcement through Implicit Rewards
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f938d6e4-c9ab-470b-bbbe-fa16b4fd35a4 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Multi-agent systems: A survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2afc82b-2a1a-495a-a462-851c017dea1c · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c323eb-4fe7-4925-8e5c-cd209bf84eb6 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ebec6f0-8c81-4910-8bcf-7fee5a73e4fa · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446c71b1-1cc7-47ac-bd20-23c2c81e9da5 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL How we built our multi-agent research system
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b0014d-7d06-4d5c-aa20-005f30600bee · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13b30e1-d8ff-4535-975a-5062af49d861 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Skywork Open Reasoner 1 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85af062-d1d0-499d-b050-0154e419afa7 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f539eb-1a58-4bed-a208-5a9535f3dc05 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0fcd6a6-61df-4133-afe5-f7e729f6796b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569bcbb2-9b69-42af-8c16-e24c18f1e10a · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwen2.5-Coder Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba5575f-73c2-4c8c-b042-4dddfe5900d2 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7cd166-2d4f-424c-b04a-ff638f752448 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56debdf8-3f47-45cf-abac-8b59149eb863 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e843709-2b3a-4e48-b8a1-38af653cd782 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f375e221-b36a-49db-b778-fcbd77e91648 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Reveal: Self-evolving code agents via iterative generation-verification, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b88e872-4291-459d-b147-7b7809f15a36 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491219c8-5731-4cf3-9ead-c526dd628817 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Sequence-level knowledge distillation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d480276-2117-48db-bed9-e2b3747df3f8 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Natural questions: a benchmark for question answering research
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6955b5-0238-40d6-8a10-fabf57a2607e · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Camel: Communicative agents for "mind" exploration of large language model society
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35eff444-9688-4119-b2fc-d4120c69eaff · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45b9d34-84d0-4829-ab88-b378f61b7e81 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7497d46-9b24-4489-a3a3-658154dbf6b9 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0824de45-a673-4ba2-a2a3-3a113de407a4 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToRL: Scaling Tool-Integrated RL
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d780e1ec-75b0-4e6b-a57a-888b514102d9 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Competition-level code generation with alphacode
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af51948-67fd-470a-bb5f-24d2555930fa · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Let's Verify Step by Step
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56168054-a414-4018-b0d5-23cf642948e2 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Inference-time scaling for generalist reward modeling
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daeccf5a-4e54-4b1d-99e2-96006c7d481b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Decoupled Weight Decay Regularization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b74efe-b908-4be7-8a9d-f20ce2219763 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9526066-f0ec-4fa6-9041-eb7abad91742 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3dda23f-b55c-45f0-a2ee-e71e1aa333ff · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Gaia: a benchmark for general ai assistants
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f83a6e-a7c5-4102-801f-ee5284d61ecb · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL American invitational mathematics examination (aime) 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e45e29-ac31-4ee9-a15a-c1cf165fd34a · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL American invitational mathematics examination (aime) 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd56ec39-9d2e-4dd0-9d70-fd0e6f2ab48b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Codeforces
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57664f53-66ae-40ee-bf40-62956e47ff95 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Humanity's Last Exam
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce665dea-d3f5-4d73-860c-e3fdff5debb1 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Measuring and Narrowing the Compositionality Gap in Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7c568a-a605-4617-984f-dbfdab6eaa81 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToolRL: Reward is All Tool Learning Needs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2447ea-9700-4597-8a65-008dda813d42 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05afe5d6-a392-4423-9300-045722a72b91 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwen2.5 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0b6d47-7d03-47a1-b67e-6530e49c7431 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ‘smolagents‘: a smol library to build great agentic systems.https://github.com/huggingface/smolagents, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3067b16b-3f58-4d30-8537-c6cba9e32084 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21dd24e-8f78-4877-99fa-aac28be7692b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL HybridFlow: A Flexible and Efficient RLHF Framework
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6813e7ef-1462-4ad8-84b5-999c92740970 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL TaskCraft: Automated Generation of Agentic Tasks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad93aa0b-f9c9-45de-b47f-7ae4d9c8160f · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5138e2-3aec-4a06-852f-36d9df1b15c8 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e531ab-4bc4-429e-996f-fb502cc01dd8 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b63bb7a-bc4c-493c-a671-6eb18354746d · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agent kb: Leveraging cross-domain experience for agentic problem solving
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fc01ef-b2d6-40a4-bf48-752f4cd7f1bf · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997d1e1f-332b-40b3-8ef6-f4b208246bd0 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e68af04-f22e-4ed1-9c4a-d3c81c4ded6d · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Verl-tool: A version of verl to support tool use, 2025
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd42928b-5df2-4c13-9dca-a8e7060461e5 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978d58b2-0f67-45da-9586-f4a534701f94 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Musique: Multihop questions via single-hop question composition
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91c3373-7184-4626-9b61-c92451bd1110 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Otc: Optimal tool calls via reinforcement learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a1e444-233c-42af-bd0b-a92e31b51857 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff451199-ebf8-41dd-a9cf-d0eae2555210 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Corag: A cost-constrained retrieval optimization system for retrieval-augmented generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e9dadf-6486-4031-b452-4f976bd78665 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Chain-of-thought prompting elicits reasoning in large language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d45d17-24ac-4d7b-9ab3-b79be03868be · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2f730c-502f-49ba-afd5-8c9677e23336 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5141165-c921-4a66-8436-6ff4afbefcc0 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL WebDancer: Towards Autonomous Information Seeking Agency
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad09509b-22a5-476b-9da2-a04b43cfa779 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Simpletir: End-to-end reinforcement learning for multi-turn tool-integrated reasoning.https://simpletir.notion.site/report, 2025
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265c1ce6-2293-4363-80f6-774048b00e63 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4dd1e80-a7f9-410c-8c91-6a50e0b59175 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL React: Synergizing reasoning and acting in language models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7267780-6bf5-4d8c-954b-797bca6902e2 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55570e4e-d987-4a9d-890f-b5f8c7a3f39c · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Auto-RAG: Autonomous Retrieval-Augmented Generation for Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5acf09dd-76c4-4cb9-8a9a-2a43bec7bf88 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976deefd-64fb-4a82-a025-15911957bbc9 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Flowmind: automatic workflow generation with llms
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cecdfa7-9a73-4d04-8df6-fbb414618969 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL EvolveSearch: An Iterative Self-Evolving Search Agent
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdecd40-5aca-4a54-9bbc-091346685d9a · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL AFlow: Automating Agentic Workflow Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e140002-0657-44be-bb0f-5d16e054be44 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Process vs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab162f6-2a18-4c07-9bd7-a58fc8dc8939 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594b156b-bb91-45a3-a408-e3da399ca22e · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0758f8d-b297-46e9-8cc1-d30b5fcd370d · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OpenResearcher: Unleashing AI for Accelerated Scientific Research
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e83de9-b3f6-4826-8813-6dc0726be362 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e66a4e0-1668-4f1f-8d1c-035c6775baf3 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agents: An Open-source Framework for Autonomous Language Agents
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d81235d-624b-42cd-af34-ed281e30b754 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Symbolic Learning Enables Self-Evolving Agents
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea47f24-507f-4a7c-aacb-ab5965079b03 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL OAgents: An Empirical Study of Building Effective Agents
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc80da3-53d8-4761-b32d-2fd4cf67c913 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Scaling Test-time Compute for LLM Agents
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692098ca-827b-401b-8f0c-3d73f7049b76 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249683ff-7b58-4757-983b-819eff462a8a · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd27c36-f15d-4f28-8e14-c1db3961abfa · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60e1ba7-42e8-4d3f-a083-84864fc73b46 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873d6a4c-b7d7-498f-ba4f-5dfc30521fd8 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905295ae-3d75-4b20-9f22-39e4d96e6773 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0919bf46-5b23-41f7-90d4-74524349be45 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL An archive of all existing APOD pages (current date through
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8e6feb-dc51-4b6c-aead-2bb9fbc372e6 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06be36e4-0a7b-4720-932d-6f6bc3e8dac4 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b93abf2-062e-4341-a327-adacede45e6c · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL 1, 2015 (Credit: NASA/Bill Ingalls)
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e796f871-e906-41de-aff2-01168f04413e · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL </observation> Step 3 <think> Step 1 of the task is to identify the NASA Astronomy Picture of the Day (APOD) from the first week of August 2015 showing city lights on the horizon
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36753cd6-d66e-4f7a-bae2-1df7277ade9b · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Marquette had a population of 20,629 at the
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43afb520-cf77-4d5f-8a29-315a88ecbeff · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856fc924-0d4e-4a7a-a21b-72606be3a7a2 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c50ade7-1dde-42f3-8be2-eaef972b8c75 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL “Back in the 1600’s he set up several missions, including
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51893399-7f40-4f42-8504-10392a842094 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e14bd8-408b-471c-ba65-a08968c95961 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Completed in 1894, the Marquette Building brings Chicago’s early history to life in an artistic and elegant setting
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c25442-ad78-4c32-9ada-b5f4885910b9 · outbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Unresolved cited work
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdc812e-5f22-4df7-b1b8-81b57d114aa0 · inbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7eaf6c-332a-4b96-a4a8-af5f03ecb7d2 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f01437bd-012f-4fca-8e25-4f829e0cd832 · inbound
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54dbf78-fc94-44f5-8dbc-82d62a9e7f40 · inbound
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fcd7cc45-b888-4cd7-b8ac-6651321df2f3 · inbound
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 45d13e58-7e23-4823-89be-325b8f3bb308 · inbound
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c6f6425-b17e-4b45-aa4c-911cbb85099b · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61c22ce0-a9eb-4916-b90d-779314c5c6a1 · inbound
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c4865a5-0294-4186-b4c5-e2124955073d · inbound
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e23cc51-8da2-4836-a10b-219e68569a74 · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee513b98-3b27-4c01-8302-c504f50af3bd · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b11d48-501a-4101-8a2e-0be63c69dd0f · inbound
Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1c0f03c-60f7-42d7-9d7c-12631d1e739a · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d80bed98-4296-410e-bd21-44b6724818ba · inbound
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7dae01a7-445c-4944-b479-069ec03d4324 · inbound
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9aa5ba37-f362-4f0e-8428-9d85470c07af · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 262
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b1b9892c-b075-46dc-8e03-96c7a9adf5cc · inbound
Mathematical methods of reinforcement learning Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.