Pith. sign in

REVIEW 15 cited by

More Agents Is All You Need

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05120 v2 pith:7S6Y425P submitted 2024-02-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords agentsllmsmethodagentagentforestavailablebenchmarkscode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We find that, simply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiated. Also, this method, termed as Agent Forest, is orthogonal to existing complicated methods to further enhance LLMs, while the degree of enhancement is correlated to the task difficulty. We conduct comprehensive experiments on a wide range of LLM benchmarks to verify the presence of our finding, and to study the properties that can facilitate its occurrence. Our code is publicly available at: https://github.com/MoreAgentsIsAllYouNeed/AgentForest

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 21 citations worldwide. Full citation record

  1. State-dependent error correlations shape voting thresholds in committees of AI agents

    cs.CY 2026-07 accept novelty 6.0 of 10

    Correlated errors among AI voters create an irreducible committee-error floor, and using state-dependent correlation estimates improves held-out k-of-n threshold selection.

  2. Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Diversity metrics used to select LLM ensembles are largely capability proxies; after control, only a modest pairwise co-failure association with majority-vote gain remains.

  3. Streaming Communication in Multi-Agent Reasoning

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    StreamMA introduces streaming communication in multi-agent reasoning to reduce latency via pipelining and improve effectiveness by leveraging reliable early steps, with closed-form analysis and a step-level scaling law.

  4. When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    A disagreement-guided routing framework dynamically selects among resolution, voting, and rewriting strategies for test-time scaling, delivering 3-7% accuracy gains with lower sampling cost on mathematical benchmarks.

  5. Effective Strategies for Asynchronous Software Engineering Agents

    cs.CL 2026-03 conditional novelty 6.0 of 10

    CAID, a manager-driven multi-agent system using git worktrees, commits, and merges, improves long-horizon SWE success by roughly 14–27 absolute points over single-agent baselines.

  6. When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines

    cs.MA 2026-03 conditional novelty 6.0 of 10

    In a 42-task controlled comparison, selecting the best candidate with judge panels beats MoA-style synthesis in every task, and a crossover threshold explains when team diversity helps.

  7. Tacit Coordination of Large Language Models

    cs.GT 2026-01 conditional novelty 6.0 of 10

    Across 20+ open-source LLMs, tacit coordination in focal-point games is often at or above human levels, with systematic failures on cultural and numerical salience that culture prompts partially fix.

  8. Efficient Leave-one-out Approximation in LLM Multi-agent Debate Based on Introspection

    cs.MA 2025-05 reject novelty 6.0 of 10

    IntrospecLOO uses a single extra prompting round to approximate leave-one-out contribution in LLM debates, but the empirical evidence is weak and one case study contradicts the method's claimed behavior.

  9. Transition from Statistical to Hardware-Limited Scaling in Photonic Quantum State Reconstruction

    quant-ph 2026-03 unverdicted novelty 5.0 of 10

    Classical shadow tomography on integrated photonics shows a sharp transition from statistical O(M^{-1/2}) error scaling to a hardware-limited floor set by unitary spectral distortions.

  10. Integrating Traditional Technical Analysis with AI: A Multi-Agent LLM-Based Approach to Stock Market Forecasting

    cs.CE 2025-06 reject novelty 5.0 of 10

    An LLM-based multi-agent system that identifies Elliott Wave patterns achieved 44-89% directional accuracy on six US stocks, with deep reinforcement learning backtesting improving accuracy on most but not all test cases.

  11. An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring

    cs.MA 2025-05 conditional novelty 5.0 of 10

    A credibility-scoring framework for multi-agent LLM systems, learning agent trustworthiness on the fly and weighting outputs accordingly, improves accuracy under adversarial conditions in some benchmarks.

  12. SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection

    cs.SE 2025-05 conditional novelty 5.0 of 10

    SecVulEval provides a statement-level C/C++ vulnerability benchmark with context; state-of-the-art LLMs achieve only 23.83% F1 on locating vulnerable statements with correct reasoning.

  13. Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

    cs.CR 2025-10 conditional novelty 4.0 of 10

    A backward-propagation scoring scheme over a signed temporal DAG can identify malicious agents in LLM multi-agent systems and cut their communications, improving defended accuracy by 3–7 percentage points in the autho...

  14. Token-Operations-Oriented Inference Optimization Techniques for Large Models

    cs.SE 2026-06 unverdicted novelty 3.0 of 10

    The paper introduces a four-layer technical architecture for token-operations-oriented inference optimization in large models and reviews key technologies and industry status at each layer.

  15. ElliottAgents: A Natural Language-Driven Multi-Agent System for Stock Market Analysis and Prediction

    cs.CE 2025-07 reject novelty 3.0 of 10

    A natural-language multi-agent system combines Elliott Wave pattern detection with LLM dialogue and DRL backtesting to generate stock predictions, supported only by small selected experiments.

Pith tools