REVIEW 15 cited by
More Agents Is All You Need
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We find that, simply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiated. Also, this method, termed as Agent Forest, is orthogonal to existing complicated methods to further enhance LLMs, while the degree of enhancement is correlated to the task difficulty. We conduct comprehensive experiments on a wide range of LLM benchmarks to verify the presence of our finding, and to study the properties that can facilitate its occurrence. Our code is publicly available at: https://github.com/MoreAgentsIsAllYouNeed/AgentForest
Forward citations
Cited by 15 Pith papers
-
State-dependent error correlations shape voting thresholds in committees of AI agents
Correlated errors among AI voters create an irreducible committee-error floor, and using state-dependent correlation estimates improves held-out k-of-n threshold selection.
-
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
Diversity metrics used to select LLM ensembles are largely capability proxies; after control, only a modest pairwise co-failure association with majority-vote gain remains.
-
Streaming Communication in Multi-Agent Reasoning
StreamMA introduces streaming communication in multi-agent reasoning to reduce latency via pipelining and improve effectiveness by leveraging reliable early steps, with closed-form analysis and a step-level scaling law.
-
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling
A disagreement-guided routing framework dynamically selects among resolution, voting, and rewriting strategies for test-time scaling, delivering 3-7% accuracy gains with lower sampling cost on mathematical benchmarks.
-
Effective Strategies for Asynchronous Software Engineering Agents
CAID, a manager-driven multi-agent system using git worktrees, commits, and merges, improves long-horizon SWE success by roughly 14–27 absolute points over single-agent baselines.
-
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
In a 42-task controlled comparison, selecting the best candidate with judge panels beats MoA-style synthesis in every task, and a crossover threshold explains when team diversity helps.
-
Tacit Coordination of Large Language Models
Across 20+ open-source LLMs, tacit coordination in focal-point games is often at or above human levels, with systematic failures on cultural and numerical salience that culture prompts partially fix.
-
Efficient Leave-one-out Approximation in LLM Multi-agent Debate Based on Introspection
IntrospecLOO uses a single extra prompting round to approximate leave-one-out contribution in LLM debates, but the empirical evidence is weak and one case study contradicts the method's claimed behavior.
-
Transition from Statistical to Hardware-Limited Scaling in Photonic Quantum State Reconstruction
Classical shadow tomography on integrated photonics shows a sharp transition from statistical O(M^{-1/2}) error scaling to a hardware-limited floor set by unitary spectral distortions.
-
Integrating Traditional Technical Analysis with AI: A Multi-Agent LLM-Based Approach to Stock Market Forecasting
An LLM-based multi-agent system that identifies Elliott Wave patterns achieved 44-89% directional accuracy on six US stocks, with deep reinforcement learning backtesting improving accuracy on most but not all test cases.
-
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
A credibility-scoring framework for multi-agent LLM systems, learning agent trustworthiness on the fly and weighting outputs accordingly, improves accuracy under adversarial conditions in some benchmarks.
-
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
SecVulEval provides a statement-level C/C++ vulnerability benchmark with context; state-of-the-art LLMs achieve only 23.83% F1 on locating vulnerable statements with correct reasoning.
-
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
A backward-propagation scoring scheme over a signed temporal DAG can identify malicious agents in LLM multi-agent systems and cut their communications, improving defended accuracy by 3–7 percentage points in the autho...
-
Token-Operations-Oriented Inference Optimization Techniques for Large Models
The paper introduces a four-layer technical architecture for token-operations-oriented inference optimization in large models and reviews key technologies and industry status at each layer.
-
ElliottAgents: A Natural Language-Driven Multi-Agent System for Stock Market Analysis and Prediction
A natural-language multi-agent system combines Elliott Wave pattern detection with LLM dialogue and DRL backtesting to generate stock predictions, supported only by small selected experiments.
Discussion (0). Sign in to comment.