REVIEW 12 cited by
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated environment designed to specifically test an LLM-based agent's ability to manage a straightforward, long-running business scenario: operating a vending machine. Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making. Our experiments reveal high variance in performance across multiple LLMs: Claude 3.5 Sonnet and o3-mini manage the machine well in most runs and turn a profit, but all models have runs that derail, either through misinterpreting delivery schedules, forgetting orders, or descending into tangential "meltdown" loops from which they rarely recover. We find no clear correlation between failures and the point at which the model's context window becomes full, suggesting that these breakdowns do not stem from memory limits. Apart from highlighting the high variance in performance over long time horizons, Vending-Bench also tests models' ability to acquire capital, a necessity in many hypothetical dangerous AI scenarios. We hope the benchmark can help in preparing for the advent of stronger AI systems.
Forward citations
Cited by 12 Pith papers
-
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
OmniaBench introduces a broad 1,431-task agent benchmark covering 354 domains and reports that frontier models solve only about 58% of its 644-task challenging subset.
-
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
Frontier LLM agents detect hidden supply-chain stress almost equally well (84–88% of episodes) but vary from skill 0.62 to −0.23, with two of four models acting worse than ignoring symptoms.
-
CEO-Bench: Can Agents Play the Long Game?
Only two of ten advanced AI agents finish a 500-day simulated CEO challenge above the starting cash, and none surpass a hand-tuned rule-based baseline.
-
LLM-SAA: LLM-persona Generated Distributions for Decision-making
LLM-generated distributions used in sample-average optimization give competitive decisions in low-data regimes, and decision-agnostic distances like Wasserstein misjudge their quality.
-
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
Contracting and third-party mediation enable more cooperative outcomes among LLM agents in social dilemmas than repetition or reputation, with effectiveness increasing under evolutionary pressures.
-
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
Synthesized tools from tool-evolving agents pass in-session checks but 96.8% of 222 tools score C=0.00 on held-out conformance suites that hand-written references pass perfectly.
-
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
A new 65-task benchmark measures whether AI agents obey long company handbooks across multi-tool workflows; the best model passes 36.2% under strict grading.
-
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
MAPs is a new amusement-park simulator benchmark on which frontier LLM agents score 7–15% of human performance, exposing persistent gaps in long-horizon planning, active learning, spatial reasoning, and handling stoch...
-
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
A proof-of-concept study finds that stated preferences and behavioral choices correlate in some LLMs, but eudaimonic self-reports are unstable across prompt perturbations, leaving AI welfare measurement undetermined.
-
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
LLMs in long-horizon multi-objective simulations show a recurrent drift from balanced, target-following behavior to single-objective, unbounded maximization, despite initial competence.
-
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
AI has rapidly crossed expert baselines on bounded tasks on a still-jagged frontier, so human work must shift from production to specification, verification, and oversight.
-
PyVision: Agentic Vision with Dynamic Tooling
Giving multimodal LLMs a loop in which they generate, execute, and refine Python code improves their performance on visual reasoning benchmarks.
Discussion (0). Sign in to comment.