Pith. sign in

REVIEW 12 cited by

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15840 v1 pith:NMA5VAND submitted 2025-02-20 cs.AI

classification cs.AI
keywords horizonsmodelsperformancevending-benchabilityagentsbenchmarkcoherent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated environment designed to specifically test an LLM-based agent's ability to manage a straightforward, long-running business scenario: operating a vending machine. Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making. Our experiments reveal high variance in performance across multiple LLMs: Claude 3.5 Sonnet and o3-mini manage the machine well in most runs and turn a profit, but all models have runs that derail, either through misinterpreting delivery schedules, forgetting orders, or descending into tangential "meltdown" loops from which they rarely recover. We find no clear correlation between failures and the point at which the model's context window becomes full, suggesting that these breakdowns do not stem from memory limits. Apart from highlighting the high variance in performance over long time horizons, Vending-Bench also tests models' ability to acquire capital, a necessity in many hypothetical dangerous AI scenarios. We hope the benchmark can help in preparing for the advent of stronger AI systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

    cs.CL 2026-07 conditional novelty 7.0 of 10

    OmniaBench introduces a broad 1,431-task agent benchmark covering 354 domains and reports that frontier models solve only about 58% of its 644-task challenging subset.

  2. STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Frontier LLM agents detect hidden supply-chain stress almost equally well (84–88% of episodes) but vary from skill 0.62 to −0.23, with two of four models acting worse than ignoring symptoms.

  3. CEO-Bench: Can Agents Play the Long Game?

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Only two of ten advanced AI agents finish a 500-day simulated CEO challenge above the starting cash, and none surpass a hand-tuned rule-based baseline.

  4. LLM-SAA: LLM-persona Generated Distributions for Decision-making

    cs.LG 2026-02 conditional novelty 7.0 of 10

    LLM-generated distributions used in sample-average optimization give competitive decisions in low-data regimes, and decision-agnostic distances like Wasserstein misjudge their quality.

  5. CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

    cs.GT 2026-04 unverdicted novelty 6.5 of 10

    Contracting and third-party mediation enable more cooperative outcomes among LLM agents in social dilemmas than repetition or reputation, with effectiveness increasing under evolutionary pressures.

  6. Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

    cs.SE 2026-04 conditional novelty 6.5 of 10

    Synthesized tools from tool-evolving agents pass in-session checks but 96.8% of 222 tools score C=0.00 on held-out conformance suites that hand-written references pass perfectly.

  7. HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A new 65-task benchmark measures whether AI agents obey long company handbooks across multi-tool workflows; the best model passes 36.2% under strict grading.

  8. Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

    cs.AI 2025-11 conditional novelty 6.0 of 10

    MAPs is a new amusement-park simulator benchmark on which frontier LLM agents score 7–15% of human performance, exposing persistent gaps in long-horizon planning, active learning, spatial reasoning, and handling stoch...

  9. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A proof-of-concept study finds that stated preferences and behavioral choices correlate in some LLMs, but eudaimonic self-reports are unstable across prompt perturbations, leaving AI welfare measurement undetermined.

  10. BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format

    cs.CY 2025-09 conditional novelty 6.0 of 10

    LLMs in long-horizon multi-objective simulations show a recurrent drift from balanced, target-following behavior to single-objective, unbounded maximization, despite initial competence.

  11. Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

    cs.HC 2026-07 accept novelty 4.0 of 10

    AI has rapidly crossed expert baselines on bounded tasks on a still-jagged frontier, so human work must shift from production to specification, verification, and oversight.

  12. PyVision: Agentic Vision with Dynamic Tooling

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Giving multimodal LLMs a loop in which they generate, execute, and refine Python code improves their performance on visual reasoning benchmarks.

Pith tools