Pith. sign in

Paper Citation Record · LEDGER

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2502.15840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.15840 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:27.304146Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e917c6d9-ae85-4aba-8155-8e4637ba99a5 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.485174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:c575d9853567bd4313e25b15caeba6884ed2e7becbd1a436d60b3590c08e6b2c

Observation b4ffdfc9-7646-4a7d-b84c-26b6d3f8d008 · inbound

PyVision: Agentic Vision with Dynamic Tooling cites this paper.

PyVision: Agentic Vision with Dynamic Tooling Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:27.304146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:27.304146Z digest=sha256:7f93aa923ddabd8811034bdb1bc4bff9ce357d02fc0af29b7a9cc88171ec1906

Observation 4877a63d-1dff-4ce2-b53b-1d5ed6dc5a8d · inbound

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format cites this paper.

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:10.799080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:10.799080Z digest=sha256:2b61df7ef6a2fa4543d59e5f610136cea444894cae57cf68394b0f18b18b1875

Observation 32327f00-ae59-4aab-aae6-910381a70a9f · inbound

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare cites this paper.

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:55.517898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:55.517898Z digest=sha256:9f2348175344bbb443b2bffa7334c37c6d81e4f695c17b16a395b5dabbbfbd88

Observation c124a8f3-92c1-4ae3-b4e8-b0464d687c5b · inbound

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions cites this paper.

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:35.290693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:20:35.290693Z digest=sha256:df96e67cbae2c04db1e6836b50f00e1dfcd75a5179bd1eba1016ce78bd1f0dfe

Observation 7ee1af49-dc9c-4beb-9424-7cf3aa09b1e8 · inbound

LLM-SAA: LLM-persona Generated Distributions for Decision-making cites this paper.

LLM-SAA: LLM-persona Generated Distributions for Decision-making Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:46.074199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:46.074199Z digest=sha256:2cc137cecdf1d61eb07af980e6c3959a409e4256423497aa8083aae96e0b0663

Observation 9329eed0-28e2-4aaf-9c6b-9b73b2a1c756 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:57:24.662243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:2f297fed55b70be3e92e292dedb5e507ed1b2a4d3478a714a701759f3453bb2c

Observation 80edd497-0d36-49de-9428-7ab2ca384f26 · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:41.129911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:c3e41245abdab9b69d931ffbd27ca1b4e852e30e5245aef7ae95ba8b247f8d6f

Observation b5035e18-9b49-4d02-810c-bb368d148b38 · inbound

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents cites this paper.

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T15:09:29.436835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:09:29.436835Z digest=sha256:c4300353acc4dc57f81d62b4622d05d71cef1fce107a58696712ca6b3a46f1f6

Observation 70d12f98-1c10-4ba0-8fd5-7f5d50dd68f3 · inbound

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models cites this paper.

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:03:20.367050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:02:12.555644Z digest=sha256:037a68da9be4b8c65ca52137f144c335f2de0c852f771ad6d812b7252f16651e

Observation ed5787e0-d2e7-485d-893e-4613887c86c6 · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.844569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:faa0d4d790be7f81ee2ac77d9bdf1cecbe180b39a73dd40e8e6564d15b01ca92

Observation 44113a13-f4f7-4fcf-af13-fdf52409b9ab · inbound

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration cites this paper.

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:53:08.341871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:49:15.486503Z digest=sha256:7be28f79fb8be8b7633ac435df530583aacd62041f5f1845f846fd95d2b1f5a5

Observation 5e7f2a29-2cc6-41d3-9f64-8cf0c5e25773 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.701078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:27:02.965559Z digest=sha256:879216a22a18c4ea66eb3ef2a4a6cebfcc08ec7eb3e6c5bba33a06b4e9927ef8

Observation 094ade2b-8ef5-4766-b959-46212130cab2 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T19:45:34.692776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:45:34.692776Z digest=sha256:aa323c8d3659038b3190546b8280b26ce60125e46e3ce5d39ceb9139b9f54c83

Observation fcfb4317-89f8-4629-8821-98b268dfbbcf · inbound

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations cites this paper.

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:07.696827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:14:01.097146Z digest=sha256:f49a63c498e50d098d5690fde5f619ec1f9a3e671921d7581c072c8ed2e886d7

Observation 7f937a47-e7c0-46c2-9d87-c97643217658 · inbound

CL-bench Life: Can Language Models Learn from Real-Life Context? cites this paper.

CL-bench Life: Can Language Models Learn from Real-Life Context? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:27.025477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T09:42:33.635866Z digest=sha256:3077ac858786e05e3a39e4e6824e2980392ca7674c0e774be4661ce839df15ff

Observation f5b9b71b-94e5-4585-bf4d-7f13bc5405a0 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:47.873226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:e11667fbe16896a1127eb1f9ffe959c469ca77400eee76da7611b14188d11ce0

Observation cf28f79f-6def-4060-99f5-446639afa09f · inbound

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces cites this paper.

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:28:19.054534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:25:48.728238Z digest=sha256:da78904a9b4f78812bef37518b3d3a7d9a92cd5ef92c64e6fa7fa67b2756a6ea

Observation 5ad775ec-7465-4ab4-8e7b-85ec66537527 · inbound

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use cites this paper.

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:16:00.246835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:03:36.851403Z digest=sha256:1925c80d3ee47d2be519562260be723325e4affdf8699cac38d7b5bc12c716b7

Observation 33b4960e-f1b0-44ee-9365-1bb344d6df77 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.693981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:b9680eb28befd1aaf644795a4b392c48b3cdc0ed21b581294a248e1c7b9fa514

Observation ee76460d-f7df-49d9-a7a4-fdc5bcfc724e · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:16.111772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:16.111772Z digest=sha256:c4144b0e500bf5261359095ffc84361e7304bb95188be9d2e5e876a14f366eab

Observation a76bf20e-fa30-4e2f-83b3-ea5bd853df8e · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.490566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:f284d579eb22e8f8e181eb27dcaefc09df7e33350469572d45246503e70c7a72

Observation 0746d08a-99fe-4367-bc8d-99bb79d794f7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.915566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:ed3284b43df3b8d9b5d58656ed8e7d3e012265920d84a66c87d4aed94b8a5e6d

Observation 51fbd126-8dbb-4a01-ab59-49c10fb4055d · inbound

Graph-Enhanced Large Language Models for Spatial Search cites this paper.

Graph-Enhanced Large Language Models for Spatial Search Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T06:46:29.757188Z digest=sha256:880876795a1355556d39271fec6d420e9b4afd1bb59927f4dc094d19a20c1583

Observation 3526cecc-8a5a-425d-b370-31f8a8788bdf · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.678930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:f9d3f4d574e0f761c5a1d0c3384a8d6ddd88a241f1275167bb3ac42da0a25f99

Observation d05376fd-dcb6-4499-8295-d0052e97d8be · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:4843c90972cc2ae2be8667be1e4bd5fabfbd565369553e22ce4246104b4cb3b9

Observation 1e3dff5a-2ea5-48cd-8372-785d5005be33 · inbound

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle cites this paper.

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:05.920354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:05.920354Z digest=sha256:924d4e879c4804e90fca74d1e33f2a77e7d9113f7db793b59f8e417d65967a89

Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · inbound

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios cites this paper.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.686510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.686510Z digest=sha256:1842b897c32c76489e7b4a8e00fee7403beb2695fc2248af17a35f1750f51528

Observation 854d0ebe-bbf0-4d65-87de-47cbcbb6674e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.191363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.191363Z digest=sha256:e761de4eaeaceb6902cc77df4ea90dfcc92ec76c887fb0d072f7570e147cc498