Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2502.15840.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:27.304146Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation e917c6d9-ae85-4aba-8155-8e4637ba99a5 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4ffdfc9-7646-4a7d-b84c-26b6d3f8d008 · inbound
PyVision: Agentic Vision with Dynamic Tooling Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4877a63d-1dff-4ce2-b53b-1d5ed6dc5a8d · inbound
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32327f00-ae59-4aab-aae6-910381a70a9f · inbound
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c124a8f3-92c1-4ae3-b4e8-b0464d687c5b · inbound
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee1af49-dc9c-4beb-9424-7cf3aa09b1e8 · inbound
LLM-SAA: LLM-persona Generated Distributions for Decision-making Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9329eed0-28e2-4aaf-9c6b-9b73b2a1c756 · inbound
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80edd497-0d36-49de-9428-7ab2ca384f26 · inbound
GLM-5: from Vibe Coding to Agentic Engineering Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5035e18-9b49-4d02-810c-bb368d148b38 · inbound
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d12f98-1c10-4ba0-8fd5-7f5d50dd68f3 · inbound
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed5787e0-d2e7-485d-893e-4613887c86c6 · inbound
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44113a13-f4f7-4fcf-af13-fdf52409b9ab · inbound
NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e7f2a29-2cc6-41d3-9f64-8cf0c5e25773 · inbound
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 094ade2b-8ef5-4766-b959-46212130cab2 · inbound
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcfb4317-89f8-4629-8821-98b268dfbbcf · inbound
Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f937a47-e7c0-46c2-9d87-c97643217658 · inbound
CL-bench Life: Can Language Models Learn from Real-Life Context? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5b9b71b-94e5-4585-bf4d-7f13bc5405a0 · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf28f79f-6def-4060-99f5-446639afa09f · inbound
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ad775ec-7465-4ab4-8e7b-85ec66537527 · inbound
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33b4960e-f1b0-44ee-9365-1bb344d6df77 · inbound
CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee76460d-f7df-49d9-a7a4-fdc5bcfc724e · inbound
CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76bf20e-fa30-4e2f-83b3-ea5bd853df8e · inbound
Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0746d08a-99fe-4367-bc8d-99bb79d794f7 · inbound
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51fbd126-8dbb-4a01-ab59-49c10fb4055d · inbound
Graph-Enhanced Large Language Models for Spatial Search Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3526cecc-8a5a-425d-b370-31f8a8788bdf · inbound
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d05376fd-dcb6-4499-8295-d0052e97d8be · inbound
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3dff5a-2ea5-48cd-8372-785d5005be33 · inbound
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · inbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 854d0ebe-bbf0-4d65-87de-47cbcbb6674e · inbound
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.