Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:35.262071Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 9 inbound Pith citation observations for arXiv:2412.21033.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:35.262071Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:54.275568Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T05:51:24.590137Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6899f4e-af3b-4b1c-9a30-0f0ebda2d7fd · outbound
Plancraft: an evaluation dataset for planning with LLM agents Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a6e7a3-bdb3-4633-8b20-6b3202ab33c4 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Video pretraining (vpt): Learning to act by watching unlabeled online videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ffdfbd1e-9f5e-424c-ae64-f3751cd43797 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Baby AI : First steps towards grounded language learning with a human in the loop
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb9af781-5bdc-42d0-b432-8c2ee8eca8ab · outbound
Plancraft: an evaluation dataset for planning with LLM agents Mind2web: Towards a generalist agent for the web, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcaafdb6-5669-4b4d-b95e-669cdd2d95a4 · outbound
Plancraft: an evaluation dataset for planning with LLM agents The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ddbc39-c668-4e58-a37b-218ed7e79878 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Minedojo: Building open-ended embodied agents with internet-scale knowledge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d2ea49-d04b-4549-9ab2-085fda3ae948 · outbound
Plancraft: an evaluation dataset for planning with LLM agents MineRL: A Large-Scale Dataset of Minecraft Demonstrations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef094eb-3095-4944-b64a-279ead6ec346 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Exploring the capacity of pretrained language models for reasoning about actions and change
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69b3124f-1bbe-4667-b15a-6e8e7662b8c8 · outbound
Plancraft: an evaluation dataset for planning with LLM agents The fast downward planning system
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db0f65c-f67f-42e7-b71b-22a5e3cfb518 · outbound
Plancraft: an evaluation dataset for planning with LLM agents LoRA: Low-Rank Adaptation of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339902b5-51f6-4cc1-b3e5-ead8166a9eaa · outbound
Plancraft: an evaluation dataset for planning with LLM agents Inner Monologue: Embodied Reasoning through Planning with Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78bec82-42db-4698-88be-c9715c2f0f62 · outbound
Plancraft: an evaluation dataset for planning with LLM agents The malmo platform for artificial intelligence experimentation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334b2164-3887-47e6-86c4-be78344f1904 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Position: LLM s can’t plan, but can help planning in LLM -modulo frameworks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d9561c3-a4ab-45b1-a64c-9e52a305d5e1 · outbound
Plancraft: an evaluation dataset for planning with LLM agents ACPB ench hard: Unrestrained reasoning about action, change, and planning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 078b8185-56bd-4951-89f5-cc1b9010685b · outbound
Plancraft: an evaluation dataset for planning with LLM agents Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c072f7f4-658f-4ecb-99b7-25cf3fda2706 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Benchmarking Detection Transfer Learning with Vision Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555f29c1-caf5-44a2-a03e-2c87edf7aac0 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a246f7b0-38bf-4ab1-b99a-10f34eb346cf · outbound
Plancraft: an evaluation dataset for planning with LLM agents AgentBench: Evaluating LLMs as Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962a9973-7623-474f-9fe5-6d44f7cd8606 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Howe, Craig A
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf260aa-6547-40dc-8624-10217e91f32f · outbound
Plancraft: an evaluation dataset for planning with LLM agents GAIA: a benchmark for General AI Assistants
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d5c23b-1742-4ff0-a903-716ee56be25e · outbound
Plancraft: an evaluation dataset for planning with LLM agents WebGPT: Browser-assisted question-answering with human feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb18fd8-7ebc-4a8c-a102-06938619207b · outbound
Plancraft: an evaluation dataset for planning with LLM agents GPT-4 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ca0f0e-10e2-4c33-9437-895551410b4e · outbound
Plancraft: an evaluation dataset for planning with LLM agents TALM: Tool Augmented Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2865db62-cda1-4847-8197-ae4521bf585c · outbound
Plancraft: an evaluation dataset for planning with LLM agents Gorilla: Large Language Model Connected with Massive APIs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c53212ee-6663-47e8-883b-bfc414e84533 · outbound
Plancraft: an evaluation dataset for planning with LLM agents The Multi-Agent Reinforcement Learning in Malm\"O (MARL\"O) Competition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ceb89d-cca3-41c6-a076-5178223b2734 · outbound
Plancraft: an evaluation dataset for planning with LLM agents ADaPT: As-Needed Decomposition and Planning with Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1152107-ef53-4331-a093-fcc67dabf01a · outbound
Plancraft: an evaluation dataset for planning with LLM agents Virtualhome: Simulating household activities via programs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aff9efb-81e9-47f5-8a39-58823c8b1925 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Toolformer: Language models can teach themselves to use tools
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 116c7831-baaf-4084-a3dd-af5a46b4c68a · outbound
Plancraft: an evaluation dataset for planning with LLM agents Narasimhan, and Shunyu Yao
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c12d3bed-6c15-4c79-a0ec-4a930b70fe5d · outbound
Plancraft: an evaluation dataset for planning with LLM agents ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cbccd52-50dc-427f-916f-7ec7d2b8b666 · outbound
Plancraft: an evaluation dataset for planning with LLM agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 055f60ee-d2ce-4e03-9aa1-de7187c82301 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Gemma 3 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081fc9bc-764e-4035-9d14-32dde4d12d35 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 91aa0402-1940-4531-8be3-14587ccb4094 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4f0be2a-6056-470e-9f2c-4e36370ee0a7 · outbound
Plancraft: an evaluation dataset for planning with LLM agents ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4e836e-958f-492f-ac7b-beeb7ee13d4f · outbound
Plancraft: an evaluation dataset for planning with LLM agents Le, Ed H
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f02204f5-6697-41a8-9983-636805a31499 · outbound
Plancraft: an evaluation dataset for planning with LLM agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b14b8f-aa06-4a74-8cd8-af1181c01011 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b5cc90-0735-4706-a9e7-014ee022779f · outbound
Plancraft: an evaluation dataset for planning with LLM agents The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad43041a-4c60-44c8-a1c1-03b4c1f24634 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Agentgym: Evolving large language model-based agents across diverse environments, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622c2cef-fc5a-47e2-8f75-c69add1c58cb · outbound
Plancraft: an evaluation dataset for planning with LLM agents T ravel P lanner: A benchmark for real-world planning with language agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab05e995-3cba-4319-ad51-5ce72c237bac · outbound
Plancraft: an evaluation dataset for planning with LLM agents Qwen3 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a114be2-1e77-43af-a2bc-2cda882c73f4 · outbound
Plancraft: an evaluation dataset for planning with LLM agents GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743ffc22-c821-4684-8ded-a37454626c15 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd52a878-8b99-4885-82ed-2115bc8ae527 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Tree of thoughts: Deliberate problem solving with large language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9fd08eaf-e80b-4421-b36c-ec56d382f42e · outbound
Plancraft: an evaluation dataset for planning with LLM agents ReAct : Synergizing reasoning and acting in language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284316ec-9da3-40f3-bcb9-955f90fb8d78 · outbound
Plancraft: an evaluation dataset for planning with LLM agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591c7327-c25b-4752-87d7-325f51abd00a · outbound
Plancraft: an evaluation dataset for planning with LLM agents @esa (Ref
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b796486c-0b2f-4657-948b-394d834a896a · outbound
Plancraft: an evaluation dataset for planning with LLM agents Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bc7824-9550-4c8d-b846-9b683a9cd7a0 · outbound
Plancraft: an evaluation dataset for planning with LLM agents Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abff24fd-b60f-419c-a4d9-d0d3e9d60792 · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey Plancraft: an evaluation dataset for planning with LLM agents
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff96bea1-04fc-4465-bc89-20013780566c · inbound
Memory in the Age of AI Agents Plancraft: an evaluation dataset for planning with LLM agents
Reference 294
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c672ca42-ac22-4414-974b-925c85ff5c05 · inbound
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? Plancraft: an evaluation dataset for planning with LLM agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf88a171-f6f1-45dc-8b52-bb38ba31f581 · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Plancraft: an evaluation dataset for planning with LLM agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cdc6800e-bb63-4bf2-a048-6976b23682f7 · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Plancraft: an evaluation dataset for planning with LLM agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9acd0518-47c9-4290-bf37-7442fdd9cd96 · inbound
TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems Plancraft: an evaluation dataset for planning with LLM agents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89f93662-bec2-4466-8eec-8de955794d7a · inbound
Object-Centric Environment Modeling for Agentic Tasks Plancraft: an evaluation dataset for planning with LLM agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a7032f-82e0-41ce-8f6d-08bafddaeae7 · inbound
Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations Plancraft: an evaluation dataset for planning with LLM agents
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19a1ebac-1b72-4a5f-b1d8-97e85909ecdc · inbound
AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration Plancraft: an evaluation dataset for planning with LLM agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.