Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:40:02.915482Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 5 inbound Pith citation observations for arXiv:2507.20395.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:40:02.915482Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:39:59.590925Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:59:43.323576Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99541be7-f270-4b91-aed6-9cbdbb8bbbcc · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d21d79d-f7d3-4b02-82d3-c505c491a83e · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41bb9291-c22c-4056-b142-7a36ac46df0c · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 434ccbdb-126d-429c-b277-edef3801563a · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models illusion of thinking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f50341c9-d0b7-4d17-ac02-04a6bfeb0dda · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36f47d57-c02c-48d9-9ad0-6130d5b6d901 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c22597bd-79f8-4566-8333-8d41d4303dd3 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17eaa702-14d6-46c0-b172-226e24bd879a · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d31b4ae-c2ec-4b68-988f-4c949a17cf66 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Advances in Embodied Navigation Using Large Language Models: A Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ecded3-8aff-4c95-9681-4f77893aac7c · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Exploring and Improving the Spatial Reasoning Abilities of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f793967-6f8e-417a-9031-75ebb38f585e · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fcea007a-56e2-4d04-8b0a-8875684c0d63 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0996122-b4b5-4084-ac0b-27e838eba294 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23523d3c-1ce1-4d82-a7d2-8e4b7cfd036d · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b9318bd-40fe-461a-8ce3-25d15b80fcc6 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Choose a direction: north, south, east, or west
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 71c1483e-9315-4cdc-9085-a1149bcdf460 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45dc4ac-9566-44a9-a228-fe49abb6f589 · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9adacda-781c-4bf2-975b-d2bcac3bfe2a · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2525298-d22f-4b43-ad67-4caf1691b83f · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa740c9c-f980-4591-a60b-95ef6745e02f · outbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · inbound
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bfd5544-d885-4fae-a784-63ccd80d8a5a · inbound
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a1c1c7b-fa67-4df1-a8dd-2da3ab196fa4 · inbound
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3db04d6-9c50-4606-b4a3-b4e5bc6bb440 · inbound
Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c66416b3-7501-44d1-9685-c2c624d4d082 · inbound
Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.