Pith. sign in

Paper Citation Record · LEDGER

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 5 inbound Pith citation observations for arXiv:2507.20395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20395 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:40:02.915482Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:39:59.590925Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:43.323576Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · outbound

This paper cites MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:dbdb59a80ad55d81c62048560202c5f4cff6dc84747c307574ce0911d0ca8e70

Observation 99541be7-f270-4b91-aed6-9cbdbb8bbbcc · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:09.594740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:39:59.694094Z digest=sha256:9bb1a4f1c4c91b4fda5be9c09e59ada58933e4d0b96ac00f70df0a848487b9d5

Observation 0d21d79d-f7d3-4b02-82d3-c505c491a83e · outbound

This paper cites We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.444915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.204151Z digest=sha256:5edf8c29a59556ae57d76697a12ab043149b9119a3cf85f5bf0ba2b9d1bd6ae7

Observation 41bb9291-c22c-4056-b142-7a36ac46df0c · outbound

This paper cites The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.254906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.343327Z digest=sha256:2785f7094a83eb02cc5e10822e6d69149d7ce4866093e649c82a4977f48a35c8

Observation 434ccbdb-126d-429c-b277-edef3801563a · outbound

This paper cites illusion of thinking.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models illusion of thinking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:07.924848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.487347Z digest=sha256:009c581ade39fa2836e22d43b56491aacfc94b35599d0e73dc99d58e2d889053

Observation f50341c9-d0b7-4d17-ac02-04a6bfeb0dda · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.194752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.765301Z digest=sha256:8cbfb5efe04e23804149fd9bcc8e365bca7a9e40344ec433cb18b82a67ccc856

Observation 36f47d57-c02c-48d9-9ad0-6130d5b6d901 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.673377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.958377Z digest=sha256:e7169908eaa261014e9024ae4f5a30ee91b8f386b1ddff75a7e2028554c4e62d

Observation c22597bd-79f8-4566-8333-8d41d4303dd3 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.114830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.114830Z digest=sha256:69741f0bc20baf25530bb2adc7cdd8fe21a642c4272cf0dbd0fa8372dfee2d43

Observation 17eaa702-14d6-46c0-b172-226e24bd879a · outbound

This paper cites Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.425301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.425301Z digest=sha256:4b0a78f36fcfa787acb84f101839a9f1f02bed769e174ea21cd551d82eddb0f5

Observation 4d31b4ae-c2ec-4b68-988f-4c949a17cf66 · outbound

This paper cites Advances in Embodied Navigation Using Large Language Models: A Survey.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Advances in Embodied Navigation Using Large Language Models: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.594827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.594827Z digest=sha256:b4ce7802fe3ee6e296eb1469ab62af46e8eea5def2a359db69eca06e23266cef

Observation f5ecded3-8aff-4c95-9681-4f77893aac7c · outbound

This paper cites Exploring and Improving the Spatial Reasoning Abilities of Large Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.783017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.783017Z digest=sha256:f2c2c7d56e1756c8e57f44d1539be6f561fc5e25be1b7f3e7323f890c6532e40

Observation 1f793967-6f8e-417a-9031-75ebb38f585e · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.244842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:02.184839Z digest=sha256:fc2c15cbef8983e3a86ec507eba8c39e89c6b8a5737fac442997f0195a75079c

Observation fcea007a-56e2-4d04-8b0a-8875684c0d63 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.837369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:02.384763Z digest=sha256:07d9b1d64894de213f6b3939b24f77715e7ffbd537fe7fc299abb3fd2be1ea29

Observation d0996122-b4b5-4084-ac0b-27e838eba294 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.374835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:02.604877Z digest=sha256:8f3a526840136d5e12125a99fedbf3ac5acdb80e0881ffa11d21fa86d63540f2

Observation 23523d3c-1ce1-4d82-a7d2-8e4b7cfd036d · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:04.934759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:02.759852Z digest=sha256:88e5178efb899e76836acf1ddf8c6aad31a2eeb0695a52badfeba801d8bd064a

Observation 7b9318bd-40fe-461a-8ce3-25d15b80fcc6 · outbound

This paper cites Choose a direction: north, south, east, or west.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Choose a direction: north, south, east, or west

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:04.577069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:02.915482Z digest=sha256:2f505feb0e069f2e9ece7d55aa54f860c7ab5ff6240e043bd35c9f7129f11f09

Observation 71c1483e-9315-4cdc-9085-a1149bcdf460 · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.974813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.974813Z digest=sha256:d61543a113355f308d580f61083140f4818f314e6af06a97311a4c9b142a3d7a

Observation c45dc4ac-9566-44a9-a228-fe49abb6f589 · outbound

This paper cites TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:09.194750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:39:59.837529Z digest=sha256:6e3703e7677c1e82ea15befc0a9912d94ed35ec6a912ba4f3a9b5bf3d27426e6

Observation a9adacda-781c-4bf2-975b-d2bcac3bfe2a · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.565032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:40:00.634891Z digest=sha256:1add507ffa0608b1d99844258f9c2d657d96b4d877874ee073f6da72fd686e30

Observation c2525298-d22f-4b43-ad67-4caf1691b83f · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:08.774753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:39:59.995856Z digest=sha256:0ca95b3b79c9649b35c1c247e3e34be35caddec1e18f641e60fafa9d75777e3a

Observation fa740c9c-f980-4591-a60b-95ef6745e02f · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.294846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.294846Z digest=sha256:74a328414b5ee1639eeb6deddaad6618dc4cf81a73a728fef9407ecaa5d198a4

Pith citing papers

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · inbound

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models cites this paper.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:dbdb59a80ad55d81c62048560202c5f4cff6dc84747c307574ce0911d0ca8e70

Observation 2bfd5544-d885-4fae-a784-63ccd80d8a5a · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:18.991223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:f114e910d7956f210b8a32ed1be86ad02ef5b6d92437fead1e155695b2c013db

Observation 8a1c1c7b-fa67-4df1-a8dd-2da3ab196fa4 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:27.126726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:bccca4941036e0348411cd88a3deaa1d8d858ad401023c5f29a7798bfcfa7d4a

Observation b3db04d6-9c50-4606-b4a3-b4e5bc6bb440 · inbound

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning cites this paper.

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:23.712055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:02:14.499758Z digest=sha256:ea5ef641abc35d74edb3a13635842f58b732e3977aa3f91a1825260c88751b21

Observation c66416b3-7501-44d1-9685-c2c624d4d082 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:43.325606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:cca61081c06e100366371a351cf833989c86df1aad3f03f0b16c516f5b9f356b