Pith. sign in

Paper Citation Record · LEDGER

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 6 inbound Pith citation observations for arXiv:2507.20395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20395 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:40:02.915482Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:26:38.234273Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:43.323576Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · outbound

This paper cites MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:53f714b9b60939c150c6bd9a13c3f1e6fd0c8d0074563e4bdf76df1ad1235054

Observation 99541be7-f270-4b91-aed6-9cbdbb8bbbcc · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:09.594740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:39:59.694094Z digest=sha256:237a6005b7a06678b9daacb40a3aa8c2af3d1c39873d07ca8c20a9835ba82f5b

Observation 0d21d79d-f7d3-4b02-82d3-c505c491a83e · outbound

This paper cites We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models We focus on the configurationthatprovidesthemostchallengingyet fair assessment of spatial reasoning capabilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.444915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.204151Z digest=sha256:8d42a3b14b56d25be65aad892f224c320611863ce2564cfed9835dbc173f48c6

Observation 41bb9291-c22c-4056-b142-7a36ac46df0c · outbound

This paper cites The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models The results re- veal significant variations in spatial reasoning capa- bilities and provide insights into how these abilities transfer across languages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:08.254906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.343327Z digest=sha256:38512c43965a4b2a4ea567bf0d89303c3901f9379c085d74ead8193825977577

Observation 434ccbdb-126d-429c-b277-edef3801563a · outbound

This paper cites illusion of thinking.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models illusion of thinking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:07.924848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.487347Z digest=sha256:3a1adef09281838e819bcd2609424da4bf0aa3729f5bf91536cf6b47520e1561

Observation f50341c9-d0b7-4d17-ac02-04a6bfeb0dda · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.194752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.765301Z digest=sha256:2b30eccd05425e2b601cd6bba5ea8076a125fafd996a96a7a619049902f5984e

Observation 36f47d57-c02c-48d9-9ad0-6130d5b6d901 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.673377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.958377Z digest=sha256:de9a15b748eb949d6e5dcf157ad99acc6b4d633ce197ba545700dc5f11bed65c

Observation c22597bd-79f8-4566-8333-8d41d4303dd3 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.114830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.114830Z digest=sha256:ff0f3f10744d4e2b4c8ca83edc102d05da2986f55fe417af2dc1a2bf2f30789e

Observation 17eaa702-14d6-46c0-b172-226e24bd879a · outbound

This paper cites Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.425301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.425301Z digest=sha256:b9838b8dcfd97ede54ec0c367e868937a1565aaaf89ef38bb1f59c7b947218e8

Observation 4d31b4ae-c2ec-4b68-988f-4c949a17cf66 · outbound

This paper cites Advances in Embodied Navigation Using Large Language Models: A Survey.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Advances in Embodied Navigation Using Large Language Models: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.594827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.594827Z digest=sha256:140cb736e52adc12d1ec68155c161e7cf701f95b2ebfe4f5ef70085bb71388d2

Observation f5ecded3-8aff-4c95-9681-4f77893aac7c · outbound

This paper cites Exploring and Improving the Spatial Reasoning Abilities of Large Language Models.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.783017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.783017Z digest=sha256:ec0dfc5eedd6fa9d73354277ce814088c82697483ee05af0e35c0a75e4fb4abd

Observation 1f793967-6f8e-417a-9031-75ebb38f585e · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:06.244842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:02.184839Z digest=sha256:eed8eafcaeec2e22d73c6cb5bbc6390762c968d5e3272ba02dd39e431dd09618

Observation fcea007a-56e2-4d04-8b0a-8875684c0d63 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.837369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:02.384763Z digest=sha256:489399eb3cf392102a7cc87f4a55726a391f386b2c566178774e14839e9f73c5

Observation d0996122-b4b5-4084-ac0b-27e838eba294 · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:05.374835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:02.604877Z digest=sha256:a8ca9f28afd26b9f11c138f515100f5af2a38498fbfb5767ad243b8bdf454715

Observation 23523d3c-1ce1-4d82-a7d2-8e4b7cfd036d · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:04.934759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:02.759852Z digest=sha256:bf817f0593428c083443993eea01c8b89586988af93416abfa490402c8d700b0

Observation 7b9318bd-40fe-461a-8ce3-25d15b80fcc6 · outbound

This paper cites Choose a direction: north, south, east, or west.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Choose a direction: north, south, east, or west

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:04.577069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:02.915482Z digest=sha256:18cb4e24b978efc2e8dced314d6df4fac971ffd7c8323058fd24ee0dd1152fab

Observation 71c1483e-9315-4cdc-9085-a1149bcdf460 · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.974813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.974813Z digest=sha256:66a67e56b8a57db249a8ce0a20cc903c83b9662ccdd7edfe90371899b7fa77ae

Observation c45dc4ac-9566-44a9-a228-fe49abb6f589 · outbound

This paper cites TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models TextWorld (Côté et al., 2018) offers text-based navigation but in richly described environments that provide substantial contextual cues

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:40:09.194750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:39:59.837529Z digest=sha256:b076f943d5ea1ef0b257e967f9b40c79b47974523b1c48182133a0cf3b9dcb42

Observation a9adacda-781c-4bf2-975b-d2bcac3bfe2a · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:07.565032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:40:00.634891Z digest=sha256:e3f82f112335b65954d6ce340e8a7f7a85052d6b247d6fca80cc27361a99abb1

Observation c2525298-d22f-4b43-ad67-4caf1691b83f · outbound

This paper cites an unresolved cited work.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:40:08.774753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T13:39:59.995856Z digest=sha256:ef8e50028f8abe9f5ceced2bd9c849a6ee4771539f1da7dfb815ad5f8c459a28

Observation fa740c9c-f980-4591-a60b-95ef6745e02f · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:40:01.294846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:40:01.294846Z digest=sha256:998dd89ce99543138cc3a010b6464ea5530a58a7bae88c19e8cb582ce7bd93a8

Pith citing papers

Observation e40e9f1c-26bb-4f59-a470-b31aae587614 · inbound

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models cites this paper.

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:59.590925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:59.590925Z digest=sha256:53f714b9b60939c150c6bd9a13c3f1e6fd0c8d0074563e4bdf76df1ad1235054

Observation 60cfa622-2f54-4a99-803c-23c5588ef30d · inbound

Orchestrator: Active Inference for Multi-Agent Systems in Long-Horizon Tasks cites this paper.

Orchestrator: Active Inference for Multi-Agent Systems in Long-Horizon Tasks MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:38.234273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:38.234273Z digest=sha256:aa18e8f1bdc1d34420707b06c429a2313044930c814301b40352db64b213ac79

Observation 2bfd5544-d885-4fae-a784-63ccd80d8a5a · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:18.991223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:55bf1e43b188ea5e09e96fb5df37445171865be5ace0bb5e797b305bca1d5555

Observation 8a1c1c7b-fa67-4df1-a8dd-2da3ab196fa4 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:27.126726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:bf6bc100684f513b4ef56c1a6bb13ed03e17930a15f9c02b94a3f53787c99c69

Observation b3db04d6-9c50-4606-b4a3-b4e5bc6bb440 · inbound

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning cites this paper.

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:23.712055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:02:14.499758Z digest=sha256:7d9d43bce02791028ecb1753989b8335f34f94b2715aafd171aa262f0e001224

Observation c66416b3-7501-44d1-9685-c2c624d4d082 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:43.325606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:348ce61cfd17e52c1ccde6efb1c9906ed980b9ec7279a61b8289eb6c76257786