Pith. sign in

Paper Citation Record · LEDGER

TextQuests: How Good are LLMs at Text-Based Video Games?

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 7 inbound Pith citation observations for arXiv:2507.23701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23701 v3

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:32:33.689457Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:20:37.943280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:55:01.576018Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cc5aa1a-e6e6-45a3-b825-0342e541670f · outbound

This paper cites This is a fake!.

TextQuests: How Good are LLMs at Text-Based Video Games? This is a fake!

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:32:34.013674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:32:33.689457Z digest=sha256:fa4c74d6af06fd8f28d4f9398702bdf0b32b974d0c1f6cc71518505ab802ee78

Observation 2c8caff6-8dcb-45c4-9c85-80b1ad948731 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

TextQuests: How Good are LLMs at Text-Based Video Games? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.646495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.646495Z digest=sha256:eae60f50bbb4d0bc1fba5136c19b7ca8d335cab937422753fe759ae095ab3daf

Observation b9a2ed50-0cf9-4000-8aeb-05c722437768 · outbound

This paper cites Interactive Fiction Games: A Colossal Adventure.

TextQuests: How Good are LLMs at Text-Based Video Games? Interactive Fiction Games: A Colossal Adventure

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.650545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.650545Z digest=sha256:bdaa1f8c7b88e2ffb9454ea2f3dc506b7736fa2cef4f87b5c5bc91568a3d9aae

Observation 6931e197-79b0-4ee8-955f-8f83aa2a1dbf · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

TextQuests: How Good are LLMs at Text-Based Video Games? Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.654873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.654873Z digest=sha256:b5dcd927f70edd059df9c7fb7b17a9177cbb591e107689f7d9ea27aad5642df3

Observation b427c8d8-807f-4f2c-8c40-6eb97c5780bf · outbound

This paper cites Humanity's Last Exam.

TextQuests: How Good are LLMs at Text-Based Video Games? Humanity's Last Exam

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.662216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.662216Z digest=sha256:bfed8e98305fde3c7b3ab1698b923219f014821684527d14485c9f82dee3a763

Observation d032d437-81f3-4623-9314-d492177e97ed · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

TextQuests: How Good are LLMs at Text-Based Video Games? GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.665707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.665707Z digest=sha256:bf5649f35025b22c1ce1e99aafded503c06a627c451a63ef41e107b514859cfc

Observation 0a513bc0-9837-42b3-8f72-4faf142f9ccd · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

TextQuests: How Good are LLMs at Text-Based Video Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.669131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.669131Z digest=sha256:7e6403dbcd570520ee0feabf82e54b370c3f132cc19de9d97111609db7bbe2a2

Observation cf4fa4a3-6d43-4ab8-b363-b39db566f4f3 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

TextQuests: How Good are LLMs at Text-Based Video Games? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.675992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.675992Z digest=sha256:cc401ffb9f62c2977075d89abecc04bd61dfb2f3ccea4030a3ce5ec5c8f7e6ca

Observation a53f149a-7e84-4e6d-a52a-cb9a4b22006b · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

TextQuests: How Good are LLMs at Text-Based Video Games? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.679463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.679463Z digest=sha256:0ab5180202a2b3e825f4c359819b4162325fe84390c6d716a0247f2af5625f37

Observation cf6e84a8-dbe5-4d56-bff1-3a6c33cbfc4b · outbound

This paper cites Keep CALM and Explore: Language Models for Action Generation in Text-based Games.

TextQuests: How Good are LLMs at Text-Based Video Games? Keep CALM and Explore: Language Models for Action Generation in Text-based Games

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.682799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.682799Z digest=sha256:59a36572c367aaf4ce51209566efbcf7ce67de2a223a3fe9b96ced503c0b2315

Observation d7ad646c-81c8-41ae-b2fd-771f9ec33677 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

TextQuests: How Good are LLMs at Text-Based Video Games? $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.686110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.686110Z digest=sha256:ace8af4bef973081a39bc3b92ef3e0d2863acae5996619c2044a57686c5560ca

Observation 936c68e0-6504-4eca-9798-42321d74d541 · outbound

This paper cites an unresolved cited work.

TextQuests: How Good are LLMs at Text-Based Video Games? Unresolved cited work

Reference 1983

Resolution
unresolved
raw_fallback, observed 2026-08-06T10:32:34.025160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:32:33.672526Z digest=sha256:4a4e67361c2a969cf4894a53e861faf9931a2895465adbc573c311ded1dbc2fb

Observation 9dc0c724-c700-4f6c-81ad-33c8d9d18991 · outbound

This paper cites Graph Constrained Reinforcement Learning for Natural Language Action Spaces.

TextQuests: How Good are LLMs at Text-Based Video Games? Graph Constrained Reinforcement Learning for Natural Language Action Spaces

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:32:33.929689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:32:33.638115Z digest=sha256:874359dfa7d6b2b197e82f0d44d00453f984d3dfc44fd4c4bdfde6431521bf2e

Observation f1e4e1fa-c232-434e-93ac-b54be87693bb · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

TextQuests: How Good are LLMs at Text-Based Video Games? GAIA: a benchmark for General AI Assistants

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.658758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.658758Z digest=sha256:4a9818882a318f4ac4f9b9d462fb10f1b225cc178b3a39f09a9f80ce8dee15af

Observation 38731489-54f2-4d90-a894-9f3335f618d0 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

TextQuests: How Good are LLMs at Text-Based Video Games? LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.642657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.642657Z digest=sha256:c9faad6758c1b0ab5dddb0225bf547ad00c36a204a636dfdda55264c7a82431d

Observation 1d61648b-c074-4e2d-b1ed-e04f00aabbe7 · outbound

This paper cites Prithviraj Ammanabrolu and Matthew Hausknecht.

TextQuests: How Good are LLMs at Text-Based Video Games? Prithviraj Ammanabrolu and Matthew Hausknecht

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.633746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.633746Z digest=sha256:250f78d3df3a00b078df518adcc8bea0ed471fc659ef8a90c453ac7c299cfc7c

Pith citing papers

Observation 45cd63f8-f908-43a4-b57f-8de2389e77a0 · inbound

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions cites this paper.

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:37.943280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:20:37.943280Z digest=sha256:6ccc1cc98fdd064e3262ffd35bf04f47809791b897de07dfd2d1f30a6112024d

Observation 6e43c9bd-e4ab-4596-ace1-ce68d015e3ad · inbound

RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents cites this paper.

RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:02.428353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:16:34.750723Z digest=sha256:01bb58996ce641b3e0e72bef493e80597de93e95b0b87af4fc930b968204605c

Observation 6f946362-47eb-4357-acd2-754e05e1d385 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.100599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:02ad4e2ae5274ebe42cb9f3a5afd890c0b8ebef439a45102799d1e9a8108f0ae

Observation 18599c8e-3c0d-43e6-829e-f524f7c865e1 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.910724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:421ad0f981f844be96c0dd06a36f5f69f3aca337eb07ea3bb1954a2f65c89c96

Observation eb0af504-3f2e-4f4a-95c6-464ffaad9468 · inbound

Muse Spark Safety & Preparedness Report cites this paper.

Muse Spark Safety & Preparedness Report TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:55:01.578039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T19:49:14.463992Z digest=sha256:b3b9084e24fa6fb5e52c8c4869fbf818fb39b82ff4c9cf437444c30ac29c15c2

Observation dfdc32e5-0e37-4fb2-8617-9d77630073e2 · inbound

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI cites this paper.

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T16:36:34.700284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:36:34.700284Z digest=sha256:e1758ce9372d5489bb1da04a59eafbe4232bdc4d8ce2a423218b7ddbd327c559

Observation c4489fd4-0148-4f49-af23-cc7ebf6ddca2 · inbound

Rushes: A Human Preference Dataset for Pluralistic Alignment cites this paper.

Rushes: A Human Preference Dataset for Pluralistic Alignment TextQuests: How Good are LLMs at Text-Based Video Games?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T09:29:01.220666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:29:01.220666Z digest=sha256:1eedc2a8f152977f215f9a8aefc8ef34b554514349b7ca5cec482b9e9b0e0397