Pith. sign in

Paper Citation Record · LEDGER

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.08160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08160 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:24:03.892023Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c978dad-8c8e-4668-9f37-f53f8112b1b7 · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Towards a Human-like Open-Domain Chatbot

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.785496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.785496Z digest=sha256:e565d771a94148483ae1be049b9a9b72a8e2bc7b59b90488733aee9e1f99dca3

Observation c6a18712-fc71-414b-ad6f-898742112cba · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Prometheus: Inducing fine-grained evaluation capability in language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.357075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.802957Z digest=sha256:79790670a7a9aede1ab8d5bf587d80755ca1c96c5132c9048f492843e8b6c107

Observation 36f5f8fc-770d-46e7-afd7-39ad98dcd032 · outbound

This paper cites Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.341954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.808280Z digest=sha256:d45d69990f40eae5576ed5f04984a906852a4b00c498f1b10374eeccb459fafd

Observation 68b75013-c94b-4c02-ac03-8690e3070cb1 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.813314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.813314Z digest=sha256:863cdfa969014ad83c55e1ecce8a9150ddc3e9b52478ee2fda1319d1bec8aea8

Observation 1607fd6b-28bb-41ab-a7b5-e57cda87559e · outbound

This paper cites Player-driven emergence in llm-driven game narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Player-driven emergence in llm-driven game narrative

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.308832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.823848Z digest=sha256:0ac52a2455666e409102780e14e94319e925d8494b1aabc861fc266359eddbdb

Observation 35975542-e343-4935-81f1-f10bde27460e · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives LaMDA: Language Models for Dialog Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.860945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.860945Z digest=sha256:c87e21aeaf4611de74ab3c21463296485b1d705440657063e812f2d12fc35b21

Observation b551a7ed-98ce-4ea7-8349-3f2efffd131a · outbound

This paper cites Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.199992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.872129Z digest=sha256:860aea499572fffa1b9c8ff089a3e7d62a789ff9440b9f5e83675d8a547d0239

Observation 0440e59e-1a60-4e6d-ab2d-0535c010162d · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Agentless: Demystifying LLM-based Software Engineering Agents

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:24:03.876817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.876817Z digest=sha256:b259046de04a9db0fdafa615e33f0cb0090c9bcffff81188925add807224c3c3

Observation 4b20e77d-dc52-46a8-a0c7-f6ffda46cd4c · outbound

This paper cites Qwen3 Technical Report.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.881671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.881671Z digest=sha256:be85628e5ecd96710bea18e39d5093646743dd6ab1d4034f5ba099c493b99b3e

Observation bec70c5d-a753-4fa3-96f0-94225235b708 · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Score: Story coherence and retrieval enhancement for ai narratives

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.887012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.887012Z digest=sha256:9f99624d37de6c5d0c8be67689e55bbdfdb23649c809e48c94bbb417a5d39816

Observation 134ab024-c815-485d-b717-19c6318be0a9 · outbound

This paper cites Interesting.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Interesting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.182847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.892023Z digest=sha256:536bf55c6d51e9cf473039527e5524e4fe1dae37b2918653b86cfd910eb6463d

Observation ba62213a-19f4-414a-bb3f-09d14ba10607 · outbound

This paper cites Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.325603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.819040Z digest=sha256:7fc1556599982f8c95edc89df895a6db07ad9fbfb227037c12e6ffd32d28e0bb

Observation d4cb5e96-c8f0-437a-9a00-e67905585a6a · outbound

This paper cites What makes a good conversation? how controllable attributes affect hu- man judgments.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives What makes a good conversation? how controllable attributes affect hu- man judgments

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.257169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.839707Z digest=sha256:35060706d51e3e150c194693146f54cac9ed71d28772869ca2a8ed4cee022f54

Observation 54b0f562-912b-44ee-b09f-f9bd56f901b5 · outbound

This paper cites Plot- machines: Outline-conditioned generation with dynamic plot state tracking.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Plot- machines: Outline-conditioned generation with dynamic plot state tracking

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.275910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.833990Z digest=sha256:b34747aaf1c9ddcebc35300befe3c5a5d4e400593606e6e416dfe7263c33032c

Observation c7051bb3-133e-4697-b26d-bb735d28ebf8 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Kimi K2.5: Visual Agentic Intelligence

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.850635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.850635Z digest=sha256:b97cdd5ea85fd4b3d1e54ce94bad4f4aad7df3e84a399907f12d42e0b5a97eb6

Observation 8425bebe-4687-4561-8b98-44732b6b51ad · outbound

This paper cites Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.844992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.844992Z digest=sha256:5c8bf7d27eb257803822de777a1bc60846561ef283df2c369c9aca644fcf2a1b

Observation b50cad13-f1d4-4bbc-ae47-7c4e9bea53ef · outbound

This paper cites Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.222720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.867337Z digest=sha256:ce03108d7339d5e02fc94ee393a328b96d20c5587d4907b63cc0afad2259d4ff

Observation a74d2641-a08f-4694-a266-9d1200787481 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.791845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.791845Z digest=sha256:3c87545b3a20a1d9c76a6f5aaa4126c1cf2c3e0a17c3785a27937aa33eabce93

Observation 8743f3e3-b051-4735-9fe1-29456f6b55b6 · outbound

This paper cites Red teaming language models with language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Red teaming language models with language models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.292768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.828593Z digest=sha256:f4bfa54f69680cd7d9a01d592b9b959fce477ed2144253a1a7a60d01716607d8

Observation 5464d4e5-f611-466d-8cb5-7484c48e3df1 · outbound

This paper cites GPT-4o System Card.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives GPT-4o System Card

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.797463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.797463Z digest=sha256:06b0df703ca88c14a282332d7ad8f7fb5933b0d74ff75641bd5da017859ee038

Observation b3ec1cb2-5686-4179-a1b6-ab08ad2558fa · outbound

This paper cites T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.239975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:24:03.856261Z digest=sha256:e6614fb1a181f1ba873dd1c72ebac34ca57477ee61d92265d6baace8f9c6143e

Pith citing papers

No inbound Pith citation observations are available.