Pith. sign in

Paper Citation Record · LEDGER

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

As of 19 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2608.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06663 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.758160Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d91c039-ad49-4de0-8523-8e0ec040dd6f · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Measuring AI Ability to Complete Long Software Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.467237Z digest=sha256:a074109b3ee5e58d49d126a1734565b205d8fe892d95b48c7e1d43dade140036

Observation 09b0131f-8e67-4138-8ced-bd0416dd7465 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Cannot Self-Correct Reasoning Yet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.471245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.471245Z digest=sha256:714764e190c4cde10061ea374f741cb625033a1610dd7a2602a0a817450e622c

Observation 9a2f02d2-5ee1-4303-83f9-b10ec55959bb · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.474552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.474552Z digest=sha256:0175738bac45a482ae014734ccd9cc5b015946928d4143a4ba6364366dc58475

Observation c637a348-385d-4bdf-9277-48b0c9453c64 · outbound

This paper cites The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.477854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.477854Z digest=sha256:2d0bbfc54eaf9251d9cb0e509c1703c40c051312d48f06e3c999cb6bba177f5d

Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.480747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.480747Z digest=sha256:9a1c4fe9b2efffb29111b2257c0a5bf46d59578e3c26657b523ea084d7b7f1e7

Observation 79e13dfe-c8ed-4052-a90e-8d56c9c6b9c3 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Understanding the planning of LLM agents: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.483951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.483951Z digest=sha256:03f2ca01b62e24747808c459faeb93184790bfd859f53b626bd8686bbc43ea27

Observation 9b045725-a423-42aa-9b11-3b0314362d1d · outbound

This paper cites LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.487410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.487410Z digest=sha256:32958aebe8d51a09aad178ad7456899eda7141f49cf7fb0ace9cac188f9635b1

Observation 49569b95-0a7b-4b37-888f-b7de3ab610f8 · outbound

This paper cites A Survey on the Memory Mechanism of Large Language Model based Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on the Memory Mechanism of Large Language Model based Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.490120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.490120Z digest=sha256:b9218b17c398554f2852663505f792f243fc39395b834768dcb761e14fce46b9

Observation 5d276526-6a02-481d-8c49-59bec7991584 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.493279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.493279Z digest=sha256:a9e2a3ce5aec34007193fdf6386242694be83713f5216e83d492ca0c0261137c

Observation 05fff376-c46b-40a1-ac08-6742aabd3304 · outbound

This paper cites Large Language Model-Brained GUI Agents: A Survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Model-Brained GUI Agents: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.495876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.495876Z digest=sha256:c3d644cc21a5771fc53e4c3f7b398a74aff4d98d88999c4555082a274e74560f

Observation e3c4aaa9-5303-41b1-9df1-04d0f9fb8a74 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.498394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.498394Z digest=sha256:16505455fe31dd90e821e76148a11417350906fb8aa14a995f4742f9904a0add

Observation 5aa6c2f5-88bd-4b5d-8583-c74358440bb0 · outbound

This paper cites A Survey on (M)LLM-Based GUI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on (M)LLM-Based GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.501038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.501038Z digest=sha256:be39f0acecbba39b24d4c987deaf4ed475a062027614c0b139f4a3a19e953546

Observation 79898d1e-c85f-4e7c-ac4c-01028f7fb264 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.503946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.503946Z digest=sha256:cdba8534e148dc09a9f58fcc09eefa06cacf429a10093622483d158e084c95bf

Observation 416228f4-611e-4317-816d-b6627376714b · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.506528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.506528Z digest=sha256:2bbe05e92f2710d8ad6d62e1df7595356c418b06b90691079835ffdd3294dc29

Observation 495abdf0-ba25-461b-9bfa-39dccef6eade · outbound

This paper cites Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.509270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.509270Z digest=sha256:d9a7abdce67c8b4882feb1e278e7bba39d8c0166c98b700a80a8fa617b379ba7

Observation 7611fd6d-329b-4548-a39b-882f21ab0d74 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Sutton, Doina Precup, and Satinder Singh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.512030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.512030Z digest=sha256:ad329cfa8d4da827a3a1ec8c0c9ca625c1dc95f99cb53634627aa7a3b2be2193

Observation c59fb5cb-4685-4af3-ad59-040a0e204946 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.514691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.514691Z digest=sha256:682043e3db1328cf5037b77588e95a34e2010eca169a1143d5e4547aaac04ef2

Observation 8f23b146-909e-479d-ab6b-2f57f65a232f · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.517185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.517185Z digest=sha256:05c59547f86ef78895b58e0efe4b62dee47de04dd7176e7c36fa11a1ff4c3bfc

Observation 0bc5a8c1-5005-47b7-a175-aabf23646121 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.520046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.520046Z digest=sha256:652b1ca3a211736dd9242a3335d54fa8f1fd5091287a357511ee3f6122349fff

Observation 805e7fd4-dd42-47f6-b2ea-372260a6d7cd · outbound

This paper cites PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.522839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.522839Z digest=sha256:cb2cf1e02e4739ac16677b2888342671ca6c0d866e5b32a6ab0bdd19da879bbd

Observation 1debe8c0-46e5-41b1-90a9-54954f4e4cb5 · outbound

This paper cites Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.525639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.525639Z digest=sha256:39d21dcd95c2f60160a8c44b73475a5234ac80921a9b6dd83cdd1e897cedd849

Observation 09e9ca7e-7693-4537-99a7-81361682208e · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.528245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.528245Z digest=sha256:0697f674e0bdd08acc88a43b6f9b4431a198c4eeb569d29be2b3d60fdde0ec5d

Observation 0e3d99fc-a984-40f2-af62-86f7a7d69fcc · outbound

This paper cites PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.531161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.531161Z digest=sha256:4801339c1ed0a2556d10a9f861ff2c5659ee7cb31fb91ca4702ed8bec24284f7

Observation 80a59d1c-9c40-4e89-9904-6c726f451504 · outbound

This paper cites Harnesses for Inference-Time Alignment over Execution Trajectories.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Harnesses for Inference-Time Alignment over Execution Trajectories

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.533892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.533892Z digest=sha256:fb04169b0ba8cc5270e30fc476108da54e9f4d2f7a290b703db8ed81185b4cf1

Observation e5b00e30-16ae-4bab-9c74-c69743b834ed · outbound

This paper cites PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.536440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.536440Z digest=sha256:af049a27a4fd5569d796b41a3d2c6dfc7f65df9a81d040b6bd229a2d54ccf0c9

Observation d32cb368-2985-4bdb-ba00-13b110e59b19 · outbound

This paper cites ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.539470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.539470Z digest=sha256:261a1801fa4d4fb0a62be85991274464ab0b01b57621edf9d2e95f6fbd1959ad

Observation 58ba2e0e-129a-4325-9afb-b97ab8098bce · outbound

This paper cites Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.542342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.542342Z digest=sha256:71e1422a9d27deffc8c6b7c5df125d0c1fd6d37f68c3294d9f2f07d480a2431e

Observation e2076d76-067d-49b0-b556-169498d713be · outbound

This paper cites The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.545080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.545080Z digest=sha256:d0ae2b0246fd7c510a23d41d4378566ed7f41da24f6297fa7ac25bc0b74bd51c

Observation 4d9aaf94-2e8c-48c3-9c7b-37f5cb63231d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.547554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.547554Z digest=sha256:4d388615906a8953d29edc153d06abd01bf0084a0da63d1e7ec15d30b5152aef

Observation 9b1e3e64-304b-4287-bd15-f678d3f508d2 · outbound

This paper cites CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.550199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.550199Z digest=sha256:047cb6c3a16040b580af7996517ad08e0d3660a89c3b1f062d5c3ed7c8083284

Observation b223cb23-1271-4f9b-a6a2-139b36034026 · outbound

This paper cites SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.553083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.553083Z digest=sha256:4770be9692b19fe815181461992a45a67cfda108dd26dac1ae736c4189ab6e61

Observation f3ed3bc5-cb75-4177-b4a5-89dd8e52bbe8 · outbound

This paper cites W ALL-e: World alignment by rule learning improves world model-based LLM agents,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents W ALL-e: World alignment by rule learning improves world model-based LLM agents,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.555658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.555658Z digest=sha256:2b92ac13dcec51cdf3b8b46a1579b8a395d90925a2163c9b853bc3ed3356515f

Observation 33743249-9f5b-4ae8-bcc4-d377b3b30ab2 · outbound

This paper cites Laird and Corey Clark.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Laird and Corey Clark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.561984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.561984Z digest=sha256:582e1b721d313ad76e26526fd2e7e36cc3a32b9c3e12813269485c02245fde09

Observation dd8b4f10-3641-48e2-acb3-f6ea52dcf7db · outbound

This paper cites MobileDreamer: Generative sketch world model for GUI agent, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MobileDreamer: Generative sketch world model for GUI agent, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.564416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.564416Z digest=sha256:c77d8e35ab12a9fbc47fba6995e5fc3c3843c7f41ba087487891ef77fefb5c34

Observation 0fdff7de-756d-46cb-b1a5-3f3f35638a9b · outbound

This paper cites ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.567093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.567093Z digest=sha256:cc930f168674e41b30838837177f27b6a9bb564d8c7a54fba03fccf2a8d012b8

Observation 304c4095-52cc-4303-8b03-781ba31759b7 · outbound

This paper cites AgentEvolver: Towards efficient self-evolving agent system, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentEvolver: Towards efficient self-evolving agent system, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.570037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.570037Z digest=sha256:5d67c65e6585019db6fffd3b1a0b52386ad6e6db89e7cca904115c4854ff1941

Observation a2ae34fd-a085-46a5-9538-a4e50a4f620d · outbound

This paper cites Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.572747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.572747Z digest=sha256:025217f6eca3f31662ba16fdcfecf680b1457a6e985e5980c3fc7b51bdcc577c

Observation af5c6012-b77a-43f6-98a7-b8f0140b360f · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lost in the Middle: How Language Models Use Long Contexts

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-10T23:01:35.575978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.575978Z digest=sha256:1b34134f8456b46e902beca2976793a6dd2b269cad21a5750c0ec36ee588376c

Observation 42ccb7fc-bc1a-44fa-9e25-e720b5f6ca86 · outbound

This paper cites Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.578943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.578943Z digest=sha256:9f04f9164175b9d79ff450ef48cf57c69344fe1634dfb474b5584b24a523f4f5

Observation 028e1dcd-ab43-4f06-921e-e162afc1fea9 · outbound

This paper cites Git Context Controller: Manage the Context of LLM-based Agents like Git.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Git Context Controller: Manage the Context of LLM-based Agents like Git

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.581794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.581794Z digest=sha256:6076b789c141f586b58735e8936320f6e28ac93b8f9d273b4dd7a2c6a120aaf6

Observation 21703a96-ef58-4aba-924f-ca12aa1fea38 · outbound

This paper cites Scaling long-horizon LLM agent via context-folding, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Scaling long-horizon LLM agent via context-folding, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.584602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.584602Z digest=sha256:2ccb2cec42f4fbc93e80a8d250c15b52fbfe2f6084a7214a6dcb6d3d6f00f5f8

Observation 27a8c101-fa00-4eed-94a8-fa0f4fb79d25 · outbound

This paper cites ACON: Optimizing Context Compression for Long-horizon LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ACON: Optimizing Context Compression for Long-horizon LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.587099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.587099Z digest=sha256:188c703df249661301e33a9eb997dff6f9a67fbbee567bc8aa67f9c9f8fead72

Observation b5c86055-aa32-468c-95fe-eb38d94e9afe · outbound

This paper cites Diagnosing and Mitigating Context Rot in Long-horizon Search.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Diagnosing and Mitigating Context Rot in Long-horizon Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.589850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.589850Z digest=sha256:e27fbc65a12b42cffd310df6e71b1284f768d23fa6ffedbaf9ee212d83939bcf

Observation 8bc86b62-d07b-43a3-9b2b-15d3a3ad1905 · outbound

This paper cites Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.592719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.592719Z digest=sha256:d6efb216e9398181c0938e7864d86613dfacc9f31d18850da7c31a3e0df37f79

Observation 416faf43-b239-422d-952c-75b229eb1b1b · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.595485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.595485Z digest=sha256:e4a2dae1af6cc7251330cb6825af8a62303b2cb35c154430e767068888a047c2

Observation c9386a00-719c-4362-9d2b-47e1ba67ab20 · outbound

This paper cites O’Brien, Carrie J.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents O’Brien, Carrie J

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.598199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.598199Z digest=sha256:e91c0d1c78f36fd467bd1f434bc6b1a34c1bcc88a5a998f3418a944cfb820afd

Observation 70f1a38c-65e0-4521-ae94-a5d7bf04b98e · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.600632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.600632Z digest=sha256:b2537f97002c8d6ddb8956899064aec3dcbb4b3f47144e698462b28120951e0e

Observation edd121b7-034a-4807-bff7-c0f3132f3d96 · outbound

This paper cites Graph-based agent memory: Taxonomy, techniques, and applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Graph-based agent memory: Taxonomy, techniques, and applications

Reference 48

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.195059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.603257Z digest=sha256:0397aa684f8a69b1e597cfed1c710facc3b58067b8331936218c3d1322c41fb5

Observation a691d46d-79fa-4b54-aa69-31ff878f77e2 · outbound

This paper cites WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.605771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.605771Z digest=sha256:e60aa14a9b6a90ed8b12c810d1b32d156f320d2dd5919b5a04cb64bcd44ce1c3

Observation 3ad3c7ab-96a1-4af4-b83f-cfbd3fa8e789 · outbound

This paper cites When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.608319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.608319Z digest=sha256:76e7fe5733b7b97b08de1893e016ba8ba179772decedaef1bb4ba0c73a4c90c8

Observation 5a0df324-3392-4164-9585-507b372a872e · outbound

This paper cites MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.610938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.610938Z digest=sha256:31c41e291c4f92bdb62ee8ad9f4db113321f6edcabc4d23ac9f4cf9bc0405f9d

Observation c9a1bb9a-942a-4db0-8214-6ed790172bda · outbound

This paper cites FadeMem: Biologically-inspired forgetting for efficient agent memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents FadeMem: Biologically-inspired forgetting for efficient agent memory

Reference 52

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.134245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.613747Z digest=sha256:96d83d15b1e5d98dde2d2815e10d30878a663f42f23de660460fa8b9cdeac5ae

Observation f72a3db0-9a42-4860-b3e8-e028ae15a2fa · outbound

This paper cites MemPO: Self-Memory Policy Optimization for Long-Horizon Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemPO: Self-Memory Policy Optimization for Long-Horizon Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.616146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.616146Z digest=sha256:6ddea6ee9ca9c080b3773d3bfafdcb93e66d5a302cd68568d2d8ff47cb95a4b7

Observation 1c123bf1-e8e8-47e0-a896-32a2a2db64e0 · outbound

This paper cites Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.619058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.619058Z digest=sha256:391f820180dcbd724f773e0e1f55d80829df22d7d77456c2abdab886f4d51418

Observation f03b548b-bc5a-4e8c-9968-3d6b5dfc9262 · outbound

This paper cites Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.621794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.621794Z digest=sha256:3b492799131df532c74226da569894c5e5729b7fc5c5b15676270df4a0a9bea0

Observation 78ed0683-b93e-444e-9d7e-ff74ac0c4d53 · outbound

This paper cites Forensic Trajectory Signatures for Agent Memory Poisoning Detection.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Forensic Trajectory Signatures for Agent Memory Poisoning Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.624528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.624528Z digest=sha256:bd995401638cb9e1785ea3983e8c3e54f21c1c8d6f5cdfbcd0ebffd432b15c5b

Observation fc37d485-bad3-4b9f-be01-ba609ce4a7be · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.627306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.627306Z digest=sha256:8e5dbd80661b62d6a7c7b24e9e453441fc3b9a72a1b291b5204ad51c00469280

Observation 1a45079d-5852-4f74-bab1-0229a88e15c3 · outbound

This paper cites Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.630273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.630273Z digest=sha256:dc51a1873859941642a3068c9640e23aebda273afd74b6f2d7880c7efa00d2d3

Observation 1f63ca0e-1696-4977-8787-6c7a4866e897 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.633179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.633179Z digest=sha256:ea9b0756f0b08a6de484a393cd7b1a910f2495a0de41deef9506f83e72b35453

Observation 3287544e-72b4-4d57-b3ce-b2c028532371 · outbound

This paper cites HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.636179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.636179Z digest=sha256:bef89e31f4868706df350c4f919f77821da854b852582cb460ca207298e650c7

Observation ec066b2d-7c9e-4b88-8e65-06046c159ad4 · outbound

This paper cites iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.638736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.638736Z digest=sha256:c187b3fde7c44d886827961cb03147b6e3454c6358b8a16e8bf56ce0647bdacb

Observation 45abc842-3186-4af0-bd7e-a95067b0d78f · outbound

This paper cites SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.641693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.641693Z digest=sha256:86c2b11bd9fa72fd5282f369a05fd227eaea7e82f96ed8066c031cb9456e1c6c

Observation 901dfc73-dfa5-4c2d-843c-eef1836d35ea · outbound

This paper cites xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.647556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.647556Z digest=sha256:cb987626c49f249a30aea257f1e3acce7c2bc7743e1daaa550571f963edf6615

Observation 84dea266-5519-4e0f-8279-6bd3287d25fb · outbound

This paper cites Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.651564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.651564Z digest=sha256:02b82b38c9a96acd5bbab79629bbf2c11f74a501b77c00f6b4eef635f85cd0a6

Observation 691c522d-49e5-43a6-b8d7-ccb8cc017c5f · outbound

This paper cites Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.654095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.654095Z digest=sha256:9ec6ca795eac02b882767a0a9d97d91b680be4af036f0fb6e16adeeae88b8227

Observation 7dd39556-31b6-416e-919d-70fc6550e6e3 · outbound

This paper cites Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.656554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.656554Z digest=sha256:38d701ed51a9cfa81cdedb7514470de53b7546a73e78e425249aa5be686122cf

Observation adb5527c-b2b1-417e-8002-2ce9847da674 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.659062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.659062Z digest=sha256:37ff87e21c8efbf5e311d6914e9a10dc7c6a7c01fd4be0d6053f3449215ed537

Observation 63528579-e520-4bbe-a426-ccf7c5517246 · outbound

This paper cites OmegaUse: Building a general-purpose GUI agent for autonomous task execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents OmegaUse: Building a general-purpose GUI agent for autonomous task execution

Reference 68

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.047704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.662924Z digest=sha256:3f4ae2db8b81e6f438d76082fb16868f43ebfa22d4c9f5a1435c619075a004e3

Observation f3c0ffee-606a-4584-97cb-e4cd47405017 · outbound

This paper cites An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.665805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.665805Z digest=sha256:132a64d9915fd269116f2180b03a421544dcb1ecc5a7f6baa35b6c66f7080a17

Observation 42aa1d8f-1201-41e9-9a3e-02e021fd7c3e · outbound

This paper cites A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.875866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.668706Z digest=sha256:0320046c1451afe055a1c4aa4006c3a388afbf755bf566279fa0e6f41ea01293

Observation add4d526-bcf5-4272-b7cf-7238890e9cf5 · outbound

This paper cites Governing AI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Governing AI Agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.671459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.671459Z digest=sha256:2454ee5f310254766f8d5494df596a8a7bcebb37e01ba9a9f3cd615051aead93

Observation c9ea7510-3b8c-49a0-852c-a570ba15ded3 · outbound

This paper cites The AI Agent Index.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The AI Agent Index

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.674505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.674505Z digest=sha256:20b22d8fe581d59fbd5af08a9372889d1c3e07aa9cdda0f18be9af4995b01660

Observation 5a9985a9-5b84-4abf-861f-88c391b0aa02 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Reflexion: Language agents with verbal reinforcement learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.677546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.677546Z digest=sha256:9f0f2db044f39244371b6dbeb487a15d9ec8c7c4e086f0559e07ba18bc80b69c

Observation db5d97df-474b-4f8f-b5c5-96f76572501f · outbound

This paper cites Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.849405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.680054Z digest=sha256:1f61281cd8456ee1432546e481332271868db79529e4245023c83e73151cccf2

Observation 5753611a-4298-44c7-bd09-4acea5a5c805 · outbound

This paper cites Large Language Models Can Self-Correct with Key Condition Verification.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Can Self-Correct with Key Condition Verification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.682869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.682869Z digest=sha256:4f818fcaa30c25122c2a23f41e987a3ce1ed5506fe22a693220195e1194c0c0b

Observation 307c0cdc-04a0-48be-a362-52d169bf18da · outbound

This paper cites Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.685622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.685622Z digest=sha256:96e9efa7fe01bf397d70c2fdfa69ee0636c80d5286b9a04e86d38ea6feace6c0

Observation 24a7b499-63cf-473e-8a30-198ba7f86666 · outbound

This paper cites CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.688258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.688258Z digest=sha256:bc09548a004193ac24e2d567579ec00cde75e7e284e5f6b178cb182396e2ea6f

Observation ac8685f8-54ae-4ee3-9784-32b8027f87f0 · outbound

This paper cites SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.691627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.691627Z digest=sha256:ae1b60cba85e95bea569b380d95bbdc79c3711476bc826f8d2fcfad4e5f66bb5

Observation 69cfafb1-596d-403e-bda3-fd2f5f9a17cb · outbound

This paper cites Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.694532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.694532Z digest=sha256:e68e770fb1aeb924aafc17cedafabf713d404d9ea8967faba1fdc2d59c5a5662

Observation a1ee9040-cad6-44c8-914d-f865b460ce2f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.697245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.697245Z digest=sha256:f8a523df7dbc457d2a26080c9bb0d366f04bb5a5d7c4c9afa64719f4520feb39

Observation 0011074b-293d-4309-a6bf-df0031a71baa · outbound

This paper cites ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.701005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.701005Z digest=sha256:8ef4f0b7d38a0b6aa1ad887d9e04947ebedd53f300ef003526b974b383172a10

Observation e3fd0de4-9bc6-48f0-84fa-fad8f6ca66fb · outbound

This paper cites SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.703441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.703441Z digest=sha256:2f9fd4690a5dae5ab7432555ac5dfd0469bd800563e4ca2f63f9f09a0854ccf5

Observation 5b467a41-9e0c-4344-bb23-a87414d2db54 · outbound

This paper cites Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.706348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.706348Z digest=sha256:044ec2bc6f7c0d56857dca920d31dc69ad1f07380f3aa4ff03d3a2d8924a4b56

Observation 341cc47c-06cd-4f9e-9f2f-7efe851364ed · outbound

This paper cites Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.708864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.708864Z digest=sha256:8ee538dd004c637c00588faf7a358bbeaf9b6b4a713977839faebd97fa83c466

Observation c6470f50-8445-4a83-8d26-7cb5846d6b56 · outbound

This paper cites Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.558636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.712150Z digest=sha256:ccca199015248d38ce5bb8ef7fa6551e4f03e8fa8d577924c83cf7a86380cc47

Observation 18100c34-618d-4f6f-a317-ab6985c78114 · outbound

This paper cites MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:35.979608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.714955Z digest=sha256:6b2f3cd4c4266cb33ac7b599c43f733409cdf4c56c8916b5fee2fb10ceb97a77

Observation 67b78875-b23f-4f73-99ed-7715ecf8436c · outbound

This paper cites Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.718280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.718280Z digest=sha256:2a33429839c78a5904d8e3fddf565b8e0a28c593e8b1e8bb12b6de733393105c

Observation 6f24c853-84c5-4563-a314-5993871c2b10 · outbound

This paper cites Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.721043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.721043Z digest=sha256:196df0eecbd8242e880c15e8a5f7cd5e08d7cb423d74db174d489f117bb4bf51

Observation 1a06d184-2b6d-4231-b3cf-177e3d10968f · outbound

This paper cites Let's Verify Step by Step.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Let's Verify Step by Step

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.723496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.723496Z digest=sha256:b77803063360689c934b27b108b47e474e2f7fe5cfd34e92b23accebdfb100c3

Observation d12b5d82-c357-4464-89c1-dcb0d289388e · outbound

This paper cites Entropy-regularized process reward model, 2024.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Entropy-regularized process reward model, 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.726497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.726497Z digest=sha256:c35556606590c5fb52bfc7086625b9e08820ab0352cc4df3b259184583b8c022

Observation f48e0268-e156-475a-ab03-ac7aadb5c4f9 · outbound

This paper cites GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.728986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.728986Z digest=sha256:1a02723ff54e81633a1fb45be90189d344e9dd69f7d0c95315ac877f0b279a74

Observation f4d8f352-658f-4f61-951f-23098439453f · outbound

This paper cites GRPO is Secretly a Process Reward Model.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GRPO is Secretly a Process Reward Model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.731504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.731504Z digest=sha256:11f94ebcd965c9d421d1ff6c592c19101fa3aa53a49c03f03060f8dea88aade1

Observation abcf2ec5-4a78-4c06-90b6-c33f2e7f18b6 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.734381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.734381Z digest=sha256:7538f2e3c3efa0b7780578fe5412de6a8ca72b577a91c917f19bbdccb7c1d6bb

Observation f6761413-445d-484c-8e28-5896ccde7a29 · outbound

This paper cites AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.737168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.737168Z digest=sha256:41dbbc13957d5bac912e46ce298683a63efbc5dad7081d5074a3d3a8b316a5ea

Observation f5be3098-e6dc-41b4-a776-5bec79c7e34c · outbound

This paper cites SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.739622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.739622Z digest=sha256:ea6200d1d715ba722aeb2eda6d3795a6951dfdbe6bcfdc1c5d90e838b839e07d

Observation 5d856c48-6af7-4ea0-9cb8-12146571136b · outbound

This paper cites Agentic rein- forcement learning for search misaligns instruction-tuning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agentic rein- forcement learning for search misaligns instruction-tuning, 2025

Reference 96

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:01:37.234083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.742845Z digest=sha256:4a6ac5e14761f57e7a39c90d7bd34590f74360fa259705d75f15c3ba25741ef2

Observation 2169a623-151b-4748-b08f-1e94f0cacc5f · outbound

This paper cites Self-evolving LLM agents with in-distribution Optimization.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Self-evolving LLM agents with in-distribution Optimization

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.167467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:01:35.745579Z digest=sha256:011fd104ad61550c83494cb25f2842385ed40167e115f6c11de1f5b46c4897a5

Observation 1b0478b4-4a35-464f-a68a-9a73bf9eea4b · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.749355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.749355Z digest=sha256:014d7fd942163c26afdbf30376706f8e87dc4eb5ba30adc7dedd1a3e09acad98

Observation c40f7a68-d53b-4b54-92b5-d969088ad129 · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.755009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.755009Z digest=sha256:be965b7c0b406134b873ea106e6a6d8e017d3022357f7ef864e17778830d3a81

Observation eabd52b4-b678-4be2-8d51-d1687327794f · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.758160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.758160Z digest=sha256:d84443ed841fb9a33b18fd7b5f9eb5722e2be65ee207abac474f8d9cb24795dd

Pith citing papers

No inbound Pith citation observations are available.