Pith. sign in

Paper Citation Record · LEDGER

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.25152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25152 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T00:06:02.229517Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba35ee0c-a4db-435c-9a67-67491996ea4a · outbound

This paper cites From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.121624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.121624Z digest=sha256:c3865ecc58659f5fef1c8b154e29e2134b78f83589cfc6a21affed0ca03799df

Observation 1ba18b0d-4372-4b9e-9146-a5d73a93a127 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.584002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.584002Z digest=sha256:95bf259085a2b364b01eedf167a4d85ed75dabcd62af54f240c8798318f3bfbe

Observation bb4a6551-b82f-4d42-a567-e2e50ff28a51 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.702622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.702622Z digest=sha256:a59246d0c9ead9e8ca1522e3be9fc8b9e76c07e04f28a635787d44e98105358a

Observation c3296c37-e76f-455f-9bad-bc9a8554d11a · outbound

This paper cites ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.839451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.839451Z digest=sha256:041a717d7acdc4faa28970ca7d5ff142f4f9a7512cba9e7b5a6d3316bfefffe7

Observation c371ef5f-2110-42c9-b13b-004fbc0ba343 · outbound

This paper cites Deep Re- inforcement Learning from Human Prefer - ences,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Deep Re- inforcement Learning from Human Prefer - ences,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.907437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.907437Z digest=sha256:a50918758b27ecf1b63eb9582307c5ee65b7db1705645a862b30b4e39931c8ff

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.011035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.011035Z digest=sha256:8d1099af4f3f79f6e95406316bc111d492123a54e26a8cabd43dd38082793ed0

Observation df82bd36-55ea-4ca2-a8aa-20f4bd708f46 · outbound

This paper cites Scal - ing Laws for Reward Model Overoptimiza- tion,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scal - ing Laws for Reward Model Overoptimiza- tion,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.086887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.086887Z digest=sha256:ae43f5dbc8be83635e89ff124f770ea7420b4561a95b0199d653a57a859a176f

Observation 9d69b39d-f819-499e-8920-194b40d65375 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Large Language Models Cannot Self-Correct Reasoning Yet,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.137455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.137455Z digest=sha256:389399e4791b1ca14c2841c1329ea4cb4d952b0a0bbd4a3baf1097e91ed5f8f2

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.192018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.192018Z digest=sha256:42fdad9a726dc5e20c64fa103f51adbf704555e6687df7653c73cc9b0ea97297

Observation 7f926ccb-1d5e-4c05-9e73-e560f5900863 · outbound

This paper cites Controlled Exper- iments on the Web: Survey and Practical Guide,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Controlled Exper- iments on the Web: Survey and Practical Guide,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.347221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.347221Z digest=sha256:d6949b2da5dc2e06095d559c3153c7ac5a065b7f2470a935b26ef3494bb0a271

Observation 8aba963f-f187-4d10-82b5-07710c08c0d8 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scalable agent alignment via reward modeling: a research direction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.562062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.562062Z digest=sha256:d897557736f28a8acf19fce25ad39984127d944c795efcc9637446d63f9bd4da

Observation 7f0ccc12-c41a-48c8-98fb-179d494630ee · outbound

This paper cites Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.611392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.611392Z digest=sha256:6d69461a410d3be2865cd1bf0fc78cce9aa7f4181594ef219ae6f3176f63c57e

Observation b7441adf-2310-4c39-991d-6033c8b791bc · outbound

This paper cites Let’s Verify Step by Step,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Let’s Verify Step by Step,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.693551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.693551Z digest=sha256:b7e33d90e1a6929a975e024141ce699dc97fc8e3fa752b46d25a49beb063cd51

Observation 9a0517dc-6130-461b-893c-de4ab8ddb17b · outbound

This paper cites Retrospec- tive Progress-Aware Self-Refinement for LLM Agent Training,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Retrospec- tive Progress-Aware Self-Refinement for LLM Agent Training,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.752157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.752157Z digest=sha256:609dc6dd25eb267323da1d64470b86df93307974b803b7d0e1735620b8b08aac

Observation effd6ed5-f217-46d1-9037-895159e7c9f4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Self-Refine: Iterative Refinement with Self-Feedback,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.814656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.814656Z digest=sha256:1df33a9edee89d6d1a3109c909db676e1f32764ee2dd69b29624bafc62981a23

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.865901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.865901Z digest=sha256:4e5d073bee73e641f31bb21460da8bc3a6635abc150b2d92370d75212dfb55b9

Observation c60514e2-ca06-4acd-b210-8d964cfcb204 · outbound

This paper cites Training Language Mod- els to Follow Instructions with Human Feedback,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Training Language Mod- els to Follow Instructions with Human Feedback,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.931998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.931998Z digest=sha256:cf01387e30fac241911f1ef4f97c57a9e5e7d0138a8376258664c1ce8b6f86a1

Observation 4b084f1d-6454-4fbe-bb3b-74e0c7ff3cf0 · outbound

This paper cites The Effects of Reward Misspecification: Map - ping and Mitigating Misaligned Models,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops The Effects of Reward Misspecification: Map - ping and Mitigating Misaligned Models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.033058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.033058Z digest=sha256:d3cb0926700fc2010d0a4dd13bd8c07f0a1fe82667d8d3e71b2c703479a2bbf0

Observation 0e0dc9f1-da99-491f-9b26-d10498874b10 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.103945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.103945Z digest=sha256:973836aac0224278cc4b8968f44ce969927bacfd529d7865ec25cacde3d398b9

Observation 06b0bd32-9eb6-4620-8bb7-18cbf470cd0a · outbound

This paper cites LLM Evaluators Recognize and Fa- vor Their Own Generations,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops LLM Evaluators Recognize and Fa- vor Their Own Generations,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.201590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.201590Z digest=sha256:154a2ba257c5a96780cfea359738680f39bac84a3782a15bcad549ab6936a33d

Observation 37e51bc7-4821-4415-a6bb-a23c9b3c4d78 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reflexion: Language Agents with Verbal Reinforcement Learning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.307077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.307077Z digest=sha256:797628ab34b2e1d7677cbd24087f8b4251ff78bf1048dfd4eff62dbc915540ce

Observation e5889ad6-d8e6-4686-bbd1-cc028bf4d6d3 · outbound

This paper cites Defining and Char- acterizing Reward Hacking,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Defining and Char- acterizing Reward Hacking,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.395358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.395358Z digest=sha256:381ab2687dd30ea7f7cea757c48cc01985a2430ea421680ccf38b4d2ce663763

Observation 300eacf1-8137-437f-891e-c2b8b8d2333a · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Solving math word problems with process- and outcome-based feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.464945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.464945Z digest=sha256:d75694ea88b86f90ba4b179c4b446879b3e71a5b9b34077e3a9da3da19997270

Observation 06b51709-4ae3-4445-b4ac-b6c889540733 · outbound

This paper cites The Verification Horizon: No Silver Bullet for Coding Agent Rewards.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops The Verification Horizon: No Silver Bullet for Coding Agent Rewards

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.599726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.599726Z digest=sha256:7cbf95a6f34e464e90b3a462aa5eca6cbc0f01809e7544eadcef9d2134d10e40

Observation 4cac049e-4bbc-4de3-acb2-ad957a09c817 · outbound

This paper cites Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.663729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.663729Z digest=sha256:11b5de77aefcd6651a2174aadd5bbbd47fd18ccdc813d19c483238a7cde7506d

Observation 12ca15e8-2dee-4054-b2cd-3432047aef7e · outbound

This paper cites SWE-agent: Agent- Computer Interfaces Enable Automated Software Engineering,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SWE-agent: Agent- Computer Interfaces Enable Automated Software Engineering,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.801069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.801069Z digest=sha256:22683cd68dc2ec5b327784695b1dc1a9691cbafa964f998e681a930b4393fdb7

Observation ce56bc70-e032-48e9-b970-c352c3180329 · outbound

This paper cites ReAct: Syn- ergizing Reasoning and Acting in Language Models,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops ReAct: Syn- ergizing Reasoning and Acting in Language Models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:01.918907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:01.918907Z digest=sha256:f1071018a17d05bafa2fc58457b2e6479f32b403ebe40082fe92588b0db2b3ea

Observation 4bd76697-74f8-4787-b2da-f22913978e9a · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:02.005364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:02.005364Z digest=sha256:874c80fdd70aae78a5268fd17589da0e1b59d69b164ea07ea3a42caca045d09b

Observation 46f31ed6-b2f3-450c-81d9-0cda687c4198 · outbound

This paper cites SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:02.127727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:02.127727Z digest=sha256:6fd49585e7e5b7910ddc47ee073d5b30cbafdc5b01053c438063a11922d86fdb

Observation 3e525af4-cf19-4908-a32b-a58f072f4bd4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:02.229517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:02.229517Z digest=sha256:9c38e083d05e2d727d06ebcce5e056f53b84dc5a17838d5b8f24fbc9f9e8105a

Observation bb047815-d530-49fb-9be7-bb482fd10599 · outbound

This paper cites Kohavi, D.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Kohavi, D

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.404153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.404153Z digest=sha256:cf046910af88e79a551449269affb272f04360fd3117866d2afa1e24c7bf0330

Observation 5a7bf308-a0e1-4d9b-9ac3-1e8ecb116fa5 · outbound

This paper cites Building Effective AI Agents,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Building Effective AI Agents,

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.408979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.408979Z digest=sha256:f35908f07e999fe6cc3e4d3dd8d69bceffe9fa052227eb4ab4441c3e06cab76c

Observation ceaeaecc-b479-4e83-98fa-9f6965791359 · outbound

This paper cites GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.959934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.959934Z digest=sha256:81626eaab14443168e3e27087afd88f0a61762ed32e014507148627edbb3718a

Observation dbb15f01-af43-4cde-b83f-623c58ad06e1 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SWE-bench: Can Language Models Resolve Real-World GitHub Issues?,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.294890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.294890Z digest=sha256:cda3820fce8c5b54187cfedf35cfaac59077e17845b85400884dc9f2270d9069

Observation 9de3934e-92c4-4e34-bcee-718312eaf032 · outbound

This paper cites Specification Gaming: The Flip Side of AI Ingenuity,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Specification Gaming: The Flip Side of AI Ingenuity,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.478706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.478706Z digest=sha256:572ec607f7f659aedfa6587bb390d99834ab4387482a2930b6386981783dbe80

Observation a7683982-5cd8-4534-b471-02c2786afd19 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.643757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.643757Z digest=sha256:afbd23e7e2bb9bd7b82dd388279544d95b62a72a80083d26e4245ffea1678243

Observation 290563ce-728e-4d18-917d-dcad6186e0c5 · outbound

This paper cites Reward Hack- ing in Language Model Agents: Revisiting AI Safety Gridworlds,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reward Hack- ing in Language Model Agents: Revisiting AI Safety Gridworlds,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.783791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.783791Z digest=sha256:28ded6ce509480f5a2cb637876fb74c0eb9abff8da50a2b29b1bbbcb12d2adab

Observation 27e3f52e-3632-40c0-a007-5f08b151d0ac · outbound

This paper cites Effective Harnesses for Long- Running Agents,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Effective Harnesses for Long- Running Agents,

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.450472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.450472Z digest=sha256:937b230bd49f8aaf1ff3ef86ac12cfe6ca1839ebeac1bbd2a3854488747236b3

Observation d1030f6f-d2ce-434f-9943-3b5aac86714d · outbound

This paper cites Reward - HackingAgents: Benchmarking Evalua - tion Integrity for LLM ML-Engineering Agents,.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reward - HackingAgents: Benchmarking Evalua - tion Integrity for LLM ML-Engineering Agents,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.528265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.528265Z digest=sha256:96cbbda64d84ac55f51143f64cd0477f81ef8dbaff774d4e23350e7c6a8b946c

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:59.249687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:59.249687Z digest=sha256:61d0e61edc024e8d18840424091fd8d7ffcedc318ccd682e8bb94aa594712647

Pith citing papers

No inbound Pith citation observations are available.