Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2410.03859.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:02:34.168723Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2e097ae9-4959-45ce-8376-d8fb03871ec8 · inbound
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0e35a0-7b30-4dce-9429-40a7af1b2d2e · inbound
CodeV: Issue Resolving with Visual Data SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffbf8629-6d21-48bc-9669-a5c5c19939bb · inbound
AutoPresent: Designing Structured Visuals from Scratch SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613aeedc-62b2-4968-98c3-4c0dbec072ea · inbound
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45939f3d-620a-46b6-85e1-1f2ce27f1231 · inbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9efe54ac-3208-402c-b903-77316b956f01 · inbound
Develop AI Agents for System Engineering in Factorio SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b24daf-d81e-4144-88ee-048b1ed30fc9 · inbound
The AI Agent Index SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752093fe-c71d-4d22-825c-a9376f7e5804 · inbound
Agentic Bug Reproduction for Effective Automated Program Repair at Google SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fbf4a3-8ef4-4e54-be6e-a2e33c7b005d · inbound
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569e9854-eada-48dd-817c-1d3e64809d0a · inbound
KernelBench: Can LLMs Write Efficient GPU Kernels? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0182c42b-11f8-4bfa-9994-6139a69d8496 · inbound
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b52647ef-c141-469c-be87-3b84e5cfc1b1 · inbound
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7b2022-6b02-4d15-b491-9a50cc08c345 · inbound
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c65b9d81-33ba-4eb1-ae19-eae98d3585b1 · inbound
SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe719371-92cc-4834-a8a2-9e075993215e · inbound
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae7f4e5-42dd-471d-b023-015998974436 · inbound
Multilingual Multimodal Software Developer for Code Generation SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3789f81f-780c-4954-b3d2-29b734b7aab6 · inbound
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0521b7dd-b328-4137-8a6e-81d155a0e877 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ad34d2a1-8cf0-47ca-92de-de061c9c6bc5 · inbound
Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c816a82b-f825-466f-9d70-8f6afb83e9bc · inbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 95d2e266-5e50-487d-bc5f-4125e7126727 · inbound
How can we assess human-agent interactions? Case studies in software agent design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1039a3cb-282b-4028-9ccb-226845f49d14 · inbound
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 66cdd8a2-8d5a-4c27-bc17-cbdd02e06549 · inbound
Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3ebc9cc-6ce8-49bf-a4b2-957fe8b7779d · inbound
Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e31f8ef-6b5a-44c4-8f5d-6d244db49fe9 · inbound
AlphaEval: Evaluating Agents in Production SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d13fd6b-bae3-476c-9a8b-788fc466c5fd · inbound
Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd71bb2e-2a45-426d-a00d-50930f3ba827 · inbound
The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5356b90d-e7b9-4a66-a319-e9d47b33bc26 · inbound
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 085178c1-03f3-4b6f-956a-f03900be35ec · inbound
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7a779983-db03-424a-b2f2-10e14520202f · inbound
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 41f1a7c4-cf12-47f7-ad24-997f0c3ee081 · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b5bdc529-81b9-402c-8587-8a1d835b3434 · inbound
ElasticMem: Latent Memory as a Learnable Resource for LLM Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ccc90743-de5f-4748-b3b5-bdb7d7df450e · inbound
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 33d5f3d5-9b40-45c6-b172-dab231b14e54 · inbound
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bb2696f8-df19-4e7c-90af-a0223aacb4d4 · inbound
RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38fa32d3-1c5a-4cff-b2b0-c8185b33f454 · inbound
Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 19d84601-11dd-49dc-aa1c-5df8de9866df · inbound
What makes a harness a harness: necessary and sufficient conditions for an agent harness SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b70e4fc8-eb15-4695-8bf6-f249c480ea16 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 282
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8d02e1bb-f5eb-4d85-a741-0e422d144021 · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3c8df488-9523-4581-9792-dc4616f0332a · inbound
Dissecting model behavior through agent trajectories SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a038349-a69d-4a13-9a3b-b18c1670b277 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f80db8c-a37a-4106-98e8-666af8c28a14 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41139610-86fb-4ddb-b748-6ef934e23745 · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e12773f3-bcc7-4168-9eae-0b672d7bcb66 · inbound
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 59593680-b4e8-4ac0-adfe-660c1024b766 · inbound
Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c9d3edf-dcc3-4f85-8f1b-787e72c9847a · inbound
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f4605d47-9cd8-480f-9cc4-c8952400f1ea · inbound
LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c428be79-3082-463a-a539-fe217ca6d6cf · inbound
MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474521f1-6871-4316-9674-9fed3c956337 · inbound
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c247b1-12a5-4bb1-97ce-5f3d7f76cb5e · inbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da65c2c5-ff8e-4cb8-8987-d510c8d416a4 · inbound
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.