Pith. sign in

Paper Citation Record · LEDGER

SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2410.03859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03859 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:02:34.168723Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2e097ae9-4959-45ce-8376-d8fb03871ec8 · inbound

TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved? cites this paper.

TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:02:34.168723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:02:34.168723Z digest=sha256:09e6309b678cb830b8ac6c884ce85a2a00a564e3b41a9b0a2bbf40c207548a4c

Observation 7b0e35a0-7b30-4dce-9429-40a7af1b2d2e · inbound

CodeV: Issue Resolving with Visual Data cites this paper.

CodeV: Issue Resolving with Visual Data SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:39:29.949421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:39:29.949421Z digest=sha256:93194a574dd4aecf2bfc4570775c26dfef23a3ded08eddeabe46f33cad25be77

Observation ffbf8629-6d21-48bc-9669-a5c5c19939bb · inbound

AutoPresent: Designing Structured Visuals from Scratch cites this paper.

AutoPresent: Designing Structured Visuals from Scratch SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:31.279289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:31.279289Z digest=sha256:9bd6032a85f727760fc34f1a090e3ca35fd54c3e618c8caeb8e47fa633d9167c

Observation 613aeedc-62b2-4968-98c3-4c0dbec072ea · inbound

DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale cites this paper.

DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:43:57.468297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:43:57.468297Z digest=sha256:32b2beb5fb85e37f1c815c1020049d6b8d1506d9d5a05c29215e610362753b12

Observation 45939f3d-620a-46b6-85e1-1f2ce27f1231 · inbound

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems cites this paper.

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:23.602920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:23.602920Z digest=sha256:042fd6bd6c0c0a877b1e39d7e51790f4cc2b9d5410d86a794375ec5a721d221f

Observation 9efe54ac-3208-402c-b903-77316b956f01 · inbound

Develop AI Agents for System Engineering in Factorio cites this paper.

Develop AI Agents for System Engineering in Factorio SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T15:09:20.607919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:09:20.607919Z digest=sha256:bcec7fdd29f97f1388a7d485a9da51f3519134b47fdb04264e1760dd80e1c697

Observation d7b24daf-d81e-4144-88ee-048b1ed30fc9 · inbound

The AI Agent Index cites this paper.

The AI Agent Index SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:36.189446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:36.189446Z digest=sha256:210d8cada033b030f7a83349e4d3c731e22f54ef1d5a937a89c4174dfd7729de

Observation 752093fe-c71d-4d22-825c-a9376f7e5804 · inbound

Agentic Bug Reproduction for Effective Automated Program Repair at Google cites this paper.

Agentic Bug Reproduction for Effective Automated Program Repair at Google SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:01.842226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:27:01.842226Z digest=sha256:0680270d71c7b95bce0d69316018b8cb07f4549dcb74495b738df3dc78172a68

Observation 54fbf4a3-8ef4-4e54-be6e-a2e33c7b005d · inbound

SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering cites this paper.

SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:11:57.706162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:11:57.706162Z digest=sha256:8cbb77873cb9e74979dddc63e508e711ad5267258ef1a2085fbd353d2b1b8a54

Observation 569e9854-eada-48dd-817c-1d3e64809d0a · inbound

KernelBench: Can LLMs Write Efficient GPU Kernels? cites this paper.

KernelBench: Can LLMs Write Efficient GPU Kernels? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:55:02.102197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T16:55:01.976356Z digest=sha256:ac2c1bc8704be0c72f0efcf8931ba3a60458f285f662d038157e41ff0a361303

Observation 0182c42b-11f8-4bfa-9994-6139a69d8496 · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:27:56.346568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:ff4d49f421c5b655988147e8dc02e918ea8fb83cd34300c29d048e2d95ed2d4c

Observation b52647ef-c141-469c-be87-3b84e5cfc1b1 · inbound

Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement cites this paper.

Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:00.917961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:00.917961Z digest=sha256:205eb8d325ae214424ca25bd7917ca7826b6dbcef0bbbaa979adb88cb0f693aa

Observation bb7b2022-6b02-4d15-b491-9a50cc08c345 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:57:16.238337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:d806c92d1fe61517e4d0806d85183f5c4dd44c076fdc1756bf8f87e31a4907a5

Observation c65b9d81-33ba-4eb1-ae19-eae98d3585b1 · inbound

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design cites this paper.

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:15.130136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:15.130136Z digest=sha256:c2a19f3239dc133dc95d43b4814b7e4f65e8506e4dfe9222cab17722ba31c7b8

Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.468272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.468272Z digest=sha256:c2367a4b2728ed94297bc833b95e84316f22c4a107622a9314568e12cc5d0fe2

Observation fe719371-92cc-4834-a8a2-9e075993215e · inbound

Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows cites this paper.

Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:31.618617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:31.618617Z digest=sha256:6a677b619cb3dc841b9ebc5f412c226c4aaa83190ce4b9da3ff69da0fc2d5836

Observation 9ae7f4e5-42dd-471d-b023-015998974436 · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.357081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.357081Z digest=sha256:92afbac42d3a1fbe7155454d3c7b560f45c9a83daea6032c8fbeff1539052fad

Observation 3789f81f-780c-4954-b3d2-29b734b7aab6 · inbound

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering cites this paper.

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:38:55.375974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T17:38:55.159673Z digest=sha256:b6726c2122273c0ef41a5a1fcd74829f0534ce718a730ae631690cd5e9dc4b14

Observation 0521b7dd-b328-4137-8a6e-81d155a0e877 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.799098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:9a19418045382c55708ebfc2b2f0610f7112a6b63098f995112bee6ab0283ad7

Observation ad34d2a1-8cf0-47ca-92de-de061c9c6bc5 · inbound

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair cites this paper.

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:29:20.689824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:29:20.689824Z digest=sha256:a59d93a886a07fc805f7daae77c3c63a6f6a6b32f5f33a19b72b6f7a563e8e11

Observation c816a82b-f825-466f-9d70-8f6afb83e9bc · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.774713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:f3c6cfaa4f708535a0ef25067de2350512b7625f5815cf0f5c697416f58521cf

Observation 95d2e266-5e50-487d-bc5f-4125e7126727 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:40.623619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:40.623619Z digest=sha256:50cbd215e02f53f34497e2175bf4ea2004898a5f8dde4770fa4964d8a8b1e866

Observation 1039a3cb-282b-4028-9ccb-226845f49d14 · inbound

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios cites this paper.

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:28:24.528054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T20:24:40.939455Z digest=sha256:9dcb9db8552e317d69960bc62483b8f3bb5bcecb91ded6e87ab4e474f82638e8

Observation 66cdd8a2-8d5a-4c27-bc17-cbdd02e06549 · inbound

Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI cites this paper.

Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:11:01.384484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T14:09:33.786576Z digest=sha256:9c6979de80b06cd693cbc75699d3b717d8a3ac43151cda27a2fc0417364c251c

Observation d3ebc9cc-6ce8-49bf-a4b2-957fe8b7779d · inbound

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective cites this paper.

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T23:31:17.421699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:31:17.421699Z digest=sha256:566d9308208cc21326132d418378db03579a7e232d2681db1bc020eada4b7e5a

Observation 6e31f8ef-6b5a-44c4-8f5d-6d244db49fe9 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.877213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:c60376b3d3cd4a6d8d37099dfd10d33b21badf98bafdebd215da4feb2e354690

Observation 7d13fd6b-bae3-476c-9a8b-788fc466c5fd · inbound

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering cites this paper.

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:56:26.738521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T13:26:26.648410Z digest=sha256:1a5ec5877888a641f68e3ff48950c9a4b81c1f42142d6dd3870fa2143b31f5e5

Observation bd71bb2e-2a45-426d-a00d-50930f3ba827 · inbound

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents cites this paper.

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:39.708557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T18:30:21.856024Z digest=sha256:a5b81cf8ffc403851e69a647a2b4b3416a1d3d68122eca85ec624fafe4b859b9

Observation 5356b90d-e7b9-4a66-a319-e9d47b33bc26 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:21:26.539858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:3fa4da8c0c2e239ee3354da6d0f8ae0058383e65f6b263e4ed157684599fc41b

Observation 085178c1-03f3-4b6f-956a-f03900be35ec · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:32:59.504330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:640c16c2adf570dfc13021596f5995935363871a3baf60f9318cf204cf304791

Observation 7a779983-db03-424a-b2f2-10e14520202f · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.348124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:0b190f5a36dc0a029c5507b65dd0bce5b0a75854651b0da9adbd21fa9491e4ee

Observation 41f1a7c4-cf12-47f7-ad24-997f0c3ee081 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:43.750853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:3ceaf2ba3688168a70a5605427d57a79b110dad582dbf56d7be5df588f7da0fe

Observation b5bdc529-81b9-402c-8587-8a1d835b3434 · inbound

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents cites this paper.

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:22:51.863309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T23:06:57.377183Z digest=sha256:7e94115a11676a8fc35318a181b9f4cc1df3eec3a570026943f0cfa87203e13c

Observation ccc90743-de5f-4748-b3b5-bdb7d7df450e · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:42:36.256892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:a0e6dab4f501c344e5379f9ca9caff5b4a6a715061ccb7401dee8bd16b56f065

Observation 33d5f3d5-9b40-45c6-b172-dab231b14e54 · inbound

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations cites this paper.

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.804536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T18:44:12.893087Z digest=sha256:276d14a307b3140fd80da8a8261f73dc062a5817e89cdb7c76ad5c905e2b56b4

Observation bb2696f8-df19-4e7c-90af-a0223aacb4d4 · inbound

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions cites this paper.

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:29.215914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T09:56:36.860369Z digest=sha256:66e757b7a7fd1159ad595732079795fba1adb7a5ee9320a777330b896a460c8c

Observation 38fa32d3-1c5a-4cff-b2b0-c8185b33f454 · inbound

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement cites this paper.

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:27:04.283039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T00:26:22.041924Z digest=sha256:e0536e2a64e299bd6669544041ee31a9b50f8c372e21399687de63c453a40d39

Observation 19d84601-11dd-49dc-aa1c-5df8de9866df · inbound

What makes a harness a harness: necessary and sufficient conditions for an agent harness cites this paper.

What makes a harness a harness: necessary and sufficient conditions for an agent harness SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-27T15:21:00.801996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T15:15:57.372858Z digest=sha256:cd92402c475370d23e1a71b47789097e934f379e7e02bddfacf9e4a940c03f7c

Observation b70e4fc8-eb15-4695-8bf6-f249c480ea16 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 282

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.639775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:b86c9584677553913db4379a4517f98d82cabc98790f3885c0d8f1893a19cb33

Observation 8d02e1bb-f5eb-4d85-a741-0e422d144021 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:02.881675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:65c97533cc3b22a5901a9340df32aeb2fb6b04d7730271fe67e9485efaf5c2d0

Observation 3c8df488-9523-4581-9792-dc4616f0332a · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.805422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:a1fc841ff68cb53ede3c21b5fecf8893e5a7464dd36d5bba6be8bd5f579e6700

Observation 8a038349-a69d-4a13-9a3b-b18c1670b277 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:58.945492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:019f8fcb38a2df2af21b4bda0066539b20ce9d574c04022f45685e88a628963c

Observation 2f80db8c-a37a-4106-98e8-666af8c28a14 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.963144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.963144Z digest=sha256:27263c6c123e4a2bb2146acbd6e94af58038a6f0cb9c4df7ba760ce46da78b6d

Observation 41139610-86fb-4ddb-b748-6ef934e23745 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.236521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:d04f10707455d3c63d025d45ba9c913361790daeffd0d3475db5462a8ad69864

Observation e12773f3-bcc7-4168-9eae-0b672d7bcb66 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.939382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:8e49bedb532621b3705067f0510eebd9d7180958326f7c90fb2c06ce1a4197d3

Observation 59593680-b4e8-4ac0-adfe-660c1024b766 · inbound

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution cites this paper.

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.769919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T20:15:00.517571Z digest=sha256:c2958ce3f8bd9f74afbdc4d333ef1b6d5e88217b804dc5122b8c3ce7d475dcf3

Observation 6c9d3edf-dcc3-4f85-8f1b-787e72c9847a · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.506209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:3b677c2dbe640f8703b88f8c4f68b86e7463a888fa3c0a90561f1efc8097ef2d

Observation f4605d47-9cd8-480f-9cc4-c8952400f1ea · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:46.348120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:46.348120Z digest=sha256:11b8da485d559e8bd2111917edfd4f2739536fad5d1fef8d5f09b49024634b66

Observation c428be79-3082-463a-a539-fe217ca6d6cf · inbound

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification cites this paper.

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T18:29:10.457483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:29:10.457483Z digest=sha256:a9c00bfb65cd426e525662156a4f167929633dbb4e14633a83cd25a47c210bfb

Observation 474521f1-6871-4316-9674-9fed3c956337 · inbound

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports cites this paper.

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:40.658366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:02:40.658366Z digest=sha256:1f2f18ba167e9f7791eff9243dc383093037ab134a413c1f8dee18f32ac9ddc7

Observation a1c247b1-12a5-4bb1-97ce-5f3d7f76cb5e · inbound

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents cites this paper.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.765027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.765027Z digest=sha256:e8ceae3176f87bc44686c65b2a9a9ddd497d895e99743895e5c1c99e95b73bfa

Observation da65c2c5-ff8e-4cb8-8987-d510c8d416a4 · inbound

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring cites this paper.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.005760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.005760Z digest=sha256:04fb2c7d2db90840f1a3212d0ba72aaec11a542a2384500526e244927c644ac2