Pith. sign in

Paper Citation Record · LEDGER

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.09802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09802 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:33:05.202398Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved45
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0671fe7f-6128-41ee-a384-16e3157ea835 · outbound

This paper cites Advances in neural information processing systems , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:04.983947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:04.983947Z digest=sha256:7bfaa477592979b945b521df990625b117b4dc8b4b40929d42c30d849c1fc09f

Observation ba4065f9-2d11-4e8b-b181-cbb4471e3da7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:04.987588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:04.987588Z digest=sha256:cbc9a6ae88da7a853a6604f80e5fd123a9404be64008791b72e4e5a552aff321

Observation d8b8dc27-fd28-41e8-8f99-2d590b709fe5 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Measuring Coding Challenge Competence With APPS

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:04.991376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:04.991376Z digest=sha256:28aca6ee1acd29172ea84182a14173a903456e2353352d8f786c413cae4764fe

Observation 7f6dfb29-f4cd-485a-afc5-f8a68a7cbc57 · outbound

This paper cites Program Synthesis with Large Language Models.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:04.994874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:04.994874Z digest=sha256:4a5458d45121cc0d1683a6a4c5ce32364958ac7a13186c0aa639c0b1092e0b14

Observation 1b7e1822-6621-44a2-aef2-1dee6c9af36b · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:04.998387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:04.998387Z digest=sha256:258eed7d8bcd6fb7e3dd3c1b4eff813a40d63f7101779a33ca951cfd9cc9a065

Observation ba8f3c37-19fd-4e78-98f1-23491d8a0227 · outbound

This paper cites International Conference on Learning Representations , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring International Conference on Learning Representations , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.002065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.002065Z digest=sha256:45b6f9ce409b0968222229b669f5aec22b0563bc4e999daeff3237286ae4f638

Observation da65c2c5-ff8e-4cb8-8987-d510c8d416a4 · outbound

This paper cites SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.005760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.005760Z digest=sha256:b166ddc0b32cd5e9f5f3dca93182e104fcd7c9637d8edb872c01293f43a978be

Observation 984a445d-8033-4f82-ba73-4789abb31792 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.009583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.009583Z digest=sha256:a6ef75d89139bb9902d36e1d1e92f9edb394ab8b59e4f924b05ca638512c8fd2

Observation e61cd33a-7aa9-438d-9865-eb2f6cb247b9 · outbound

This paper cites SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.013212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.013212Z digest=sha256:336327e6dcf79711809671314ce04aab7cb5b131fa2a99da6443c130d3ec2357

Observation c50d99b2-2016-4d90-9ff2-20ba87f1c09d · outbound

This paper cites arXiv preprint arXiv:2505.20411 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2505.20411 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.016766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.016766Z digest=sha256:715517128934c28af064b157a46fabfe9031072417548ec53f30b6b2f6d86dad

Observation c0d96752-88ba-4c31-8c76-00e5ddae7d3e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Advances in Neural Information Processing Systems , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.272400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.019960Z digest=sha256:a5e0c3a32733df365ef4f82d9335e8875bf9d4d1abfd7c8c3855ae9d8e71e4eb

Observation 42608fbb-d9c5-4640-92ea-c7c511186ed0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.262675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.023134Z digest=sha256:361f312f7c714e0fb58922f31a3155bcf72980393c068e6c7aacab2d1782cfa8

Observation 979c20a9-a755-42df-b7f3-8fced91b6e1c · outbound

This paper cites arXiv preprint arXiv:2506.10954 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2506.10954 , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.026501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.026501Z digest=sha256:d5a22740fdb0650c12f9a0df103e524d471447430518a73e704f2911bf5d99b8

Observation 621873e8-03d4-4472-893d-c4d838f58423 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.253300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.029600Z digest=sha256:67697ad2f357262cafeee8d231073ee728f18406378f2666e80e6d11fc745a29

Observation 4d37490d-44e5-4d0e-aa94-90a9af43c851 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.032815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.032815Z digest=sha256:d541ebca51d2a300945b5e44f6b31168388d341ef30ce8d20b7d88eb8afdf2af

Observation d4aaed29-35da-48fc-be1e-9e68ab3a9048 · outbound

This paper cites RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.036133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.036133Z digest=sha256:1b46ac72e6c1aff2a984544ed72411c1b8f69ab1ee296d4373def78219cee95c

Observation d6a52527-4614-416a-8710-bcebae9d8cc6 · outbound

This paper cites arXiv preprint arXiv:2602.03712 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2602.03712 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.039582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.039582Z digest=sha256:25e4640bb6c1a60cb78f086954762c75c0c60b640e4b678bde53c8da9a6e82ff

Observation 25e89b9f-932b-4eac-8061-950416898b20 · outbound

This paper cites arXiv preprint arXiv:2511.04824 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2511.04824 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.042611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.042611Z digest=sha256:0caa73bbaf12f526e7726a105d4f188c5169482490d909d8086fc4a54a7c7897

Observation ad755029-fadb-4876-9b01-0b3fe4063600 · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring IEEE Transactions on Software Engineering , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.242884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.045569Z digest=sha256:7d32cfff61a6e22d7e295dc89b36535aca631706203d55cc972817f617640c36

Observation 8cb2ddda-2b35-443f-a163-d793dba49ea9 · outbound

This paper cites IEEE Transactions on Software Engineering , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring IEEE Transactions on Software Engineering , year=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.232370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.048552Z digest=sha256:f81205080e0828f6bb7ef99eee9c6506e909c2fc2c0d2fbec41000759ce2bfc0

Observation 40df2ad2-d5a7-4e39-996e-e87e448198d1 · outbound

This paper cites Advances in neural information processing systems , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Advances in neural information processing systems , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.051716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.051716Z digest=sha256:3447a289d8439aac85d1312062053f203d6f01498aff703b10feb535b16c7f40

Observation 01a76607-3310-4624-8f98-49138acca355 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.054754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.054754Z digest=sha256:1b41dd3a1bc67167a8c4b2a1aa3335696d07cd0d6fd6e17e6bdad43d72599482

Observation 02f6759b-a952-4f8f-ba1f-98ad898b5125 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.058082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.058082Z digest=sha256:cd0ca1792df814d0beb3c7e0923caea28da25dd7116988c47d0d51477eac9216

Observation ece6602b-a382-473c-8e8a-48fcc6dc7e16 · outbound

This paper cites CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.061353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.061353Z digest=sha256:3f5d092421a95c1edc018d3402fcc6b8977a379e59cecd4cd31f70f05752ad1b

Observation 2a255c00-b754-4ee2-8dce-90c2a3be19dd · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.215860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.065007Z digest=sha256:568bbd410d04e741283bd1e557b7f9ef57ea4f7abfaecf830928e1205519e945

Observation db050c1e-b44e-40e8-8c10-cd973fec65ee · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics , pages=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Proceedings of the 31st International Conference on Computational Linguistics , pages=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.205727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.067844Z digest=sha256:153563dec0a62718daedd5a20e61d382d2cf50610ad4896ee4bea806759a0485

Observation 5318006a-30e2-46d7-993f-25f879d27782 · outbound

This paper cites an unresolved cited work.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Unresolved cited work

Reference 27

Resolution
parse uncertain
no resolver link, observed 2026-08-11T10:33:05.070967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.070967Z digest=sha256:0acd8506d012e27fbefdd41ccfc07046835e67925808f15ebeac9732632d854b

Observation 443b405a-d8e2-45b9-a100-ab8637ec457f · outbound

This paper cites 2025 , howpublished =.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2025 , howpublished =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.190967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.074298Z digest=sha256:16d7c8569ebb2094d0463a81b94cf41724aeff9cad2f4734a431fde24a8913e8

Observation dfec91d2-8b0a-408f-bc2c-059ff393977d · outbound

This paper cites 2026 , howpublished =.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2026 , howpublished =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.181601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.077625Z digest=sha256:3cd68bbf22e631b44e46a307bea4526d9cc1717db96cb0a527e4024dab025044

Observation 72f292ff-db2f-4c91-89c1-74bc5d4d15f1 · outbound

This paper cites How We Compare Model Quality at.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring How We Compare Model Quality at

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.172541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.080506Z digest=sha256:51547b1f70d1e2bca475e83083df2804b8b6c4f4b2d8d97c6362709843ecaa2e

Observation b186b56c-43b8-455a-a9c2-40ec02736268 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Measuring AI Ability to Complete Long Software Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.083332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.083332Z digest=sha256:8b226ca015fbf8deb8c0996bcdcc141c28c70a83f1b8aa03ac020a67ddc6e698

Observation e4842b13-56af-4084-b7e1-253247be46b4 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.086937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.086937Z digest=sha256:56c9a3f620536b33e3798e199e2e362c89d740d92950d8b33cf012fc38a2c90a

Observation 8a5184c6-c769-47c3-9a81-73ee291514b3 · outbound

This paper cites Qwen3-Coder-Next Technical Report.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Qwen3-Coder-Next Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.090287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.090287Z digest=sha256:374f5fb3f82d1053a753d13002ec1503961a517c87d0737a1523b2e1ecd3fd80

Observation 731ffd53-be13-4f03-a94d-8e327639f055 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.093659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.093659Z digest=sha256:cd9ec3c6dada3496c423c58f1bce06202c6b5fb1cbb881aefbd19606c6726813

Observation 8230e813-5a2d-47ef-ae72-f0b568d5a1dd · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring GLM-5: from Vibe Coding to Agentic Engineering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.097132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.097132Z digest=sha256:6c34bb211b0e186e2de5df7adc5269c8da3ad664caedd4818ffb49f657d3a519

Observation 6c5135ea-a973-4a49-8c48-c8e5152a42b1 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Kimi K2: Open Agentic Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.100868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.100868Z digest=sha256:0f9f7343ed4024c2692ba75cf1c4db0ea8082d7c3ee5077894d451c829667b56

Observation 406a1acd-ab51-467a-afd4-b8e6e7699bd8 · outbound

This paper cites OpenAI GPT-5 System Card.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring OpenAI GPT-5 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.104283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.104283Z digest=sha256:1667a94883f1645788d22dc9ed9fdd42a83ec6245c5d0f9f0abfde58c816d01e

Observation 1748a498-1845-4bda-a114-f5f25eecee9a · outbound

This paper cites 2025 , howpublished =.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2025 , howpublished =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.156322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.108179Z digest=sha256:47a40fa7b5499e26f32159483acdfabef8fb55c18400b2325a9151ccb73ad5c5

Observation 34c565ed-c927-413d-9fc7-8565157b2451 · outbound

This paper cites 2026 , howpublished =.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2026 , howpublished =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.111771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.111771Z digest=sha256:a8680a87f8b42ccee92aaf2405eb8778c5fc07568cab46fcd10296a7865c1722

Observation 459d61d5-ea68-421a-9083-f6c51ca8b05d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Kimi K2.5: Visual Agentic Intelligence

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.114887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.114887Z digest=sha256:d91c2865a460614586e43e1e6417390a9eafcaaf10a4105216b4bc00fb33ea78

Observation de7e9c81-7baa-48c6-8698-31fe1f9475c3 · outbound

This paper cites The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.118894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.118894Z digest=sha256:4af543ea99f0c32a6b647a8a09b01a4195f2d6acfa3983aa9c25619752769323

Observation 9bb98ca7-5641-4c58-8757-34fb9cbf61fc · outbound

This paper cites International Conference on Learning Representations , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring International Conference on Learning Representations , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.122258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.122258Z digest=sha256:878f929fa452ea85b613faddbd2cdec598b7091a4a826dc29a751e334dce9eef

Observation ffdf1e1a-ef57-4b7b-b636-7aeff17513f7 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Agentless: Demystifying LLM-based Software Engineering Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.125611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.125611Z digest=sha256:4dbb7cbffd748f60ef597fef325f85dc51e8fe1ae89133f9f3b6dd2616d90f15

Observation 8cc28ad2-853d-46d2-bbfd-95ee42cf3fb0 · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Forty-third International Conference on Machine Learning , year=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.131821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.129227Z digest=sha256:cad1b0381697a8ef5d0427825eb0530c2db682973c3b11e09f3f875fcedb64ef

Observation fedb4a3e-bb25-4824-bec6-3ab8f9c0ed4f · outbound

This paper cites ACM Transactions on Software Engineering and Methodology , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring ACM Transactions on Software Engineering and Methodology , year=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.121704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.132829Z digest=sha256:0bb963818e6d594a3202a4e0e64df6010108ceeb579823c440382615ca0d1f24

Observation 911cf3f4-d1a3-43e4-b5bb-1a855598c1dd · outbound

This paper cites arXiv preprint arXiv:2511.21788 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2511.21788 , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.136319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.136319Z digest=sha256:0f8e95d7284de71f57651684dd25dcfdc0ad8bf977fecc80df1861714b8d5d59

Observation 50db8eec-8615-47d7-b8b2-ac6cf7d6aab1 · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring IEEE Transactions on Software Engineering , volume=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.111818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.139162Z digest=sha256:215483e039c4613bde1e695910f7008444cfba89fd5e3400d975d5d1bb94151d

Observation a70a37d2-dabb-4c9b-9b66-d9a51647561f · outbound

This paper cites SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.142802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.142802Z digest=sha256:ca4802f9be3e2b411ebfaf404922393800a4745e191c924743da3578a54eb68f

Observation 0133102d-002d-46d6-8134-d2c2cc1674d9 · outbound

This paper cites arXiv preprint arXiv:2508.05988 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2508.05988 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.146584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.146584Z digest=sha256:062b39b98f8938b2a5f43f0056c75f856c0925a9ee1cf6afd99d999776ffe4c2

Observation 6956c2d6-5a6d-47dc-9e2d-c31fc1b6cabf · outbound

This paper cites Proceedings of the ACM on Software Engineering , volume=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Proceedings of the ACM on Software Engineering , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.100412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.149726Z digest=sha256:35d1cad6302cbb1310d01b73d9f1130eb0c67db9799c6ff45c894d78d7657ab1

Observation 86c062fe-b170-47c3-816b-b40cb2bcd596 · outbound

This paper cites GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.153067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.153067Z digest=sha256:560289864f2ea6cbc91880f17d1ce1640c73bea7d3f6df2a3a8aced229b45da1

Observation 1a80bc5a-a564-4cac-917f-6b6d4ffe40bd · outbound

This paper cites Dockerless: Environment-Free Program Verifier for Coding Agents.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Dockerless: Environment-Free Program Verifier for Coding Agents

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T10:33:05.459249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.156297Z digest=sha256:48ac71a266ab865ff3c6f1a9b8735db0eb8eada2b8519caa63c0d9266a78eca3

Observation b3298fe7-8a0e-4ee9-b842-f1631abf9daa · outbound

This paper cites SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.159719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.159719Z digest=sha256:990f777a469e5d2b8c8e88741aae1d7330313b7c000414d200383468b86006d0

Observation b7a9525b-0127-4587-8ccf-24c3de93cbc0 · outbound

This paper cites 2026 56th Annual IEEE International Conference on Dependable Systems and Networks (DSN) , pages=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2026 56th Annual IEEE International Conference on Dependable Systems and Networks (DSN) , pages=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.089954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.162960Z digest=sha256:0d017d28ce2583208e3aee67f6499c11d802cbaff263161eac6731421a755a49

Observation 721849b7-af00-4ab7-b0fb-945c222faace · outbound

This paper cites , author=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring , author=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.079551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.165940Z digest=sha256:503000229de41895edad8a982e1bf306be2fbdf93660ed97e534d3486a379b9d

Observation 68561b30-bfd6-4eca-b5a7-f71b14a3cece · outbound

This paper cites DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.168600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.168600Z digest=sha256:5db93b07191f9d0232aa1d4cda9286517c681d19875fbc8505ed57b8b320e346

Observation c39e42ed-83c4-4a40-90f6-032e842bc57f · outbound

This paper cites arXiv preprint arXiv:2602.09892 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2602.09892 , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.171640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.171640Z digest=sha256:7d4b61d77eb56649a2e0421f87e304171e795c524e11fb2e6b022ee70ebff474

Observation 2d2557dc-c8bd-4b1c-8229-0e96dbda32ca · outbound

This paper cites arXiv preprint arXiv:2603.13023 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2603.13023 , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.174132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.174132Z digest=sha256:dab614930fcad110b353482a53a6d484ddb1e7a9edb2e14c943a0268c739a04f

Observation f289ac6f-bd83-4ed4-b5b4-87df9830f8c8 · outbound

This paper cites SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.176471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.176471Z digest=sha256:33be5e9b18b8b6912dead4b145d9de30fcd979d88480851842d78ab1c4eaff67

Observation 37b1c540-47f4-44c3-b6b8-3fbf94f47fb2 · outbound

This paper cites 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE 2025) , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE 2025) , year=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.069679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.179259Z digest=sha256:b2dbcb3e18c6197d785b7a2b52d3951ae661398923691c2f25b4b137255143d8

Observation 482b1260-b00e-4acb-a488-0cebbb20a2f8 · outbound

This paper cites 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE 2026) , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE 2026) , year=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.059866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.182108Z digest=sha256:ce1b0e9eeae5108e3508e5e8ed912417e3edd2a804540c60019dc5009950c46f

Observation b455b31f-4f54-4dea-8092-a73530b9f2e8 · outbound

This paper cites SWE-QA: Can Language Models Answer Repository-level Code Questions?.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-QA: Can Language Models Answer Repository-level Code Questions?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.185551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.185551Z digest=sha256:a7d377dc47ebd63767ed889baaa78fdefbaa197b7e9719e62ec1e48c75340887

Observation a5c3626a-90c9-42ee-85f2-10a2eb169b66 · outbound

This paper cites 2025 IEEE/ACM 40th International Conference on Automated Software Engineering (ASE) , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 2025 IEEE/ACM 40th International Conference on Automated Software Engineering (ASE) , year=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:33:06.049571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:33:05.189108Z digest=sha256:3bd6386e5c1025c4a6cccfe859f94137cbfe1edcf75fbc52afca37b79fc3e27c

Observation 3feab7a6-395e-4bdc-8421-cbf47fc53339 · outbound

This paper cites CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.192539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.192539Z digest=sha256:c11c871f7f1be4f0d0e09c259a89312b0945a2c5cdb11f541787431938b9f8a8

Observation 8ef10c72-4675-4346-ad97-1eaa3499ffc7 · outbound

This paper cites arXiv preprint arXiv:2507.23361 , year=.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring arXiv preprint arXiv:2507.23361 , year=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.195925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.195925Z digest=sha256:4411c42f298fd0c7e9117b16afa18560c9be3fcae5e1a694b9fc4ee21da5b522

Observation e7360d33-bf91-44ff-a54d-97619353261a · outbound

This paper cites SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.199145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.199145Z digest=sha256:c12e2c8a566a8f85ec00e970e92956d50e1ebcb4edeab1c8882b3fbcae633844

Observation de7094cb-f340-4247-b194-167c88acfb46 · outbound

This paper cites SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.202398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.202398Z digest=sha256:9ae56c7b32776c646956d982427907ec9afa6a5294208e1b0c53fa7e443cef77

Pith citing papers

No inbound Pith citation observations are available.