Pith. sign in

Paper Citation Record · LEDGER

CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2404.03543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.03543 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:26:32.137744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T09:56:51.748338Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65ce7920-4ac8-47e2-987a-56bd2a8ff0f8 · inbound

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks cites this paper.

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:26:32.137744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:26:32.137744Z digest=sha256:b4daa0b397dd725f71bf43ea5a4568e123ad2eaf20966538c33ef3c95f17c494

Observation 65491e5d-6e7b-496d-837a-10ed59f3db6d · inbound

Coding Triangle: How Does Large Language Model Understand Code? cites this paper.

Coding Triangle: How Does Large Language Model Understand Code? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:32.548428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:32.548428Z digest=sha256:ca767c5424090f50f2c204fb6c67066381760c18f7963ba8954965faed6d9080

Observation 97c5f416-faaf-41ee-9f3b-f77d71468709 · inbound

SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? cites this paper.

SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:54:56.877947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:54:56.877947Z digest=sha256:eaab4c593d174d5ef289565bbba08420ec30c21f5e4c5fb11b58fdb333044c14

Observation fa5a5a28-0ac1-4775-851b-8c5c793c757c · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:26.902628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:26.902628Z digest=sha256:d8fffcb517a234a715854660634b4b367d47b122cfa043845c212f315c8d5423

Observation 72a0df32-de3b-4480-a656-caafb8363ce2 · inbound

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models cites this paper.

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:28:36.592404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:28:36.592404Z digest=sha256:452b0d5b346ffad1f3fbeb795ed4e1e0efb5fc3f8ac1e2d0dff051ed82405b22

Observation 965dca21-848b-4ea9-bf07-7bb7c331be2c · inbound

LLMs Corrupt Your Documents When You Delegate cites this paper.

LLMs Corrupt Your Documents When You Delegate CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:47.657487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T09:47:21.966292Z digest=sha256:b943769c5c873592e84aea61d64bcc3c10428238cc6110002eb5f053cf84f7c4

Observation a55143d8-6a5f-4653-bee1-5fdb49768ba1 · inbound

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? cites this paper.

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:56:12.440718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:03:35.146431Z digest=sha256:4219870cba2b28236dea174c4b85c91b3f13de3200898c044faf5ce55a9883cb

Observation 1a81398d-e544-471b-afe4-6ccada41dc78 · inbound

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? cites this paper.

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:07.540349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:44:55.943351Z digest=sha256:c2add1a5eca9b3e35d5671ec3b642178db6ca2085933ea5baa2a8930f867e6aa

Observation b71b1bbb-7251-4113-96c5-9d1106809ec3 · inbound

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code cites this paper.

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:56:51.749829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T05:21:34.772199Z digest=sha256:a698f05a9dd17d47e1e467e98dfcd40d9dbe069ba065162d72c030e5f68273e2

Observation e41cf113-8493-4b92-ac55-2bfbd5a3899c · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:0cc8e0d086e7e8f7e7a6ba74058716fee578b2d4cc46ab76b5af507a244840ef

Observation 46bceb00-e12b-4605-9632-16673c4d605a · inbound

SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training cites this paper.

SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:42.284447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:42.284447Z digest=sha256:80dd8f12de3afe485a1d75497a2df651ac85815e452f68080c05dc1a2dcd0ebe

Observation b049bfd3-e361-4861-89e0-d91457d9ee0a · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:26.829811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:26.829811Z digest=sha256:8b83e84792307a4d17cf7835259044323bd31ad0a4fc9a4e14518bf062ebcbc0