Pith. sign in

Paper Citation Record · LEDGER

CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2404.03543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.03543 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:26:32.137744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T09:56:51.748338Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65ce7920-4ac8-47e2-987a-56bd2a8ff0f8 · inbound

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks cites this paper.

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:26:32.137744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:26:32.137744Z digest=sha256:ba1839f0c945c8322a6dfffc82a739d116af679c4aa596963710c265950149f9

Observation 65491e5d-6e7b-496d-837a-10ed59f3db6d · inbound

Coding Triangle: How Does Large Language Model Understand Code? cites this paper.

Coding Triangle: How Does Large Language Model Understand Code? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:32.548428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:32.548428Z digest=sha256:6167a879fa8ecafc46029bab98d3da002339176546cfc1cd67f3ad6c9f230957

Observation 97c5f416-faaf-41ee-9f3b-f77d71468709 · inbound

SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? cites this paper.

SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:54:56.877947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:54:56.877947Z digest=sha256:d9be92e71b13bdd5c9b207550b11e0649f998fc444a9dcf9019868f2da562101

Observation fa5a5a28-0ac1-4775-851b-8c5c793c757c · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:26.902628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:26.902628Z digest=sha256:cd11138841190db20cea9353b03d00369e462b92b3135662885c20a6d133af80

Observation 72a0df32-de3b-4480-a656-caafb8363ce2 · inbound

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models cites this paper.

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:28:36.592404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:28:36.592404Z digest=sha256:9a8c51c8753785da55a33ce43cac8abe5a4d70f202f4c6cf4b650ff35a1ccaad

Observation 965dca21-848b-4ea9-bf07-7bb7c331be2c · inbound

LLMs Corrupt Your Documents When You Delegate cites this paper.

LLMs Corrupt Your Documents When You Delegate CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:47.657487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T09:47:21.966292Z digest=sha256:5e5a66648e991a757f48c95cc7d7145c09bfa1cf5db55a1dbcf22cdc67e70254

Observation a55143d8-6a5f-4653-bee1-5fdb49768ba1 · inbound

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? cites this paper.

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:56:12.440718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:03:35.146431Z digest=sha256:4742fc163666e410b247b811f501ada6293f08bd305881099d6b72f533710b08

Observation 1a81398d-e544-471b-afe4-6ccada41dc78 · inbound

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? cites this paper.

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:07.540349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T10:44:55.943351Z digest=sha256:214d8e3b913a1c4f3f0f99e2aa9ad8025ddf182ff1f2a3254b443686492ef1ed

Observation b71b1bbb-7251-4113-96c5-9d1106809ec3 · inbound

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code cites this paper.

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:56:51.749829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:21:34.772199Z digest=sha256:47360422f53b6e80df65c592f1a3b75e6eae7d956daaf031dae0cbe4c850806a

Observation e41cf113-8493-4b92-ac55-2bfbd5a3899c · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:3749acc336c38a0ac89a92e2fe7a849ac6d28e2e8401019ef3885bca5e4f63d6

Observation 46bceb00-e12b-4605-9632-16673c4d605a · inbound

SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training cites this paper.

SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:42.284447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:42.284447Z digest=sha256:daf16ee699094dc98276173366a733b839c33ec84a778ac2841dfb28f48afd5e

Observation b049bfd3-e361-4861-89e0-d91457d9ee0a · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code CodeEditorBench: Evaluating Code Editing Capability of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:26.829811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:26.829811Z digest=sha256:53b3e41c0f798ac142e508cbc69f53bee01a49f3a50d76bcc88afb722c6e69da