Pith. sign in

Paper Citation Record · LEDGER

Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2403.19114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19114 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:10:37.247818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:34:27.439389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e90f8d23-761a-4087-99c1-79c4c7dbde82 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.129426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:e7a353d50e340e319b26eff9b6f151730e61f0399a13a518e2f3cda0ad84d2ca

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · inbound

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities cites this paper.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:25de121e27b469d1ddb50e7ff67719371d97f31578651e6e3129ff23f4bc55aa

Observation 96d3e6c8-3e6f-4a8a-82ff-6496f9855eee · inbound

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar cites this paper.

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.670461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.670461Z digest=sha256:bfaff10cf4d1c08f1bb20319cd5fff28a2e3d7e8c8e59ddde6c2aa1b54a3769c

Observation 774795ee-b9cd-4c71-b77b-31099f857353 · inbound

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment cites this paper.

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:37.247818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:37.247818Z digest=sha256:70a46711b559b31d26a2981b8ab36fc9db6bcce7c7a990af2137845532e927fc

Observation d2961752-69f7-4042-b9f0-395cb5a083b5 · inbound

Combining TSL and LLM to Automate REST API Testing: A Comparative Study cites this paper.

Combining TSL and LLM to Automate REST API Testing: A Comparative Study Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:38.548821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:38.548821Z digest=sha256:b003b2b464c1c22a6ee81404bededcce4ddb514355dd19329dda334d2edf764c

Observation 774f1bd7-85c3-42f8-b5e7-8a24ba23a3bd · inbound

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review cites this paper.

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:50:16.506793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T18:49:01.097179Z digest=sha256:8b94c0c3be7290d78b1a5afe8ece60aa05515134cb7481e849fc38343776d147

Observation 4effb991-0933-4ee7-a21f-50f16cdbac37 · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.205216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:e2737f69bf68e1c2fcfb96c53c0bfe9e027068329c671d750cd5dea9d609e778

Observation 6824477f-af6f-4a10-ad7f-2b49a9c628a8 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.013236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:b2aec00ecd573a8633426fda782837e325f97a67adc452bc72349021e19aa66c

Observation b7683286-aac9-48aa-a80a-85bb01283e22 · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 56

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T08:34:27.441561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:69e0f01f6837c0aa5248e46742ce4fcc6c8deee12c237a0d289115f65d3d0546

Observation 6945d7be-c82c-4611-9500-12bf084d04fc · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T04:56:20.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:56:20.002453Z digest=sha256:f53ceedef2a016642d1612ff5f9a416106d3955599b474aaef5d090986f6e7db

Observation 8440dad5-8fea-4ce1-9b80-024b1dfbed09 · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:43:20.935793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:43:20.935793Z digest=sha256:7784b61e30393ccff473b3345fe616653aea50cf475dc4024deb18633453d0f2