Pith. sign in

Paper Citation Record · LEDGER

Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2403.19114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19114 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.725599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:34:27.439389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e90f8d23-761a-4087-99c1-79c4c7dbde82 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.129426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:862aa9c51f39fba43f539109f66a9b37b51b68fa9b638a369992fc931a64fa2f

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · inbound

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities cites this paper.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:d3e611b1a18724e0f39d563d601b576a265bf4f36d08732691b99e4118dcfc19

Observation 96d3e6c8-3e6f-4a8a-82ff-6496f9855eee · inbound

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar cites this paper.

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.670461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.670461Z digest=sha256:9edb1eb6e7e98a1b11863c909dfd7414db3f08f80dd7ef190f098a55f610ce63

Observation 774f1bd7-85c3-42f8-b5e7-8a24ba23a3bd · inbound

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review cites this paper.

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:50:16.506793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T18:49:01.097179Z digest=sha256:e8733cd5c9df086aab02ec0eac5fdd140c30859fc4a957d8a5570b1e80ef43b0

Observation 4effb991-0933-4ee7-a21f-50f16cdbac37 · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.205216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:5428ee9cc2157db11c2f59ec1785e599deb0f132da305c18bfd093584dc09ec4

Observation 6824477f-af6f-4a10-ad7f-2b49a9c628a8 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.013236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:27a1ca4d2ac4f69d45cebdc7a45cb474f4511a94ad7e9c16a55cad144560eef6

Observation b7683286-aac9-48aa-a80a-85bb01283e22 · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 56

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T08:34:27.441561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:e5e33f8d564b421923fd528a57b03f7bd7d02da46595c568e45f7d13b4d65395

Observation 6945d7be-c82c-4611-9500-12bf084d04fc · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T04:56:20.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:56:20.002453Z digest=sha256:59b4d8d051cdea58ffc0547c4a1629a9e846236ca02aa1d3f1c78327546157c8

Observation 8440dad5-8fea-4ce1-9b80-024b1dfbed09 · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:43:20.935793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:43:20.935793Z digest=sha256:15a47d84ab0a841286e73371d6adda8a4da57189b512fd845b139369c4a34cfd