Pith. sign in

Paper Citation Record · LEDGER

On Leakage of Code Generation Evaluation Datasets

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2407.07565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07565 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:19:26.247246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:01:17.068967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 400562a5-4b06-4464-a87d-5e112fa05651 · inbound

CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit cites this paper.

CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit On Leakage of Code Generation Evaluation Datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:19:26.247246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:19:26.247246Z digest=sha256:b4c12fec26b32ec46660f9d9f71e13bd27efee2aa7fd593ce171257f0ff3a864

Observation 5ac72d58-a6b3-490b-a98e-ff9070808409 · inbound

If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs cites this paper.

If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs On Leakage of Code Generation Evaluation Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:46:25.087729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:46:25.087729Z digest=sha256:efce9e7e9a1c75ca81f93811cfda3554d7297e77c20590430634a14f41b3788a

Observation 2058e9b4-805b-45fe-90bc-afaf688edf9b · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation On Leakage of Code Generation Evaluation Datasets

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.674987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.674987Z digest=sha256:f131c23c7e963f59b4af13c88e3d6988e99beb8af671218bebe5aeabe383f6b0

Observation 7a0ac1f2-210a-47c6-9806-17bf56a4be43 · inbound

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models cites this paper.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models On Leakage of Code Generation Evaluation Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.683526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.683526Z digest=sha256:8b860a2d82662ccde46b9d59dbd88575f1076037ef724116610751fe747db7ed

Observation e008c9ec-d0dd-42a5-b23d-e94421bec760 · inbound

Disproving Program Equivalence with LLMs cites this paper.

Disproving Program Equivalence with LLMs On Leakage of Code Generation Evaluation Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.846372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.846372Z digest=sha256:bef19fa301e153baf2f23e78aeae9ee7e208f27fc69e3411f4917e1f4204ee69

Observation 3385d443-8d5b-4b22-b902-f41d71ab10d9 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis On Leakage of Code Generation Evaluation Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.323144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.323144Z digest=sha256:738378124fc8c8492d3f0cc188d1b9c397773c904147c89c2c2c139a55895b7d

Observation 515d0685-b149-4a0f-a75b-7d8e1e6cfc92 · inbound

Evaluating and Improving Large Language Models for Competitive Program Generation cites this paper.

Evaluating and Improving Large Language Models for Competitive Program Generation On Leakage of Code Generation Evaluation Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:28.369986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:28.369986Z digest=sha256:90ea93f3d39875cd9e4c5c4a82352f9c8a57dcdddf6f9fb7ddd05e5bc64626e1

Observation aee95852-047f-4726-848b-90c4e94a2711 · inbound

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation cites this paper.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation On Leakage of Code Generation Evaluation Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.436986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.436986Z digest=sha256:2828991a4569f7a89aa31208ecce4e8511f99cc07c3d0e76ade434584f2d4350

Observation 32ae7065-6e4d-4ea8-b263-25f7340ca3d5 · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners On Leakage of Code Generation Evaluation Datasets

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:01:17.071706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:69fb69688821d177b49cca88117b2e1a9c761d138e2c73a057423ad39f4f3b14

Observation d5d6489f-1323-46d6-80d3-670dc0c6d0e6 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code On Leakage of Code Generation Evaluation Datasets

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:11.090250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:f80ab6a564f0068a643e9179efe9dca055a890d41e25c9d67ef46886864caedc