Pith. sign in

Paper Citation Record · LEDGER

NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.04520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04520 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:56.926699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:42:36.222980Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa7df653-a6fc-4082-bd95-810bf5c01e09 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.715752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:944478646c7571c1e4281694dc52b38982b316562983ddd6995f971f0eadf1b3

Observation feece394-c6f3-426e-bf88-a55688ba4792 · inbound

Seed-Coder: Let the Code Model Curate Data for Itself cites this paper.

Seed-Coder: Let the Code Model Curate Data for Itself NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:56.926699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:56.926699Z digest=sha256:e6a89d55d99670f012a59198d57ad6b74d491dacb3dcfa5dcfab95a85ea67257

Observation 852763eb-388f-4f56-a45b-c50e763b21d6 · inbound

IFEvalCode: Controlled Code Generation cites this paper.

IFEvalCode: Controlled Code Generation NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:34.139423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:44:34.139423Z digest=sha256:5b71d00c9405caa671ad89f2c31df01e4231b15e3ee5164f0ce50c12227ce0a4

Observation 683694ca-d5c2-46a4-ac32-93202e1c86ed · inbound

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference cites this paper.

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:17:50.333004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T16:17:50.215935Z digest=sha256:16319e2fb39a05ef0e7d4a90d045d198f71a5e0866c900c8c4be2e7a057be8d7

Observation 71b73a86-1486-4778-a990-b970b84c1c16 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:28.467608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:28.467608Z digest=sha256:1bfc2b2e0288d74ff683cae361d750e305dd85695ca9ea21525b61eceb41e8a3

Observation 04c5bdb8-117f-4e82-9696-6c2fc7fe2736 · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.224566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:2b93dbb190a197b218d10ee030dbad2a3a45d21067f2fbe05ea0ed9b0f4b86a4