Pith. sign in

Paper Citation Record · LEDGER

HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.06918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06918 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:58.505581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:05:10.822591Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c618f911-c028-42e7-98f4-592bb192423a · inbound

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility cites this paper.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:26.001569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:26.001569Z digest=sha256:63ec661928afc15426eaca694bc92de91e6b8b312d49957d40b99ae26b7a950b

Observation 42976470-38cb-4eec-936a-bfec9dca0d8a · inbound

OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution cites this paper.

OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:58.505581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:27:58.505581Z digest=sha256:dce1a81f58bc1269644038dc7d9891c8d70976a6db6b06b7de9ab1b8b29c6408

Observation a016940c-4a33-4e26-af3d-37407ca86c7d · inbound

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation cites this paper.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.535532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.535532Z digest=sha256:eae7ebbd7bc3c87eff7b4a4153e7a9a9829316034226f6249601ef3406529f11

Observation 2a1ede90-41be-470a-9b78-4a28d183bb75 · inbound

ChatModel: Automating Reference Model Design and Verification with LLMs cites this paper.

ChatModel: Automating Reference Model Design and Verification with LLMs HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:44.471461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:48:44.471461Z digest=sha256:a7ec3bdf6986c236a8abfce9bd3f37c6dd796af836076abbd1d1b69aeb2d117e

Observation 1c8d4fd0-0b17-4e1a-81da-a90ea6095dd3 · inbound

Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation cites this paper.

Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:05:10.866913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:05:10.431735Z digest=sha256:6f9be7aba92a44ab2f14c8106440c8e8dd40bd4a5eff4b529413e0a9e4bdb6fb

Observation e83dd308-7d55-4fc3-8487-4444bf1107d0 · inbound

CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems cites this paper.

CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T05:31:41.388721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:31:41.388721Z digest=sha256:672b8200437b7fabdd886cb28ae19c995ba25b37c8f4d7ed40e46628ea482c96

Observation d93ba1a4-9c8d-4fb5-87e6-87fc00b28b0e · inbound

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents cites this paper.

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:55.790898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:55.790898Z digest=sha256:4f09a1d3b0d03a7e91888616f98275fc82d335c9d24e32cbe268018e3108279d