Pith. sign in

Paper Citation Record · LEDGER

Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2410.09997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09997 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:46.990851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:36:47.499426Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 37340459-c733-4faf-be3e-941b902b7aa6 · inbound

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments cites this paper.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.990851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.990851Z digest=sha256:c14bd8490885abb54e065ca8258e4e0665d2b6810f64d1119c45951306eb5947

Observation 07e8ff9b-35bf-4fd3-9922-edc2abeed53d · inbound

Neurosymbolic Repo-level Code Localization cites this paper.

Neurosymbolic Repo-level Code Localization Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:27:51.726148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:26:02.980824Z digest=sha256:02c69f1933b04012c7227361401a1165adda78494af96f9742ecad8ed11d21ed

Observation 678c4f60-7473-4c21-9e20-9eae3c4e2e64 · inbound

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification cites this paper.

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:27.728041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:23:01.482449Z digest=sha256:0dd99da955ed6d61b9ed00fddc0ffbbb5c40ab4dd3c582bff7d8d58342f99ba9

Observation 76cfa23d-448c-403b-8cd2-f22288a74bc8 · inbound

Knowledge-Enhanced Agentic Vulnerability Repair cites this paper.

Knowledge-Enhanced Agentic Vulnerability Repair Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.501532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T08:34:20.374717Z digest=sha256:1f398f849ca3264e9ded61892ea880d679324301f4c8410102a74c53f405c537