Pith. sign in

Paper Citation Record · LEDGER

Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.11470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.11470 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:10.690842Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 936321a4-c108-462a-bd8c-590b38ec447c · inbound

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics cites this paper.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.690842Z digest=sha256:b29c66890ef310738b532354ad980152fc826dfc15423497cd503eff8d07af67

Observation d2902e08-b1bb-4dfd-8b28-1bae3b13dc2d · inbound

MetaLint: Easy-to-Hard Generalization for Code Linting cites this paper.

MetaLint: Easy-to-Hard Generalization for Code Linting Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.872408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T04:07:31.283348Z digest=sha256:73290e9fb66de04101fb55e4197a42f593321c201480699b9723506e61ff3cea

Observation e031d502-458c-4d91-bdb1-dc320af3c388 · inbound

Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness cites this paper.

Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T18:40:05.431137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:40:05.431137Z digest=sha256:5a4526898b343a8d9fd990862f6651991470641f78fde437924600a1d3d7a05f

Observation d1d97fc7-ad59-4126-a863-8e9ed0db2a6e · inbound

Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software cites this paper.

Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:41:00.347065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:39:42.391102Z digest=sha256:1e57e0afe60bf3ee2776f3d28dc2cbb2169b5539acc2bd55cee2336347821ca5

Observation dac42bef-5b25-4d36-b2f1-8906c18b8a89 · inbound

A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code cites this paper.

A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T06:49:11.634098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:49:11.634098Z digest=sha256:75b4bb6c87f70e61b4d1c9abd50c18789ae547fc544eae306b40ef0115c5d0dd

Observation b722956f-ca0a-477b-8cc0-a3193382c056 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:10.766568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:f927bd6a11bf29a0234875aa6cea790e6d75cc1b36a906e7aadc0258f5e52e46

Observation 8743e843-2620-47c1-afb5-b5b82e95057e · inbound

"Like Taking the Path of Least Resistance": Exploring the Impact of LLM Interaction on the Creative Process of Programming cites this paper.

"Like Taking the Path of Least Resistance": Exploring the Impact of LLM Interaction on the Creative Process of Programming Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:42:30.623598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T17:40:19.994225Z digest=sha256:2c4055e4a54bacb70cf1c5de5b7802bb44938917eec69e45f6b9c99e8c2d101a

Observation d49ddea4-c397-4d87-bc76-a4c43edbef35 · inbound

Subjective Code Preferences in Experts and Large Language Models cites this paper.

Subjective Code Preferences in Experts and Large Language Models Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:01.812357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:18:27.778928Z digest=sha256:25b394313eaf0dd7853f53e3dd2185a14525be2ac2694c77137d02f37070ee97

Observation 8a716987-56df-4609-b807-15fdae804c12 · inbound

LLM vs. Human Unit Tests: Fault Detection on Real Python Bugs cites this paper.

LLM vs. Human Unit Tests: Fault Detection on Real Python Bugs Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:37:27.555356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T18:07:02.131365Z digest=sha256:8a54aa167c8d76e33a11b611703859f6b2bc7e1b7201ee9baaea427aac5b82f0

Observation 4f3dcc34-e852-46e0-b251-62dd46f8d9c2 · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 206

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:57:06.898375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:0c6f1d9ab1e29cfeaba9f17fac7060f467f482627a17e6327311ba4f381d7b78

Observation de004fe9-a825-4661-90da-7af60fb827bf · inbound

CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation cites this paper.

CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:07.898036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:07.898036Z digest=sha256:d1d5609df3062c8b15cd5846d0998ab79bee7fd10439ab94ff1385a5fa42b75b