Pith. sign in

Paper Citation Record · LEDGER

Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.02395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02395 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:14.361719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T19:56:10.801171Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec566a92-e435-4bc6-8eb7-23f91e20f791 · inbound

LLM Performance for Code Generation on Noisy Tasks cites this paper.

LLM Performance for Code Generation on Noisy Tasks Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.375530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:43.375530Z digest=sha256:cf6f99608479f8b432830803e0b0ca80c8e235b94c4801cced0f3c07a877b327

Observation 9d273f42-c329-43ba-bd73-61fc2f03955b · inbound

CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs cites this paper.

CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:14.361719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:14.361719Z digest=sha256:e29be3f47cc38c0ff3fba035b6b07bec88f2bcb16c68510aac25e8fdac15f407

Observation 2c725ddf-2f77-46cf-8740-b74455af6572 · inbound

Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation cites this paper.

Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:56:17.536958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:56:17.536958Z digest=sha256:d4642f029678b0cd2a7d5cee77b304f279d8a6ed6944dc9fbbe9052b3542be68

Observation 0cfa7f88-6cea-4a69-8188-34069b47c5d3 · inbound

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions cites this paper.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:52.828980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:52.828980Z digest=sha256:5d8c4810db0d74d8f81bd2c02b50a424c728355b0104c8ad3bb6e3c72db056f4

Observation 31ac015f-04da-469f-8178-46ff07505fa0 · inbound

Secure Code Generation at Scale with Reflexion cites this paper.

Secure Code Generation at Scale with Reflexion Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:50:49.625959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:50:49.625959Z digest=sha256:85ddbb3bc1db52ac2d53043b86a3b2af4b1b8b7a41bdd07fb0c71796ba19ba75

Observation f292a6ca-192e-4192-b4e7-dc142ce8f940 · inbound

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting cites this paper.

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:08:13.072442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T18:04:44.543311Z digest=sha256:7b1e3bacbfb84d9ab7acea765cf0bd14b0d515e18375e280cec2a79f1bc2594e

Observation 3a468b82-4d6b-4f18-9857-85c3e7ea6785 · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:44.638936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:44.638936Z digest=sha256:06c30b0f30b7867e5574553a14122d8eedf2bddab70c6d814a252dfd90f3f461

Observation 453e945b-dba5-472b-91ce-450fa6c4044c · inbound

A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories cites this paper.

A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:49:33.951503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:48:29.311237Z digest=sha256:6e580e00da935579e1790caaaccc9935669cd429ea6579d4586f11bdde61f4cc

Observation 2bfa60fa-e387-4c62-80ef-dea2b0e8cba1 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.021556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:5b1d1f790c01638120e7c3ac59d1405b968d781a6d7c790d4008c9fd0baf5e62

Observation 026c2599-6751-47d2-8c72-78782ab782ae · inbound

Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025) cites this paper.

Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025) Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.512231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:38:43.363709Z digest=sha256:b8f39d66bde3dad7119a01c99b13c08cf1d92307c141d8f4ae8f5d044f37abf8

Observation 092fec05-4353-4187-b4a8-4812da081a03 · inbound

What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants cites this paper.

What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.802989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:57:02.677603Z digest=sha256:a015a8f2a6e64969d9603b82fd52cec4f8ffc3ea7ece254cc077023a52bf5c9a