Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2312.12575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12575 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:02:43.645372Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:43.061376Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8ad7b8cf-605d-492d-86b9-08d4f961b81a · inbound

Integrating Artificial Open Generative Artificial Intelligence into Software Supply Chain Security cites this paper.

Integrating Artificial Open Generative Artificial Intelligence into Software Supply Chain Security LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:43.645372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:43.645372Z digest=sha256:a740fcf42f9c729c530d3cd1efe7f062181aa09b296fd3803ff904fd569058e8

Observation 15390e56-20ed-4974-80f4-af460d0c14df · inbound

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants cites this paper.

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:57:17.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:55:14.887237Z digest=sha256:350bcea15e791f59ee258070e20d8442f8b2222b9d6045132e9284a24cc441d9

Observation a4f1c9de-b850-44b7-9fb2-c4ceaaf81f90 · inbound

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond cites this paper.

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:16.105652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:16.105652Z digest=sha256:88e4faf7c23879ca1489f6800e1ddd8136f7f32cd79a4f69b218d5ffc9ce118c

Observation f450ee49-39b8-4ee6-82a8-6824a1947f21 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.983497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.983497Z digest=sha256:93914230fd791fb1a43b1d19b50dbcf6eb5d1142f34001f53321115e1a4d766c

Observation 3ba98a77-e4f3-4c0b-8aeb-54fdb87972f6 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:28.904690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:28.904690Z digest=sha256:632d28ba0d79e89bc8faceed263a68c7ab516543a803f0448f30a36629c02fe2

Observation 466ba651-e18b-424e-bade-bc8ad19b0068 · inbound

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis cites this paper.

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:20.536909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:20.536909Z digest=sha256:6de74ff9ea6f333fbe83cebb5a4df536ddc3a5afa398edef942a366c9419fc03

Observation 58e5ca18-62e4-43de-ba8b-c3b41903be0e · inbound

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software cites this paper.

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.926724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:07:44.158768Z digest=sha256:c713179c6b0401b1ad03d1d195404e917602e80c87dcae9b9d34534f99d792ff

Observation 99891289-6814-44ea-83bb-a4ada86260e3 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.750390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:3fed563d5945082ca07591acb9dae49eaad131b9a9ea9bee248dfa8de8dddda4

Observation e280e68f-1216-4b7d-9d5b-f34c9e5e8d36 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:43.063064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:7ddf1f5040fb040b0eeff132937864f7fe225229675a637f976e5e1e201cb75a