Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2312.12575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12575 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:01:16.105652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:43.061376Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 15390e56-20ed-4974-80f4-af460d0c14df · inbound

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants cites this paper.

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:57:17.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:55:14.887237Z digest=sha256:fda8c4a6350d8eea5230fb4f0966b22886fb733d007d51c10033510256752788

Observation a4f1c9de-b850-44b7-9fb2-c4ceaaf81f90 · inbound

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond cites this paper.

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:16.105652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:16.105652Z digest=sha256:492a403194b65af8370decd76a28f556a88901c4e43112fb43e74c6867f5c45c

Observation f450ee49-39b8-4ee6-82a8-6824a1947f21 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.983497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.983497Z digest=sha256:8ca610a7e462fab1ac7f1b9eae19a130100c6bef5f3dfeca082d89a7f9c9d046

Observation 3ba98a77-e4f3-4c0b-8aeb-54fdb87972f6 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:28.904690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:28.904690Z digest=sha256:f159fae188e5c04c8fec1a157c664bd421f2095e2d3eff9367a29c9b76843f3a

Observation 466ba651-e18b-424e-bade-bc8ad19b0068 · inbound

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis cites this paper.

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:20.536909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:20.536909Z digest=sha256:0b4ab69f7d7c172a6ff12e46044a7bb22945d8814524ea24c81ab6ad4b6ce653

Observation 58e5ca18-62e4-43de-ba8b-c3b41903be0e · inbound

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software cites this paper.

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.926724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:07:44.158768Z digest=sha256:731e0f3d1a09c47eb443696806e461ec66a2d976fbb08a6b89d0240ba5fbcae1

Observation 99891289-6814-44ea-83bb-a4ada86260e3 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.750390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:58350470e74473c5dc017b8bbd3cd7863861230bc17321ee54136dc353837a70

Observation e280e68f-1216-4b7d-9d5b-f34c9e5e8d36 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:43.063064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:c8834f16e98e9394280ab3c09266a890aec1608bf1a2973c9bf794a59f43af9c