Pith. sign in

REVIEW 13 cited by

LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16185 v4 pith:WGBIFC4U submitted 2024-01-29 cs.CR cs.AIcs.SE

classification cs.CRcs.AIcs.SE
keywords vulnerabilityllmsknowledgereasoningevaluationllm4vulncapabilitiescontext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large language models (LLMs) have demonstrated significant potential in various tasks, including those requiring human-level intelligence, such as vulnerability detection. However, recent efforts to use LLMs for vulnerability detection remain preliminary, as they lack a deep understanding of whether a subject LLM's vulnerability reasoning capability stems from the model itself or from external aids such as knowledge retrieval and tooling support. In this paper, we aim to decouple LLMs' vulnerability reasoning from other capabilities, such as vulnerability knowledge adoption, context information retrieval, and advanced prompt schemes. We introduce LLM4Vuln, a unified evaluation framework that separates and assesses LLMs' vulnerability reasoning capabilities and examines improvements when combined with other enhancements. To support this evaluation, we construct UniVul, the first benchmark that provides retrievable knowledge and context-supplementable code across three representative programming languages: Solidity, Java, and C/C++. Using LLM4Vuln and UniVul, we test six representative LLMs (GPT-4.1, Phi-3, Llama-3, o4-mini, DeepSeek-R1, and QwQ-32B) for 147 ground-truth vulnerabilities and 147 non-vulnerable cases in 3,528 controlled scenarios. Our findings reveal the varying impacts of knowledge enhancement, context supplementation, and prompt schemes. We also identify 14 zero-day vulnerabilities in four pilot bug bounty programs, resulting in $3,576 in bounties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. Geometric quantification for nonlinear deformation in knitted fabrics

    cond-mat.soft 2026-04 unverdicted novelty 7.0 of 10

    A geometric quantification framework reconstructs yarn centerlines and fabric surfaces from sparse knit data and partitions large deformation into stitch reorientation, loop bending, surface bending, and dilation.

  2. Directed Symbolic Execution for Vulnerability Discovery: An LLM-Guided Approach in KLEE

    cs.SE 2026-07 conditional novelty 6.0 of 10

    An LLM-guided KLEE searcher that steers paths toward marked vulnerable code and exits unmarked loops found 87 unique sanitizer-confirmed violations across 11 programs, ahead of 13 baselines.

  3. DREA: Decoupled Reasoning and Exploration Agents for Repository-Level Vulnerability Detection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    DREA improves repository-level vulnerability detection by coupling an LLM planner that forms hypotheses with a cheap local explorer that gathers cross-file evidence, lifting paired accuracy from 19-26% to 30-42% at mu...

  4. Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A tri-grounded multi-agent harness improves precision and auditability of Android malware behavior reports over fixed LLM pipelines and frontier coding agents.

  5. ShadowProbe: Language-Extensible Detection of Hidden Algorithmic Complexity Vulnerabilities

    cs.CR 2026-07 conditional novelty 6.0 of 10

    ShadowProbe detects language-extensible algorithmic complexity vulnerabilities from hidden library costs via static screening, context recovery, LLM input synthesis, and measured runtime growth.

  6. A Mixture of Linear Corrections Generates Secure Code

    cs.CR 2025-07 conditional novelty 6.0 of 10

    An inference-time mixture of linear correction vectors, derived from linear probes on LLM hidden states, improves the security and functionality of code generated by Qwen2.5-Coder and CodeLlama models.

  7. FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction

    cs.CR 2025-06 conditional novelty 6.0 of 10

    An LLM-driven pipeline extracts and CWE-classifies 27,497 smart contract vulnerabilities from 6,454 audit reports, producing a benchmark that exposes the poor performance of existing detection tools.

  8. Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges

    cs.AI 2025-06 reject novelty 6.0 of 10

    A benchmark and agent for CTF solving, but the agent's retrieval database appears to contain the answers to the test challenges, undermining the reported improvements.

  9. Larger Is Not Always Better: Exploring Small Open-source Language Models in Logging Statement Generation

    cs.SE 2025-05 conditional novelty 6.0 of 10

    A fine-tuned 14B small open-source model with LoRA and RAG outperforms larger proprietary LLMs on automated Java logging statement generation in AL-Bench point estimates.

  10. Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A LLaMA-3.1-8B model trained with continual pre-training, supervised fine-tuning, and direct preference optimization reports state-of-the-art accuracy and F1 for smart contract vulnerability detection and explanation.

  11. Adaptive Plan-Execute Framework for Smart Contract Security Auditing

    cs.CR 2025-05 reject novelty 5.0 of 10

    SmartAuditFlow claims 100% detection on a standard smart contract benchmark and all 13 tested CVEs via a plan-execute LLM workflow, though the supporting evaluation has major reproducibility and validation gaps.

  12. LLM-BSCVM: An LLM-Based Blockchain Smart Contract Vulnerability Management Framework

    cs.CR 2025-05 conditional novelty 4.0 of 10

    An LLM-based six-agent framework reports 91% detection accuracy on smart contract benchmarks, with a 5.1% false positive rate and a partially successful automated repair stage.

  13. Generative AI for Internet of Things Security: Challenges and Opportunities

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey that catalogs 33 GenAI-for-IoT-security works through the MITRE ICS mitigations lens, with three small case studies on adapting LLMs to IoT incident response and security question answering.

Pith tools