REVIEW 4 cited by
Exploring the Limits of ChatGPT in Software Security Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have undergone rapid evolution and achieved remarkable results in recent times. OpenAI's ChatGPT, backed by GPT-3.5 or GPT-4, has gained instant popularity due to its strong capability across a wide range of tasks, including natural language tasks, coding, mathematics, and engaging conversations. However, the impacts and limits of such LLMs in system security domain are less explored. In this paper, we delve into the limits of LLMs (i.e., ChatGPT) in seven software security applications including vulnerability detection/repair, debugging, debloating, decompilation, patching, root cause analysis, symbolic execution, and fuzzing. Our exploration reveals that ChatGPT not only excels at generating code, which is the conventional application of language models, but also demonstrates strong capability in understanding user-provided commands in natural languages, reasoning about control and data flows within programs, generating complex data structures, and even decompiling assembly code. Notably, GPT-4 showcases significant improvements over GPT-3.5 in most security tasks. Also, certain limitations of ChatGPT in security-related tasks are identified, such as its constrained ability to process long code contexts.
Forward citations
Cited by 4 Pith papers
-
Geometric quantification for nonlinear deformation in knitted fabrics
A geometric quantification framework reconstructs yarn centerlines and fabric surfaces from sparse knit data and partitions large deformation into stitch reorientation, loop bending, surface bending, and dilation.
-
ACSE-Eval: Can LLMs threat model real-world cloud infrastructure?
ACSE-Eval benchmarks LLMs on threat-modeling 100 AWS architectures and finds that GPT-4.1 and Gemini 2.5 Pro lead threat identification, while all models score below 50% on exact CWE and ATT&CK classification.
-
SoK: Towards Effective Automated Vulnerability Repair
This SoK benchmarks ten automated vulnerability repair tools across C/C++ and Java and concludes that learning-based methods are strong on synthetic benchmarks but trail non-learning methods on real-world vulnerabilit...
-
RTL-Breaker: Assessing the Security of LLMs against Backdoor Attacks on HDL Code Generation
RTL-Breaker shows that fine-tuning an LLM on a small fraction of poisoned Verilog samples can make it generate malicious RTL code when a specific trigger appears, while still passing VerilogEval functional checks.
Discussion (0). Continue with ORCID to comment.