Pith. sign in

REVIEW 2 cited by

Software Testing with Large Language Models: Survey, Landscape, and Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.07221 v3 pith:XFKP2W57 submitted 2023-07-14 cs.SE

classification cs.SE
keywords softwarellmstestinglanguageusedanalyzesareacommonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained large language models (LLMs) have recently emerged as a breakthrough technology in natural language processing and artificial intelligence, with the ability to handle large-scale datasets and exhibit remarkable performance across a wide range of tasks. Meanwhile, software testing is a crucial undertaking that serves as a cornerstone for ensuring the quality and reliability of software products. As the scope and complexity of software systems continue to grow, the need for more effective software testing techniques becomes increasingly urgent, making it an area ripe for innovative approaches such as the use of LLMs. This paper provides a comprehensive review of the utilization of LLMs in software testing. It analyzes 102 relevant studies that have used LLMs for software testing, from both the software testing and LLMs perspectives. The paper presents a detailed discussion of the software testing tasks for which LLMs are commonly used, among which test case preparation and program repair are the most representative. It also analyzes the commonly used LLMs, the types of prompt engineering that are employed, as well as the accompanied techniques with these LLMs. It also summarizes the key challenges and potential opportunities in this direction. This work can serve as a roadmap for future research in this area, highlighting potential avenues for exploration, and identifying gaps in our current understanding of the use of LLMs in software testing.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Comparative Evaluation of Large Language Models for Test-Skeleton Generation

    cs.SE 2025-09 reject novelty 4.0 of 10

    In a single-prompt study on one Ruby class, DeepSeek-Chat produced the most maintainable RSpec test skeletons, while GPT-4's high method coverage was undercut by incorrect RSpec conventions.

  2. The role of large language models in UI/UX design: A systematic literature review

    cs.HC 2025-07 conditional novelty 4.0 of 10

    A systematic review of 38 studies finds LLMs are increasingly integrated across the UI/UX design lifecycle, with prompt engineering and human oversight as key practices.

Pith tools