Evaluating AI reasoning and prompt engineering in automated test case generation: A comparative study of GPT-4o, O1 models, and human QA

  • Zubair, Muhammad Azlaan
  • Bouchelligua, Wided
  • Danish, Sufyan
  • Ahmad, Shabir
  • Ksibi, Amel
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The increased use of AI in SQA requires a strict analysis of AI models in TC generation. This paper examines the factuality of the advanced reasoning of Open AI (O1 series) being the distinguishing factor between Open AI and non reasoning models GPT 4o and human Quality Assurance (QA) engineers. To address this, we evaluated the performance of 12 experimental groups: human testers, eight variants of GPT 4o which used a variety of prompt engineering techniques (e.g., chain of thought, few shot, self consistency), and three variants of the O1 model in the generation of TCs on 11 software applications. The 3682 TCs were tested with the mixtral-8x7B-32768 model according to four criteria: Coverage, Clarity, ENC and NFC. The descriptive statistics indicated that there was consistently high coverage and high clarity among the groups but there was a significant difference in the capture of edge, negative and non-functional scenarios. Statistical analyses using both parametric and non parametric methods identified significant differences across several TC quality metrics depending on prompt ing strategy and underlying model architecture. Interestingly, AI models not only produced more TCs but also had a better QTQ ratio than human engineers under the conditions of receiving little contextual information. Finally, reasoning models have more structured outputs but may not yield significant performance improvement given the increased computational costs. Enhanced prompt engineering enables non-reasoning models to offer an efficient and effective alternative for TC generation, thereby underscoring the potential for AI integration in optimizing software testing processes. To facilitate replication and further research, an extensive dataset and fully documented codebase are available at the GitHub repository provided in the Contributions section.

키워드

Generative AIAutomated TC generationSoftware testingPrompt engineeringQuality assuranceLarge language models
제목
Evaluating AI reasoning and prompt engineering in automated test case generation: A comparative study of GPT-4o, O1 models, and human QA
저자
Zubair, Muhammad AzlaanBouchelligua, WidedDanish, SufyanAhmad, ShabirKsibi, Amel
DOI
10.1016/j.asoc.2026.115708
발행일
2026-09
유형
Article
저널명
Applied Soft Computing Journal
201