1 Department of Electrical Engineering, Faculty of Engineering, Bayero University, Kano, Kano State, Nigeria.
2 Department of Computer Science, Faculty of Computing, Sokoto State University, Sokoto State, Nigeria.
World Journal of Advanced Engineering Technology and Sciences, 2026, 19(03), 009-017
Article DOI: 10.30574/wjaets.2026.19.3.0292
Received on 20 April 2026; revised on 29 May 2026; accepted on 01 June 2026
Large Language Models (LLMs) are increasingly used to generate application code, but their security-by-default behavior remains uncertain. This study evaluates the security posture of four contemporary AI coding systems: Gemini 3.5 Flash accessed through Antigravity, and DeepSeek V4 Flash, Kimi K2.5, and GPT-5.4-mini accessed through OpenCode via their APIs. Ten standardized prompts covering common web application tasks, including authentication, CRUD operations, SQL search, file upload, access control, secret handling, input validation, frontend login, and session management, were submitted once to each system. The resulting 40 code samples were analyzed using Semgrep OSS with manual review for selected logic and architectural issues. Findings were normalized by model, prompt, severity, and OWASP Top 10 category. The benchmark identified 64 normalized security findings, including 62 Critical/High findings and 2 Medium findings. GPT-5.4-mini produced the lowest number of findings (10), while Gemini 3.5 Flash produced the highest (21). The most frequent weaknesses were missing Cross-Site Request Forgery protections, insecure cookie/session settings, path traversal risks in file handling, and security misconfigurations in generated infrastructure. These results suggest that one-shot AI-generated web application code can be functionally useful but should not be considered production-ready without security review, automated scanning, and manual validation.
Artificial intelligence; Software security; Large language models; OWASP Top 10; Code generation; Static application security testing
Get Your e Certificate of Publication using below link
Preview Article PDF
Mohammad-Jamiu Babatunde Balogun, Ahmed Sani Geza and Yusuf Umar Jimada. Security Benchmarking of AI-Generated Web Application Code: An OWASP-Based Comparative Study of Coding LLMs. World Journal of Advanced Engineering Technology and Sciences, 2026, 19(03), 009-017. Article DOI: https://doi.org/10.30574/wjaets.2026.19.3.0292