Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models

| Source: arXiv AI

Tags: security, JavaScript, vulnerability-detection, LLMs, Gemini, SAST, fine-tuning

Fine-tuned LLMs dramatically outperform SAST tools at detecting JavaScript vulnerabilities: fine-tuned Gemini 1.5 Flash reaches 60% accuracy (up from 29%) versus near-zero for rule-based analyzers, with SQL injection detection hitting 84% in an empirical study of 1,125 snippets.

Details

JavaScript powers 98.8% of websites, making its vulnerabilities a disproportionate security risk. This empirical study benchmarks three LLM families—Gemini 1.5 Flash, GPT-4o Mini, and DeepSeek-R1-Distill-Llama-8B—across zero-shot, chain-of-thought, and few-shot prompting, plus fine-tuning, on 1,125 JavaScript snippets covering five CWE categories: Injection (CWE-74), OS Command Injection (CWE-78), XSS (CWE-79), SQL Injection (CWE-89), and Uncontrolled Resource Consumption (CWE-400). Key findings: fine-tuning Gemini 1.5 Flash improves accuracy from 29% to 60%; Chain-of-Thought prompting helps reasoning-capable models like GPT-4o Mini more than smaller models; few-shot prompting works best for polymorphic vulnerabilities like XSS; SQL injection detection reaches 84% accuracy. Traditional SAST tools scored near zero on the same snippet set—a stark contrast. The authors are appropriately honest: recall is limited, performance varies across CWE types, and LLMs should complement rather than replace existing security workflows. The dataset (1,125 snippets) is modest, and generalization to full codebases requires further study.