Benchmarking 13 AI Models on Known CVE Detection
This post evaluates 13 AI models on their ability to rediscover 26 known CVEs from the GitHub Advisory Database. Key findings show that GPT-5.6 achieved the highest recall at 88.5% (23/26), but the most expensive models did not always justify their cost. Pooling multiple runs of a cheaper model (pass@3) often outperformed a single pass of a flagship model, while open-weight models like GLM-5.2 and the newly released Kimi K3 showed strong performance, with Kimi K3 matching frontier models at a fraction of the price. https://www.aikido.dev/blog/benchmarking-ai-models-known-cves