Benchmarking 13 AI Models on Known CVE Detection

This post evaluates 13 AI models on their ability to rediscover 26 known CVEs from the GitHub Advisory Database. Key findings show that GPT-5.6 achieved the highest recall at 88.5% (23/26), but the most expensive models did not always justify their cost. Pooling multiple runs of a cheaper model (pass@3) often outperformed a single pass of a flagship model, while open-weight models like GLM-5.2 and the newly released Kimi K3 showed strong performance, with Kimi K3 matching frontier models at a fraction of the price. 

https://www.aikido.dev/blog/benchmarking-ai-models-known-cves

Comments

Popular posts from this blog

Prompt Engineering Demands Rigorous Evaluation

Open-SPDD proposes an open framework for Spec-Driven Development workflows

OWASP ASVS 5.0 Released - Key Updates and What You Need to Know