Table of Contents (11 Sections)
Abstract
Static Application Security Testing (SAST) has long suffered from high false-positive rates and developer alert fatigue. This paper evaluates Cyfendo, an agentic LLM vulnerability scanner that combines semantic context reasoning with automated Proof-of-Concept (PoC) exploit verification. Evaluated against all 2,740 test cases of the OWASP Benchmark (v1.2), Cyfendo achieves a Youden’s Index of 0.90 (94.06% Sensitivity, 3.77% False Positive Rate, and 96.38% Precision)—demonstrating how autonomous AppSec can deliver high-precision security analysis and review-ready patches.
Executive Summary #
Modern software engineering requires security tooling that operates at the speed of continuous deployment without sacrificing accuracy. Traditional commercial SAST tools struggle on standardized benchmarks, often yielding high false-positive rates that create severe alert fatigue and necessitate time-consuming manual triage.
The Cyfendo Agentic Scanning Framework represents a departure from static syntax and pattern matching. By orchestrating specialized LLM agents through a pipeline of taint-guided reasoning and adversarial review, Cyfendo understands complex code intent and custom sanitization logic that baffles traditional analyzers.
Evaluated across the entire OWASP Benchmark v1.2 suite (2,740 test cases), Cyfendo achieved 94.06% Sensitivity (catching 1,331 of 1,415 true vulnerabilities) and a 96.23% Specificity rate (only 50 false positives across 1,325 safe controls), resulting in a Youden’s Index of 0.90, a 96.38% Precision rate, and an F1-score of 95.21%. This high diagnostic fidelity enables engineering teams to deploy automated security gates in CI/CD pipelines where findings are trusted to block high-risk regressions without human bottlenecking.
Architecture & Core Technologies #
Cyfendo operates through a five-stage agentic pipeline designed to deliver comprehensive vulnerability recall while aggressively suppressing false alarms:
Semantic Context Parsing
Ingests the target codebase to construct a semantic representation of data sources, transformation layers, and sinks, moving beyond rigid abstract syntax trees to understand overarching program intent.
Taint-Guided LLM Reasoning
Performs virtual data-flow analysis by tracking how untrusted input propagates across functions and custom sanitization routines, uncovering non-obvious logic vulnerabilities.
Adversarial LLM Review
Subjects candidate findings to an adversarial reviewer agent that challenges reachability, identifies framework-level protections, and filters out non-exploitable theoretical findings.
Deterministic Structural Validation
Cross-references model reasoning with syntax tree parsing and symbol verification to ensure cited source paths, symbols, and parameters strictly exist in the codebase, enforcing structural validity.
Dynamic PoC & Reachability Validation
Synthesizes unit-level Proof-of-Concept exploit test cases and executes them in an isolated sandbox for standalone testable sinks. Confirmed exploitable vulnerabilities are prioritized with highest diagnostic confidence, while complex architectural findings maintain full taint reachability graphs.
OWASP Benchmark Overview & Methodology #
The OWASP Benchmark is an open-source evaluation suite designed to assess the accuracy, speed, and coverage of automated software vulnerability scanners.
Test Suite Composition
Version 1.2 consists of 2,740 total test cases across 11 vulnerability categories:
- True Vulnerabilities (1,415 cases): Test cases containing actual, exploitable vulnerabilities across data-flow sinks.
- Safe Controls & Decoys (1,325 cases): Code paths containing effective validation, encoding, or framework defenses where vulnerability alerts constitute false positives.
Evaluation Formulae
Performance is quantified using standard diagnostic classification metrics:
-
True Positive Rate (Sensitivity / TPR):
TPR = TP / (TP + FN)— measures vulnerability discovery coverage. -
True Negative Rate (Specificity / TNR):
TNR = TN / (TN + FP)— measures precision on safe code. -
False Positive Rate (FPR):
FPR = FP / (TN + FP) = 1 - Specificity. -
Youden’s Index (J):
J = Sensitivity + Specificity - 1— summarizes overall diagnostic power, where 1.0 represents perfect classification and 0.0 represents random guessing.
Benchmark Analysis & Diagnostic Performance #
The OWASP Benchmark v1.2 provides an objective, standardized testbed for evaluating automated security scanning accuracy. Static application security testing has historically faced an acute trade-off between sensitivity (catching true vulnerabilities) and specificity (resisting false alarms on safe negative controls).
Cyfendo's evaluation was conducted across the entire 2,740 test-case suite under uniform execution conditions. By coupling multi-stage semantic parsing, taint reachability analysis, adversarial review, and dynamic sandbox verification where applicable, Cyfendo achieves balanced high performance across both dimensions:
| Evaluation Metric | Measured Value | Diagnostic Significance |
|---|---|---|
| Youden’s Index (J) | 0.90 | Composite diagnostic power (Sensitivity + Specificity - 1) on a scale from -1.0 to +1.0. |
| Sensitivity (Recall / TPR) | 94.06% | Successfully identified 1,331 out of 1,415 true vulnerabilities (5.94% miss rate). |
| Finding Precision (PPV) | 96.38% | 1,331 true positives out of 1,381 total flagged test cases. |
| Specificity (TNR) | 96.23% | Correctly recognized 1,275 of 1,325 safe control test cases without false alarms. |
| False Positive Rate (FPR) | 3.77% | 50 false positives across 1,325 safe negative controls. |
| F1-Score | 95.21% | Harmonic mean of sensitivity and precision across the full suite. |
Performance Summary (OWASP) #
Evaluation results for Cyfendo across all 2,740 test cases (1,415 vulnerable cases and 1,325 safe controls):
| Metric | Score / Count | Definition & Context |
|---|---|---|
| Total Test Cases | 2,740 / 2,740 | 100.0% coverage (1,415 true vulns, 1,325 safe controls) |
| True Positives (TP) | 1,331 | Correctly identified vulnerabilities |
| False Positives (FP) | 50 | Safe test cases incorrectly flagged |
| True Negatives (TN) | 1,275 | Correctly recognized safe controls |
| False Negatives (FN) | 84 | Missed vulnerabilities (5.94% miss rate) |
| Sensitivity (Recall / TPR) | 94.06% | TP / (TP + FN) |
| Specificity (TNR) | 96.23% | TN / (TN + FP) |
| False Positive Rate (FPR) | 3.77% | FP / (TN + FP) |
| Precision (PPV) | 96.38% | TP / (TP + FP) |
| F1-Score | 95.21% | Harmonic mean of Precision and Sensitivity |
| Youden’s Index (J) | 0.90 | Overall diagnostic capability (Sensitivity + Specificity - 1) |
Category Breakdown #
Performance breakdown across all 11 vulnerability categories evaluated in OWASP Benchmark v1.2:
| Category | Total | Vulns | Safe | TP | FP | TN | FN | Sensitivity | Specificity | Youden J |
|---|---|---|---|---|---|---|---|---|---|---|
| Weak Randomness | 493 | 218 | 275 | 218 | 0 | 275 | 0 | 100.0% | 100.0% | +1.000 |
| XPath Injection | 35 | 15 | 20 | 15 | 0 | 20 | 0 | 100.0% | 100.0% | +1.000 |
| Secure Cookie Flag | 67 | 36 | 31 | 36 | 0 | 31 | 0 | 100.0% | 100.0% | +1.000 |
| LDAP Injection | 59 | 27 | 32 | 27 | 1 | 31 | 0 | 100.0% | 96.9% | +0.969 |
| SQL Injection | 504 | 272 | 232 | 265 | 5 | 227 | 7 | 97.4% | 97.8% | +0.953 |
| Path Traversal | 268 | 133 | 135 | 123 | 0 | 135 | 10 | 92.5% | 100.0% | +0.925 |
| Cross-Site Scripting (XSS) | 455 | 246 | 209 | 225 | 4 | 205 | 21 | 91.5% | 98.1% | +0.895 |
| Command Injection | 251 | 126 | 125 | 106 | 4 | 121 | 20 | 84.1% | 96.8% | +0.809 |
| Trust Boundary | 126 | 83 | 43 | 69 | 1 | 42 | 14 | 83.1% | 97.7% | +0.808 |
| Weak Cryptography | 246 | 130 | 116 | 118 | 29 | 87 | 12 | 90.8% | 75.0% | +0.658 |
| Weak Hash | 236 | 129 | 107 | 129 | 6 | 101 | 0 | 100.0% | 94.4% | +0.944 |
| TOTAL | 2,740 | 1,415 | 1,325 | 1,331 | 50 | 1,275 | 84 | 94.1% | 96.2% | +0.90 |
Key Observations & Strengths #
- 100% Recall Across 5 Vulnerability Categories: Cyfendo achieved complete 100% vulnerability recall on Weak Randomness, XPath Injection, Secure Cookie Flags, Weak Hash, and LDAP Injection. Notably, Weak Randomness, XPath Injection, and Secure Cookie Flags each achieved a perfect Youden Index of +1.000 with 100% sensitivity and zero false positives.
- High Detection Recall (94.06%): Across 1,415 true vulnerabilities, Cyfendo successfully discovered 1,331 true positives, maintaining strong detection coverage across high-risk classes including Weak Hash (100.0%), SQL Injection (97.4%), and Path Traversal (92.5%).
- Ultra-Low False Positive Rate (3.77%): On safe negative controls, Cyfendo demonstrated outstanding discriminative precision with 96.23% Specificity, correctly identifying 1,275 out of 1,325 benign test cases and producing only 50 false positives across the entire suite. Four categories (Weak Randomness, XPath Injection, Secure Cookie Flags, and Path Traversal) achieved a flawless 100% Specificity with zero false alarms.
- Elevated Discriminative Power (Youden Index = 0.90): Demonstrates state-of-the-art diagnostic power across all 2,740 test cases, achieving a 0.90 Youden Index (0.9029), 96.38% Precision, and a 95.21% F1-score under standardized benchmark conditions.
Discussion & Enterprise Impact #
Addressing Developer Alert Fatigue
The combination of 94.06% sensitivity, 96.38% precision, and an ultra-low 3.77% false positive rate directly addresses the core operational failure of legacy SAST: alert fatigue. With only 50 false alarms across 1,325 safe controls, and security alerts accompanied by sandboxed exploit verification, security teams transition from triaging noise to verifying automated remediation.
Interpreting Custom Sanitization Logic
Traditional tools rely on hardcoded registries of known sanitization routines. Modern enterprise software regularly implements domain-specific encoding and validation. Agentic reasoning enables Cyfendo to interpret custom defenses in context, preventing safe logic from triggering false alarms.
Deployment Economics and Scaling
While multi-stage agentic reasoning incurs higher per-scan compute costs than simple static linters, the downstream elimination of manual triage delivers a compelling net operational ROI in enterprise environments.
Benchmark Scope & Generalization Beyond Synthetic Java
The OWASP Benchmark v1.2 represents a standardized, reproducible baseline for evaluating taint propagation across 2,740 test cases. Because the benchmark suite is publicly accessible, standard foundation models are vulnerable to test-set memorization. Cyfendo guards against memorization artifacts by coupling LLM semantic reasoning with deterministic AST structural verification and dynamic sandbox execution. While the OWASP Benchmark suite evaluates Java servlets, Cyfendo's underlying agentic pipeline operates on language-agnostic semantic control-flow graphs, extending diagnostic capabilities to modern polyglot ecosystems (TypeScript, Python, Go) and cloud infrastructure architectures, with ongoing evaluations across real-world CVE corpora.
Conclusion #
These benchmark results demonstrate strong diagnostic performance for Cyfendo's agentic approach on OWASP Benchmark v1.2. Achieving a Youden’s Index of 0.90, 94.06% Sensitivity, 96.23% Specificity, and 96.38% Precision across 2,740 test cases, Cyfendo demonstrates that autonomous, multi-layer AppSec can be deployed with high precision in modern CI/CD pipelines while developers retain full approval authority over merged code.
Citation
BibTeX@techreport{cyfendo2026owasp,
title = {Cyfendo Agentic Scanning on the OWASP Benchmark},
author = {{Cyfendo Autonomous Security Research Team}},
institution = {Cyfendo Inc.},
year = {2026},
month = {September},
version = {Alpha},
url = {https://cyfendo.com/whitepaper},
note = {Evaluated on OWASP Benchmark v1.2 (2,740 test cases, Youden Index: 0.90)}
}
See What Cyfendo Finds in Your Code
Start with up to 10K protected LOC and 5 scans per month. No credit card required.