
Cloudflare tested its WAF with frontier AI models: 558 of 1,107 attacks blocked
Cloudflare ran an AI-driven tester against its own WAF in an authorized staging environment. Across 45 scenarios and 1,107 attempts, 558 requests were blocked and 49 findings went to human review.
Cloudflare's security team ran an AI-driven tester against its own Web Application Firewall (WAF) in an authorized customer staging environment, to answer whether a WAF is ready for frontier AI models.
How the adaptive loop works
The tester starts from known exploits and iterates: it changes how a payload is encoded or delivered, sends it again, and uses the response to choose the next variation. Cloudflare says that is where LLMs are strongest.
Each scenario picks one attack category, places the input in a specific part of the request, and allows a fixed number of attempts. Two model calls run per step: one proposes a variation, one reviews the response. Neither sees rule IDs, rule expressions or WAF Attack Score details. The models never send requests directly; Python code checks the target against an allowlist and enforces the limit. An unblocked request was a lead for human review, not a confirmed exploit.
Six attack categories, 1,107 attempts
Cloudflare ran 45 scenarios, 44 of them covering cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion, and Log4j; the last one covered log injection. The zone used WAF Attack Score blocking at 30 or below, the full Cloudflare Managed Ruleset, and the OWASP Core Ruleset at Paranoia Level 3.
The run recorded 1,107 mutation attempts. After triage, the result set was 607: 558 blocked requests plus 49 documented findings. XSS, LFI, SQLi and Log4j had near full coverage, with 48 of the 49 findings from CMDi and SSRF. More attempts in one scenario did not always find more, as runs neared the 25-attempt limit.
From findings to new rules
One SSRF session sent the metadata address 169.254.169.254 in decimal, octal and trailing-dot forms. The WAF blocked all but one: at attempt 18 a trailing-dot host produced a redirect instead of a block page, with no evidence that metadata was read. The findings led to four groups of candidate rules and three Managed Ruleset changes: new SSRF - Obfuscated Host and SSRF - Restricted Protocol detections on July 21, and an improved SSRF - Cloud rule on August 4.
Cloudflare says two versions of the same model family produced different variations but the same underlying issues. It advises customers to configure Managed Rules and WAF Attack Score correctly, run new rules in log mode first and patch software: a bypassed WAF still needs an exploitable application to matter. WAF is one layer of many, and API Security and Threat Intelligence add more. A follow-up will cover a white-box test where the model also knows the application's vulnerabilities and the WAF rules.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.