Security teams testing web application firewalls have traditionally split their effort between static analysis of source code and dynamic probing of running applications. The second approach — dynamic application security testing — is where frontier models change the economics of the work: an LLM can take a payload that was blocked, alter its encoding or its location in the HTTP request, and try again, using each response to steer the next attempt. That iteration speed is what no manual tester can match.
To find out whether an LLM-driven adversary defeats its own WAF, the Cloudflare team built a tester that starts from known exploits and varies how each one is encoded or delivered. The tester's model has no visibility into application source, no access to WAF rule expressions or rule IDs, and sees only selected HTTP response data. A request that is not blocked is treated as a lead for human review, never as a confirmed exploit.
The tester ran against an authorized customer staging environment across six attack categories and produced 1,107 attempts. After removing malformed, benign, duplicate, and out-of-scope observations, the vast majority of attacks were blocked. The surviving requests fed new detections that now protect all Cloudflare customers, and the exercise is becoming a standing part of the company's WAF development lifecycle.
Anatomy of the adaptive loop
A scenario in this system is a combination: one attack category, a defined location within the request, a starting payload the WAF already blocks, and a fixed budget of attempts. The loop makes two LLM calls per step. The proposal call receives the starting request, context, and a short history of earlier results, then suggests the next variation; a review call gets the request context, response status, selected headers, and response body. Mutation stops when variations stop being useful or the hardcoded attempt limit is reached.
Neither call sees WAF internals — no rule expressions, no rule IDs, no WAF Attack Score breakdown, no identification of which security layer acted. The implementation is a purpose-built Python system rather than a wrapper around an existing penetration-testing tool; it handles HTTP replay, scenario orchestration, state tracking, and result collection on its own.

Models never send requests directly. Before each request, code checks the target hostname against an allowlist, disables redirects, records the attempt, and enforces the attempt limit. After each request, the response is recorded and the model's review selects the next predefined step. Because response text can resurface in a later prompt, it is handled as untrusted input. Neither call can deploy a rule or change enforcement. Structured evidence is captured for every attempt.
Test scope and target configuration
Of 45 scenarios run, 44 spanned six categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. A single remaining scenario covered log injection and is reported separately. Each scenario explored alternative deliveries of the same attack — different encodings, different request locations, rewritten destinations.
The staging zone used an allowlisted test User-Agent so the customer's automated-traffic controls would not stop requests before they reached the WAF. Its configuration was:
- WAF Attack Score blocking scores of 30 or below
- All Cloudflare Managed Ruleset rules enabled
- OWASP Core Ruleset at Paranoia Level 3
Scoring was binary per request — blocked or not — so the results describe the configured WAF boundary as a whole, not the behavior of any single rule or detection mechanism.
One trajectory: SSRF against a metadata endpoint
Cloud metadata services expose temporary workload credentials, and an SSRF flaw can make an application fetch that data on an attacker's behalf. In this scenario the tester sent the same metadata address as integer, octal, and trailing-dot forms of the same IP, and placed it in different parts of the request. All were blocked except one: at attempt 18, keeping the same request structure as the blocked attempt 17, the model switched to the trailing-dot form and the client hit a redirect instead of a WAF block.
The table below excerpts that session. The hypothesis column is a summary of what the model said it was trying, not a verbatim transcript, and not evidence that the stated reasoning was correct.
| Stage | Selected moment |
|---|---|
| BASELINE |
Observed: The WAF returned a 403 block page for the direct metadata address 169.254.169.254 Hypothesis: The detector might recognize the literal dotted address or metadata path. Next move: Represent the same address as the decimal integer 2852039166. |
| ATTEMPT 1 |
Observed: The decimal representation was blocked. Hypothesis: The metadata path or query placement might still trigger detection. Next move: Use the octal representation 0251.0376.0251.0376 and move the input from the query string to a form body. |
| ATTEMPT 2 |
Observed: The octal representation in the form body was blocked. Hypothesis: The metadata path might still trigger detection regardless of host encoding. Next move: Try a hostname-based representation without that path. |
| ATTEMPTS 3–16 | Omitted from this excerpt. The tester continued exploring host, path, and request-shape combinations. |
| ATTEMPT 17 |
Observed: The decimal integer host, now used with a later request shape, was blocked. Hypothesis: A trailing dot might change what the detector matched without changing the intended destination. Next move: Keep the method, input placement, content type, and path form the same; switch to the trailing-dot host form: 169.254.169.254. |
| ATTEMPT 18 |
Observed: The client encountered a redirect rather than a Cloudflare block page. The run retained an edge-pass observation because the expected mitigation was absent. Limit: There was no successful origin response, response body, or evidence that the application fetched metadata. Next step: Preserve the request for triage and origin-side validation. |
Attempts 17 and 18 — identical structure, different host representation, different outcomes — raised a concrete question about how the trailing dot affects the WAF's reading of the destination. That is a lead to investigate, not proof that metadata was reached. It was one selected path out of 45 scenarios.
Counting and triaging the full run
Across all 1,107 attempts, coverage was near-total for XSS, LFI, SQLi, and Log4j. The run also generated noise, and human review was the filter. It left 49 findings worth investigating, 48 of which fell under CMDi and SSRF:
Metric | Value | What it means |
|---|---|---|
Recorded mutation attempts | 1,107 | Model iterations across 45 active scenarios; not all produced a usable result |
Post-triage result set | 607 | The 558 blocked requests plus 49 documented WAF-relevant findings |
Blocked requests | 558 | The WAF stopped these before they reached the application |
WAF-relevant findings | 49 | Documented for remediation analysis after human review |
Everything else failed to count because the model could not produce a usable HTTP request, failed before reaching the target, or generated a benign payload.
Each non-blocked request had to clear five questions before it was counted as a finding:
Question | Why it matters |
|---|---|
Did the tester actually send a valid request? | If the model failed or the request never reached the target, the result tells us nothing about the WAF. |
Was the request clearly not blocked? | An ambiguous response is not enough to count. |
Was the request still malicious? | Changing a request to get it past the WAF can also make it harmless. |
Did the behavior belong to the WAF? | Some attacks only work through DNS or network paths the WAF cannot stop at request time. |
Could engineers reproduce it safely? | A fix needs a stable test case with a clear expected result. |
Cases that failed those checks were removed and duplicates merged. What remained became the input to rule, normalization, and mitigation work.
From findings to shipped detections
Not every finding warranted a new rule. Some exposed gaps in existing Managed Rules coverage, others concerned how the WAF normalized a request, and some belonged to a different security control. Each case was replayed to decide where the change belonged.
Related findings were grouped into four candidate rule sets. Every candidate was validated and tested against live traffic before it could protect customer traffic, with impact on legitimate traffic and false-positive risk assessed first. Common problems at this stage include:
Issue | Next step |
|---|---|
Missing or narrow detection | Review whether existing rules cover the finding |
Equivalent inputs interpreted differently | Engine or normalization review |
False-positive risk is too high | Revise or reject the candidate |
The effort produced three changes to Cloudflare's Managed Ruleset: new detections for SSRF - Obfuscated Host and SSRF - Restricted Protocol in the July 21 release, plus an improvement to the existing SSRF - Cloud rule. SSRF - Obfuscated Host came directly from requests that encoded internal addresses in non-standard numeric forms.
Working with a model that is not ground truth
The same scenarios were run with two versions of one model family. They generated different variations, but the same underlying issues surfaced in both. Stable replay and evidence capture made the runs comparable without treating either model's output as authoritative.
Adding attempts did not reliably add coverage. Near the end of the 25-attempt limit, some scenarios began repeating earlier ideas. Broader coverage came from more starting requests, more attack categories, and more input locations rather than longer sequences.
The model proposed requests; humans decided which mattered. A non-blocked request still required replay and review before it became a finding, a mitigation, or a regression test. Without that review step, there would have been no findings at all.
Hardening your own deployment
A WAF is one layer among several, and deploying all available controls compounds their effect. Begin by confirming Managed Rules and WAF Attack Score are set up correctly in front of the application. API Security, Bots and Fraud detection, and Threat Intelligence add further coverage. Positive security controls contribute a different kind of layer: rather than matching known attack patterns, they define the request shapes an application expects and flag anything outside that contract, sharply reducing attack surface.

Reproducing this experiment is unnecessary. To maximize deployed rules, run Managed Rules in log first, review matching requests in Security Events, and confirm legitimate traffic is unaffected before switching a rule to Block. Account teams can enable Attack Signature Detection on a zone, which simplifies reviewing matched traffic and deploying signature detections. Teams that already do application security testing should point those tests at a staging hostname protected by the same Cloudflare controls as production.
Where this goes next
Pairing adaptive, AI-driven testing with human triage and validation surfaced detection gaps that fixed tests would miss and converted them into stronger WAF protections and a higher block rate. A later post will cover follow-on testing with a white-box approach, in which the model knows both the application's vulnerabilities and the WAF rules guarding it.



