Cloudflare Tests WAF Against Frontier AI Models, Finds Gaps and Fixes Them

Cloudflare's internal test with frontier AI models reveals both the strength of its WAF and specific vulnerabilities that were quickly patched.

Last Updated: September 29, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: September 29, 2026

September 29, 2026, (Inside AI) — Cloudflare has revealed the results of an internal experiment testing its Web Application Firewall (WAF) against frontier AI models. The company built an adaptive system that uses large language models (LLMs) to generate and mutate attack payloads, aiming to see if the WAF could withstand AI-driven attacks. The test, conducted on an authorized customer staging environment, involved 1,107 attempts across six attack categories. The vast majority of attacks were blocked, but the exercise uncovered specific gaps that led to three updates in Cloudflare's Managed Ruleset.

The experiment underscores a growing concern in cybersecurity: as AI models become more capable, they can automate the discovery of vulnerabilities at a speed and scale unmatched by human hackers. Cloudflare's proactive approach aims to stay ahead of this threat by using the same technology to harden its defenses. The findings provide a rare glimpse into how AI can be leveraged both offensively and defensively in the ongoing battle to secure web applications.

Cloudflare's WAF tester works by starting with a known exploit and then iteratively changing how it is encoded or delivered. The LLM proposes variations, and a second review call analyzes the response to decide the next step. The system runs without access to WAF internals, simulating an external attacker. "The model generated requests. We decided which ones mattered," the company stated, emphasizing the role of human review in triaging results.

The test targeted six attack types: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The WAF was configured with Attack Score blocking scores of 30 or below, all Managed Rules enabled, and the OWASP Core Ruleset at Paranoia Level 3. After running 45 scenarios, the system generated 1,107 attempts. XSS, LFI, SQLi, and Log4j had near full coverage, meaning almost all variations were blocked. However, 49 findings emerged, with 48 belonging to CMDi and SSRF.

Read: OpenAI Releases GPT-6 Astra Model With Critical Cybersecurity Capability

One notable SSRF scenario involved a cloud metadata address. The model tried different representations of the IP, such as integer, octal, and trailing-dot forms. The WAF blocked all but one: a trailing-dot form that resulted in a redirect rather than a block. This specific case led to a new detection rule for obfuscated hosts. "The SSRF - Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms," Cloudflare explained.

Cloudflare also learned that more attempts within a single scenario did not always yield better results. Some scenarios began repeating earlier ideas near the 25-attempt limit. Broader coverage came from testing more starting requests, attack categories, and input locations. The company ran the same scenarios with two versions of the same model family, and both produced different variations but surfaced the same underlying issues. This consistency allowed for reliable comparison without treating either model's output as ground truth.

The findings were not immediately turned into rules. Each non-blocked request underwent human review to rule out false positives, malformed requests, or benign payloads. Only then were they considered for mitigation. Cloudflare grouped related findings into four sets of candidate rules, validated each, and tested them against live traffic before deployment. This careful process ensured that new rules would not disrupt legitimate traffic.

For customers, Cloudflare recommends ensuring that Managed Rules and WAF Attack Score are correctly configured. They also suggest deploying additional layers such as API Security, Bots and Fraud detection, and Threat Intelligence. Positive security controls, which define expected request shapes, can further reduce attack surface. Cloudflare advises running Managed Rules in log mode first, reviewing matches in Security Events, and confirming no impact on legitimate traffic before switching to block mode. The company also offers an Attack Signature Detection feature that simplifies reviewing matched traffic.

Read: OpenAI and 100 Others Warn Window to Defend Against AI Attacks Is Narrowing

This experiment is part of Cloudflare's broader effort to integrate AI into its security development lifecycle. By combining adaptive AI-driven testing with human triage, the company aims to find detection gaps that fixed tests might miss. In a future post, Cloudflare plans to share results from a white-box approach where the model knows both the application's vulnerabilities and the WAF rules. This ongoing research highlights the evolving arms race between AI-powered attacks and defenses, with Cloudflare positioning itself at the forefront.

More from Inside AI

  • Features, Interviews, Press Releases

    Beyond Transcripts: Modulate Secures $25M to Scale Frontier Audio-Native AI Architecture Against Monolithic LLMs

    September 28, 2026
  • Cybersecurity AI

    Cloudflare Launches Adaptive Security Framework to Counter AI-Driven Attacks

    September 29, 2026
  • Cybersecurity AI

    Cloudflare Tests WAF Against Frontier AI Models, Finds Gaps and Fixes Them

    September 29, 2026
  • AI In Business

    India’s Bharat-VISTAAR AI Platform Targets Farm Advisory for Smallholders

    September 29, 2026
  • AI In Business

    HBR Research Reveals AI Intensifies Work and Erodes Judgment

    September 29, 2026
  • AI Policy & Regulation

    Hawley Bill Would Make AI Companies Liable for Rogue Agents

    September 29, 2026
  • Agentic AI

    Meta’s Muse AI: User Control vs. Autonomy, Says VP Vishal Shah

    September 29, 2026
  • Artificial Intelligence (AI)

    Harvard AI Tutor Study Sparks Viral ‘Universities Are Obsolete’ Claims

    September 29, 2026
  • AI Safety

    OpenAI Proposes Safer AI Training With Three-Pillar Plan

    September 29, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Join Our Newsletter Community

Subscribe

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Advertise with us
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital