Cloud Firewalls Span 0% to 99.95%; AI Agent Traffic Makes the Choice Matter More

Bar chart comparing 2026 cloud network firewall security effectiveness scores from 0% to 99.95% across nine products

TL;DR · 30-second read

The Short Version

An independent testing group ran nine cloud firewalls through the same set of real attacks. A cloud firewall is the software gatekeeper that screens traffic before it reaches a company’s systems in the cloud.

Results ranged from blocking nearly everything to scoring zero. The built-in firewalls from Amazon and Microsoft scored zero on the test’s protection measure. Four products from specialist security companies scored above 99.7 percent.

The lesson: having a firewall is not the same as being protected. That matters more as companies let automated artificial intelligence assistants pass information to each other inside their networks.

On September 29, 2026, CyberRatings.org, a non-profit cybersecurity testing organization based in Austin, Texas, published its 2026 Cloud Network Firewall (CNFW) results, with testing carried out by NSS Labs. Nine products were evaluated. Four earned a Recommended rating, and each scored 99.7% or higher in Security Effectiveness: Fortinet FortiGate, HPE Juniper Networking vSRX, Palo Alto Networks VM-Series and Versa Networks Next Generation Firewall.

Check Point CloudGuard posted the highest score at 99.95% but was rated Neutral because of its above-average price per Mbps. AWS Network Firewall and Microsoft Azure Firewall each scored 0% and were rated Caution. Google Cloud Platform NGFW Enterprise (77.45%) and Cisco Secure Firewall Threat Defense Virtual (66.10%) were also rated Caution. The reports are free on the CyberRatings.org website.

Executive Summary

The headline finding is the spread. Products sold into the same category and tested under identical conditions landed anywhere from 0% to 99.95% on the same security measure. In practice, the label “cloud firewall” says little about what a product will stop. The specific product, and how it handles disguised and encrypted traffic, does.

The results also split along a line many cloud buyers care about. The firewalls built into cloud platforms were far less consistent than dedicated third-party products, which were deployed on Amazon Web Services for the test. Native firewalls are the path of least resistance for many teams because they are already inside the cloud account. CyberRatings’ advice is to treat them as a baseline, not a finish line, and to validate them before relying on them as a primary control.

The organization also flagged a forward-looking concern. AI models and agentic workflows, meaning software agents that take actions and call other agents, are expected to consume more network resources and require larger, more complex firewall policies. That includes more “east-west” traffic moving between internal systems. As that traffic grows, how well a firewall inspects it becomes a larger share of an organization’s actual exposure.

Native Convenience Versus Tested Protection

Of the nine products, the four rated Recommended all came from dedicated security vendors, and all cleared 99.7% Security Effectiveness. The three native offerings from the major cloud platforms scored 0% (AWS), 0% (Microsoft Azure) and 77.45% (Google Cloud). The split is not purely native versus third-party, though. Cisco’s virtual firewall, a third-party product, scored 66.10% and was also rated Caution. The more accurate reading is that results vary widely by product. Among the natives, none reached the level of the top group.

This matters because native firewalls have real structural advantages. They are billed through the same cloud account, managed through the same console and require no separate virtual appliance. For a team under deadline, that often decides the purchase. The test results suggest that convenience and protection are separate questions, and buyers should answer both rather than assume the first settles the second.

The ratings also weigh cost. Check Point scored highest of all at 99.95% but received a Neutral rating because of its above-average price per Mbps, a measure of cost per megabit per second of inspected throughput. A top score with a middling rating shows the ratings are a value judgment, not only a ranking of effectiveness. Buyers with different budgets or throughput needs may reasonably weigh that trade-off differently.

Why AI Agent Traffic Raises the Stakes of the Pick

The case for the title rests on three facts from the results. First, CyberRatings expects local AI models and agentic workflows to consume more network resources and to drive “larger, more complex policies, including east-west agent-to-agent traffic.” East-west traffic moves between systems inside a network, as opposed to north-south traffic entering or leaving it. Agents that call tools, query data stores and hand work to other agents generate exactly this kind of internal chatter. That means more flows, and more rules governing them, pass through whatever firewall sits in the path.

Second, the test showed how much depends on evasion handling. Evasions are techniques that disguise malicious traffic so inspection misses it. NSS Labs used 5,004 evasion attacks spanning 43 techniques across OSI Layers 3, 4 and 7: the network addressing layer, the transport layer that manages connections, and the application layer where web and API traffic lives. The organization notes that a single successful evasion can enable an entire class of exploits or malware. A firewall that is strong on average but weak against one evasion technique is therefore not partly protective against that class. It may not be protective at all.

Third, CyberRatings tells buyers to require TLS inspection, the ability to decrypt and inspect encrypted traffic, because a firewall that cannot do this leaves a substantial blind spot. Agent-to-agent and API traffic is commonly encrypted. Put together, more internal traffic, more complex policies and a large share of encrypted flows make the gap between a 99.7% product and a 0% product more consequential. The groups most affected are platform and security teams running AI workloads in public clouds, particularly those that have defaulted to the firewall their cloud provider supplies. The test did not measure AI traffic directly. The conclusion is an extrapolation from how the tested weaknesses interact with the traffic patterns CyberRatings expects agents to create.

Reading a 0% Score Carefully

A 0% Security Effectiveness score is striking, and it deserves precise interpretation rather than amplification. It reflects performance under one specific methodology, NSS Labs Cloud Network Firewall Test Methodology v4.1. That methodology used real-world exploits, evasions, malware, false-positive samples and sustained enterprise traffic loads. It measures how a product performed as configured for that test, at that point in time. It does not describe every possible deployment, and cloud providers frequently update their services.

CyberRatings itself points to that dynamism. Its first recommendation to buyers is to test regularly, because continuous delivery pipelines let vendors ship fixes quickly but can also introduce regressions that quietly erode protection. The same logic cuts both ways: a low score today can improve, and a high score is not permanent. The organization’s advice to hold vendors accountable, and to treat a lack of transparency as a red flag, applies to every product on the list, including the Recommended ones.

The practical takeaway is modest and concrete. Organizations relying on a native cloud firewall as their primary network control should validate exploit blocking, malware blocking, evasion resistance and TLS inspection against their own requirements. They should not treat the presence of the service as evidence of protection.

Background

CyberRatings.org is a non-profit membership organization that publishes independent ratings of security products, with NSS Labs generating its test results and reports. NSS Labs methodologies evaluate products against real-world attacks, evasions, malware and false-positive samples under sustained traffic loads. The resulting ratings of Recommended, Neutral and Caution combine protection results with value considerations such as price.

Organizations running workloads in public clouds generally choose between two kinds of network firewall. One is the native service offered by the cloud provider, such as AWS Network Firewall, Azure Firewall or Google Cloud NGFW. The other is a third-party vendor’s virtual firewall deployed inside the cloud account. Native services integrate closely with the provider’s billing and management tools. Third-party products often mirror features that security teams already run in their own data centers.

Sources

Source: CyberRatings.org and NSS Labs Announce 2026 Cloud Network Firewall Test Results, CyberRatings.org’s announcement of its 2026 comparative test of nine cloud network firewalls, conducted by NSS Labs.