Part 2 of the AI firewall walkthrough. What Palo Alto and Fortinet claim against what the independent tests show, the four evasions our red team used against the firewall we bought, the IoT problem at the scale we found it, and the three questions to ask any “AI-powered” vendor before signing.
In Part 1, I walked through the ML-powered firewall we chose for a 6,000-seat enterprise after a four-week proof of concept: six lanes, ten model families, a Cobalt Strike beacon blocked in ten minutes. The picture is clean on a slide. The year that followed was messier. Attackers adapt. Enterprise networks are noisy. Models trained in a vendor lab degrade against traffic that looks nothing like the training set. This part covers what the leading vendors claim and what the independent tests show, the four techniques the red team used to get past the firewall, what the IoT problem looked like once we could see it, and the three questions I now put to every vendor before a purchase order.
Key Takeaways
- Vendor numbers are mostly volume numbers. The figure that matters is the false-positive rate on your own traffic, and only a four-week proof of concept gives you that.
- Our red team got past the firewall four ways: domain fronting, a shadow IoT device, adversarial file padding, and slow-drip exfiltration. Every fix was policy or architecture, not a better model.
- The enterprise had 23,000 connected devices where the inventory said 9,000. Segmentation fixed that; no model could have.
- An AI firewall needs roughly one engineer per 8,000 to 15,000 endpoints to tune it. Without that person it is a box that blinks.
What Is Inside
Vendor Numbers, Translated
A handful of vendors dominate the firewall market, and each publishes its own AI claims. Three sets of numbers came up in every meeting during our evaluation. Here is what each one measures, and what it looked like from inside the proof of concept.
Free to use, share it in your presentations, blogs, or learning materials.
Palo Alto Networks: “2 billion security events processed per day” across its AI-enhanced firewall platform. That is real, and it is a volume number, not an accuracy number. Volume is useful context, since the model has seen a lot of traffic, but it says nothing about precision or recall on your network. Palo Alto has also claimed a 42% improvement in unknown-threat detection from PAN-OS 9 to PAN-OS 10. That one is measurable, if you run before-and-after tests on the same traffic corpus. We did, on our own corpus, and got 34%. Close enough to be honest marketing.
Fortinet: FortiGuard AI delivers “zero-day signatures in seconds” instead of the minutes to hours of the old cloud delivery. Also real. The pipeline squeezes the gap between a sandbox verdict and a deployed signature to single-digit seconds. The caveat: it only matters if the sandbox sees the file. Targeted malware that never reaches a sandbox never gets a signature, however fast the delivery. Our September installer incident, from Part 1, is the case for this claim. The February beacon is the case against it.
Independent benchmarks (NSS Labs, ICSA Labs, MITRE ATT&CK Evaluations) publish detection rates on specific threat corpora. Results move year to year and vendor to vendor, but the pattern holds. Signature detection above 90% on known threats. ML detection between 40 and 65% on genuinely new threats. Combined rates of 85 to 95% with both layers on. No vendor hits 100%. The gap between 95% and 100% is where the expensive breaches live, and it is the gap the datasheets never mention.
The single most useful number in a vendor evaluation is the false-positive rate under your traffic, not theirs. Vendors demo on clean corpora. Your network is not a clean corpus. Our two finalists flagged 0.3% and 1.1% of real sessions in week one, and that one figure decided the purchase more than any detection score did. Run a proof of concept on at least four weeks of real traffic before buying, because the behavioural models need ten days just to learn what your normal looks like.
Four Ways the Red Team Got Past the Firewall
After the purchase, the enterprise’s red team spent a quarter trying to beat the firewall we had just installed. They got through four ways. Each one has been seen in the wild in the last 18 months, and each needed a different fix.
Free to use, share it in your presentations, blogs, or learning materials.
1. Domain Fronting and CDN Abuse
Domain fronting is the trick where the TLS SNI header and the HTTP Host header point at different domains. The firewall sees one, a legitimate CDN-hosted site. The real destination is something else. For years this let attackers tunnel command-and-control through CDN edges the firewall could not block without also blocking real business. The big CDNs have removed fronting support (Google Cloud and AWS CloudFront both did between 2018 and 2020), but it still works on a long tail of smaller CDNs and self-hosted edges. Our red team found one in an afternoon. Policy was the fix: CDN egress is now blocked for the two network zones that never need it, and decrypted and inspected everywhere else.
2. IoT Shadow Devices at Scale
A 5,000-person enterprise has 15,000 to 40,000 connected devices. Printers, cameras, badge readers, thermostats, smart TVs, factory sensors, medical devices, personal phones. The security team knows about maybe 60% of them. Everything else is a shadow device. So the red team plugged a Raspberry Pi into a meeting-room network port, gave it the traffic profile of a conference phone, and ran it for two weeks. Our IoT model clustered it with the phones and nobody looked twice. ML fingerprinting narrows the gap; it cannot close it, because a device that behaves like a sanctioned one is, to the model, a sanctioned one. The fix is segmentation. IoT lives on its own VLAN with a restrictive egress policy, and the meeting-room ports now need 802.1X. Any model that tries to separate good IoT from bad IoT on a flat network is fighting a battle it cannot win.
3. Adversarial Inputs Against the Models
The models themselves can be attacked. An attacker who can observe how the firewall reacts can craft inputs that sit just under the decision boundary. Research since 2020 has shown this against malware classifiers, URL classifiers, and flow classifiers. For file classifiers, small byte-level changes (padding, inserting harmless sections) can drop a confident-malicious score to benign without changing what the malware does. For URL classifiers, small lexical changes can move a phishing URL just outside the detected range. Our red team took a known-bad executable, padded it with 2 MB of benign library code, and the inline CNN score fell from 0.94 to 0.31. Ninety seconds later the sandbox still caught it on detonation, which is the point of having both. Countermeasures exist, adversarial training and model diversity, so an attack that works on one model fails on another. Most production firewalls have not done either rigorously, and ours has not.
4. Slow-Drip Exfiltration
The LSTM and autoencoder models that catch behavioural anomalies need enough samples to see a pattern. An attacker who moves 10 MB a day for 120 days might never trigger a behavioural alert, even though the cumulative 1.2 GB is significant. The red team ran a slower version: 6 MB a day to a fresh cloud bucket. The firewall never flagged it. What caught it, on day 19, was the cumulative volume cap per new destination that we had added after the Cobalt Strike test in Part 1. Neither the cap nor the destination-reputation check is machine learning. Both are rules, informed by ML signals, and they caught what the models missed.
The Two We Could Not Fix
Two red-team results have no clean answer, and I would rather say so than pretend. First, the padded executable. Padding fools the inline CNN, the vendor knows it, and the honest mitigation is that the sandbox catches the file 90 seconds later. Ninety seconds is enough for a dropper to run. The enterprise now blocks execution of unsigned binaries on endpoints, which moves the problem off the firewall entirely. The second is the four-hour beacon with heavy jitter. No timing model catches a beacon that talks six times a day, and the volume cap does not trigger on a few kilobytes. What caught the intern’s laptop in February was destination reputation ageing: a domain registered 30 days earlier, never seen by anyone else, talking to one host. That is threat intelligence, not machine learning, and it is the oldest trick in the book.
The IoT Problem, at the Scale We Found It
IoT gets its own section because it is the largest AI-resistant problem in enterprise firewalls today, and because the proof of concept made it visible for the first time. What we found at the enterprise, and what the industry numbers say:
- The inventory listed 9,000 devices. The first fingerprinting pass found 23,000, which works out to just under four per employee. The industry range is three to eight.
- About 1,400 of those devices still had default credentials, and 600 had no patch path at all, because the vendor was gone or the contract forbade updates.
- The Mirai botnet family still accounts for 30 to 60% of compromised-IoT traffic seen by threat-intelligence feeds, six years after the original campaign.
- The worst cases were the building-management controllers and two lab instruments: frozen operating systems, no updates permitted, full network access on the subnet they were plugged into.
ML fingerprinting identifies devices well when the model has seen that device class. HDBSCAN clusters by behaviour; a random forest names the cluster. When an unknown device appears, the model returns low confidence on every known class, and the firewall either flags it for review or applies a default “unknown IoT” policy. The policy decides what happens next, not the model. Ours is “deny all internet egress from unknown IoT until classified”, which is aggressive and has so far been right every time.
The real answer to IoT visibility is not better ML. It is network architecture. IoT goes on a segmented VLAN with explicit allow-lists for outbound destinations: firmware servers, the cloud APIs the device genuinely needs. Everything else is denied. That took the enterprise’s network team four months, it was the least glamorous work of the year, and it removed more risk than the firewall did.
How the Segmentation Project Actually Went
The plan was simple on paper. Three new VLANs: known IoT, unknown IoT, and building systems. Each with its own egress allow-list. Each device moved once it was classified. The first week moved 4,000 devices and broke the badge readers, the same badge readers from the cut-over weekend, because the door controllers talk to a cloud service nobody had documented. Week two moved the printers and broke scan-to-email. Week three moved the cameras and broke nothing, which everyone found suspicious until we checked. By month four, 21,000 of the 23,000 devices sat on a segmented VLAN. The remaining 2,000 are the lab instruments and building controllers, which live on a fourth VLAN with no internet egress at all and a jump host for the two vendors who service them.
What the firewall contributed to that project was the inventory and the classification. What it could not contribute was the decision about what each class of device is allowed to reach. That took a spreadsheet, the facilities team, and eleven vendor phone calls. The model told us what was on the network; people decided what it was allowed to do, and that split is the whole story of AI in security.
What Happens Without a Human to Tune It
A pattern I have now seen at three organisations: buy an AI firewall, deploy it in default mode, never tune it. The models ship with thresholds chosen for vendor benchmarks. Those thresholds are almost always too loose for a regulated environment and too tight for a high-traffic commercial one. Untuned, the firewall either misses threats or blocks real traffic faster than the help desk can answer the phone. Our month-one alert count of 1,800, of which the SOC could act on 300, is what untuned looks like.
The minimum staffing for a genuine AI-firewall deployment is one full-time security engineer per 8,000 to 15,000 endpoints. That engineer tunes policies, reviews false-positive reports, allow-lists real applications, works flagged alerts, and coordinates with the SOC. The enterprise hired that person in month three, after two months of the network team doing it badly in their spare time. Within eight weeks of the hire, the alert count fell from 1,800 to 240 a month.
The vendors will not tell you this. A firewall without a tuning engineer is a box that blinks. Somebody has to decide what normal looks like in your network, and the model cannot.
Three Questions for Any AI Firewall Vendor
These are the three questions that separated the vendors with real ML from the vendors with marketing ML during our evaluation. I now ask them on every call.
Free to use, share it in your presentations, blogs, or learning materials.
Question 1: “How often do your models retrain, and on what data?”
A good answer: monthly to quarterly for the main threat models, on data from the vendor’s global sensor network plus customer verdicts given under consent. A bad answer: “our models are very accurate”, which dodges the question. A worse answer: “our models do not need retraining”. Threat patterns shift constantly. A model that never retrains is a model that is quietly getting worse. One of our vendors answered in an hour with dates. The other took nine days, and the answer was vague when it came.
Question 2: “What is your false-positive rate on independent benchmark corpora?”
A good answer: specific numbers from NSS Labs, ICSA, or MITRE ATT&CK Evaluations. A bad answer: “we do not publish benchmark numbers because they help attackers”. That is a red flag. Independent benchmark participation is table stakes for a serious vendor. If a vendor does not take part, assume the model would not survive the test.
Question 3: “Show me a customer case where your ML caught a threat signatures missed, with the features that drove it.”
A good answer: a detailed technical narrative, possibly under NDA. A bad answer: a glossy one-pager about “significant improvements”. A vendor that cannot explain how its ML caught one specific thing does not have a trustworthy ML story. The best vendor engineer we met walked us through the feature contributions on a real detection, the feature engineering behind them, and the retrain that produced the current model. That 40-minute conversation told us more than four weeks of datasheets.
What the Year’s Numbers Looked Like
Twelve months after cut-over, the numbers the enterprise’s security lead put in front of the board were these. Blocked threats per month: about 41,000, of which signatures caught 81%, ML 11%, and policy 8%. Sandbox verdicts: 4,300 files detonated, 140 malicious, 131 of them unknown to the signature feed at the time. ML alerts to the SOC: down from 1,800 a month to 240, with a confirm rate that rose from 13% to 19%. Mean time from alert to ticket: 22 minutes, down from four hours under the old firewall, mostly because the new alerts are worth reading. Incidents that reached a laptop or server: two, the trojanised installer and the intern’s beacon, both contained inside a day.
Two numbers went the wrong way. Help-desk tickets about blocked sites rose 30% in the first quarter and took until month five to fall below the old baseline. And the firewall’s own licence cost is 2.4 times the old one. The board asked whether the second number was worth it. The honest answer was that the installer catch alone would have cost more than the licence, and that the segmentation project the firewall made possible was worth more than either. Nobody on the board asked about detection rates. They asked what got in, and the answer was two things, both caught.
What I Would Tell a CISO Before Buying
Five things, in the order they would have saved us time. First, run the proof of concept for four weeks and ignore week one. Second, score the false-positive rate on your own traffic above every detection number. Third, budget the tuning engineer into the purchase, because the firewall without that person is the old firewall with a higher licence fee. Fourth, treat the IoT inventory as the first deliverable of the project and plan the segmentation before the cut-over, not after. Fifth, write the volume cap per new destination and the unknown-device egress rule on day one. Both are rules. Both caught things the models did not. Neither is in any datasheet.
What the Security Team Still Owns
The pattern across this whole series has held. Models do the mechanical work. Humans do the judgement, the tuning, the exceptions, and the accountability. Firewalls are no different. A SOC analyst working an ML-flagged alert still does what an analyst did 15 years ago: pivot through the data, correlate across sources, decide whether the alert is a real compromise, run the playbook. What has changed is the volume and the character of what reaches the analyst. Fewer alerts, and sharper ones, because the ML filter removed the obvious noise.
The firewall engineering team owns something else: the deployment itself. Where the firewall sits, what traffic it sees, how it connects to the rest of the security stack, who has policy authority. None of these are decisions the AI can make. A firewall placed wrong sees the wrong traffic. A firewall without SIEM forwarding or SOAR hooks produces alerts nobody acts on, which is exactly what happened with the second vendor’s beacon test in Part 1. These are human problems, and they are where security programmes live or die.
Closing the Series
This is the last article in a four-domain series on what AI actually does: lending, banking, insurance, and firewalls. Eight articles, four projects, one lesson that never changed. AI in production is a collection of specialist models stitched together by policy, not a single intelligent system making decisions. The marketing is louder than the architecture. The real work is in the plumbing, the feedback loops, and the human judgement that sits above the models. In every one of the four projects, the fixes that mattered most were a rule, a report, a threshold, or a contract clause. Not once was it a better model.
If this series has been useful, the next one covers what happens when AI is asked to do something it is not suited for, and how to recognise those cases before the project ships. The lending, banking, and insurance pairs are linked below.
Related Reading
References
- Palo Alto Networks, Introducing 4th Generation ML-Powered NGFWs
- Tech Today Global, Next-Generation Firewall AI Revolution, 2025
- MITRE, ATT&CK Evaluations
- MDPI, Detection of Malware by Deep Learning as CNN-LSTM Techniques in Real Time
- Palo Alto Networks, AI-Powered Next Generation Hardware Firewall
Frequently Asked Questions
Combined signature-plus-ML detection from independent benchmarks (NSS Labs, ICSA, MITRE ATT&CK) lands between 85% and 95% for leading vendors. ML alone catches 40 to 65% of genuinely new threats. No vendor reaches 100%, and the gap between 95% and 100% is where expensive breaches happen.
Yes, in four main ways: domain fronting and CDN abuse, shadow IoT devices on flat networks, adversarial inputs crafted just under the model’s decision boundary, and slow-drip exfiltration under the window the behavioural models need. Each has a different countermeasure, mostly in policy and network design rather than in ML.
Leading vendors retrain the main threat models monthly to quarterly, on global sensor data plus customer verdicts given under consent. A vendor that cannot answer with dates, or that claims its models never need retraining, is hiding a problem.
Yes. Plan on one full-time security engineer per 8,000 to 15,000 endpoints. That person tunes policies, allow-lists real applications, works alerts, and coordinates with the SOC. Without that role the firewall will not reach the detection rates on its datasheet.
Shadow IoT. Enterprises average three to eight connected devices per employee and can inventory about 60% of them. The rest run old firmware, often with default credentials. The fix is network segmentation with explicit allow-lists for IoT egress, not better ML.
At least four weeks on a mirror of real traffic, with the same test cases fed to each vendor. The behavioural models need about ten days of your own traffic before their baselines settle, so anything shorter measures the vendor’s lab defaults rather than the product.
Sometimes. Padding a known-bad executable with benign code can drop an inline CNN score from confident-malicious to benign without changing what the file does. The cloud sandbox usually still catches it on detonation, which is why a serious deployment runs both layers rather than trusting either alone.
