Get a free audit

Field Notes / Threat Intelligence

AI Agents Are Transacting on Your Website. Your Penetration Test Still Assumes a Human.

Automated traffic grew eight times faster than human traffic in 2025. Traffic from AI agents, the kind that log in, fill forms and pay, grew 7,851%. Post-login compromise attempts hit 402,000 per organisation. HUMAN Security's 2026 benchmark puts numbers on a shift most penetration test scopes have not caught up with. Here is what an attacker does with each figure, and what a test must now cover: sessions, step-up, your own allowlists, loyalty logic and the checkout behind the API.

Author
Red Team Partners
Read
12 MIN READ
Filed
18 Sep 2026
An abstract field of connected autonomous agents, the software clients now browsing, logging in and paying on commercial websites without a person watching.

01 The traffic your scope assumes is human

HUMAN's platform processed more than one quadrillion interactions in 2025. The split matters more than the total. Training crawlers took 67.5% of AI traffic, scrapers 31.9%, and agents about 1.7% HUMAN Security 2026 . That last slice is the one that grew 7,851%, and the one that acts. An agent navigates pages, compares products, manages an account and completes a purchase with nobody watching the screen.

Where the agents go is the part to read twice. 77% of agentic activity landed on product and search pages. 8.8% was on account pages, 5% on authentication flows and 2.3% on checkout HUMAN Security 2026 . Two years ago a checkout completed by software was a fraud signal by definition. Today the same checkout can be a shopper's assistant. HUMAN's own line: same behaviour, different intent.

Three companies generate the bulk of it. OpenAI's bots accounted for 69% of AI-driven traffic, Meta 16% and Anthropic 11%, and over 95% of AI traffic concentrated in retail, media and travel HUMAN key findings . If you run an online shop, a booking engine or a customer portal in the UK, your access decisions about three operators shape most of your automated traffic. That is the reason an attacker wants to look like one of them.

For a tester the consequence is plain. A scope that says "web application, authenticated, two roles" tests the site a human sees. It never asks what happens when a client that does not sleep, does not mistype and holds dozens of sessions at once reaches the same endpoints. It never touches the rules you wrote to let the good agents through. Those rules are now part of your attack surface.

02 The attacker has moved past the login page

Overall account takeover volume fell by more than 30% in 2025. Post-login compromise attempts more than quadrupled, to an average of 402,000 per organisation HUMAN Security 2026 . HUMAN reads the two figures together: login defences are working, so attackers have moved to what happens after a legitimate login. They abuse session tokens, change account settings and walk through weak step-up controls to keep access. EMEA-sourced account takeover traffic exceeded 13% of login attempts, against under 3.5% globally HUMAN Security 2026 . A UK login page sits in the region taking the heaviest fire.

An attacker with a stolen or replayed session token does not need your password page. They test what the token still unlocks. Change the recovery email. Add a card. Redeem a balance. Each action passes on its own. HUMAN's guide makes the point with one example: a single password change looks harmless unless it follows dozens of login retries. The signal lives in the chain, and most controls inspect one link at a time.

A test must now cover the session after authentication, in full. Does the token survive a password change? Can it be replayed from a second device, a second country, a second user agent? Which sensitive actions demand a fresh factor, and can that step-up be skipped by calling the API behind the form instead of the form itself? We chain those questions on every authenticated engagement, because 402,000 attempts a year against the average organisation is a rate, and rates find gaps.

03 Your allowlist is the door

HUMAN's Satori threat intelligence team compared the declared identity of AI crawlers with behavioural and infrastructure signals. A significant portion of requests claiming to be ChatGPT, Mistral and Perplexity bots did not come from those operators' infrastructure HUMAN Security 2026 . The attacker copies a user-agent string and inherits the trust you extended to the real crawler: the robots.txt permission, the rate-limit exemption, the WAF rule someone added so the AI search bot would stop tripping alerts.

robots.txt was never a control. RFC 9309 describes rules a crawler chooses to honour RFC 9309 . The exemptions you attached to those names are controls, and they are public. The attacker reads the file, adopts the names, and enjoys the exemptions.

The same research found publicly exposed OpenClaw gateways, self-hosted agent front ends, doing three things you would not want done to you. They generated synthetic referral traffic with fabricated social-media UTM parameters, the tags your marketing team uses to attribute conversions and spend. They ran high-velocity directory brute-forcing against web applications. And they were targets for infostealer malware adapted to lift agent configuration secrets HUMAN Security 2026 . HUMAN's conclusion is that tools like OpenClaw lower the skill threshold: a person with no security training now runs attacks that used to need hands-on knowledge.

A test must attack every exemption you have granted on declared identity. We take your robots.txt and your WAF allowlist as the first input, present ourselves as each permitted crawler in turn, and measure what changes: rate limits, CAPTCHA, geo rules, bot scoring. Then we do the same with the UTM parameters your marketing team trusts, to see whether a fabricated campaign tag lands unfiltered in the dashboard your board reads. The fix the industry is converging on is cryptographic. HTTP Message Signatures under RFC 9421 let an agent prove who it is RFC 9421 . Until your stack verifies signatures, treat every user-agent string as an attacker's claim.

04 Carding, fake accounts and loyalty points at machine speed

Checkout interactions blocked grew more than 20% from 2024 and 250% from 2022 HUMAN Security 2026 . Satori documented one case that should rewrite every e-commerce scope. An AI agent added 11 cards and made 6 payment attempts across two sessions. When the cards failed, it pivoted to redeeming loyalty points HUMAN Security 2026 . The agent was verified. It was still misused. Who the agent is answers half the question. What it is allowed to do answers the other half.

Fake accounts detected per organisation grew 89% in 2025, on top of 259% the year before, and HUMAN describes them as infrastructure for incentive abuse, fraudulent orders and fake reviews HUMAN Security 2026 . Scraping attempts approach 20% of all traffic at the median and exceeded 61% for heavily targeted organisations. Scraping volume grew 47% year on year and 138% since 2022 HUMAN Security 2026 .

The attacker's playbook follows the numbers. Create accounts in bulk to farm sign-up credit. Chain refund and loyalty logic across those accounts. Use an agent to test stolen cards, and when the cards decline, cash out through whichever balance the site still honours. Pull the catalogue and the prices every hour.

A test must cover the logic, because the fingerprint will not save you. HUMAN notes that most fraud and chargeback tools rely on device fingerprints, IP reputation and cardholder history, and that those signals are weak or missing when traffic arrives through an agent HUMAN Security 2026 . So we test account creation at volume: what stops account number 500? We test the loyalty and refund state machine: can points be redeemed from a session whose cards were declined seconds earlier, and can a refund be routed to a card that was never charged? We test the agent-facing API and the checkout behind it: does the API enforce the same card-addition and retry limits as the form? And we test the price and stock endpoints to measure how much one client can take before anything reacts.

05 How we scope an agent-aware test

An agent-aware test is a standard authenticated web application test with six additions, run in a fixed order. The order matters because the first item changes the result of everything after it.

  1. Your own exemptions. robots.txt, WAF allowlists, rate-limit carve-outs, analytics filters. We impersonate what you trust and record what relaxes.
  2. Post-login state. Token lifetime and binding, replay across devices, step-up on the six sensitive actions, and API parity with the user interface.
  3. Volume paths. Account creation, promotion codes, loyalty accrual and redemption, refunds, all at machine cadence. Where the site slows us down, and where it never does.
  4. Checkout and agent-facing APIs. Card-addition limits, payment retry limits, and the pivot paths available after a decline.
  5. Marketing data integrity. Fabricated UTM parameters and referrer spoofing into your analytics, to show whether an outsider can move your attribution numbers.
  6. Scraping economics. How much of your catalogue, pricing and availability one client can pull per hour before a control reacts.

Every finding carries a request log, a timestamp and a fix. The report to the board opens with three lines. What an agent can do on your site today that a human customer cannot. What that costs you, in chargebacks, points liability and misattributed marketing spend. What closes it, and by when. Below that sit HUMAN's four priorities as a status column: AI traffic exposure audited, allowlisting on declared identity removed, an AI traffic governance policy in place, and a fraud stack that can tell a trusted agent from a bot HUMAN Security 2026 . A CISO reads the first page in four minutes. An IT lead hands the appendix to a developer and each item is reproducible.

06 If your clients run shops or portals

Only 13% of UK businesses ran a penetration test in the last twelve months DSIT CSBS 2025/26 . If you are a managed service provider, most of the other 87% who run an online shop, a booking engine or a customer portal are on your client list. Their platforms sit inside the 95% of AI traffic concentrated in retail, media and travel. Their marketing agencies ask you to allow the AI crawlers. Their fraud tooling was built for human buyers. When the chargeback wave or the loyalty-point drain arrives, the client rings you.

Three actions cover most of the exposure. Inventory which clients hold exemptions granted on declared identity, in robots.txt, in the WAF and in the CDN, and who asked for each one. Check whether any client's sensitive actions can complete on a stale session or through an API call that skips the step-up shown in the interface. Then make the agent-aware test a named line in the managed service, because the client's board is about to read the same numbers you just did. The Cyber Security and Resilience Bill gives that board a second reason to ask.

07 What to do this quarter

Pull your robots.txt and your WAF allowlist and write down every name that earns an exemption. Then list the actions a logged-in user can take that move money, points or contact details, and check each one for a fresh factor, on the form and on the API. Have an independent team run the six steps above against your shop or portal at machine cadence, fix what opens, and re-test. When your board asks what the agents on your site can do, you will answer with a request log.

That is the work we do. The same CREST-certified test covers the standard web application scope and the six agent-aware additions, with every finding human-verified, a fix beside each one, and a year of re-tests so the proof does not age while the traffic keeps growing. Get a free audit and see which of your own exemptions an attacker would borrow first. If you are an MSP and want to offer this to your clients under your own brand, start at our partner page.

References

Sources

  1. HUMAN Security. The CISO's Guide to AI and Agentic Traffic, 2026. Based on the 2026 State of AI Traffic & Cyberthreat Benchmark Report (Human Defense Platform telemetry, January to December 2025). humansecurity.com
  2. HUMAN Security. Measuring the AI-Driven Internet with the 2026 State of AI Traffic & Cyberthreat Benchmark Report: key findings. humansecurity.com
  3. IETF. RFC 9421: HTTP Message Signatures. February 2024. rfc-editor.org
  4. IETF. RFC 9309: Robots Exclusion Protocol. September 2022. rfc-editor.org
  5. Semrush. AI search visitors convert 4.4x better than traditional organic visitors (2025 study), as reported by HUMAN Security. Semrush's own summary of the finding. semrush.com
  6. Adobe. 85% of consumers who tried agentic shopping said it improved the experience, as reported by HUMAN Security in the CISO's Guide to AI and Agentic Traffic, 2026. humansecurity.com
  7. Gartner. Prediction that by 2028 AI agent machine customers will replace 20% of interactions at human-readable digital storefronts, as reported by HUMAN Security in the CISO's Guide to AI and Agentic Traffic, 2026. humansecurity.com
  8. Department for Science, Innovation and Technology / Home Office. Cyber Security Breaches Survey 2025/2026. Official statistics, 30 April 2026. gov.uk