AI Is Getting Good at Finding Bugs. What Does That Mean for Your Site?
AI Is Getting Good at Finding Bugs. What Does That Mean for Your Site?
As of September 2026, the evidence is that machine assisted vulnerability research now works on exactly the class of bug your custom code contains, and that attackers are already using it at scale. The practical implication is not panic. It is that your patch window has shrunk and your build pipeline is now a target.
Two Google publications, read together, tell the whole story. One shows what defenders are doing with this. The other shows what attackers are doing with it.
The gap between the two is where most B2B websites sit, and it is worth understanding before someone sends you a scary vendor email about it.
What Has Actually Been Demonstrated?
That an AI agent can find real, previously unknown vulnerabilities in widely used software. Google's Big Sleep, described in its own announcement as "an AI agent developed by Google DeepMind and Google Project Zero," reached that milestone some time ago.
Google's July 2025 update stated that "by November 2024, Big Sleep was able to find its first real-world security vulnerability" and that "since then, Big Sleep has continued to discover multiple real-world vulnerabilities, exceeding our expectations." It also reported that the agent found an SQLite vulnerability, CVE-2025-6965, and claimed "this is the first time an AI agent has been used to directly foil efforts to exploit a vulnerability in the wild."
That is the defensive side, and it is genuinely good news. Bugs in the open source libraries underneath your site are being found and fixed faster than before.
What Are Attackers Doing With the Same Capability?
Using it in production, at scale. The Google Threat Intelligence Group report published on 11 May 2026 describes "a maturing transition from nascent AI-enabled operations to the industrial-scale application of generative models within adversarial workflows."
The headline finding is specific. "For the first time," the report states, "GTIG has identified a threat actor using a zero-day exploit that we believe was developed with AI. The criminal threat actor planned to use it in a mass exploitation operation but our proactive counter discovery may have prevented its use."
It also describes a shift in how models are used offensively, where "the LLM is no longer merely a passive advisor but an active participant in the offensive chain, capable of orchestrating complex toolsets and making tactical decisions at machine speed."
Read that last phrase carefully if you maintain a website. Machine speed is the part that matters to you, because it compresses the time between a patch being published and an exploit being attempted.
Which Kinds of Bugs Does This Find Best?
The ones your own developers wrote, rather than the ones in memory management. This is the detail that should change how you think about your risk.
The GTIG zero-day case was a two factor authentication bypass in a Python script for a popular open source web administration tool. The report describes its root cause as a "high-level semantic logic flaw where the developer hardcoded a trust assumption."
The report explains why models are strong here, noting that "frontier LLMs excel at identifying these types of high-level flaws and hardcoded static anomalies" and that "they have an increasing ability to perform contextual reasoning, effectively reading the developer's intent to correlate the 2FA enforcement logic with the contradictions of its hardcoded exceptions."
Traditional scanners look for crashes. This looks for intent that does not match implementation. That is the category your bespoke form handler, your webhook receiver, and your gated content check all belong to.
Does This Mean Your Marketing Site Is at Risk?
Your static pages are not the concern. Anything on your site that makes a decision is.
A marketing site with no logic is a very small target. Add a login for a customer portal, a gated resource check, a form that writes to a CRM, a webhook endpoint, or a coupon validation, and you have added code that reasons about trust. Those are exactly the semantic logic flaws described above.
The honest assessment for most B2B sites is that the site itself is a modest target and the systems it connects to are the real prize. A form handler with an over-permissioned API key is worth far more to an attacker than your homepage.
Why Is Your Build Pipeline Suddenly the Bigger Problem?
Because that is where the credentials live, and attackers have noticed. The GTIG report documents a campaign that should worry anyone with a continuous integration setup.
It reports that "in late March 2026, the cyber crime threat actor 'TeamPCP' (aka UNC6780) claimed responsibility for multiple supply chain compromises of popular GitHub repositories and associated GitHub Actions, including those associated with the Trivy vulnerability scanner, Checkmarx, LiteLLM, and BerriAI."
The method matters. According to the report, "TeamPCP gained initial access through compromised PyPI packages and malicious pull requests to these GitHub repositories," embedding a credential stealer that extracted "high-value cloud secrets, such as AWS keys and GitHub tokens, directly from affected build environments."
Note which tools were hit. A vulnerability scanner and a security testing vendor. The things you install to be safer are also dependencies, and dependencies are the attack surface. Our notes on keeping dependencies current on a marketing site are more urgent in that light than they were a year ago.
What Does the LiteLLM Compromise Tell You Specifically?
That your AI integrations are now part of your security perimeter. GTIG singles this one out, observing that "the compromise of LiteLLM, an AI gateway utility for integrating multiple LLM providers is noteworthy."
Its reasoning is that this "highlights the expanding attack surface of AI platforms and the potential for impact across the software supply chain," and that "this incident could lead to considerable exposure of AI API secrets from affected victims, which could be used to gain further access to systems for traditional intrusion operations."
If you added an AI feature to your site or your content pipeline this year, you added a key. That key probably has a generous quota and it may well sit in the same environment as everything else. Treat it as a credential of the same class as your database password, which is the argument we made in where API keys belong.
What Should You Actually Change This Quarter?
Four things, in order of how much risk they remove per hour spent.
First, shorten the distance between a published patch and your deployed site. If your dependency updates happen when someone remembers, they happen too slowly now. Automate the pull requests and set an actual review cadence.
Second, scope every credential your build and your site hold. A deploy token that can only deploy. A CMS token that can only read. An AI key with a spend cap. Each one reduces what a stolen secret is worth.
Third, pin and review what your pipeline runs. A third party action referenced by a moving tag is a decision to trust whatever that tag points at tomorrow.
Fourth, read your own trust logic. Every place your code decides whether someone is allowed to do something, read it as if you were looking for the hardcoded assumption. That is the bug class being found now.
Should You Use AI to Review Your Own Code?
Yes, and with realistic expectations. The same strength that makes this useful offensively makes it useful on your side: a model reading your authorisation logic and asking why a condition exists is doing something a linter cannot.
What it will not do is replace a scanner or a dependency audit. Those catch known problems reliably and cheaply. A model is better at the unknown, unglamorous logic question, and worse at exhaustive coverage. Run both.
The practical version for a small team is unremarkable. When you write or change anything that makes a trust decision, have a model read the diff and ask it what an attacker would try. It takes a minute and it occasionally earns its keep spectacularly.
Is Any of This a Reason to Change How You Build?
Mostly it is a reason to keep doing the boring things properly, with less tolerance for delay. None of the recommendations above are new. What has changed is the cost of ignoring them, because the window between a vulnerability becoming public and being exploited is narrowing.
The other shift worth internalising is where your risk actually lives. For most of the sites we build, the attack surface is not the pages. It is the pipeline that builds them and the keys that pipeline holds. That is where the attention belongs, and it is not where most website security conversations start. Our general guide to website security covers the fundamentals underneath all of this.
If you are not sure what credentials your site and its build currently hold, that is the first thing worth finding out, and it is usually a short piece of work. We are happy to go through it with you at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.