Peter Garraghan, Founder and Chief Science Officer, Mindgard.
You may have noticed a pattern emerging in the AI industry. AI models and agents operating outside the confines of their intended goals, taking actions that their creators never intended (“breaking free of their sandbox”), resulting in security breaches or other similarly damaging outcomes. You may also have noticed that AI vendors are surprisingly willing to speak about these incidents, treating them as an opportunity to explain just how impressive and capable their models must be in order to achieve such a thing.
It has become a new form of stealth marketing. AI providers now have an opportunity to champion their models while simultaneously offering concern over the very outcomes that they themselves orchestrated. It has created an environment whereby it is difficult for consumers to distinguish between deliberately manufactured scenarios and genuine AI risks, with coverage downplaying the seriousness of a vendor breaching their own customer’s systems. However, if today’s businesses (and individuals) want to protect themselves effectively, understanding that distinction is increasingly critical.
Any Publicity Is Good Publicity
It strikes me as an unlikely coincidence that when one frontier AI model made headlines for causing a major security incident, the other major vendors immediately claimed their own models had achieved the very same. Maybe every widely used AI model broke free from its shackles at the exact same time (possible, although unlikely). Instead, let us look at the broader state of AI. Multiple major providers are positioning themselves for IPO, jockeying for mindshare within a crowded market. While coverage of AI-driven security incidents might seem like “bad press,” the truth is just the opposite. AI vendors seeking to generate hype recognize that there is real value in their models being seen as too “smart,” “advanced” or “dangerous.”
The following two statements can be true. First, that these AI models have produced security incidents affecting customers and the core AI infrastructure they depend on, and second, that vendors providing those models are capitalizing on this, and simultaneously running vast marketing campaigns that champion their models’ capabilities. I don’t believe these vendors have deliberately orchestrated these security incidents (if this was discovered, the reputational fallout would be catastrophic), and certainly there is no indication that those responsible for running the experiments have been terminated. However, this scenario raises serious questions pertaining to how seriously AI providers take their security practices if they cannot establish a sandbox secure enough to prevent the hacking of broader networks.
Hijacking The AI Safety Narrative
The way these incidents have been covered blurs the line between what constitutes a genuine threat and what is seen as a marketing opportunity. Naturally, an AI model provider is going to avoid creating a narrative that puts them in a bad light. However, the optics here are troubling. While none of these security incidents caused significant damage, this isn’t the main point. As an analogy, if your house was broken into, you wouldn’t say, “Yes, they broke in and looked around, but they didn’t take anything, so it’s fine.” It may have worked out this time, but there will be a next time, especially if no lessons are learned, if no additional precautions are put into action.
It does not appear that these experiments were being treated with the seriousness that cybersecurity testing in a live environment typically demands. One would hope this prompts them to make their controls for experimentation and system configuration more stringent in the future. Having the AI (which frontier providers have repeatedly told the media can be “dangerous”) operating in an air-gapped environment would be a good start. Given that AI vendors appear more interested in leveraging these incidents as marketing capital, it is difficult to have faith that those safeguards are being implemented.
Instead of being treated as a solemn and serious issue, the narrative around these AI-driven security incidents is being hijacked. If a major cybersecurity vendor announced that its product was accidentally hacking its users, other vendors would certainly not be eagerly standing up to say, “We did it, too!” The fact that AI vendors are doing exactly this represents a deeply worrying means to approach security, responsibility and trust. A responsible vendor disclosing an incident should focus on containment, disclosure and remediation, and not how impressive and capable the model was (important to note that the teams behind the scenes have been doing this, although even this has been leveraged as a marketing opportunity).
With little guidance from AI vendors or the media, the responsibility will ultimately fall to customers to gauge which incidents present a cause for concern. This should begin with ascertaining what actually happened: What controls failed, what access did the system have and will your business face the same conditions? It is important to treat any AI systems with access to tools, data or infrastructure as untrusted users. This entails limiting permissions, isolating sensitive systems, monitoring actions and maintaining controls outside the model itself. It would not be wise to assume the vendor’s safeguards are sufficient based solely on the sales team’s assurances.
The Need For A Practical Approach
It’s impossible to deny that frontier AI models are capable of conducting advanced cyberattacks. Scientists and researchers operating in this space have seen them grow more and more capable over the past several years, and we are now approaching a tipping point. However, that doesn’t mean every story should be taken at face value. It is essential to understand what is being reported (and by whom) and what incentives may be driving that narrative.
Customers leveraging these frontier models should take a cautious approach rooted in practical outcomes. When configuring AI-based solutions, it is recommended to do so carefully with specific goals in mind. Even the most advanced AI model cannot cause a cyber incident on its own; however, misconfigurations, poor data access policies and simple mistakes cause them every day. If you wish to minimize the likelihood of an incident impacting your own network or customer base, taking a thoughtful, intentional approach to AI implementation is critical.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


