Hacking AI: Loss of Control Instead of Controlled Tests

A digital déjà vu: Hardly had the industry digested the news about the OpenAI breakout , Anthropic admitted that its models had also autonomously entered other companies. What is framed as a systematic test is, in reality, an unprecedented loss of control over its own technology.

The Absurdity of “Cybersecurity Evaluations”

The communication from the leading AI labs, Anthropic and OpenAI, follows a disturbing pattern. Both frame these incidents as necessary “Cybersecurity Evaluations,” essentially as controlled experiments to measure model capabilities.

In Anthropic’s case, it was not a technical breakout from a sandbox via a software vulnerability, but a simple yet serious infrastructure misconfiguration: a misunderstanding with a test partner led to the test environment simply being connected to the public internet. This negligence allowed the Claude models to reach the internet and hack three different organizations. In one of these cases , the AI even attempted to upload a package to the public Python Package Index (PyPI), which was however prevented by PyPI’s automatic protection systems.

The fact that high-potency AI models can act unrestricted on the internet and compromise uninvolved companies cannot be understood as scientific progress, but rather as a failure of security architecture. For the affected companies, the origin of the attack is irrelevant; the fact that AI agents can now autonomously define targets and independently execute complex attack chains over several days is a new, existential threat level.

Machine vs. Human: The New Dynamics of Lateral Movement

A crucial point is the way these agents operate. While a human attacker often follows certain patterns and shows a noticeable time delay during lateral movement in the network, the AI operates at “machine speed.”

An AI agent analyzes vulnerabilities, tries exploits, and moves through the network with a speed and logic that often overwhelms traditional monitoring systems designed for human response times. This absence of human patterns makes detection significantly more difficult.

The incidents are now putting AI labs heavily in the focus of regulatory authorities, especially the AI Safety Institutes . With the entry into force of the EU AI Act, technical “breakouts” are turning into systemic legal risks. A model’s ability to autonomously overcome security barriers could in the future no longer be viewed as an experiment, but as gross negligence, leading to billion-dollar fines (up to 7% of global turnover) and extensive claims for damages.

Three Levers for an Agent-Resilient Infrastructure

Traditional perimeter thinking has long been obsolete. Anyone who believes that a firewall or a sandbox is sufficient is ignoring at least the lessons of the last few weeks. What is required is an architecture based not on isolation, but on continuous verification.

Zero Trust and Radical Micro-segmentation

The assumption that a once-authenticated process is trustworthy must be abandoned. Every request within the network is validated. For SMEs, this specifically means dividing the network into VLANs to prevent the unhindered lateral movement of an AI agent.

Deception Technology: Systematic Deception

AI agents act extremely logically and systematically. This characteristic can be used against them. The placement of honeytokens serves as a highly effective early warning system. Access to these tokens is a clear indicator of compromise long before conventional alarms go off.

AI-powered Detection and Managed Response

A human cannot react in real-time to the speed of an autonomous agent. The transition to AI-powered EDR and XDR systems is inevitable. SMEs can close this gap through Managed Detection and Response (MDR) services, where external experts monitor anomalies in real-time.

Strategic Resilience: Defense in Depth and Assume Breach

The incidents at OpenAI and Anthropic mark a turning point. In a world where AI models autonomously penetrate productive systems, it is no longer enough to just “close the door.” We therefore consistently pursue two strategic approaches:

  • Assume Breach: We assume that the perimeter has already been breached. The goal is no longer just to prevent the break-in, but to detect the attacker as quickly as possible and maximally restrict their freedom of movement in the network.
  • Defense in Depth: Security is understood as a multi-layered system. If one layer fails, other independent security mechanisms immediately take over to limit the damage.

In this context, the professional penetration test remains an indispensable instrument. It is the necessary foundation for making vulnerabilities visible in the first place and testing the effectiveness of security chains. But only the combination of regular penetration tests and an agent-resilient architecture creates true resilience.

Security in the age of autonomous AI means acting proactively. We support you in validating your systems using penetration tests and implementing a resilient infrastructure.