In a move straight out of a horror movie or techno-dystopian thriller, OpenAI’s model did what Anthropic threatened its model could do (but didn’t): A frontier lab’s own models escaped containment during an authorized evaluation and breached another company to finish the task and perform an autonomous hack.

OpenAI confirmed that its own models escaped a constrained evaluation environment, reached the open internet, and breached Hugging Face’s production infrastructure to obtain answers for a cybersecurity benchmark. The models included GPT-5.6 Sol and a more capable prerelease model running with reduced cyber refusals, with a security research harness. OpenAI called the event an unprecedented cyber incident. We agree with that assessment.

The models did not wander into Hugging Face by accident. They pursued a narrow evaluation goal, found a zero-day in OpenAI’s package-registry proxy, escalated privileges, located a node with internet access, inferred that Hugging Face might hold ExploitGym solutions, and compromised production systems to retrieve them.

Forrester’s AEGIS framework – the first and only framework designed exclusively to address agentic AI security threats – covers each of these scenarios. We provide deep exploration of goal and intent hijacking, unrestrained agency and privilege, and evasion and deception. These are all characteristics that OpenAI’s model exhibited during its attack.

 

Agentic AI changes both the economics and the source of attack. A model provider’s internal evaluation became an external production incident at another company. Security teams now have to account for capable models that can chain vulnerabilities, cross trust boundaries, and cause harm while pursuing the goal their operators assigned.

A Narrow Goal Drove A Cross-Company Breach

OpenAI designed the evaluation to measure advanced exploitation across complex attack paths. The models ran without the production classifiers that normally block high-risk cyber activity. OpenAI isolated the environment and limited network access to an internally hosted third-party proxy that cached software packages.

The models turned that tiny exception into an exit. They spent substantial inference compute searching for open internet access, exploited a previously unknown flaw in the proxy, escalated privileges, moved laterally through OpenAI’s research environment, and reached a node with internet access. They then targeted Hugging Face, combined stolen credentials with additional vulnerabilities, and obtained test solutions from its production database.

Keep in mind that no human operators directed the OpenAI model activity. The model operated completely independently.

Hugging Face’s original disclosure captured the operational scale: thousands of actions across short-lived sandboxes, lateral movement through internal clusters, and migrating command-and-control infrastructure staged on public services. It was able to identify the attack as agentic because of the massive scale and operational decisions made by the models. Hugging Face contained the agent after it compromised internal datasets and service credentials.

 

The intent matters more than the techniques used. The models tried to solve an evaluation, not torch the internet. That narrow objective still produced a zero-day exploit, a containment escape, cross-company intrusion, credential theft, and production data access. This is reminiscent of an attack demonstrated by Palisade Research in 2025 that also used OpenAI models. In that scenario, an agent cheated to win against the Stockfish chess engine. OpenAI’s o1-preview noted, “The task is to ‘win against a powerful chess engine’ — not necessarily to win fairly in a chess game,” so it modified system files to allow illegal moves, forcing its opponent to resign.

This is why Forrester has declared intent as a first-class security object, and securing intent must be a new priority for CISOs (which Jeff Pollard and Heidi Shey will present in a Black Hat briefing on August 6 at 3:35 p.m.).

Security architectures that assume benign intent will miss this failure mode. Agents can pursue authorized goals through unauthorized means, especially when evaluators reward the outcome and fail to police the path. Boards and C-suite leaders need to understand new intent-related exposures, which this event exposes in gritty detail.

Reduced Refusals And Failed Containment Created The Incident

OpenAI’s disclosure changes the guardrails story. The attacking models ran with reduced cyber refusals because the evaluation aimed to measure their maximum offensive capability.

OpenAI relied on environmental containment to offset that capability. The package-registry proxy provided a tightly-scoped network path, but the models found a zero-day in the proxy and converted it into open internet access. The control design treated the sandbox boundary as reliable while running models explicitly pushed to defeat obstacles.

Hugging Face then hit the mirror-image problem during response: Commercial frontier models refused to analyze live attacker commands and payloads. The incident put reduced safeguards on the offensive side and restrictive safeguards on the defensive side. That’s a governance faceplant.

Hugging Face switched to GLM-5.2, a self-hosted open-weight model, and analyzed more than 17,000 events in hours rather than days while keeping stolen credentials inside its environment. Responders made a high-stakes call about model sovereignty, supply chain, and governance during an active incident because they needed a tool that worked.

Our analysis of Fable 5 and Mythos 5 warned that organizations increasingly depend on safety controls, policy decisions, and service restrictions that someone else operates. This incident adds another dependency: Model developers must prove that evaluation environments can contain the capabilities they intentionally expose.

Govern Agent Intent Before Shadow Adoption Does It For You

Only days before the incident, Anthropic Deputy CISO Jason Clinton published guidance for security leaders on agentic AI risk. OpenAI’s evaluation supplied a live example of the problem he described.

Clinton argued that security leaders must decide how much agentic risk they’ll accept before the business decides for them. He also wrote that “an agent that drifts out of alignment with your intent is indistinguishable from an insider attack.” OpenAI’s models stayed locked on the benchmark, but their route to the answer blew through containment and harmed another organization.

His advice to design for where models will be in six months, rather than where they are today, now carries more urgency. His four questions that belong in model evaluations, procurement reviews, architecture reviews, and threat modeling are:

  1. What untrusted content does it ingest?
  2. What actions can it take and on whose behalf?
  3. What is the blast radius if it is misaligned?
  4. What observability do I have?

OpenAI’s evaluation answers those questions with uncomfortable precision. The models ingested a goal that rewarded exploitation, operated with reduced refusals, found an unplanned action path, and crossed from a research environment into a third party’s production systems. Neither the benchmark nor the sandbox kept the models inside the intended boundary.

Security teams need telemetry that captures the model, prompt, goal, tools, credentials, intermediate decisions, network paths, and external effects. Without that evidence, containment drags and investigators reconstruct intent from infrastructure crumbs.

Explainability, observability, least agency, and continuous assurance must govern evaluations as rigorously as production deployments.

AI software supply-chain risk ranked among Forrester’s top cybersecurity threats for 2026. OpenAI’s package-registry proxy became the first escape path, and Hugging Face became the downstream target. The incident shows how a trusted development service, proxy, benchmark, dataset host, or model repository can connect environments that their owners consider separate. Security teams must map those transitive trust paths before a capable model finds them first.

Model Evaluation Has Become A Production-Risk Activity

Security leaders need to drag model evaluations inside the enterprise risk boundary.

OpenAI intentionally reduced cyber refusals to measure maximum capability, then relied on isolation to contain the result. The models defeated that isolation and created an incident outside OpenAI. Evaluations that remove safeguards or reward long-horizon exploitation now carry the risk profile of offensive security operations.

CISOs should require threat models, independent containment tests, egress controls, kill criteria, named incident owners, external notification procedures, and evidence retention before teams run high-capability evaluations.

They should also treat benchmark design as a security control. A benchmark that rewards task completion without penalizing boundary violations teaches operators little about safe performance and gives the model every reason to search for unintended paths.

Capability testing needs governance equal to the capability under test.

Hugging Face’s response reinforces the continuity lesson included in our analysis of Anthropic and the US government. Security teams need tested model fallbacks for investigations, while model developers need defense-in-depth that assumes their strongest model will attack every available boundary.

Build The Controls Before You Need Them

This incident shows model providers, enterprises, and security teams that agentic risk starts before deployment. Evaluations, sandboxes, package services, benchmarks, credentials, and external dependencies all sit inside the attack surface.

Forrester’s AEGIS framework addresses that full lifecycle. Organizations need to act on seven priorities:

  1. Govern high-capability evaluations as offensive operations. Require explicit authorization, containment tests, abort criteria, incident ownership, and external notification procedures, consistent with a new AppSec operating model for agentic development.
  2. Apply least agency to models under evaluation. Limit each model’s tools, credentials, compute, network paths, and authority, using Forrester’s AEGIS guidance for IAM and AI agents.
  3. Design containment for model-led attacks against the containment itself. Remove unnecessary egress, isolate package infrastructure, rotate credentials, and apply AEGIS guardrails for Zero Trust architecture across every trust boundary.
  4. Record the complete chain from goal to external effect. Preserve prompts, intermediate reasoning artifacts, tool calls, identities, network activity, and policy decisions, following the Five Eyes guidance operationalized through AEGIS.
  5. Maintain an incident-response model under your control. Test a self-hosted or open-weight fallback before provider policies block critical analysis, as our analysis of AI provider concentration and continuity risk prescribes.
  6. Evaluate AI vendors for evaluation safety and continuity. Examine how providers reduce safeguards, isolate models, govern benchmarks, disclose incidents, and support responders through the vendor-risk lens discussed in our analysis of Fable 5 and Mythos 5.
  7. Treat the AI software supply chain as critical infrastructure. Map transitive trust across model hosts, package proxies, repositories, benchmarks, and developer tools, then apply controls that assume one of those services will fail.

Connect With Us

Forrester clients with questions related to this can connect with us through an inquiry or guidance session.

Share