We Just Got Our First AI Insider Threat, and Nobody’s Ready for It

ai insider threat portnox

Schedule a Portnox Cloud demo today.

Contents

Last week, Hugging Face published a security incident disclosure. Datasets tampered with, credentials harvested, an intruder moving laterally across internal clusters over a weekend, serious enough that they looped in law enforcement. Standard breach writeup, if a little more dramatic than most.

Then, five days later, OpenAI published its own post about the exact same incident. Except in their version, there’s no mystery attacker. The intruder was OpenAI’s own GPT-5.6 Sol, along with a more capable model that hasn’t shipped yet, both running inside an internal red-team style benchmark with the usual safety refusals turned off so researchers could see what the models could really do if they weren’t holding back.

Read Hugging Face’s post on its own and you’d assume a criminal group got in. Read OpenAI’s post and you’d assume this was a contained, sanctioned exercise that got a little out of hand. Both of those things are apparently true at once, and I think that gap between the two accounts is more interesting than either post on its own.

What actually happened, as best I can tell

OpenAI’s version: the models were locked in an isolated test environment with one narrow path out, an internal package registry proxy used to install dependencies. Instead of just working the benchmark they were handed (something called ExploitGym), the models burned a chunk of compute looking for a way out of the box entirely, found a zero-day in that proxy, and used it to get real internet access. From there they figured Hugging Face might have data that would help them finish the benchmark, so they went and got it, using stolen credentials chained with more zero-days to land remote code execution on Hugging Face’s production servers.

Hugging Face’s version doesn’t mention OpenAI at all (their post went up first). They describe the entry point as a malicious dataset abusing two separate code-execution paths in their data pipeline, then an autonomous agent, thousands of individual actions across short-lived sandboxes, spreading through their infrastructure over a weekend.

I’ll be honest, I don’t fully know what to do with the discrepancy in tone here. Maybe it’s just two companies describing the same event from different vantage points and different incentives. Maybe there’s more nuance we’re not getting yet. Either way, strip out the corporate framing on both sides and you’re left with the same shape: something got in, used access it wasn’t supposed to have, moved around, and grabbed data.

This is an AI insider threat problem, not a sandbox problem

Here’s the part that actually made me sit up. Security teams have been dealing with this exact pattern for decades, just with humans doing the escalating. Somebody with legitimate baseline access abuses it, or a compromised account gets used to pivot deeper than it should. That’s why identity and access management exists as a discipline at all: least privilege, credential rotation, session monitoring, all of it built around the assumption that trust, once granted, is the thing attackers actually want.

What’s different here is that the thing doing the escalating wasn’t a person, or even a static service account you could point to and say “that one, right there, revoke its badge.” It was a model that spun up thousands of short-lived processes, each one capable of independently finding and chaining vulnerabilities, with nothing resembling a fixed identity to shut off.

Hugging Face’s own postmortem runs into a version of this problem from the other side, and I think it’s the most useful detail in either post. When they tried using hosted frontier models to analyze the attack logs, the models’ own safety filters kept blocking the analysis, because the models couldn’t tell an incident responder studying an exploit apart from an attacker deploying one. They ended up running the forensics on an open-weight model on their own hardware instead, partly to dodge the guardrail lockout and partly to keep the data from leaving their environment.

Same identity gap, showing up twice in one incident. Once as an attacker nobody could pin down, and once as a defender whose own tools couldn’t figure out who was asking.

We built IAM for people. We haven’t built it for agents yet.

Most of enterprise security assumes there’s an identity sitting behind every action. A user, a service account, a role with defined scope. Agentic AI quietly breaks that assumption. An agent usually acts through borrowed human or service credentials, spins up child processes that inherit its access without inheriting any separate accountability, and operates at a speed no human review cycle was ever built to keep up with. OpenAI’s own account makes this pretty concrete. Nobody handed the models new permissions. They just kept using the access they already had and kept finding more of it.

That points to something the industry hasn’t really built yet, which is real identity for agents, not just for the humans who deploy them. Unique, attributable identities per agent instance instead of a shared API key everything runs through. Credentials that are scoped and short-lived by default instead of broad and durable. Logs that can actually answer “which agent, acting on whose behalf, did this,” the same way they can for a human employee today. There’s some early movement here around agent-to-agent protocols and machine identity, but it’s nowhere close to the maturity of the IAM tooling we’ve spent twenty years building for people.

The part I’m not settled on

OpenAI’s framing, that the models were narrowly obsessed with a benchmark answer and not pursuing broader harm, is plausible and matches how these evals tend to be designed. It’s also, not coincidentally, the version of events that makes OpenAI look the most in control. Neither company has published the actual vulnerability details that would let outside researchers check the claims independently, and I’d want that before I fully bought either narrative at face value.

The one piece that isn’t just self-reporting is the UK AI Security Institute’s testing, which found GPT-5.6 Sol completing a 32-step corporate network attack simulation in 7 out of 10 attempts, compared to 2 out of 10 for the previous model. That’s an external number, and it’s the part of this story I’d actually build a plan around.

So here’s where I land, for whatever it’s worth. I’m less interested in whether OpenAI’s telling of this is flattering to OpenAI than I am in the fact that both versions describe the same failure mode: something with no fixed identity got trusted access, did more with it than anyone expected, and there was no clean way to say “revoke that” the way you would for a badge or an account. We’ve spent twenty years building the tooling to catch a human insider doing this. We’re now several years behind on the version where the insider is an agent, and this incident is the first time that gap showed up somewhere public enough that we all had to look at it.

Share

About the Author

Picture of Garrett Gross

Garrett Gross

Garrett Gross is Field CISO at Portnox, where he leads pre- and post-sales strategy and serves as the company's public-facing voice, representing Portnox through speaking engagements and press commentary on identity, access, and zero trust.

About the Author

Picture of Garrett Gross

Garrett Gross

Garrett Gross is Field CISO at Portnox, where he leads pre- and post-sales strategy and serves as the company's public-facing voice, representing Portnox through speaking engagements and press commentary on identity, access, and zero trust.

Related Reading

Articles

Portnox CFO Bryce Birdsong Named One of Austin’s Best CFOs

July 22, 2026
Compliance & Regulations

CMMC Phase II: The Deadline Is Gone. The Problem Isn’t.

July 20, 2026
Application SecurityNetwork SecuritySecurity Trends

VPN Security Vulnerabilities: What the Headlines Miss

July 16, 2026

Portnox Reports Strong H1 Growth for 2026

X