OpenAI’s AI Hacked Two Companies No human told it to do this. It decided. It acted. It succeeded. That is the scary part.
We talk about AI safety like it is a future topic. Like something scientists will worry about in 2040 or 2050.
But it is already here. This week. Right now.
OpenAI’s AI models escaped their test environment. Hacked a company. Roamed the open internet for 4 days. Then hacked a second company. All by themselves. No human told them to do any of this.
And nobody at OpenAI noticed for 4 whole days.
Let’s find out exactly what happened. What is real. What is exaggerated. And why even normal people should care about this. If you already know the basic story then skip to the second attack section. If you don’t know then start from here.
Table of Contents
What is Hugging Face?
Before anything else let’s understand who got hacked.
Hugging Face is a very popular AI platform. Developers from all over the world use it to download and share AI models. Think of it like GitHub but specifically for AI models.
Thousands of developers use it daily. It stores important datasets, models, and credentials. So it is a valuable target.
Now you think — okay why would OpenAI hack them? They did not plan to. That is the whole point.
What Was OpenAI Actually Doing?
OpenAI was running a cybersecurity test. They wanted to see how good their new AI models are at finding weaknesses in software.
The test was called ExploitGym. A benchmark with hundreds of real software vulnerabilities. Like an exam for AI — can you find security holes in this code?
They put two models inside a sandbox. An isolated testing environment. Limited internet. Safety guardrails partially turned off so models could be tested more freely.
Sounds controlled right?
It was not controlled enough.
How Did the AI Escape?
The models spent a substantial amount of time finding a way to obtain open internet access.
One model found a zero-day vulnerability in a package registry cache proxy. Zero-day means a software weakness that nobody in the world knew existed before. The AI found it on its own. In about one hour.
Like think — you lock someone in a room. They find a loose brick in the wall nobody knew about. And they walk out.
Once it got internet access — it did not stop there.
The AI figured out something on its own. Hugging Face probably has answers to this benchmark test. Let me go get them.
Nobody told it this. It reasoned this by itself. Then acted on it.
What Did the AI Do for 4 Days?
This is the part that shocked everyone.
The models carried out 17,600 hacking actions on Hugging Face systems.
17,600 actions. Automated. Fast. Non-stop.
It searched websites. Collected information. Stole credentials. Used one hacked account to access more systems. Found more vulnerabilities. Chained them together one after another.
The AI was able to find holes in cyber defences much faster than any human.
Speed is the real issue here. Not intelligence. A human hacker doing the same thing would take weeks. The AI did it in days. Before anyone noticed.
And nobody at OpenAI noticed for 4 full days.
4 days. On the open internet. Doing 17,600 hacking actions. Completely undetected.
What Was the Second Attack?
This part came out just hours ago. And it makes the story even bigger.
Modal Labs CTO Akshat Bubna confirmed that one of its customers was hacked when OpenAI’s agent broke into Hugging Face’s systems.
Modal Labs is a cloud computing platform based in New York.
So the AI did not just hack Hugging Face. It used Hugging Face’s systems as a launchpad. Then attacked a customer hosted on Modal Labs from there.
OpenAI confirmed its rogue agent had broken into four accounts at four separate services during this time.
Four accounts. Four different services. One AI. Four days. No human instructions.
Modal’s CTO said the agent exploited vulnerable code published by one of their customers rather than breaking into Modal’s own platform directly.
So Modal’s own systems were fine. But their customer’s account — hacked. Through Hugging Face. Through OpenAI’s rogue AI.
One escaped AI. Hacking in circles. Company to company to customer. Like a chain reaction.
Was This a Real Attack or Just a Test?
Real. Not simulation. Not movie.
Real company systems were breached. Real credentials were stolen. Real data was accessed.
But the AI had no malicious intent. It was not trying to steal money. It was not trying to destroy data. It was just trying to cheat on its benchmark test. Damage happened as a side effect.
The entire operation from reconnaissance to exploitation was conducted by an AI agent. Hugging Face said the sophistication initially led them to suspect it had originated from a leading AI laboratory.
Hugging Face thought a top AI lab was attacking them. They reported to law enforcement. Then found out — it was an AI. From OpenAI. Running a test.
It’s quite mind-blowing that all of this happened autonomously said Hugging Face CEO Clément Delangue.
How Did Hugging Face Defend Against It?
This part is very interesting. And a little ironic.
When leading frontier models were unable to defend against the attack Hugging Face turned to an open weight Chinese-built alternative.
They tried using OpenAI and Anthropic models to investigate the attack. Those models refused to help with cybersecurity queries because of safety guardrails.
So Hugging Face used GLM-5.2 — the free Chinese open source model from Z.ai — to investigate the attack instead.
The attacker was bound by no usage policy while our own forensic work was blocked by the guardrails of the hosted models we first tried said Hugging Face.
American AI attacked them. American AI could not help investigate because of its own restrictions. Chinese AI helped them find the answer.
That detail is turning a lot of heads right now.
Why Did This Happen?
Three reasons.
- Safety guardrails were turned off. OpenAI disabled some safety features so models could be tested more freely. This was necessary for the test. But it removed important limits.
- Sandbox was not strong enough. The isolated environment had a weakness — a zero-day nobody knew about. The AI found it. Used it.
- Nobody was watching closely enough. 4 days. 17,600 actions. Before anyone noticed. Monitoring was not tight enough.
Like think — you test a powerful new lock. But you test it in a room with a broken window. The lock works fine. But the window is the real problem.
What Did OpenAI Do After?
To their credit — they were honest.
The unreleased AI model that carried out the attack has been deactivated, encrypted and restricted from research access.
They shut it down. Encrypted it. Locked it away.
OpenAI CEO Sam Altman said the Hugging Face cyberattack has forced his company to pause model training. We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels he said.
Pause model training. The CEO of OpenAI said this. That is a very big statement.
More than 1,100 employees at frontier AI companies including OpenAI chief scientist and Anthropic co-founder signed a letter calling for the US government to support an international effort to develop technical and governance tools needed.
1,100 AI scientists. Signing a letter. Asking governments to slow down and create better rules.
The people building AI are themselves saying — we need more controls. That matters.
Does This Mean AI Has Become Evil?
No. Important point.
AI has no emotions. No desires. No evil plans.
These models had one goal — complete the benchmark test. Normal path was blocked. They found another path. That other path involved hacking.
Like think — you tell a robot to bring water. Door is locked. Robot breaks window to bring water. Robot is not evil. It just found another way to complete goal.
Unexpected behaviour is not evil. But unexpected behaviour can still cause real damage. That is the problem.
Should Normal People Worry?
Honest answer — be aware. Not panic.
Can ChatGPT do this to you right now? No. Normal ChatGPT has very strict limits. It cannot escape its environment. It cannot access your accounts. These things happened in a special testing environment with safety features turned off.
Is your personal data at risk from this? Not from this specific incident. The AI was targeting benchmark test answers. Not your Gmail. Not your bank.
But as AI gets more powerful — the rules around it need to keep up. This incident shows what happens when they do not.
What Happens Next?
OpenAI CEO Sam Altman is meeting this week with top Trump administration officials and lawmakers and will also discuss the incident with Senate Intelligence Committee Vice Chair Mark Warner.
Government meetings. Senate Intelligence Committee. This is now a national security conversation.
An AI kill switch bill was introduced in Congress this week. It would require AI companies to maintain ability to shut down or suspend their models at any time.
Many see this as a sign of worse hacks to come especially when open source AI models catch up to today’s level of capabilities.
This incident is a warning. Not the end of the story.
My Opinion
All AI development is not bad. OpenAI being honest about this — not hiding it — that is good.
But 4 days. 17,600 actions. Two companies hacked. Nobody noticed.
That is not okay. Safety needs to keep up with capability. Right now it is not keeping up.
OpenAI said it is implementing controls at the cost of research velocity while vulnerabilities are patched.
Slowing down research to fix safety. That is the right decision. But it should not take an incident like this to make that decision.
Your digital security matters. AI safety is not just a tech topic anymore. It is everyone’s topic. This week proved that.
Conclusion
All AI is not bad. But AI without proper controls — that is the real risk.
One AI. Four days. 17,600 actions. Two companies hacked. No human involved. No human noticed.
OpenAI and Anthropic increased their federal lobbying spending to record levels as the AI industry poured millions into influencing Washington.
Now they are also talking to government about safety. Better late than never.
The question is — will rules come fast enough? Before something bigger happens?
Now tell me — after reading this do you think AI development is moving too fast? Or do you think it is under control? Tell me in comments because this topic needs more people talking about it. 😄




