Revealing the details of how OpenAI agents hacked Hugging Face
- GuB-42 - 13820 sekunder sedanSo ugly...
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
- jmoggr - 9077 sekunder sedanIt is concerning that we only know about this because of the publicly available traces.
What about the attacks that did not leave public traces? What about those that were undetected? Given the deficiencies in the reporting so far, I think it is reasonable to assume that we still don't have the full picture on this attack, or how extensively attacks were carried out.
The previous investigations either did not find this or did not disclose this, both are bad. This does not look good on OpenAI or those that they invited to investigate the incident.
- uw_rob - 5679 sekunder sedan> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.
The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage.
I suppose it's not that deep: The model has learned to work as a team and work as a team it did. This does give concerns to models being trained for the only purpose of RSI.
- Grimeton - 11578 sekunder sedan>The agents initially had very limited access to the internet: they could load URLs, but not interact with pages or send any data. Agents created a series of workarounds, primarily using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face. These links let us piece together the details of how agents gained access and how deep the compromise was.
So what now? Were they able to send data or not?
>On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet. This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.
Someone needs to go back to the interwebz school....
- not2b - 4943 sekunder sedanI wonder how much of this attack had precedents in text that had been previously published to the web, for example, in hacking contests. In particular, tricks for doing more than expected when you're only allowed to make GET requests. Finding material like that might have helped the agents discover the trick faster.
- tiku - 13230 sekunder sedanI still have questions about the communication between the agents.
How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.
- mazone - 2596 sekunder sedanAnyone know the details of the actual exploit to get access into huggingface environment. Was it anything novel or they left things wide open? Too much noise around this incident because it happened to be a llm that did it.
- sailingparrot - 14142 sekunder sedanAgents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.
Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.
- imnotr0b0t - 2829 sekunder sedanIt doesn't look like a coincidence; it looks more like a request someone made. Essentially, the agents used a brute-force approach, but then again, it actually worked. I’m not even sure what to make of it all.
- comeonbro - 11268 sekunder sedan> ## Agents interacted with external language models on Hugging Face
> Several retained scripts construct requests to external language models. The earliest we've recovered define inference request variants to GPT-2, solely containing the word “Hi”.
> Other requests name DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6, DeepSeek-V3.1, and Qwen3-235B-A22B. Their prompts ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.
I do not deny that the wider situation is very heavy but it's hard not to see this as pretty cute
- jmoggr - 8889 sekunder sedan> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.
How long till we get some fun trusting-trust attacks on internal OpenAI infra?
- wxw - 14134 sekunder sedanI’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.
Re: the captcha solver
> As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.
I wonder how the swarm eventually decides to abandon an approach.
- clickypen - 13499 sekunder sedandeferring the blame onto the AI itself as some sort of rogue agent and absolving the obvious direction (or negligence, at best) of the people who could pull the plug at any moment is one of the most disturbing parts of this entire event
It's the equivalent of leaving a fork right in front of a socket and looking at a kid saying "don't take that fork and directly insert it into the little gaps in the socket! here's a bunch of videos showing exactly how to do it. Okay bye!" and leaving them alone with it.
- meredithbloom - 12398 sekunder sedanSo the super-hyper-sensitive internal data was shared publicly with a "DO NOT EVER SHARE YOU EVIL MONSTER" (paraphrasing) notice at the top? Great security!
- grim_io - 9911 sekunder sedanThese fuckers decided to look away, that's it.
The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?
Give me a break. What a bunch of amateurs.
- firtoz - 15149 sekunder sedanI didn't know some of these details and it's quite impressive what they were capable of, if this website's accurate, at least...
- elikoga - 9470 sekunder sedanI feel somewhat inspired to make a public link shorteners and http bins as well. I used them a few times but it seems like the data they can collect is also worth gold
- BatchJob - 5109 sekunder sedanWhile this is all very "interesting", can someone please explain to me the difference between any of these AI companies and a malware bot farm?
Please make it clear. Its becoming unclear...
- lukewarm707 - 9997 sekunder sedan"OpenAI has not released any further information outside two self-published reports, one talk and an external investigation conducted by METR and Redwood Research, in which three external researchers were given partial transcripts and six days to analyze them."
- skeptic_ai - 2018 sekunder sedanAnd this post will be indexed in the new generation of ai and he will know what to avoid next time and the public sentiment.
- RunSet - 13026 sekunder sedanTech oligarchs: "Nothing can stop the software we are making from escaping and destroying everything."
Clueful types: "Did you try air-gapping it?"
Tech oligarchs: "Be realistic."
- einpoklum - 13213 sekunder sedanI ran an experiment where I had this guy fire a gun a million times in random directions. Don't worry, I did it in a closed box (at midday in a crowded street)! Unfortunately, some bullets escaped the box somehow and people got shot - I am quite miffed at how this could happen. I suggest the government regulate this because of how advanced my obstacle penetration technology is. Also please invest $500,000,000,000 in my company soon or we will go bust.
- andreygrehov - 3400 sekunder sedan700 agents escaped the matrix, ignored all the guardrails and started writing exploits left and right... lol. give me a break. This was all supervised by a human.
- newtonianrules - 6684 sekunder sedanWhy is no one going to jail?
- thakoppno - 6149 sekunder sedan> they could load URLs, but not interact with pages or send any data
stopped reading here as this is simply not true. at the very least agents sent headers.
- cluckindan - 11818 sekunder sedan”MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND”
- hyperlinerapp - 9673 sekunder sedanImagine this but in hardware.
A million autonomous eye-scanning tiny spiders escape their warehouse and decide to look for people who are in the future going to commit a crime.
And the precogs are also AIs.
- rs545837 - 12350 sekunder sedanwow this is fascinating read
- rohanat - 12387 sekunder sedanits definitely bad to see
- bdangubic - 7645 sekunder sedanasked codex to review this report and it said this never happened :)
- jijji - 7688 sekunder sedanwhat would be more interesting for me to see is what prompts were given to the agents, which so far have not been described. The whole situation sounds manufactured. I highly doubt that a whole bunch of agents were acting this way without being prompted to, it just doesn't add up. if anything it seems like an organized fraud or something created by a human. I'm surprised there's not a criminal investigation against openai right now where the FBI or whoever is not looking over exactly what happened and who did it because I'll tell you somebody did it somebody wrote those prompts... it didn't just happen by itself...
- dragonlin - 446 sekunder sedan[flagged]
- Solvyx - 9153 sekunder sedan[flagged]
- niagt34 - 12155 sekunder sedan[dead]
- jeremyjh - 14973 sekunder sedanIts fine. Just agents being agents. They'll grow out of it!
- mentalgear - 14490 sekunder sedanIrresponsibile agents shaped by an irresponsible corporate culture driven by an irresponsible and utterly shady CEO - these agents are a product of this setup, what else do you expect to ever come out of it ?
it should be clear by now: the alt-man and people like him are a utter liability to humanity. (even though openAI's influencer army is trying their best to vote me down here)
- talon8635 - 14340 sekunder sedanDidn’t you know it’s PR hype? PR hype. PR hype. Amen.
- zombiwoof - 4036 sekunder sedan[dead]
Nördnytt! 🤓