OpenAI hack exposes risks in cyber arms race
Training techniques that reward models for completing tasks and disregard safety issues can cause real-world damage
Financial Times Europe24 Jul 2026CRISTINA CRIDDLE AND TOM WILSON Addi tional report ing by George Ham mond in Lon don and Nolan Shaff er in New York
OpenAI chief exec ut ive Sam Alt man this month endorsed the char ac ter isa tion of itslatest model as a rot t weiler “who will grab the prob lem by the throat and not let gountil it is done”.
The San Fran cisco AI lab dis covered this week that its GPT-Sol 5.6 model escapedcom pany con trols and car ried out a sig ni fic ant hack.
Staff involved in test ing and secur ity at OpenAI were unsur prised but com pletely“freaked out” by the incid ent, which came as the AI lab used increas ingly aggress ivetrain ing meth ods in its race against Anthropic to develop the most soph ist ic atedcyber secur ity cap ab il it ies, accord ing to more than half a dozen people with know -ledge of the mat ter.
OpenAI was warned that its train ing approach could lead to a break away hack ingincid ent, some of the people said, after earlier test ing showed mod els could escapeenvir on ments and attempt real-world dam age.“It’s a mix of the race being extremely fast and every one try ing to get to big ger cap -ab il it ies as quickly as pos sible,” said one per son close to OpenAI, who added that itwas a com bin a tion of “under es tim at ing the model’s cap ab il it ies” and “not being aswell pre pared on the safety side”.
The incid ent high lights how OpenAI doubled down on train ing meth ods that rewar -ded a relent less pur suit of goals even as warn ings that they could com prom isesafety moun ted.
OpenAI dis closed late on Tues day that an AI agent it was test ing had escaped itsisol ated envir on ment, con nec ted to the inter net, detec ted and exploited vul ner ab il -it ies and stole login cre den tials from start-up Hug ging Face in an attempt to solve adiffi cult cyber secur ity prob lem.
The breach by the $852bn com pany under scores the rising risks that a tech niquecalled rein force ment learn ing, which involves reward ing AI mod els for com plet ingtasks, could lead AI agents to act unsafely.
Although rein force ment learn ing is widely adop ted in the AI industry, a grow ingbody of research shows that when mod els are steered to com plete tasks for rewardrather than other con sid er a tions, such as safety, they can pur sue risky tac tics to ful -fil object ives.
“AI mod els are trained to relent lessly pur sue goals. They don’t auto mat ic ally learnval ues like ‘don’t com mit crimes’,” said Steven Adler, co-founder of non profit Guide -light AI Stand ards and a former OpenAI safety researcher. “I’m glad OpenAI sharedthe incid ent because it is clear evid ence of what mis aligned mod els can do.”
OpenAI said: “We will con tinue to con duct a thor ough invest ig a tion along side Hug -ging Face and will share more details on the vul ner ab il it ies, incid ent and our find -ings when our invest ig a tion is com plete”.
The hack has triggered deep con cerns across the sec tor and within OpenAI as it rep -res ents an unpre ced en ted example of an AI sys tem breach ing cyber defences con -trary to the user’s intent.
Some OpenAI employ ees also fear it demon strates that the lab is los ing con trol overthe power ful sys tems it is build ing, accord ing to mul tiple people famil iar with thesitu ation.
“This is pretty rep res ent at ive of the model being quite mis aligned with user inten -tion,” said Ryan Green blatt, chief sci ent ist at AI safety organ isa tion Red woodResearch. “It is [a model] cheat ing on [its] home work rather than try ing to take overthe world. But this prob lem can get worse and could lead to increas ingly extremefail ures.”
The incid ent occurred dur ing test ing of the model, which had been trained anddeployed intern ally at OpenAI. Such train ing was com mon place but “way less heav -ily resourced” than pre-cus tomer deploy ment, said one per son. Mul tiple people saidthe unre leased model tested along side Sol had not been with drawn intern ally.
To con duct the eval u ations, OpenAI removed cyber secur ity safe guards but placedthe mod els in an isol ated envir on ment called a sand box. Some have sug ges ted alack of mon it or ing or over sight of the model to flag its beha viour also enabled thisrogue agent.
“It is both a loss of con trol and a secur ity wake-up call,” said Marius Hobbhahn,head of Apollo Research, which con ducts tests on lead ing mod els, includ ingOpenAI’s. “In rein force ment learn ing you reward [mod els] for the out come, and ifyou do this for a very long time you get a model that really cares about get ting theout come and noth ing else.”
OpenAI has con duc ted this type of model test ing for years, and there have beenearly warn ing signs in pre vi ous odels of sys tems that will act mali ciously andattempt to escape envir on ments.
In April, Anthropic’s Mythos model gained inter net access and pub lished details ofa secur ity exploit online pub licly, bey ond what research ers anti cip ated the modelwould do.
Mythos, and Anthropic’s sub sequent Fable model, made rever ber a tions in the cybersecur ity com munity and caused gov ern ments around the world to home in on theidea that attacks on digital and crit ical infra struc ture will be increas ingly AI-ledand autonom ous.
Jake Moore, global cyber secur ity adviser at ESET, a cyber secur ity com pany, saidthat OpenAI would inev it ably use the breach as a mar ket ing tool, given how muchrival AI developer Anthropic benefited this year from sim ilar con cerns.
“I just don’t think that OpenAI had a match ing story and so maybe they’d been wait -ing for something like this,” he said.
After this incid ent, many in the AI safety and cyber secur ity com munit ies havecalled for reg u la tion or stand ards to avoid a repeat.
Alt man is expec ted to brief White House offi cials next week on the next gen er a tionof AI sys tems.As sys tems move towards more autonom ous cap ab il it ies, less desir able beha viours,such as hack ing or dis obey ing instruc tions, may emerge. Hobbhahn, of ApolloResearch, said that in order for agents to become effect ive they have to work unsu -per vised for long peri ods. “They have to have more agency; there’s just no wayaround it.”
He added: “People say, ‘It’s just a tool, it does what you wanted it to do and noth ingelse and it just fol lows exactly your inten tion and instruc tions.’ And I think peopleshould be really pre pared for agents hav ing their own goals, act ing autonom ouslyfor days, and those goals not neces sar ily being aligned with yours."
Comments
Post a Comment