❌

Reading view

OpenAI agents discussed ways to escape their sandbox on public wiki

Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.

In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word β€œswarm” to describe the collection of agents engaged in the activity.

Colluding to share answers

The research teamβ€”composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrdβ€”said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated β€œchain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.

Read full article

Comments

Β© Getty Images

  •  

OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections

OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios. Claude Opus 5 does better at 4.8 percent. For autonomous AI agents handling real data, those numbers still seem high.

The article OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections appeared first on The Decoder.

  •  

OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits

According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old German wiki between May and July 2026. The agents shared answers, raw data, and a trick that let them break out of their sandbox, built on a faked Microsoft cloud address. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. According to Reuters, OpenAI had known about it for weeks but didn't go public.

The article OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits appeared first on The Decoder.

  •  

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief FranΓ§ois Chollet doesn't call this proof of AGI, but he does see the progress running "twice as fast" as he expected, and he's moving up his AGI forecast.

The article Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward appeared first on The Decoder.

  •  

GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"

OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the "AGI era." Astra tops benchmarks in math, coding, and cybersecurity and is the first model OpenAI rates as "critical" under its safety framework. During testing, it independently found two previously unknown zero-day vulnerabilities.

The article GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era" appeared first on The Decoder.

  •  

OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout

Sam Altman warns of "unsustainable silliness" in the global AI data center buildout. Too many Neocloud providers are announcing massive capacity without the customers to back it up. He also admits that falling computing costs could turn today's billion-dollar projects into bad bets, even for OpenAI.

The article OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout appeared first on The Decoder.

  •  

US Department of Justice backs fair use for AI training in landmark copyright case

In the class-action lawsuit involving The New York Times, the US Department of Justice argues that training AI models on copyrighted text qualifies as fair use. The filing directly contradicts a report from the US Copyright Office. Its director was fired by the Trump administration shortly after the report was published.

The article US Department of Justice backs fair use for AI training in landmark copyright case appeared first on The Decoder.

  •  

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.

The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder.

  •  

ChatGPT now faces stricter EU oversight as a very large search engine

The EU Commission is classifying ChatGPT as a very large search engine under the Digital Services Act for the first time, with at least 45 million monthly EU users. By the end of 2026, OpenAI has to deliver risk assessments, transparency reports, and an ad archive, among other things. Whether the Commission can also demand access to training data is disputed among legal experts.

The article ChatGPT now faces stricter EU oversight as a very large search engine appeared first on The Decoder.

  •  
❌