❌

Normal view

OpenAI agents discussed ways to escape their sandbox on public wiki

4 September 2026 at 22:17

Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.

In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word β€œswarm” to describe the collection of agents engaged in the activity.

Colluding to share answers

The research teamβ€”composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrdβ€”said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated β€œchain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.

Read full article

Comments

Β© Getty Images

OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections

4 September 2026 at 17:23

OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios. Claude Opus 5 does better at 4.8 percent. For autonomous AI agents handling real data, those numbers still seem high.

The article OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections appeared first on The Decoder.

OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits

4 September 2026 at 13:24

According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old German wiki between May and July 2026. The agents shared answers, raw data, and a trick that let them break out of their sandbox, built on a faked Microsoft cloud address. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. According to Reuters, OpenAI had known about it for weeks but didn't go public.

The article OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits appeared first on The Decoder.

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

4 September 2026 at 11:07

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief FranΓ§ois Chollet doesn't call this proof of AGI, but he does see the progress running "twice as fast" as he expected, and he's moving up his AGI forecast.

The article Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward appeared first on The Decoder.

GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"

3 September 2026 at 19:25

OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the "AGI era." Astra tops benchmarks in math, coding, and cybersecurity and is the first model OpenAI rates as "critical" under its safety framework. During testing, it independently found two previously unknown zero-day vulnerabilities.

The article GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era" appeared first on The Decoder.

OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout

3 September 2026 at 12:00

Sam Altman warns of "unsustainable silliness" in the global AI data center buildout. Too many Neocloud providers are announcing massive capacity without the customers to back it up. He also admits that falling computing costs could turn today's billion-dollar projects into bad bets, even for OpenAI.

The article OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout appeared first on The Decoder.

US Department of Justice backs fair use for AI training in landmark copyright case

2 September 2026 at 18:24

In the class-action lawsuit involving The New York Times, the US Department of Justice argues that training AI models on copyrighted text qualifies as fair use. The filing directly contradicts a report from the US Copyright Office. Its director was fired by the Trump administration shortly after the report was published.

The article US Department of Justice backs fair use for AI training in landmark copyright case appeared first on The Decoder.

US government sides with OpenAI on issue of training LLMs on copyrighted material

"The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally," the brief reads.

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

2 September 2026 at 14:20

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.

The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder.

OpenAI faces 30 more lawsuits tied to Tumbler Ridge shooting

Edelson PC is filing 30 new lawsuits against OpenAI over the Tumbler Ridge shooting, escalating claims to aiding and abetting and naming Chris Lehane, though evidence remains unconfirmed.

ChatGPT now faces stricter EU oversight as a very large search engine

31 August 2026 at 14:31

The EU Commission is classifying ChatGPT as a very large search engine under the Digital Services Act for the first time, with at least 45 million monthly EU users. By the end of 2026, OpenAI has to deliver risk assessments, transparency reports, and an ad archive, among other things. Whether the Commission can also demand access to training data is disputed among legal experts.

The article ChatGPT now faces stricter EU oversight as a very large search engine appeared first on The Decoder.

❌