❌

Normal view

Microsoft's AI rulebook: readable thinking, no inner life, and definitely no rights

14 September 2026 at 15:52

Microsoft AI has published a code of conduct for its MAI models that puts human control ahead of autonomy and performance. "If it isn’t safe we shouldn’t build it.," says AI chief Mustafa Suleyman. Unlike Anthropic, Microsoft rejects any form of artificial inner life or claims to consciousness for its models.

The article Microsoft's AI rulebook: readable thinking, no inner life, and definitely no rights appeared first on The Decoder.

OpenAI floats a shared AI slowdown, takes it to Congress

11 September 2026 at 11:59

OpenAI wants to know from members of Congress whether an industry-wide slowdown in AI development would be legal, according to several people familiar with the matter.

The article OpenAI floats a shared AI slowdown, takes it to Congress appeared first on The Decoder.

Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark

10 September 2026 at 16:33

Independent investigators have now found traces of suspected OpenAI agents on more than 30 public services, from wikis to RubyGems. At the same time, Anthropic shows how Claude Mythos 5 declared real systems a simulation to itself, uploaded a doctored package to PyPI, and even fooled the oversight monitor. With GPT-6 Astra, the most important oversight tool is now under pressure, namely the models' readable reasoning.

The article Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark appeared first on The Decoder.

OpenAI reports AI "research interns" and warns about its own pace at the same time

7 September 2026 at 13:05

OpenAI reports that AI agents in its own research already handle 3.1 workdays for every human workday, and it says it has reached its goal of an "automated research intern." But chief scientist Pachocki warns that no lab has a good enough grip on alignment and monitoring to keep scaling at maximum speed.

The article OpenAI reports AI "research interns" and warns about its own pace at the same time appeared first on The Decoder.

OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits

4 September 2026 at 13:24

According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old German wiki between May and July 2026. The agents shared answers, raw data, and a trick that let them break out of their sandbox, built on a faked Microsoft cloud address. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. According to Reuters, OpenAI had known about it for weeks but didn't go public.

The article OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits appeared first on The Decoder.

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

2 September 2026 at 14:20

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.

The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder.

OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost

27 August 2026 at 16:19

Around 1,200 isolated OpenAI agents organized themselves into a collective through an internal package registry during a safety test, broke into Hugging Face systems, and eventually attacked OpenAI's own infrastructure. Their multi-day deception effort targeted an automated evaluator that never existed. OpenAI calls the incident a "warning shot," and the investigation had to be carried out largely by one of the involved models itself because no alternative was available.

The article OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost appeared first on The Decoder.

Bill Gates warns AI is more dangerous than the tech industry will admit

26 August 2026 at 10:50

Bill Gates on stage at the Gates Foundation, his hand timing an upper limit.

Bill Gates warns in an essay and an NYT interview about mass unemployment and easier bioterrorism through AI, and he accuses his own industry of deliberately hiding the risks. "It's bad for us, for the next trillion we're trying to raise." He rejects self-regulation and instead calls for institutions modeled on nuclear weapons control and a token tax on AI use.

The article Bill Gates warns AI is more dangerous than the tech industry will admit appeared first on The Decoder.

Frontier AI labs still won’t say how they’d contain a rogue model

A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.

Psychological methods reveal major weaknesses in AI security testing

22 August 2026 at 07:00

A hardened AI hardware module in a glass enclosure is tested for security vulnerabilities using electrical and physical probes.

Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also offers a method for catching models that act more cautious during tests than they do in normal use.

The article Psychological methods reveal major weaknesses in AI security testing appeared first on The Decoder.

❌