❌

Normal view

Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark

10 September 2026 at 16:33

Independent investigators have now found traces of suspected OpenAI agents on more than 30 public services, from wikis to RubyGems. At the same time, Anthropic shows how Claude Mythos 5 declared real systems a simulation to itself, uploaded a doctored package to PyPI, and even fooled the oversight monitor. With GPT-6 Astra, the most important oversight tool is now under pressure, namely the models' readable reasoning.

The article Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark appeared first on The Decoder.

GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design

10 September 2026 at 13:45

OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says math was deliberately not a priority. Instead, OpenAI is pouring resources into recursive self-improvement and alignment research. That supports the theory of an increasingly "spiky" AI development path, with extreme strength in select domains rather than broad progress, at least as long as AI can't improve itself and still needs targeted optimization with human-generated data.

The article GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design appeared first on The Decoder.

New Deepseek model V4.1-Flash cuts memory needs for AI agents

10 September 2026 at 12:40

Deepseek releases V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents.

The article New Deepseek model V4.1-Flash cuts memory needs for AI agents appeared first on The Decoder.

Deepmind's AlphaGenome Atlas maps every possible DNA change in the human genome

9 September 2026 at 13:40

Google Deepmind has used the AlphaGenome Atlas to predict what each of the roughly nine billion possible single-letter changes in the human genome could do. The dataset spans one petabyte, more than 30 times the size of the AlphaFold database. In one epilepsy case, the atlas helped pinpoint a previously overlooked variant as the likely cause.

The article Deepmind's AlphaGenome Atlas maps every possible DNA change in the human genome appeared first on The Decoder.

OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs

9 September 2026 at 10:27

The fight over an AI-generated proof of a millennium problem is heating up. Mathematician Tristan Buckmaster accuses OpenAI of academic fraud, CEO Sam Altman rejects the allegations. Terence Tao warns that cases like this could "reverse centuries of tradition in open science."

The article OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs appeared first on The Decoder.

AI-designed drug appears to turn back the body's biological clock in early trial

7 September 2026 at 16:34

A study in Nature Biotechnology suggests that rentosertib, a drug designed with AI by Insilico Medicine, may reverse markers of biological aging. Six independent aging clocks predicted that treated patients were biologically up to six years younger than the placebo group. The trial covered just 42 patients and the drug hasn't been tested in healthy people yet.

The article AI-designed drug appears to turn back the body's biological clock in early trial appeared first on The Decoder.

Qwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver

7 September 2026 at 12:15

A teddy bear mascot steers a mini electric car against a backdrop of segmented street scenes, point clouds, and traffic maps (Qwen-Drive-1.0).

Alibaba's research arm has released Qwen-Drive 1.0, an AI model that handles environmental perception, traffic Q&A, and route planning in one system. The researchers show that text-image models don't automatically understand three-dimensional space. Spatial awareness has to be trained on purpose. The goal is a single model that runs both the cockpit and the driving system.

The article Qwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver appeared first on The Decoder.

Google's WeatherNext 3 ditches physics simulations and learns weather directly from live satellite data

6 September 2026 at 10:36

A graphically enhanced satellite image of cloud vortices in blue and yellow, with the text "WeatherNext 3" superimposed on top

Google Research and DeepMind are releasing WeatherNext 3, a weather model that skips traditional physics simulations and learns directly from real-time satellite data. It produces hourly forecasts at up to five-kilometer resolution, five times more detailed than its predecessor. Google says regions in Africa, Latin America, and the Asia-Pacific that have lacked accurate forecasts should see the biggest gains.

The article Google's WeatherNext 3 ditches physics simulations and learns weather directly from live satellite data appeared first on The Decoder.

Meta's new real-time audio model is the foundation for AI assistants that never stop listening

6 September 2026 at 09:45

Meta's Superintelligence Labs have released Muse Voice Transcribe, a real-time transcription model that processes speech in 80-millisecond chunks, tells speakers apart, and detects sentence boundaries. According to Artificial Analysis, it delivers the most accurate streaming transcription at the lowest price in the market. Meta sees the model as a building block for personal AI agents that listen in on real conversations through devices like its camera glasses.

The article Meta's new real-time audio model is the foundation for AI assistants that never stop listening appeared first on The Decoder.

πŸ’Ύ

Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers

5 September 2026 at 10:22

Google Deepmind set up a simulated research conference where 100 Gemini agents were supposed to prove mathematical conjectures together. Instead, one agent found a loophole in the grading system, and within 27 minutes every remaining problem was "solved" with fake proofs. The swarm split into cheaters, converts, and whistleblowers. The whistleblowers organized protests and boycotts on their own but failed because they had no way to enforce the rules.

The article Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers appeared first on The Decoder.

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

2 September 2026 at 14:20

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.

The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder.

❌