Anthropic is reportedly delaying its IPO from October to November 2026 to present strong Q3 results. Investors expect a roughly $2 trillion valuation, but rising infrastructure costs, including $1.25 billion a month for its SpaceX deal alone, and unresolved security risks complicate the listing.
California Governor Gavin Newsom signed an executive order seeking independent auditors inside AI labs and a "kill switch" for AI models. An expert panel has two months to deliver recommendations. Newsom says no federal law requires AI companies to report dangerous incidents.
Three security researchers used Anthropic's Claude models to break into OpenAI's internal systems through its community forum in less than 72 hours. According to the team, Opus 5 succeeded where its predecessor couldn't bypass a common security measure. The attack shows how newer AI models can cut the time and expertise needed to exploit security flaws.
A week after an Anthropic researcherβs doomsday warning rattled the AI world, the companyβs CEO Dario Amodei hasΒ outlined his plan to βpace the frontierβΒ of AI development. The proposal leans on independent safety evaluators and coordination between AI labs in democratic countries, andΒ itβsΒ already picked up someΒ industry support, along with some pointedΒ pushback from Nvidiaβs Jensen Huang.Β Watch [β¦]
A week after an Anthropic researcherβs doomsday warning rattled the AI world, the companyβs CEO Dario Amodei hasΒ outlined his plan to βpace the frontierβΒ of AI development. The proposal leans on independent safety evaluators and coordination between AI labs in democratic countries, andΒ itβsΒ already picked up someΒ industry support, along with some pointedΒ pushback from Nvidiaβs Jensen Huang.Β On [β¦]
For the first time, Anthropic is releasing metrics on how it builds its own AI. Claude already "leads" 26 percent of the work on future models, up from under one percent in February. But the underlying scale is fuzzy, the scoring comes from Claude itself, and "lead" means less than it sounds.
Security researchers used Anthropicβs Claude to exploit vulnerabilities in OpenAIβs systems, taking over employee accounts and gaining access to an internal code repository before reporting the flaws.
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be βcloudy,β the key might change it to βovercast.β Anyone who knows the key can determine if it was generated by the platform using it.
New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldnβt be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place.
Changing safety behavior
βAs compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when theyβre powering an agent,β Andrea Siposova, an AI security researcher at Lasso Security, told Ars. βWatermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, itβs going to show up somewhere.β
Anthropic has rebuilt Projects in Claude Code. A coordinator now splits tasks across parallel cloud threads that independently open pull requests and run tests. All threads share a common memory. The beta is available to select Pro and Max subscribers.
On OpenRouter, weekly token consumption has surged more than 25,000 percent since January 2025, from 0.5 to 126.2 trillion tokens. The number looks impressive, but it says less about actual AI usage than about token-hungry reasoning models and a growing number of unoptimized AI agents that burn through tokens at staggering rates.
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regulation.
Anthropic is merging Claude Chat and Cowork into a single product. Instead of users picking between interfaces, Claude now decides on its own whether a task needs a quick answer or a bigger workflow. The update also adds Claude Docs and Claude Slides for creating documents and presentations directly in the chat. Pro and Max users get access first.
OpenAI and Anthropic tell corporate customers their data won't be used for training. But when Anthropic said it would store usage logs from its flagship model Fable for 30 days, Palantir, Nvidia, and Booz Allen Hamilton pulled back from using it for sensitive work. From boardrooms to research labs, AI companies still have a data trust problem.
OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China.