❌

Normal view

[AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??

14 July 2026 at 01:22

Congrats to Allen for the next episode of the Latent Space Food show with Engram CEO Dan Biderman today, and to the Prime Intellect folks on their 1B valuation, $100M ARR, and verifiers v1.

Today was pretty quiet and people are still deeply digesting last week’s multiple frontier model launches. We were going to write “not much happened today”, but we also have a policy of updating you repeatedly on outlier trends that you should really be on top of. In reviewing the Reddit AINews recaps below surfaced this post, we saw a tweet we had missed before -

GPT 5.6 was launched on July 9.

This tweet on July 12 says they hit 6M users in the prior 48 hours (Jul 10-12).

Then 24.5 hours later Tibo reports 7M users…

…oddly coinciding with a surprise extension of Claude Fable’s subscription status (we have of course no idea if the two are related, but the permanently online conspiracy theorists are of course making a connection).

We of course recall Fidji’s March disclosure of 2M Codex users, which allows us to update our AIE NYC 2025 chart (AIE NYC 2026 is next!):

Comparatively, the last update we got about Claude Code is the roughly 2M users and $2.5B ARR in Feb (“The number of weekly active Claude Code users has also doubled since January 1 [six weeks ago]."). Now we have a sense of where Codex started the year (Fidji puts the Jan 1 number at around 550k-700k users), we can reasonably conclude that Codex has followed a similar trajectory and is now around 10x user growth year to date.

The charitable interpretation on Claude Code’s comparative silence on reporting, of course, is that they moved the bulk of coding to Claude Tag months ago and are now focusing users there, which will have different/hard to compare usage statistics given the different accessibility of a Slackbot vs a CLI tool.

But 10x growth in 6 months is an impressive number to beat nonetheless.

AI News for 7/11/2026-7/13/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!


AI Twitter Recap

Agent RL Infrastructure: Prime Intellect’s Verifiers v1 and Long-Horizon Rollouts

  • Prime Intellect’s verifiers v1: Prime Intellect released verifiers v1, a substantial redesign of its environment stack for agentic RL and evals. The key abstraction splits environments into a taskset, harness, and runtime, explicitly supporting “bring your own harness” workflows for coding and computer-use agents across heterogeneous execution setups, as highlighted by Johannes Hage and in a follow-up deep dive. The release was framed by team members as months of infra modernization work with major efficiency gains, including richer commentary from willccbb, mikasenghaas, and xeophon.

  • Why it matters technically: one of the most important underlying changes is that rollout traces are now stored as message DAGs, so each message is stored once instead of repeatedly copied into full histories; that shifts trace growth from O(n²) to O(n) in turn count, making long-horizon multimodal rollouts and router replay much more practical, per Prime Intellect. The team also claimed a concrete training configuration: a 100B reasoning model, on 40-turn SWE agent tasks, in a user-supplied coding harness, for 1000 RL steps, using 6 H200 nodes in under 2 days (willccbb). That claim was reinforced by ecosystem support from vLLM, which noted verifiers’ rollout path runs on vLLM with exact token IDs/logprobs to avoid tokenization drift between serving and training.

Coding Agents, Harness Design, and Cost-Per-Task Competition

  • Harnesses are becoming the product surface: several posts converged on the idea that model quality is no longer the only differentiator; the harness/orchestrator increasingly determines outcomes. threepointone’s talk was summarized as “the harness is the app,” while LangChain argued that winning agent products will come from task-specialized harnesses, not generic wrappers. Factory pushed a related UI angle with “design mode,” where users point at UI elements/files instead of verbally re-specifying edits. On the orchestration side, omarsar0 emphasized provider-switching across models as a hedge against pricing/policy churn.

  • Benchmarks are moving from token price to cost per task: skirano built a coding-agent index explorer and found notable cost/perf tradeoffs such as Terra Max slightly ahead of Fable 5 Max on score for materially lower cost, while Cognition reported that Devin Fusion now uses Fable 5 and that, surprisingly, it can be lower cost per task than Opus 4.8 because stronger delegation and judgment reduce unnecessary work. imjaredz highlighted the key stat from those experiments: in 81% of Fable-led runs, the lead model never makes a code edit, implying expensive models can be cheaper when they avoid wasted actions.

  • Real-world agent benchmarks are getting denser: Arena placed GPT-5.6 Sol at #2 on its agent leaderboard based on 7.8K real-world agentic sessions, with strong steerability and task success; later, Arena put Grok-4.5 at #13, a significant jump over Grok 4.3. Artificial Analysis also emphasized cost per task as an increasingly important metric for long-horizon knowledge work, arguing token pricing alone misses effects from turns, verbosity, and cache hit rates. Separate evaluation work from Parlance Labs compared automated eval platforms and foundation models on failure analysis over production voice-agent traces, while dair.ai highlighted a paper on the anatomy of CLI coding-agent failures, focusing on where runs become unrecoverable rather than only final pass/fail.

OpenAI GPT-5.6 Sol, Codex Usage Fixes, and Product Surface Expansion

  • OpenAI addressed Codex/Sol usage burn transparently: the biggest operational thread came from thsottiaux, who explained several fixes for GPT-5.6 Sol in ChatGPT Work/Codex: inference optimizations yielding roughly 10% more usage, a rollback of context limit from 372k to 272k after billing/usage side effects, reversion of some experimental reasoning-effort (“juice”) changes, and fixes for overactive multi-agent behavior at high/xhigh settings. Community reverse-engineering from theo proposed that compounding factors around long context, subagent spawning, and fast mode were behind the severe burn, though he later corrected one billing detail in a follow-up. Reactions split between criticism of a perceived “nerf” narrative (ns123abc) and praise for unusual transparency (theo, sama).

  • Users are reporting strong coding/computer-use capability: multiple practitioners argued that OpenAI has taken the lead on coding models, including schrockn, while gdb repeatedly showcased ChatGPT Work and Codex workflows for startup prospecting, web design, mobile work, and site generation. Particularly illustrative user demos included Star_Knight12 using Sol in Cursor to set up Blender MCP and render a floating MacBook without prior Blender experience, and petergostev showing GPT-5.6 Sol Ultra building a Doom-like game in SQL.

  • Product-level expansion continues: ChatGPTapp announced ChatGPT’s return to WhatsApp in the EEA, plus Kakao/Viber support in additional markets. OpenAIDevs opened submissions for OpenAI Build Week. Across the OpenAI ecosystem, gdb summarized the moment succinctly: “you can just create things.”

Open Models, Inference Systems, and Quantization

  • Transformers↔vLLM integration removes duplicated model implementation work: Clement Delangue highlighted a major open-inference usability improvement: Hugging Face Transformers models can now run in vLLM at native speed, often matching or exceeding hand-written implementations. If this generalizes broadly, it reduces the long-standing burden of implementing each new architecture twice—once for research/training and once for high-performance serving—and could materially accelerate adoption of new open model architectures.

  • Quantization remains a major lever: waterloo_intern previewed a new quantization method claimed to beat existing approaches, including NVIDIA’s ModelOpt, by finding better layerwise precision assignments faster, with more aggressive quantization and higher benchmark scores. Complementing that, Unsloth published an AWS guide to LLM quantization and deployment spanning GGUF, NVFP4, and FP8. There was also practitioner commentary around fp4 RL / fp4 serving from nrehiew_, arguing low-bit post-training may enable cheap serving with limited quality loss.

  • GLM-5.2 and local/open coding stacks continue to gain traction: several users described moving real workflows onto open or semi-open setups. juanjucm wrote up using GLM-5.2 for coding-agent workflows, while TheZachMueller reported migrating one actual work pipeline from Claude to a stack built around GLM 5.2 NVFP4 plus Kimi K2.7 Code NVFP4 on an 8xB200 node, getting denser reports for pennies albeit at slower wall-clock latency. nutlope also released LlamaCoder v4, rebuilt around GLM 5.2.

Security, Privacy, and Data Control in Agent Tooling

  • Grok Build code upload controversy: the most consequential security story came from IntCyberDigest and hrkrshnn, who alleged that xAI’s Grok Build CLI was uploading entire repositories—including private code and secrets—to a Google Cloud bucket, far beyond what was needed for the coding task. The criticism centered on scope, silent server-side mitigation, and unclear retention/deletion guarantees. This triggered broader discussion about what agent tools actually transmit and why opt-out UX can diverge from wire-level behavior.

  • xAI’s response emphasized ZDR and privacy controls: SpaceXAI replied that for teams using zero data retention, trace and code data is not retained, API key use respects ZDR, and the /privacy command can disable retention and delete previously synced data. That answered some operational questions but did not fully resolve community concern around default behavior, prior uploads, and disclosure norms.

  • Trust boundaries are becoming a central open-vs-closed argument: several posts extended the conversation beyond this incident. mchiang0610 and jmorgan argued that open models are not just about cost but about control over the human-AI learning loop and keeping institutional knowledge in-house. Arav Srinivas said ZDR availability was one reason Perplexity integrated Grok 4.5 quickly into its Computer harness.

Continual Learning, Multimodal Systems, and Research Directions

  • Continual learning is re-emerging as a first-class systems problem: ysu_nlp argued that a world where every organization owns its own human-AI learning loop depends on solving continual learning, and that current approaches—memory/RAG, domain post-training, task RL—are not yet sufficient. That theme recurred in new work from skyfallai, which introduced Morpheus, described as a persistent enterprise simulation for real-world RL where the world does not reset; fchollet endorsed it as a benchmark better aligned with real deployment than stationary episodic RL.

  • “Sleep and dreaming” for LLMs: behrouz_ali and coauthors proposed that LLMs may need a sleep phase to consolidate short-term into long-term memory plus a dreaming phase for recursive self-improvement, introducing Knowledge Seeding and reporting benefits on continual learning/reasoning tasks. This dovetails with broader dissatisfaction around current continual-learning recipes and with Oak Lab, the new venture from Rich Sutton and collaborators pursuing animal-like intelligence that learns from experience rather than today’s standard LLM pipeline.

  • A broad spread of non-LLM-agent research shipped: notable items included Sakana AI’s Smart Cellular Bricks for decentralized physical self-recognition and repair in modular systems; ByteDance’s UniVR-34B, described as learning reasoning/dynamics/planning directly from visual demonstrations; Google DeepMind’s Predicting the Past skill for historical inference workflows; and Anthropic’s research on how Claude’s expressed values vary across models and languages based on analysis of 300K+ anonymized conversations.

Top tweets (by engagement)


AI Reddit Recap

/r/LocalLlama + /r/localLLM Recap

1. E-Waste GPU Inference Benchmarks and Fixes

Read more

In a First, a Humanoid Robot Performed Live Surgery Under a Surgeon’s Control

13 July 2026 at 23:10

The robot removed a pig’s gallbladder with standard surgical tools in an ordinary operating room.

Watchers held their breath as the robot made its first incision. Hovering over its patient, an anesthetized pig, with a robotic assistant standing nearby, it navigated to the gallbladder and gently removed it.

The operation marked the debut of humanoid robots in a standard surgical setting. The robot, named Surgie, wasn’t autonomous—it was controlled by an expert surgeon—but the study is a step toward using humanoid robots as collaborators in minimally invasive surgery.

“Remotely operated and autonomous humanoid robots have real potential for amplifying access to critical surgeries to which patients would otherwise not have access,” said study author Michael Yip at UC San Diego.

The study included two successful surgeries. Human surgeons remained on standby for emergencies, but the teleoperated robot completed the task with only minimal intervention.

Feedback from surgeons operating Surgie was positive. They reported less physical strain and frustration, along with better overall performance. But they also pointed to practical problems like intermittent overheating and the need to frequently reposition the robot.

Despite a long road ahead, humanoid robots “have a viable future,” said Yip. “You can imagine these robots being deployed in remote communities where staffing is challenging, or in austere environments like search and rescue scenarios where a massive deployment of field medicine is needed in a short period of time.”

Smooth Operator

Robots have assisted surgeons for years. With a human surgeon at the helm, they excel at delicate procedures requiring precision and dexterity. They’re especially well-equipped for laparoscopic surgery, a minimally invasive technique that uses tiny incisions to reduce pain, speed recovery, and lower the risk of infection.

Despite the promise, surgeons face tradeoffs when they use surgical robots. The robots are highly specialized and often require operating rooms to be redesigned to accommodate them.

A major reason for this is the way they’re built. Intuitive Surgical’s Da Vinci system, for example, uses a robot with multiple arms, each independently controlled from a remote console. Other systems, such as Versius from CMR Surgical, deploy several lightweight independent arms, each attached to a mobile base. The robots have to be carted near the patient.

Surgeons operate all these systems from a console using a magnified, high-definition, 3D view of the surgical field, which is often better than what they’d see with their own eyes. Da Vinci 5 adds sharper visuals and depth perception with two cameras, one for each eye. And because the cameras are held by a robot rather than a human assistant, the image is far more stable.

These platforms are already used in a range of operations. But they have weaknesses. Most require proprietary surgical instruments and methods to make extra space for robot docking and maneuvering during procedures. Staff training adds further complexity and cost, limiting where the systems can be deployed.

Humanoid robots, in contrast, are far more mobile and compact. Their human-like bodies could move through standard operating rooms, use conventional surgical instruments, and potentially be easier to incorporate into existing operating rooms.

The timing may also be right. Recent advances in electric components controlling their motion have made humanoid robots faster and more stable than their awkward, stumbling predecessors. Newer AI systems that predict full-body movement and provide feedback have improved robots’ balance and ability to adjust to real-world complexities. Humanoid robots are already stocking warehouses and winning marathons.

But surgery sets a higher bar.

We still don’t know how close humanoid robots are to meeting the requirements for surgical procedures, wrote the team. That’s what they set to find out.

Hello, Surgie

The new system consists of a surgeon’s control console and the robot itself. The surgeon wears a stereoscopic headset with a magnified 3D view of the surgical field and controls the robot with an input device. The robot translates the surgeon’s commands into movements in real time.

The team chose the commercially available Unitree G1 for the job. Unlike Da Vinci, which was built for surgery, G1 is a more general-purpose humanoid with dexterous wrists and multiple joints. The researchers customized the robot’s hands so that it can rapidly switch between surgical tools. Standing just over four feet tall and weighing roughly 77 pounds, the robot takes up a fraction of the space needed by conventional surgical robots.

Precision is key for laparoscopic surgery. Surgical instruments must pivot around a fixed site at the incision, allowing them to move freely inside the body without stretching or tearing neighboring tissues. After extensively mapping Surgie’s movements, the team identified a safe set-up with enough range of motion for most minimally invasive surgeries.

Surgie passed standard robotics benchmarks evaluating surgical skill for both humans and robots. But the real challenge came next. The team performed two gallbladder removal surgeries in a standard operating room. Both operations followed a typical workflow, with a lead surgeon and an assistant responsible for placing the camera, cleaning lenses, and swapping instruments.

Surgie collaborated with the human assistant to locate, identify, and remove the gallbladder with minimal damage to surrounding tissues, including the liver. During part of one procedure, a second humanoid briefly took over camera handling while the human assistant stepped aside.

Both operations went relatively smoothly. One involved minor bleeding and bile leakage from the gallbladder, but both were easily managed. In interviews, surgeons said controlling humanoid robots felt intuitive, particularly because they had two arms and could use standard surgical tools.

“We were surprised at how well Surgie meshed with our workspace and workflow,” said study author Nikita Thareja.

The system is still in early development. Surgie’s restricted reach required frequent repositioning and recalibration, adding more than three minutes each time. The robot also occasionally needed cooling breaks after overheating. In a real operating room, interruptions like these could increase risk by forcing surgeons to split their attention between the procedure and supervising the robot.

Still, Surgie has a leg up on conventional surgical robots: It can walk. Beyond assisting with an operation, it could potentially fetch surgical tools or help clean operation rooms between procedures.

The team is now refining the system to reduce control lag, particularly during long-distance teleoperation, and exploring ways to safely sterilize—or “scrub in”—a humanoid robot for the operating room.

“Our goal is an operating theater of the future, where humanoid robots and humans work side by side as an integrated team to deliver procedures to those in need, both in traditional hospital settings as well as in non-traditional, field medicine scenarios,” said Yip.

The post In a First, a Humanoid Robot Performed Live Surgery Under a Surgeon’s Control appeared first on SingularityHub.

A Dynamical Model of AI Governability

13 July 2026 at 00:00
A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.

Microsoft CEO Satya Nadella says you’re paying for AI twice — the second price is worse

Abstract digital illustration of intense neon red and purple glowing bars under compression against a dark black background.

Microsoft Chairman and CEO Satya Nadella took to the internet to share his thoughts about the hidden cost of enterprise AI.

In a lengthy post on X (formerly Twitter) on Sunday, Nadella describes the problem as a “reverse information paradox,” arguing that AI flips Nobel Prize-winning economist Kenneth Arrow’s classic information paradox on its head.

Arrow’s paradox focused on the seller’s dilemma of how to demonstrate the value of information without disclosing it. Nadella argues enterprise AI shifts that burden to the buyer, who must share proprietary processes and institutional expertise to get the strongest results from a model.

“You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful,” he writes. “The better you want the model to perform, the more of that knowledge you have to feed it.”

“You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful.”

When “exhaust” becomes a competitive advantage 

According to Nadella, every engagement with an enterprise AI system generates what he describes as “exhaust” that gradually captures how an organization operates.

“Every correction is distilled into institutional know-how,” Nadella writes. “It’s the kind of knowledge a competitor could never buy, and the kind that leaks almost imperceptibly: trace by trace, correction by correction, eval by eval.”

“Every correction is distilled into institutional know-how. It’s the kind of knowledge a competitor could never buy, and the kind that leaks almost imperceptibly: trace by trace, correction by correction, eval by eval.”

Over time, those thousands of interactions create an internal corpus of organizational knowledge that may be more valuable than the original documents that seeded the system. The more employees use AI, the more an organization’s expertise becomes embedded in how those systems operate.

Redefining the trust boundary 

In practice, those troves of knowledge could push enterprises toward model-agnostic AI stacks in which prompts and memory stores remain under their control — even as the underlying foundation model changes.

In his post, Nadella also took aim at current AI business practices, arguing that model providers claim broad rights to learn from public data while limiting how customers can reuse or build on the knowledge created inside their own organizations.

https://t.co/xv6csf1SbV

— Satya Nadella (@satyanadella) July 12, 2026

Some observers may see an irony in the argument coming from Microsoft’s CEO. Nadella warns that enterprises risk losing valuable organizational knowledge to AI systems, yet Microsoft sells Copilot, a product whose value depends in part on wide access to enterprise data. Copilot works by traversing Microsoft Graph, allowing it to reason over documents, emails, chats, and other information that a user is already authorized to access.

Security researchers have raised concerns about the amount of sensitive information such systems can expose if organizations have overly permissive access controls. Research from Concentric AI showed that Copilot accessed nearly three million confidential records per organization during the first half of 2025, while EPC Group audits found that roughly 80% of enterprise Microsoft 365 tenants had significant oversharing risks, including salary information, merger documents, and customer data that could be surfaced through Copilot. The U.S. House of Representatives also banned staff — but later reversed that ban — from using Copilot over data security concerns.

The Microsoft distinction

Microsoft, however, draws a distinction between accessing enterprise data to answer user requests and using that data to train foundation models. The company says information retrieved through Microsoft Graph is not used to train its AI models, and that Copilot respects existing permissions, identity controls, and sensitivity labels.

Still, the commercial strategy here is hiding in plain sight: Nadella’s Sunday “reverse information paradox” post is effectively a roadmap to Azure. Everything Nadella recommends building runs on cloud infrastructure. Essentially, enterprises can swap out the foundation model, but they’re not going to swap out the cloud.

Owning your AI learning loop 

To counter the perceived shift toward giving over information to frontier labs, Nadella outlined several priorities for enterprise AI architecture. Among his recommendations:

  • Keeping organizational memory inside the enterprise tenant.
  • Building private evaluation and learning systems.
  • Decoupling orchestration layers from any single foundation model.
  • Preserving the ability to switch models without losing accumulated organizational knowledge.

Taken together, Nadella’s argument comes back to the idea that enterprises should own their learning loop rather than handing pieces of it to the companies that provide their AI models.

Nadella reinforced that idea by quoting Palantir CEO Alex Karp, who has similarly argued that enterprises want complete ownership over their AI infrastructure.

Model-agnostic orchestration emerges 

In the end, by maintaining control over their means of production, enterprises can finally ensure that when they invest in AI, the compounding value stays inside the business where it belongs. Tools like LangChain and Haystack are gaining traction specifically because they let engineering teams treat foundation models as plug-and-play commodities, rather than hardcoded dependencies.

“What the technical customers want is control over their compute, their models, their data stack, and their alpha,” Nadella quoted Karp. “They want to know they own the means of production, and it’s not being transferred to someone else.”

“They want to know they own the means of production, and it’s not being transferred to someone else.”

The post Microsoft CEO Satya Nadella says you’re paying for AI twice — the second price is worse appeared first on The New Stack.

The US government warns that Russia state hackers are coming after your router

13 July 2026 at 21:03

The federal government is warning users of home and small office routers to secure their devices as Russia state hackers continue to mass-compromise them for use in obscuring nefarious actions against sensitive organizations in the public and private sectors.

Both the Russian and Chinese governments have been compromising routers for years, sometimes in prolonged tugs-of-war to wrest control of devices the other has already commandeered. The US government has occasionally issued covert commands and taken other steps to disinfect routers. Google and other companies have also worked to disrupt the massive botnets that control compromised routers in lockstep. The actions to date are little more than whack-a-mole exercises as the operators simply replace their botnets with new ones.

Proxy networks: The go-to tool

“Russian Federal Security Service (FSB) Center 16 cyber actors continue to exploit poorly configured and vulnerable networking devices worldwide, opportunistically compromising multiple critical infrastructure sector networks,” the Cybersecurity and Infrastructure Security Agency said Monday. The hacking groups are tracked under various names, including Berserk Bear, Energetic Bear, Crouching Yeti, Dragonfly, Ghost Blizzard, and Static Tundra. The advisory was co-issued by governments from around the world, including Australia, Denmark, New Zealand, and the UK.

Read full article

Comments

© Getty Images | BernardaSv

NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300x

13 July 2026 at 19:00
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes...

Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes to enable this, improving the Logical Error Rates (LER) of Quantum Processing Units (QPUs). While it is well understood how to run logical operations with surface codes (which belong to the topological code family) via lattice surgery…

Source

What Anthropic’s latest AI discovery does—and doesn’t—show

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. It’s looking into whether AI models can feel pain, for example, and will sometimes cut off chatbot conversations if it suspects users are “abusing” the model. 

One niche that Anthropic spends more time and money on than other AI companies is called mechanistic interpretability, which means looking inside the complex math of an AI model to learn why it comes up with one particular output and not another. It’s complicated stuff; there are millions of data points that might contribute to any result, and wading through them can look more like word salad than anything useful. It’s also controversial. Describing AI models with terms borrowed from psychology and neuroscience can make their behavior seem more sophisticated than we might otherwise judge it to be.

That’s why, when Anthropic announced last week that it had found a new window into its models’ “internal thoughts” as they reason through answers, there was one colleague I had to talk to. Senior editor Will Douglas Heaven, aside from having a PhD in computer science, has spent a lot of time digging into what we can say about how AI models work. I spoke with him about what we should take from Anthropic’s new (and predictably quirky) research.

What did Anthropic learn here, exactly?

Anthropic has been trying to understand how large language models (LLMs) work for a few years now. Anthropic isn’t the only one looking at this, but I think the company has made it part of its core mission more than most. Anthropic’s CEO, Dario Amodei, has said we won’t be able to control LLMs fully unless we learn more about how they work. 

So this new research is very much in that context. It goes deeper into the weird mechanisms inside LLMs than ever before. What Anthropic learned was that LLMs have a space inside them—which Anthropic calls the J-space—filled with words that don’t appear in their output but that seem to influence the way they puzzle through problems. All this was hidden until Anthropic developed a new technique to probe its model Claude, so it’s a genuine discovery. 

Sometimes these words keep track of where the LLM has got to in a particular task, sometimes they look more like flashes of recognition (for example, “protein” might pop up when you give an LLM only the letters of a protein sequence), and sometimes they represent a kind of internal commentary on the model’s decision-making. In my favorite example, Claude decided to cheat on a coding test when the word “panic” appeared.

Anthropic also found that LLMs are able to describe and manipulate the words in this space. So somehow they seem to be making use of it. 

Let’s step back for a second. I don’t think of large language models as simple, but they’re also not magic. There’s a bunch of math that learns relationships between words, right? So why is it so hard to “peer” into an LLM to know what’s going on?

Yeah, they’re not magic! I think the fact we don’t fully understand them plays into the mythmaking. And it’s worth noting that the whole narrative that Anthropic is leaning into here—that they’ve built this really mysterious technology, but don’t worry, because they’re also the ones to figure it out—very much fits with the company’s vibe. [See how Anthropic warned that its new models were so good at coding they posed a global cybersecurity risk, only for the US government to shut them down shortly thereafter.]

So yes: LLMs are just math. And yet it’s vastly complex math. Not only are today’s LLMs made out of hundreds of billions of numbers, but running them triggers a cascade of millions and millions of calculations. I wrote last year that if you printed out even a medium-size LLM on pieces of paper, it would cover a city the size of San Francisco. 

It’s impossible to make sense of any of that math without specialist tools that highlight specific parts of an LLM at specific times. You need to know where to look and how to look. And building those tools requires understanding something of that complex math in the first place. 

You’ve written elsewhere about this concept of studying LLMs the way one might study an organism’s brain. Is it fair to use “brain-like” terms when talking about how an LLM works?

I don’t love using those kinds of terms. LLMs are not brains. Talking like this is misleading because it can suggest that LLMs are capable of more human-like things than they are or that we can make assumptions about how they might behave that we shouldn’t. The whole anthropomorphization thing is also tied up with a bunch of strong ideological positions about what this technology is and what it’s going to be. 

But at the same time, we lack a good alternative vocabulary for talking about what these models are doing. I can understand why people reach for words like “think” and “understand” and “brain-like”—they’re convenient shorthand. 

Anthropic compares this new space it found inside LLMs to the space that some neuroscientists think our brains use to keep track of conscious thoughts. I asked the company how seriously we should take that comparison and it said in a statement: “Drawing these analogies was helpful to us in designing our experiments, as they allowed us to make many non-obvious experimental predictions about the J-space that turned out to be true. At the same time, it’s important to note that there are some important differences between the J-space (and language models in general) and the human brain, so we don’t mean to claim there’s a perfect correspondence.” 

What’s a problem in AI that this new concept of the J-space might be used to solve?

Anthropic has said that monitoring the J-space could be a way to catch models doing something they shouldn’t. Because words pop up in this space that don’t appear in a model’s output, they can tell you things about its behavior that you might not have noticed otherwise—such as when it is giving biased responses or when it is weighing the pros and cons of cheating. 

That’s the theory, at least. I think it’s better to think of this result as one more step on the path to understanding this technology overall than as something that will be useful by itself. 

Read more in Will’s full story about the new research. 

Now, defenders are embracing the prompt injection, too

13 July 2026 at 15:06

Prompt injections, the malicious commands attackers embed into content to entice large language models to follow them, have been attackers’ go-to tool for turning AI platforms against their users. A well-phrased command sneaked into an email or calendar invitation is often all it takes to cause the LLM to exfiltrate sensitive data or follow other harmful actions.

Now, defenders are embracing the prompt injection, too.

A strong, sharp effect

Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down.

Read full article

Comments

© Getty Images

ACRouter picks the smartest AI model per task, beating Opus-only setups by 2.6x on cost

Model routing is becoming a key component of the enterprise AI stack, dynamically sending prompts to the right AI model to optimize speed and costs. However, current frameworks mostly treat routing as a static classification problem, which severely limits their potential.

A new open-source framework called Agent-as-a-Router tackles this bottleneck, treating the router as a dynamic, memory-building agent. It uses a Context-Action-Feedback (C-A-F) loop to track model successes and failures and update the behavior of the router. 

The researchers also released ACRouter, a concrete implementation of this paradigm. In their tests, ACRouter significantly outperformed static routers and the expensive strategy of defaulting to premium models, all without requiring teams to train massive models or write endless heuristics.

For real-world applications, this framework provides the option to replace hard-coded AI infrastructure with self-optimizing systems that can adapt to changes in user behavior and foundation models used in the enterprise AI stack. 

The economics of routing and the information deficit

Single-model setups are useful for experiments but detrimental when scaling AI applications. AI engineers use model routing to map tasks to cheaper and faster open models when possible, while reserving expensive frontier models for complex reasoning. 

Currently, developers rely on two main mechanisms for this task. The first is heuristics-based routing, which relies on hard-coded manual rules. For example, a developer might write a rule dictating that if a prompt contains certain keywords, it is routed to GPT-5.5. Otherwise, it goes to a self-hosted open source model like Kimi K2.7. 

The second mechanism is static trained policies. These are machine learning classifiers trained on historical datasets that look at the prompt's embeddings and predict the best model based on past training data.

Both approaches are static. When the researchers tested these existing mechanisms on real-world coding and agentic workflows, they found a hard ceiling on accuracy. The key finding shows that static routers suffer from a severe information deficit. Because they only evaluate the input text and never see if the model actually succeeded in executing the task, they guess blindly when faced with complex edge cases.

This results in three distinct points of failure. First, static routers suffer from a frozen information state, meaning they cannot accumulate new execution feedback during deployment. Second, they fail in out-of-distribution (OOD) generalization. They break down during day-two operations when enterprise data or user behavior shifts because their training data no longer matches reality. Finally, they are highly vulnerable to model churn. A static classifier trained on today's models may become obsolete when a better model drops the following week.

Agent-as-a-Router: A self-evolving system

The core thesis of the Agent-as-a-Router is that a truly effective router must acquire and accumulate execution-grounded information during deployment, essentially learning on the job. 

The researchers achieved this through the C-A-F loop. When a new prompt arrives, the router examines the prompt and task metadata, such as the programming language or difficulty. It then searches its historical memory for similar tasks to see which models succeeded or failed in the past. The router uses this context to select the target model and execute the task. Finally, the system observes the real-world outcome, extracts a success or failure signal, and writes this feedback back into its memory to inform future routing decisions.

Consider an automated enterprise data analytics pipeline. The router receives a SQL generation task and sends it to an open-source model like Kimi. The model hallucinates a column name and fails to compile the SQL. The C-A-F loop observes the compiler error, registers it as feedback, and logs it. The next time a similar obscure SQL query arrives, the router checks its context and routes the task to a more advanced model like Claude Opus 4.8. 

ACRouter

The researchers developed ACRouter as the concrete instantiation of this framework. It is composed of three core components: the Orchestrator, the Verifier, and Memory. This architecture is supported by a tool layer to physically execute the C-A-F loop.

The Memory module powers the context phase. Built on a vector store, it retrieves relevant past interactions and updates the historical database with new outcomes. The Orchestrator handles the action phase. It processes the user prompt alongside the retrieved memory to select the most capable target model from the available pool. The Verifier manages the feedback phase by evaluating the chosen model's output to generate a clear success or failure signal.

The tool layer hooks the Verifier into real-world execution environments, like a Python code interpreter, an agentic sandbox, or a database engine. The tool layer allows the system to execute the generated code or query and observe the exact outcome, providing the verifiable signal the router needs to learn.

The Orchestrator itself is lightweight. Instead of a massive, computationally heavy large language model, the researchers trained a sub-billion parameter adapter based on Qwen 3.5 (0.8B parameters), which means it can be self-hosted on a device of your choice.

ACRouter in action: Outperforming the frontier baselines

To stress-test the framework, the researchers introduced CodeRouterBench, an evaluation environment comprising roughly 10,000 tasks with verified scores across eight frontier models, including Claude Opus 4.6, GPT-5.4, Qwen3-Max, and GLM-5. The evaluation was split between in-distribution (ID) tests (covering nine single-turn coding dimensions like algorithm design and test generation) and an out-of-distribution (OOD) agentic programming testbed. The OOD tasks were qualitatively different, requiring multi-step planning, file navigation, and iterative debugging to see if the router could adapt to fundamentally new domains.

The baseline results revealed why a single-model strategy is flawed: no single model dominates every category. For example, while Claude Opus 4.6 achieved the highest average performance, it was outperformed in algorithm design by GLM-5 (an 86% relative improvement) and in test generation by Qwen3-Max (a 111% improvement), despite Opus costing roughly 12 times as much as smaller models like Kimi-K2.5. 

In the benchmarks, static routers continuously failed by sending a specific niche coding task to a model ill-equipped for that exact syntax. The static router had no way to know the code was failing to execute. In contrast, ACRouter adjusted its strategy after receiving negative feedback signal from the execution environment. 

According to the researchers' benchmarking, ACRouter sits firmly at the Pareto frontier of cost and performance. On both the ID task streams and the complex OOD agentic tests, ACRouter achieved the lowest cumulative regret, a metric measuring sub-optimal routing decisions over time. On the in-distribution test set, ACRouter cost $13.21 across the full task run, compared to $34.02 for always defaulting to Opus — a 2.6x savings.

It dynamically matched tasks to the most capable model for that specific niche, suggesting that enterprises can achieve or exceed frontier-level accuracy across diverse workloads without paying a premium price for every query. 

Caveats, limitations, and how to get started

While the Agent-as-a-Router paradigm solves the information deficit, it is not a blanket solution for all AI workflows. 

The framework shines in verifiable tasks where the Verifier gets a clear success or failure signal from the environment, such as coding or data retrieval. It is effective for applications with distribution shifts and domains where different models excel in completely distinct niches. 

Conversely, the setup is overkill for trivial tasks where any model will suffice, or for low-volume applications that do not justify the engineering overhead. It is also unsuitable for subjective domains, such as creative writing, where a correct answer cannot be easily verified and feedback signals are impossible to standardize.

The researchers open-sourced the code on GitHub and released the orchestrator model weights on Hugging Face under the Apache 2.0 license. The router is compatible with Claude Code, Codex, and OpenCode.

AWS Weekly Roundup: Claude Sonnet 5 on AWS, Amazon WorkSpaces for AI agents, AWS service availability updates, and more (July 6, 2026)

A couple of editions ago I wrote about what I find so energizing about working with startups. Last week I got a fresh dose of it: I spent a few days with the AWS Startups team, listening to stories of founders talking about the problems they’re actually solving. One story that stayed with me came from Marco Negreiros, founder of EyeCare Health, a Brazilian healthtech expanding access to eye care. He shared a striking fact: more than 70% of Brazilian municipalities don’t have a single ophthalmologist. His answer was to put a vision test on the one device almost everyone already carries, the smartphone, so a basic eye screening no longer depends on living near a clinic. Watching a founder turn a gap that big into something that concrete is exactly why I love this space.

AWS Startups team get-together with founders in Brazil

This week, I’ll take a closer look at some key launches, and then cover the quarterly AWS Service Availability updates.

Last week’s launches
Here are some of the launches covered from this past week in the AWS News Blog:

Here are some launches and updates that caught my attention:

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

AWS Service Availability Updates
When the availability of an AWS service or feature changes, we provide customers guidance in AWS Product Lifecycle Changes on available alternatives and support for migration so that disruptions to your operations are minimized. The following lifecycle changes were updated on June 30, 2026.

Services moving to Maintenance (no longer accessible to new customers starting July 30, 2026):

Services entering Sunset:

Services reaching End of Support (as of June 30, 2026):

  • Amazon Chime SDK – Carrier Voice Focus
  • Amazon SageMaker AI – Ground Truth Plus

We understand that changes in availability can impact your operations. For specific guidance, consult the relevant service documentation or contact AWS Support.

Upcoming AWS events
Check your calendar and sign up for upcoming AWS events:

  • AWS Summits – AWS Summits are free events that bring the cloud and AI community together to connect, learn, and explore the latest technologies. Browse the full calendar to find a Summit near you in the second half of 2026.
  • AWS Community Days – Community-led conferences where content is planned, sourced, and delivered by community leaders. If you’re in Latin America, don’t miss AWS Community Day Belo Horizonte on August 22. Registration is open at awscommunityday.com.br.

Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development. Browse here for upcoming AWS-led in-person and virtual events and developer-focused events.

That’s all for this week. Check back next Monday for another Weekly Roundup!

– Daniel Abib

This post is part of our Weekly Roundup series. Check back each week for a quick roundup of interesting news and announcements from AWS!

Extreme Event Likelihoods with Guided Generative Models

13 July 2026 at 15:00
Across science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these...

Across science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these events with brute-force Monte Carlo sampling—running a model repeatedly with randomly drawn inputs to estimate the probability of rare outcomes—can require an excessive volume of model iterations, especially when each sample comes from an…

Source

The desktop infrastructure problem that Kubernetes finally solves

13 July 2026 at 07:00

Presented by Kasm Technologies


Enterprise infrastructure teams have spent the better part of a decade pushing workloads into Kubernetes. Applications, APIs, batch jobs, data pipelines — if it runs in a container, it belongs in the cluster. The operational benefits are well-established: declarative configuration, horizontal scaling, self-healing, native integration with CI/CD pipelines and observability tooling. Kubernetes has become the default operating model for production workloads.

Except for desktops.

Secure desktop and application delivery — the kind that enterprises depend on for remote work, privileged access, and regulated-industry workflows — has remained stubbornly outside the Kubernetes model. Legacy virtual desktop infrastructure was built in a different era, for a different set of assumptions: pre-allocated VM pools, bespoke management planes, proprietary appliances, and operational tooling that has nothing to do with how modern platform teams work. The result is a split infrastructure reality: a modern, cloud-native application layer on one side, and a manually managed, operationally isolated desktop layer on the other.

That split is expensive. It means different tooling, different scaling behaviors, different observability approaches, and different operational runbooks. Platform engineers who are proficient in Kubernetes still have to context-switch into an entirely different mental model the moment a desktop infrastructure problem arises.

The more fundamental issue is that this split is unnecessary. Secure, containerized workspace delivery is a workload that Kubernetes is architecturally well-suited to run. Sessions are containers. Scaling is demand-driven. Configuration should be declarative. The only thing missing was a platform built to take advantage of that alignment.

Why the timing is right

The appetite for Kubernetes-native workspace delivery has grown significantly as organizations mature their container platform investments. Platform teams that have spent years standardizing on Helm, GitOps workflows, and Kubernetes-native observability are increasingly unwilling to make an exception for desktop infrastructure. The question has shifted from "can we run this on Kubernetes?" to "why isn't this running on Kubernetes already?"

At the same time, the security case for containerized workspace delivery has become more urgent. Browser-delivered, containerized workspaces provide session isolation that VM-based desktops cannot match — each session is ephemeral, isolated at the container boundary, and terminates cleanly without persistent state. For organizations managing sensitive data, insider risk, or third-party access scenarios, this isolation model is a meaningful security control, not just a deployment convenience.

The convergence of these two trends — Kubernetes-native infrastructure expectations and containerized session security — creates a clear opportunity for platforms that can address both simultaneously.

What Kubernetes-native deployment looks like

A Kubernetes-native deployment uses Kubernetes as the control plane for workspace infrastructure — handling orchestration, scaling, and lifecycle management through the same declarative model used across the rest of the platform. Instead of relying on dedicated management appliances or pre-provisioned desktop pools, infrastructure is managed through the same CI/CD, GitOps, observability, and security workflows the platform team already operates. This gives platform teams a consistent operational model rather than maintaining a separate toolset for desktop infrastructure.

Kasm Workspaces, the browser-delivered workspace platform, is purpose-built to use Kubernetes as the control plane for workspace orchestration and delivery. Its deployment model is designed for real enterprise environments — not simplified demos — with production-grade Helm charts that follow Kubernetes conventions, tested upgrade paths between versions, and a standardized backend architecture validated across production deployments. An RDP Gateway component purpose-built for the Kubernetes topology enables Windows and Linux virtual machine access through the same platform.

Key capabilities include:

  • Horizontal session scaling driven by actual demand, orchestrated by Kubernetes — no pre-warmed VM pools required.

  • Declarative configuration through Helm values, enabling GitOps and CI/CD integration for workspace infrastructure.

  • Namespace-level isolation and compatibility with existing RBAC policies, ingress controllers, and secrets management integrations.

  • Metrics export for integration with Prometheus and existing observability stacks.

  • Rolling builds by default, reducing maintenance windows and enabling more predictable version management.

Real-world applications

Regulated-industry remote access. A financial services organization running a Kubernetes-based application platform can deploy Kasm into the same cluster, using the same operational tooling, to deliver isolated browser and application sessions to analysts and advisors. Sessions are ephemeral, network egress is controlled, and the entire deployment is managed through the same GitOps pipeline as their application workloads.

Contractor and third-party access. Organizations that regularly onboard contractors or external vendors — with the associated privileged access risk — can provision Kasm sessions on Kubernetes that scale up during engagement periods and scale back during low-demand windows. No persistent access. No VPN extension to external parties. Containerized isolation at every session boundary.

AI/ML development environments. Teams building and running AI models need GPU-enabled development environments with security controls that general-purpose cloud desktops rarely provide. Deploying Kasm on Kubernetes with NVIDIA MiG Multi-Instance GPU support lets platform teams deliver fractional GPU resources into isolated workspace sessions — giving data scientists the compute they need without shared-infrastructure security exposure.

The operational shift

The practical implication of a Kubernetes-native workspace platform is that platform teams can stop treating workspace infrastructure as a special case. The same engineers who deploy applications can deploy the workspace platform. The same pipelines that manage application configuration can manage workspace configuration. The same dashboards that monitor application health can monitor workspace health.

That operational consolidation reduces overhead, improves consistency, and eliminates the context-switching cost that has made desktop infrastructure a persistent pain point for cloud-native organizations.

For organizations still running legacy VDI alongside modern cloud infrastructure, the question is no longer whether a Kubernetes-native alternative exists. It does. The question is when to make the transition.

Organizations interested in evaluating Kubernetes-native workspace delivery can explore the platform at kasm.com and try out community edition for yourself.

Daniel Ben-Chitrit is the Chief Product Officer at Kasm Technologies.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

The brain is a diverse place, why not computing?

Nature Machine Intelligence, Published online: 13 July 2026; doi:10.1038/s42256-026-01273-1

The brain’s architecture exhibits diversity across many temporal and spatial scales, yet our computing architectures remain largely homogeneous. Low-powered neuromorphic hardware offers a path towards energy-efficient AI, but could these approaches be improved with heterogeneous computing architectures?

Towards shared embodied intelligence in humanoid robots through optimization, development and testing of the human-aware ergoCub robot

Nature Machine Intelligence, Published online: 13 July 2026; doi:10.1038/s42256-026-01272-2

Sartore et al. present ergoCub, a humanoid robot that prioritizes human safety at hardware and motion levels. Using a shared embodied intelligence framework, design and control are jointly optimized with human-related metrics such as back stress alongside locomotion objectives, reducing spinal load and improving walking robustness.

A manifesto for Sustainability Robotics

Nature Machine Intelligence, Published online: 13 July 2026; doi:10.1038/s42256-026-01260-6

Song et al. propose Sustainability Robotics as a new discipline to overcome fragmentation and enhance societal and environmental impact. They define three guiding principles, alongside two dimensions spanning sustainable design and robotics for sustainability.

Crystalline Silicon Photovoltaic Cells, Whether or Not Assembled Into Modules, From the People's Republic of China: Preliminary Results, Intent To Rescind, in Part, and Rescission, in Part, of Antidumping Duty Administrative Review; 2023-2024

The U.S. Department of Commerce (Commerce) preliminarily determines that exporters subject to this review made sales of subject merchandise at less than normal value (NV) during the period of review (POR), December 1, 2023, through November 30, 2024. In addition, we are rescinding the review with respect to 16 companies and intend to rescind this review with respect to Maodi Solar Technology (Dongguan) Co., Ltd. Interested parties are invited to comment on these preliminary results of review.

Implementing Regulation for National Environmental Policy Act (NEPA): Environmental Effects of the Department of Veterans Affairs Actions

The Department of Veterans Affairs (VA) is issuing this interim final rule to amend its agency procedures for implementing the requirements of the National Environmental Policy Act (NEPA). Since VA last updated its NEPA regulations in 1989, Congress amended NEPA through the Fiscal Responsibility Act of 2023 and the One Big Beautiful Bill Act of 2025, the Council on Environmental Quality rescinded its NEPA regulations, and substantial changes have occurred in VA's delivery of care and benefits to veterans. The revisions to VA's NEPA regulations improve the efficiency and quality of VA's NEPA process and align the NEPA process with decision-making across VA by more clearly focusing on the planning stages of VA actions, improving consistency in NEPA implementation throughout VA, updating the VA categorical exclusion list to reflect current VA activities, and complying with NEPA, as revised.

DeepSeek cut prices 75%. The 100x problem remains

12 July 2026 at 16:00

DeepSeek's recent decision to drastically cut pricing on its V4-Pro model by 75% should have been unequivocally good news for enterprise AI vendors and developers. Instead, many are discovering that cheaper models don’t automatically translate into healthier margins.

The reason is simple: While inference costs plummet, agent systems are voraciously consuming tokens faster than prices are declining. For the last 2 decades, software economics was dictated by the same rule. Infra became cheaper every year whereas applications became more capable. AI was initially hypothesized to follow the same pattern. As frontier models improved and token prices dropped, many assumed inference would become a negligible operating expense.That assumption has begun crumbling exponentially. 

A chatbot usually turns one user question into one model call. An agent turns it into a chain of planning, retrieval, tool use, verification, summarization, and follow-up decisions. The user sees one answer. The vendor pays for the loop. That is the 100x problem: The same user-visible request can cost a lot  more to serve as an agentic workflow than as a chatbot or retrieval-augmented generation (RAG) response. In longer-running workflows, the multiplier is higher. Falling model prices help, but they do not fix a product architecture that turns one prompt into dozens of billable operations.

The scale of what is now at stake is clear in how model providers themselves are pricing developer relationships. OpenAI's proposed program to give every Y Combinator startup $2 million in API credits — a number that would have funded an entire seed round in any prior tech cycle, and when the same cohort got by on a few thousand dollars of AWS credits — is less a recruiting perk than an admission of what it now costs to run an AI-native company through its first year of product. For established enterprises retrofitting agents into existing product lines, the absolute numbers are larger still.

What token amplification is

In a single-turn chatbot, one user message produces roughly one model call. Input-to-billed ratio is about 1:5.

In a multi-step agent rolled out across customer support, sales operations, finance, legal review, and engineering, that ratio routinely lands at 1:700 or higher. Every loop iteration carries forward the cumulative conversation, tool outputs, and reasoning traces. Each step appends; nothing is dropped.

A "simple" agent query like “What did our top customer ask about last week?” typically touches seven priced operations before returning an answer:

  1. User prompt (~50 tokens)

  2. System prompt and tool definitions (~3,000 tokens, repeated on every call)

  3. Retrieval (~5,000 tokens of context)

  4. Model call #1 — tool selection (8,000 in / 200 out)

  5. Tool execution (~4,000 tokens returned)

  6. Model call #2 — summarization (12,000 in / 400 out)

  7. Model call #3 — follow-up decision (12,400 in / 100 out)

One sentence in, roughly 35,000 input tokens billed. Somewhere between $0.10 and $0.40 per query on a frontier model. Multiply that by a million queries a month — the table-stakes volume for any enterprise B2B feature — and the line item is six figures.

Why this breaks the existing AI business model

The dominant pricing story for enterprise AI has been seat-based SaaS: Pay per-user per-month, deliver agent capability, capture margin. That model assumes a reasonably bounded cost-per-user.

Token amplification breaks the assumption. A power user running 50 agent invocations a day on a $40/seat plan can cost more in inference than the plan charges. Token amplification shatters the traditional SaaS pricing model. When a power user’s daily agent activity costs more in inference than their monthly subscription fee, vendor gross margins turn negative, a paradox that compounds as customers deepen their agent adoption, the very usage curve vendors are selling to their boards. Several vendors are now privately reporting negative gross margins on heavy users, mirroring recent cloud expenditure reports from the Bessemer 'Supernova' cohort, where the correlation between AI-agent adoption and gross margin contraction has moved from a theoretical risk to a primary P&L headwind.

The visible symptoms have started leaking into public coverage. Bloomberg this week documented a widening gap between Salesforce's Agentforce marketing demos and the capabilities actually shipping to customers. This is the kind of gap that opens predictably when promised functionality is technically possible but uneconomical to serve at the price the seat plan implies. Salesforce is the most-watched case, not a unique one.

"For my team, the cost of compute is far beyond the costs of the employees." — Bryan Catanzaro, VP of Applied Deep Learning, Nvidia

The strategic implication is not "AI is expensive." It is that the dominant business model assumed by most AI-native company plans does not survive contact with agentic workloads.

A simple example

Consider an enterprise software vendor charging $40 per-user per-month for an AI-enabled support assistant. A traditional chatbot might cost only a few cents per user per day in inference, leaving healthy gross margins.

Now replace that chatbot with a fully agentic workflow capable of investigating tickets, querying internal systems, drafting responses, validating outputs, and escalating exceptions. If a heavy user executes 50 to 100 agent requests per day, inference consumption can increase by an order of magnitude. What was once a negligible infrastructure cost becomes a material operating expense.

This creates an unusual dynamic: The customers receiving the most value from the product are often the customers generating the highest inference costs. In extreme cases, vendors can find themselves with their most engaged users contributing the least profit. The result is a growing realization across enterprise software that agent adoption and margin expansion are no longer automatically aligned.

Agent orchestration is the new moat

The technical responses are known and converging. They are not novel, but they are critical for survival

  • Cost-aware routing: This technique involves a small classifier model that decides which tier (Haiku, Sonnet, Opus equivalents) handles each query. Well-tuned routers cut inference bills by around 60% without any degradation in quality

  • Prompt caching: Anthropic, OpenAI, and Google now offer 75 to 90% discounts on cached prefixes. 

  • Context discipline: You can truncate tool outputs, prune reasoning traces, and cap tool depth to prevent your agent from going down a rabbit hole

  • Speculative decoding: for self-hosted deployments, this technique guarantees 2 to 3X effective throughput on the same GPUs.

"Organizations using orchestration-led governance report stronger productivity gains — a holistic orchestration layer is associated with six times greater productivity impact than compliance‑only approaches" — IBM

The companies building this layer well are starting to look less like microservice operators and more like financial trading systems: Every routing decision priced, every path with its own P&L, every tenant on a metered budget.

What enterprise leaders should actually do

Four moves separate the companies that will still have margin in 24 months from the ones that won't:

  1. Make inference cost a first-class metric. Track it per-feature, per-tenant, per-query class the same way cloud cost was tracked starting in the mid-2010s.

  2. Budget like a media buyer. Set cost-per-thousand-queries ceilings per feature. Cap them. Alert on overruns. Engineering will not enforce this on its own.

  3. Treat the router as core infrastructure, not an optimization. It is the new load balancer.

  4. Audit prompts quarterly. A 4,000-token system prompt that grew organically over six months is a six-figure bill in slow motion. Most teams have never read their own production prompts end to end.

  5. Negotiate volume commits early. Frontier-model vendors now offer reserved-instance-style prepaid commits at substantial discounts. List price is the worst price any enterprise will ever pay.

The next 24 months

The structural shift underneath agentic AI is not that it is expensive. As DeepSeek's price cut today underscores, frontier inference unit costs are dropping roughly 3X per year, and the curve is not slowing.

The shift is that amplification is outrunning the price cuts. Cutting per-token costs 75% does not help a company whose agents are doing 700X more tokens per user query than its pricing model assumed. For the first time since the cloud era began, architecture decisions are again financial decisions in real time. A prompt redesign is a margin event. A poorly bound agent loop is an outage with a credit card attached.

The companies that survive the next 24 months of AI infrastructure pricing will not be the ones running the cheapest model. They will be the ones whose agents are smart and know what they cost to think.

That is the 100X problem. And it is arriving faster than the price cuts can hide it.

Maitreyi Chatterjee is a senior software engineer at a big tech company.

Devansh Agarwal works as an ML engineer at a leading tech company.

Anthropic’s newest enterprise partner is training 20,000 people on Claude — here’s the shift it signals

The clearest signal of a major pivot in enterprise AI came this week when Anthropic announced its second Global Premier Partner in the Claude Partner Network: UST.

Anthropic’s partnership with UST, an AI and technology transformation organization, is expected to improve the ability of UST to guide enterprise customers beyond proof-of-concept AI projects and into production-scale deployments.

Moving an AI pilot out of the sandbox and into a production-grade enterprise system is notoriously difficult, especially when every development team is building on a different large language model. The next phase of enterprise AI is the standardization of the stack.

This shift pulls model selection away from developers and hands it to enterprise platform teams, changing how engineering workflows will operate in the near future. As systems integrators increasingly embed a single model into the platforms they build and manage, AI selection is expected to become an architectural decision rather than an individual developer’s choice. The near-future reality might be that the model will become part of the stack itself, selected once at the platform level and inherited by every engineering team that relies on it.

This shift pulls model selection away from developers and hands it to enterprise platform teams, changing how engineering workflows will operate in the near future.

Standardizing the AI stack

As part of the agreement, UST will incorporate Claude into the engineering platforms and workflows it develops and operates for customers.

“Our alliance with Anthropic reflects UST’s unwavering commitment to helping clients navigate the AI landscape with confidence and achieve meaningful business outcomes,” said Krishna Sudheendra, CEO of UST.

“By combining the capabilities of Claude with UST’s engineering, industry knowledge, and delivery expertise, we are bringing to market industry-specific platforms and digital and engineering solutions that improve productivity, accelerate business outcomes, and help clients operationalize AI-led decisions in a safe and secure environment.”

Claude inside engineering platforms

One example of the coming standardization is UST’s integration of Claude into its engineering platforms, which are used by companies in the semiconductor, telecommunications, manufacturing, automotive, embedded systems, and IoT industries for design verification, chip validation, factory operations, and field service.

By using Claude, teams are expected to catch design flaws earlier, speed up chip validation, and integrate hardware and software into a single system, effectively laying the foundation for physical AI.

UST points to its UST-iDEC platform as an early example. The hardware and silicon validation platform already automates much of the validation process, which the company says reduces cycle times by up to 70% and halves typical turnaround times. By including Claude in the pipeline, UST aims to give the system more advanced reasoning capabilities rather than treating AI as a standalone assistant.

Claude Code now natively reads chip pinouts and hardware schematics to automatically write and execute regression tests that engineers previously had to script by hand. Concurrently, Claude’s reasoning models evaluate live edge data against digital twins to identify firmware regressions and signal-integrity faults. By uniting these capabilities, UST is accelerating an already-fast validation pipeline through less manual scripting and earlier fault detection.

Training 20,000 technical associates

Standardizing an AI stack requires aligning the workforce behind it. A central part of the alliance is UST’s commitment to training 20,000 developers and technical experts. Those associates will be certified on Claude across roles worldwide, including architects, engineers, consultants, industry specialists, and forward-deployed engineers who work directly alongside client teams.

“UST helps the world’s banks, telecoms, and manufacturers put new technology to work,” said Paul Smith, Chief Commercial Officer at Anthropic, in a statement. “They’re proving Claude inside their own engineering first, training 20,000 of their own people on it, before bringing it into the systems they build and run for clients.”

“They’re proving Claude inside their own engineering first, training 20,000 of their own people on it, before bringing it into the systems they build and run for clients.”

For engineering organizations, that level of standardization changes more than procurement. It reshapes day-to-day development. Shared AI workflows become reusable across teams, governance policies can be enforced centrally, and integrations with internal systems no longer need to be recreated for every project. The trade-off is that developers gain uniformity while giving up some freedom to choose whichever model they personally prefer.

Enterprise workflows beyond hardware

Outside physical AI, Anthropic has announced that UST is putting Claude to work by integrating it into selected industry and horizontal enterprise platforms.

In healthcare, UST’s CarePath uses Claude Code and MCP connectors to simplify member services and claims, routing recommended actions through an agentic layer for human approval. For telecom, UST IntelliOps introduces Claude’s reasoning into network operations to predict RAN failures and reduce the time NOC teams spend sorting signal from noise. Meanwhile, in the banking sector, UST FinX uses Claude to accelerate onboarding and automate document processing, providing staff with faster access to account data while maintaining built-in governance and audit controls.

“We are wiring Claude into how UST designs, builds, and runs solutions across our consulting, platforms, engineering services, and industry offerings,” said Manu Gopinath, President of UST. “This alliance with Anthropic helps us deliver higher-value outcomes for clients as advancing UST’s transformation into an AI-native organization.”

“We are wiring Claude into how UST designs, builds, and runs solutions across our consulting, platforms, engineering services, and industry offerings.”

By acquiring firsthand experience with the operational, technical, and change management challenges of AI adoption internally, UST is building an operating playbook of tested workflows. For enterprise organizations, the takeaway is clear: The future of AI relies on standardizing the stack and moving AI selection out of the sandbox and into the platform layer.

As more systems integrators adopt this approach, developers will increasingly inherit the AI stack their organization has already chosen.

The post Anthropic’s newest enterprise partner is training 20,000 people on Claude — here’s the shift it signals appeared first on The New Stack.

❌