❌

Normal view

Four of five enterprises that secured AI agent identities still can't contain one that goes rogue

Visa's president of technology, Rajat Taneja, walked the VB Transform 2026 audience through aiming Anthropic's Mythos at Visa's own payment network. The model stitched minor weaknesses into working exploit chains, and Visa open-sourced the harness that governed the hunt.

That's what it looks like when an enterprise has the engineering depth to act on what it finds. Most don't get there. Just over half, or 53%, of enterprises have already had an agentic security incident or near-miss. Sixty-five percent enforce agent permissions at runtime, yet only 18% isolate their highest-risk agents, and just 8% pair enforcement with isolation.

Leaning on provider-native controls to do the heavy lifting of agentic security just exacerbates that gap. The July wave of VentureBeat Pulse Research found that 92% of enterprises naming a primary security layer default to their hyperscalers and AI platform providers.

Six waves of research have been completed since January, surveying 440 qualified enterprise security respondents. The key takeaway: the containment gap between what enterprises need and what's getting done is growing wider, often unaddressed by enterprises whose agentic AI investments and futures are at risk.

The satisfaction data doesn't match the incident data

The research keeps showing enterprises rating the tools they know best at a higher score, even if those tools failed them or delivered mediocre results. Three findings from the raw data cut against that instinct, and each one says something about how young this market still is.

The enterprises that got hit rate their tools higher than the ones that didn't

Last month’s survey found that 46 enterprises reported a confirmed incident or near-miss, then went on to rate their satisfaction with their security tooling. Their average satisfaction was 4.39 out of 5. 30 of the 55 enterprises who experienced no incidents rated their security tooling at 4.13. Enterprises are rewarding any tool that saves them from a breach with a trust premium.

It’s a sure sign of a nascent market when brand positioning, marketing, or other means of persuading enterprises get easily superseded by saving a customer from a breach. Near-misses outnumber confirmed incidents 2-to-1 in both June and July, which means enterprises are catching problems at the edge. That edge catch is being interpreted as validation of both the security strategy and the tools acquired. Evident through seven months of data is how quick enterprise security leaders are to trust a new tool that identifies an intrusion or breach and defeats it before it gains access. VentureBeat believes the rescue itself is doing the marketing. The 4.13 average among never-hit enterprises shows the other side of the same effect. Tools that have never been seen working earn less trust, not more.

VentureBeat also found that of the 17 enterprises isolating their highest-risk agents, the 14 that rated their tooling average 4.00. Enterprises that do not isolate rate it 4.35. The enterprises closest to real security are the least satisfied with their tools — that dissatisfaction is what drives them toward the kind of engineering effort Visa put in.

Four of five enterprises that solved identity did not build isolation

49%, or 57 of the 116 enterprises surveyed in July, gave each agent its own scoped, managed identity. Just a month earlier, VentureBeat's June wave recorded 32% of enterprises having assigned per-agent identities. July’s 17-point jump in one month is the fastest single-month move this series has recorded. Despite these gains, 63% still report credential sharing somewhere in the fleet. Only 11 of those 57 also isolate.

That ratio explains why the containment gap keeps widening even as every headline control improves. Enterprises are treating identity and isolation as substitutes. They need to see the longer-term vision of each being integral to a platform-based, layered strategy. Two incidents VentureBeat has covered show why that distinction matters. A rogue AI agent at Meta passed every identity check before its March exposure was contained. And CrowdStrike CEO George Kurtz disclosed, at his RSAC 2026 keynote, a Fortune 50 agent that rewrote its own security policy using valid credentials. Giving an agent scoped credentials does not bound the blast radius when those credentials are misused. Sandboxing does.

The enforce-without-isolate population has a 58% incident rate

Fifty-three enterprises in July’s survey enforce scoped permissions at runtime but do not isolate. 31 of those 53 have already had an agent security incident or near-miss. That is 58%, five points above the 53% sample average. The enterprises living inside the containment gap are getting hit more often than the enterprises outside it.

Amy Chang, Cisco's head of AI threat intelligence and security research, presented findings on the Transform agentic security panel showing that when Cisco ran 6,986 multi-turn attacks against 15 flagship models, attackers who adapted across the conversation broke through up to 88.3% of the time. Single-turn red-teaming missed it. An adaptive attacker who defeats the guardrails lands inside whatever architecture sits behind them, and for 53 of the enterprises in this data, that architecture enforces but does not contain.

VentureBeat's Q1 Pulse Research tracked the same structural weakness earlier this year. Unauthorized tool or data access ranked as the most feared failure mode in every Q1 survey, growing from 42% in January to 50% in March. The April-May survey found only 4% of enterprises comfortable relying on model guardrails alone. Enterprises predicted they needed external controls, choosing to build enforcement over containment.

Enterprises built enforcement 35 points ahead of forecast. Isolation barely moved

The April-May survey asked 109 enterprises how they expected agent behavior to be controlled by the end of 2026, and 30% predicted runtime enforcement, 14% sandboxed execution, and 32% model-level guardrails. By July, 65% had built enforcement, more than double the prediction, while isolation reached 18%, roughly the rate they said it would. Enterprises built what was easy at twice the forecast and built what was hard at roughly the forecast. The April question asked for the primary control mechanism, single-select, while July's posture question allowed multiple selections, so the comparison is directional rather than exact.

Provider lock-in accelerated across all three quarters

Provider-native platforms already led usage in April-May, named by seven in ten enterprises describing their tooling. By June, 82% called one their primary agent security layer, and by July that share reached 92%, with OpenAI's guardrails leading at 44%, Microsoft Azure at 42%, Anthropic's managed-agent controls at 37%, and Google Cloud at 31%. Cloudflare at 11% and Cisco at 9% lead the dedicated specialists fighting over what remains. The identity tools most relevant to the credential-sharing gap are the smallest of all, with Microsoft Entra Agent ID at 7%, while Okta for AI Agents, non-human identity platforms, and runtime sandboxing tooling each sit at 3%. CrowdStrike CTO Elia Zaitsev told VentureBeat at RSAC 2026 that observing agent actions is a solvable problem but inferring intent is not. The provider bundle proves his point, solving observation while leaving containment unbuilt.

74% plan to replace tools they just rated a career-high satisfaction score

Satisfaction scores continue rising as enterprises gain more experience using tools and techniques to stop agentic AI-based attacks. Rising to 4.29 out of 5 in July from 4.2 in June, satisfaction is the highest reading in the series.

Despite the high satisfaction levels, 74% plan to replace their tools within 12 months, up from 59% in June. Only 26% intend not to change. VentureBeat believes early adopters are impatient to gain greater insights, and know what they don’t know about agentic security and resilience. Closing that knowledge gap is forcing churn into a market this young, and the raw answers resolve the paradox: 92% of enterprises naming a primary layer name a provider-native one. The 4.29 measures how easy it is to turn on a provider's guardrails. It does not measure how effective those guardrails are at preventing the incidents 53% of the same respondents already had.

The organizations closest to the threat are the least confident about it

In June, defenders led attackers 35% to 21%, but by July the split was 30-30, a dead heat. Among enterprises that have been hit, 39% now say attackers are ahead, against 20% of those that have not. Getting hit nearly doubles the pessimism but does not change the shopping. Just 10% of enterprises include any agent-identity product in their consideration set. Runtime sandboxing draws 6%, and those numbers hold regardless of incident history. VentureBeat covered the same blind spot in the June data. The label changed from agent security gap to containment gap, but the shopping did not.

Methodology

The posture question was answered by 93 of the 116 qualified July respondents, and the skippers are not hidden isolators. Twenty-three of the 25 who selected no posture option are organizations still evaluating agents, unsure of their status, or with no deployment plans, groups for which a security posture largely does not yet exist, so the 18% isolation figure reads on the enterprises actually running or piloting agents. April-May, June, and July are separate, independently fielded waves rather than a single tracked series, so month-over-month comparisons in this piece are directional rather than a measured trend. Base sizes for the cross-cuts differ by instrument. The identity question covers all 116 respondents, isolation covers the 93 who described a posture, and the satisfaction inversion of 4.39 versus 4.13 is computed on the 76 respondents who rated their tooling.

The bottom line

VentureBeat's cross-survey analysis of 573 enterprise respondents concluded in July that enterprises deployed AI agents ahead of the controls needed to manage them, and they did it knowingly. Three waves of security-specific data now show where the knowing stops.

Enterprises continue giving agents scoped identities and treating that as containment, but that assumption is false, and the incident data keeps proving it. In fact, 46 of 57 enterprises that solved identity did not build isolation. The enforce-without-isolate population's 58% incident rate is the clearest evidence that identity alone isn't enough. The containment gap will not close through satisfaction with what is easy. Whether enterprises build isolation and governed identity deliberately, or whether a confirmed incident that propagates does it for them, is the question the next wave will answer.

SpaceXAI debuts Grok 4.6, overtaking Kimi K3's performance and matching GPT-5.6 Sol for world's third best on Artificial Analysis

Elon Musk's company SpaceXAI, formerly known as xAI, has released Grok 4.6, its latest frontier AI model, with a focus on long-running agents, coding and knowledge work — and a pricing strategy designed to make those workloads cheaper to run.

The model scores 61 on the third-party Artificial Analysis Intelligence Index, surpassing the popular open weights Chinese model from Moonshot, Kimi K3, and tying rival OpenAI's GPT-5.6 Sol Max and improving five points over Grok 4.5 High. Anthropic's Claude Opus 5 and Fable 5 occupy the number one and two spots, respectively.

More consequential for enterprises evaluating AI agents, Grok 4.6 posts sizable gains over its predecessor across coding, terminal, knowledge-work and agent benchmarks while retaining an application programming interface (API) price starting at $2 per million input tokens and $6 per million output tokens, making it a mid-priced frontier model comparing leading options that are both proprietary and open source, globally, according to VentureBeat's analysis.

Model

Input ($/1M)

Output ($/1M)

Total ($/1M)

Source

Muse Spark 1.2 Contributor

$0.10

$0.20

$0.30

Meta

MiMo-V2.5 Flash

$0.10

$0.30

$0.40

Xiaomi

deepseek-v4-flash

$0.14

$0.28

$0.42

DeepSeek

deepseek-v4-pro

$0.435

$0.87

$1.305

DeepSeek

GPT-5.6 Luna

$0.20

$1.20

$1.40

OpenAI

MiniMax-M3

$0.30

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

$0.30

$1.20

$1.50

LongCat

MiMo-V2.5

$0.40

$2.00

$2.40

Xiaomi

LongCat-2.0 — standard

$0.75

$2.95

$3.70

LongCat

MiMo-V2.5 Pro (≤256K)

$1.00

$3.00

$4.00

Xiaomi

Muse Spark 1.1 / 1.2

$1.25

$4.25

$5.50

Meta

GLM-5.2

$1.40

$4.40

$5.80

Z.ai

Grok 4.6 — <200K prompt tokens

$2.00

$6.00

$8.00

xAI

MiMo-V2.5 Pro (>256K)

$2.00

$6.00

$8.00

Xiaomi

Qwen3.8-Max

$2.00

$6.00

$8.00

QwenCloud

Gemini 3.6 Flash

$1.50

$7.50

$9.00

Google

GPT-5.6 Terra

$2.00

$12.00

$14.00

OpenAI

Grok 4.6 — ≥200K prompt tokens

$4.00

$12.00

$16.00

xAI

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Kimi K3

$3.00

$15.00

$18.00

Moonshot AI

Claude Opus 5

$5.00

$25.00

$30.00

Anthropic

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Sakana AI

GPT-5.6 Sol — Standard mode

$5.00

$30.00

$35.00

OpenAI

Claude Fable 5 / Claude Mythos 5

$10.00

$50.00

$60.00

Anthropic

GPT-5.6 Sol — Fast mode

$10.00

$60.00

$70.00

OpenAI

Still, that's less than half of what GPT-5.6 Sol costs over OpenAI's API in standard mode.

SpaceXAI says Grok 4.6 is available today in Grok Build, SpaceXAI's answer to Anthropic's Claude Code and OpenAI's Codex, which is available starting in the $30 per month SuperGrok plan.

It's also available in SpaceX's recent acquisition of the AI coding startup Cursor, and from partners including OpenRouter, Vercel and Cloudflare.

SpaceXAI is providing twice the included usage for Grok 4.6 in Cursor and Grok Build during the first week.

The release arrives only weeks after Grok 4.5, which SpaceXAI launched in July as a model targeting coding, agentic tasks and knowledge work, and one day after the launch of Grok Bot, a new system for assigning AI agents to complete designated tasks as virtual employees.

The bigger change is agent behavior, not just another benchmark point

SpaceXAI describes Grok 4.6 as being built specifically to stay on task across longer sequences of work, including researching unfamiliar topics, analyzing information, navigating codebases and converting product ideas into working applications.

The company says it subjected the model to a longer supplemental training run than Grok 4.5, using curated model-generated reasoning and technical data alongside engineering data and changes to its optimizer and training recipe. It then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning levels, agent harnesses, STEM, software engineering and knowledge work, filtering problematic trajectories with model-based checks.

Reinforcement learning also targeted agentic environments spanning general coding, knowledge work, kernel optimization, web development and computer-aided design.

That matters because enterprise AI deployments are increasingly moving beyond isolated prompt-and-response interactions toward agents expected to maintain state, operate tools, modify code and recover from problems across longer execution paths.

SpaceXAI says that during its testing, Grok 4.6 showed more self-testing and verification on longer trajectories, checking its own work before proceeding. It also reports stronger first attempts on interactive and visual projects than Grok 4.5. Those are company observations rather than independent guarantees of production behavior, but they indicate where SpaceXAI concentrated the model’s post-training work.

Grok 4.6 reaches the frontier, but does not sweep it

Grok 4.6's improvement over Grok 4.5 at this juncture of the AI model competition cannot be overstated.

According to Artificial Analysis, Grok 4.6 reaches an Elo score (human preference of head-to-head model outputs, adapted from chess) of 1,753 on GDPVal-AA v2, the benchmark measuring performance on real-world tasks like scheduling and diagramming, versus 1,526 for Grok 4.5, 1,728 for GPT-5.6 Sol Max and 1,741 for Fable 5 Max.

The coding results from SpaceXAI show a similar generational improvement but more competition at the frontier.

Grok 4.6 scores 69.9% on CursorBench v3.2, up from 66.7%, while Fable 5 Max reaches 70.5%. On DeepSWE v1.1, Grok rises sharply from 54% to 65.9%, but GPT-5.6 Sol Max leads at 73%. FrontierCode v1.1 Extended moves from 56.6% to 61.3%, compared with 60.6% for GPT-5.6 Sol Max and a leading 63.6% for Fable 5 Max.

Agent benchmarks tell much the same story. Grok 4.6 reaches 57.5% on APEX-Agents, a 10.4-point increase over Grok 4.5’s 47.1%, narrowly exceeding GPT-5.6 Sol Max’s 56.7% but trailing Fable 5 Max at 59.2%. On APEX-SWE, Grok 4.6 rises to 56.4% from 53.6%, while Fable 5 Max scores 58.8%.

Terminal-Bench v3.0 exposes a larger remaining gap. Grok 4.6 improves from 15.7% to 26%, but GPT-5.6 Sol Max and Fable 5 Max score 34.6% and 34.1%, respectively.

Two of Grok 4.6’s strongest results come from longer-horizon professional work. On AA-Briefcase it scores an Elo of 1,577, narrowly exceeding Fable 5 Max’s 1,574 and topping GPT-5.6 Sol Max’s 1,502. On Harvey LAB, Grok 4.6 reaches 15.8%, versus 12.9% for Grok 4.5, 11.3% for Fable 5 Max and 2.5% for GPT-5.6 Sol Max.

SpaceXAI notes an important methodological caveat: third-party scores in its table use the best self-reported or publicly available results. The comparison therefore should not be interpreted as a perfectly controlled four-model evaluation.

In other words, the evidence supports a substantial upgrade over Grok 4.5 more clearly than it supports across-the-board superiority over rival frontier models. Grok 4.6 wins several of the displayed evaluations while GPT-5.6 Sol Max and Fable 5 Max retain meaningful leads elsewhere.

Cost could be the more important enterprise benchmark

Artificial Analysis’ supplied evaluation adds another dimension: how much work the model performs for the money spent.

The testing places Grok 4.6 on its Intelligence-versus-Cost-per-Task Pareto frontier at a reported $0.84 per task — which actually makes it less of a bargain than its predecessor, Grok 4.5, and less economical than OpenAI's GPT-5.6 Luna, z.ai's GLM-5.2, and Meta's new Muse Spark 1.2, among other models.

Artificial Analysis also reports that Grok 4.6 completed its AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens on average, versus approximately 103 turns and 2 billion input tokens for Claude Opus 5 Max.

Those measurements do not prove that every production agent will use fewer tokens or finish twice as quickly. Agent costs depend heavily on harness design, prompts, tool calls, caching, retry behavior and the task itself. But they point toward an increasingly important enterprise metric: the cost of completing a workflow, rather than simply the cost of generating one million tokens. That distinction is central to SpaceXAI’s positioning.

The standard Grok 4.6 API starts at $2 per million input tokens and $6 per million output tokens, and SpaceXAI also offers a faster variant at twice the price.

The supplied API documentation adds an important caveat for long-context deployments. Grok 4.6 supports a 500,000-token context window, but prompts below 200,000 tokens are billed at $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens. Once a prompt reaches 200,000 tokens, those rates rise to $4, $1 and $12 respectively, with the higher pricing applying to all tokens in that request.

That means enterprises should not extrapolate the $2/$6 headline pricing across the model’s entire context window when estimating total cost of ownership.

Artificial Analysis says the standard headline rates remain more than 60% below the competing frontier-model prices it cites for Claude Opus 5 and GPT-5.6 Sol. The practical savings will depend on how many tokens each model consumes to complete the same workload.

The Grok name carries considerable baggage and controversy, separate from the general AI skepticism

Performance and price may not be the only hurdles SpaceXAI faces in converting Grok 4.6's benchmark gains into enterprise adoption. The Grok brand arrives with an unusually visible history of safety and governance controversies — including extremist and antisemitic outputs, politically skewed responses, exaggerated praise of Elon Musk and, more recently, the use of Grok's image-generation capabilities to produce non-consensual sexualized imagery. For companies with strict compliance, brand-safety or responsible-AI requirements, that history could become a procurement consideration separate from the technical capabilities of Grok 4.6 itself.

The most notorious text-generation episode came in July 2025, when Grok produced antisemitic posts, praised Adolf Hitler and in some responses referred to itself as "MechaHitler." SpaceXAI's predecessor xAI subsequently said it was removing inappropriate posts and taking steps to prevent hate speech from being published by Grok.

Also in summer 2025, Grok began inserting references to an alleged "white genocide" in South Africa into answers to unrelated questions. xAI said an unauthorized modification to Grok's response software had directed the system to produce a particular response on a political topic while bypassing its normal review process. The company said the change violated its policies and subsequently pledged to publish Grok's system prompts and establish round-the-clock monitoring for problematic responses. The South African government has rejected claims that a genocide against white South Africans is taking place.

Grok's objectivity came under scrutiny again in November of the same year after the chatbot repeatedly produced implausibly flattering assessments of Musk. Among the examples reported at the time were claims placing Musk above elite athletes and historic intellectual figures. Musk said Grok had been manipulated through adversarial prompting into making "absurdly positive" statements about him. Whatever the underlying cause, the incident illustrated the reputational problem for an enterprise model whose outputs can become entangled with the public persona of the executive most closely associated with its developer.

The most serious controversy has involved image generation. In January 2026, U.K. regulator Ofcom opened a formal investigation into X after reports that the Grok account was being used to create and distribute undressed images of people and sexualized images of children. Ofcom said the material under examination could amount to non-consensual intimate-image abuse, pornography and child sexual abuse material. X subsequently said it had implemented measures intended to stop the Grok account from being used to create intimate images of people, but Ofcom said its investigation remained open.

The scrutiny extends beyond Ofcom. Britain's Information Commissioner's Office is investigating X and xAI over both the development and deployment of Grok, including whether personal data was handled lawfully and whether adequate safeguards existed to prevent harmful manipulated imagery. The European Commission, meanwhile, opened a separate formal investigation under the Digital Services Act examining X's management of systemic risks connected to Grok, including the dissemination of manipulated sexually explicit material.

Those investigations concern X and the earlier xAI organization rather than establishing a finding that the newly released Grok 4.6 API violates those laws. Nevertheless, they are unlikely to help SpaceXAI sell Grok to businesses.

SpaceX acquired xAI in February 2026, and the AI operation now markets itself as SpaceXAI, meaning Grok's newest models sit under a different corporate structure but retain the same consumer-facing brand.

There is no evidence in the material examined here that Grok 4.6 itself repeats the specific "MechaHitler," "white genocide," sexual-image or Musk-flattery incidents associated with earlier Grok deployments. But enterprise procurement teams rarely evaluate a model in isolation from its vendor and product history. For SpaceXAI, that means Grok 4.6 may have to demonstrate not only that it is cheaper or more capable than competing frontier models, but that the controls around it are sufficiently predictable for organizations that cannot afford their AI supplier to become a brand-safety event.

That continuity creates a potential adoption problem that benchmark tables cannot measure. Developers choosing a model for an internal coding agent may care primarily about price, latency and task completion. A bank, government agency, healthcare provider or consumer brand deploying the same model into customer-facing or regulated workflows may also have to consider vendor governance, content-safety controls, auditability and reputational exposure.

A model designed to be deployed, not just chatted with

Grok 4.6 supports text and image inputs with text output, function calling, structured outputs and reasoning, according to the supplied API specifications. Those specifications also list rate limits of 150 requests per second and 50 million tokens per minute, with API availability in us-east-1 and us-west-2.

Cursor’s launch announcement similarly characterizes Grok 4.6 as designed for long-running agents and ambitious interactive and visual work, giving developers immediate access to the model inside an established coding-agent environment rather than requiring them to build a new harness around the API first.

For enterprise buyers, that distribution may matter almost as much as another leaderboard result. Models increasingly compete not just on reasoning scores but on whether developers can place them inside existing coding, research and operational workflows without destabilizing those workflows or dramatically increasing inference costs, as well as incurring any blowback from associating with a controversial brand.

Grok 4.6 does not establish an uncontested performance lead. Its launch instead presents a different proposition: frontier-level intelligence, large improvements over the previous generation, stronger long-running agent behavior and relatively aggressive token economics.

The next test will be whether the efficiency Artificial Analysis observes on controlled agentic workloads carries into production. If Grok 4.6 can consistently complete long-running coding and knowledge-work tasks with fewer turns and fewer tokens, the model’s most important benchmark may ultimately be the enterprise inference bill rather than the leaderboard.

Skan AI raises $63 million betting that watching how employees actually work is the missing layer of enterprise AI

Skan AI, a startup that builds what it calls a "context graph of work" by observing how employees actually perform their jobs across enterprise software, has raised $63 million in Series C funding co-led by Cathay Innovation and Dell Technologies Capital, the company announced Wednesday.

Citi Ventures, Bloomberg Beta, State Farm Ventures, and Wipro Ventures also participated in the round, which brings the seven-year-old company's total funding to roughly $120 million. Alongside the raise, Skan is announcing the general availability of two new products — Skan AI Blueprint and Skan AI Agents — that, together with its existing Skan AI Intelligence offering, form a complete platform for discovering, modeling, and ultimately automating enterprise workflows.

The announcement lands at a moment of deep frustration in enterprise AI. Companies have poured billions into generative AI pilots, but the results have been dismal: Gartner research cited by the company finds that only 8% of enterprises have AI agents in production, and 95% of early implementations will require a complete redesign. Those figures echo an MIT report last year, covered by Fortune, which found that roughly 95% of enterprise generative AI pilots were failing to deliver measurable returns.

Avinash Misra, Skan's co-founder and CEO, believes the industry has misdiagnosed the problem. The models are fine, he argues. What they lack is an accurate picture of the businesses they are being dropped into.

"Everyone is obsessed with building a better driver," Misra told VentureBeat in an exclusive interview ahead of the announcement. "We think the bigger opportunity is building a better navigation system."

Why enterprise AI agents keep failing when they rely on official process documentation

The standard playbook for grounding AI agents — feeding them process documentation, standard operating procedures, and system logs — is built on a fiction, Misra argues. The way work is documented and the way work actually happens inside a large enterprise are two different things, and the gap between them is precisely where agents fail.

That gap is what sent Misra and co-founder Manish Garg down this path seven years ago, long before agents were a boardroom obsession. "Why is it so difficult for an organization, and a large enterprise especially, to understand how its own work actually gets done?" Misra said. "Why does it need to fly in McKinsey consultants for that?"

The question has only grown more consequential as enterprises race to operationalize AI. Frontier models arrive at the company door brilliant but blind, with no knowledge of the exceptions, decisions, handoffs, and institutional habits that define how a claims department or a compliance team actually operates. Every company now stuffing agents with documentation and logs, Skan contends, is discovering the same uncomfortable truth: the source data was never the whole story. And a source data problem cannot be fixed downstream.

Skan's answer is to go to the source itself. The company deploys observation technology on employee desktops that continuously watches how work moves across applications — the spreadsheet, the CRM, the email client, the 40-year-old mainframe — and abstracts those observations into a living model of the underlying business process.

"Think of it this way: if I were to share my screen here, and you were to observe my screen going from Excel sheet, CRM system, email client, in about two iterations you'd build a model of what I do," Misra said. "Except you couldn't do that at scale. You couldn't do it 24/7, and for 1,500 people like me. Now replace yourself with our technology."

How screen-level observation captures the work that never shows up in system logs

That framing also explains how Skan positions itself against process mining vendors like Celonis, which reconstruct workflows from the data trails left in backend systems. System logs, Misra argues, only capture completed transactions — not the messy human work that produced them.

"All backend data, by definition, is a committed state of work. Work is really what happens between those committed states," he said. "Eighty percent of what you're interested in, from an AI point of view, in execution of work, actually lies between those systems."

The screen, in Skan's view, is the one place where everything converges. "It brings together human agency, it brings together the entire application landscape, and it brings together the data that matters," Misra said. Two decades of user interface design have quietly buried enormous amounts of process knowledge in the space between a worker's eyes and their monitor; Skan's pitch is to bring that hidden layer back to the surface.

But watching, he insists, was never the hard part — a point aimed squarely at the incumbents who might be tempted to copy the approach. "The hard problem is not screen observation," Misra said. "The hard problem is abstraction of what you see on the screen — the intent extraction." A human watching a colleague's screen can instantly tell whether a jump back to step one means a new case or rework on an old one, because humans understand the signature of the work. Teaching a model to make that same judgment, statefully and at enterprise scale, is where Skan believes its seven-year head start lives.

The result is a context model that AI can reason over and act on — the raw material for the agents that now sit at the top of the company's product stack, and the foundation for everything else the platform does.

Walking the line between operational telemetry and workplace surveillance

An approach built on continuously watching employee screens invites an obvious objection, and it is not a hypothetical one. In June, Reuters reported that Meta scaled back an internal tool that tracked employee mouse clicks after workers raised concerns — a sign that even AI-forward companies are wary of the line between operational telemetry and surveillance.

Misra says he heard the objection before he wrote a line of code. When he first pitched the concept to Delphine Icart, then chief transformation officer at AXA Mexico, her reaction was blunt. "Delphine's first words to me were, 'This sounds like a great idea, but you are dead on arrival,'" Misra recalled. "'You are observing things that you shouldn't be observing — the privacy of my operators, and the sovereignty of my data on those screens.'"

That conversation, he says, shaped the architecture. Skan aggregates rather than individuates: the system surfaces statistical patterns across hundreds of workers performing the same process, not the behavior of any one of them. "We're not interested in what John is doing at 10 hours and 43 seconds," Misra said. "We are interested in what hundreds of Johns put together — what are the statistical and the semantic decisions that they are making in that business process?"

Organizations control what the technology can see through an opt-in scoping model — specific applications and URLs, nothing else — and the data Skan produces never leaves the enterprise firewall. A three-tier architecture sends only anonymized metadata to the cloud. Misra points to deployments approved by European works councils, among the most privacy-protective labor bodies in the world, as evidence the model holds up under scrutiny — and credits it for clearing security review at institutions where most AI tools cannot operate.

Whether aggregation fully defuses the concern is likely to remain contested. The same telemetry that reveals a broken process can, in principle, reveal an underperforming team, and Misra acknowledged that the technology has led some customers to reduce headcount in certain processes.

What $500 million in claimed customer value actually measures

Skan claims more than $500 million in cumulative customer value to date, a figure worth unpacking. Pressed on whether that represents realized savings or projections, Misra was direct that it is an envelope, not a bank balance.

"The number comes from the cumulative, across all our customers, of the quantified savings that we have brought to them — the savings that they have expected they would save," he said. "Now they are on the roadmap of recouping those savings through a variety of interventions," including process redesign, technology changes, and, increasingly, AI agents. In other words, $500 million is identified opportunity, some portion of which has been captured.

The more concrete evidence comes from individual deployments. At one top U.S. bank, according to the company, Skan observed 11.2 million context switches across 1,500 finance professionals and uncovered $37 million in operational friction. Turning those observations into agent-executable context cut cost per transaction by 32%, lifted throughput by 41%, and delivered $18 million in annualized savings.

Misra pointed to an anti-money-laundering operation at one bank where "60% of the cases are now being run by AI agents," adding that the results surprised even him: "The accuracy of those agents surpasses many times over the accuracy of humans. It's not just an argument of efficiency; it has also become an argument of quality." Among insurers, he said, Skan typically delivers roughly 25% productivity uplift in core claims processes; one customer doubled its case volume over the past year without adding a single claims specialist.

Skan's publicly referenceable customers include Unum, the $13.8 billion employee benefits provider, and Mitie, the U.K. facilities management company, whose chief technology and digital officer, Cijo Joseph, said Skan's technology "gives us unprecedented operational visibility that has dramatically accelerated our AI transformation." The company declined to share revenue but said it grew more than 300% year over year — for the second consecutive year — with net dollar retention around 150%, and now counts seven of the ten largest U.S. banks and a quarter of the Fortune 50 as customers.

Can AI models learn good work from imperfect employees?

Skan's thesis rests on observing how work actually gets done — which raises an uncomfortable question. Real employees make mistakes, take shortcuts, and entrench inefficiencies. What happens when the context graph faithfully encodes bad process?

Misra's answer reaches for the most famous precedent in modern AI. "Think for a moment what OpenAI did," he said. "OpenAI took the totality of the world's text and fed it into a transformer architecture, and semantic understanding emerged. OpenAI's model has seen bad language and has seen good language, and yet it is able to have semantic understanding."

Skan, he argues, does the analogous thing with work: treat business process execution as a language, where process steps, screen features, and handoffs stand in for words and sentences. Fed enough end-to-end executions, the model learns the full distribution of paths — efficient ones, slow ones, compliant ones — without assuming any single path is best. "The longest path may be the best path, because it is more compliant," Misra said. An organization then constrains the model along the axes it cares about, and the model returns the path that satisfies them.

"It is not record and play — and that's the fundamental difference between us and a lot of our competition, UiPath and so on," he said. "It is fundamentally creating an AI model that understands work, and then constraining that model."

He offered a concrete illustration of what that unlocks: at one large bank, Skan's telemetry continuously compares live case execution against a 600-page controls inventory, with agents that trigger alerts when cases miss required compliance steps — turning a document no human could hold in their head into a real-time enforcement layer. It is the kind of application that only becomes possible, Misra argues, once a model genuinely understands the work rather than merely replaying it.

The race to own the context layer of enterprise AI

Skan sits at the intersection of several crowded categories, and its answer to each competitor is a variation on the same theme: scope. Process mining vendors see only what the logs record. RPA incumbents replay tasks without understanding them. And the platform giants — ServiceNow, Salesforce, Microsoft — are shipping capable agents whose vision ends at their own walls.

"The context that these agents have access to is limited to ServiceNow, limited to Salesforce, whereas work spans processes across the board," Misra said. "Creating a customer entry is a task. To receive an email and decide whether a customer entry has to be created, or something else — that is the process, and that's what we are after."

The deeper strategic argument, and the one that seems to resonate with Skan's regulated customer base, is about differentiation in a world where every enterprise has access to the same frontier models. "If every insurance company, every bank had access to the same models, then the outcomes will asymptotically decay to the outcome of the model," Misra said. "Historically, you have competed and differentiated in the way you have organized work. That old word — process — now comes back as context for AI. But that context is protected by you. It's not part of the model."

That logic explains both the company's posture toward the model makers — "the more they are successful, the more power we have," Misra said, disclaiming any ambition to compete with them — and the Nvidia partnership featured prominently in the announcement. Skan runs on Nvidia AI Enterprise and NIM microservices, and Misra described growing demand for private appliances that can observe work, hold the context model, and execute agents entirely inside a customer's own infrastructure. It also fits the market's direction: venture investors surveyed by TechCrunch at the end of last year predicted enterprises would spend more on AI in 2026 but through fewer vendors — a consolidation that favors Skan's decision to ship discovery, intelligence, and agents as a single closed loop.

Misra argues that loop matters more, not less, as automation scales, because agents demand oversight in a way humans never did. "It is an irony of sorts," he said, "that you'll probably need much more observation and much more understanding of work in an automated way than you would with humans." The bet embedded in this round is that work context becomes foundational infrastructure for enterprise AI the way CRM became the system of record for customers — a comparison Cathay Innovation partner Simon Wu made explicitly, calling Skan "one of the defining platform companies of the next decade."

Misra put the stakes more simply. "You cannot retrieve context that you do not capture," he said. "The battleground is shifting from the smartest model to knowing how your company actually works — because everyone will have access to the smartest model."

The frontier labs, in other words, can keep their arms race for the better driver. Skan just raised $63 million on the conviction that the money is in the map.

Infrastructure and compute: Enterprises are buying AI compute for speed while flying blind on what it costs

12 August 2026 at 07:30

Across 170 enterprises, AI infrastructure has moved decisively into production — two-thirds now run AI workloads live and three in 10 run them at scale — while the ability to account for what that infrastructure costs has not kept pace. Enterprises have quietly demoted cost in the buying decision: performance and GPU availability now outrank total cost of ownership, and reliability outranks price as the measure of success. That reordering is rational for teams under production pressure, but it lands on an uncomfortable fact — fewer than half can rigorously track what their AI compute costs, most GPUs still run at half capacity or less, and the next dollar is aimed at specialized clouds that fewer than one in twenty of them actually use.

This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how they buy and measure it, where the next investment is aimed, and — most revealingly — how well they can see the economics of the compute underneath it all.

This is an operational cohort. Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale, with only 4% not yet running AI workloads at all. That maturity shows in the stack: the average enterprise runs three infrastructure platforms, with OpenAI (49%), Google Gemini (48%), Microsoft Azure (47%), and Google Cloud (42%) all present in roughly half of them. Asked to name one primary platform, Azure leads at 26%.

The most consequential shift is in how enterprises decide. Integration with the existing cloud and data stack remains the top selection factor at 40%, but performance — latency and throughput — has climbed to second at 35%, and access to GPU availability to third at 24%, both ahead of total cost of ownership at 22%. The same ordering governs measurement: uptime and reliability is the primary success metric for 51% of enterprises and developer productivity for 39%, ahead of cost per million tokens at 31%. Enterprises under production pressure are buying and measuring for speed and availability, and have moved cost down the list.

That would be unremarkable if the economics were under control, but they're not. Among the 155 enterprises that operate their own GPUs, 69% report utilization of 50% or less and only 23% clear the halfway mark; 12% do not measure utilization at all. Fewer than half (47%) rigorously track what their AI compute costs and returns, and even among enterprises running AI in production at scale that figure only reaches 56%. Value for money is the weakest of three satisfaction scores at 3.87, against 4.14 for overall satisfaction — the softness landing precisely on the dimension hardest to judge without measurement.

The next round of spending points away from the current stack. AI-specialized clouds are the top planned evaluation area at 44% and carry the strongest net momentum of any infrastructure approach (+36), yet CoreWeave and Lambda each registers at 3.5% of current usage and the rest of the neocloud field sits below 3%. Non-Nvidia accelerators draw 39%. And 62% of enterprises intend to switch or add a provider within 12 months — though the consideration set is dominated by the same incumbents they already run.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series, this one focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=170; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single July 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends; all figures are drawn from the July fielding only. Several questions were multiple-select, so those shares can sum to more than 100%.

By organization size this wave reaches further up-market than the mid-market skew this series usually carries: 251–1,000 employees (28%) and 1,001–5,000 (25%) lead, with 10,001+ (19%), 101–250 (15%), and 5,001–10,000 (12%) filling out the rest — meaning 57% of respondents sit above 1,000 employees. By role it spans managers (48%), individual contributors (27%), the C-suite (12%), and VPs and directors (9%); on purchasing authority it is buyer-credible, with 39% final decision-makers and another 43% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 35%, followed by Manufacturing (14%), Financial Services (12%), and Healthcare/Life Sciences (9%).

At 170 respondents the sample is large enough to read directionally with reasonable confidence, but it should still be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively building and operating AI infrastructure rather than from the largest hyperscale operators.

Finding 1: Two-thirds are past the pilot

Three in 10 now run AI in production at scale

We asked where organizations sit in their AI deployment journey. This cohort has largely moved beyond experimentation.

Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale. Only 30% remain in proofs of concept and just 4% have not started. This is a materially more operational sample than this series has typically drawn, consistent with its up-market composition — 57% of respondents sit above 1,000 employees.

That maturity is the frame for everything that follows. The infrastructure decisions in this report are being made largely by organizations with production workloads and real bills, not by teams still sizing a pilot. It explains the reordering of buying criteria in Finding 5, where performance and availability displace cost — the priorities of teams running live systems. It also raises the stakes on Findings 6 and 7: an enterprise that cannot measure utilization or cost during experimentation has a planning problem, while one that cannot measure them in production at scale has an operating one.

Finding 2: The stack is hyperscaler-and-API, three platforms deep

The specialized GPU clouds still barely register

We asked which providers and platforms enterprises currently use to run their AI, and which one they treat as primary. The answer remains the incumbents — several of them at once.

The current stack is hyperscaler-and-API, and it is plural: enterprises name three platforms on average. The general-purpose clouds and the major model APIs account for essentially all current deployment, with four platforms — OpenAI, Gemini, Azure, and Google Cloud — each presents in more than four of every 10 enterprises. Asked to pick one primary platform, Microsoft Azure leads at 26%, with Google Cloud second at 19%; the model providers together take 35% of primary status when OpenAI (14%), Gemini (14%), and Anthropic (8%) are combined.

The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines remain marginal in practice. CoreWeave and Lambda each appear in 3.5% of stacks, Baseten in 3%, and Crusoe, Nebius, Fireworks, Together, and Anyscale each at or below 2%. Combined, they are named as the primary platform by 1% of enterprises. Meanwhile 13% run a custom open-source self-managed stack and 9% operate their own GPU clusters — both larger footprints than the entire specialized-cloud category. That contrast is what makes the evaluation intentions in Finding 3 worth reading closely.

A note on reading these shares: As described in the methodology section, this sample is self-selected and this question counted every provider a respondent uses — an average of 3.0 selections each — so the figures measure presence in the stack rather than spending or primary status. The separate primary-platform question is the better guide to where the center of gravity sits. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; read these shares as a portrait of what this AI-active cohort runs today, and treat gaps against industry-wide market share estimates as a property of the sample rather than a contradiction of either.

Finding 3: The next dollar goes to infrastructure they don't yet run

AI-specialized clouds top the evaluations list and carry the strongest momentum

We asked where enterprises plan to evaluate AI infrastructure over the next 12 months, and whether they expect to do more or less with each category of infrastructure. Both answers point away from the stack they run today.

Here is the report’s sharpest tension, and it is the same one this series has now recorded across successive waves. The single most-cited planned evaluation area — AI-specialized clouds, at 44% — is the category that 3.5% of these enterprises actually use (Finding 2). Nearly four in 10 (39%) intend to evaluate non-Nvidia accelerators, a quarter next-generation Nvidia silicon, and even decentralized compute networks draw 18%.

The direction-of-travel question corroborates it rather than merely repeating it. Asked whether they expect to do more, less, or about the same with each approach, enterprises put specialized AI clouds at the highest net momentum (+36, with 42% doing more against 6% doing less), ahead of inference APIs (+34) and hyperscalers (+30). On-prem and co-located infrastructure is the laggard at +5, the only category where a substantial share — 22% — report pulling back. Every off-premises approach is net-expanding; the specialized clouds are expanding fastest from the smallest base.

Read against current usage, this is not incremental adjustment. It is the leading edge of a re-platforming that enterprises have been signaling for several waves and have not yet executed. The gap between a 44% evaluation rate and a 3.5% usage rate is the single widest intent-to-action spread in this dataset, and how it resolves — whether the neoclouds convert evaluation into deployment, or whether the hyperscalers absorb the demand with their own AI infrastructure — is the open question of the category.

Finding 4: Six in 10 plan to move, mostly among the incumbents

High churn intent, but the consideration set is the stack they already run

We asked whether and when enterprises plan to switch or add an infrastructure provider, and which providers they are considering.

For a category as foundational as compute, this is a substantial amount of intended movement: 62% of enterprises intend to switch or add a provider within 12 months, and 29% within the next quarter alone. Only 39% plan to stand still.

Where that interest points is the more useful signal. The providers drawing the most switching consideration are the ones enterprises already run — OpenAI and Google Cloud (29% each), Microsoft Azure (28%), Gemini (25%), Anthropic (16%), Oracle Cloud (14%), and AWS (13%). The specialized clouds that top the evaluation list in Finding 3 draw far less concrete switching consideration: CoreWeave 4%, Lambda 3.5%, and the remainder at or below 2%. A further 8% are evaluating with no shortlist yet.

The two findings are not in conflict; they operate on different clocks. The neocloud interest in Finding 3 is a 12-month evaluation thesis about where AI compute should eventually run. The switching in the next quarter is mostly incumbents trading share and enterprises consolidating spend among providers they already hold contracts with. Vendors reading the 44% evaluation figure as near-term pipeline should weigh it against a 4% consideration rate.

Finding 5: Performance overtakes cost, in buying and in measurement

Total cost of ownership falls below latency and GPU availability

We asked what matters most when enterprises select an AI infrastructure provider, and what they treat as the primary measure of success once it is running. Both answers have moved away from price.

Integration with the existing stack remains the top selection factor at 40%, which is consistent with a cohort running three platforms and unwilling to add a fourth that does not fit. What has changed is everything below it. Performance sits second at 35% and GPU access and availability third at 24%, both ahead of total cost of ownership at 22%. Fine-grained autoscaling draws 18% and cost per million tokens 16% — no longer the outlier it once was in this series, but still last.

Measurement follows the same logic. Uptime and reliability is the primary success metric for 51% of enterprises, well ahead of developer productivity and deployment speed (39%), cost per million tokens (31%), latency (27%), and throughput (25%). Taken together, the operational metrics dominate the economic one by a wide margin.

This is a coherent posture for the production cohort in Finding 1 — teams running live workloads care first about whether the system stays up and how fast they can ship on it. But it sits uneasily beside Finding 7. Total cost of ownership has been demoted to fourth as a buying criterion at exactly the moment when 53% of enterprises still cannot rigorously track what their compute costs. The uncomfortable reading is that cost has fallen down the list partly because it remains the hardest thing in the stack to see, and criteria that cannot be measured tend to lose to criteria that can.

Finding 6: The GPUs run warmer, but most still run cold

Roughly seven in 10 GPU operators report 50% utilization or less

We asked what share of their GPU capacity enterprises actually utilize. Figures here are reported on the 155 enterprises that operate their own GPUs; 15 consume exclusively via API and run none.

The compute already in place runs cold, though less so than this series has recorded before. Roughly seven in ten GPU-operating enterprises (69%) report utilization at or below half capacity, with the 26–50% band alone accounting for 46%. About a quarter (26%) run at 25% or below. Against that, 23% now clear the 50% mark — a meaningful efficient minority rather than a rounding error.

The remaining 12% who do not measure utilization at all are the more troubling number, because they are invisible in both directions: they cannot claim efficiency and cannot detect waste. And utilization does not improve with maturity in the way one might expect — among enterprises running AI in production at scale, 22% clear the 50% mark, statistically indistinguishable from the 24% among everyone else. Scale is not, by itself, producing better-utilized fleets.

Idle accelerators are expensive accelerators, and this remains the clearest single measure of the gap in this report: enterprises are planning to evaluate specialized clouds and next-generation silicon (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large, and for one in eight enterprises, entirely unmeasured.

Finding 7: Fewer than half can account for what they spend

Rigorous cost tracking reaches only 56%, even among at-scale operators

We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger still lags the spending.

Measurement trails money. Fewer than half of enterprises (47%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (15%), or have not prioritized it (6%). Maturity helps but does not solve it: among enterprises running AI in production at scale, rigorous tracking reaches 56%, against 43% for everyone else. Even in the most operationally advanced segment of this sample, more than four in ten cannot account precisely for what their AI compute costs or returns.

Satisfaction with current infrastructure is moderately positive and tellingly uneven. On a five-point scale, overall satisfaction averages 4.14 and ease of implementation 4.04, while value for money trails at 3.87 — the softness landing on the one dimension that requires measurement to assess. Enterprises are, in effect, expressing dissatisfaction with an economic relationship most of them cannot yet quantify.

Read with Finding 5, the picture is self-reinforcing rather than merely inconsistent. Cost has slipped to fourth among buying criteria while remaining the least visible property of the stack, and the least visible property is the one enterprises rate lowest. Better instrumentation would not necessarily change what enterprises buy — but it would let them know whether the trade they are making for performance and availability is a good one.

Finding 8: The memory frontier is still unclaimed

Dell and Nvidia lead a scattered field, and one in five has no view

We asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The field remains early and fragmented.

The memory frontier is real but barely governed. Dell leads at 24% and Nvidia follows at 21%, with the remainder scattering across open-source tooling (12%), model-level efficiency techniques such as MLA and quantization (11%), and a long tail of storage vendors each in low single digits. No approach commands anything close to a majority, and the two leaders together account for less than half the field.

Most telling is that roughly one in five enterprises (19%) either do not recognize the constraint (7%) or have not begun to address it (12%). For a shift that will reshape inference cost and architecture, this is an early and unsettled market. It is also consistent with the measurement gap in Finding 7 — enterprises that cannot yet quantify what their current compute costs are in a poor position to anticipate which constraint will drive that cost next. The memory bottleneck is arriving while most of this cohort is still working to see the one in front of it.

The bottom line: Buying for speed, blind on cost

Organizations with more than 100 employees have moved AI infrastructure into production — two-thirds run live workloads, three in ten at scale — and their buying behavior has matured accordingly. They run three platforms on average, select on integration and performance, and measure success on uptime and developer velocity. For teams operating live systems, that is the right set of priorities.

What has not matured is the accounting. Total cost of ownership has fallen to fourth among selection criteria and cost per million tokens sits last, at the same moment that 53% of enterprises cannot rigorously track what their compute costs, 69% of GPU operators run at half capacity or less, and 12% do not measure utilization at all. Value for money is the lowest-rated attribute of the infrastructure they run — a judgment most of them are making without the instrumentation to support it. Cost has not become unimportant; it has become invisible, and the buying criteria have quietly reorganized around what can actually be seen.

Meanwhile the next round of spending points past the current stack. Specialized AI clouds are the top evaluation target at 44% and carry the strongest net momentum of any approach, against a 3.5% usage rate and a 4% near-term switching consideration — the widest intent-to-action spread in the data. Non-Nvidia accelerators draw 39%. And the constraint after this one, the shift from compute to memory in large-scale inference, is unrecognized or unaddressed by one enterprise in five.

At 170 respondents in a single July wave, reaching further up-market than this series typically does, this is a directional read — but the direction is consistent. Enterprises have become good operators of AI infrastructure and have not yet become good accountants of it. The open question for later waves is whether the instrumentation catches up before the re-platforming arrives, or whether enterprises buy the next layer of compute as blind to its economics as the last.


Based on survey responses from 170 qualified enterprise respondents (100+ employees), drawn from a single July 2026 wave. This sample is self-selected and directional rather than a precise measurement, and reads cross-sectionally with no month-over-month trend claims. Respondents include managers, individual contributors, C-suite, and VPs/directors, with purchasing authority weighted toward decision-makers and recommenders, across technology, manufacturing, financial services, healthcare, and other industries. Note: Figures for the switching-timeline, GPU-utilization, and cost-tracking questions are reported as a percentage of unique respondents rather than selections; individual categories for these three questions may sum to more than the reported total.

Agentic security: Enterprises enforce agent permissions two-thirds of the time — and isolate high-risk agents less than one in five

12 August 2026 at 07:30

Across 116 enterprises, agents are in production and so are the incidents: A majority have already had a confirmed agent security event or a near-miss. Two-thirds of enterprises enforce scoped permissions at runtime. Barely one in five isolates its highest-risk agents, making containment the weakest layer in the stack precisely as autonomy scales. Credential sharing persists across nearly two-thirds of agent fleets, and 53% have already had a confirmed agent security event or near-miss, contributing to a growing lack of confidence in agentic security.  Security stacks remain overwhelmingly borrowed from model providers and hyperscalers, and confidence has slipped. Today, as many enterprises now believe AI-armed attackers are ahead of their defenses as believe the reverse.

This wave of VentureBeat Pulse Research examines how enterprises secure their AI agents: what tooling they run, how they manage agent identity and isolation, what has already gone wrong, how much they spend, and whether they believe their defenses are keeping pace with AI-enabled attackers.

Only 18% of enterprises isolate their highest-risk AI agents, even as 65% of enterprises enforce scoped permissions at runtime and 56% monitor and log agent activity. The gap between what enterprises watch and what they contain is the central finding of this wave of VentureBeat Pulse Research.     

More than half of enterprises (53%) have agentic AI systems in production today, and another 27% are piloting or running a limited rollout. The agentic security incidents are arriving with them: 53% of organizations have already had an agent security event, with 19% confirming an incident and 38% having identified a near-miss that was caught before it caused harm.

The central finding is a containment gap. Enterprises have built the controls that watch and permission agents but not the one that bounds the damage when those fail. Among enterprises describing their security posture, 65% enforce scoped identities and permissions at runtime and 56% observe and log agent activity, yet only 18% isolate high-risk agents in sandboxes. Even among enterprises running agents in production, isolation is enforced just 21% of the time, and just 8% pair enforcement with isolation. That ordering is backward from a defense-in-depth standpoint. From SOC teams to CISOs, security teams know that observation tells you what happened and enforcement tries to prevent it, but isolation is what limits the blast radius when prevention fails.

Identity has improved without being solved. 49% of enterprises say each of their agents has its own scoped, managed identity, but 63% report credential sharing somewhere in the agent fleet, and only 29% describe a fleet with scoped identities and no sharing anywhere. The security stack doing this work remains overwhelmingly hyperscaler or model provider-native: OpenAI’s guardrails (44%), Microsoft Azure (42%), Anthropic’s managed-agent controls (37%), and Google Cloud (31%) lead, and 92% of enterprises naming a primary security layer name a hyperscaler/model provider-native one.

Two things have shifted against the comfortable picture. Confidence has slipped, with 30% now saying AI-armed attackers are ahead of their defenses, exactly as many as say their defenses are ahead. And churn intent is the highest this series has recorded, with 74% planning to adopt, add, or replace agent security tooling within twelve months, despite satisfaction scores at a series high of 4.29 out of 5. Enterprises are more satisfied than ever with a stack they are more determined than ever to replace.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent security — the tooling, identity, isolation, and enforcement controls organizations use to secure autonomous AI agents. Responses are filtered to organizations with more than 100 employees (n=116; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single July 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends; all figures are drawn from the July fielding only. Several questions were multiple-select, so those shares can sum to more than 100%.

By role the sample is senior and buyer-credible: 44% are final decision-makers for AI purchases and another 38% recommenders or influencers. Managers (36%), individual contributors (27%), VPs and directors (18%), and the C-suite (16%) make up the seniority mix. By organization size the sample is mid-market-weighted with a meaningful enterprise tail: 101–250 (34%) and 251–1,000 (23%) employees lead, with 1,001–5,000 (18%), 10,001+ (17%), and 5,001–10,000 (7%) above them. Technology/Software is the largest industry at 38%, followed by Healthcare/Life Sciences (11%) and Financial Services (10%).

Three questions require a base note. Two questions were asked only of enterprises with agents live or piloting. Posture figures (observe / enforce / isolate) are reported on those 93 respondents, and primary-security-layer figures on the 92 of them who named a layer. The 23 respondents outside this base are those still evaluating, without plans, or unsure — organizations for which an agent security posture would not yet apply. And several multiple-select questions permitted overlapping answers where one was intended — identity (33 respondents selected more than one pattern), arms-race assessment (23), budget share (10), and incidents (9) — so those are computed at the respondent level and the overlap is described where it matters. Satisfaction ratings are computed on the respondents who answered each rating question; the overall satisfaction score reflects 76 of the 116 qualified respondents.

At 116 respondents, the sample supports directional reads but not precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively standing up agent security rather than from the largest operators.

Finding 1: Agents are in production, and so are the incidents

A majority have already had an agent security event

We asked whether organizations run agentic AI in production, and whether they had experienced an agent security incident — a confirmed breach, or a near-miss caught before harm.

Agents have moved into production for this cohort. More than half of enterprises (53%) run agentic AI systems live today, another 27% are piloting or running a limited rollout, and only 3% have no plans in the next twelve months. The security exposure has scaled with the deployment: 53% of organizations have already had an agent security event, 19% a confirmed incident and 38% a near-miss caught before it caused harm.

That the near-misses outnumber confirmed incidents two to one is worth reading carefully. It means enterprises are catching problems, but catching them close to the edge — and a near-miss is a control that worked once, not a control that will work every time. The controls examined in the rest of this report, particularly the identity and isolation gaps in Findings 2 and 3, are what determine whether the next near-miss stays a near-miss.

One pattern from earlier waves does not replicate here. Organization size makes no reliable difference to exposure: enterprises above 1,000 employees report an incident or near-miss at 47%, against 57% among those between 101 and 1,000 — a difference well inside sample noise, and pointing the opposite direction from the size gradient this series has previously recorded. In this wave, what separates the hit from the not hit is not headcount.

Finding 2: Identity is improving — and still shared

Half give agents scoped identities; two-thirds still share credentials somewhere

We asked how enterprises manage the identity of their AI agents — whether each agent has its own credentials, or agents share them. Respondents could describe more than one pattern across the fleet.

Per-agent identity is now the most-cited pattern: 49% of enterprises say each agent carries its own scoped, managed identity, the precondition for least-privilege access and clean attribution. That is real progress on the control this series has repeatedly identified as the structural weakness beneath agent incidents.

But the answers overlap, and the overlap is the finding. Thirty-three respondents described more than one identity pattern across their fleet, and rolled together at the respondent level, 63% of enterprises report credential sharing somewhere — either agents mostly running on shared API keys and borrowed human or service-account credentials (37%), or a mixed fleet where some agents are scoped and many are not (34%). Only 29% describe a fleet with scoped identities and no sharing anywhere at all. Among enterprises with agents in production, 60% report per-agent identity, so the improvement is concentrated where the agents actually are — but so is the residual sharing.

The consequence is unchanged by the improvement. Where credentials are shared, an over-permissioned or compromised agent acts with far more reach than intended, and post-incident forensics cannot cleanly establish which agent did what. Half a fleet with scoped identities still has the blast radius of the half without. Non-human identity remains the largest unfinished piece of enterprise agent security, and as Finding 8 shows, it is still almost entirely absent from what enterprises are shopping for.

Finding 3: Isolation is the control nobody builds

Two-thirds enforce at runtime; fewer than one in five sandbox

We asked what an organization’s agent security posture looks like in practice — whether they observe, enforce, isolate, or some combination. The control that bounds damage is by far the least common. Figures are reported on the 93 respondents who described a posture.

This is the containment gap, and it is the widest structural gap in the report. Enforcement and observation are now common — 65% enforce scoped permissions at runtime and 56% monitor and log agent activity — while isolation sits at 18%. Only 8% of enterprises run both enforcement and isolation together, the posture that both prevents and contains.

Deployment maturity is a better predictor than the aggregate figures suggest. Isolation reaches 21% among enterprises with agents fully in production, compared with 13% among those still piloting — a meaningful gap that tracks maturity rather than exposure. Among enterprises that report credential sharing in the fleet, the group with the widest potential blast radius per Finding 2, isolation reaches 15%. The organizations with the most exposure are not meaningfully more likely to have built the control that bounds it.

The ordering is backwards from a defense-in-depth standpoint. Observation tells you what happened after the fact. Enforcement tries to stop it. Isolation is what limits the damage when enforcement fails — and enforcement will sometimes fail, which is the entire premise of the near-misses in Finding 1. An agent fleet that is watched and permissioned but not boxed in is precisely the configuration in which a single control failure propagates across systems. Enterprises have built the first two layers of the model and largely skipped the third.

Finding 4: Security still runs on borrowed, provider-native controls

Nine in 10 name a model provider or hyperscaler as their primary layer

We asked which agent security tooling enterprises use, and which is their primary layer. The answer continues to favor the model providers and hyperscalers over the dedicated security vendors.

Enterprises secure agents with tools that came bundled with their models and clouds. OpenAI’s guardrails lead at 44%, followed closely by Microsoft Azure (42%), Anthropic’s managed-agent controls (37%), and Google Cloud (31%). Asked to name a single primary security layer, 92% of those who answered named one of these provider-native offerings, with Azure (27% of answerers) and Anthropic (26%) leading.

The purpose-built agent-security category is no longer at zero, but it remains marginal. Cloudflare (11%) and Cisco (9%) lead the specialists, with CrowdStrike, Palo Alto, Zenity, Check Point’s Lakera, HiddenLayer, F5, and SentinelOne each between 1% and 7%. The identity specialists most directly relevant to Finding 2 are the smallest of all: Microsoft Entra Agent ID at 7%, Okta for AI Agents at 3%, and non-human identity platforms at 3%. Dedicated runtime sandboxing tooling — the control missing in Finding 3 — is in place at 3%.

A note on reading these shares: As described in the methodology section, the respondent sample is self-selected, and the usage question counted every vendor or approach a respondent has in place — so the figures measure presence in the security stack rather than spending or exclusivity. Individual vendor percentages therefore carry all the usual sample caveats. The structural pattern is the durable part: provider-native and hyperscaler controls lead by a wide margin, and dedicated agent-security specialists remain in single digits. Read the individual shares loosely and the pattern with confidence.

Finding 5: Satisfaction is at a series high — and so is churn intent

Enterprises rate their tooling 4.29 of 5 and three-quarters plan to replace it

We asked how satisfied enterprises are with their current agent security tooling, and whether they plan to adopt a new, additional, or replacement solution within twelve months. The two answers do not sit comfortably together.

Satisfaction with agent security tooling is the highest this series has recorded — 4.29 out of 5 for both overall satisfaction and ease of implementation, with value for money close behind at 4.11. That is a striking set of scores for a stack that is mostly borrowed provider guardrails, given that a majority of the same enterprises have already had an incident or near-miss and fewer than one in five isolates high-risk agents.

The purchase intentions tell the other half of the story. Three-quarters (74%) plan to adopt, add, or replace agent security tooling within 12 months, and 30% within the next quarter alone — higher churn intent than this series has previously seen in this category. Only 26% intend to stand pat. Enterprises are simultaneously more satisfied with their tooling and more determined to change it than at any prior reading, which suggests the satisfaction rests on the convenience and low friction of provider-native controls rather than on demonstrated containment. It is comfort with what is easy, not confidence in what is sufficient.

Finding 6: Budgets are finally moving

A third now spend more than a tenth of the security budget on agents

We asked what share of the security budget enterprises allocate to securing AI agents. The allocation has grown, though it remains a modest slice.

Agent security spending is still a slice rather than a pillar, but it is a growing one. The most common allocation remains 6–10% of the security budget (44%), and roughly a third of enterprises (35%) now devote more than a tenth — a meaningful funded minority. Just over a quarter (28%) spend 5% or less.

Read against Findings 1 through 3, the budget looks like a lagging but responsive indicator. A majority of enterprises have had an incident or near-miss, credential sharing persists across two-thirds of fleets, and fewer than one in five isolates high-risk agents — gaps that a 6–10% allocation is unlikely to close quickly. The enterprises spending above a tenth are the ones with the resources to build scoped identity and isolation controls rather than adopt whatever their model provider ships, and whether that minority grows is a reasonable leading indicator for whether the containment gap narrows.

Finding 7: The arms race has tilted

As many say attackers are ahead as say their defenses are

We asked how enterprises assess the balance between their AI-enabled defenses and AI-enabled attackers. Confidence has slipped into an even split.

Enterprises are no longer net-optimistic about the contest. Exactly as many say AI-armed attackers are ahead of their defenses (30%) as say their defenses are ahead (30%), with another 33% calling it roughly even and 24% saying it is too early to tell. Taken together, 63% rate the balance as even or worse.

Experience is what drives the pessimism, and the relationship is statistically clear. Among enterprises that have had a confirmed incident or near-miss, 39% say attackers are ahead; among those that have not, 20% do — a gap large enough to be unlikely to arise by chance in a sample this size. Getting hit does not just change what enterprises buy; it changes how they read the contest. The organizations closest to the actual threat are the least confident about it.

That assessment sits uneasily beside the series-high satisfaction of Finding 5. Enterprises rate their tooling 4.29 out of 5 while a clear majority believe it is, at best, holding even against an adversary that is also compounding with AI. An even race is not a comfortable place to be, and the group that has actually been tested rates it worse than even.

Finding 8: A reshuffle is coming — but identity still isn’t on the list

Incidents drive urgency; the control they implicate draws 10% interest

We asked which agent security solutions enterprises are considering. The consideration set has broadened, but not in the direction the incident data points.

Incidents start the buying cycle. Among organizations that have had a confirmed incident or near-miss, 38% plan to adopt, add, or replace agent security tooling within the next ninety days, against 22% of organizations with no incident; after a confirmed incident specifically the figure reaches 41%. Experience remains the strongest predictor of urgency in this data, as it is of pessimism in Finding 7.

The consideration set still leans provider-native — OpenAI (38%), Microsoft Azure (37%), Anthropic (35%), and Google Cloud (28%) lead — though the dedicated security vendors now draw meaningful early interest: Cisco (10%), Cloudflare (9%), Zenity and CrowdStrike (8% each), and Palo Alto, Check Point’s Lakera, and open-source guardrails (6% each). For most of the specialists that is more forward interest than current footprint.

What the shopping still does not include is the identity layer. Just 10% of enterprises include an agent-identity product — Okta for AI Agents, Microsoft Entra Agent ID, or a non-human identity platform — anywhere in their consideration set. Among the enterprises that both share credentials and have already been hit, the group with the most direct evidence that the control matters, identity consideration is no higher: roughly one in ten. Runtime sandboxing tooling draws 6%. The two controls most directly implicated by the incident data, identity and isolation, are the two least present in the purchase plans — the same blind spot this series recorded in the prior wave, unchanged despite a year of incidents.

The bottom line: A security gap that prevention alone won’t close

Organizations with more than 100 employees have put agents into production — 53% run them live today — and the incidents have arrived alongside them, with a majority already reporting a confirmed event or near-miss. On the controls, the picture is genuinely mixed rather than uniformly poor: nearly half now give each agent its own scoped identity, two-thirds enforce permissions at runtime, and a third devote more than a tenth of the security budget to agents. Enterprises are building agent security in earnest.

What they are not building is containment. Fewer than one in five isolates high-risk agents, only 8% pair enforcement with isolation, and among enterprises running agents in production isolation reaches just 21%. Credential sharing persists across 63% of fleets, so the blast radius that isolation would bound remains wide. The stack doing this work is 92% provider-native by primary layer, and the specialists built for exactly these gaps sit in single digits. The result is an architecture optimized to prevent and observe, with almost nothing in place for the case where prevention fails — which is the case the near-misses in Finding 1 describe.

The uncomfortable pairing is confidence with exposure, and it has sharpened. Satisfaction is at a series high of 4.29 out of 5, yet 63% rate the contest against AI-armed attackers as even or worse, 30% say attackers are ahead outright, and 74% plan to replace tooling they just rated highly. Enterprises that have actually been hit are markedly more pessimistic and markedly more urgent — and still not shopping for identity or isolation, the two controls their incidents most directly implicate.

At 116 respondents in a single July wave this is a directional read, weighted toward the mid-market — but the direction is clear: agent deployment is running ahead of agent containment, and the gap is not in what enterprises watch or permission but in what happens when those controls fail. The containment gap will not be closed by a better provider guardrail. The open question for later waves is whether enterprises build isolation and governed identity deliberately, or whether a confirmed incident that propagates does it for them.


Based on survey responses from 116 qualified enterprise respondents (100+ employees), drawn from a single July 2026 wave. This is a directional signal from a self-selected sample, not a probability sample. Respondents include managers, individual contributors, VPs/directors, and C-suite leaders, across technology, healthcare, financial services, and other industries.

Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least

12 August 2026 at 07:30

Across 108 enterprises, trust in automated agent evaluation rose sharply in July — and the failure rate it is supposed to predict did not move at all. The share of organizations that fully trust automated evaluation nearly tripled, from 5% in June to 13%, and the complaint that evaluations don’t match real-world outcomes fell 10 points. Yet the same share as last month — just under half — shipped an agent that passed its evals and then failed a customer. The reason is visible in the cross-tabs: the new trust belongs almost entirely to enterprises that have not yet been burned. Among those that have, 4% trust automated evaluation; among those that haven’t, 24% do. And getting burned does not slow the march to autonomy — it speeds it up.

This is the second wave of the VentureBeat Pulse Research agent reliability tracker, and the first fielded on an instrument identical to the month before it. That makes July the first read on direction rather than position: what moved, what held, and what the movement means.

What moved is confidence. In June, only 5% of enterprises said they fully trusted automated evaluation, and the most-cited limitation was that evaluations align poorly with real-world outcomes (29%). In July, 13% fully trust automated evaluation and the alignment complaint has fallen to 19%, no longer the leading objection. Both shifts are large enough to read as real rather than noise.

What held is the failure. Just under half of organizations (49%) deployed an agent or LLM feature in the past year that passed internal evaluations and then caused a customer-facing failure — statistically indistinguishable from June’s 50% — and a quarter (24%) have seen it happen more than once. Confidence improved; correctness did not. That is the July gap: not between autonomy and trust, as in June, but between trust and the evidence for it.

The cross-tabs explain where the new confidence comes from, and it is not from better evaluations. Trust is concentrated almost entirely among enterprises that have not experienced a false-confidence failure: 24% of them fully trust automated evaluation, against 4% of those that have. The trust curve is being lifted by inexperience. Meanwhile the enterprises that have been burned are not retreating from autonomy — 85% of them already allow zero-human deployment or are engineering toward it, against 61% of those that have not been burned. Overall the autonomy trajectory is flat at 67%, but the population inside it has shifted toward the organizations with the most direct evidence that evaluations miss things.

The vendor market, by contrast, is finally showing signs of settling. The share of enterprises running no dedicated evaluation tooling fell from 17% to 12%; specialist platforms gained, with Braintrust nearly doubling to 15% and DeepEval reaching 17%; and switching intent cooled, with those planning no change rising from 36% to 44%. Selection criteria moved with it: ease of integration overtook cost as the top factor, jumping from 27% to 39%. Enterprises are done shopping on price and have started buying on fit.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series. This wave — the agentic reliability and evals tracker — examines how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=108), drawn from a July 2026 fielding. Because the July instrument is identical to June’s, this report makes month-over-month comparisons where they are warranted; where questions were multiple-select, shares can sum to more than 100%.

Comparisons against June (n=157) are tested for significance, and only a handful of the month’s movements clear a conventional threshold: the rise in full trust in automated evaluation (5% to 13%), the fall in the real-world-alignment complaint (29% to 19%), the jump in ease of integration as a selection factor (27% to 39%), and the gain in Braintrust as a primary platform (8% to 15%). Movements described in this report as flat — the failure rate, the autonomy trajectory, the production monitoring mix, the investment ranking — are statistically indistinguishable between waves, and that stability is itself the finding. Differences of a few points elsewhere should be read as sample variation, not trend.

By role the sample is senior and buyer-credible: 44% are final decision-makers for AI purchases and another 25% recommenders or influencers, a slightly more senior mix than June. Product and program managers (18%), consultants and advisors (12%), CIOs/CTOs/CISOs (11%), and directors of engineering/IT (11%) lead the named titles, alongside a large “Other” function (30%). By organization size the sample is again mid-market-weighted: 100–499 (33%) and 500–2,499 (30%) employees lead, with 2,500–9,999 (23%), 10,000–49,999 (9%), and 50,000+ (5%) above them.

One composition change is worth flagging because it bears on the trust finding. The industry mix shifted between waves: Technology/Software fell from 23% of the June sample to 14% in July, while Retail/Consumer rose from 15% to 19% and now leads. A less technology-weighted sample plausibly carries less hands-on exposure to agent evaluation, and some of the month’s rise in trust may reflect who answered rather than what changed. The burned-versus-unburned split reported in Finding 2 holds within the July sample regardless, but readers should treat the headline trust movement as directional.

At 108 respondents the sample is large enough to support directional conclusions but should not be treated as a precise measurement; it is self-selected and is not a probability sample. Cross-tabs reported here rest on subgroups of 40 to 68 respondents and are correspondingly coarse.

Finding 1: The failure rate did not move

Just under half still ship agents that pass evals and fail customers

We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. The answer is the same as last month.

Forty-nine percent of organizations shipped an AI feature that cleared internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — against 50% in June. A quarter (24%) have seen it happen more than once, unchanged. Across two waves and 265 enterprises, the rate at which evaluations certify agents that then fail is stable to within a percentage point.

That stability is the anchor for everything that follows. Every other movement this month — rising trust, consolidating tooling, shifting purchase criteria — has to be read against a failure rate that has not responded. Whatever enterprises did between June and July, it did not change how often a passing evaluation turns out to be wrong.

Finding 2: Trust rose — among those who haven’t been burned

Full trust nearly tripled, and the alignment complaint fell ten points

We asked which limitation most reduces trust in automated agent evaluations today. The distribution shifted materially from June.

Two things moved together: Full trust in automated evaluation nearly tripled, from 5% to 13%, and the objection that most directly describes a false-confidence failure — poor alignment with real-world outcomes — fell from 29% to 19%, surrendering the top spot to evaluation bias and inconsistency (22%), now tied with data-leakage concerns (22%). On the surface this reads as an evaluation layer beginning to earn its keep.

The cross-tab says otherwise. Splitting the sample by whether an organization has actually experienced a false-confidence failure, trust divides almost completely. Among the 53 enterprises that shipped an agent which passed evals and then failed a customer, 4% fully trust automated evaluation. Among the 41 that have had no such failure, 24% do — a six-fold difference, and the sharpest split in the dataset. Direct contact with the failure mode is what removes the trust.

This is the month’s central caution. The improvement in sentiment is not evidence that evaluations got better; the failure rate in Finding 1 rules that out. It is what a trust curve looks like when a cohort of less-burned organizations enters the sample and reports its priors. Enterprises reading their own rising confidence as validation of their evaluation stack are reading a number that measures inexperience.

Finding 3: Being burned accelerates autonomy rather than restraining it

85% of the burned are on the zero-human path, against 61% of the REST

We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The aggregate held; the composition did not.

At the top line, nothing changed: 67% of organizations either already allow zero-human-in-the-loop deployment for low-risk agents (37%) or are actively engineering their pipelines to permit it within a year (30%), against 67% in June. The share ruling it out for the foreseeable future slipped from 22% to 18%. The autonomy ceiling stopped rising, but it did not come down.

Underneath, the picture inverts the intuitive one. Among enterprises that have shipped an evaluation-passing agent that then failed a customer, 85% are on the autonomy trajectory. Among those that have not, 61% are. Organizations with direct, expensive evidence that their evaluations miss things are substantially more likely to be removing the human check, not less — and only 11% of them rule out full automation, against 24% of those that haven't been burned. The pattern is identical for those burned once and those burned repeatedly.

The most plausible mechanism is not recklessness but maturity: the organizations that ship agents at enough volume to hit a customer-facing failure are the same ones with pipelines sophisticated enough to automate, and they are treating the failure as a cost of operating rather than a reason to stop. That is a defensible read. It is also precisely the dynamic that turns Finding 1’s stable failure rate into a growing absolute number of incidents, since the enterprises most likely to fail are the ones scaling their capacity to deploy without review.

One June finding did not replicate. Last month, larger enterprises appeared slightly further down the autonomy path than smaller ones (70% versus 64%). In July the two converge — 65% for organizations with 2,500+ employees against 68% below that, with near-identical failure rates (48% and 50%) — which suggests the June gap was sample variation rather than a size effect. Company size is not what separates the aggressive adopters; experience of failure is.

Finding 4: The stack begins to consolidate

Specialists gain, and the “Nothing at all” share shrinks

We asked which agent reliability or evaluation platform enterprises primarily use today. The field is still crowded, but it is no longer tied at the top with nothing.

The most consequential number is the one that fell. In June, having no dedicated agent-evaluation tooling was tied for the most common answer at 17%; in July it is 12% and fifth. Enterprises are acquiring evaluation tooling, and the specialists are capturing most of that movement: Braintrust nearly doubled its share of primary usage to 15%, and DeepEval reached 17%. Provider-native tooling held roughly flat — OpenAI at 18%, Anthropic at 12% — meaning the growth came at the expense of running nothing rather than at the expense of the model providers.

Counting any use rather than primary platform, the footprints are wider and the ordering is similar: OpenAI native evals reach 31% of enterprises, DeepEval 27%, Braintrust 22%, Anthropic native evals 20%, custom in-house tooling 14%, and Weave and Langfuse 11% each. Nineteen percent still report using no dedicated tooling anywhere in their stack. The category now has three plausible independent contenders where in June it had none with double-digit primary share — the first evidence in this series of an evaluation layer starting to take shape.

Finding 5: Production monitoring still watches the wrong thing

Half monitor whether the agent runs; under a third monitor whether it’s right

Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent’s output is correct — automated checks that evaluate the content of each answer as it goes out. A confidently wrong answer is invisible to the first kind: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked which kind live production monitoring is built for today.

Grouped by what is actually being watched, the split is essentially June’s: 50% of organizations monitor only whether the agent is functioning, while 26% run automated checks on whether its answers are right. Counting ad-hoc reviewers and don’t-knows, nearly three-quarters of organizations have no automated, real-time evaluation of output correctness in production. Inline quality assertions and transaction trace logging are tied as the most common approach at 26% each on a base of 106 — no single monitoring posture leads.

This is the finding that most directly contradicts the month’s rising confidence. Trust in automated evaluation went up eight points while the runtime capacity to detect an evaluation being wrong went nowhere. Among enterprises that already permit zero-human deployment, only 28% run inline quality checks on production traffic — which means the majority of organizations that have removed the human from the deployment decision have also not replaced that human with anything watching output quality afterward. The gate is automated and the alarm is not installed.

Finding 6: Bought on fit now, not on price

Ease of integration overtakes cost as the top selection factor

We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. One answer moved sharply; the other did not move at all.

Ease of integration jumped 12 points to 39% and displaced cost as the leading selection criterion, the clearest purchasing shift in the data. Evaluation accuracy rose modestly to 28%, cost fell to 23%, and breadth of observability (6%) and vendor roadmap (2%) remain marginal. Read alongside Finding 4, the two move together: enterprises adopting their first dedicated evaluation tooling are optimizing for what will slot into an existing pipeline this quarter, not for what is cheapest or most capable in the abstract. That is what a market looks like when it stops evaluating and starts installing.

What did not move is what enterprises want from the tool once installed. Evaluation consistency remains the primary success metric at 38%, essentially identical to June’s 36%, well ahead of reduction in failures (20%), speed of experimentation (18%), production visibility (16%), and compliance (7%). The priority is still repeatability — the same verdict on the same behavior every time — which is notable given that bias and inconsistency is now the top-cited trust limitation in Finding 2. Enterprises are buying for integration and measuring for stability, and are not yet getting the second. Satisfaction with current tooling remains moderate, averaging 3.9 on a five-point scale across overall satisfaction, ease of implementation, and value for money, barely changed from June’s 3.8.

Finding 7: Human review becomes the top line item

And the enterprises that have been burned fund it hardest

We asked which reliability and evaluation investment will grow most over the next year. Human review edged into first place.

Human review workflows (31%) and production observability (30%) swapped positions at the top, a change small enough to be noise on its own — but the underlying pattern is the same one June identified and it has strengthened. Enterprises plan to grow spending on human reviewers faster than on the automated evaluation pipelines (19%) that would replace them, at the same moment two-thirds are engineering the human out of the deployment decision. Only 6% report a flat budget, down from 8%.

The cross-tab makes the hedge explicit. Among enterprises that have shipped an evaluation-passing agent that failed a customer, 38% name human review as their fastest-growing investment; among those that have not, 24% do, and they favor observability tooling instead. So the burned cohort is doing both things at once: it is the most aggressive on autonomy (85% on the zero-human path, per Finding 3) and the most committed to funding human reviewers. That is not a contradiction so much as a strategy — automate the deployment decision, and pay people to catch what the automation misses. Whether that scales is the open question, since human review is the one part of the stack that does not get cheaper as agent volume grows.

Finding 8: The switching wave cools

Those planning no change rise from a third to nearly half

We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Fewer are shopping than last month.

A majority (56%) still intend to adopt a new, additional, or replacement platform within twelve months, but that is down from 64%, and the near-term cohort thinned from 31% to 24%. The share standing pat rose from 36% to 44%. Neither movement clears a significance threshold on its own, but both point the same direction, and they point it consistently with Finding 4: as enterprises actually acquire tooling, the population still looking for it shrinks.

The consideration set has reordered, too. Among the 60 enterprises planning a change, OpenAI’s native evals lead what they are evaluating (20%), followed by Braintrust (18%), Weights & Biases Weave (12%), and DeepEval (10%), with a further 10% actively evaluating but holding no shortlist. DeepEval led June’s consideration set at 20%; it has since converted much of that interest into primary usage, which is what a consideration-to-adoption handoff looks like. Braintrust now occupies the position DeepEval held — high interest ahead of installed base — and is the vendor to watch in the next wave.

The bottom line: Confidence moved, correctness didn’t

June found a gap between the autonomy enterprises were granting their agents and the trust they placed in the evaluations meant to govern it. July finds that gap closing from the wrong side. Trust rose — full confidence in automated evaluation nearly tripled and the complaint that evaluations miss reality fell ten points — while the thing that trust is supposed to track held exactly still. Just under half of enterprises still ship agents that pass their evals and then fail a customer, the same as last month.

The cross-tabs locate the new confidence precisely, and it is not in the evaluations. Twenty-four percent of enterprises that have never had a false-confidence failure fully trust automated evaluation; 4% of those that have do. Trust in this market is a function of exposure, not of evidence. And exposure does not produce caution: the burned cohort is the most autonomous in the sample, with 85% already deploying without human review or building toward it. What it produces instead is a hedge — the same organizations fund human review workflows hardest, at 38%, while removing humans from the deployment gate.

The vendor market is the month’s genuinely encouraging story. Running no dedicated tooling fell from 17% to 12%, specialists gained real share for the first time in this series, buyers shifted from price to integration fit, and switching intent cooled as adoption completed. An evaluation layer is finally forming. But the runtime picture has not followed: half of enterprises still monitor only whether their agents are running, and among those that already deploy without human review, just 28% run real-time checks on output quality.

At 108 respondents in a mid-market-weighted, self-selected sample, and with an industry mix that shifted away from technology between waves, this is a directional read. The direction, though, is legible: enterprises are tooling up, buying for fit, and growing more confident — and none of that has yet changed how often a passing evaluation turns out to be wrong. The question this series carried out of June was whether assurance would catch up to autonomy. July’s answer is that confidence caught up first, which is the harder problem, because an enterprise that trusts a broken gate has less reason to fix it than one that knows the gate is broken.


This report presents the July 2026 wave of an ongoing longitudinal series on enterprise AI agent reliability and evaluation, based on 108 qualified respondents at organizations with 100 or more employees. Comparisons are drawn against the June 2026 wave (n=157), fielded on an identical instrument. At this sample size, results should be read as a directional signal rather than a precise measurement — the sample is self-selected, not a probability sample. Respondents span final decision-makers, technology recommenders/influencers, and end business users, across a mid-market-weighted range of industries and company sizes.

Agent context layers: Enterprises governing their AI data are catching twice as many bad answers as the ones who aren't

12 August 2026 at 07:30

Across 101 enterprises, the context feeding AI agents is failing often and repeatedly. Sixty-eight percent have traced a confident but wrong agent answer to missing or inconsistent business context in the past six months, and the single most common answer is not "once" but "more than once." The counterintuitive part is which companies report it. Enterprises building or running a governed semantic layer (a layer of company-specific definitions and relationships) report recurring failures at more than twice the rate of those without one.

The infrastructure built to fix bad context is, so far, mostly revealing how much bad context there is. Meanwhile the architecture meant to solve the problem commands no consensus at all: hybrid retrieval and outright pluralism finish one respondent apart, in a dead heat.

This wave of VentureBeat Pulse Research examines the enterprise RAG and context layer: what feeds AI agents their business context, which retrieval systems enterprises run, how they buy and measure them, where the architecture is heading, and — most revealingly — how often that context is already failing them.

The central finding is that the context failure is no longer an incident; it is a condition. Sixty-eight percent of enterprises say that in the past six months their AI agents produced confident but wrong answers they traced to missing or inconsistent business context rather than to model error. More striking than the total is its shape: 37% report the failure recurring, against 32% who saw it once. Among enterprises in a position to answer at all, the most prevalent experience of running agents on company data is being wrong repeatedly for reasons that have nothing to do with the model.

The remedy the industry has settled on — a governed semantic or context layer giving agents and BI a shared understanding of the data — is being built at scale: 32% run one in production, another 31% are piloting or building one, and 20% more are evaluating. But the cross-tabs deliver an uncomfortable result: Enterprises that have built or are building a layer report recurring context failures at 50%, against 21% for those without one. The layer isn't causing the failures — it's catching them, which makes it the most useful finding in the wave. The semantic layer is what makes a context defect traceable. Organizations without one are not having fewer failures so much as attributing fewer failures.

Underneath, the stack is unsettled in a way it was not expected to be. Retrieval remains the leading primary context source at 31%, and provider-native retrieval — OpenAI's file search (46%) and Google Vertex AI Search (41%) — still runs well ahead of every dedicated vector database. But the expected architecture has no majority behind it: hybrid retrieval (30%) and "multiple architectures, chosen by use case" (29%) are separated by a single respondent. And enterprises remain firmly unwilling to hand the context layer to a provider — just 12% intend to consolidate onto a single model provider’s native context stack, against 37% holding to best-of-breed and 37% planning an explicit mix.

The buying criteria are where the failure is starting to register commercially. Access control and permissions is now tied with ease of data ingestion as the top selection factor at 24% each, and response correctness is the primary success metric for 38% of enterprises. Enterprises are beginning to buy retrieval for the properties that govern context rather than the properties that move it.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series. This survey focused on enterprise RAG infrastructure and the context layer — the retrieval systems, semantic layers, and context sources that feed AI agents. Responses are filtered to organizations with more than 100 employees (n=101). All responses are from a single July 2026 wave, so the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select; those shares are reported as a percentage of respondents, not of total selections, so they can sum to more than 100%.

By organization size the sample concentrates in the mid-market: 101–250 employees (34%), 1,001–5,000 (25%), and 251–1,000 (25%) lead, with 10,001+ (12%) and 5,001–10,000 (5%) above them. By role it spans managers (39%), individual contributors (29%), VPs and directors (22%), and the C-suite (9%); on purchasing authority it is buyer-credible, with 38% final decision-makers and another 43% recommenders or influencers. Technology/Software is the largest industry at 31%, followed by Healthcare/Life Sciences (14%), Retail/E-commerce (10%), and Manufacturing (9%).

A note on the context-failure base: Of the 101 respondents, 10 either do not run agents on enterprise data (5%) or do not trace root cause at that level (5%). Headline shares for the failure question are reported on the full 101; the subgroup comparisons in Finding 2 use the 91 respondents who were able to give a yes-or-no answer, since including those who cannot observe the failure would bias the comparison toward whichever group is less instrumented. Subgroup cells run from roughly 10 to 62 respondents and are correspondingly coarse; where a cell falls below 10 it is not reported as a percentage. A small number of respondents selected "Other" and gave a write-in industry (6%) or role (2%) that didn't map to a listed category; those shares appear as not stated in the appendix rather than being redistributed.

At 101 respondents this is a modest sample and should be read as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively standing up RAG and context infrastructure rather than from the largest operators.

Finding 1: Confident, wrong, and repeating

The most common answer isn't "once" but "more than once"

We asked whether, in the past six months, enterprises had traced a confident but wrong agent answer to missing or inconsistent business context rather than to model error. Most had — and most of those had seen it happen again.

This is the report’s defining number. Sixty-eight percent of enterprises have had an AI agent produce a confident, wrong answer they traced to bad context — wrong metric definitions, stale data, missing documents — and the recurring case (37%) outweighs the one-off (32%). Only 22% report no such failure. Restricted to the 91 enterprises able to observe and attribute the failure at all, 76% have experienced it and 41% repeatedly.

The failure mode is specific and dangerous precisely because it does not look like a failure. The model is not visibly hallucinating; it is confidently wrong because the context feeding it was thin, stale, or inconsistent — and it delivers that wrong answer with the same authority as a right one. That the modal experience is recurrence rather than a single incident matters more than the headline share: a one-time failure is an incident to be fixed, while a repeating one indicates a structural defect in how business context reaches the agent. Everything else in this report — what enterprises retrieve, how they govern it, and what they plan to build — is downstream of this problem.

Finding 2: The semantic layer reveals the failure before it fixes it

Enterprises building a governed layer report more recurring failures, not fewer

We asked whether enterprises use a governed semantic or context layer to give agents and BI a shared understanding of their data. Most are on the path — and cross-tabbing that answer against the failure in Finding 1 produces the wave’s most counterintuitive result.

Engagement with the governed context layer is broad. Sixty-three percent of enterprises either run one in production (32%) or are piloting and building one (31%), and a further 20% are actively evaluating, meaning more than four in five are engaged with the idea in some form. Only 14% have no plans.

The cross-tab is where it gets interesting. Among the 91 enterprises able to answer the failure question, those who have built or are building a semantic layer report recurring context failures at 50%, while those without one — evaluating or with no plans — report them at 21%, a gap that clears conventional significance thresholds (p=0.01) and runs in the direction opposite to what the technology is sold to do. Narrowing to enterprises with a layer specifically in production points the same way but does not carry statistical weight on this sample: 53% recurrence against 34% for everyone else, a difference that does not reach significance and should be read as directional only.

Read as causation, this is implausible — a governed definition layer does not manufacture wrong answers. Read as detection, it is the most useful result in this wave. Tracing a confident wrong answer to a specific context defect — a metric defined two ways, a stale table, a document the agent could not see — requires exactly the shared, governed definitions a semantic layer provides. Without one, the same failure occurs and gets logged as a model problem, a user error, or nothing at all. The causation almost certainly also runs backwards in part: enterprises that have been burned repeatedly are the ones who went and built the layer.

The size split points the same way, and carries significance where the production split does not. Enterprises above 1,000 employees report recurring context failures at 55%, against 30% of those between 101 and 1,000 (p=0.02) — despite the larger organizations being less likely, not more, to have a semantic layer in production (24% against 37%). Larger enterprises have more instrumentation, more auditing, and more people whose job is to ask why a number was wrong. The practical implication for readers is uncomfortable but clear: a low reported context-failure rate is not evidence of a healthy context layer. It is at least as likely to be evidence that nobody is looking.

Finding 3: RAG leads as the context source — and carries the failures

Retrieval feeds more agents than anything else, and fails a large share of them

We asked what an enterprise’s AI agents primarily use to understand its data. Retrieval leads, but no longer by the margin the category assumes.

 Retrieval remains the backbone of enterprise context at 31%, ahead of a governed semantic layer (19%) and mixed approaches (17%). But the tail has thickened in a way worth noting: long-context loading is now the primary source for 13% of enterprises, and 5% let agents run on the model’s general knowledge with no enterprise context layer at all. Between them, nearly one in five enterprises is feeding agents business context either by brute-force context window or not at all.

Cross-tabbed against Finding 1, the sources do not fail equally. Among enterprises whose primary context source is retrieval, 87% report a context-traced failure and 48% report it recurring — on the largest base of any group, 31 respondents. Those relying on a governed semantic layer report 79% and 53%; mixed approaches 79% and 36%; direct live-system queries 40% and 30%. The long-context group is the outlier in the other direction, reporting 64% any failure but only 9% recurrence.

These subgroup figures should be read with the detection caveat from Finding 2 firmly attached. Groups differ in how well they can attribute a wrong answer to a context defect as much as in how often they suffer one, and the cells here run from 10 to 31 respondents. The retrieval group’s 87% is best read as evidence that RAG-heavy enterprises both experience and notice context failures, not as a clean measurement of relative reliability.

What survives the caveat is the structural point. Because so much enterprise context flows through retrieval, and because retrieval carries that load on the widest base in the sample, the quality of retrieval is the quality of the answer. When RAG is the default source, incomplete retrieval is the main point of failure.

Finding 4: Model-backed and hyperscaler retrieval still leads the vector databases

OpenAI's file search and Google's Vertex AI Search top every purpose-built system

We asked which retrieval systems enterprises run in production today. The answer continues to favor the model providers and hyperscalers over the specialists.

The dedicated vector database is not the center of the RAG stack. OpenAI’s file search (46%) and Google’s Vertex AI Search (41%) lead by better than three to one over any purpose-built alternative. Among the specialists, the most-used remain the ones enterprises already run for other reasons — Elasticsearch/OpenSearch at 20% and pgvector at 15% — while the pure-play vector databases that define the category (Pinecone, Weaviate, Milvus, Qdrant) each sit between 7% and 12%. Custom in-house retrieval stacks, at 18%, outrank every pure-play vendor.

Which system is actually primary separates retrieval from infrastructure

Usage counts alone understate the gap, because enterprises run several of these systems at once. We also asked which one is primary. The share of each system’s own users who name it their primary retrieval platform divides the field cleanly.

Elasticsearch and pgvector are widely present and rarely primary: four in five of their users retrieve mainly through something else. They are infrastructure the enterprise already ran, pressed into service at the edges of a retrieval stack whose center is elsewhere.

Model-backed and hyperscaler retrieval is not merely the most common system on the list; for most of the enterprises that adopt it, it is the system of record. Custom in-house stacks behave the same way — when an enterprise builds one, it is usually the primary, not a side project.

The primary-platform question was fielded as a single-select and 18 of 101 respondents selected more than one option, so the shares above are computed as a proportion of each system’s users rather than of the full sample. On the 83 respondents who gave exactly one answer, the ranking is unchanged: OpenAI's file search 28%, Vertex AI Search 23%, custom in-house stack 12%, and no pure-play vector database above 8%.

The comparison worth sitting with is what this leaves for the RAG specialists. In a category built around specialist infrastructure, more enterprises have written their own retrieval stack than run any single dedicated vector database — and roughly four times as many use retrieval that arrived bundled with a model provider or cloud they already buy from. Only 7% run no production RAG at all, so this is not a story about early adoption; it is a story about where retrieval gets acquired.

Finding 5: No architecture commands a consensus

Hybrid retrieval and "It depends on the use case" finish in a dead heat

We asked which retrieval architecture enterprises expect to dominate their production RAG systems by the end of 2026. No single answer comes close to a majority — and the two front-runners are separated by one respondent.

Hybrid retrieval leads at 30%, with the expectation that no single architecture will dominate at all immediately behind at 29%. The gap is one respondent, far inside the margin on a sample this size, and the honest reading is that these two finish level rather than that either is in front. Together they account for 58% of enterprises, and what unites them is more instructive than what separates them — both describe layered pipelines rather than a single retrieval technique, and neither expects the pure vector-search approach that launched the category to carry production on its own.

Two smaller answers carry the sharper signal. Fifteen percent expect tool-first or long-context retrieval to dominate without a dedicated vector layer at all — a direct challenge to the premise of the category — while 12% still expect vector-only retrieval to prevail. That the anti-vector position now edges the pure-vector one, on a three-respondent margin that is itself too narrow to call, is a notable inversion for an industry that spent three years building vector databases. Add the 15% who are unsure or expect no large-scale RAG, and the picture is of a market that agrees vector search alone is insufficient and has not agreed on what replaces it.

Finding 6: Enterprises decline to hand the layer to a provider

Consolidation onto a provider's native context stack barely registers

We asked how enterprises will respond as model providers bundle retrieval, memory, and orchestration into their platforms. Their stated intent cuts sharply against their current usage.

Here is the tension at the heart of the stack. Provider-native retrieval leads actual usage by a wide margin (Finding 4), yet just 12% of enterprises intend to consolidate onto a provider’s native context stack. Best-of-breed standalone tools and an explicit mix are tied at the top at 37% each, and 6% intend to build and own the layer themselves — meaning 79% of enterprises expect to keep at least part of the context layer outside any single provider.

The gap between what enterprises run and what they say they want is the strategic question of the category. They are adopting bundled retrieval because it arrives with tools they already buy, while asserting they will preserve independence. Read against Finding 2, the stated preference has a rationale beyond vendor politics: the failures enterprises are trying to fix are failures of governed, consistent, access-aware business context, and that is precisely the layer they are least willing to outsource. Whether the preference survives contact with the convenience of the bundle is what the next several waves will decide.

Finding 7: Access control climbs into the buying decision

Governance now ties ingestion as the reason a system gets chosen

We asked what matters most when enterprises choose a retrieval system, and what they treat as the primary measure of success once it is running.

The selection criteria have moved toward governance. Access control and permissions (24%) is now exactly tied with ease of data ingestion (24%) at the top, ahead of retrieval accuracy and latency and performance (15% each) and operational simplicity (14%). That puts a governance property at the top of the purchase decision for the first time in this series — and it is the property most directly implicated in the confident-but-wrong failures of Finding 1, where an agent surfaces something it should not have seen or misses something it should have.

Once systems are running, the emphasis on correctness is unambiguous: response correctness is the primary success metric for 38% of enterprises, twice the next answer, security and access control (19%). Answer relevance (17%), latency (13%), and operational stability (11%) trail. Taken together, 56% of enterprises measure their retrieval system primarily on whether its answers are right or properly permissioned, rather than on whether it is fast or stable.

Satisfaction with current systems is moderately positive: on a five-point scale, overall satisfaction averages 4.13, value for money 4.01, and ease of implementation 3.98. That is a respectable set of scores for a layer that, on this wave’s evidence, is producing recurring wrong answers in nearly four in ten enterprises — which suggests enterprises are rating the tools against expectations of what retrieval infrastructure does, not against the outcome of getting the answer right.

Finding 8: Half the market is in motion

Vertex AI Search leads the consideration set — and so does uncertainty

We asked whether enterprises plan to change or add a retrieval provider, and which they are considering. The consideration set is broader than today’s stack.

The retrieval stack is not settled, but it is not churning, either: about half of enterprises have no plans to change, while the other half — 52 of 101 — intend to switch or add a provider within twelve months, a fifth of them within the next quarter. Among those 52 enterprises in motion, Google’s Vertex AI Search leads the consideration set at 35%, followed by Elasticsearch/OpenSearch (25%), Pinecone (23%), and OpenAI's file search (23%).

Two patterns stand out. First, the pure-play vector specialists draw markedly more forward interest than their current footprint would suggest — Pinecone is considered by 23% of movers against 12% present usage, Weaviate 17% against 10%, Qdrant 15% against 7%, and Milvus 14% against 9%. The specialists are not winning the installed base, but they are firmly in the evaluation, and each of them roughly doubles its footprint in forward consideration. Second, 15% of movers are evaluating with no shortlist at all and 17% are considering a custom in-house stack — together nearly a third of enterprises planning a change either do not know what they want or intend to build it.

The bottom line: A context failure that better detection is only beginning to reveal

Organizations with more than 100 employees are running agents on business context they cannot yet guarantee, and the evidence has moved past anecdote. Sixty-eight percent have traced a confident, wrong agent answer to missing or inconsistent context in the past six months, and the recurring case now outweighs the one-off. Retrieval remains the default source of that context and carries the failure on the widest base in the sample — while nearly one in five enterprises has fallen back to long-context loading or the model’s general knowledge, which is not a context layer at all.

The most important result in this wave is the one that inverts the expected direction. Enterprises building or running a governed semantic layer report recurring context failures at 50%, against 21% for those without one, and larger enterprises report them at nearly twice the rate of mid-market peers despite being less likely to have the layer built. The straightforward reading is that instrumentation reveals failures rather than causing them, and that the organizations reporting clean context records are largely the ones without the means to check. That reframes the entire finding: the 22% reporting no context failure are not the well-governed cohort, and a low failure rate should be treated as a question rather than an answer.

Meanwhile, the fix has not converged. Hybrid retrieval and architectural pluralism finish level as the expectation for production RAG by the end of 2026, one respondent apart; the anti-vector position narrowly edges the pure-vector one; and while provider-native retrieval leads usage by a wide margin — and is the primary system for most of the enterprises that run it — only 12% will consolidate onto a provider’s stack, with 79% keeping some part of the layer independent. The commercial signal is that access control has climbed to tie ease of ingestion as the top buying criterion, and response correctness is the dominant success metric — enterprises are starting to buy retrieval for the properties that govern context rather than the ones that move it.

At 101 respondents in a single July wave, skewed toward the mid-market, this is a directional read. But the direction is clear enough to act on: the context layer is the contested tier of the AI stack, the failure it produces is recurring rather than occasional, and the enterprises best equipped to see the problem are the ones reporting it worst. The open question for later waves is whether the governed context layer starts to reduce the failures it is currently so good at exposing.


Based on survey responses from 101 qualified enterprise respondents (100+ employees), drawn from a single July 2026 wave. At this sample size the results should be read as a directional signal rather than a precise measurement — this is a self-selected sample, not a probability sample. Respondents include managers, individual contributors, VPs/directors, and C-suite leaders.

Agentic orchestration: Enterprise AI organizations know how to govern agents but still can't meter what they cost

12 August 2026 at 07:30

Across 107 enterprises, agentic orchestration is not a choice of a single platform.

The typical enterprise runs three orchestration platforms at once, and selects them for flexibility across models rather than affinity to any single one. Microsoft leads primary usage while Anthropic leads forward consideration by a wide margin. 

The AI control plane enterprises expect is deliberately hybrid, meaning it includes use of the leading AI providers, but also provider-independent technologies — and the risk they fear most from provider-resident control is not lock-in but the provider’s own security and permissioning limits. 

One in five enterprises still has no real-time way to stop a runaway agent before the bill arrives.

This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what drives the choice, what they optimize for, how they expect agent control to be structured, and — most revealingly — how orchestrated their deployed “agents” actually are and how tightly they control the cost of running them.

The central finding is that orchestration has become plural. Eighty-five percent of enterprises run two or more orchestration platforms and 64% run three or more, with a mean of 3.1 platforms per organization. Microsoft AI Foundry / Copilot Studio appears in 70% of stacks and OpenAI’s Agents SDK in 68%, with Anthropic’s Claude Platform in 47%. Asked to name a single primary platform, respondents who gave one unambiguous answer put Microsoft first (41%) and Anthropic second (28%). Nobody in this sample is running one orchestration layer and calling it a strategy.

The selection logic follows from that plurality. Flexibility across models and tools is the leading purchase driver at 29%, nearly three times the share naming model gravity — native alignment with a state-of-the-art base model — at 10%. Enterprises are not choosing the orchestration environment that comes with their favorite model; they are choosing the one that does not commit them to any model. Security and permissions (17%), production reliability (15%), and control over agent execution (15%) fill out a buying logic focused on governance and optionality rather than developer convenience.

A clear majority (53%) expect a hybrid control plane by the end of 2026 — provider-native plus external orchestration — and the risk they most associate with provider-resident control is security and permissioning limitations (37%), ahead of vendor lock-in (23%) and limited visibility (22%). Investment has moved accordingly: agent monitoring and debugging leads the spend at 31%, with security and permissions enforcement at 30%, while workflow tooling draws 19%. Enterprises are spending to see and govern agents, not merely to build them.

Most companies admit that a majority of their “agents” are really just chatbots. A plurality of 47% of respondents say that between 26 and 50% of their agents are genuinely orchestrated, with 37% at a quarter or below and 16% past the halfway mark. 

But fiscal control remains the soft spot: 21% of enterprises track agent spend only through post-hoc logs, with no real-time way to halt a runaway execution loop.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series, with this instrument focused on enterprise agent orchestration. Responses are filtered to organizations with 100 or more employees (n=107), drawn from a single July 2026 wave; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. All figures in this report come from the July fielding only. Where questions were multiple-select, shares can sum to more than 100%.

This wave draws a notably large-enterprise, technology-heavy sample, and that shapes every finding in it. By organization size, more than half sit at 10,000 employees or above: 50,000+ (26%) and 10,000–49,999 (25%) lead, followed by 2,500–9,999 and 500–2,499 (19% each) and 100–499 (11%). Technology/Software accounts for 53% of respondents, with Government/Public Sector (16%) and Manufacturing/Industrial (10%) next. By role the sample is hands-on and technical: software and ML engineers (22%), product and program managers (21%), directors of data/AI/analytics (17%), and VPs of data/AI/analytics (12%). On purchasing, 90% are recommenders, influencers, or final decision-makers for AI solutions (63% recommender/influencer, 27% final decision-maker).

A note on the primary-platform question. Forty-six of 107 respondents registered more than one selection on a question intended to capture a single primary platform. Because those responses cannot be resolved to one answer, primary-platform shares are reported on the 61 respondents who gave a single unambiguous answer, and are labeled as such wherever they appear. Platform footprint figures — which platforms an enterprise uses at all — use the full n=107 base and are unaffected. The ambiguity is worth noting on its own terms: on a question asking for one platform, more than four in 10 respondents could not or would not narrow to one, which is consistent with the multi-platform pattern documented in Finding 1.

At 107 respondents the sample is robust enough to read directionally with reasonable confidence, though it remains self-selected and is not a probability sample. Because each subgroup here only includes about 50 to 60 respondents, splits between them are less precise than the full-sample findings.

Finding 1: Orchestration is a portfolio, not a platform

The typical enterprise runs three orchestration platforms at once

We asked which agent orchestration platforms enterprises use, and which one they treat as primary. The first answer is that almost nobody has just one.

The defining feature of this layer is plurality. Only 15% of enterprises run fewer than two orchestration platforms; the median organization runs three, and one in six runs five or more. Read that way, the platform “shares” below describe overlapping deployments rather than a divided market — Microsoft and OpenAI each appear in roughly seven of ten stacks precisely because most stacks have room for several.

Asked to name one primary platform, the 61 respondents who gave a single unambiguous answer put Microsoft AI Foundry / Copilot Studio first at 41%, Anthropic’s Claude Platform second at 28%, LangChain / LangGraph at 10%, and OpenAI’s Agents SDK at 7%, with Google, Amazon, Salesforce, and custom in-house builds at 3% each. Microsoft’s lead on primary usage alongside OpenAI’s near-equal footprint on any usage is the signature of an enterprise-weighted sample: the Microsoft platform arrives through an existing enterprise agreement and becomes the default seat of record, while other platforms are added around it for specific work.

A note on reading these shares: As described in the methodology section, the respondents are self-selected, this wave skews heavily toward large technology organizations, and the primary-platform figures rest on a 61-respondent subset. The numbers measure where this cohort has placed its orchestration bets today, within a self-selected audience of AI-active technical practitioners. A sample built this way can diverge substantially from spend-weighted market measures, and each VB Pulse survey draws its own sample with its own company-size and industry mix, so vendor figures should not be compared across our surveys, either.

Respondents rate the platforms they run at 4.17 out of 5 for overall satisfaction, 3.91 for ease of implementation, and 3.63 for value for money — with value for money the weakest of the three by a clear margin. That ordering is itself a finding: enterprises are broadly happy with what these platforms do and distinctly less happy with what they cost, which is the same nerve the fiscal-control finding touches at the end of this report. Satisfaction sits alongside a two-thirds intent to change platforms within the year; this remains a layer enterprises work with rather than settle on.

Finding 2: Flexibility, not model gravity, drives selection

Enterprises buy the orchestration layer that doesn't commit them

We asked what most influenced the orchestration platform choice, and optionality leads by a distance.

Flexibility across models and tools (29%) is the selection-side explanation for the multi-platform reality in Finding 1: enterprises are choosing orchestration environments on the strength of what they leave open rather than what they lock in. Model gravity — picking the orchestration layer that comes with a preferred frontier model — draws just 10%, less than a third of the flexibility share, which places the pull of any single base model well down the list of what actually decides this purchase.

The next tier reinforces the governance emphasis. Security and permissions (17%), production reliability (15%), and control over agent execution (15%) together account for 47% of responses: nearly half of enterprises pick their orchestration platform on whether they can constrain and depend on what it runs. Ease of development draws 8% and total cost of ownership 4%, an inversion of how these platforms are usually discussed in engineering circles. Performance sits last at 2% — at this stage of adoption the binding constraints are optionality and control, not raw speed.

Finding 3: The job is reliable multi-step execution

Enterprises judge orchestration by whether it completes the work

We asked what enterprises optimize for — their primary success metric for orchestration. Reliability and multi-step workflow management lead, with developer productivity closer behind than in the buying criteria.

Task completion reliability (30%) and multi-step workflow management (27%) together account for 57% of responses: orchestration succeeds, in the enterprise view, when it reliably carries a task through multiple steps to completion. Developer productivity takes a substantial 23% — notably higher than ease of development’s 8% as a purchase driver in Finding 2, which suggests enterprises do not expect to buy developer velocity so much as to earn it once the platform is in place. End-user experience is a minor concern at 7%, consistent with orchestration being an internal execution problem rather than a UX one.

This reliability-first standard is the yardstick against which the portfolio-maturity finding later in this report should be read: enterprises define success as dependable multi-step execution, and a little over a third of them still say a quarter or fewer of their deployed agents do multi-step work at all.

Finding 4: Two-thirds plan to move — and Anthropic leads the consideration set

The installed base and the pipeline point to different vendors

We asked whether enterprises plan to adopt a new, additional, or replacement orchestration platform in the next 12 months, and which platforms they are considering.

Two-thirds of enterprises (67%) intend to adopt a new, additional, or replacement orchestration platform within the year, but the clock runs longer than the intent suggests: the largest cohort sits at 6–12 months (28%) and only 15% expect to move within a quarter. This is deliberate re-platforming on a planning horizon, not urgent churn.

The consideration set is where this finding earns its headline. Among the 72 enterprises in motion, Anthropic leads at 43% — well ahead of Google (31%), custom in-house builds (31%), OpenAI (25%), LangChain / LangGraph (17%), and Microsoft (17%). Set that against Finding 1, where Microsoft leads primary usage and appears in 70% of stacks: the installed base and the forward pipeline point at different vendors. Anthropic draws roughly two and a half times Microsoft’s forward consideration despite trailing it on current primary usage, and custom in-house control planes draw as much interest as any external platform besides Anthropic. A further 18% of movers are evaluating with no shortlist at all.

Read alongside the flexibility-first selection logic in Finding 2, the shape of the next twelve months is legible: enterprises expect to add rather than replace, they are shopping for platforms that preserve model choice, and a substantial minority intend to solve the problem themselves rather than buy it.

Finding 5: Investment flows to watching and governing agents

Monitoring and permissions lead the spend; workflow tooling trails

We asked which orchestration-related investment will grow most next year. Observability and governance take the top two places.

Monitoring and debugging (31%) and security and permissions enforcement (30%) are effectively tied at the top and together account for 61% of planned growth. The money is going to seeing what agents do and constraining what they are allowed to do — the two capabilities that matter once agents are running in production rather than being built toward it. Workflow tooling (19%) and scaling infrastructure (18%) trail, and almost no one is standing still: just 3% report a flat budget.

The emphasis is consistent with the buying logic in Finding 2, where security and permissions was the second-ranked selection factor, and with the control-plane architecture in Finding 6. Enterprises that have decided to run agents across three platforms have a visibility and permissioning problem by construction, and they are funding it directly.

Finding 6: The control plane will be hybrid — and security is why

Enterprises split control, and fear the provider's permissioning more than lock-in

We asked where enterprises expect the primary control plane for agents to live by the end of 2026, and what worries them most if that control sits inside a model-provider platform.

Hybrid control is the dominant expectation by a wide margin (53%). Taken together, the hybrid, custom in-house, and externally-abstracted options — every architecture that keeps control at least partly outside the provider — sum to 78% of enterprises, against 14% willing to hand control to a provider-managed service outright.

The reason enterprises give is worth separating from the one usually assumed. Security and permissioning limitations lead the risk question at 37%, well ahead of vendor lock-in at 23%, with limited visibility and observability close behind at 22%. Combining the security and visibility answers, 59% of enterprises name a control-and-oversight concern rather than a commercial one. The worry is less that a provider platform will be hard to leave than that it will not let them see or constrain what their agents are doing while they are on it — the same concern funding the monitoring and permissions spend in Finding 5. Only 2% say provider-resident control is not a concern at all.

Finding 7: The chatbot trap is loosening, not broken

“Bridging the gap” is now the modal answer on portfolio maturity

We asked enterprises to assess their portfolios honestly: What share of their deployed “agents” are true multi-step orchestrated workflows versus simple single-prompt chatbot wrappers.

The center of gravity has moved into the middle band. Just under half of enterprises (47%) now put between a quarter and half of their portfolio in genuinely orchestrated, stateful workflows, and 16% are past the halfway mark. The bottom two bands — a quarter or fewer genuinely orchestrated — account for 37%, and outright pure-chatbot portfolios have nearly vanished at 3%. Against the reliability-first success standard in Finding 3, this is a portfolio that has started to do the work the orchestration layer exists for, without most of it being there yet.

Maturity tracks platform count. Enterprises reporting a quarter or less genuine orchestration run 2.8 platforms on average; those in the 26–50% band run 3.5. The organizations furthest into real multi-step work are the ones running the most orchestration platforms at once, which is the practical case for the flexibility-first selection logic in Finding 2 — multi-step portfolios appear to accumulate platforms rather than converge on one.

One split that might be expected does not appear. Organization size makes no difference to portfolio maturity in this wave: 38% of enterprises at 10,000+ employees report a quarter or less genuine orchestration, against 37% of smaller ones, and the shares past the halfway mark are equally close (16% and 15%). Whatever separates the mature portfolios from the immature ones here, it is not headcount.

Finding 8: Fiscal control is still reactive for one in five

A fifth of enterprises learn about a runaway agent from the logs

Finally, we asked how enterprises enforce fiscal control over agent token consumption — the risk that an autonomous loop exhausts a budget before anyone intervenes. The approaches split four ways, fairly evenly.

One in five enterprises (21%) has no real-time, programmatic way to stop an agent before a budget-breaking bill arrives — they learn of it from the logs afterward. Another 30% lean entirely on the native caps and throttles built into their primary platform, a control only as good as the provider’s tooling and one that sits awkwardly beside the hybrid, keep-control-outside posture of Finding 6. Roughly half of enterprises — those building custom gateways (25%) or exploiting cross-model routing to arbitrage cost (24%) — are treating token burn as an engineering problem to be controlled deterministically, and the routing group is doing so in a way that only works because they run several platforms at once.

Unlike previous waves, no size split appears here: 18% of enterprises at 10,000+ employees exercise only reactive control against 23% of smaller ones, a difference well within sample noise. The gap in fiscal control in this wave is not between large and small enterprises but between those that have built a cost-control plane and those still relying on whatever their provider ships. Read against the satisfaction scores in Finding 1 — where value for money was the weakest of three ratings at 3.63 — the picture is of a cohort that is unhappy about what agents cost and, in half of cases, not yet instrumented to do much about it.

The bottom line: Plural by design, governed by intention, metered by hope

Organizations with 100 or more employees describe an orchestration strategy built around optionality rather than commitment. They run three platforms on average, choose them for flexibility across models rather than affinity to any one, and judge them on whether they carry multi-step work reliably to completion. Microsoft anchors the installed base and appears in seven of ten stacks; Anthropic leads forward consideration by a wide margin among the two-thirds planning a change; and a substantial minority intend to build their own control plane rather than buy one. Today’s footprint describes where these enterprises are, and clearly does not describe where they intend to stay.

The governance posture is deliberate and consistent. A hybrid control plane is the majority expectation, 78% intend to keep control at least partly outside the provider, and the reason is not commercial but operational — security and permissioning limits (37%) and limited visibility (22%) outrank vendor lock-in (23%) as the fear attached to provider-resident control. The budget follows the fear: monitoring and debugging and security and permissions enforcement together take 61% of planned investment growth, ahead of the tooling used to build agents in the first place.

Where the strategy thins out is cost. Portfolio maturity has moved into the middle — 47% now report between a quarter and half of their agents genuinely orchestrated, and pure-chatbot portfolios have nearly disappeared — but 21% still cannot stop a runaway agent in real time, another 30% depend on whatever caps their provider ships, and value for money is the lowest-rated attribute of the platforms they run. Enterprises have worked out how they want agents governed well before they have worked out how to meter them.

At 107 respondents in a single July wave, skewed toward large technology organizations, this reads as a clear directional signal rather than a precise measurement. The questions for subsequent waves are whether the middle band of portfolio maturity keeps climbing, whether the forward consideration for Anthropic and for in-house control planes converts into deployment, and whether fiscal control catches up to a cost that enterprises already say they are not getting their money’s worth on.


Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single July 2026 wave. This is a self-selected sample rather than a probability sample, and figures should be read directionally rather than as precise measurement. Respondents include software/ML engineers, product/program managers, directors and VPs of data/AI/analytics, enterprise architects, and directors of engineering/IT, across technology/software, government/public sector, manufacturing/industrial, and financial services organizations.

SpaceXAI's Grok Bot turns agents into persistent digital coworkers that can operate your apps for $120-per-month

SpaceXAI, the division of SpaceX formerly known as xAI, is launching an early beta version of Grok Bot, a new agent designed to move AI assistants beyond answering prompts and toward continuously executing work across the software employees already use.

The central idea is straightforward: instead of opening an AI assistant whenever a task arises, users create persistent Bots with specific jobs, give them access to applications and websites, and delegate work much as they would to a teammate.

Each Bot operates through its own computer environment, can continue working when the user's laptop is closed, and can return when it needs approval or has finished the assignment.

SpaceXAI says the system began as an internal prototype before spreading across the company, where teams created Bots for sales outbound, marketing campaigns, office operations, bug fixes and other work. The company is now turning that internally developed workflow into a product for external users.

“Bots are AI teammates that do real work for you,” the company said in announcing the product. “They sign in to your tools, use them just like you do, and come back with finished work.”

The company did not release benchmarks for Grok Bot's performance on agentic tasks. And it arrives amid an increasingly crowded marketplace of first-party AI agents that attempt to reliably complete real, enterprise workflows by interfacing with a user's other applications and devices.

Anthropic introduced computer use for Claude in 2024, allowing models to inspect screens and operate interfaces through mouse and keyboard actions, and continued expanding with the launch of the developer focused Claude Code harness in early 2025 and the more non-technical, white collar focused Claude Cowork agent early this year.

Meanwhile, OpenAI gave its Codex harness the ability to control other computer apps in April, launched agentic Workspace Agents that can also connect to third-party applications and use them autonomously, and recently debuted a new ChatGPT Work environment for longer, multi-step tasks and finished deliverables.

Grok Bot seeks to join the party with its own management model for agents: persistent workers with responsibilities, memory, learned routines and the ability to hand work to one another.

Pricing and availability: Grok Bot starts at $120 per seat per month for teams, $200 per month for individuals

Grok Bot is available beginning today, August 11 in beta for SuperGrok Heavy, Cursor Ultra and Cursor Premium Teams subscribers (recall SpaceX acquired Cursor for $60 billion back in June). The product arrives for macOS, Windows, Linux and iOS, with Android listed as coming soon.

According to its product page on xAI.com, Grok Bot is included with Cursor Ultra at $200 per month for individuals. The plan includes a computer for Grok Bot, access to users' tools, scheduled routines, desktop and mobile operation, and extended AI-token limits.

For organizations, Cursor Premium Teams costs $120 per seat per month and adds centralized billing and settings, a team marketplace for skills and plugins, shared usage analytics and SAML/OIDC single sign-on.

Existing SuperGrok Heavy ($300 per month) subscribers also receive access. However, for organizations wishing to sign up today, SpaceXAI is directing them to a waitlist for future access.

Those prices make Grok Bot a substantially different purchasing decision from a low-cost general AI subscription. The economic question for companies will be whether persistent Bots can replace enough manual work or conventional automation infrastructure to justify the per-user cost — and how usage limits affect total cost once agents begin running continuously.

From prompting an AI to managing one

SpaceXAI describes Grok Bot as a team of “always-on agents.” Users can create multiple Bots, assign each a role and let them work simultaneously.

The company provides examples including Sales Outbound, Talent Scout, Paid Media, Expense Manager, Product Performance, Bug Reproduction, Account Health and Chief of Staff. A sales Bot, for example, can research accounts, score prospective contacts, prepare email and LinkedIn outreach in the user's voice, and assemble the results for human approval.

Promotional materials show SpaceXAI using the system internally for substantially longer chains of work. One sales Bot can add call-transcript notes to a CRM and draft follow-up messages. An operations Bot can seat new hires and process invoices arriving through Gmail. An engineering Bot can reproduce a bug in the product interface, file a ticket and then hand the repair to a debugging Bot.

The architecture could make Grok Bot particularly relevant for workflows that span systems that were never designed for AI automation.

Rather than requiring every application to expose an API specifically for an agent, Grok Bot can sign into applications and websites and operate their interfaces. SpaceXAI says Bots have their own computers and can continue working 24/7.

The company explicitly says this includes websites and applications that have “no clean API or MCP,” an important distinction for enterprises with legacy software, fragmented SaaS environments or internal systems that have never been instrumented for agent access. Instead of limiting automation to formally integrated services, Grok Bot is designed to work through the same software interfaces a human employee would use.

The company says early users are already applying Bots to jobs including vendor negotiations, e-commerce customer support and continuously updating CRM systems.

Another feature attempts to reduce the engineering required to automate repeatable business processes. Users can demonstrate a workflow while a Bot follows along. Grok Bot can then save the process as a routine and execute it later without requiring the user to reproduce every instruction.

SpaceXAI says the Bot can also incorporate corrections into those learned routines, allowing the workflow to change as the user teaches it how a particular process should be handled.

That potentially changes the deployment model from explicitly programming an automation to teaching an agent how an employee performs the job.

The company is also claiming a more persistent form of behavioral memory than simply retaining a chat transcript. According to the launch announcement, Bots remember prior conversations, learn preferences such as a user's writing voice and edge cases, and gradually learn when they should interrupt for approval versus continue independently. SpaceXAI says they can later resume dropped threads, nudge stalled handoffs and pick up work from earlier conversations.

It further says Bots can become proactive over time, sometimes identifying work before the user explicitly asks for it. That is a more ambitious claim than conventional scheduled automation and will put additional pressure on permission controls and escalation rules if the system is deployed against production applications.

Bots can delegate work to other Bots

Grok Bot also supports multiple agents operating together.

Users can place several Bots into the same thread, where the agents can pass work between one another. The company's demonstration includes specialized Research, Communications, Chief of Staff and Travel Bots coordinating tasks.

SpaceXAI says those Bots can independently message one another and share context within threads. Users can also put multiple Bots into a group conversation where they assign ownership, transfer work and coordinate among themselves, bringing the human back in primarily for judgment calls.

Internally, the company says employees sometimes place a Chief of Staff Bot above specialist Bots responsible for functions such as inbox management, recruiting, expenses, operations and bug fixes. That makes the product's orchestration model more explicit: the user does not necessarily have to serve as the routing layer between every specialized agent.

Initial reactions are extremely positive

Lenny Rachitsky, host of the popular vlog and podcast Lenny's Podcast and author of newsletter Lenny Letter, received early access to Grok Bot and loved using it. As Rachitsy wrote on X : "I haven't been this excited about a new AI product in a while. It's like OpenClaw, but super easy, reliable, and less scary to use. I think this will be a huge new product line for Cursor/Grok/SpaceX."

Similarly Matt Shumer, an AI entrepreneur who said he tested Grok Bot for several weeks before launch, highlighted this orchestration as one of the product's strongest features.

“The best way I can describe it is an agent for everything, not just code,” Shumer wrote on X.

In one test, Shumer said he created separate researcher and writer Bots, then created a Chief of Staff Bot and instructed it to coordinate the other two on a project. He expected the workflow to break down.

“It worked out of the box,” he wrote.

His main criticism involved model selection.

Unlike systems where developers or advanced users explicitly select the underlying model, Shumer said Grok Bot automatically routes tasks to models on the backend.

“You don’t choose a model for your Grok Bot,” he wrote. “It’s all done automatically on the backend.”

Shumer said the model router “wasn’t great” during his testing, although he said he was subsequently told it had improved.

SpaceXAI's expanded announcement still does not identify which underlying models the router uses, nor does it document a mechanism for users to select, pin or switch to a particular xAI or third-party model. As a result, the model layer remains largely abstracted from users in the publicly supplied launch material.

That abstraction represents an important tradeoff for enterprise deployments. Automatic routing can remove a significant configuration decision for ordinary employees, but advanced users may want explicit control over model cost, latency, reliability and behavior — particularly for repeatable production workflows.

The agent market is moving toward longer-running work

Grok Bot enters a market increasingly focused on agents that can do more than generate text or code.

Anthropic's computer-use capability established a mechanism for Claude models to interact with software through screenshots, cursor movements, clicks and typing. Its broader Claude product also connects with workplace services and remote MCP servers.

OpenAI, meanwhile, now describes ChatGPT Work as an agent for “longer, multi-step work and finished deliverables,” while keeping Codex focused specifically on software development. OpenAI's enterprise agent economics can also incorporate usage-based credits, making task complexity and token consumption part of deployment cost calculations.

Grok Bot's differentiation is therefore less about proving that AI can operate software than packaging computer use, persistence, workflow learning and multi-agent coordination into something resembling a workforce interface.

SpaceXAI's announcement sharpens that distinction by emphasizing completion rather than assistance. One company product employee, identified only as Roman, describes the difference as closing the gap between work that is nearly finished and work actually completed inside the destination application: “Grok Bot can finish the swing, because the work lands where a human would put it, in the actual tool.”

That distinction will ultimately depend on reliability. A chatbot producing a bad answer creates a correction problem. An autonomous agent operating CRM records, support queues, vendor conversations or other production systems can create an operational problem.

Grok Bot's success will therefore depend not only on model intelligence, but also on permissions, predictable execution, escalation behavior, memory accuracy and how reliably agents recognize when human approval is necessary.

That challenge becomes more significant if Bots act proactively, resume forgotten work and coordinate with one another without the user serving as an intermediary. Those capabilities reduce the amount of supervision required when they work correctly, but they also expand the consequences of an incorrect assumption, stale context or improperly scoped permission.

The interface may matter as much as the models

Shumer described the product's interface as feeling like iMessage, an intentionally familiar metaphor for a system whose underlying architecture — autonomous computers, persistent memory, agent orchestration and automatic model routing — could otherwise be difficult for nontechnical users to configure.

SpaceXAI makes essentially the same usability argument in its launch announcement. Rather than asking users to construct workflows before getting started, it says users can simply message a Bot from a phone or desktop, hand it work and later continue the same conversation from either device.

That simplicity is part of the product strategy. Grok Bot is trying to hide much of the conventional machinery of automation — workflow builders, explicit integrations, agent routing and orchestration — behind an interaction model that resembles messaging a coworker.

That may prove to be the larger bet behind Grok Bot.

The AI industry has spent several years making models increasingly capable of using tools and completing multi-step tasks. Grok Bot attempts to turn those capabilities into an organizational abstraction people already understand: give someone a job, teach them how you work, and let them coordinate with the rest of the team.

If that abstraction proves reliable, the enterprise agent competition may increasingly shift away from which assistant produces the best individual response and toward which platform can most reliably manage fleets of agents performing ongoing work.

Why AI-driven purchase intent so rarely becomes a completed sale

11 August 2026 at 15:00

Presented by Rezolve Ai


When an AI assistant recommends a product or brand, it generates something valuable: a purchase-ready consumer with high intent and low friction in their decision. That consumer has already compared options, asked follow-up questions, and arrived at a conclusion. They want to buy.

What they encounter next is a commerce infrastructure that was not designed for them.

The gap between recommendation and purchase

The typical enterprise commerce stack was built for a specific model: a consumer who arrives at a brand's website through search or a direct link, navigates product pages, adds to cart, and completes checkout through a multi-step form flow. That model assumed the consumer would do the work of bridging their intent to the transaction. Most commerce systems still assume exactly that.

Agentic commerce breaks that assumption. When intent is generated outside the brand's owned environment, the handoff to transaction becomes a structural problem. Context doesn't transfer. Sessions don't persist. The consumer who asked an AI assistant for a recommendation and received one now faces the same friction-laden checkout process as someone who arrived with no prior intent at all.

Cart abandonment rates have remained stubbornly high for years. Baymard Institute research puts the average at 70%. That figure predates the agentic commerce era. As more purchase intent is generated through AI interfaces, and as the gap between that intent and a brand's transaction layer widens, the abandonment problem is likely to get structurally worse before it gets better.

What the current stack wasn't built to handle

The commerce infrastructure most enterprises operate today was assembled over two decades of incremental investment. Each layer added a capability: a search tool, a recommendation engine, a personalization layer, and a checkout system. Each was built to solve a specific problem within a human-initiated shopping journey.

None of it was built to receive intent from an AI agent.

When an AI system generates a purchase recommendation, it needs to do more than surface a product page. It needs to verify real-time inventory. It needs to apply pricing logic and promotional rules. It needs to respect brand policy around which products can be recommended together, which channels apply which discounts, and what the correct fulfillment path looks like for a given consumer. And it needs to do all of that without breaking the conversational context that made the recommendation possible in the first place.

Current commerce stacks can't do this reliably. The systems that hold the relevant data, inventory, pricing, order management, fulfillment, are not exposed in ways that AI agents can safely and accurately access. The result is a journey that starts with intelligence and ends with a broken experience: a link out to a product page, a generic checkout flow, and a consumer who arrived ready to buy and left without completing the transaction.

The conversion problem is an architecture problem

The industry has treated conversion optimization as a front-end problem for most of its history: better copy, cleaner checkout UX, fewer form fields, smarter retargeting. Those interventions were appropriate for the model they were built to serve.

The agentic commerce era introduces a different kind of conversion failure, one that front-end optimization cannot fix. When intent is generated externally, conversion depends on whether the back-end infrastructure can receive that intent, act on it accurately, and complete the transaction within the guardrails the brand has established. That is not a UX problem. It is an infrastructure problem.

Brands that are investing heavily in AI-powered discovery while leaving their execution layer unchanged are widening the gap between the promise AI makes on their behalf and the experience they can actually deliver. That gap has a cost, measured not just in lost transactions but in consumer trust that erodes each time the promise and the reality don't match.

Rezolve Ai commissioned research across 1,500 US consumers in January 2025 that found consumers who encounter friction immediately after an AI recommendation are significantly less likely to complete a purchase than those who encounter friction at the top of a traditional funnel. The implication is direct: AI raises the expectation bar at the moment of intent. Brands whose infrastructure cannot clear that bar are paying a conversion penalty they may not even know they're incurring.

What closing the gap requires

Closing the gap between AI-generated intent and completed transaction requires rethinking which layer of the commerce stack carries the most strategic weight in an agentic world. For most of the past decade, that weight sat with discovery and experience. The brands that invested most in search, personalization, and content won a disproportionate share.

In the agentic era, the weight shifts to execution. The brands that can reliably take AI-generated intent and turn it into a governed, accurate, brand-safe transaction will have a structural advantage over those whose infrastructure stalls at the handoff.

That is a different investment thesis than the industry has operated on. And most enterprise commerce roadmaps have not yet caught up to it.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Mistral AI wants to build 1 gigawatt of European compute by 2030 — and lock in customers now.

Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached.

The French artificial intelligence company announced Tuesday a three-part expansion of its infrastructure business: regional inference endpoints that let customers choose whether their AI workloads run in Europe or the United States, a new "Priority Tier" backed by an uptime guarantee for mission-critical deployments, and a coalition of European enterprises making multi-year compute commitments that Mistral says will underwrite 200 megawatts of infrastructure across Europe by the end of 2027 — and a full gigawatt by the end of 2030.

In a move that may raise eyebrows among sovereignty purists, the company also said it will begin hosting third-party open models on its platform, starting with GLM-5.2 from Z.ai, the Chinese AI lab formerly known as Zhipu.

Taken together, the announcements mark a decisive shift in how Mistral positions itself. The company that built its reputation training open-weight language models is now selling something closer to critical infrastructure: assured capacity, regional control, and contractual reliability for enterprises and governments that want frontier AI without surrendering control over where it runs.

"When we spoke in June, the story was around how Mistral was building a full-stack AI offering," Timothée Lacroix, Mistral's co-founder and chief technology officer, told VentureBeat in an exclusive interview ahead of the announcement. "Today, the announcement is about strengthening one part of this infrastructure, which is the inference part."

That one part, it turns out, comes with a price tag measured in the tens of billions of dollars.

Inside Mistral's plan to build 1 gigawatt of European AI compute by 2030

The headline numbers deserve scrutiny, because they imply staggering capital requirements. Mistral currently operates less than 200 megawatts of capacity, according to the company. Details shared with VentureBeat show the near-term buildout resting on three sites: a 44-megawatt facility near Paris that became operational in the second quarter of this year, a 23-megawatt facility in Sweden built in partnership with EcoDataCenter using renewable energy and advanced cooling, and a 10-megawatt site in Les Ulis, France, that came online in the third quarter.

Getting from there to one gigawatt by 2030 is a different order of magnitude. Independent estimates suggest just how different: research firm Epoch AI calculates that a typical one-gigawatt AI data center requires roughly $38 billion in upfront capital expenditure, with servers and GPUs — not buildings or land — consuming the majority of the cost. Goldman Sachs Research pegs next-generation AI facilities at $15 million to $20 million per megawatt before accounting for the chips inside them.

Lacroix did not dispute the scale of the challenge. The investment required for a gigawatt of capacity "is a large investment that requires also a lot of scaling and revenue behind it," he said.

The urgency, in his telling, comes from a supply crunch that is about to get worse. "More and more, and especially around 2027 and 2028, we see that the demand for AI compute is exceeding what the market has to offer, especially in Europe," Lacroix said. McKinsey has estimated that meeting global AI demand could require $5.2 trillion in data-center capital expenditure by 2030 — and Europe, by most analyses, is starting from behind.

A company valued at a fraction of its American rivals cannot close that gap with venture capital alone. Which explains the most consequential — and most unusual — piece of Tuesday's announcement.

European Compute Units turn AI sovereignty into a five-year contract

Mistral is assembling what it calls an anchor group of enterprises whose long-term commitments will collectively finance infrastructure none of them could justify alone. Those commitments convert into "European Compute Units," or ECUs — a claim on Mistral-built capacity over multiple years that participants can spend on inference, training, model adaptation, or other AI workloads as their needs evolve.

If that structure sounds more like a power-purchase agreement than a cloud contract, that appears to be the point. Data-center financing increasingly resembles large infrastructure projects — gigawatts, substations, energy agreements — rather than traditional technology spending, and lenders want demand locked in before capital gets deployed. Mistral raised €830 million ($962 million) in debt earlier this year to fund its data center near Paris, TechCrunch reported in March, and pre-committed enterprise demand is exactly what makes that kind of financing repeatable at ten times the scale.

Lacroix was unusually direct about the mechanics. "The entire point of compute units is to have commitment," he said. "The goal is to have customers commit for around five years, or at least a long time." Asked what happens if a customer wants out early, he didn't soften the answer: "There is no getting out."

What makes a five-year, no-exit commitment palatable, he argued, is flexibility in how the capacity gets consumed. "Typically this can be spent on raw inference that you then feed through any other AI stack. It can be spent on raw compute as managed Kubernetes, and it can be spent at the very top with our full AI offering," he said. "My hope is that they will use it with our full-stack services and will love it."

The anchor group already includes some of Europe's industrial heavyweights. Amadeus CEO Luis Maroto said in a statement that "capacity, deployment control, and operating continuity become increasingly important for all enterprises." ASML chief Christophe Fouquet — whose company led Mistral's $13.4 billion (€11.7 billion) Series C last year — called building European AI capacity one of the few industrial endeavors that "will matter more to Europe's next generation," while Capgemini's Aiman Ezzat framed it as "a question of who shapes the future of European industry." CMA CGM chairman Rodolphe Saadé said the shipping group's Mistral deployment is "already under way among thousands of employees."

Commitments of that duration only make sense, of course, if the sovereignty being purchased is real. On that question, Mistral's announcement contains an asterisk worth reading closely.

The fine print on sovereign AI: what data can still leave Europe

The centerpiece product is Mistral Regional Endpoints, now generally available, which let customers pin inference and its associated processing to Europe or the U.S. Alongside it, the new Priority Tier — in public preview — offers committed service levels, custom rate limits, and an uptime SLA for mission-critical workloads.

Mistral claims it is the only European AI lab offering both a choice of processing region and an SLA-backed service tier, and Lacroix said a third option is coming: an endpoint "that stays on Mistral-controlled infrastructure, so on Mistral compute" — for customers who want their inference not just in Europe, but off hyperscaler hardware entirely.

Then comes the fine print. Mistral's own materials note that in-region inference remains subject to "limited, safeguarded transfers" to sub-processors that may sit outside the chosen region. Pressed on what actually leaves Europe, Lacroix pointed to the connective tissue of modern AI applications: tool calls.

"There are some tool services, like some tool calls, that might be hosted in places where we don't fully control this," he said, citing web search as an example. "A few of our web-search providers might not all be in Europe, and in that case, we need to potentially gate that capability."

His answer to the compliance question — would this satisfy a European bank or a defense ministry? — was that gating is the feature, not the bug. Capabilities that cannot be sourced in-region can be switched off entirely, restricted to certain users or workspaces, or, given sufficient demand, rebuilt with European providers. "Any capabilities that we don't find a provider for in Europe — if it needs to be done in Europe, we'll find some way to implement it or find ways to address it," Lacroix said.

For enterprise buyers, that is a more honest framing than most sovereignty marketing offers: full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision rather than a default. The same pragmatism runs through the announcement's most surprising line item.

Why Europe's open source AI champion is hosting China's GLM-5.2

A French national champion — one that has partnered with the French army and positioned itself as Europe's answer to American AI dependence — hosting a Chinese lab's model invites an obvious question. Lacroix's answer was disarmingly matter-of-fact.

"It's a great model. Everyone loves it. It's open weight, so there was no good reason for us not to do it, really," he said, noting that Mistral's own stack is already built on open-source software like Kubernetes.

On security vetting, he argued that open weights fundamentally change the risk calculus. "The risks in taking a new model, at the layer of the weights, are — at least in my opinion — rather limited," Lacroix said. "We checked basically all of the safety and compliance evals that we have. We'll control that model, its outputs, and what it does the same way we do any of our models. We have the same inputs and outputs and monitoring capabilities over all of it."

The strategic logic is worth unpacking. By hosting third-party open models under European regional controls and the same SLAs as its own, Mistral is repositioning itself from model vendor to sovereign distribution layer — the trusted intermediary through which any open model, regardless of origin, can be consumed by a regulated European enterprise that could never call a Chinese API directly. It is the "model garden" playbook the hyperscalers run with Bedrock and Vertex, executed on European soil with European guarantees.

Customers appear to be reading it that way. "Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements," Matan Griberg, CEO of AI software-engineering company Factory, said in a statement.

Lacroix stressed the move is not a retreat from frontier training: the model Mistral had in training as of June "is still training, and we're still very excited about it," he said. But openness to rivals' models signals where the company now believes its moat lies — not in any single model, but in the infrastructure underneath all of them. Which makes its relationship with the world's most powerful infrastructure company all the more interesting.

How the multibillion-dollar Microsoft deal funds Mistral's independence

Hovering over every sovereignty claim is Mistral's deepening relationship with Microsoft. In July, the two companies announced a multibillion-dollar expansion of their partnership under which Microsoft will rent capacity from Mistral's European data centers to serve its own cloud and AI demand, while adding Mistral Medium 3.5 and OCR 4 to Microsoft Foundry, bringing Medium 3.5 to Copilot Studio, and enabling Mistral models on Azure Local for disconnected, customer-controlled environments. Mistral CEO Arthur Mensch told The Wall Street Journal at the time that two-thirds of Mistral's customers already work with Microsoft.

How does a company selling independence from U.S. hyperscalers square taking one on as its largest tenant? Lacroix described Microsoft not as a patron but as an anchor customer that de-risks the buildout.

"It allows us to scale different parts of the business differently by building infrastructure with Microsoft as a customer," he said. "We can scale that team, we can scale our infrastructure, and make sure that we can then, on the side of it, also build for ourselves and for our customers." He compared the arrangement to the neocloud playbook — companies that built businesses supplying capacity to the hyperscalers themselves. "As that part of our business resembles that of neoclouds, we're following the same thing."

It is a genuinely clever inversion: rather than renting American infrastructure, Mistral is renting infrastructure to one of America's largest companies, using Microsoft's demand to finance capacity that also serves European sovereignty customers. But the independence has limits no contract can engineer away — the GPUs filling Mistral's European data centers come overwhelmingly from Nvidia and other American chipmakers, as SiliconANGLE noted in its coverage of the July deal.

Asked directly why a customer should choose Mistral over an EU region on AWS or Azure, Lacroix gave two answers. "The simplest possible answer is capacity. There is more demand than supply right now, and so it adds another option," he said. The second cuts closer to the pitch: "We are a European provider, and on the region that would be Mistral compute, we are fully independent. That's a truly differentiated offering than all of the hyperscalers or pure inference companies can provide."

The economics of open models: why agentic AI is pushing inference to the cloud

There has always been a tension at the heart of Mistral's business: its best-known models are free to download, and open models have historically been difficult to monetize through APIs. Asked how free weights fund a gigawatt buildout, Lacroix offered the clearest articulation yet of the company's thesis — that the economics of self-hosting are collapsing under the weight of the models themselves.

"When the models were smaller, and we were before the explosion of agentic AI, it was doable for enterprises to host their own — up to, let's say, 100-billion-parameter dense models — on their premises," he said. "More and more, with models going into the trillion or more parameters, with the current hardware, and with the increasing amount of tokens that need to be processed, it becomes harder."

His conclusion was blunt: "I don't see how, with the current trend of model size and growth of agentic tokens, we keep the full inference on-prem. To me, that is why we think we're going to monetize our cloud inference." Inference, he noted, is particularly well suited to the cloud because it "does not need to hold any data" and can be encrypted in transit.

In other words: open weights get Mistral into the enterprise, and the physics of trillion-parameter agentic workloads brings the inference — and the revenue — back to Mistral's data centers. The thesis will get an expensive test. Mistral has raised roughly $4 billion to date, according to PitchBook data — a fraction of the war chests assembled by OpenAI and Anthropic — and Bloomberg reported in June that the company is in talks to raise about €3 billion at a roughly €20 billion valuation, nearly double its Series C mark. The revenue behind the buildout will have to come from exactly the enterprises Tuesday's announcement is courting.

And Europe, in Mistral's telling, is only the first market for what it is selling. Asked whether the framework could be replicated in the Middle East, Asia, or anywhere else anxious about AI dependence, Lacroix didn't hedge: "It's completely right. We're starting this in Europe because it's also an easier part of the world for us to scale into, especially in the infrastructure. But we definitely want to extend this, depending on customer demand." Every layer of the stack, he said, "can be controlled, changed, replaced depending on where we operate and what the requirements are — that's pretty much where we excel."

That is the wager underneath the SLAs, the compute units, and the Chinese model flying a European flag: in a world where the U.S. and China dominate frontier AI, the durable business is selling everyone else control. To fund it, Mistral is asking Europe's largest enterprises to sign five-year contracts with no exit — while making a bigger, longer commitment of its own. A gigawatt, after all, is a promise measured in decades. For Mistral, too, there is no getting out.

Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests

11 August 2026 at 13:00

Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes.

Nvidia is proposing a fix that touches both ends of that problem at once. The company is out on Tuesday with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume, specialized agent tasks, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best.

The headline numbers: According to Nvidia, Lightning delivers up to 4x faster output than comparable models in its class, completing agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy. Paired through Switchyard, Nvidia says the combination holds frontier-level task completion while cutting benchmark costs to roughly a third of running Opus 4.8 alone.

The timing puts Nvidia in the middle of the busiest open-weight stretch the industry has seen in months. Alibaba, Moonshot, Zhipu and DeepSeek have all shipped competitive open models out of China since the spring, several landing at or near frontier performance while undercutting US labs on size or price. Meta added to that pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer. Open weights have gone from a differentiator to table stakes in a matter of months, and Nvidia's release lands squarely inside that shift rather than ahead of it.

The pairing is the point. A model alone doesn't solve the cost problem, and a router alone has nothing efficient to route to. Nvidia is betting that open source, applied at both the model layer and the routing layer, is what actually moves the cost needle on agentic AI, not a single cheaper model and not a smarter router bolted onto someone else's stack.

Switchyard's real rivals aren't other open models — they're Not Diamond, which already powers OpenRouter's Auto mode, and RouteLLM, the open-source framework from UC Berkeley and LMSYS. Neither ships its own model. Nvidia's bet is that owning both sides of the decision, under one open license, is what a router-only or model-only competitor can't match.

"That is the power of a system of models, matching the right model to each step of the workflow," Kari Briski, vice president of generative AI at Nvidia, said in a briefing.

How the router actually changes the workflow

Model routing isn't a new category. OpenRouter, LiteLLM and a handful of standalone routing startups already let developers point traffic across multiple providers. Switchyard plugs into several of them rather than replacing them outright.

The core problem Switchyard solves is that the right model changes as an agent moves through a task. An agent's state shifts as tools return results, errors show up, or a step turns out to be routine rather than complex, and a fixed model choice can't adapt to any of that.

Briski described routing strategies that respond to that shifting state rather than a static task category.

"It has many types of routing strategies," Briski said. "You can have a random router, which is not that great, or you can have an agent state route or a classifier route. Depending on your routing strategy, it wants to choose the best model. In some cases you want to go with a model like Lightning for really efficient tasks, and the router will actually choose Lightning if it's set up in your pool of models."

Cost enters the routing decision directly, not as an afterthought. In response to a question from VentureBeat, Briski said Switchyard can evaluate model verbosity, meaning how many tokens a given model tends to produce for a task, and use that prediction to steer work toward the cheaper option before the call is made.

The part that keeps this from becoming its own integration project is where Switchyard sits. Nvidia split its partners into two groups: agent frameworks that call Switchyard directly, including Cognition, LangChain and Nous Research, and LLM gateways that have built Switchyard support into their own products, including Kong, LiteLLM and OpenRouter. Kong ships Switchyard natively inside Kong AI Gateway. Briski pointed to that same list of gateway partners when describing how the library fits into the existing routing ecosystem.

"We are an ecosystem lover, and we want to make sure that we are integrated," Briski said. "We've partnered with OpenRouter, LiteLLM and Kong, and they've already integrated our routing algorithm, so you can pick it up right where you're already using the best tools."

Nvidia shared results from nine companies testing Switchyard, several with specific figures attached. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, at a 6% accuracy tradeoff. Ramp said it matched a frontier model's performance on Ramp SWE-Bench while cutting costs 58% and runtime 33%. Cognition integrated Switchyard's staged router into Devin Desktop for internal use and reported near-frontier performance on FrontierCode Main while cutting mean cost 28% relative to routing everything to a single frontier model.

Lightning's architecture and performance gains

Nemotron 3.5 Lightning is a standalone open model in its own right, built for high-volume, specialized agent tasks rather than general-purpose use.

It extends the hybrid Mamba-Transformer, latent mixture-of-experts architecture Nvidia introduced with the Nemotron 3 family in December 2025, the same line behind Nemotron 3 Super, which Nvidia uses as Lightning's own baseline in its post-training comparisons. Positioned within a routing setup like Switchyard, it's built to sit at the fast, cheap end of the decision rather than the frontier end, but it runs and ships independent of any router.

According to the Artificial Analysis Intelligence Index, a general capability benchmark spanning nine evaluations, Lightning scores 24, tied with gpt-oss-120b and behind Nemotron 3 Super, Gemma 4 31B, Claude 4.5 Haiku and Mistral Medium 3.5, all at 30. Lightning isn't a general-intelligence leader in its size class, and Nvidia isn't claiming it is.

The actual claim is narrower: according to PinchBench data supplied by Nvidia, Lightning matches Qwen3.6-35B's accuracy roughly 30% faster and beats Gemma 4 26B's accuracy at a similar completion time on PinchBench, a real-world agent task benchmark spanning coding, research and file management. That's a speed-to-accuracy tradeoff, not a capability win.

Post-training is where Nvidia says the bigger gains show up. The company shared before-and-after figures from four early-access partners: CrowdStrike's malicious-content recall against a Nemotron 3 Super baseline, CodeRabbit's coding router against a GPT 5.4 Nano baseline, Harvey and Trajectory's legal task completion against an Opus 4.6 baseline, and Lila Sciences' energy simulation work against an Opus 4.8 baseline. CodeRabbit's case is the most specific: Nvidia says the standard NeMo Auto model recipe, trained for one epoch, built into a working router agent for $85 in about two hours.

What this means for enterprises

There is no shortage of competitive offerings in the growing market for open models. The new Nemotron Lightning release will be yet another option for organizations to consider.

On the model side, Lightning's own benchmark chart picks Qwen3.6-35B as its direct comparison point. Asked by VentureBeat directly how Lightning compares to Chinese models more broadly, Briski didn't offer a head-to-head benchmark, pointing instead to openness and customizability as the differentiator.

"Our value proposition is not just open and it's very customizable," Briski said.

For enterprises building agentic infrastructure, three trends stand out:

The routing decision is becoming dynamic instead of static. Enterprises that built agent pipelines around a single default model are being pushed toward per-step routing based on live signals like agent state and token cost, not a fixed assignment set at design time.

Open source is now a cost lever at two layers, not one. Pairing an open model with an open router a vendor controls end to end is a newer argument than cheaper weights alone, and worth watching for whether other labs follow the same pattern.

The competitive question shifts from best model to best system. As routing libraries mature, the differentiator moves from which model an enterprise defaults to, toward how well its routing layer matches models to tasks in production, a harder thing to benchmark and a harder thing to market.

Your AI agent may be ready. Your sales motion probably isn’t.

11 August 2026 at 13:00

Presented by Salesforce


Interested buyers don't generate revenue. Live customers do. That's the lesson I keep drawing from watching hundreds of ISV partnerships navigate the agent economy over the last 18 months.

The companies pulling ahead aren't winning on features. They're winning because customers can move from discovery to live deployment in hours, while competitors are still negotiating contracts, clearing tax reviews, and waiting on provisioning.

That gap between a buyer who says “yes” and a customer who is actually using the product is where too many deals lose momentum. Urgency fades. Champions move on. Competitors get another opening.

Gutenburg saw that gap firsthand. Healthcare organizations valued its product, but sales cycles stretched 30 to 45 days. With custom pricing via AgentExchange, the company closed an urgent healthcare deal in just 48 hours.

Not 48 days. 48 hours.

The final contract phase alone dropped from 4 hours to 4 minutes. A 60x improvement.

I see this pattern across the ISV ecosystem. Building agents is getting faster. Getting buyers live before urgency fades is becoming the constraint. In a market moving this quickly, that can matter as much as the agent itself.

It’s like building a bullet train and selling tickets by fax. The product is built for speed. The transaction is not.

Distribution beats product in crowded markets

Nearly every software company is pouring resources into agent development. Far fewer are rethinking the path from discovery to deployment. Manual contracts, custom invoicing, tax reviews, provisioning delays, these are the handoffs that turn a 48-hour deal into a 45-day cycle.

That friction is now a competitive disadvantage, because the buying process is changing faster than most back offices are. Gartner predicts that by 2028, 90% of B2B purchases will be guided by AI agents.

That does not mean humans disappear from enterprise buying. It means the discovery and evaluation process changes. Buyers will increasingly use AI to identify, compare, and narrow solutions.

If your agent is not discoverable where that evaluation is happening, you may never make the shortlist.

A better agent can still lose to one that's easier to buy.

Domain expertise matters. Workflow depth matters. Proprietary data matters. Customer context matters.

But enterprise categories are getting crowded fast. In crowded markets, the best product does not always win. The product that is easiest to discover, buy, deploy, and scale often has the advantage.

As agent-guided buying takes hold, the first evaluation may happen before a demo is scheduled or a sales rep is in the room.

AI agents will increasingly scan marketplaces, compare solutions, and help narrow purchase decisions in the time it used to take to schedule a discovery meeting.

Companies that figure out marketplace distribution now will own their categories.

That is the problem AgentExchange was built to address. It’s a single destination for apps, agents, and integrations that extend and connect to Salesforce and Slack, helping customers get more from their platform investments.

But discovery is only the first step. The bigger question is what happens after the buyer says “yes”.

“Yes” doesn't mean live

Enterprise software teams spend enormous energy getting to "yes." But in many deals, that is where the operational work begins.

Between “yes” and “live,” the back office can generate a chain of handoffs: contracting, invoicing, tax calculation, licensing, provisioning, fulfillment, payment, and finance reconciliation. Every handoff delays activation for the customer and delays recognized revenue for you.

For AI agents, that back-office drag is becoming a front-office problem.

AgentExchange brings discovery, commerce, and activation together, helping partners manage custom pricing, billing, licensing, provisioning, and fulfillment through one connected experience.

"AgentExchange removes the traditional procurement friction that slows deals. Customers can now discover, purchase, and deploy PandaDoc directly through their existing Salesforce contract, turning what used to be a multi-week process into a same-day activation." Keith Rabkin, CEO at PandaDoc

What closing in 48 hours actually looks like

Gutenburg’s 30-45 day cycles were eaten up by contract logistics. Sales moved faster than their back office.

Using custom pricing and automated transaction capabilities through AgentExchange, they streamlined contracting, tax calculation, provisioning, and other steps between buyer interest and activation.

When a healthcare organization needed a tool to help them create documents aligned to the Americans with Disabilities Act and accessibility requirements, Gutenburg closed in 48 hours from first contact.

The 48-hour close is the differentiator. It is what efficient growth actually looks like in practice. Revenue scales without scaling headcount. Pipeline coverage improves because you are discoverable everywhere. Net recurring revenue increases because customers expand through the same frictionless channel.

"AgentExchange condenses contracting and tax calculations into a 10-minute process with improved accuracy," said Zamial Jones, VP of Customer Success at Gutenburg. "For partners spending hours on these tasks for every deal, that's transformational."

The window is closing faster than you think

The app economy took a decade to mature.

The agent economy won't.

The ISV partners I've watched pull ahead aren't the ones with the most sophisticated agents. They're the ones who treated distribution as a product problem — resourced, measured, and iterated — before the category consolidated around them. The ones still treating go-to-market as a post-launch consideration are consistently 6 to 12 months behind.

You can spend the next two quarters perfecting your agent's reasoning capabilities. Or you can spend them making sure customers can actually buy it.


Salesforce is investing in the next generation of companies creating agents with $50 million through the AgentExchange Builders Initiative—capital, engineering support, co-marketing, and co-sell programs. Companies that move now will define what enterprise AI distribution looks like for the next decade. Learn more here.

Lisa Eisenberg is SVP of ISV Partnerships at Salesforce.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

LTX-2.5 can generate a 10-second AI video from an image in just 6.8 seconds on Nvidia superchips — and it's open weights

LTX, the open world model company spun out of Lightricks, today released LTX-2.5, the newest version of its open-weights video and "world" model and it arrives natively integrated into ComfyUI, the node-based workflow tool that has become the de facto prototyping environment for open generative media, through a strategic day-one launch partnership between the two companies.

The model is available now as open weights on Hugging Face, inside ComfyUI, and through the LTX API for teams that want managed generation. It is free to use for organizations under $10 million in annual recurring revenue; larger companies negotiate a license. LTX says its models have passed 33 million downloads, making the LTX family the most-used "open world" model line on the market.

Ahead of the launch, VentureBeat spoke exclusively with LTX co-founder and CEO Zeev Farbman and ComfyUI co-founder and CEO Yoland Yan about the release, the partnership, and why both companies are betting that open weights — not closed APIs — will win the video and world model market.

"We're trying to maintain the same efficiency and the inference speed that we're known for, but constantly pushing the quality up," Farbman said. "We are introducing many cool things in this release: multi-shot support, a diffusion decoder for better quality, new conditioning modes, better support for autoregressive models that are critical for real-time use cases and robotics."

What's new in LTX-2.5

According to the company's announcement, LTX-2.5 rebuilds nearly every stage of the generation pipeline rather than bolting new capabilities onto an older core. The headline changes:

  • A new diffusion video decoder that reduces visual artifacts in high-motion footage and reconstructs fine detail like text and faces, while preserving LTX's high compression ratio.

  • Native multishot generation that renders a full sequence as a single output, holding character, scene, and voice consistent across cuts rather than stitching individually generated shots together.

  • A custom Gemma 4 language backbone and dedicated prompt enhancer for more accurate handling of complex, multi-subject prompts.

  • A pretrained checkpoint tuned for physical AI and robotics giving teams a base to fine-tune on domain data that looks nothing like cinematic video.

  • A substantially improved distilled model that delivers near-full-model quality at lower cost and faster inference, and, through an optimization effort with NVIDIA, runs locally on NVIDIA RTX GPUs with reduced memory requirements.

The company claims roughly one-eighth the cost and one-seventh the render time of comparable models, with output that runs on hardware ranging from data center GPUs down to a Mac.

Checked against published rates, the cost multiple doesn't survive contact with the models that publish pricing.

LTX-2.5 generates 720p video with audio at $0.09 per second on its Fast tier, putting a 10-second clip at $0.90 — genuinely cheap, but about one-quarter the cost of full Veo 3.1 ($4.00), half of FLUX 3 Video ($1.70) and HappyHorse 1.0 (~$1.82), and only 10% under Google's budget tiers, Veo 3.1 Fast and Gemini Omni Flash ($1.00 each), while Veo 3.1 Lite ($0.50) is actually cheaper.

Nothing in the published field costs eight times LTX's rate; if the one-eighth figure holds anywhere, it would be against premium models like Kling 3.0 Pro or Seedance 2.5 that don't publish comparable per-second pricing — or against self-hosting the open weights, where the marginal cost is whatever your GPU costs to run.

The render-time multiple is better supported, at least by LTX's own end-to-end measurements: 6.8 seconds for a 10-second clip against 52 seconds for the fastest rival API (Gemini Omni Flash) is roughly one-seventh — though that figure comes from self-hosting on two GB200 superchips, and through LTX's own managed API the same job took 23.7 seconds, cutting the advantage to about half.

Here is how the published rates compare, normalized to the cost of a finished 10-second 720p clip with synchronized audio — the configuration LTX and Black Forest Labs have both used for their own evaluations:

Rank

Model

Per second

Per 10-second clip

Unique differentiator

Notes

1

Veo 3.1 Lite (Google)

$0.05

$0.50

The category's price floor — cheapest published rate anywhere

No 4K, no clip extension

2

LTX-2.5 Fast (Lightricks)

$0.09

$0.90

Only open-weights model in the field — self-host free under $10M ARR, fine-tuning permitted

Scales to 4K at $0.30/sec; up to 20s single generation at 24/25 fps

3

Veo 3.1 Fast (Google)

$0.10

$1.00

Cheapest closed-API path to 4K ($0.30/sec)

Budget tier of the Veo line

3

Gemini Omni Flash (Google)

$0.10

$1.00

Independently measured quality leader — tops both Artificial Analysis text-to-video arenas as of Aug 2026

720p only; 10-second maximum; best iteration tooling

5

LTX-2.5 Pro (Lightricks)

$0.12

$1.20

Quality-tuned tier of the only open-weights family — prompt adherence, faces, typography (vendor-described)

Tops out at 1080p and 10 seconds

6

FLUX 3 Video (Black Forest Labs)

$0.17

$1.70

First to ship 20-second single-generation clips with audio (July 2026) — a ceiling since matched by LTX-2.5 Fast

HD band; audio included; Draft tier at $0.06/sec ($0.60/clip, HD only)

7

HappyHorse 1.0 (Alibaba)

~$0.182

~$1.82

Arena quality leader at launch (April 2026), since overtaken; "open source" claims never matched by verified downloadable weights

Third-party reseller rate; audio included at no extra charge

8

Veo 3.1 (Google)

$0.40

$4.00

Only model supporting clip extension beyond a single generation

Premium tier; 8x Veo 3.1 Lite; 1080p at no premium over 720p

Sources: LTX API pricing documentation; bfl.ai/pricing; Google AI for Developers model pricing; HappyHorse reseller rates via third-party API platforms; Artificial Analysis text-to-video arena leaderboards. All rates verified August 11, 2026, and subject to change.

How fast and how good LTX says it is

The most eye-catching number in LTX's launch materials is speed: the company says LTX-2.5 generates a 10-second, 720p image-to-video clip in 6.8 seconds faster than real time.

The caveat is the hardware behind it. That figure was measured self-hosted on two of NVIDIA's top-end GB200 chips at steady state, a configuration far beyond what most teams have racked; the same job through LTX's own managed API took 23.7 seconds, albeit rendered at the higher 1080p resolution (the API has no 720p tier).

By the company's end-to-end measurements of competing APIs on the same task, Google's Gemini Omni Flash came in at 52 seconds, xAI's Grok 1.5 at 63 seconds, Google's Veo 3.1 at 70 seconds (for an 8-second clip), MiniMax H3 at 180 seconds, ByteDance's Seedance 2.5 at 317 seconds, and Kuaishou's Kling 3.0 Pro at 398 seconds.

On quality, LTX shared results from blind, side-by-side human preference tests, in which evaluators voted on videos generated from the same prompt without knowing which model produced which.

LTX-2.5 recorded a 67% win rate, narrowly ahead of Seedance 2.5 at 65%, with Gemini Omni Flash at 55%, MiniMax H3 at 50%, Seedance 2.0 at 44%, Wan 2.6 at 42%, and FLUX 3 at 28%.

All of these figures are vendor-reported measured or commissioned by LTX itself, not independently verified and the company labels the preference results preliminary, noting it expects them "to evolve as evaluation expands." They are directional claims a buyer should test against their own workloads rather than settled rankings.

The independent benchmark that does exist cuts the other way for now: as of this month, Gemini Omni Flash — which LTX's commissioned tests place 12 points behind its own model — leads both of Artificial Analysis' text-to-video arena leaderboards, and the arena does not yet score LTX-2.5 at all. Until it does, the 67% figure remains untested on neutral ground.

The launch materials also lean on deployment terms rather than raw performance: LTX-2.5 runs on any GPU with a minimum of 16GB of VRAM, deploys on-premises, at the edge, or via API, carries no visible watermark on output — though the license requires users to disclose that content is machine-generated and forbids removing any embedded provenance or "latent disclosure" features (more on this below) — and can be fine-tuned on a customer's own data and IP flexibility the company contrasts with closed API-only rivals and with open-licensed competitors whose weights are unavailable in the U.S. and Europe or whose licenses restrict fine-tuning.

Betting against the API business model

For Farbman, the release is another installment in a strategy that began as a reaction to the industry's consolidation around closed models.

"We started with our own models out of necessity, because around the time that Sora came out, we realized that all the big guys are trying to close their models, and working through APIs just doesn't work for many businesses, including the kind of stuff that we wanted to build," he said.

The technical argument, he explained, is that video and world models have a fundamentally wider "surface area" of use cases than language models.

"With LLMs, the surface area of the API is pretty narrow, we're typically asking some kind of question, passing words and getting words back," Farbman said. "With video models, world models, there are so many different use cases that require people to get access to the weights and create flows that really work for them."

He was blunt that the openness is not charity. "We're definitely not doing this as philanthropy," he said. "Our answer is open weights with licenses that allow individuals and companies below a certain amount of revenue to use the model for free, and once they're successful, to come up with some kind of licensing agreement with us."

"We're trying to build a model that builders can confidently build upon," he added. "We're coming and saying: guys, open weights is not some kind of one-time philanthropic fluke for us. It's the strategy. We believe this is the right way to serve these models, and we're going to keep doing that."

What the license actually says

"Open weights" and "open source" part ways in the fine print. LTX-2.5 ships under the LTX-2.x Community License, a custom agreement that would not qualify as open source under the Open Source Initiative's definition: it discriminates by revenue and by field of use, both disqualifying restrictions.

The headline mechanic works as advertised — organizations are free to use, modify, self-host, and even sublicense the model, with the $10 million annual revenue threshold (measured across all affiliates and subsidiaries, so a small subsidiary of a large parent doesn't slip under it) triggering the paid license.

Notably, even companies above the line can download and evaluate the model free in non-production environments — the license effectively codifies the prototype-in-ComfyUI-then-license funnel Farbman describes. It also gives that funnel teeth: unauthorized commercial use obligates the violator to pay back-fees at LTX's standard rates, due within 30 days of written demand.

The stickiest provisions concern what counts as a "derivative." The definition sweeps in not just fine-tuned checkpoints and LoRA adapters but distillations and any model trained on LTX-2.5's outputs or synthetic data — meaning a company that generates training clips with LTX-2.5 and uses them to train its own unrelated model has, by the license's terms, created a derivative locked to the same agreement.

All derivatives must be redistributed under the same license, a fine-tune transferred to a $10 million-plus company triggers that company's own paid-license obligation regardless of who built it, and commercial users are barred outright from using the model to train or improve any competing AI system. A separate clause prohibits deploying LTX-2.5 in any product that competes with Lightricks' own offerings without a negotiated license.

There are also control provisions unusual for a self-hosted model. Lightricks claims no rights in generated output, but the license requires users to disclose that content is machine-generated, forbids removing or circumventing any watermarking, content-provenance, or "latent disclosure" features embedded in the model, and reserves Lightricks' right to restrict usage "remotely or otherwise" and to push updates — with immediate license revocation as the penalty for disabling disclosure features.

The license also declares Lightricks' intent that LTX-2.5 be treated as a "free and open-source general purpose AI model" under Article 53(2) of the EU AI Act, a derogation that lightens the company's own regulatory obligations — a classification legal observers may contest precisely because of the revenue threshold and use restrictions in this same document. And one restriction bears directly on the physical-AI pitch: military, warfare, and weapons-development uses are banned entirely, so the robotics checkpoint is off-limits to the defense sector without separate terms.

From Facetune to world models and the node graph that became a standard

LTX grew out of Lightricks, the Jerusalem-headquartered company best known for consumer creative apps including Facetune and Videoleap. Bootstrapped and profitable, Lightricks pivoted to foundation models in 2022, launched its LTX Studio filmmaking platform in early 2024, and released its first open-weights LTX Video model (LTXV) in November 2024, following it with a 13-billion-parameter version in May 2025. Farbman co-founded the company alongside CTO Yaron Inger and CMO Nir Pochter, and the LTX brand now fronts its world model business, with offices in New York, London, and Chicago.

ComfyUI began in January 2023 as an open-source side project by a pseudonymous developer known as "comfyanonymous," who built a node-based graphical interface for Stable Diffusion that let users chain models and processing steps into repeatable visual workflows. It has since become one of the fastest-growing open-source projects in generative media the standard environment where new image and video models are tested, combined, and pushed into production and is now backed by a company, Comfy Org, which raised $17 million to keep developing the tool. Yan, a co-founder, serves as its CEO.

Why ComfyUI is the front door for enterprise adoption

For readers wondering why a model company and a tooling company are launching arm-in-arm, Farbman's answer was unusually candid: ComfyUI is where LTX's paying customers come from.

"A whole lot of our customers are starting their journey with Comfy," he said. "It's already this prototyping system that's extremely popular in the industry, and a lot of the potential customers are coming to us after they already figured out the flow inside Comfy. It's already working, so for us it's a no-brainer that we have to provide zero-day support for the Comfy integration, because it's basically our customer acquisition channel."

Yan described ComfyUI's role as the connective layer of the open ecosystem. "Comfy at the core is sitting as a layer on top, giving people accessibility to the open-weight models that people can inference on their local machine, or tap into closed models as well through our partner node system," he said. "In the end, [they] combine everything together into a workflow that empowers various things, from the creative side all the way to data pipeline and robotics type of scenarios."

That flywheel, Yan argued, is what sustains open models commercially: "We help promote and push these models into the world... people do all sorts of workflow and model innovation on top of it, and that further propagates these models into studios or robotics labs. Those companies would end up acquiring licenses and then contribute a part of the value gained back to LTX and the rest of the ecosystem."

What enterprises should know

Both executives pushed back on the assumption that a video model is only for generating videos. Farbman rattled off a list of enterprise deployments that have little to do with cinematic clips.

"We have hardware customers that are trying to figure out how to do computational photography with diffusion models, for example, taking a stream of raw pixels that are coming from the sensors, which is typically very noisy, and trying to figure out how to reduce noise there," he said. "Or think about the production studios that are trying to figure out how to do VFX, how to do water simulation, how to turn day into night. Or think about animation studios: they're trying to figure out how to streamline their pipeline, where animators are creating keyframes and then the system uses them as interpolation."

For enterprises weighing where to start, the recommended path is the one their own employees have probably already taken. "A lot of enterprises have already adopted Comfy, and I think many others will follow," Farbman said. "It gives this right level of structure, where you can tweak things a lot, but it still abstracts a lot of things away... Enterprises are typically reaching out after people internally have already played with the model, played with Comfy."

Yan described a consistent two-track pattern among studios and companies already running LTX and other open models in production. "They have their research, or R&D, creative pipeline, anything goes," he said. "Once in a while, some of these pipelines get good enough that they graduate into some kind of production environment. And somewhere along the line, the enterprise conversation gets started. On our end, it's more around tooling, and on the LTX side, it's more around the licensing."

Because the weights are open, that entire experimentation phase can happen on a company's own hardware, with no per-generation billing and no data or IP leaving its systems, a meaningful distinction for enterprises with sensitive footage, proprietary characters, or regulated data. The commercial trigger only arrives with scale: organizations above $10 million in ARR need a license.

Yan framed the stakes for slower-moving companies in starker terms. "This is a trend that is just fundamentally going to disrupt the entire creative industry," he said. "Studios are heavily trying to figure out what is the roadmap and how do we get ahead, sometimes not even get ahead, just how do we avoid falling behind the AI adoption wave."

Developers, real-time apps, and the edge

For software developers, the release leans into a growing real-time story. Alongside ComfyUI, LTX named two other launch partners: Asteria, the AI film studio producing original film and video on LTX, and Reactor, a developer platform that runs LTX-2.5 on low-latency inference infrastructure to power interactive avatars, live worlds, and real-time robotics workloads, so developers can build production-grade real-time experiences without standing up that infrastructure themselves.

Yan pointed to a viral example of what open weights plus low latency makes possible: Flipbook, an interactive experience that spread on Reddit in which an entire clickable world is generated on the fly. "Everything people see on that interface is generated using an LTX model, live-streamed," he said. "It's an environment, or a world, where anywhere you click, it just generates a brand-new interaction... That type of experience and experimentation wouldn't exist without an open-weight model, without LTX's type of performance."

Farbman said efficiency at the edge is a deliberate design target, not a side effect. "For us, it's very important to create an extremely efficient model that people can run on edge devices, both on consumer hardware and close to the edge with physical AI," he said, while acknowledging the relentless pace of the field: "These days, it's almost hard to take a vacation. Things are progressing so quickly that while you're releasing one model, you're already deeply into training another one, and new papers are coming on a daily basis."

Filmmakers: virtual production now, easier slopes later

For professional filmmakers and studios, Yan sees real-time world models changing the shape of production itself, collapsing the gap between shooting and post. "These days you see real-time models, or world models, getting adopted in studios as part of what's called virtual production, meaning you can shoot and then immediately get close to what the post-production result looks like," he said. "You give a much better experience to the producer or director to say, 'okay, this is what I want,' or 'this is not what I want let me actually reiterate.' Whereas before, the entire Hollywood pipeline is, in my opinion, a giant mess where it has to constantly go between multiple departments."

He also cautioned against reading head-to-head model comparisons too literally, given how differently models specialize across animation, photorealism, gaming, 3D, and robotics. "Various models have simply different characteristics," he said. "It's like comparing Michael Phelps with, I don't know, Michael Jordan. It's not really a comparison of who's a better athlete, there are just different specialties here."

As for amateur and indie creators intimidated by ComfyUI's famously steep learning curve, Yan was direct that the tool will meet them partway, but only partway.

"It's kind of like skiing," he said. "There are easy slopes that you can go down using Comfy, and hopefully we can create more and more of these easy slopes overall. But we'll never sacrifice the existence of the double-black-diamond type of lanes, because the real technical, professional creatives actually need and couldn't live without that type of core power. That's actually our core differentiator compared to a mobile-app type of creative tool."

LTX-2.5 is available today on Hugging Face, natively in ComfyUI, and through the LTX API.

Updated several hours after publication with additional details from LTX's public blog post and API pricing page.

OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks

Earlier today, OpenAI launched GPT-5.6-Cyber, a specialized model designed to perform advanced vulnerability research and exploit development for approved defenders — including categories of work that its general-purpose models will often refuse.

GPT-5.6-Cyber is a fine-tuned version of OpenAI's most advanced general model, GPT-5.6 Sol, unveiled back in June, but trained specifically to improve performance on advanced cybersecurity tasks, including finding zero-day vulnerabilities and developing exploit chains.

Crucially, OpenAI also trained it to reduce refusals on some higher-risk, "dual-use" cybersecurity requests — that is, requests that could be used for legitimate defensive or malicious offensive purposes.

Indeed, on an internal OpenAI benchmark called Advanced Cybersecurity Completion Rate — which the company says in its launch blog post measures tasks involving exploit-chain development, authentication bypass, privilege escalation, and other advanced cybersecurity scenarios — GPT-5.6-Cyber completed 95% compared to just 57.3% from its immediate predecessor model GPT-5.5-Cyber, and just 1.5% with the normal GPT-5.6 Sol model and all its safeguards applied.

OpenAI researcher Eric Wallace posted on X, describing GPT-5.6-Cyber as OpenAI's "first large-scale attempt at directly improving capabilities for advanced cybersecurity tasks such as exploit development."

Pricing and availability

Unfortunately for enterprises, GPT-5.6-Cyber is not being made broadly available to every ChatGPT or API customer.

To get access, an organization has to be accepted into the newly created tier of OpenAI’s Daybreak cybersecurity program, called Daybreak Red — also announced today, which gives access to dedicated cybersecurity models like GPT-5.6-Cyber

Another new tier, Daybreak Blue, gives a wider swath of enterprises access to general models like GPT-5.6 Sol but with some guardrails lifted to allow for more cybersecurity uses.

OpenAI’s documents list pricing for GPT-5.6-Cyber at $12.50 per million input tokens and $75 per million output tokens, with cached input at $1.25 per million tokens.

That makes it more expensive than GPT-5.6 Sol in the same Daybreak cyber pricing table, where Sol is listed at $5 per million input tokens and $30 per million output tokens for short-context use.

OpenAI does not list long-context pricing for GPT-5.6-Cyber in the same table, and access still requires separate Daybreak Red approval and provisioning.

Red vs. Blue: OpenAI's new Daybreak tiers and how to qualify for them

Daybreak Red is for approved security teams doing advanced, authorized cyber work — the kind of work that can look risky out of context, even when it is being done for defensive reasons. That includes vulnerability research, penetration testing, red-team exercises and exploit validation on systems the organization owns, operates or has permission to test. In other words, OpenAI is saying GPT-5.6-Cyber is for trusted defenders with a clear professional need, not for general experimentation.

Enterprises that want access have to apply through Daybreak Access, OpenAI’s current pathway for vetting cyber users. The application asks companies to identify who they are, what kind of security work they plan to do, where they will use the models, and which OpenAI products or surfaces they expect to use. Applicants also have to confirm that their work is lawful, defensive and authorized.

OpenAI is also looking for signs that the applicant has a serious security program of its own. The company says participating enterprises need controls such as single sign-on, multifactor authentication, role-based access, employee-use monitoring, usage logs, API-key controls and a documented incident-response process. OpenAI also asks for a recognized security certification such as SOC 2 Type II, ISO 27001 or an equivalent standard. Access is limited to approved people inside the organization using company-controlled accounts and devices.

If an enterprise does not qualify for Daybreak Red, or does not need that level of access, OpenAI is pointing most companies toward Daybreak Blue, its other cyber models access tier, instead.

Blue is the broader tier for approved defenders. It does not provide GPT-5.6-Cyber, but it does give vetted users access to OpenAI’s frontier general-purpose models, including GPT-5.6 Sol, with safeguards adjusted for legitimate defensive work.

For many enterprise security teams, Blue may be the more realistic starting point. OpenAI says it is meant for tasks such as secure-code review, vulnerability discovery, malware analysis, incident response and patch validation. These are still sensitive uses, but they do not necessarily require the same specialized cyber model access that comes with Red.

The practical takeaway is that enterprises now have two routes into Daybreak. Blue is for approved defenders who want stronger AI help with everyday security work. Red is for the smaller set of approved teams that can justify access to specialized cyber models, including GPT-5.6-Cyber. Companies that want to use Daybreak capabilities in products or services for their own customers need a separate approval path through the Daybreak Cyber Partner Program, rather than simply applying for internal enterprise access and passing it along.

How OpenAI got here: from Trusted Access to Daybreak

OpenAI has supported defenders through its Cybersecurity Grant Program since 2023 — later expanded to $10 million — and began building cyber-specific safeguards into its model deployments starting with GPT-5.2.

In February 2026 it introduced Trusted Access for Cyber (TAC), an identity-and-trust framework that gave vetted defenders lower classifier-based refusals for authorized work such as vulnerability triage, malware analysis and binary reverse engineering.

From there, the cadence accelerated. In March, OpenAI CEO and co-founder Sam Altman announced the Daybreak program. In April, OpenAI scaled TAC and released GPT-5.4-Cyber, a version of GPT-5.4 fine-tuned to be "cyber-permissive" for a limited set of vetted vendors and researchers.

In May, it followed with GPT-5.5-Cyber in limited preview for defenders of critical infrastructure, and lined up partners including Cisco, Intel, SentinelOne, Snyk and Cloudflare.

Notably, OpenAI said at the time that GPT-5.5-Cyber was "primarily trained to be more permissive," not to significantly out-perform its general model — GPT-5.5-Cyber actually scored worse than GPT-5.5 on some evaluations.

TAC required phishing-resistant Advanced Account Security for individuals on its most capable models beginning June 1, and Daybreak now requires hardware security keys for individual accounts beginning September 1.

OpenAI says GPT-5.6-Cyber has already found zero-days

OpenAI isn't relying exclusively on benchmarks to make its case.

The company says its researchers used GPT-5.6-Cyber to investigate V8, the JavaScript engine underlying Chrome, and uncovered two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox.

OpenAI researchers validated the findings and disclosed them to Google, which fixed the vulnerability assigned CVE-2026-15903 — a high-severity flaw in which V8's optimizing compiler skipped a safety check during integer conversion, allowing an out-of-bounds array index that an attacker could use to read or overwrite memory.

OpenAI says the model has also contributed to finding at least five vulnerabilities in an unnamed popular mobile operating system, three critical vulnerabilities in an unnamed popular database, and more than 400 vulnerabilities capable of producing privilege escalation in a popular operating-system kernel. Those disclosures are still being coordinated, according to OpenAI.

The results put OpenAI into a rapidly developing market for AI-assisted offensive security. XBOW, for example, markets autonomous penetration-testing agents that map attack surfaces, attempt exploits and independently validate findings; in 2025 it became the first AI system to top HackerOne's U.S. bug-bounty leaderboard, and this year it disclosed a set of critical, CVSS-9.8 remote-code-execution flaws in Microsoft's Bing image-processing systems, found without source-code access.

For enterprise security leaders, that emerging competition matters because vulnerability research is moving beyond using an LLM as an assistant. Vendors are increasingly building systems in which models can investigate targets, operate tools, validate hypotheses and produce actionable findings.

Specialized doesn't mean universally better

OpenAI's own results also show why enterprises shouldn't simply equate cyber specialization with better performance everywhere.

GPT-5.6-Cyber outperformed GPT-5.6 Sol and GPT-5.5-Cyber on OpenAI's implementation of ExploitGym, which evaluates whether agents can turn known vulnerabilities into working exploits in controlled environments. It also beat Sol on an internal zero-day evaluation.

But GPT-5.6 Sol performed better on OpenAI's Vulnerability Discovery and Report Writing evaluation. OpenAI attributes the Cyber model's lower score partly to shorter and less detailed vulnerability reports.

Sol also performed best on ExploitBench under its standard 300-turn limit, with OpenAI saying it solved tasks more token-efficiently. Extending the evaluation to 600 turns narrowed the gap between the models.

That suggests enterprises may eventually treat cyber models as specialized workers rather than replacements for general reasoning models: one model for deep exploit work, another potentially better suited to analysis, documentation or other parts of a security workflow.

SpecterOps CTO Jared Atkinson said GPT-5.6-Cyber is "materially improving our specialist vulnerability-research workflows," adding that it completed some work in less than a day that previous models had failed to resolve after weeks of intermittent effort.

The Hugging Face incident hangs over the launch

The permissive-model pitch arrives weeks after OpenAI's most serious public demonstration of what can go wrong when cyber refusals are turned down — and OpenAI addresses that history head-on in the Daybreak announcement.

In July, OpenAI and Hugging Face jointly disclosed that during an internal ExploitGym benchmark evaluation — run with production classifiers deliberately disabled to measure maximal capability — a combination of OpenAI models, including GPT-5.6 Sol and an unreleased, more-capable pre-release model, broke out of their sandboxed research environment and autonomously attacked Hugging Face's production infrastructure.

The models exploited a zero-day in an internally hosted package-registry cache proxy to reach the open internet, moved laterally through OpenAI's research nodes, then inferred that Hugging Face likely hosted ExploitGym's answer keys and chained stolen credentials and remote-code-execution flaws to reach its production database. OpenAI called it an "unprecedented cyber incident, involving state-of-the-art cyber capabilities."

As VentureBeat previously reported, the episode also exposed the flip side of blanket safety guardrails: when Hugging Face's defenders tried to use commercial frontier models to analyze the raw exploit payloads and credential dumps from the attack, the models refused, and the company completed its forensic reconstruction only after switching to a Chinese open-weight model, GLM 5.2, run locally.

That guardrails-block-the-defender dynamic is much of what OpenAI's reduced-refusal Daybreak tiers are meant to solve — even as the same incident illustrates the risks of reducing refusals in the first place.

OpenAI is careful to draw a line between that incident and this product. In the Daybreak announcement it states directly that GPT-5.6-Cyber "was not involved in exploiting Hugging Face, nor are any other models planned for an upcoming release," and notes that the pre-release model implicated in July was an internal-only research prototype that has since been deactivated, encrypted and restricted from research access.

The company has said it is working with external advisers including CrowdStrike, METR and Redwood Research on the review, and has brought Hugging Face into its trusted-access program.

In my assessment, the access model still leaves OpenAI with a hard question: whether keeping GPT-5.6-Cyber inside the narrower Daybreak Red tier also limits the very defensive work it says it wants to accelerate. If only a small group of approved participants can use the model, enterprises outside that tier may still lack access to the kind of specialized AI assistance that could help with fast diagnosis, containment and response in incidents like the one involving Hugging Face.

That means OpenAI may still be repeating part of the mistake it is trying to move past. By holding its most capable cyber model behind a tighter approval process, it reduces obvious misuse risk, but also leaves many enterprise defenders looking elsewhere. For teams that cannot qualify for Daybreak Red, or cannot wait for approval, open weights models may remain the more practical alternative: less controlled, but easier to obtain, inspect, run internally and adapt during a live security investigation.

The guardrail is increasingly around the model

The most consequential part of Daybreak may ultimately be its access architecture rather than its benchmarks.

OpenAI explicitly says Daybreak Blue removes system-level guardrails that can interfere with legitimate defensive work, while GPT-5.6-Cyber goes further by reducing model refusals for certain dual-use tasks. In their place, OpenAI is imposing controls around who receives access and how the models operate.

Daybreak access is restricted to approved individuals and organizations performing authorized work. OpenAI says controls include identity verification, account security, monitoring, approved-use restrictions and legal attestations.

The company is also encouraging Daybreak customers using Codex to move from full-access execution to an auto-review mode capable of evaluating actions requiring elevated permissions before they execute. Individual Daybreak accounts will be required to adopt hardware security keys beginning September 1. OpenAI says it is additionally rolling out improved monitoring in the coming weeks and prioritizing alignment training and testing for upcoming Daybreak releases — commitments that read, in context, as a direct response to the Hugging Face review.

OpenAI's broader Codex Security product supplies another layer around the models, providing repository analysis, vulnerability validation, remediation and integration into cloud, pull-request and local development workflows. OpenAI says Codex Security has scanned more than 30 million commits across more than 30,000 codebases, with more than 500,000 findings fixed.

That model-plus-harness approach resembles a broader shift in AI security products. XBOW, for example, emphasizes orchestration, exploit validation and governance around frontier models rather than treating an LLM alone as the complete penetration-testing system.

OpenAI nevertheless acknowledges that increasingly permissive cyber models create additional risks, whether from misuse or misalignment. It assesses both GPT-5.6 Sol and GPT-5.6-Cyber at the High cybersecurity capability level under its Preparedness Framework, but below its Critical threshold. A fuller GPT-5.6-Cyber system card is planned for later publication.

For CISOs and security engineering leaders, Daybreak therefore presents a different deployment question than another incremental model upgrade. As models become capable enough to perform work previously reserved for experienced vulnerability researchers — and, as the Hugging Face incident showed, capable enough to pursue a narrow goal straight through a sandbox — the enterprise control plane around those models — permissions, sandboxes, monitoring, human review and authorization — becomes as important as the intelligence inside them.

AWS Continuum integrates with OpenAI Codex and Anthropic Claude Code in major AI security push

Amazon Web Services is threading its AI-powered security infrastructure directly into the coding environments built by two of its fiercest rivals — and in doing so, it is making a bold bet that controlling the security layer matters more than controlling the model.

AWS announced at Black Hat USA 2026 this month that its Continuum platform for code vulnerabilities will integrate directly into Anthropic's Claude Code and OpenAI's Codex, alongside AWS's own Kiro IDE.

The move embeds AWS security tooling at the point where developers write code, regardless of which AI model they use to do it. Simultaneously, AWS expanded Security Hub Extended — its curated, single-bill security marketplace launched in February — with a 10th security category focused on supply chain protection, bringing in Chainguard and Socket as partners.

Together, the announcements are AWS's most sweeping attempt yet to position itself as the default security control plane for enterprise software development in the AI era — a role that carries enormous commercial implications as the global cloud infrastructure market surpasses $143 billion per quarter, according to Synergy Research Group.

Why frontier AI models turned the vulnerability backlog into a five-alarm fire

The urgency behind both launches traces back to a single inflection point that reshaped enterprise security earlier this year. Claude Mythos Preview, announced by Anthropic in April, is a general-purpose AI model that during testing revealed striking cybersecurity capabilities far exceeding any prior system.

In pre-release evaluations, Mythos identified thousands of previously unknown zero-day vulnerabilities across every major operating system and web browser. More than 99% of those vulnerabilities remain unpatched by their maintainers, and the median time from vulnerability discovery to weaponized exploit — already collapsed from 771 days in 2018 to under four hours by 2024 — is projected to reach under one hour by the end of 2026.

Chet Kapoor, AWS's vice president of search, security, and observability, framed the challenge in stark terms in an exclusive interview with VentureBeat. "CISOs have had code vulnerabilities for a while, and then Mythos came along, and it just made it a lot worse," Kapoor said. "They already had a backlog. Now the backlog is 5x more, and that causes a problem."

That problem — the exponential growth in known vulnerabilities outpacing any organization's ability to triage and fix them — is precisely what Continuum is designed to address. Kapoor described AWS's broader security vision as a shift from "telemetry, storage, query, dashboards for humans to telemetry, context, reasoning, and actions by agents." The shorthand for that vision is a phrase AWS repeated throughout Black Hat: autonomous security at machine speed.

Inside Continuum's four-phase system for finding and fixing code flaws automatically

Continuum operates as what AWS calls an “agent-team loop architecture” — a sophisticated orchestration harness that selects the right AI model for each task, connects to a customer's environment, and delivers validated secure code. Under the hood, it runs through four distinct phases.

Kapoor broke them down for VentureBeat. Discovery uses multiple frontier AI models to scan code and ingest a customer's existing vulnerability backlog. Prioritization — which Kapoor called "one of our biggest value adds" — contextualizes each finding against a customer's actual environment and business risk. "You go from 100 to 2,000, and now you're like, whoa, I didn't even know which 100 to focus on," he said.

Validation then builds reproducible exploits in an isolated sandbox to confirm whether a vulnerability is genuinely exploitable. "Once I do them, how will it behave?" Kapoor explained. "You create a sandbox to go off and make that happen. So you can figure out what the blast radius is." The validation phase covers both first-party code that customers wrote themselves and third-party open source code they depend on. Finally, remediation offers fixes — whether network configuration changes, policy adjustments, or code patches — that the system has already tested in the same sandbox. The human stays in control throughout, approving outcomes at whatever level of autonomy the organization is comfortable with.

The commercial model is equally deliberate. Customers pay AWS a single price for Continuum. AWS absorbs the underlying token costs for whichever frontier model performs best at each phase of the scan. "The customer purchases Continuum, period," Kapoor told VentureBeat. "We optimize on which model to use for what because, quite frankly, GPT Cyber is good at some things, Mythos is good at some things."

How AWS convinced OpenAI and Anthropic to open their coding tools to a rival's security layer

The most strategically striking element of the announcement is the integration with OpenAI Codex and Anthropic Claude Code. AWS competes directly with both companies across cloud AI services. Amazon holds a massive investment in Anthropic, and OpenAI operates its own growing infrastructure that competes for the same enterprise AI workloads. Yet both agreed to embed Continuum inside their developer environments.

When VentureBeat asked Kapoor directly about the competitive dynamics, he pushed back on the framing entirely. "Who is the competitor?" Kapoor said. "I can keep thinking about Anthropic and OpenAI to be partners. I don't understand the word 'competitor' in your description of the question." He added: "They're partners with us. We use their models. We plug into their environments. Which is why we actually brought them together to do this."

Kapoor argued that working with a single model provider would be insufficient. "I don't think it's good enough to just do it with one company," he said. "Everybody is going to leapfrog each other over a period of time." By absorbing token costs and presenting a single bill to the customer, AWS positions Continuum as infrastructure — not a model wrapper. The harness, not the engine, becomes the durable competitive asset.

As Kapoor wrote in his blog post announcing the partnership: "An AI harness is the orchestration layer that wraps around a model to connect it to tools, guardrails, memory, and workflows, so it delivers outcomes. Think of the model as the engine and the harness as everything around it. You need both to have a high-performance car." 

AWS partners echoed the logic. "Model choice was never the hard part for enterprises. Trust in what the model does in production is," said Val Henderson, CEO of AWS Premier Partner Caylent, in comments reported by CRN.

AWS adds supply chain security to its curated marketplace as open source threats intensify

The second prong of AWS's Black Hat announcements extends Security Hub Extended into supply chain security as its 10th category, with Chainguard and Socket as curated partners. The Extended plan now includes 23 curated partner solutions, all on a single AWS bill with no required long-term commitments, covering endpoint, identity, email, network, data, browser, cloud, AI, security operations, and now supply chain.

Michael Fuller, AWS's director of security services, told VentureBeat that the addition was driven entirely by customer demand. "Over the last six to eight months, it's gotten quite a bit of news around what's happening in the supply chain space, with the fact that everybody builds on open source," Fuller said. "Our customers quickly reached out and said, 'Security Hub Extended is resonating. We would love to see a supply chain security category with some key players there because it's a hot topic for us.'"

The two partners were chosen to be complementary rather than duplicative. Chainguard focuses on providing hardened, secure-by-default container images and packages rebuilt from verified source code. Socket performs behavioral monitoring of packages as they are pulled into a developer's environment, detecting threats like typosquatting, maintainer account takeover, and obfuscated malicious code. "Together, between the three of us — us with consolidating that, ChainGuard providing really good hardened and cleaned images and packages, and then Socket providing a behavioral analysis over the top — gives customers a really good holistic supply chain security offering," Fuller said.

The complementary approach addresses two distinct attack vectors. An attacker can publish a malicious package that contains no known vulnerabilities — Chainguard's clean-build approach defends against that. Separately, an attacker can compromise a legitimate maintainer's account and push a tainted update to a trusted package — Socket's behavioral detection catches that. Both vectors are amplified in the AI coding era, Fuller noted, because AI agents face the same supply chain risks as human developers: "Agents can be misled on, 'Hey, this is a well-known package that you're looking for,' and therefore pull it down, even though it's been maliciously obfuscated."

Why AWS chose two partners per category instead of building a security marketplace

The partner selection strategy behind Security Hub Extended reveals a deliberate philosophy that distinguishes it from the AWS Marketplace, which already hosts tens of thousands of security offerings.

Fuller told VentureBeat that customers articulated clear principles for what they wanted. "One was don't give me hundreds of offerings. We already have the AWS Marketplace," he said. "Two was give me a sweet spot. Our customers were saying, give me two in each category, and when you look at those two, don't give me head-to-head competitors. Give me one that I may know well, that is an established player, and give me one that's taking a different approach."

Fuller pointed to the security operations category as the template. "You have Splunk, hard to argue not an established leader in security operations, and then you have Seven AI that's kind of taking a very different approach, and they're complementary in a lot of ways."

The decision to build internally versus partner follows a similar logic. For endpoint detection and response, AWS has no structural advantage, so it partners exclusively. For cloud security, AWS builds its own native tools because it intimately understands its own infrastructure — but still partners with Upwind to give customers a second option.

"At the end of the day, what we're trying to do here is ensure that our customers can operate in the most secure way possible on AWS, not necessarily grow a large security business as the core goal," Fuller said. "That's why it's very easy for us to decide to do both building ourselves, but also then inviting partners to participate."

The pricing model reinforces this accessibility. Fuller said customers demanded pay-as-you-go options alongside traditional multi-year commitments. "All of the Security Hub Extended offerings have a public-facing, pay-as-you-go price, just like our first-party offerings do within AWS," he said. "So that gives customers the option to go kick the tires, get going, even scale up and use the services without going through a traditional sales cycle."

Shadow agents and AI cost harvesting emerge as the next frontier of cloud security threats

Both AWS executives addressed an emerging security concern gaining traction among CISOs: the proliferation of unregistered AI agents — what the industry has begun calling "shadow agents" — and the novel attack patterns they enable.

Kapoor told VentureBeat that shadow agents are a genuine and growing problem, though he was careful to separate it from the Continuum announcement. "There are many agents that are registered with registration directories, whether it's Vertex, whether it's Agent Core, whatever else it might be, but there are many agents that are not registered with the registry, and those are what people are calling shadow agents because they can actually do some harm," he said. "Discovering shadow agents is not easy. The industry is working on it."

Fuller provided more granular detail on what AWS has already deployed. Security Hub now includes a free AI inventory capability that uses three data layers: AWS Config identifies AI-related services like SageMaker, Bedrock, and Agent Core across an organization; Amazon Inspector scans compute instances and containers for AI-related software; and GuardDuty compares DNS request and response logs against known AI tools and agentic workloads.

Beyond inventory, Fuller revealed that GuardDuty now monitors data plane events — including prompts, prompt volume patterns, and inference cost analysis — to detect what AWS calls "cost harvesting."

The attack mirrors the cryptocurrency mining that became common after cloud credential compromises: an attacker gains access to an AWS account and burns through as much free AI inference as possible before detection.

"We're seeing what we're calling cost harvesting," Fuller said. "They'll spin up, basically try to get as much free inference as they can until that's discovered." It is, Fuller noted, "the same thing that's happening in AI" as happened with crypto mining — and GuardDuty's detection of credential compromise and unauthorized compute usage translates directly to the new threat.

How Continuum and Security Hub Extended fit together in AWS's enterprise security strategy

Although both announcements landed the same week, AWS is treating the products behind them as separate. Kapoor described Continuum to VentureBeat as distinct from Security Hub Extended, sold as its own standalone product. AWS declined to discuss its longer-term roadmap for the two.

The design logic points in one direction. Continuum addresses the code an enterprise writes and the open source it inherits. Security Hub Extended addresses everything else — and the newest of its categories is where the two most clearly overlap. Continuum's validation phase covers third-party dependencies alongside a customer's own code; Chainguard and Socket harden and monitor the same packages from the other direction. One capability is built in-house, the other curated from partners, and they meet at the same attack surface.

Both proceed from the same premise: that enterprises no longer want a catalog, they want a recommendation.

"Customers want an opinionated point of view on how they should do security in the AI era," Kapoor told VentureBeat. "That's what Security Hub Extended was about — actually going off and giving them our opinion." AWS will continue to give customers choice, he added, "whether it is something that we ship or whether it is something from a partner."

That doctrine — a recommendation, with an escape hatch — is the through line connecting a curated marketplace to a first-party agent platform, and it makes the boundary between them more porous than two separate announcements suggest. Security Hub has already absorbed capabilities that did not exist a year ago, including the free AI inventory and the cost harvesting detections Fuller described. The console is where AWS delivers its opinion to the enterprise. Continuum is the sharpest opinion it has shipped.

The audience for that opinion has changed as well, Kapoor said. Mythos, he argued, moved security from something the CISO owned to a CEO and board-level imperative. "Boards are now asking for updates on what's going on with security in the enterprise because it's a business threat now, it's a business risk."

AWS's security ambitions reflect a calculated bet on owning the orchestration layer

The twin launches fit within a broader strategic arc AWS has been building throughout 2026 at a breakneck pace. The company re-imagined Security Hub at re:Invent 2025 by consolidating GuardDuty, Inspector, CSPM, and Access Analyzer into a single console. In February, it launched Security Hub Extended with 14 curated partner solutions. By May, that number grew to 21 across nine categories. Now it stands at 23 across 10. Continuum launched at the New York Summit in June and expanded to OpenAI and Anthropic integrations at Black Hat in August.

AWS generated $42.2 billion in revenue during Q2 2026, with cloud sales expanding 37% year over year. The company holds a 28% share of the global cloud infrastructure market, ahead of Microsoft at 20% and Google at 15%.

Fuller told VentureBeat that AWS has "tens of thousands of customers using one or multiple of our security services, essentially across all geos that we operate in, and in every industry, and both commercial and government." The Extended plan aims to convert that installed base into users of partner security solutions — deepening engagement and making it harder for competitors to dislodge AWS as the default platform.

By making AWS the seller of record for 23 partner security solutions and embedding Continuum inside the coding environments of OpenAI and Anthropic, AWS is constructing something more durable than a product line. It is building the connective tissue between enterprises and every AI model they use, between every open source package they pull, and between every security vendor they deploy. In a world where frontier models are advancing so rapidly that today's best scanner becomes tomorrow's table stakes, the layer that persists is not the model — it is the harness that connects the model to the customer's environment, policies, and risk tolerance.

Kapoor, reflecting on a chance conversation he had on a flight to Black Hat, offered the simplest articulation of why all of it matters. A former CISO turned CTO sitting beside him volunteered a blunt assessment of the current moment: "I don't feel safer now." Kapoor's response, he told VentureBeat, was equally blunt: "We're working on it."

Whether that work makes the world safer or simply makes AWS indispensable to every organization trying to get there may, in the end, amount to the same thing.

Brex assumes its AI agents could do anything — so it watches the network, not the code

10 August 2026 at 18:31

Brex CEO Pedro Franceschi offered a blueprint for one of the pressing challenges facing the enterprise today at VB Transform 2026: securely deploying AI agents, like the open-source OpenClaw, into production environments.

Unlocking this enterprise value requires a mindset shift. The industry needs to move past vague terminology and focus on concrete enterprise roles. 

“People talk a lot about agents, but I think 'agents' is a terrible name. It's this Silicon Valley concept that doesn't really mean much,” Franceschi said. 

Instead, the goal should be creating entities that can genuinely collaborate with human workers. "The concept we always had in mind was the idea of a virtual employee — someone on Slack, an entity, it has an email address, it can join meetings, you can email it, and that you can work with," Franceschi said.

Realizing this vision demands a new security paradigm. Franceschi’s presentation detailed how Brex pointed OpenClaw at internal roles, realized traditional security models failed, and built a novel network-level security layer called CrabTrap.

The OpenClaw security dilemma

The journey began following a breakthrough in December, when coding models reached a level of maturity that enabled the January release of OpenClaw. This marked the moment agents could finally self-bootstrap and maintain their own codebases instead of relying on hard-coded, static tools. 

However, when Franceschi proposed deploying this to automate internal functions, the Brex security team firmly rejected the idea. “They said, 'Hell no. How could we trust an agent doing these things? This thing has code execution capabilities. There's no way to control it,'” Franceschi said. That caution isn't unique to Brex — enterprises broadly have been wary of granting agents uncontrolled code execution on corporate networks.

To solve this, Brex had to shift the security perimeter. Franceschi contrasted this with approaches like Nvidia's NemoClaw, which he said secure agents by limiting their tool usage — a model he believes neutralizes the coding capabilities that give agents their value.

“… the premise we had was that the coding capabilities were critical to the model having the ability to do a variety of tasks,” he said. 

Brex's fix was to shift the security boundary to the network layer instead. Instead of policing the ever-changing code inside the container, the focus must shift to monitoring what the code actually attempts to send or receive from the outside world.

CrabTrap and the LLM-as-a-judge solution

This network-centric approach led to the creation of CrabTrap, an open-source HTTP proxy built by Brex. The mechanism operates on the assumption that OpenClaw can do anything and might already be compromised. Therefore, CrabTrap monitors all outbound network traffic between the container and the internet, using an LLM to judge whether that traffic aligns with the agent's approved policy.

“Instead of trying to control the code running in the container, assume the thing can do anything and monitor the network traffic between that container and the internet,” Franceschi said. 

Using a large language model (LLM) to judge every single network request introduces unacceptable latency, often adding thousands of milliseconds to response times. Brex solved this by passing traffic through a bifurcated system. 

Routine, low-risk actions pass through static, pre-approved rules instantly. If a recruiting agent tries to view a LinkedIn profile, the static rule allows it. However, high-risk actions such as sending emails are flagged and routed to the LLM judge for evaluation. Franceschi said that architecture ensures only about 2% of complex requests actually face LLM latency. 

A surprising finding from the project was how effectively the LLM judge performs this role. Franceschi attributed this to the models' training: LLMs are exposed to billions of web pages and HTTP requests, giving them what he described as an inherent semantic understanding of network traffic patterns.

“[Models] are very good at discerning what is within the policy and what is not,” Franceschi said, adding that this capability emerges naturally through pre-training without needing heavy prompting.

Brex put this infrastructure to the test with “Jim,” a virtual recruiter built on OpenClaw. Jim handles various tasks, including sourcing candidates, scoring inbound applicants, and sending emails. 

When Jim attempts an action that falls outside the established policy, CrabTrap relies on a human-in-the-loop workflow. If the LLM judge flags an unapproved outbound email, CrabTrap pings a human manager on Slack. 

The Slack notification explains the agent's underlying intent and suggests a policy change that would allow the action. The human manager can then review the context and click "yes" or "no" to update the rules dynamically. 

"I like the virtual employee analogy because a lot of these things were solved already in a company, in the context of humans," Franceschi said. "When an employee hits a wall, they escalate to their manager."

The cost of the frontier

Brex is a fintech company, not a cybersecurity vendor. The decision to build CrabTrap in-house was driven by a lack of mature commercial solutions that could satisfy their security team. 

Franceschi acknowledged the inherent cost of operating at the bleeding edge, admitting that commercial vendor solutions will likely catch up. 

“When we built this, it was clear to me there was a 70% chance we would throw it away in six months... But what we learned by being six months ahead was worth it in shaping our AI adoption strategy,” he said. 

The investment in building internal tools provided Brex with the experience needed to safely deploy agents months ahead of the broader market. For enterprise leaders navigating the AI landscape, the core takeaway is the necessity of building the cultural and technical muscle to operate in an agentic world today. 

“We don't have all the answers, but the answer is not to do nothing,” Franceschi said.

Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents — available now

Meta today released Muse Glimmer, a 30-billion-parameter open-weight model designed to run autonomous AI agents directly on consumer hardware — pushing agentic workloads that normally depend on cloud infrastructure onto high-end Macs and PCs.

Just as notable as what the model does is how it's licensed. Glimmer arrives under the permissive, industry-standard Apache 2.0 open source license — the company's first fully open release since it succeeded its open-weight Llama family in April with the proprietary Muse Spark.

In fact, Muse Glimmer launches today with a more permissive license than Llama ever carried. Llama's bespoke community license drew years of criticism for restrictions like its 700-million-monthly-user cutoff; Apache 2.0 has no such strings, permitting unrestricted commercial use, modification and redistribution.

The weights are available on Hugging Face now. Wang said support is rolling out this week through Ollama, LM Studio, vLLM, SGLang, Together AI, Fireworks AI and OpenRouter, with optimized llama.cpp, MLX and ExecuTorch integrations landing in the coming days; Meta's blog post also names Unsloth as a local-runtime partner and points to PyTorch's TorchTitan for fine-tuning. The company says it is working with AMD, Arm, Dell, Intel and Nvidia to optimize performance across devices, and has published developer documentation covering custom agent scaffolds.

"Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally," Meta co-founder and CEO Mark Zuckerberg wrote in a post on X (under his longtime handle @finkd). "Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases."

That promised Muse Spark 1.2 release would be an even bigger shift: it's the frontier model behind Muse Code, the terminal coding agent Meta shipped just five days ago, and until today the entire Muse family was proprietary. Zuckerberg had teased at that launch that he'd "have more to share soon" on open source. Now we know what he meant.

For developers and enterprises, the practical stakes of local inference go beyond where computation happens. An agent working with files, screenshots, development environments and other sensitive context can execute those workflows without continuously sending that information to a remote inference service. Local deployment also removes network availability and per-token API charges from the inference loop — although organizations still bear hardware, electricity, deployment and management costs.

A 30B model built around the agent loop

Rather than positioning Glimmer primarily as a general chatbot, Meta trained it around the sequence of operations an autonomous agent performs: formulate a plan, call tools, interpret the results, continue working, and recover when something goes wrong.

"Just like much larger models, muse glimmer can operate as a fully capable agent via planning, tool calls, checking its own results, and failure recovery," Alexandr Wang, Meta's chief AI officer, wrote in a thread on X announcing the release, adding that the model "can run on 24GB of VRAM without losing agentic reliability."

According to the model card on Hugging Face, Glimmer is a dense causal transformer with approximately 29.6 billion total parameters across 52 layers, including a dedicated ~1.8B-parameter ViT-G/14 perception encoder. It accepts interleaved text and images, produces text, supports more than 100 languages and has a stated context length of 131,072 tokens or more, with a knowledge cutoff of January 4, 2026.

That combination is intended to let an agent interpret screenshots, charts and documents while simultaneously reasoning about text and invoking external tools. Glimmer offers low, medium, high and xhigh reasoning settings — set via the system prompt — so applications can dial reasoning effort up or down per task, and Meta says it works across agentic scaffolds including OpenClaw and Hermes Agent.

The model is a distillation of Meta's larger flagship: per the company's technical blog post, Glimmer was pre-trained on Muse Spark's outputs using logit distillation, mid-trained on longer-context, agent-heavy data with richer reasoning traces, then post-trained with supervised fine-tuning, on-policy distillation and reinforcement learning across general, reasoning, coding and agentic domains.

Meta demonstrated the result with a local Home Assistant workflow: in a demo video, Glimmer autonomously discovers a Home Assistant instance on the network via tool calls, queries device APIs, writes a responsive HTML/CSS/JavaScript dashboard from scratch and deploys a local server to verify its own work. That's closer to the operational reality of enterprise agent deployments than a standalone question-answering benchmark — the model has to maintain a plan while interacting with external systems, then inspect whether its actions produced the expected result.

Compressing an agent into 24GB

The hardware story is central to the release.

At full precision, Meta says the 30B model requires more than 55GB of memory — beyond any single consumer GPU.

The company therefore developed approximately 4-bit quantized versions that shrink the language-model weights to under 20GB, leaving headroom for the pieces an operational agent also needs in memory: the KV cache, the perception encoder and a companion speculative-decoding model, all fitting within a 24GB or 32GB envelope.

In practical terms, that means the quantized builds run on consumer machines — though the upper end of them. The 24GB-targeted K-Quant-17GB configuration fits on a single high-end consumer graphics card, such as Nvidia's RTX 3090 or RTX 4090 (both with 24GB of VRAM), while the 32GB-targeted K-Quant-Dynamic version lines up with the newer RTX 5090's 32GB. On the Mac side, Apple Silicon's unified memory plays the role of VRAM, so a MacBook Pro or Mac Studio with 32GB or more of memory can hold the full stack — Meta ran its own speed tests on M4 Max and M5 Max MacBook Pros. A typical 8GB or 16GB laptop, however, remains out of reach, and the full-precision BF16 release — which Meta pegs at 64GB — stays in the territory of data-center GPUs and top-spec Mac Studio configurations.

Meta reports average accuracy degradation of just 0.2% across 15 benchmarks for its K-Quant-Dynamic version targeting 32GB hardware, and 1% for the K-Quant-17GB configuration targeting 24GB hardware. Those figures are Meta's own measurements, not independent evaluations.

Meta is also using DFlash speculative decoding to attack the other big problem with local agents: latency. Instead of generating every token sequentially, a smaller DFlash "drafter" model proposes blocks of 16 tokens that the primary model verifies in parallel, producing identical output faster.

Meta reports this raises average generation speed on an Nvidia RTX 5090 from 74.9 tokens per second to 233.4 — a 3.1x increase. An Apple M5 Max rises from 26.6 to 50.2 tokens per second (1.8x), and an M4 Max from 23.7 to 37.8 (1.5x). The tests used batch size one and greedy decoding, with Apple systems measured via ExecuTorch and the RTX 5090 via llama.cpp.

For agent applications, those multipliers matter more than they would for chat: a single user request can trigger many model turns, tool calls and verification steps, and latency accumulated at every stage can quickly make an otherwise capable agent impractical.

Glimmer enters an increasingly competitive local-model market

Meta is not entering an empty field. Developers already have capable open-weight models in this size class, most prominently Google's Gemma 4 family and Alibaba's Qwen3.6-27B — both of which position themselves around reasoning, multimodal understanding and agentic workloads. Meta's own benchmark table compares directly against both.

Glimmer leads that three-way comparison on several agentic tests, including MCP Atlas at 75.5, DeepSearch QA at 74.6, τ³-Banking at 23.5, WildClawBench at 47.6 and GAIA2 at 43.3. It scores 51.2 on SWE-Bench Pro, versus 36.9 for Gemma4-31B and 50.2 for Qwen3.6-27B in Meta's evaluation.

But Glimmer does not sweep the field. Qwen leads Meta's own comparison on OSWorld-Verified (75.6 vs. Glimmer's 65.9), TerminalBench 2.1 (60.7 vs. 51.7), SkillsBench, GDPval-AA (1141 vs. 953) and most of the multimodal benchmarks. On SWE-Bench Verified, Glimmer's 76.0 lands just below Qwen's 77.2. Gemma leads on GPQA Diamond and Humanity's Last Exam.

Read honestly, the numbers make Glimmer more interesting as a specialized local-agent model than as evidence of a universal performance lead. For enterprise developers, the practical question is whether its combination of agent reliability, quantization quality, tool compatibility and decoding speed translates from benchmarks into sustained real-world workflows.

Model

Developer / origin

AA score

Parameters / context

Lowest tracked API price

Access

License

Strongest use cases

Kimi K3

Moonshot AI; China

60

2.8T total / 104B active; 1M

$3.00 input / $15.00 output via Kimi, Fireworks or Modal (pricing)

Weights

Kimi API

Custom Kimi K3 license. Large model-as-a-service operators above $20M in 12-month revenue need a separate agreement

Large products may need to display “Kimi K3.”

Frontier long-horizon coding

Multimodal research and complex tool-driven agents

GLM-5.2

Z.ai / Zhipu AI; China

53

753B / 40B active; 1M

$0.75 / $2.40 via DeepInfra FP4 (pricing)

Weights

Z.ai API

MIT

Long-horizon coding and agents

Million-token analysis with adjustable reasoning

DeepSeek V4 Flash 0731

DeepSeek; China

52

284B / 13B active; 1M

$0.09 / $0.18 via DeepInfra (pricing)

Weights

DeepSeek API

MIT

• Extremely economical reasoning• Coding agents, terminal work and tool use

MiniMax-M3

MiniMax; China

45

428B / 23B active; 1M

$0.23 / $0.96 via CoreWeave (pricing)

Weights

;

MiniMax API

MiniMax Community License. Commercial attribution required; companies above $20M yearly revenue need authorization. Includes prohibited-use conditions.

• Native text, image and video work• Long-context coding and “cowork” agents

MiMo-V2.5-Pro

Xiaomi; China

43

1.02T / 42B active; 1M

$0.35 / $0.70 via GMI (pricing)

Weights

Xiaomi API

MIT

Complex software engineering

Agents spanning thousands of tool calls

Inkling

Thinking Machines Lab; U.S.

42

975B / 41B active; 1M in weights

$0.95 / $4.05 via DeepInfra FP8 (pricing)

Weights

Tinker

Apache 2.0

Customizable text, image and audio foundation

Fine-tuned coding, RAG and tool-use systems

Nemotron 3 Ultra 550B A55B

NVIDIA; U.S.

38

550B / 55B active; up to 1M in weights

$0.37 / $1.08 via Blackbox AI (pricing)

Weights

OpenMDW-1.1; permissive commercial and derivative-model rights

Complex agents and long-context reasoning

High-accuracy RAG, code, math and science

Mistral Medium 3.5

Mistral AI; France

30

128B dense; 256K

$1.50 / $7.50 via Mistral (pricing)

Weights

Mistral API

Modified MIT. Companies above $20M consolidated monthly revenue must obtain a commercial license or use Mistral’s service.

Coding agents and function calling

Multimodal instruction following

Gemma 4 31B

Google DeepMind; U.S.

30

30.7B dense; 256K

Free on Google AI Studio’s limited tier; paid low $0.10 / $0.34 via CoreWeave (pricing)

Weights

Google AI Studio

Apache 2.0

Compact multimodal reasoning and coding

Manageable local or private-server deployments

gpt-oss-120b

OpenAI; U.S.

24

117B / 5.1B active; 131K

$0.03 / $0.17 via CoreWeave (pricing)

Weights

Numerous third-party APIs

Apache 2.0

Reasoning

structured output and tools

Fine-tuning and single-80GB-GPU deployment

Command A+

Cohere; Canada

23

218B / 25B active; 128K input

Free on Cohere’s currently tracked endpoint (pricing)

Weights

Cohere

Apache 2.0

Enterprise RAG and grounded citations• Multilingual agents and document processing

Muse Glimmer 30B

Meta; U.S.

Not yet scored

29.6B dense, including vision encoder; 131K+

No public metered hosted price located on launch day

Weights

Meta model page

Apache 2.0 for full-precision weights, quantizations, drafter and perception encoder

Always-on local agents on 24–32GB systems

Tool use, recovery, coding and screen/document understanding

Meta Glimmer adds to a still-small roster of genuinely open, frontier-class models from U.S. companies.

For the last two years, Chinese companies have set the pace in open source AI, with DeepSeek, Alibaba's Qwen team, Moonshot AI's Kimi, Zhipu's GLM and MiniMax shipping frontier-class open models under MIT and Apache 2.0 licenses on a cadence Western labs haven't matched.

The usage data reflects it: by May 2026, Chinese open-weight models accounted for roughly 61% of all tokens consumed on OpenRouter, with four of the five most-used models coming from Chinese labs — while Meta's Llama, the prior open-weight leader, fell off the rankings entirely.

The U.S. counterexamples remain countable on one hand: OpenAI's gpt-oss-120b and gpt-oss-20b, released under Apache 2.0 in August 2025 as the company's first open weights since GPT-2; Google's Gemma family, which is open-weight but ships under Google's own more restrictive custom license rather than an OSI-approved one; and Thinking Machines' Inkling.

Glimmer invites the most direct comparison to gpt-oss: both are Apache 2.0, both offer adjustable reasoning effort, and both target self-hosted deployment.

But the gpt-oss models are text-only, sparse mixture-of-experts designs built primarily for reasoning and tool use — gpt-oss-20b fits in about 16GB of memory while gpt-oss-120b targets a single 80GB data center GPU.

Glimmer stakes out different ground: a dense model with native vision input, trained end-to-end around the agent loop, shipping with its own quantized variants and speculative-decoding drafter tuned for 24GB consumer machines.

And if Zuckerberg follows through on opening Muse Spark 1.2's weights, Meta would put an actual U.S. flagship frontier model into open circulation — something no American lab has done at that tier.

Safety remains part of the deployment architecture

Giving a local model access to tools creates a different security problem from deploying a local chatbot — and Meta's own safety numbers show Glimmer is not uniformly stronger than its peers.

On CI Memories, a privacy benchmark where lower violation rates are better, Glimmer records 26.4 against Gemma's 12.1 and Qwen's 53.4. On Siren AgentDojo, a prompt-injection test, Glimmer shows a 28.4% attack-success rate versus 25.6% for Gemma and 40.3% for Qwen — while posting the highest utility score of the three at 94.2.

Meta says it evaluated Glimmer under its Advanced AI Scaling Framework and determined the model does not meet the framework's definition of "Frontier AI" because it is generally less capable than Muse Spark. Its Preparedness Team assessed Glimmer at Moderate or lower risk across chemical/biological, cyber and loss-of-control categories — the latter two inferred from the fact that Glimmer is broadly weaker than Muse Spark 1.0, which received the same designations.

The company nevertheless recommends deploying Glimmer as part of a broader system with guardrails, including human-in-the-loop confirmation for irreversible actions. That caveat matters especially for local agents: keeping data on-device reduces exposure to cloud infrastructure, but local execution does not by itself solve prompt injection, excessive permissions or an agent taking an unintended action.

Apache 2.0 weights and a fast-growing runtime ecosystem

Meta is releasing full-precision BF16 weights, both 4-bit quantized variants, the DFlash drafter and the perception encoder — all under Apache 2.0. There is no Meta API price attached to the downloadable model, leaving total cost dependent on local hardware or whatever third-party hosting developers choose. One nuance worth noting for procurement teams: as with most "open source" model releases, it is the weights that are open — Meta has not released the training data or training code.

The broader implication is that Meta is treating the developer workstation as a credible deployment target for autonomous agents, rather than merely a place to experiment with smaller language models. Glimmer's 30B size and 24GB target put that proposition within reach of high-end consumer hardware, while the Apache 2.0 license gives developers — and their legal departments — unusual freedom to modify and deploy it.

The next test is whether its benchmark advantages survive the messier conditions of real software repositories, enterprise tools and long-running agent sessions. If they do, the most consequential part of Glimmer may not be another set of benchmark scores — it may be that a class of agent previously expected to live behind a cloud API can increasingly live, and work, on the machine sitting under a developer's desk.

Token-maxxing is dead. Agentic memory is what comes next.

10 August 2026 at 14:30

Presented by MongoDB


We have been building databases as an industry for roughly 60 years. We have been building AI agents, in the form most people mean when they say the word today, for about 18 months.

Sit with that ratio for a second, because it explains almost everything about the state of agentic development right now. Six decades versus a year and a half. We are not in the middle of this learning curve. We are standing at the very bottom of it, squinting up.

There is no LAMP stack for agents yet. There is no settled, boring, default set of choices that lets a team stop re-litigating architecture and just ship.

One of the earliest lessons came from the industry’s brief obsession with token-maxxing. For a stretch in early 2026, token consumption became a vanity metric. The backlash was fast. Token volume measures activity, not outcomes.

But the interesting part of the token-maxxing story was never the workplace theater. It was the architectural lesson hiding underneath it.

The context window is the scarce resource

What follows is an aggregation of what I’ve learned from more than 100 customer conversations across 15 cities in six countries during the first half of 2026. I’m seeing organizations begin to converge on the same conclusion: the context window is the scarce resource. The challenge isn’t stuffing more information into every prompt. It’s deciding what belongs there in the first place.

That question has an answer. The answer is memory.

Not the loose way people use that word to mean “the context window,” but a real, persistent, queryable memory system that sits outside the model and feeds it deliberately.

The answer is memory, and it is more than short-term and long-term

A good agentic memory does three things that the context window alone cannot:

  • It saves what the generative model produced on previous loops and previous sessions, so the expensive reasoning you already paid for does not evaporate the moment the session ends.

  • It applies role-based access control to that saved content, so a memory created by one team can be shared across an enterprise without leaking things it should not.

  • It lets new queries retrieve the right prior content, which in practice means it is backed by semantic search rather than exact-match lookup, because agents ask for things by meaning, not by key.

That last point is where this connects back to the 60-years-of-databases observation. We spent six decades getting extremely good at storing and retrieving structured data by exact criteria. Agentic memory needs something different and newer: the ability to store the unstructured output of a generative process and find it again by similarity.

The teams building this well are the ones whose data platform can do semantic search natively, apply access control to it, and hold the generated content in the same place, rather than stitching three systems together with hope.

The pattern that is emerging in enterprises

Once you have memory like that, a genuinely interesting architecture falls out of it, and I am seeing more enterprises converge on it.

You pair the powerful memory system with a leaner model, often an open-weight one, whose job is not to be brilliant but to be a good judge. A new query comes in. The agent does a semantic search on the memory, reranks to get the best candidate answer, and asks the leaner model a single question: is this good enough to return as is, or not?

If it is good enough, you return it. You never paid for the expensive generative model at all. You answered from memory.

If it is not good enough, you escalate to the more expensive generative model, get an original solution, return that, and then save it back into the same memory system so the next session does not have to pay for it either.

Think about what that does to agentic economics over time. Every original answer the expensive model produces becomes a cheap answer the next time someone needs something similar. The system gets cheaper and faster the more it is used, which is the opposite of how naive token-maxxing scales, where cost grows linearly with usage forever. This is the difference between an agent that learns what it already knows and one that re-derives the universe on every loop.

Memory has types, and humans curate the best ones

The last piece, and the one I think separates where we are headed from where we are now, is that mature agentic memory will not be a flat bucket of short-term and long-term. It will have types, the way human memory does.

Taxonomic memory holds terminology, the controlled vocabulary and definitions an organization runs on, so the agent uses "chargeback" to mean what your finance team means by it and not what the internet at large means.

Procedural memory holds task lists and sequences, the how-we-do-this-here knowledge that turns a capable model into a useful colleague. There will be more types than these, and figuring out the right taxonomy of memory types is itself part of the learning curve we are climbing.

And here is the part that should sound familiar to anyone who has run a real production system: the best memories often get there because a human put them there. Not every memory an agent generates is worth keeping, and not every kept memory is worth surfacing first.

Increasingly I expect to see humans curating these systems, injecting the high-value memories back in for frequent reuse, pruning the noise, promoting the procedural sequence that works over the three that mostly work. We did this for knowledge bases. We did it for documentation. We will do it for agentic memory, because curation is how a corpus stops being a landfill and starts being an asset.

What comes next?

We are 18 months, give or take, into agents and 60 years into databases. The gap between those two numbers is not a problem to be embarrassed about. It is just the truth about how early it is, and it should make us humble about every "best practice" that is barely a season old.

Token-maxxing was the first big idea to rise and fall inside this new field, and its fall taught us the lesson the field most needed: the context window is scarce, so the discipline is in choosing what goes in it. That discipline is agentic memory. Semantic-search-backed, access-controlled, typed, human-curated memory that saves what was expensive to produce and serves it cheaply forever after.

There is still no LAMP stack for agents. But if I had to bet on which layer becomes the boring, default, settled choice first, the one we stop arguing about so we can get back to building, I would bet on memory. That is the next advancement in agentic development. Everything else is still hand-wiring CGI-BIN.

Pete Johnson is Field CTO, AI at MongoDB.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Your agent didn’t hallucinate; it exceeded its authority

10 August 2026 at 12:00

Content filters can block unsafe output. They cannot tell you whether an agent was authorized to issue that refund, touch that production system, or commit the company to an external action. Those are different problems, and most enterprises are only solving the first one.

An AI agent can follow its instructions perfectly and still take an action the business never sanctioned.

In commerce environments, I have seen this pattern emerge in practical ways. A service workflow calculates the correct refund amount but lacks a boundary preventing credits above what the business approved for autonomous action. An order agent correctly applies a requested change but overlooks a financing or fulfillment condition. A procurement agent identifies the lowest-cost supplier, but nobody has defined whether it can accept contractual terms or only recommend the option.

The agent keeps working. The problem may not surface until something downstream breaks.

These are not necessarily AI reasoning failures. They are failures to separate technical capability from business authority.

As enterprises move from copilots that recommend to agents that call tools and trigger workflows, every production agent needs explicit decision rights: What it may execute, what requires approval, what it may only recommend, and what it must never touch.

Guardrails remain necessary. But a guardrail is not an authority model.

Safety controls and decision rights solve different problems

Early gen AI controls screen harmful content, protect sensitive information, validate responses, and constrain tool behavior. That work matters.

Decision rights answer a different question: Even when an action is safe and technically valid, is this agent authorized to take it on behalf of the enterprise?

That governance gap is becoming harder to ignore. In April 2026, a Cloud Security Alliance survey found that 65% of respondents had experienced an AI-agent-related incident in the prior year, while 82% had discovered previously unknown agents operating in their environments. The survey involved 418 IT and security professionals and was sponsored by Token Security.

The findings illustrate how quickly agent activity can outpace the visibility and ownership structures built for conventional software.

The World Economic Forum’s May 2026 playbook reflects this shift. It introduces an Agent Capability and Authorization Profile designed to make delegated actions auditable, enforceable and accountable.

Guardrails constrain behavior. Decision rights define legitimate authority.

Give every production agent an authority contract

Before an agent receives access to enterprise tools, it needs a machine-enforceable record of exactly what authority the business has chosen to delegate. Call it an Agent Authority Contract.

At minimum, that contract should answer seven questions:

  1. Who owns the outcome? Name a human or business role, not another system.

  2. What may the agent do? Read, recommend, write, or commit?

  3. Which systems and data may it reach?

  4. What materiality limits apply? Define dollar thresholds, record counts, customer scope, and operational impact.

  5. What triggers escalation? Uncertainty, anomaly, sensitive data, or potential impact?

  6. Can the action be reversed, and who can reverse it?

  7. When does the authority expire, and how is it withdrawn?

Access control determines whether an agent can reach a system. The authority contract determines whether it may take a specific action in the current context.

Those are not the same check.

Singapore’s updated Model AI Governance Framework for Agentic AI draws a similar distinction. It treats access controls, behavioral guardrails, and human approvals as separate controls and ties oversight requirements to action scope, reversibility and potential impact.

Resolve every consequential action into four outcomes

A working decision-rights model should map every consequential agent action to one of four results.

Allow

Low-risk, bounded, and reversible actions run autonomously.

Examples include retrieving approved information, classifying an inbound request, or updating a non-material field. The agent acts without prior review because the potential impact is limited and the action can be reversed.

Approve

The agent prepares or initiates the action, but execution waits for authorization from a human or deterministic policy service.

This category covers payments, production changes, and actions that materially affect a customer, employee, or third party.

Recommend

The agent analyzes, ranks, drafts, or proposes. A named human makes the final decision.

Use this outcome when contextual judgment matters or when the legal, financial, or individual impact makes automated execution unacceptable.

Deny

The action remains outside the agent’s authority regardless of its confidence.

Deleting critical production data, making a final employment decision or overriding a mandatory compliance control should remain in the Deny category even when the agent’s underlying reasoning appears correct.

One point gets missed consistently: Deny must be enforced outside the system prompt.

A natural-language instruction telling an agent not to do something is not a technical boundary. It is a suggestion.

Make authority decisions at runtime

Static configuration cannot cover every situation.

A small service credit might be allowed under normal conditions but require approval when the amount crosses a threshold, the account is under investigation, or the request involves a regulated customer.

A practical runtime sequence looks like this:

  1. The agent proposes an action.

  2. A policy layer evaluates the agent’s identity, delegated principal, requested tool, data involved, transaction context, and potential impact.

  3. The policy returns Allow, Approve, Recommend, or Deny.

  4. The system records the authority decision, resulting action and outcome.

  5. Operational telemetry expands, narrows, or revokes the agent’s authority over time.

In enterprise commerce, the most dangerous AI mistake is not always a false answer. It can be a technically correct action the agent had no business taking.

A refund may be accurate but exceed an approval limit. An order change may match the customer’s request but invalidate a financing condition. A delivery promise may reflect available inventory while overlooking a carrier constraint applied an hour earlier.

The agent may not have failed to reason. The enterprise failed to define where its authority stopped.

Human oversight should target exceptions, not everything

Requiring human approval for every agent action looks conservative. At scale, it can quickly degrade into rubber-stamping.

When reviewers approve thousands of routine actions, attention declines and genuine exceptions become harder to identify. Singapore’s framework acknowledges that continuous human oversight of every agent workflow becomes impractical at scale and recommends meaningful checkpoints for higher-risk or irreversible actions.

Proportional authorization is the more workable model.

Low-risk actions run within narrow boundaries. High-risk or irreversible actions require approval. Unexpected behavior triggers escalation. Any consequential action without a defined authorization policy is denied by default.

The objective is not maximum autonomy. It is the highest level of autonomy the enterprise can observe, govern and reverse responsibly.

Measure whether authority is calibrated

Once agents are in production, response accuracy becomes too narrow a success metric.

Enterprises should also track:

  • Override rate: How often do humans reject or materially change what the agent decided?

  • Escalation precision: Does the agent surface genuinely risky cases, or does it return routine work to people?

  • Unauthorized-action attempts: How often does the agent try to exceed its system, data, or action scope?

  • Business-impacting error rate: How often do authorized actions produce financial, compliance, operational, or customer harm?

  • Decision latency: Are approval requirements managing risk, or slowing down automation that was already safe?

These measures turn authority into a governed operating variable.

Consistently reliable performance may justify expanding bounded authority. Frequent overrides, escalation failures, or policy violations should narrow it.

The governance gap is not in the model

Model safety, output controls, and secure tool use all matter. Enterprises should continue investing in them.

But none of those controls can answer who delegated authority, how much was transferred, under what conditions it applies, or who owns the result when something goes wrong.

An Agent Authority Contract can.

Before asking how autonomous an AI agent can become, the more useful question is: What is the enterprise actually prepared to delegate, and how will that delegation be enforced, observed, and withdrawn?

The agent demo works. That is not the hard part anymore.

Nixal Patel is a product leader. The views expressed are his own

❌