❌

Normal view

57% of enterprises have watched AI agents be confidently wrong. The fix is an agentic context layer, but who has one?

10 July 2026 at 20:58

An enterprise AI agent answers with total confidence, but the number is wrong. Nobody catches it until someone traces it back to a stale metric definition or a document the retrieval system never pulled. The model did not fail. The context it was given did.

In the past six months, 57% of enterprises traced a confident but wrong AI agent answer to missing or inconsistent business context, and 31% said it happened more than once, according to a VB Pulse June 2026 survey of 101 qualified enterprises with more than 100 employees.

The reason is not hard to find. Retrieval over documents is the default way agents get business context for 38% of enterprises, nearly double the next closest approach. The way most enterprises choose a retrieval system compounds the problem. Ease of ingestion and operational simplicity lead the selection criteria, with retrieval accuracy running behind both. The accuracy problem only shows up after the system is already live.

There is a known fix for this, a governed context layer every agent reads from instead of guessing. Vendors are racing to roll out context platforms while most enterprises are still figuring out what it is.

75% don't have an agentic context layer yet

The context layer is meant to be a shared model of what business data actually means, built once and referenced consistently instead of re-derived by every agent that touches it. 

The VentureBeat research shows the enterprise response to that idea is broad but unfinished. Twenty-five percent of respondents run one in production. Thirty-four percent are building one right now. The remaining 41% have not started.

Among companies already building or running a governed context layer, 78% report a confident-wrong failure — an AI agent that answered with total certainty and was still wrong. Among companies with no plans to build a layer, only 20% report the same thing. Companies that already got burned are far more likely to be building the fix. Companies that haven't been burned yet see no urgency.

What governed context looks like when someone actually builds one

Every major data and AI platform vendor is now building some version of this layer, and they are not converging on the same architecture. 

  • DataHub is treating catalog metadata and years of analyst query behavior as a knowledge source, then keeping it current as a living system rather than a static wiki. 

  • Microsoft's Fabric IQ is building a business ontology that any agent, not just Microsoft's own, can query over MCP. 

  • Couchbase is pushing agent memory and context retrieval down to the edge, arguing the operational database is a more natural home for it than a search or analytics layer bolted on after the fact. 

  • Pinecone's Nexus is compiling structural logic into the metadata layer ahead of runtime, betting that agents need pre-built structure more than they need faster search.

  • Snowflake runs a two-layer system, Horizon Context for customer-managed definitions and Cortex Sense for context the platform infers on its own. 

  • Oracle's Unified Memory Core takes the opposite approach, folding vector, graph and relational data into one transactional engine so there is no sync layer left to go stale. 

  • Google's Knowledge Catalog mines query logs and usage patterns to curate semantic context automatically.

  • AWS's Context service makes the same bet, a knowledge graph that gets smarter from how agents actually use it rather than from manual re-curation.

Analysts converge on one diagnosis

The vendor approaches differ. What analysts and practitioners have told VentureBeat about the underlying problem, across a run of interviews this year, does not.

When DataHub's context layer push landed this spring, Constellation Research VP and principal analyst Michael Ni framed the stakes in blunt terms. "Whoever controls runtime context controls the AI decision layer for enterprise data," Ni said. He was equally direct about how far any single product actually gets a buyer. "Vector memory isn't business meaning, business meaning isn't governance and governance isn't execution," Ni said.

In the same interview, BARC analyst Kevin Petrie pointed to a narrower but concrete gap. Most context platforms concentrate on structured tables, he said, which give agents trusted facts but miss the harder, messier context locked in documents and unstructured content, exactly the material a business actually runs on day to day.

Stephanie Walter, practice leader for AI Stack at HyperFRAME Research, made a related point earlier this year when VentureBeat asked her about enterprise context fragmentation. 

"The market is converging on the same conclusion," Walter said. "Agents don't just need more tokens or better models. They need governed, current, low-latency context." She made a similar case in an earlier review of Pinecone's Nexus launch, careful not to overstate how new any of this is. Nexus, she said, "shifts knowledge work from runtime chaos to pre-compiled structure. But it's an evolution of RAG architecture, not a complete reinvention." 

Gartner's Arun Chandrasekaran, reviewing the same launch, offered the more forward-looking read. Agentic AI, he said, is moving from pure information retrieval toward a reasoning architecture, one where long context works as short-term memory and a vector database functions as deep storage underneath it.

The fragmentation problem shows up hardest at the practitioner level, where separate tools for retrieval, memory and access control were never built to agree with each other. Steven Dickens, CEO and principal analyst at HyperFRAME Research, put it bluntly after Oracle's AI database push landed this spring. "Data teams are exhausted by fragmentation fatigue," Dickens said. "Managing a separate vector store, graph database and relational system just to power one agent is a DevOps nightmare." 

Matt Kimball at Moor Insights and Strategy, in that same story, put the production reality more simply. Getting an agent working is not the hard part, he said. The struggle is running it in production, where the goal becomes removing the distance between data and execution rather than adding another layer on top of it.

What this means for enterprises

Here's what this adds up to for enterprises building on this layer.

Retrieval alone will not close the context gap. RAG is the default source for context in most enterprises today, and it is also the layer most closely associated with the confident-wrong-answer failure. Adding more documents or a bigger index does not fix a definition that is inconsistent across systems.

The semantic context layer is where the budget is actually moving, even where it hasn't shipped. Fifty-eight percent of enterprises are already engaged — building or in production — but only 25% have actually gotten a layer live. That gap shows where enterprises have decided to spend, not where they've arrived.

No single vendor owns the architecture yet, and that is likely to stay true for a while. Enterprises evaluating this layer should expect to integrate rather than pick a single winner, at least for the next several quarters.

The buying decision is happening this year, and it is concentrated among the companies already burned by it. Fifty-seven percent of enterprises plan to switch or add a retrieval or context platform within the next twelve months. That intent is not spread evenly. Enterprises that reported a repeat confident-wrong failure plan to switch or add a provider at roughly 81%, against 32% among enterprises that never hit the problem. The companies shopping for new context tooling right now are largely the ones whose agents already got it wrong.

The agents are already running. The context underneath most of them is still being built, and the vendor selling the fix is being chosen this year.

This data will be part of a broader conversation at VB Transform 2026 on July 14 and 15 in Menlo Park: the context gap enterprises are racing to close, and which of the emerging approaches — governed semantic layers, hybrid retrieval, provider-native bundles — actually holds up in production.

OpenAI introduces ChatGPT Work, a cloud-based AI agent that manages tasks across email, Slack and calendars

OpenAI on Thursday launched ChatGPT Work, a new AI agent embedded inside its flagship chatbot that aims to transform ChatGPT from a question-and-answer tool into an autonomous work platform capable of executing complex, multi-step tasks across users' email, calendars, code repositories, and messaging apps.

The product is powered by OpenAI's latest flagship model, GPT-5.6, and is designed to go far beyond generating text. ChatGPT Work can gather context from connected apps, files, and workflows to produce finished documents, spreadsheets, presentations, reports, and websites. The agent takes a stated outcome, breaks it into smaller steps, and stays with complex projects for hours, completing them independently.

The launch marks OpenAI's clearest attempt yet to reposition ChatGPT as a workplace platform rather than a chatbot — and it arrives at a moment of extraordinary financial significance for the company. Last month, OpenAI confidentially submitted a draft S-1 registration statement to the SEC, initiating what could become one of the largest technology IPOs in history, with reported valuations clustering between $730 billion and $852 billion and annualized revenue that has blown past $25 billion.

In a short demonstration and conversation with VentureBeat on Friday, Ty Geri, a product manager at OpenAI who helped build ChatGPT Work, said the product's mission is to democratize the kind of agentic AI capabilities that OpenAI's internal engineering tool, Codex, has already demonstrated. "What's really exciting is we've seen how much Codex has been able to push the frontier of what we can get done with these AI tools, as opposed to just getting information or answers or guidance," Geri said. "Our internal adoption of Codex is literally an exponential curve across every single product function and every single use case."

Why OpenAI built a persistent virtual machine that works from the beach

The core architectural bet behind ChatGPT Work is a persistent cloud-based virtual machine that runs on OpenAI's servers, always available to the user regardless of which device they happen to be on. That marks a deliberate departure from competitors whose agents require a local machine to remain powered on and connected.

"What's really exciting about ChatGPT Work is that it's a virtual machine in the cloud that's always on for you, and this is available across all of our paid tiers," Geri said. "All Plus users are getting this. I think that's a very unique aspect of this."

The mobile-first aspect of the launch is something Geri described as "missing from the market." He pointed to the ability to create a website on a phone and share it with collaborators as a particularly novel capability. "Sites are new in general to Codex. They launched in Codex about a week and a half ago, but now we're launching also in web and mobile. You can create a site on your phone at the beach and share it with your friends," he said.

ChatGPT Work will roll out beginning with Pro, Enterprise, and Edu users, and will expand to Plus and Business users over the next few days. In the interview, Geri emphasized that the availability of the product to Plus subscribers — not just premium tiers — is central to OpenAI's strategy. "It's accessible to all paid plans, including Plus users, which in my opinion is a really big feat, and really part of that OpenAI mission, which is about bringing all this power to as many people," he said.

How MCP plugins connect ChatGPT Work to Slack, Gmail, and GitHub

The product relies on MCP-based plugins to connect to external services like Gmail, Google Calendar, Slack, and GitHub. When asked whether the plugin architecture is based on the Model Context Protocol standard, Geri confirmed: "These are all based on MCP." He added that connecting multiple Gmail accounts — a frequent user request — "is definitely on the roadmap."

The experience is designed to be action-oriented from the first interaction. ChatGPT Work offers a personalized onboarding flow that surfaces different suggested use cases depending on the user's role. Geri demonstrated how the system, detecting his role as a product manager, immediately suggested tasks like evaluating AI systems, building research artifacts, and managing his calendar. "You can start with a simple task like catch me up on Slack or Teams or read today's calendar," Geri said. He described a scenario where the system reviewed his calendar, identified scheduling conflicts, flagged meetings requiring preparation, and then — on his instruction — declined, accepted, or rescheduled events directly.

Users can also customize the agent by teaching it their writing style, organizing outputs into projects, and — in a lighter touch — choosing a virtual pet that accompanies them in the interface. The interface also introduces a hosted website feature that allows users to build and share interactive sites directly through ChatGPT Work, turning what would typically be a static slide deck into a dynamic, collaborative artifact. "Now we suddenly have a collaborative interface that's actually more exciting and more accessible than a slide deck, which has all these formatting restrictions," Geri said.

Scheduling 10 bug bashes at once: what agentic productivity looks like in practice

Geri's own usage of ChatGPT Work illustrates the breadth of tasks the system can handle. In the run-up to the product's launch, he needed to organize pre-release testing sessions — known internally as "bug bashes" — across dozens of features and team members.

"I just come to ChatGPT Work and say, 'Set up a bug bash for all the distinct features in ChatGPT Work. Add all the people that worked on that feature,' and it can check Slack, it can check GitHub, it can check Docs, and find a time that works for the four highest contributors to that feature," Geri said. "It went and scheduled 10 bug bashes, all coordinated across all those different people. That would have taken me 30 minutes at least."

But Geri pushed back against the characterization that ChatGPT Work is limited to rote administrative work. He described using it for analytically complex tasks like identifying the biggest causes of user churn for specific product features and generating product solutions — work he said would previously have taken months. "Things that we would have spent three months doing, we can now spend a week doing — and do much more, and make a much better product," Geri said. "Bugs that we would have found three or four weeks from now, we can now find within two days and fix for our users."

He also described handing off the tedium of product testing itself. "It used to be that even though like the most interesting part of my job is like what to test, I would actually end up having to spend most of my job doing the testing, which is like me taking a mouse and like clicking on the same thing over and over again, like five times," Geri said. "Instead, now I can define what do we want to test, and ChatGPT Work or Codex can actually go test it for me, deliver me that bug report, and then we can work on fixing that bug."

What OpenAI says about data privacy when AI reads your Slack and email

When pressed on data privacy concerns — given that ChatGPT Work pulls sensitive information from workplace tools like Slack, Google Drive, and email — Geri said privacy "is incredibly important, and the most important part of this is it's always in the user's control."

He pointed to OpenAI's existing enterprise security infrastructure, noting that "enterprise accounts have ZDR, and users can always opt out of letting their conversations help improve future models, which many users do." The comment aligns with assurances OpenAI made when it first launched ChatGPT Enterprise in August 2023, when the company wrote in a blog post that it does "not train on your business data or conversations."

The privacy question carries additional weight now because of the sheer volume of sensitive workplace data ChatGPT Work is designed to access. Unlike a chatbot session where a user voluntarily pastes text into a prompt, ChatGPT Work actively reaches into connected systems — reading Slack messages, scanning calendar invitations, pulling GitHub commit histories — to assemble context for its tasks. That represents a fundamentally different data surface area than anything OpenAI has offered before, and one that enterprise security teams will scrutinize carefully before granting access.

ChatGPT Work enters a three-way arms race with Anthropic and Microsoft

ChatGPT Work lands squarely in the middle of what has become the defining competitive battlefield in enterprise AI: the race to build autonomous workplace agents that can go beyond generating text and actually execute tasks.

The product arrives months after Anthropic took Claude Cowork out of preview and into general availability in April, bringing its AI agent to web and mobile platforms aimed at helping enterprise users monitor and manage long-running AI-driven tasks from anywhere. Meanwhile, Microsoft made Copilot Cowork generally available worldwide on June 16, built in partnership with Anthropic to move beyond chat and into execution. The three products — ChatGPT Work, Claude Cowork, and Microsoft Copilot Cowork — now compete directly for the attention of enterprise IT departments and individual knowledge workers alike.

The convergence is striking. All three products share a remarkably similar vision: a persistent AI agent running in the cloud that can break complex tasks into steps, connect to workplace tools via plugins, and produce finished outputs rather than just conversational replies. All three work across desktop, web, and mobile.

What distinguishes OpenAI's approach is its raw consumer distribution advantage. ChatGPT has reached 900 million weekly active users, and OpenAI now has 50 million paying subscribers. More than 9 million paying business users rely on ChatGPT for work, and 92% of Fortune 500 companies now use ChatGPT. By making ChatGPT Work available to Plus subscribers at $20 a month — not just Enterprise or Pro customers — OpenAI is betting that broad accessibility will drive adoption faster than any competitor can match.

OpenAI's product manager says AI is a partner, not a replacement — with a caveat

When asked about the potential impact on the labor market, Geri was careful with his framing. He declined to speak broadly about workforce disruption but offered his personal experience as a product manager whose day-to-day work has been substantially reshaped by the tool.

"My job is not to schedule bug bashes and find out who contributed to a specific feature. That's a task I do in my job, but that's not my job," Geri said. "My job is to make an amazing product." He described ChatGPT Work as "a partner" and "an extension of me, certainly not a replacement," adding: "Everybody feels far more productive than before, but is also almost working harder than before, because you get to work on all the things you want to work on as opposed to the drudgery around it."

But Geri was also careful not to minimize the sophistication of the work the agent can handle. "I also don't want to say that it's only doing mundane tasks because, like something like hill climbing retention curves on a given feature is not mundane. It's actually really hard to do," he said. The distinction matters. If ChatGPT Work were merely automating calendar invitations and expense reports, it would be a convenience tool. The fact that Geri describes it compressing three months of analytical product work into a single week suggests something with far greater implications for how teams are structured and staffed.

An IPO-bound company needs ChatGPT Work to prove enterprise AI can generate revenue

The timing of ChatGPT Work's launch is impossible to separate from OpenAI's IPO trajectory. The company needs to demonstrate that it can convert its massive consumer user base into durable enterprise revenue — a narrative that becomes significantly more compelling with a product explicitly designed around professional workflows.

OpenAI said it is generating $2 billion in revenue per month, growing four times faster than Alphabet and Meta did at comparable stages, with enterprise now making up more than 40% of revenue and on track to reach parity with consumer by the end of 2026. But OpenAI remains heavily loss-making, and the company does not expect to reach profitability until around 2030, with internal projections suggesting losses of $14 billion in 2026 alone.

The competitive dynamics are unprecedented. Anthropic filed for its own IPO on June 1 at a $965 billion valuation, setting up simultaneous public listings from the two most prominent AI startups in history. Whether both can sustain their lofty valuations under the scrutiny of public market investors will depend in large part on whether products like ChatGPT Work and Claude Cowork deliver measurable productivity gains to paying enterprise customers.

The launch also caps a product trajectory that began with ChatGPT Enterprise in August 2023, accelerated through the release of OpenAI's Operator agent in January 2025, and continued through Operator's deprecation and shutdown on August 31, 2025, when its capabilities were folded into the ChatGPT agent framework. ChatGPT Work is the consolidation of those efforts into a single, unified product — one that pairs GPT-5.6's three model variants (Sol for power, Luna for speed, and Terra for balanced everyday use) with a persistent cloud environment and an expanding library of MCP plugins.

The future of work may already be running in the cloud

When asked whether ChatGPT Work signals a shift toward a new kind of operating system — one where users interact with their computers primarily through an AI agent rather than through traditional mouse-and-keyboard interfaces — Geri stopped short of making sweeping predictions. But he hinted at the direction OpenAI sees ahead.

"Anybody who has worked with Codex or now ChatGPT Work will realize how exciting it is to interact with your environment and your computer via the agent," he said. "Especially in the desktop app, where the model has access to your entire machine and can interact with websites on your behalf — it's really able to be an extension of you and a real partner, and that certainly feels like the future."

At the end of the interview, Geri circled back to something personal. "I've never enjoyed work as much as I have in the last month using ChatGPT Work and Codex," he said — a striking admission from a product manager who, until recently, spent a meaningful share of his days clicking through the same interface five times in a row just to see if it would break. OpenAI is now asking 900 million users to believe that feeling scales. For a company weeks away from one of the largest public offerings in history, the answer to that question is worth roughly $850 billion.

Wall Street is debating the AI buildout. Enterprises just answered: 86% say their GPUs run at half capacity or less

Enterprise companies are running AI agents ahead of the controls needed to manage them — and they deployed that way knowingly. That is the central finding from VentureBeat Research's June survey of 573 technical leaders at companies with 100 or more employees, fielded across five parallel surveys of the agentic stack. 

Enterprises are now retrofitting to catch up with their own standards, and they are budgeting for it: Roughly six in 10 enterprises plan to switch or add vendors in each of five control layers within the next 12 months, and roughly a third — depending on the layer — plan to move within the quarter, the research finds.

There are five main layers where enterprises are building: identity for agents (which agent is allowed to do what, under whose credentials); evaluation of agent output (whether the work is any good); cost telemetry (what each agent costs to run); the context layer (the business data and definitions agents draw on to answer); and the orchestration control plane (the software that coordinates multi-step agent work).

Enterprises are already paying the price for deploying agents ahead of adequate control functions. Fifty-four percent of companies had an agent security incident or near-miss caught before harm in the past 12 months. Twenty-seven percent exercise only reactive control of agent spend — they learn what an agent costs when the invoice arrives, with no per-agent budget or ceiling in place.

Here are the five findings that anchor the set — one finding per layer of the tech stack — and what the data suggests doing first in each.

Expensive hardware is idle: 86% of GPU operators report utilization of 50% or less

Eighty-six percent of enterprises that run their own GPUs report utilization of 50% or less. Wall Street has spent the quarter debating whether the AI buildout is overbuilt. This is buy-side measurement, from the enterprises doing the buying, and the research says the most expensive hardware in buildings of these enterprises runs at no more than half its capacity.

The measurement gap compounds it: A minority 44% rigorously track what their AI compute actually costs and returns. Everyone else is only estimating. And the enterprise shopping process continues regardless: 45% of these enterprises say the emerging compute option they are most likely to evaluate in the next 12 months is an AI-specialized cloud (CoreWeave, Lambda, Crusoe, Nebius). However, under 2% of these enterprises report using one of these neoclouds today.

Moreover, roughly one in three companies appears to be considering a hedge against Nvidia: Asked which emerging compute option they are most likely to evaluate in the next 12 months, 32% of enterprises named non-Nvidia accelerators (AWS Trainium, Google TPUs, AMD), while 28% named next-generation Nvidia GPUs. The data suggests that enterprises should measure the utilization and per-workload cost of the GPUs they already own before committing budget to new compute — whether that's an AI-specialized cloud contract, new accelerators, or more GPUs. 

Most deployed "agents" do single-prompt work: 71% say a quarter or fewer complete multi-step tasks on their own

Seventy-one percent of enterprises say a quarter or fewer of their deployed "agents" can complete multi-step work on their own; the rest are single-prompt chatbots. Only 10% say true agents are the majority of what they run. To be sure, the respondents reported that they are in a position to know these things: 81% said they recommend or decide AI purchases at their companies.

That finding — that most agents are actually just chatbots in trenchcoats — lands amid adoption claims across the industry running well ahead of what enterprises are actually running. Gartner predicted 40% of enterprise applications will be integrated with task-specific AI agents by the end of 2026, up from less than 5% in 2025. It also warned that the most common misconception is referring to these AI assistants as agents, a misunderstanding known as "agentwashing."

Meanwhile, Zapier's enterprise survey said 72% reported deploying or testing autonomous agents; and Writer's 2026 survey has 97% of executives saying their company deployed AI agents in the past year. 

Those surveys asked whether companies have deployed something called an AI agent, and companies said yes. Our survey asked the people running those deployments a harder question: Of the agents you have in production, how many can complete a multi-step task without a person driving each step? The gap matters for two practical reasons. First, the inflated adoption figures are the benchmark boards and vendors use to pressure technical leaders into moving faster — and this data says the real bar is far lower than the headlines suggest. Second, the label determines the bill: A single-prompt chatbot with a human reading every answer needs none of the identity, evaluation, and cost controls this report covers, while a true multi-step agent needs all of them. 

66% let agents push to production on automated evals alone — or are engineering toward it. 5% fully trust those evals

Two-thirds of enterprises fall into one of two camps: 34% already allow an AI agent to push a code or system change to production based on automated evaluation results alone, with no human reviewing it, and another 33% are actively engineering their pipelines to allow that within the next 12 months. Only five percent fully trust the automated evaluations that would make that decision.

The distrust is earned. Half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year; a quarter watched it happen more than once. Asked to name the biggest weakness in their current evaluations, more enterprises chose “poor alignment with real-world outcomes” than any other answer — 29% of respondents.

And most of the checking happens before an agent ships, then stops. Once agents are live with real users, only 23% of enterprises run real-time quality checks on the answers those agents produce. Another 51% monitor system health only — uptime, request traces, and gateway logs — which tells them the agent is running, and nothing about whether its answers are right. The first move: Before removing human review from any workflow, test your evaluations against production outcomes rather than internal benchmarks, and instrument answer quality, not just uptime.

This finding is explored in more depth in VentureBeat's related coverage of the evaluation gap, which found that larger enterprises are moving faster toward zero-human deployment while also failing more often — and outlines a regression-testing framework built on production outcomes rather than internal benchmarks.

69% run credential sharing somewhere in the agent fleet — and those companies get hit far more often

Sixty-nine percent of companies allow agent credential sharing somewhere in their agent fleet during runtime – meaning multiple agents operating under one API key or service account. Those companies were far more likely to get hit: Organizations with credential sharing anywhere in the fleet experienced a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (9 of 22) where every agent has its own scoped identity. 

The takeaway for enterprises is this: Give every agent its own scoped identity, starting with the agents that touch production systems.

57% traced a confident, wrong agent answer to their own missing or inconsistent business context

Fifty-seven percent of enterprises traced at least one confident, wrong agent answer in the past six months to missing or inconsistent business context: wrong metrics, stale definitions, absent documents. Most of them watched it happen more than once.

Most enterprise companies are fixing this, even though they’ve moved forward with agent deployment already: 25% already run a governed semantic layer, or one governed definition of the business that every AI reads from, in production. However, 34% are still building one, and 41% haven't started. The takeaway: Govern the definitions your agents answer from, metrics and entities first, before scaling the agents that depend on them.

The quarter where agent technology “portability” became a priority

One more shift is worth reporting with its limits stated plainly. In our spring orchestration survey wave, the top concern about provider-controlled orchestration was security and permissioning limits (32%). By June, vendor lock-in led at roughly a third, with security limits at 28%. 

Those are two snapshots one quarter apart, and here’s one possible explanation for why portability became a top issue for enterprises. Our June survey went into market after a June 12 U.S. Commerce Department export order took Anthropic's Claude Fable 5 offline for enterprises for roughly three weeks. Meanwhile, Chinese company Z.ai released GLM-5.2's open weights under an MIT license on June 16 at roughly one-sixth of GPT-5.5's price; and Tencent's Hy3 arrived July 6 under Apache 2.0; and OpenAI previewed GPT-5.6 on June 26 to a small group of government-vetted partners, opening it broadly on July 9 after the government's review cleared. The open-weight releases in particular promise enterprises more control over their agents, and while we haven't established a causal link here, the timing is worth noting.

The posture data matches the mood: 51% now expect their primary control plane for enterprise agents to be hybrid — provider-native plus external orchestration — by the end of 2026, up from 34% in the spring survey wave. Enterprises reporting that they rely purely on provider-managed agent services fell from 12% to 7%.

Five layers, no incumbents, 12 months

The synthesis across all five surveys reveals a huge “buying” window. In each of the five control layers, 57% to 64% of enterprises plan to switch or add vendors within 12 months — 64% in infrastructure and in evaluations, 59% in agent security, 57% in retrieval and context — and 26% to 38%, depending on the layer, plan to move within a quarter. No layer has an established incumbent: The most common evaluation tooling is the model provider's built-in evals, tied with no dedicated tooling at all (17% each); 82% of respondents name provider-native or hyperscaler controls as their primary agent security layer; and provider-native retrieval leads the context technology layer (RAG, etc) as well. 

Most enterprises are defaulting today to the built-in tools that ship with the big AI platforms they already use: Anthropic, OpenAI, Google, Microsoft, and AWS. That holds true across every one of these agentic technology layers: enterprises are looking to their primary cloud and model providers to supply the guardrails, evaluations, and retrieval solutions already bundled into those providers' offerings.

Those defaults are winning on convenience, and they're also what the coming spending decisions will test. The survey didn't ask which direction that money moves — toward the platforms' built-in tools or toward the specialists challenging them — which is exactly why every contract in these five layers is worth watching over the next four quarters.

The Q3 survey wave will measure whether the enterprises made good on these budget plans: whether their agents gained scoped identities, whether evaluations got tested against production outcomes, whether GPU utilization rose, and whether the semantic layers under construction shipped.

VentureBeat will release the full Q2 reports across all five VB Pulse trackers at VB Transform, July 14–15 at Hotel Nia in Menlo Park, where we convene enterprise technical leaders building autonomous agents in production. 

Disclosure: VentureBeat produces both this research and VB Transform

Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them

Enterprise AI teams are giving agents more freedom at the same moment their confidence in automated testing is collapsing.

Half of enterprises have deployed an AI agent or LLM feature that passed internal evaluations and yet still caused a customer-facing failure — one in four more than once — according to the June 2026 VB Pulse survey of 157 qualified enterprise respondents at companies with 100 or more employees.

The sample is self-selected rather than a probability sample, so the findings should be read as directional, not precise.

But enterprises are not responding by slowing automation: 66% of respondents already permit some production deployment without human review or are building systems intended to do so within the next 12 months. Only 5% say they fully trust the automated evaluations that would make those release decisions.

That mismatch is the evaluation gap: the autonomy ceiling is rising faster than the assurance beneath it. 

It also fits a broader thesis that will be explored at VB Transform 2026: enterprises ship agents first, while the control layers around identity, evaluation, cost, context and orchestration are arriving later. The next year will be a retrofit cycle, with buyers shifting budget toward the systems that make agentic deployments governable and dependable.

Why a passing evaluation is not a working agent

Traditional software testing usually asks whether a defined input produces an expected output. Agent testing is harder because the system may choose its own sequence of steps, call tools, retrieve data, alter state and respond differently from one run to the next.

An agent can make several individually plausible decisions and still reach the wrong result. It may retrieve the correct account but update the wrong field. It may draft a valid refund request but send it without approval. It may call five tools successfully before a sixth step leaks sensitive information or leaves a workflow incomplete.

The survey shows enterprises already recognize this limitation. The most common reason for distrusting automated evaluation is poor alignment with real-world outcomes, cited by 29% of respondents. Bias or inconsistency follows at 21%, lack of explainability at 18%, and data leakage or privacy concerns at 17%.

That hierarchy matters. Enterprises are saying the score often does not predict what happens when a customer, employee or business process encounters the agent in production — not that automated scoring is too slow or expensive.

NIST makes a similar point in its Generative AI Profile: measurements gathered in controlled environments may not transfer cleanly to deployment because behavior changes with prompts, users, context and operating conditions. Its guidance calls for field testing, post-deployment monitoring and clear processes for escalating failures.

Capability is not consistency

A single successful run proves that an agent can complete a task. It does not prove that it will complete the task reliably.

Anthropic’s guidance on agent evaluation distinguishes between measuring whether a system succeeds at least once across repeated attempts and whether it succeeds every time. That distinction is essential for customer-facing or operational workflows. A model that occasionally produces an excellent answer may still be unacceptable if the same task fails unpredictably on the next attempt.

Enterprise teams should therefore treat repeatability as a first-class metric. That means running the same scenario multiple times, varying phrasing and context, testing tool failures, and measuring whether the final business outcome remains correct even when the route changes.

The evaluation set also has to evolve. Every production incident should become a permanent regression test. Customer escalations, failed tool calls, incorrect approvals and data-handling mistakes should feed back into the pre-deployment suite rather than remaining isolated support cases.

Autonomy should expand by risk, not by ambition

The survey does not imply that every agent action should require a person. Human review cannot scale across millions of low-consequence decisions.

But zero-human operation should be earned by demonstrated reliability and bounded by the consequences of failure.

Low-risk actions such as drafting internal summaries or categorizing documents can tolerate broader autonomy. Financial transactions, customer communications, code deployment, access-control changes and data deletion need stricter thresholds, repeated consistency tests, policy checks, rollback mechanisms and clear human escalation paths.

The risk isn't evenly distributed by company size, either. Larger enterprises — those with 2,500 or more employees — are moving toward zero-human deployment fastest, at 70% versus 64% for smaller companies, and they're also shipping more agents that go on to fail a customer, at 54% versus 48%. 

That is the warning for enterprise leaders. Removing the human from the loop does not remove uncertainty. Without stronger assurance, it converts uncertainty into an automated production decision.

The market will keep pushing toward greater autonomy because the economic incentive is real. The organizations best positioned won't be those that remove people fastest — they'll be the ones that treat repeatability and regression testing as seriously as deployment speed.

Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading

10 July 2026 at 18:17
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,...

Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states, communication buffers, and intermediate activations all compete for GPU high-bandwidth memory (HBM). As model size, sequence length, and batch size grow, HBM capacity often becomes the primary scaling bottleneck. This post explains how…

Source

Amazon and University of Michigan give robots a sense of touch

10 July 2026 at 17:13
From warehouse automation to surgical assistance, many real-world applications depend on robots performing delicate, contact-intensive tasks. Often missing in these situations is the sense of touch: robots need to feel the forces on their fingertips to manipulate objects effectively. Despite years of effort, robust and scalable solutions to this problem remain out of reach, especially in industrial settings. One approach has been to use vision-based tactile sensors, in which cameras embedded in soft fingertips capture contact geometry. Researchers have used this approach to estimate object shape and pose, but computing the forces that correlate most with manipulation capabilities remains a challenge. Modeling tactile shear — the forces that arise when an object slides or rotates against a sensor — is crucial for building robots that can grasp objects, use tools, and perform complex manipulation skills. Our solution, HydroShear, gives simulators the ability to accurately model tactile forces, enabling robots to learn dexterous, contact-rich manipulation policies entirely in simulation. These policies transfer seamlessly to the real world with no modification, achieving a 93 percent average success rate across four challenging tasks. Bridging the tactile reality gap Simulators for robot locomotion have found success in real-world applications because physics engines model rigid body dynamics and proprioceptive sensing well. But subtle tactile forces and shear feedback are notoriously difficult to simulate accurately. This has made it nearly impossible for tactile sensors trained on simulators through reinforcement learning to succeed when deployed on real robots. Existing tactile simulators face a fundamental trade-off. Physics-based methods like finite-element methods accurately model contact forces but are too slow for training reinforcement learning policies at scale. Faster approximations, on the other hand, oversimplify how forces build up and change during contact. They miss critical events like the moment a gripped object begins to slide or the way a soft sensor deforms over time. Modeling touch with fidelity and speed HydroShear’s key innovation is to add new capabilities to an existing physics simulation technique known as hydroelastic contact models. Called path-dependent force tracking, this approach accurately tracks how forces accumulate over a soft sensor membrane during a physical interaction. Rather than computing forces based only on instantaneous contact, HydroShear remembers the motion history of the object as it moves across the sensor. More concretely, when a robot grasps an object and moves it, different points on the object's surface come into contact with the sensor at different times. HydroShear tracks each of these contact points individually, computing how the soft elastomer deforms as the object moves. It then converts these deformations into realistic force fields, accounting for friction, slipping, and the elastomer's material properties. The simulator handles full 3-D motion — tilting and rolling as well as in-plane sliding — which is essential for dexterous manipulation. It's also GPU parallelizable, enabling efficient large-scale policy training. We calibrate HydroShear by collecting controlled real-world data with a robot arm and vision-based GelSight Mini tactile sensors. The calibration isolates four key parameters: how forces dissipate across the sensor surface, how tangential and normal forces build up, and the friction coefficient between the object and the elastomer. This systematic approach ensures that the simulator accurately reproduces real tactile feedback. Validation: From simulation to real robots We evaluated HydroShear on four contact-rich manipulation tasks, each highlighting different challenges. In all tasks, the robot perceives touch and proprioception (joint positions and gripper state) and has no access to object poses. Peg insertion: The robot grasps a cylindrical peg at an unknown orientation and must insert it into a tight socket. Because the grasp pose varies with each trial, the robot must use tactile feedback alone to detect and correct alignment errors during insertion. Bin packing: The robot inserts a cube into a target slot within a crowded bin. Neighboring cubes partially block the slot, so the robot must push through multiobject contact while sensing forces from multiple directions simultaneously. Book shelving: The robot inserts a book laterally into a shelf, with gravity pulling perpendicular to the insertion direction. The book is larger than the fingertip, producing broad contact patches that make it difficult to localize the object from touch alone. Drawer pulling: The robot pulls open a drawer while external-force perturbations are applied at random times. The robot must detect when the handle begins to slip and tighten its grip just enough to maintain hold without crushing it. We trained reinforcement learning policies entirely in simulation using HydroShear, then deployed them on a real Franka robot with GelSight Mini sensors without any modification or fine tuning. HydroShear achieved a 93% average success rate across all four tasks. We compared against two strong baselines: TacSL, which uses simplified force approximations, and FOTS, a recent learning-based method. TacSL achieved only 34% success, while FOTS reached 58 to 61%. The performance gap underscores the importance of accurate tactile shear simulation. Interestingly, the performance difference correlates directly with simulation fidelity. On tasks like peg insertion, where precise force feedback is critical, HydroShear's advantage is most pronounced. On drawer pulling, which requires detecting and reacting to slippage, HydroShear's path-dependent force tracking proves essential. Faster and less expensive Accurate tactile simulation unlocks a powerful recipe for robot learning: train policies entirely in simulation, then deploy them on real robots. This approach is dramatically faster and cheaper than learning from real-world interactions, which can damage sensors and require extensive trial and error. For warehouse automation, the approach is particularly valuable. Tasks like bin packing, sorting, and careful handling require robots to feel their way through complex interactions. HydroShear enables robots to learn such skills without extensive real-world data collection. While HydroShear yields strong results when coupled with vision-based tactile sensors like GelSight, the underlying principles could extend to other tactile modalities. We're also exploring how higher-resolution sensor simulations and more-complex object geometries could further improve performance.

Meta’s Iris push signals the next phase of AI infrastructure

Abstract digital illustration of a circuit board pattern with interconnected nodes and pathways in cyan and black, representing technology infrastructure and connectivity.

Meta is preparing to manufacture its own AI chip for the first time. According to an internal memo, the company expects production of its proprietary processor, Iris, to begin in September.

After clearing bug testing in about six weeks, the chip — reported on by Reuters — is expected to take on some of the inference work currently running on third-party GPUs, giving Meta more control over how it builds and scales its AI infrastructure.

It’s unmistakable that this could be the company’s most important move yet toward in-house silicon for AI workloads.

But anyone can see this isn’t really about the hardware. It’s unmistakable that this could be the company’s most important move yet toward in-house silicon for AI workloads. The timing, as Meta is locked in an aggressive multi-billion-dollar infrastructure race, is critical. It’s clear that the company’s CEO, Mark Zuckerberg, wants to grow into the AI titan he believes the company can be, but it’s nearly impossible when the competition controls the core infrastructure.

Custom silicon for inference

Iris is designed for a specific job inside Meta’s AI infrastructure as custom silicon optimized for Meta’s heavy workloads.  Iris expands Meta’s Meta Training and Inference Accelerators (MTIA) program, which is intended to move targeted AI inference workloads onto custom silicon.

The processor would handle workloads that drive content ranking, recommendations, and generative AI services across Meta’s family of applications, including Facebook, Instagram, and WhatsApp.

  • The MTIA 300 is already deployed in production to run ranking and recommendation inference across Meta’s platforms.
  • The 450 and 500 variants target generative image and video inference through 2027.

By shifting these high-volume inference tasks to custom silicon, Meta can lower data center costs while bypassing the traditional hardware supply bottleneck for its day-to-day operations.

Securing the AI supply chain

Meta’s modular, rapid-fire approach to custom silicon is aggressive versus traditional industry timelines. The company plans to drop a new iteration roughly every six months through 2027.

Meta is working with Broadcom to design Iris, while TSMC will manufacture the chip. But custom silicon is only one piece of the equation. Scaling AI infrastructure also requires a steady supply of memory, storage, and networking components at a time when demand for AI hardware continues to strain global supply chains.

To support that expansion, Meta has also been securing key components across its supply chain. The company has signed long-term agreements for high-bandwidth memory from Samsung Electronics, flash storage from SanDisk, and fiber-optic networking equipment from Sumitomo Electric.

The strategy mirrors similar investments by other hyperscalers. Google continues to expand its TPU program, while Amazon has developed its Trainium and Inferentia processors.

Scaling to 14 gigawatts

The Iris rollout is one component of Meta’s broader AI infrastructure expansion. The company plans to bring roughly 7 gigawatts of computing capacity online this year, then double that to 14 gigawatts in 2027. At that scale, Meta’s AI infrastructure would consume more electricity than many small countries.

At that scale, Meta’s AI infrastructure would consume more electricity than many small countries.

And, scaling AI infrastructure at this level comes with an enormous price tag. Meta has projected 2026 capital expenditures of between $125 billion and $145 billion, making it one of the largest single-year infrastructure investors in corporate history

Meta just pulled off something rare, which is essentially convincing Wall Street that spending more money is actually a good thing.

Wall Street rewards AI spending

Yet Meta just pulled off something rare: essentially convincing Wall Street that spending more money is actually a good thing. Following a trillion-dollar wipeout in tech market cap amid investor nervousness about the sheer scale of AI spending, Meta’s shares climbed roughly 8%.

With new MTIA chips planned roughly every six months through 2027, Meta is betting that vertically integrated AI hardware can deliver lower inference costs and better performance than relying exclusively on merchant silicon. By bringing chip design in-house and securing critical components across its supply chain, the company is slated to scale AI infrastructure with greater control over cost, deployment, and optimization.

The post Meta’s Iris push signals the next phase of AI infrastructure appeared first on The New Stack.

Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead

10 July 2026 at 16:41
Decorative image.There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead,...Decorative image.

There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead, along with multiple ways to apply it in NVIDIA CUDA code. A common bottleneck when writing GPU code is that GPU compute is so fast that even high-bandwidth device memory doesn’t use the GPU kernel fully. Kernel fusion addresses this by…

Source

AI Model Co-Design: Hardware-Friendly LLM Design

10 July 2026 at 16:36
AI performance comes down to three dimensions:  Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a...

AI performance comes down to three dimensions: Deployments must balance all three: High accuracy is wasted if responses are slow, and raw throughput means little if each user’s experience is laggy. Practical systems therefore optimize accuracy, throughput, and interactivity together. This post focuses on throughput and interactivity, and how model-design choices shape both without…

Source

Video Friday: A World Cup for Robots

10 July 2026 at 16:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH
Humanoids Summit Seoul: 22–23 September 2026, SEOUL

Enjoy today’s videos!

For the first time, two full teams of humanoid robots played an 11-vs-11 soccer match on hardware, bringing one of robotics’ most ambitious long-term visions closer to reality. Never before have two full-sized humanoid robot teams played a soccer game against each other.

[ RoboCup ]

Engineers at MIT and EPFL in Lausanne, Switzerland, have designed a robot that can swim underwater and flap out of the water to continue flying through the air, much like a diving bird. The robot can help scientists study the mechanics that enable these actions in aquatic aviators and may help launch a new class of aerial-aquatic drones and vehicles.

[ MIT ]

We’re excited to announce our breakthrough robotic hands for the NEO platform: hands that match or exceed human-level dexterity, strength, safety, and reliability. Designed from the ground up, these 25-DoF hands combine 25 fully actuated degrees of freedom with a tendon-driven system, rich tactile sensing, and built-in compliance. The result is a hand capable of true in-hand manipulation, precision tool use, and delicate interaction.

[ 1X ]

This match, Tech United played against IRIS at the midsize league at RoboCup 2026 in Incheon, South Korea.

[ Tech United ]

Atlas arrived pitchside at NYNJ Stadium in front of 80,000 people gathered to see Brazil vs. Norway. After performing some of the sport’s most memorable player celebrations, Atlas helped kick off the second half by delivering the match ball!

[ Boston Dynamics ]

Navigating discrete terrain such as stepping stones remains a major challenge for legged robots. Conventional approaches often rely on dense environment reconstruction from cameras or lidar, which can be affected by latency, occlusions, and significant computational overhead. We show that proximity sensors integrated into the bottom of a quadruped’s feet enable safe, terrain-seeking autonomous locomotion.

[ Paper ]

On this holiday, Digit is on grill duty. It turns out precise force control is good for more than payload handling. Happy 4th of July from all of us at Agility.

[ Agility ]

We’ve created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99 percent on tasks where previous models achieve 64 percent, completes tasks roughly 3x faster than state-of-the-art, and requires only one hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step toward our mission of creating generalist intelligence for the physical world.

[ Generalist ]

Four years at Figure.

[ Figure ]

Reachy Mini is becoming your real AI companion. The Conversation App makes it able to talk fluently with you, help you with your to-do list, remind you of important tasks, and even chat about music. Long-term memory, voice interaction, always ready to help.

[ Reachy Mini ]

Is this sort of thing now a real job for humanoid robots, then?

[ Unitree ]

Quite a story, but is it a real job?

[ EngineAI ]

If you have a cute animal logo for your research, I will always share it.

[ BIEVR-LIO ]

This is very delicate work, although the real challenge would be picking those nuts out of a jumbled bin full of randomly sized nuts, which is how most of us live our lives.

[ Sanctuary ]

Not for me, thank you, although I’m not saying that most of the other humanoid robots out there are any better looking, fundamentally.

[ UBTECH ]

Robotics professor Dr. Christian Hubicki judges robot soccer skills while knowing very little about soccer himself.

[ ORL ]

In this presentation, Brendan Schulman, vice president of policy at Boston Dynamics, outlines the critical role of government engagement in driving the success of the humanoid robotics industry. He demonstrates how legged robots like the Spot quadruped and Atlas humanoid are moving beyond factory settings to deliver real-world value in infrastructure inspection, industrial manufacturing, and public safety. Schulman highlights the intersection of AI and robotics, showcasing how large behavioral models and reinforcement learning enable robots to navigate slippery floors and autonomously avoid workplace hazards. Ultimately, he calls for a proactive national robotics strategy focused on workforce training, safety standards, and ethical frameworks to support supply chain resilience and global competitiveness.

[ Humanoids Summit ]

Google's TabFM skips per-dataset training and still predicts on tables it's never seen

The vast majority of business data is tabular — living in data warehouses, CRMs, and financial ledgers — yet building a reliable model from it still means training a new one from scratch for every dataset, then maintaining hyperparameter tuning loops, feature engineering, and retraining pipelines to fight data drift. Google Research is proposing a way around that: a new foundation model called TabFM that treats tabular prediction as an in-context learning problem instead.

It can generate predictions for a new, unseen table in a single forward pass. For enterprise developers and AI engineers, this reduces the time-to-production from weeks of pipeline engineering to a single API call.

The challenge with traditional ML

To extract reliable predictions from a gradient-boosted tree, data scientists must build and maintain complex data pipelines. They have to clean messy inputs, impute missing values, encode categorical variables into numerical formats, and engineer custom feature crosses.

Once the data is ready, they must run repetitive hyperparameter optimization loops, searching across learning rates, tree depths, subsampling ratios, and regularization grids to find the best configuration. 

Once deployed, these traditional models "incur ongoing operational debt through data drift monitoring and retraining pipelines to stay accurate," Weihao Kong, Research Scientist at Google Research, told VentureBeat.

Meanwhile, the rest of the AI industry has moved on. Generative AI models for text and computer vision have seamlessly shifted to zero-shot inference, where a model can perform a completely new task simply by being prompted with context. 

Large language models (LLMs) already excel at in-context learning, so why can't we just feed tables into an off-the-shelf LLM?

Because LLMs are trained on natural language rather than structured data, they struggle to process tables directly. First, their context limits are exhausted quickly by medium-sized tables containing just a few thousand rows and hundreds of columns. Second, LLMs suffer from tokenization inefficiency, awkwardly splitting numerical values and destroying mathematical precision. Finally, they suffer from structural blindness. When a 2D table is serialized as a 1D text string, LLMs lose track of which value belongs to which row and column as the table grows. 

"That's why, today, it is far more effective to use an LLM to write the code that handles feature engineering and calls XGBoost than to ask the LLM to read the table itself," Kong said.

What is TabFM?

To run inference with TabFM, you do not update any model weights. Instead, you take your historical examples (the training rows with their known labels) and your target rows (the new data you want to predict) and pass them to the model as a single, unified prompt. The model learns to interpret the relationships between columns and rows directly from this context at runtime.

For example, consider an enterprise analyst trying to predict customer churn. Instead of building a bespoke data pipeline and training an XGBoost model, they can simply pass a sample of historical user session data alongside a new, active session into TabFM. In one forward pass, the model returns an instant churn probability. 

TabFM overcomes the limitations of LLMs by treating the data as a grid, preserving its structural integrity without forcing it into a single-dimensional text string.

To effectively process diverse tabular structures while enabling scalable zero-shot prediction, TabFM synthesizes the strengths of earlier experimental architectures, TabPFN and TabICL. TabPFN, developed by Prior Labs, first proved that a transformer architecture could perform zero-shot classification on small tables, though it struggled to scale computationally to larger datasets. 

Later, TabICL, developed by France's National Research Institute for Digital Science and Technology, addressed this bottleneck by introducing row compression, allowing in-context learning to efficiently process much larger tables. 

TabFM combines TabPFN's deep feature contextualization with TabICL's efficient compression into a novel hybrid design built on three key mechanisms:

1. Alternating row and column attention: The raw table is first processed through a multilayer attention module that alternates across both columns (features) and rows (examples). By continuously attending across these two dimensions, the model natively captures complex feature interactions. This deep contextualization does the heavy lifting that would usually require tedious manual feature crafting by data scientists.

2. Row compression: Following this contextualization, the cross-attended information for each row is compressed into a single, dense vector representation. TabICL pioneered this by using CLS tokens to compress a row's rich information into one vector, "in contrast to TabPFN v2, v2.5, and v2.6, which attend over the full cell grid throughout the network," Kong explained. This drastically shrinks the computational footprint.

3. In-context learning (ICL): A causal Transformer then operates on this sequence of compressed embeddings. This Transformer model uses the attention mechanism of TabICL to attend over these dense row vectors, drastically reducing the computation cost and allowing the model to process large datasets efficiently.

A major selling point of TabFM is its pretraining recipe. The model was trained entirely on hundreds of millions of synthetic datasets. These datasets were dynamically generated using structural causal models (SCMs) that incorporate a wide variety of random functions. By training exclusively on synthetic SCMs, TabFM learned the fundamental mathematical priors of how tabular features interact without ingesting real-world, confidential CSV files.

TabFM in action

To test the model's capabilities, Google researchers benchmarked TabFM on TabArena, a comprehensive evaluation suite spanning 51 diverse tabular datasets across 38 classification and 13 regression tasks.

On these public benchmarks, TabFM's zero-shot predictions already match or beat heavily tuned supervised baselines. However, Google is careful to note that this does not automatically mean TabFM will universally dethrone bespoke, hyper-optimized production models on every enterprise workload.

"Instead of replacing hyper-optimized production models, the true practical business value it unlocks for lean engineering teams is velocity," Kong said. "It allows data analysts and backend engineers to instantly spin up high-quality baseline models without a dedicated data science team managing a complex lifecycle."

For advanced practitioners looking to squeeze out maximum accuracy, the research team also introduced a "TabFM-Ensemble" configuration. By running the model through 32 distinct variations and blending the results, TabFM pushes the performance even further. 

Getting started, trade-offs, and the cloud future

The shift to in-context learning for tables introduces a new economic trade-off that engineering teams must consider. 

With traditional algorithms, training is slow and expensive, but inference is lightning-fast and cheap. TabFM flips this dynamic. While training time drops to zero, inference becomes significantly heavier. Because the model must process the entire historical dataset as context during every single prediction, it requires more compute and memory at runtime. 

In this new paradigm, "traditional machine learning training becomes the 'prefill' phase (KV caching) in the context window," Kong said. While this prefill cost is steep, it is paid only once per table, and the cache is reused across subsequent queries. "The catch is prediction latency, which no amount of caching removes," Kong added. Every new prediction requires a pass through a large transformer. "Any production API requiring single-digit-millisecond response times cannot tolerate TabFM's forward-pass overhead."

For developers looking to evaluate the model today, the barrier to entry is low. Google designed TabFM as a drop-in replacement for traditional ML workflows, offering a scikit-learn compatible API (TabFMClassifier and TabFMRegressor). It natively handles mixed numerical and categorical columns, works directly with pandas DataFrames, and requires no manual ordinal encoders or numerical scalers. The library supports both JAX and PyTorch backends.

However, enterprise teams need to be aware of current limitations and licensing restrictions. The model architecture has a hard limit of 10 output classes for classification tasks, and it is optimized for tables with up to 500 features. More importantly, while Google released the underlying codebase under the permissive Apache 2.0 license, the pre-trained model weights are published on Hugging Face under a strict tabfm-non-commercial-v1.0 license. Developers can evaluate the model internally, but it cannot be deployed in commercial products yet.

Looking ahead, Google is addressing the commercial deployment friction through its cloud ecosystem. TabFM is being integrated directly into Google BigQuery, allowing analysts to run zero-shot predictions natively via an “AI.PREDICT” command. By putting foundation model inference right next to the data warehouse, TabFM could soon make complex tabular machine learning as accessible as a basic database query.

In practice, TabFM shines in rapid prototyping, high data drift environments, and small to medium-sized datasets under 100,000 rows. Conversely, teams should stick to traditional models for strict, ultra-low latency APIs, or massive tables exceeding one million rows, which currently require aggressive row sampling that degrades the foundation model's competitive advantage.

Why retrieval quality is becoming the defining challenge in AI agent architecture

Neon digital waves and scattered data particles on a dark background, representing hybrid search, data pipelines, and AI engineering infrastructure.

Agentic systems usually have two jobs: Build context, then use that context to produce an answer or action.

Many failures that look like LLM problems start in the context-building step. The answer the LLM gives is limited by the context it was given, or it finds through tool calls. If the agent model cannot find the right sources, then improving the generation model will not improve the overall system.

“Many failures that look like LLM problems start in the context-building step.”

A client, Specstory, wanted to give users the ability to ask questions from the agent’s history. For example, why a team chose Authlib for authentication and what alternatives they considered. The chatbot needs the right prior conversations, decisions, and tradeoffs from a large corpus of coding sessions. The model and system prompt help only after those chat turns have been retrieved and are in context.

If retrieval ranks implementation snippets above the discussion where the team weighed alternatives, the agent can still produce a confident answer. It may find code that imports Authlib and a few inline comments, then describe the decision based on implementation evidence rather than the actual trade-off discussion.

The same pattern showed up in an AnkiHub operator review in our private community. A request for help studying based on lecture slides only works if the agent’s tool calls retrieve the right flashcards. The hard part is not finding any related cards. A lecture on the function of the heart may match hundreds of cards. Ranking decides whether the core cards make it into context or whether the system has to raise top_k and flood the prompt.

The exact setup changes by product. The context-building step might use local search, semantic search, web or API calls, or database queries. It might be handled by an agent, a fixed workflow, or application code. The process stays the same: gather the right context, then generate from it.

For example, a coding agent runs rg, opens files, reads logs, and inspects tests before writing a patch. A research agent searches the web and internal notes before writing an answer. A study assistant searches deck facts and user context before suggesting what to learn next.

When context building fails, the symptoms look like generation failures.

Retrieval failures mimic generation bugs

SymptomRetrieval cause
HallucinationThe answer source never made it into context.
Context rotLow recall forces a high top_k, so noisy results fill the context window.
LatencyWeak retrieval leads to more tool calls, larger candidate sets, and larger context windows.

A better model helps with reasoning and writing, but it cannot give a better answer without the right context.

“A better model helps with reasoning and writing, but it cannot give a better answer without the right context.”

The Mixedbread OfficeQA-Pro Eval shows the same pattern at the benchmark scale. OfficeQA-Pro uses 89,000 pages of financial documents, dense tables, scanned PDFs, and questions that require reasoning across documents. Giving Codex better search tools reduced tool calls and improved answer quality.

A scatter plot mapping Accuracy (%) against Tool Calls for three AI configurations.

Plain-text tools like grep and rg work (ish) on flat code files. They do not work well when context lives in PDFs, tables, chat histories, multi-modal inputs, web results, and permissioned data. In those cases, the agent needs a retrieval that can combine exact terms, meaning, metadata, permissions, and ranking quality.

Retrieval needs traces and evals

Once retrieval enters the architecture, the next question is whether it finds the right information.

For that, you need traces and evals. For each retrieval step, the minimum trace is the input, the outputs, and a way to label whether each output was relevant.

A flowchart diagram illustrating a data workflow where a horizontal sequence connects four steps: Input, Tool call, Output, and Label.

For a coding agent using rg the input is the command, the output is the returned snippets, and the label says which snippets helped, which were noise, and which relevant files were missing.

For product retrieval, the step might be BM25, semantic search, hybrid search with reranking, a generated SQL query, or something else. Capture the query or arguments, the returned documents or chunks, and whether those results were helpful.

Trace each retrieval step by itself, then evaluate the full context-building pass. The local trace answers “Did this query return useful material?” The full trace answers “Did the system collect everything the model needed before generation?” If it did, failures are a generation problem. If not, it’s a retrieval problem.

You cannot know where the failure started or what to fix without traces.

Different failures need different fixes

“Improve retrieval” is too broad to be useful, as different problems require different solutions. If a relevant document is missing, the trace should show where it disappeared: query building, retrieval, filtering, ranking, or final context assembly.

A sequential flowchart which maps a five-stage pipeline—Query builder, Retriever, Filters, Ranking, and Context—with each stage pointing down to its respective failure mode.

The failed step, plus what the trace shows, tells you what change to make.

Failed stepWhat the trace showsChange to make
rg / grepA conceptual query returns literal matches while missing relevant files.Add semantic search over files or chunks, or generate better keyword queries before calling rg.
BM25The query uses the right concept but different words from the source material.Add semantic search, synonyms, or query expansion.
Semantic searchExact names, error strings, document IDs, or domain terms are missing from the results.Add a keyword or BM25 path, or boost exact term matches.
Hybrid retrievalThe relevant passage is ranked 7th, but the context only takes the top 5.Add or tune a reranker, or raise candidate top_k before reranking.

The right fix depends on what the system was trying to retrieve. A decision-history question requires the decision, the alternatives, and the chats in which the team worked through them. A study question depends on the lecture material, deck metadata, semantic matches, and the user’s study context.

The architecture

Once you trace individual retrieval calls, the full architecture has a simple shape: fan out to context-building tools, then fan in to generate the final output.

A system architecture diagram showing a RAG pipeline.

The retrieval layer might be a search engine, a vector database, an SQL query, a local file tool, a web search API, or a custom service. The pattern stays the same: build candidate context, narrow it, rank it, assemble it, then generate from it.

Give agents human search controls

Semantic search compares embeddings (numerical representations of meaning). It helps when wording differs, but most retrieval intents also depend on structured constraints. A meeting search box can use semantic search over transcripts and notes, but a useful interface also lets someone filter by person, date, project, and source. 

A finance search may need the latest filing, a specific quarter, or an official source in addition to the closest semantic match. In e-commerce, the best semantic match for “32×30 cargo pants” may be an out-of-stock product. The system still has to decide whether to hide it, return it with a backorder note, or show it so the user can check later. That product decision is a retrieval decision because it changes which candidates reach the agent.

In a chat interface, those controls are in the tool schema, query planner, or app logic. If an agent runs the search, it needs arguments for the same constraints a human would set with filters, sliders, tabs, and sort menus.

A retrieval system usually needs several controls working together:

ControlWhat it doesExample
Exact matchMatches names, IDs, error strings, quoted phrases, tickers, or product codes.Find EADDRINUSE, Authlib, or a specific SEC accession number.
Semantic matchFinds related content when the wording differs.Find the meeting where the team discussed authentication tradeoffs.
Hard filtersRemoves invalid results before ranking.Limit by tenant, permissions, person, date range, size, or stock status.
SortsOrders candidates by a structured field.Prefer the newest, latest filing, lowest price, highest rating, or recency.
RankingScores candidates based on their likely usefulness for this request.Combine semantic match, exact match, freshness, source quality, and use.
RerankingUses a slower model or scorer on a smaller candidate set.Compare the query against the top 100 candidates before returning 10.

Here, a chunk means a small piece of source content, and a candidate is a chunk returned by the first search step. Ranking is the scoring step that orders those candidates. Context assembly then selects which chunks and structured fields to include in the model prompt.

Better ranking improves precision, which means a larger share of the returned chunks is useful. If the relevant chunks are near the top, the system can pass fewer chunks to the model, use fewer tokens, reduce latency, and expose the model to less noise. If the right chunk is ranked 40th and the context only includes the top 10, the system behaves as if the retrieval missed it.

People and agents use the same basic search path: ask for results, inspect what comes back, and decide what to use. A person can skim ten search results, compare titles, snippets, dates, domains, and URLs, and decide whether the result set looks right. They can open the third result, ignore the rest, and search again with a better query. An agent usually receives a bounded set of returned documents and reasons from the context. If the right source falls below the cutoff, the agent may answer from partial context. To avoid this, the system has to retrieve more candidates, run more searches, or pass more evidence into the model.

A missed document can change what the agent searches for next. Suppose someone asks why the team chose Authlib. If the first search misses the transcript where the team compared Authlib with alternatives, the agent may search the codebase instead. It finds imports, callback handlers, tests, and maybe a comment. Then it asks follow-up questions about OAuth configuration. The context starts to look complete, but it supports the wrong answer. It explains how Authlib was used and why it makes sense in the codebase, not why the team chose it.

“The context starts to look complete, but it supports the wrong answer.”

But ranking cannot repair every search problem. If the agent failed to request the latest filing, a reranker may faithfully select an older document with a closer wording match. If the tool has no date_range, person, source_type, size, or in_stock argument, the model has to impose hard constraints in the search text and hope that retrieval infers them. A hard filter gives the model less to infer, making semantic search more reliable.

Scale changes the retrieval problem

Search systems already have tools for this: indexes, filters, facets, sorts, caching, bounded reranking, and freshness jobs. Agent systems need the same discipline.

A human might search, adjust a date filter, scan the first page, then search again. One agent request can do that many times in seconds: rewrite the query, run keyword and semantic search, inspect thin results, issue follow-up searches, fetch sources for citations, and ask for more context before answering. With many concurrent users or agents, the retrieval layer can become a bottleneck.

Humans often wait through a slow search if the result is good. Agent systems often turn slow and uncertain search into more work. When ranking is weak, teams compensate by raising top_k, running keyword and semantic searches in parallel, adding reranking, fetching more source documents, and passing larger evidence bundles to the model. That can improve answers, but it moves the cost into tokens, latency, and retrieval load. A better ranking lets the system return fewer, better candidates, rather than making every request carry a larger pile of possible evidence.

With a small corpus, you can still search comprehensively quickly and cheaply, even with fully agentic approaches. That’s what I recommend when you’re starting and don’t have much data. Don’t add complexity until you need it. But with millions or billions of chunks, every extra retrieval call, candidate, ranking pass, and returned token adds up quickly.

Multi-stage retrieval is the production shape

Most production systems should split retrieval into stages, even when the UI is a chat box.

StageWhat happensTrace question
Search argument constructionThe app or agent turns the request and state into a query, filters, and sort.Did it ask for the right content with the right constraints?
Candidate generationThe system finds plausible chunks from text, vectors, or structured data.Did the right source enter the candidate set?
FilteringPermissions and product constraints narrow what can be returned.Was the source correctly excluded or wrongly lost?
SortingStructured fields order results when order matters.Was the latest, cheapest, highest-rated, or current item surfaced?
RankingThe system scores the candidates based on their usefulness for this request.Was the source present but ranked too low?
Summary returnThe system returns only the fields the agent needs.Did the app receive usable evidence and provenance?
Context assemblyThe app selects, formats, and budgets evidence for the model.Did useful evidence get dropped before generation?
EvaluationHumans or automated checks label whether the retrieval path worked.Can the team turn the failure into a specific fix?

Each stage leaves a different repair path. If the agent chose the wrong filters, changing the embedding model will not help. If the latest document was available but the tool never sorted by date, the fix belongs in the search arguments or retrieval API. If the right source was present but below the cutoff, the fix belongs in the ranking. If the right source came back but was dropped before generation, the bug is in context assembly.

As retrieval becomes a core part of agent architecture, teams increasingly need infrastructure that can combine semantic search, exact matching, filtering, ranking, and large-scale retrieval in a single system. Depending on requirements, this may involve search and retrieval platforms such as Vespa, Elastic, or Coveo, each of which supports different approaches to ranking, retrieval, and operational scale. 

The important point is not the specific technology choice, but recognizing that retrieval quality has become a first-class engineering concern. As agent workloads grow, retrieval systems are increasingly determining the accuracy, cost, latency, and reliability of the overall application.

The post Why retrieval quality is becoming the defining challenge in AI agent architecture appeared first on The New Stack.

OpenAI, Microsoft & Anthropic agree on who runs the agent. They disagree on what you can take back.

Colorful digital static resembling TV signal noise, evoking uncertainty over how AI agents like ChatGPT Work and Claude Cowork manage control, state, and data.

OpenAI announced ChatGPT Work on July 9 and began rolling it out to Pro, Enterprise, and Edu users. It runs on the new GPT-5.6, opens a user’s local files, edits Google Workspace and Microsoft 365 documents, and carries a multi-step task through to a finished deliverable.

Reuters placed it directly against Anthropic’s Claude Cowork, and both target the same person — a non-coder who wants the power of a coding agent without the terminal. Counting Anthropic, Microsoft, Perplexity, and Amazon, five leading labs have now released an agent of this kind. The batch shows that the newest agents are organized more by their intended users than by their functions.

Based on the target user persona, four archetypes appear: the knowledge worker, the power user who self-hosts, the developer, and the enterprise. The personas often overlap since one individual can embody all three roles. Additionally, a product like Claude Code caters to both solo developers and platform teams.

Therefore, consider these deployment archetypes categorized by the main buyer, rather than strict separations. The archetype only represents what marketing promotes. Behind the scenes, each lab has almost consistently decided who owns the runtime, persists memory, manages credentials, and enforces policy.

Four archetypes, based on the user persona

Let’s analyze each of the four archetypes individually, as each differs in the level of control available to users.

The first archetype serves the knowledge worker. A vendor operates the runtime and sells the agent as a delegation to someone who lives in documents rather than code.

ChatGPT Work is the newest, alongside Claude Cowork, which now runs cloud sessions on web and mobile while keeping local-file access on the desktop; Microsoft’s Copilot Cowork, a cloud-hosted agent that executes long-running tasks inside the Microsoft 365 trust boundary; Perplexity Computer, which works across local files and Microsoft apps; and Amazon Quick, the successor to Q Business as that product closes to new customers at the end of July. The user grants access and supervises the result. In most cases, the vendor manages the runtime and persisted state, except for Perplexity’s local option.

A second archetype belongs to the power user who self-hosts. The provider controls the persistent agent process and chooses where to store the state and credentials, often on a Mac mini that has become a piece of personal infrastructure in its own right.

OpenClaw and Hermes are the reference examples, open source, and run on the operator’s own machine. Self-hosting involves managing the control plane rather than full local custody, since both options still allow access to a hosted model and the storage of credentials for external services.

Related reads:

“Microsoft has proved it can survive major changes in the tides of technology… Today, it faces another evolution in one of its core cash cows, as late-stage unicorns and AI labs alike push deeper into Office territory.”

→  Read more in Cautious Optimism

The developer gets the third archetype, whose runtime spans the IDE, the terminal, the repository, and a cloud sandbox. Claude Code, OpenAI Codex, GitHub Copilot in agent mode, and the open-source OpenCode all live here, and Amazon’s developer agent is folding into its Kiro tool. Coding-agent execution is extending from the local IDE to vendor-managed sandboxes and asynchronous cloud workers, making this the most challenging archetype to categorize clearly.

The fourth archetype is built for enterprise workflows and integration with business processes. They run an open agent framework, such as LangGraph or CrewAI, on a managed, governed runtime. ADK on the Gemini Enterprise Agent Platform, Strands on the Bedrock AgentCore, Microsoft Agent Framework on Foundry Agent Service, and Claude Managed Agents belong to this category. OpenAI’s Agents SDK can be hosted on some of these runtimes, including AgentCore, which AWS lists as one of its supported frameworks. The vendor operates the infrastructure, and the customer configures identity, policy, and retention on top of it.

PersonaRepresentative productsRuntime ownershipState and credentialsPlatform type
Knowledge workerChatGPT Work, Claude Cowork, Copilot Cowork, Perplexity Computer, Amazon QuickVendor cloud for most, with per-folder local access on someVendor persists session state, user grants scoped credentialsPackaged experience
Power user, self-hostOpenClaw, HermesThe operator’s own machineOperator chooses where state and tokens live, though inference is often externalPackaged experience, self-operated
DeveloperClaude Code, OpenAI Codex, GitHub Copilot, OpenCodeSplit across the IDE, the laptop, and a cloud sandboxRepo and local for now, drifting into hosted sandboxesSwing, moving toward platform
Enterprise, workflow-drivenADK on Agent Platform, Strands on AgentCore, MAF on Foundry, Claude Managed AgentsManaged vendor runtime, customer-configurableCustomer defines identity, policy, and retention; platform brokersProgrammable platform

The line below the personas

The persona-based approach abstracts four things: where execution runs, where state is persisted, how authority is delegated, and where policy is enforced.

Anthropic describes its design as decoupling the brain from the hands. The harness that calls Claude runs separately from the sandbox where code executes, and a session, an append-only log of every model call, tool call, and result, connects the two. Because the sandbox is kept separate from the brain, the agent can start reasoning before any container exists, and the code it runs remains far from the developer’s credentials.

The same four planes show up at the other vendors. AgentCore Runtime gives each session a dedicated microVM with an isolated CPU, memory, and filesystem, and meters compute usage. Google can route governed traffic through its Agent Gateway, where Model Armor policies inspect configured ingress and egress flows, while Agent Identity and an Agent Registry track the fleet. Microsoft assigns each hosted agent a dedicated Entra Agent ID and runs it in a per-session sandbox whose filesystem survives idle periods.

Deploying agents on a managed runtime is more like leasing a workshop than purchasing a tool… a long-term tenant installs their own locks, maintains their records, and takes their tools when the lease ends.

Deploying agents on a managed runtime is more like leasing a workshop than purchasing a tool. The landlord manages the building and supplies the power, but a long-term tenant installs their own locks, maintains their records, and takes their tools when the lease ends. The personas often conceal this difference. Currently, the vendor typically operates the runtime, so the key question is how much control the customer can still exert and what they can take away across the four planes.

Where the line falls

What differentiates each offering is not the compute operator, since vendors handle nearly all of it. It is how much of those four planes a product leaves the customer to configure and export. A packaged experience hands nearly all four to the vendor and returns supervision and a finished outcome. A programmable platform operates the infrastructure but lets the customer define identity, policy, and retention and move the code elsewhere.

Copilot Cowork shows that the two axes are separate. It is a packaged knowledge-worker experience, yet it runs on a governed enterprise platform and inherits Microsoft’s identity, compliance, and audit controls.

The persona sells the product, but the four planes decide the lock-in.

A product can be packaged on the surface and programmable underneath, which is why the personas and the control planes have to be read as different questions. ChatGPT Work makes the same point from the developer side, since OpenAI’s new desktop app folds Chat, Work, and Codex into a single surface, though OpenAI has not detailed how far the runtime or credential store are shared beneath it. The persona sells the product, but the four planes decide the lock-in.

UsecaseAgent TypeTradeoff
Delegate a knowledge task to an agent you superviseA knowledge-worker agent such as ChatGPT Work, Copilot Cowork, or Amazon QuickThe vendor typically operates the runtime and persists state, and you configure little below the surface beyond access and approval
Keep state and credentials on hardware you controlA self-hosted agent such as OpenClaw or Hermes, in a local configurationYou control the persistent process, though model inference and some tools may still be remote
Ship code changes across the IDE, repo, and CIA developer coding agent such as Claude Code, Codex, or CopilotExecution spans your tools and a cloud sandbox, so ownership is split and worth mapping before you commit
Run many governed agents with audit and identityAn enterprise runtime platform such as AgentCore, Agent Platform, Foundry, or Managed AgentsThe vendor operates the infrastructure while you define identity, policy, and retention, in exchange for coupling workflow logic to one cloud

Real deployments combine the rows rather than picking one. Teams on Foundry Agent Service commonly run open-source orchestration, such as LangGraph, for agent logic while leaning on the platform for governed execution, and Microsoft’s own hosted runtime now supports long-running personal agents like OpenClaw and Hermes with durable state. The boundary between experience and platform is not a thick, well-defined boundary, but a thin line a single system can cross, bridging rival camps.

The agent market is not being decided by open against closed. Open frameworks like Strands, ADK, and Microsoft Agent Framework are precisely what the governed runtimes are built to host.

The agent market is not being decided by open against closed. Open frameworks like Strands, ADK, and Microsoft Agent Framework are precisely what the governed runtimes are built to host. Vendors now manage the runtime across nearly all archetypes, shifting the competition to who can most effectively configure and export state, identity, and underlying policy.

If coding agents enter managed sandboxes alongside knowledge-worker agents, the developer archetype will be established on the platform side. The map will then transform into what the vendors are already outlining. When enterprise teams assess agents, their most important question should be how much of execution, state, identity, and policy they can configure and extract from the product. They should not focus on which persona it presents or whether its framework is open, as the framework no longer determines the product’s value or locking-in capabilities.

The post OpenAI, Microsoft & Anthropic agree on who runs the agent. They disagree on what you can take back. appeared first on The New Stack.

CAR T Revolutionized How We Treat Blood Cancers. Now It’s Closing In on Solid Tumors.

10 July 2026 at 14:00

Separate teams discovered the same target in solid cancers, enabling a powerful two-pronged attack on both tumors and the cells shielding them.

Cancer researchers just found a new way to take on tumors.

CAR T cell therapy revolutionized blood cancer treatment by supercharging a patient’s own immune cells to hunt down cancers. But the approach has struggled in solid cancers. These are some of our top killers—breast, lung, prostate. Roughly two million Americans are expected to be diagnosed with cancer in 2026, and over 600,000 will likely succumb to the disease.

Unlike blood cancers, solid tumors rarely share a single, universal target for CAR T cells. Even cells within the same tumor are a mishmash. Some have little or none of a target protein, allowing them to evade the engineered immune cells, survive treatment, and fuel relapse.

“Target discovery remains a considerable challenge in the development and translation of

CAR T cell therapies for solid tumors,” wrote Christopher Mount and Marcela Maus at the Massachusetts General Brigham Cancer Institute.

Now, two independent teams have converged on the same promising target: A cell-surface protein called GPNMB. In one study, CAR T cells engineered to recognize GPNMB rapidly destroyed glioblastoma—a lethal brain cancer—in tissues taken from patients and shrank tumors in mice.

A second team used a similar strategy against an aggressive soft tissue cancer to fight tumors in organoids and mice. In an early clinical trial involving a single participant, one infusion stabilized the disease for three months without serious side effects.

CAR T designers are often wary of broadly shared targets because they can trigger dangerous attacks on healthy tissue. But GPNMB is an odd duck. In addition to cancer cells, it also sits on immune cells that spur cancer growth or suppress the body’s innate ability to get rid of tumors.

“Our approach attacks both the tumor and the environment that allows it to thrive,” said Sheila Singh at McMaster, who led the glioblastoma study, in a press release. “We’re going beyond targeting the cancer alone and eliminating the immune cells that help shield it from treatment.”

Cancer Fortress

Solid cancers have plenty of tricks to outsmart CAR T cells.

Researchers make these supercharged immune cells  by extracting a patient’s own T cells and genetically engineering them to produce protein “claws” that latch onto a specific cancer target. After infusing the cells back into the body, they seek and destroy tumor cells. CAR T has transformed treatment for several blood cancers and is showing promise in autoimmune diseases and excessive heart and kidney scarring. To simplify the procedure, researchers are also exploring ways to directly transform T cells inside the body with gene therapy.

Solid cancers, however, are far tougher opponents. Unlike blood cancers, which are heavily coated with a shared target called an antigen, solid tumors are molecular patchworks. Cells within the same tumor can display different targets—or none at all—allowing some to evade a CAR T attack and trigger relapse. Many of these targets also appear on healthy tissues, raising the risk of dangerous side effects. And then there’s the tumor microenvironment: A toxic, glue-like “fortress” that hijacks immune cells and uses them to battle incoming CAR T cells.

These barriers aren’t impenetrable. Previous work enlisted  bacteria to help CAR T cells burrow into tumors. Other efforts engineered ultra-sensitive CAR T cells capable of detecting tiny amounts of a cancer target shared across multiple solid tumors.

“Recent reports of activity in several clinical trials reinforce optimism that these efforts may result in true clinical benefit,” wrote Mount and Maus, who were not involved in either study.

But these strategies require additional engineering steps, increasing complexity and cost. And most still leave one major roadblock intact: The tumor’s immune defenses.

One-Two Punch

In the glioblastoma study, the team at McMaster University scoured donated tumors for proteins that distinguished the most aggressive cancer cells. They found one standout: GPNMB. Another test of every protein dotting the cell surface confirmed it as a promising target. The protein is evident across a cancer cell’s membrane, making it readily accessible to CAR T cells.

In lab tests, CAR T cells engineered against GPNMB performed well, nearly eliminating tumors grown from patient samples and extending survival in mice.

The target turned out to be far more valuable than expected. The team soon realized that GPNMB also marked the immune cells that suppress anti-cancer drugs. CAR T cells attacked both fronts simultaneously, weakening the tumor’s immune shield and killing the cancer itself.

“Most approaches have focused on killing cancer cells alone,” said study author Shan Grewal. “Our work suggests we may also need to dismantle the immune support system that helps the tumor survive.”

The second team focused on alveolar soft-part sarcoma, a rare soft-tissue cancer that often spreads to the lungs, brain, and bones before it’s diagnosed. Treatment often comes too late.

The disease is driven by a type of “fusion” gene created when pieces of genetic material are accidentally stitched together. These genes are extremely tough to target directly. Instead, the team screened all surface proteins on the cancer cells and again landed on GPNMB as a top candidate for intervention. The protein’s levels closely tracked the activity of the fusion gene.

CAR T cells targeting GPNMB cleared tumors and prevented metastasis in mice. But because an earlier antibody drug against the protein caused severe skin toxicity in patients, the team also tested their CAR T cells in mice carrying small human skin grafts. Although inflammation initially flared, there were no signs of ongoing skin damage.

Encouraged, the team treated a patient with relapsed, metastasized alveolar soft-part sarcoma. After a single infusion, the engineered cells rapidly divided in the bloodstream and remained detectable for roughly a month. The treatment didn’t trigger skin rashes or more dangerous side effects, like cytokine release syndrome where the body mounts a hyperactive immune defense that harms healthy organs.

The treatment’s benefits outlasted the engineered cells themselves. For roughly three months, imaging tests found fewer of the small, round spots on the patient’s lungs that often signal metastatic cancer, suggesting the disease had stabilized.

A final analysis identified another roadblock: Clusters of cells that suppress the immune system and could blunt the benefits. Adding drugs to block these immune molecules boosted tumor killing in mice. Because the same kind of gene fusion drives other cancers, including kidney, the CAR T cells could have reach beyond this specific type of sarcoma.

Together, the studies underscore that the best CAR T targets might extend beyond cancer cells to expose and attack cancer’s immune cell supporters too. Finding a viable target is a delicate balancing act. Chosen well, and CAR T cells could tackle multiple drivers for cancer growth. Choose poorly, and healthy tissues could get hurt in the crossfire.

Even so, “these two studies indicate that GPNMB represents an actionable target for CAR T cell therapies in several solid tumors,” wrote Mount and Maus.

The post CAR T Revolutionized How We Treat Blood Cancers. Now It’s Closing In on Solid Tumors. appeared first on SingularityHub.

Accelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent Toolkit

10 July 2026 at 13:00
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein...

Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design. Increasingly, they’re driven end-to-end by AI agents. For an agent to run that pipeline well, every step needs to be fast and scalable: Multiple Sequence Alignment (MSA) generation, co-folding inference, serving, and multi-GPU scale-out.

Source

Ground Robots Inherit the Kill Zone

10 July 2026 at 11:00


Borys Drozhak has a vision: a front line almost free of humans, patrolled by flying drones and ground robots, and continuously monitored by AI-controlled sensor networks. And it’s not a pipe dream. Ukrainian roboticists have made major strides in that direction over the past four years. Remotely controlled ground vehicles fitted with machine guns and grenade launchers now patrol the no-man’s-land straddling the front, part of a robotic legion that has stymied Russia’s territorial ambitions so far this year.

Drozhak is a co-founder and CEO of RoverTech, which manufactures the Zmiy, one of Ukraine’s most successful ground robots. Zmiy, Ukrainian for snake, is an 800-kilogram (1,700-pound) rover, 2.15 by 1.5 meters in size, with 75-centimeter diameter wheels. The Zmiy comes in various configurations—for demining, logistics, fighting fires, firing a machine gun, or launching grenades.

According to Drozhak, the uncrewed ground vehicle (UGV) is a record-breaker among Ukrainian ground robots. It’s engineered to be nearly noiseless and emit as little heat as possible, helping it to elude Russia’s intelligence, surveillance, and reconnaissance (ISR) drones. As a result, a Zmiy rover completes on average 57 missions across the kill zone before being destroyed. The kill zone is the roughly 35-kilometer-wide swath of land that straddles the front line; its width is variable and determined mainly by the growing range of the drones.

“Usually, a UGV on the battlefield lasts about seven missions,” Drozhak says. “The Zmiy is quite a bit bigger and stronger” in comparison with most other UGVs, “and can make it back even if two of its wheels get destroyed.”

Drozhak is a software engineer turned roboticist whose story is echoed everywhere in the Ukrainian defense establishment. Before the Russian invasion, he was living a quiet life in Ireland, working for an international software development firm. He returned home shortly after the war began to help defend his homeland. Together with his friend, Vasyl Korenovskyi, who had been a mining engineer, he founded RoverTech with the goal of building robots to perform some of the most dangerous tasks in the war zone. In 2023, they rolled out their first product—the Zmiy de-miner. Earlier this year, one of RoverTech’s assault UGVs was part of a widely reported operation that forced a group of Russian soldiers to surrender without the presence of any Ukrainian troops. Such feats, Drozhak insists, are not rare on Ukrainian battlefields these days.

UGVs are the latest chapter in the military-technology race spurred by the war in Ukraine. Scores of Ukrainian startups have developed dozens of different small ground robots, each with typically multiple variants, over the past three years. They’re mostly replacing human-driven tanks and other military vehicles that used to crisscross the war zone. These remotely controlled robotic vehicles cost a few tens of thousands of dollars apiece compared to millions for a traditional tank, and they can be tweaked and modified in frontline workshops to serve the most urgent needs.

Zelenskyy Orders Up 50,000 More UGVs

In April, Ukraine’s President Volodymyr Zelenskyy signed an order for the government to procure 50,000 UGVs for Ukraine’s military forces by the end of 2026. That’s more than three times as many as the government purchased in 2025 and a massive increase from the 2,000 procured in 2024, according to defense analyst Marc C. Lange.

The rise of UGVs, Lange explains, is a direct response to the warfighting revolution ushered in by the speedy evolution of uncrewed aerial vehicles that came to define the war in Ukraine.

As the number of drones zooming above the front line rose and their range increased, the battlefield became completely transparent. Today, anything that enters the kill zone gets hit by a first-person view (FPV) kamikaze drone within minutes.

“Any armored formation, any resupply and logistics vehicle, and any manned formation anywhere near the edge of the battle area has between seconds to a low amount of minutes before it gets turned to dust,” Lange says. “The Ukrainians were losing drivers. Traditional methods of evacuating injured soldiers became impossible. That space is basically unsurvivable.”

Ukraine, suffering from a shortage of infantry, has taken that problem more seriously than Russia, which has a larger pool of fresh recruits to draw from. UGVs began ferrying supplies to troops at frontline positions in 2024. Gradually, they took over the complex and risky evacuations of the wounded, using special enclosures to protect the soldier being transported. But this year, Lange says, is “the year of the assault UGV.”

Emerging Ukrainian tactics combine UGVs with real-time reconnaissance and surveillance from aerial drones, which discover enemy troops, often under cover of night. The reconnaissance data are then used by remote operators who guide UGVs as they stalk, corner, and shoot to kill. Oleg Fedoryshyn, the head of research and design at DevDroid, another prominent Ukrainian UGV developer, said the ground robots can be controlled from as far as 100 kilometers away using Starlink connectivity, LTE networks, or mesh-networked military radio systems. The UGVs can also carry strike UAVs (uncrewed aerial vehicles), serve as communication relays for drones, or carry and launch communication relay drones that further extend the range of the attack vehicles. The UGV can lurk in position for up to one week without needing a battery charge, Fedoryshyn said, and wait for the enemy to move closer.

“It’s better than to put people there,” he notes. “A guy with a machine gun is always the first target for the enemy.”

An Ukrainian soldier adjusting an unmanned ground vehicle\u2019s machine gun. The Droid TW 12.7, by DevDroid, is shown here outfitted with a 0.50-caliber M2 Browning machine gun that can be aimed and fired by a remote operator using a tablet and an encrypted communications link.DevDroid

Fedoryshyn estimates that UGVs could eventually help cut the number of soldiers needed along the front line by 30 to 40 percent. Drozhak is even more ambitious. He envisions a future front line that’s entirely automated, relying on sensors and other systems that are only occasionally serviced by humans.

A guy with a machine gun is always the first target for the enemy.

“Right now, we need a lot of UGVs because there are people on the front line and we need to deliver supplies to them,” he says. “But we can substitute many of them with sensor systems, servicing robots, and UGVs, and then we will not need that many for logistics. At some point, we could have only robots in the kill zone.”

Ukraine, with a prewar population of around 41 million, has lost over 150,000 fighters in the war since 2022, according to estimates by the Center for Strategic and International Studies and others. Many thousands of others have been mutilated or permanently disabled. Even those who return without physical injuries suffer lasting psychological trauma. Drozhak dreams that a future robot army would put an end to the ability of autocratic regimes worldwide to brutalize their neighbors.

“There will be no need to push people on the battlefield anymore,” says Drozhak, the RoverTech CEO. “Once we achieve that in Ukraine, any country with a decent economy would be able to defend themselves just with technology.”

RoverTech’s Tarantula active-protection system, which uses acoustic and visual sensors combined with AI algorithms to detect approaching killer drones, is the first step in that direction, he declares.

“The future battlefield will rely on networks of robotic sensors and autonomous systems that can continuously monitor dangerous areas, provide early warning, and reduce the need for soldiers to expose themselves to direct threats,” he says. “Human operators will remain responsible for critical decisions, but increasingly advanced sensing technologies will help move people away from the most dangerous positions on the battlefield.”

Why UGVs Are Vulnerable

Militaries around the world were looking at UGVs prior to Russia’s 2022 invasion of Ukraine. But those were quite different, explains Samuel Bendett, a defense analyst at the consultancy CNA. They were larger, more complex, and conceived to operate in smaller numbers. The more compact forms now seen in Ukraine are the result of an evolution that paralleled that of the first-person view (FPV) attack drones. Both needed to be cheap as they don’t last long and small to be less conspicuous. Now, the West is trying to understand the overall role of UGVs in future warfare. So far, in Bendett’s view, the impact of UGVs on warfare isn’t as profound as that of the FPVs and other aerial drones.

“Not every terrain would be applicable to using a UGV,” Bendett explains. “So far, a lot fewer countries are seeking to integrate them into their combat operations than UAVs, which very much democratized the way of enabling short-range to mid-range strikes against adversaries.”

UGVs, he points out, are much more susceptible to communication disruptions than UAVs, while being less suitable for autonomous operations and swarming due to the complexity of ground terrain.

“With UAVs, communication is much easier,” according to Bendett. “There are no interferences between the ground station and the UAV save the distance, Earth’s curvature, and the radio horizon. But on Earth, there’s lots of different obstacles that interfere with radio signals.”

Most UGVs rely on Starlink as the first choice for operator control, but even that comes with problems. Starlink signals are easily disrupted by trees and buildings. And Russia, having been cut off from Starlink, is working hard to find ways to jam the system.

On top of that, Lange says, as UAV autonomy progresses, UGVs could be left behind. The reason is that UGVs are likely to remain dependent on operator communication links for some time and will therefore be vulnerable to enemy UAVs that can’t be stopped by jamming systems that still provide some protection today.

“The low production cost of strike drones will mean that UGVs will have to endure a barrage of strikes that might be too much,” Lange says. “The question is whether you can make UGVs more survivable on the front line both in terms of command and control and the actual survivability of that many strikes.”

Still, he thinks there’s “no path back from UGVs.” The idea of distributing a whole range of tasks, performed in the past by a single large and expensive tank, to a fleet of small, cheap UGVs provides more resilience against the omnipresent drones. Moreover, although many international commentators now say that Russia appears to be losing, the war grinds on—and so does the cat-and-mouse game of lethal innovation.

Proposed Information Collection; ATUS Artificial Intelligence (AI) Questions

The Department of Labor, as part of its continuing effort to reduce paperwork and respondent burden, conducts a pre-clearance consultation program to provide the general public and Federal agencies with an opportunity to comment on proposed and/or continuing collections of information in accordance with the Paperwork Reduction Act of 1995. This program helps to ensure that requested data can be provided in the desired format, reporting burden (time and financial resources) is minimized, collection instruments are clearly understood, and the impact of collection requirements on respondents can be properly assessed. The Bureau of Labor Statistics (BLS) is soliciting comments concerning the proposed new collection, the "American Time Use Survey (ATUS) Artificial Intelligence (AI) Questions." A copy of the proposed information collection request can be obtained by contacting the individual listed below in the Addresses section of this notice.
❌