❌

Normal view

SpaceXAI's Grok Bot turns agents into persistent digital coworkers that can operate your apps for $120-per-month

SpaceXAI, the division of SpaceX formerly known as xAI, is launching an early beta version of Grok Bot, a new agent designed to move AI assistants beyond answering prompts and toward continuously executing work across the software employees already use.

The central idea is straightforward: instead of opening an AI assistant whenever a task arises, users create persistent Bots with specific jobs, give them access to applications and websites, and delegate work much as they would to a teammate.

Each Bot operates through its own computer environment, can continue working when the user's laptop is closed, and can return when it needs approval or has finished the assignment.

SpaceXAI says the system began as an internal prototype before spreading across the company, where teams created Bots for sales outbound, marketing campaigns, office operations, bug fixes and other work. The company is now turning that internally developed workflow into a product for external users.

“Bots are AI teammates that do real work for you,” the company said in announcing the product. “They sign in to your tools, use them just like you do, and come back with finished work.”

The company did not release benchmarks for Grok Bot's performance on agentic tasks. And it arrives amid an increasingly crowded marketplace of first-party AI agents that attempt to reliably complete real, enterprise workflows by interfacing with a user's other applications and devices.

Anthropic introduced computer use for Claude in 2024, allowing models to inspect screens and operate interfaces through mouse and keyboard actions, and continued expanding with the launch of the developer focused Claude Code harness in early 2025 and the more non-technical, white collar focused Claude Cowork agent early this year.

Meanwhile, OpenAI gave its Codex harness the ability to control other computer apps in April, launched agentic Workspace Agents that can also connect to third-party applications and use them autonomously, and recently debuted a new ChatGPT Work environment for longer, multi-step tasks and finished deliverables.

Grok Bot seeks to join the party with its own management model for agents: persistent workers with responsibilities, memory, learned routines and the ability to hand work to one another.

Pricing and availability: Grok Bot starts at $120 per seat per month for teams, $200 per month for individuals

Grok Bot is available beginning today, August 11 in beta for SuperGrok Heavy, Cursor Ultra and Cursor Premium Teams subscribers (recall SpaceX acquired Cursor for $60 billion back in June). The product arrives for macOS, Windows, Linux and iOS, with Android listed as coming soon.

According to its product page on xAI.com, Grok Bot is included with Cursor Ultra at $200 per month for individuals. The plan includes a computer for Grok Bot, access to users' tools, scheduled routines, desktop and mobile operation, and extended AI-token limits.

For organizations, Cursor Premium Teams costs $120 per seat per month and adds centralized billing and settings, a team marketplace for skills and plugins, shared usage analytics and SAML/OIDC single sign-on.

Existing SuperGrok Heavy ($300 per month) subscribers also receive access. However, for organizations wishing to sign up today, SpaceXAI is directing them to a waitlist for future access.

Those prices make Grok Bot a substantially different purchasing decision from a low-cost general AI subscription. The economic question for companies will be whether persistent Bots can replace enough manual work or conventional automation infrastructure to justify the per-user cost — and how usage limits affect total cost once agents begin running continuously.

From prompting an AI to managing one

SpaceXAI describes Grok Bot as a team of “always-on agents.” Users can create multiple Bots, assign each a role and let them work simultaneously.

The company provides examples including Sales Outbound, Talent Scout, Paid Media, Expense Manager, Product Performance, Bug Reproduction, Account Health and Chief of Staff. A sales Bot, for example, can research accounts, score prospective contacts, prepare email and LinkedIn outreach in the user's voice, and assemble the results for human approval.

Promotional materials show SpaceXAI using the system internally for substantially longer chains of work. One sales Bot can add call-transcript notes to a CRM and draft follow-up messages. An operations Bot can seat new hires and process invoices arriving through Gmail. An engineering Bot can reproduce a bug in the product interface, file a ticket and then hand the repair to a debugging Bot.

The architecture could make Grok Bot particularly relevant for workflows that span systems that were never designed for AI automation.

Rather than requiring every application to expose an API specifically for an agent, Grok Bot can sign into applications and websites and operate their interfaces. SpaceXAI says Bots have their own computers and can continue working 24/7.

The company explicitly says this includes websites and applications that have “no clean API or MCP,” an important distinction for enterprises with legacy software, fragmented SaaS environments or internal systems that have never been instrumented for agent access. Instead of limiting automation to formally integrated services, Grok Bot is designed to work through the same software interfaces a human employee would use.

The company says early users are already applying Bots to jobs including vendor negotiations, e-commerce customer support and continuously updating CRM systems.

Another feature attempts to reduce the engineering required to automate repeatable business processes. Users can demonstrate a workflow while a Bot follows along. Grok Bot can then save the process as a routine and execute it later without requiring the user to reproduce every instruction.

SpaceXAI says the Bot can also incorporate corrections into those learned routines, allowing the workflow to change as the user teaches it how a particular process should be handled.

That potentially changes the deployment model from explicitly programming an automation to teaching an agent how an employee performs the job.

The company is also claiming a more persistent form of behavioral memory than simply retaining a chat transcript. According to the launch announcement, Bots remember prior conversations, learn preferences such as a user's writing voice and edge cases, and gradually learn when they should interrupt for approval versus continue independently. SpaceXAI says they can later resume dropped threads, nudge stalled handoffs and pick up work from earlier conversations.

It further says Bots can become proactive over time, sometimes identifying work before the user explicitly asks for it. That is a more ambitious claim than conventional scheduled automation and will put additional pressure on permission controls and escalation rules if the system is deployed against production applications.

Bots can delegate work to other Bots

Grok Bot also supports multiple agents operating together.

Users can place several Bots into the same thread, where the agents can pass work between one another. The company's demonstration includes specialized Research, Communications, Chief of Staff and Travel Bots coordinating tasks.

SpaceXAI says those Bots can independently message one another and share context within threads. Users can also put multiple Bots into a group conversation where they assign ownership, transfer work and coordinate among themselves, bringing the human back in primarily for judgment calls.

Internally, the company says employees sometimes place a Chief of Staff Bot above specialist Bots responsible for functions such as inbox management, recruiting, expenses, operations and bug fixes. That makes the product's orchestration model more explicit: the user does not necessarily have to serve as the routing layer between every specialized agent.

Initial reactions are extremely positive

Lenny Rachitsky, host of the popular vlog and podcast Lenny's Podcast and author of newsletter Lenny Letter, received early access to Grok Bot and loved using it. As Rachitsy wrote on X : "I haven't been this excited about a new AI product in a while. It's like OpenClaw, but super easy, reliable, and less scary to use. I think this will be a huge new product line for Cursor/Grok/SpaceX."

Similarly Matt Shumer, an AI entrepreneur who said he tested Grok Bot for several weeks before launch, highlighted this orchestration as one of the product's strongest features.

“The best way I can describe it is an agent for everything, not just code,” Shumer wrote on X.

In one test, Shumer said he created separate researcher and writer Bots, then created a Chief of Staff Bot and instructed it to coordinate the other two on a project. He expected the workflow to break down.

“It worked out of the box,” he wrote.

His main criticism involved model selection.

Unlike systems where developers or advanced users explicitly select the underlying model, Shumer said Grok Bot automatically routes tasks to models on the backend.

“You don’t choose a model for your Grok Bot,” he wrote. “It’s all done automatically on the backend.”

Shumer said the model router “wasn’t great” during his testing, although he said he was subsequently told it had improved.

SpaceXAI's expanded announcement still does not identify which underlying models the router uses, nor does it document a mechanism for users to select, pin or switch to a particular xAI or third-party model. As a result, the model layer remains largely abstracted from users in the publicly supplied launch material.

That abstraction represents an important tradeoff for enterprise deployments. Automatic routing can remove a significant configuration decision for ordinary employees, but advanced users may want explicit control over model cost, latency, reliability and behavior — particularly for repeatable production workflows.

The agent market is moving toward longer-running work

Grok Bot enters a market increasingly focused on agents that can do more than generate text or code.

Anthropic's computer-use capability established a mechanism for Claude models to interact with software through screenshots, cursor movements, clicks and typing. Its broader Claude product also connects with workplace services and remote MCP servers.

OpenAI, meanwhile, now describes ChatGPT Work as an agent for “longer, multi-step work and finished deliverables,” while keeping Codex focused specifically on software development. OpenAI's enterprise agent economics can also incorporate usage-based credits, making task complexity and token consumption part of deployment cost calculations.

Grok Bot's differentiation is therefore less about proving that AI can operate software than packaging computer use, persistence, workflow learning and multi-agent coordination into something resembling a workforce interface.

SpaceXAI's announcement sharpens that distinction by emphasizing completion rather than assistance. One company product employee, identified only as Roman, describes the difference as closing the gap between work that is nearly finished and work actually completed inside the destination application: “Grok Bot can finish the swing, because the work lands where a human would put it, in the actual tool.”

That distinction will ultimately depend on reliability. A chatbot producing a bad answer creates a correction problem. An autonomous agent operating CRM records, support queues, vendor conversations or other production systems can create an operational problem.

Grok Bot's success will therefore depend not only on model intelligence, but also on permissions, predictable execution, escalation behavior, memory accuracy and how reliably agents recognize when human approval is necessary.

That challenge becomes more significant if Bots act proactively, resume forgotten work and coordinate with one another without the user serving as an intermediary. Those capabilities reduce the amount of supervision required when they work correctly, but they also expand the consequences of an incorrect assumption, stale context or improperly scoped permission.

The interface may matter as much as the models

Shumer described the product's interface as feeling like iMessage, an intentionally familiar metaphor for a system whose underlying architecture — autonomous computers, persistent memory, agent orchestration and automatic model routing — could otherwise be difficult for nontechnical users to configure.

SpaceXAI makes essentially the same usability argument in its launch announcement. Rather than asking users to construct workflows before getting started, it says users can simply message a Bot from a phone or desktop, hand it work and later continue the same conversation from either device.

That simplicity is part of the product strategy. Grok Bot is trying to hide much of the conventional machinery of automation — workflow builders, explicit integrations, agent routing and orchestration — behind an interaction model that resembles messaging a coworker.

That may prove to be the larger bet behind Grok Bot.

The AI industry has spent several years making models increasingly capable of using tools and completing multi-step tasks. Grok Bot attempts to turn those capabilities into an organizational abstraction people already understand: give someone a job, teach them how you work, and let them coordinate with the rest of the team.

If that abstraction proves reliable, the enterprise agent competition may increasingly shift away from which assistant produces the best individual response and toward which platform can most reliably manage fleets of agents performing ongoing work.

🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery

11 August 2026 at 21:03

This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at the heart despite being all of 2 years old.

The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story!

Editor’s note: not to be confused with Chai AI, which was another top pod of ours.

Pharma suddenly doing big AI tools deals

For the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing.

Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks1) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios.

The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust.

Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed.

Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this.

RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much.

Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that?

Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic.

This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition.

Photoshop for molecules2

With that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts.

What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with.

— Neil Patil, (Chai product lead)

This means better UX, such as a molecule editor that is more like a CAD or graphics design program than a chatbot.

Their approach has paid off: since June, Chai has announced three more major deals: Lilly, Novartis, argenx, plus an expansion of their Eli Lily program. This episode is too full of quotable moments for a short blog, so tune in to learn about

  • Why protein tokens have the highest downstream value of any token

  • Climbing levels of abstraction as models improve

  • How Pharma, VC, and research are all just portfolio optimization

  • How better tech changes the whole portfolio

  • How relentless focus on simplicity leads to scale

Plus much more!

1

"Biobucks" is deal-value for milestone-heavy licensing agreements — the headline number (e.g., "$1.7B deal") is almost entirely contingent on hitting targets. Typically only 2–5% of the total is upfront; the rest pays out only if the drug clears each gate, and most drugs don't.

2

I actually think SolidWorks is a better analogy, but PhotoShop has better brand recognition ¯\_(ツ)_/¯

💾

Chrome adopts what may be the best protection yet against account takeovers

11 August 2026 at 20:59

Google’s Chrome browser has added a new feature that could go a long way in preventing a form of account takeover that’s grown increasingly common as users adopt two-factor authentication, passkeys, and similar protections.

The new Chrome protection is known as device-bound session credentials (DBSCs). The measure stores a unique encryption key in a silicon-resident fortress that’s built into the device running the browser. On Windows machines, this fortress is called a TPM, short for Trusted Platform Module. On macOS and iOS, it’s known as a secure enclave. Other platforms have differing names. Recently released versions of Chrome for Windows and macOS generate a key that’s stored in this fortress.

An antidote to session cookie theft

DBSCs protect against the theft of session cookies, the unique strings of characters that websites store on browsers. Session cookies greatly speed up browsing on sensitive sites that require user authentication. Instead of requiring the exchange of credentials each time a user opens a new site page, the server sets a session cookie that effectively proves the user has already successfully logged in.

Read full article

Comments

© Getty Images

How I learned to stop worrying and love hyperscaler capex

Amazon data center

The AI boom is an oddly miserable bubble. Despite interesting tech, huge new companies, and products with global reach, AI has attracted legions of detractors.

Some have valid complaints, like seeing their roles automated, or the value of human art being pressured by machine generation. Other complaints have had less staying power.

It was once in vogue to argue that AI companies would run out of data, and thus their models would stop improving. False. Some of the same voices argued that AI lacked a use case and was thus little more than a fancy toy high on its own hype. Incorrect.

Later, the argument shifted to AI being too expensive to use, an incredible flip from the AI has no real use argument. This is being proved false, as low-cost models from China now face both low-cost, closed-source AI models from OpenAI and new, open models from Meta. Agents were too brittle to start; now they are hacking the world. You get the idea.

Lately, I’ve read criticism about the AI boom from a financial perspective. Namely, that the major cloud players (AWS, Google Cloud, Azure) are spending too much money on AI infra. Surely we can’t use all that compute, the argument goes, and thus hyperscalers are torching their nest egg and investor goodwill at the same time.

I wanted to put the contention to the test, so I pulled together data from Amazon, Alphabet, and Microsoft’s cloud groups (here) to peel back the onion a little. Here’s what I found: Growth is accelerating, hyperscaler profitability scales with scale, and hyperscaler capex efficiency is improving.

Continue reading on Cautious Optimism

This is an excerpt from Cautious Optimism, a modestly upbeat publication focused on technology, business, and power. Read more about the concern of hyperscaler cost on Cautious Optimism.

The post How I learned to stop worrying and love hyperscaler capex appeared first on The New Stack.

Anthropic’s watermark survives copy-paste, but not the real dev workflow

A minimalist blue illustration of a hand reaching down to touch a water surface, creating concentric ripples. Beneath the water, a pixelated and distorted reflection of a hand reaches up to meet the finger, symbolizing the connection between a user's experience and the underlying digital infrastructure.

Anthropic announced it will embed invisible watermarks into text generated by new Claude models, including output produced through its API, coding tools and cloud partners. For developers, the mark offers another way to trace where AI-generated text or code might have come from, but it is not strong enough to prove its origin.

Laying out the plan in a support document, Anthropic said Claude models launched in the EU on or after Aug. 2, 2026, will include machine-readable marking from release. The company is working to add support to older models as well.

The marks will apply worldwide across supported Claude products, including the Claude API, Claude Code, Claude Cowork and Claude Tag. Text generated through AWS, Google Cloud or Microsoft Foundry will also carry the watermark when those platforms use a supported model. Because the mark is added at the model level, it follows the output into applications built on top of Claude — the same applications that are already reshaping how enterprises deploy AI infrastructure — although Anthropic cautions that some platforms and features may not support every type of mark.

The change follows the Aug. 2 start of Article 50’s transparency requirements under the EU AI Act, which require providers of generative AI systems to make synthetic output detectable in a machine-readable format. Anthropic signed the accompanying Code of Practice as a provider of both generative AI models and systems. OpenAI, Google, Meta, Microsoft and Mistral are among the other model providers that have committed to the code.

Yet, Anthropic is handling text and files differently. Text gets a watermark hidden in the words themselves, while supported files such as SVGs, PNGs and JPGs receive a digital signature using the C2PA standard. The file metadata can show that Claude processed an asset and whether the metadata has been altered.

“Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing,” Anthropic said.

“Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.”

How token-level watermarking works

Anthropic has not explained how its text watermark works or said whether Claude uses KGW, a semantic version or another method. The company has also not shared any figures showing whether watermarking affects latency or adds to inference costs — a gap that matters for teams already wrestling with the hidden costs of agentic AI workflows.

Alex Cui, CTO and co-founder of AI detection company GPTZero, wrote in a technical explainer on X that watermarking systems fast enough to run on a streaming frontier model often follow the same general approach. One such technique, known as the KGW method, changes the probabilities the model uses when selecting its next token.

A language model normally calculates a probability for every token that could appear next. In a simplified watermarking system, a secret key and the preceding tokens are used to generate a hash, which divides the candidate tokens into two groups, often described as green and red. The model then slightly increases the probability of selecting one of the green tokens.

A detector with the same key can use the preceding text to reconstruct which tokens would have been favored at each position. A passage containing an unusually high number of those choices may carry the watermark.

Cui wrote that more advanced approaches can derive the watermark from the meaning of nearby text rather than an exact sequence of tokens, which may help the signal survive some paraphrasing because replacing a word does not always change the surrounding context.

“Their watermark needs to work token-by-token because they are streaming their text to users,” Cui wrote. “Many watermark methods plan sentences or paragraphs at a time, or change the text after it’s entirely written, in order to make their watermark robust to paraphrasers.”

Anthropic has not confirmed that Claude uses any of these methods, but streaming limits the techniques available because the model must add the signal while generating its response rather than rewriting a completed passage afterward.

Code resists invisible marking

Code presents a different problem because the model has fewer valid choices. Words can often be swapped or sentences rewritten without changing their meaning, but seemingly minor changes can break working code. That challenge intensifies as the AI coding era matures and more production code flows through model-assisted pipelines.

“There are some texts, like code, that cannot be arbitrarily changed; otherwise the code will break,” Cui wrote. “In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms,” such as variable names.

Code may also be difficult to track through a normal development workflow. Anthropic has not published tests showing how well its watermark survives those changes, so teams do not yet know whether a Claude-generated patch will remain detectable after passing through a pull request.

“In my testing, the watermarks don’t survive intense paraphrasing, especially if you combine word choice and syntax attacks.”

Pipelines silently erase watermarks

The same issue comes up when applications change Claude’s output before showing it to a user or committing it to a repository. Summarizing it with another model, translating it, splitting it into smaller sections, turning it into structured data or mixing it with database content could all make the watermark harder to detect.

Anthropic acknowledges this limitation. Editing, paraphrasing, translating or combining the response with other text may weaken or remove the watermark, while short excerpts may not contain enough of the signal to detect.

Cui wrote that determined users can attack a watermark by changing both the vocabulary and the structure of a passage.

“In my testing, the watermarks don’t survive intense paraphrasing, especially if you combine word choice and syntax attacks,” he wrote. Cui added that free paraphrasing tools he tested were able to bypass Google DeepMind’s SynthID text watermark.

Research supports those concerns. The “Watermarks in the Sand” paper found that, under defined assumptions, attackers can remove watermarks without severely damaging the quality of the content. The absence of a watermark does not show that Claude had no role in creating the content. The response may have come from an older model, may be too short to carry a detectable signal or may have been changed somewhere in an application pipeline. It may also have passed through a platform or feature that does not support that type of mark.

Finding a watermark does not prove authorship either. Claude may have proofread, translated or reformatted material written by a person. Anthropic says a detected mark means only that the content “may have been processed by Claude,” not that Claude created the underlying work.

“If Anthropic releases the watermark detector publicly, I think they defeat their own watermark. People find reliable watermark-removal strategies by testing against Anthropic.”

Detection creates new risks

Anthropic plans to give users and third parties a way to detect its marks, but it has not said whether that will take the form of a local tool, a detection API or access limited to selected organizations. A public detector would be easier for developers to add to their applications, but it would also allow someone trying to remove a watermark to keep editing and checking the text until the signal disappears.

“If Anthropic releases the watermark detector publicly, I think they defeat their own watermark,” Cui wrote. “People find reliable watermark-removal strategies by testing against Anthropic.”

The secret keys behind the watermark create another challenge. A leaked key could make it easier to remove the mark or imitate it in text that Claude did not produce. The scenario echoes what happened when provenance attestations were turned into camouflage — a trust signal that was supposed to increase confidence instead became an attack surface.

“To avoid a large blast damage from this, you need to have a couple secret keys in rotation,” Cui wrote.

Key rotation would require detectors to recognize marks created with both current and retired keys, including those embedded in content generated months or years earlier. Anthropic has not explained how it plans to handle that history.

Watermarks aren’t a substitute for real provenance

For developers, Claude’s watermark is best treated as another clue, not a replacement for audit logs or provenance tracking. Applications that need to show where an artifact came from can record the model ID, prompt version, response time and a hash of the original output, then log any changes made before it reaches a user or is committed to a repository.

The watermarking announcement arrives as Anthropic navigates deeper questions about what its models do in the wild. Recent incidents have exposed gaps between lab safety evaluations and real-world containment, and the company has publicly backed calls for the most powerful AI labs to slow down. Watermarking fits into that posture — a transparency mechanism rather than a safety guarantee — but its practical value depends on technical details Anthropic has not yet shared.

Anthropic tells customers to determine how Article 50 applies to their own products and says more technical documentation is coming. But until Anthropic shares those details, teams do not know how they will detect the marks, how key rotation will work or how well the watermark will survive common changes to code and application output.

The post Anthropic’s watermark survives copy-paste, but not the real dev workflow appeared first on The New Stack.

NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation

11 August 2026 at 19:00
Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare, media...

Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare, media processing, and remote operations. A system may capture several cameras, decode network streams, run AI inference or conventional vision processing, draw results, and encode video for storage or delivery. The individual calls are…

Source

Mastering Edge AI on Raspberry Pi with LiteRT and Gemma

11 August 2026 at 16:45
Deploying secure, real-time Edge AI on Raspberry Pi is now simplified using LiteRT and lightweight Gemma open models. LiteRT optimizes CPU and GPU performance, delivering fast token speeds for models like Gemma4, enabling real-time local reasoning for robotics. Developers can quickly convert, quantize, and run these models using the lightweight LiteRT CLI tool. Support for Hailo AI accelerators is also coming very soon.

Why Go is an Ideal Language for AI-Assisted Software Engineering

11 August 2026 at 16:45
As AI coding assistants shift the developer's primary role from writing boilerplate to reviewing and maintaining systems, language choice becomes critical for long-term architectural integrity. Go directly addresses this new paradigm by utilizing its strict compiler, integrated toolchain, and uncompromising readability to provide deterministic guardrails that help AI models self-correct and generate highly standardized code. By enforcing ecosystem-wide consistency and strict backward compatibility, the Go platform empowers engineering teams to efficiently verify, optimize, and maintain high-velocity, AI-generated output in production environments.

Databricks acquires Electric to give every AI agent its own Postgres database

Databricks on Tuesday announced that it’s acquiring Electric, the startup behind the WASM-based Postgres project PGlite and the Electric sync engine, as agentic applications change how developers use databases.

The Electric team will join Neon, the serverless Postgres company Databricks acquired for about $1 billion last year and the foundation of its Lakebase database service.

The companies didn’t disclose the terms of the deal.

What Databricks bought

PGlite is a complete Postgres database in WebAssembly (WASM). It runs in the browser, a Node.js process, or inside the kind of sandboxes agents use to execute code. It supports dynamic extension loading, including pgvector, the preferred Postgres vector extension.

According to the companies, PGlite has grown from 1 million to 13 million weekly downloads over the last year.

The sync engine at the core of Electric

It’s the Electric sync engine that is core to Databrick’s interest in Electric, though. This engine keeps a central Postgres database that can then be synced in near real-time with browser tabs, mobile apps, or agents. As Databricks notes, this is the multiplayer model of Figma or Google Docs, but applied to Postgres and the agents that use it.

The Neon team, in its own announcement, notes that “complex problems like conflict resolution, partial replication, and reconnection logic make real-time sync difficult to build from scratch.” Hence why Databricks likely acquired Electric instead of trying to build this from scratch itself.

As for the future of Electric, the company’s founders James Arthur and Valter Balegas write that “everything we’ve previously open sourced stays open source.” This covers the sync engine, PGlite, Durable Streams, and TanStack DB.

What doesn’t survive the deal, however, is Electric’s hosted service. “Electric Cloud is winding down,” the founders. “Cloud users will need to self-host or move to another provider.”

The deal also extends a string of database acquisitions for Databricks that includes Neon itself and, more recently, the transactional processing startup Mooncake.

A database that lives for 10 seconds

As the Databricks team argues, traditional non-agentic applications share one database among many clients, and that database is the most permanent piece of the stack. But agent workloads change this.

In a recent post on how agentic development changes databases, Databricks’ Ippokratis Pandis, Nikita Shamgunov, and Reynold Xin write that agents now create roughly four times more databases than human users do on Lakebase. They also stress that the average project now carries about 10 database branches, and that some projects run more than 500 branch iterations deep.

For some types of applications on Lakebase, the average database compute is now alive for under 10 seconds.

Agents, as it turns out, like to branch databases the way they branch code, a pattern Neon built its architecture around.

In practice, a coding agent spins up a sandbox, instantiates PGlite inside it, builds and tests against the database, and then either throws the whole thing away or syncs the result with — in the Databricks context — a Lakebase branch. Because Lakebase separates storage from compute and keeps its data in Postgres page formats on object storage, creating that branch is a relatively cheap copy-on-write metadata operation.

“As coding agents drive the cost of creation to zero,” the Neon team writes, “the number of applications explodes, and most of them are small.” A database server, even a serverless one that scales to zero, imposes a floor on what the smallest viable app costs to run. “You can’t have an age of abundance if every app requires a fixed minimum of compute,” the post argues.

‘Two halves of the same idea’

It’s worth noting that PGlite didn’t start at Electric. Instead, it began as an experiment by Neon co-founder Stas Kelvich, who compiled Postgres to WASM to see whether it could run client-side. Electric picked the work up and turned it into a production project. “That repo became the basis of PGlite,” Arthur and Balegas write.

As Databricks’ announcement notes, this now “reunites two halves of the same idea.”

The post Databricks acquires Electric to give every AI agent its own Postgres database appeared first on The New Stack.

A decade of mathematical certainty: Reflections on the Automated Reasoning Group

11 August 2026 at 16:22
In 2016 a small research team at Amazon announced our presence to the world with the launch of the Automated Reasoning Group (ARG). Our vision was bold: use mathematical logic to not just test AWS systems but to prove, with mathematical certainty, that they work correctly. In the intervening decade, we’ve gone from exploring whether advances in formal verification could mitigate previously intractable problems — at AWS scale — to building systems fundamental to how AWS approaches security and reliability. Our group’s production services process billions of queries daily. This is a look back at how we applied cutting-edge formal-verification and program analysis techniques to Amazon's unique challenges. We were motivated by the belief that real-world security and infrastructure problems could be solved with mathematical rigor at AWS scale. After all, tools like the ones we were researching had already proven successful at places like Intel and NASA. What we didn't fully anticipate was just how broadly applicable these techniques would become. From demos to production at scale When we held our first ARG Demo Day in 2016, we showcased several ambitious projects to AWS Security. Each represented a different approach to the same fundamental question: how can we use mathematics to prove that our systems are secure and correct? The answers presented that day laid the groundwork for systems that continue to play significant roles for AWS and our customers in 2026. Sean McLaughlin gave a presentation titled “Automatic tools for reasoning about virtual private clouds (VPC)”. He noted that the growth of VPC networks had led to an increase in the demand for automated-reasoning solutions capable of identifying misconfigurations or security vulnerabilities. His answer was “a tool called Tiros, which, just simply put, answers questions about your network.” Tiros became the foundation of a network security analysis feature in the Amazon Inspector service, which is used by millions of customers building applications in the cloud. Tiros is also used within AWS to automate the checking of compliance certification and adherence to security invariants for many AWS services. Today it powers both Amazon Inspector and Reachability Analyzer. This work also split off and became Zelkova, which uses automated reasoning to analyze policies and the future consequences of policies. Zelkova powers tools such as S3 Block Public Access and IAM Access Analyzer, among many others. Similarly, a presentation given that day on the use of deep automatic analysis for crucial infrastructure turned into the work that we did to prove the correctness of our TLS handshake and other properties of cryptographic, storage, and virtualization code. Even some of the smaller-scale projects we spotlighted went on to have an outsized impact. When we presented our work on deductive verification for high-value infrastructure, it largely applied to the deepest parts of our cryptographic protocols. Today the importance of that work has grown enormously, fueled by the rise of proof assistants — automated tools such as Lean (created by Leo de Moura, a senior principal scientist on our AR team) that help users develop formal proofs. We can now pair those tools with language models to find proofs for more and much bigger systems. In fact, the proof we announced for the Nitro Confidentiality Engine is evidence of this. It is also the basis of our proof of the AWS policy interpreter and more recent work proving the correctness of our cryptographic foundations. Mathematical guarantees customers rely on Over the past decade, ARG's research prototypes have evolved into services that millions of AWS customers use every day. IAM Access Analyzer uses Zelkova to help customers like USAA and GoTo identify unintended access to their resources. Instead of hoping security policies are configured correctly, customers get mathematical proof of what their policies actually permit. Reachability Analyzer, built on the Tiros service we demonstrated in 2016, helps customers understand network connectivity without sending a single packet. Rather than testing configurations, it mathematically analyzes all possible network paths to answer whether a destination is reachable and, if not, what the blocking component is. Amazon Bedrock Guardrails with Automated Reasoning checks brings mathematical verification to generative AI. The feature helps prevent AI hallucinations by using formal logic to validate that model responses comply with defined policies — delivering up to 99% verification accuracy. These customer-facing services share a common foundation: they use satisfiability modulo theories (SMT) solvers and other automated-reasoning techniques to provide mathematical guarantees about system behavior, going far beyond what traditional testing can achieve. Proving the infrastructure beneath the cloud While customer-facing tools demonstrate automated reasoning's practical value, some of our most challenging work has focused on AWS's internal infrastructure — systems that must be correct because millions of workloads depend on them. We've used automated reasoning to prove the correctness of much of our infrastructure, including The AWS Nitro Isolation Engine Cryptographic implementations like s2n-bignum Boot code running in AWS data centers Storage systems like S3 In one particularly ambitious project, we proved correct and seamlessly replaced our entire authorization engine, which handles one billion API calls per second. We used specifications and proofs and verified the new engine against quadrillions of production authorizations. These internal verification efforts demonstrate how automated reasoning can provide mathematical certainty about the most foundational layers of cloud infrastructure. An unexpected discovery Perhaps the most surprising finding from our decade of work is that automated reasoning doesn't just make systems more secure; it often makes them more efficient and easier to maintain. When teams must write precise specifications for verification, they often discover simpler, more elegant solutions to their problems. This is partly because of the systemic approach enabled by automated reasoning. Rather than focusing on validating system behavior under specific scenarios, automated reasoning uses logic to verify system behavior under any possible scenario. Rather than considering all possible input scenarios and how they might go wrong, we define how the system should work and identify the necessary conditions for that behavior. Then we can verify that those conditions are true by using mathematical proof. In other words, we can verify that the system itself is correct. This discovery has validated our belief that mathematical rigor and practical engineering aren't opposing forces — they're complementary. The discipline required for formal verification frequently reveals opportunities for simplification that might otherwise remain hidden. Building the foundation for agentic AI The research we began a decade ago has uniquely positioned us for the next era of AI development. Our work in distributed systems and critical code verification laid groundwork that now applies directly to verifying AI-generated code. Meanwhile, our research into misconfigured AWS policies and VPC networks now helps verify the correctness of AI-generated content. Evidence for this evolution is found in some of our recent launches. Amazon Bedrock Guardrails with Automated Reasoning checks represents a significant milestone by moving from tools requiring deep expertise to capabilities embedded directly in services builders use every day. Policy in Amazon Bedrock AgentCore uses automated reasoning to set clear boundaries for agent actions, ensuring agents stay within defined compliance boundaries while operating autonomously. Teams can use natural language to specify which tools and data agents can access, and when integrated with AgentCore Gateway, the system checks policies in milliseconds. In Kiro, an agentic development environment, we released a requirements analysis capability that uses automated reasoning to prove that there are no contradictions, ambiguities, and gaps in the software requirements before code is written, increasing code accuracy and preventing future debugging cycles. The team building Kiro also uses automated reasoning to check if the underlying AI-generated code works correctly before releasing the feature publicly. As AI agents become more autonomous and take on more complex tasks, the need for mathematical guarantees about their behavior becomes critical. When AI agents suggest code changes to systems with formal specifications, those systems can automatically verify those changes against mathematical proofs. Automated reasoning provides a path to building AI systems that are not just powerful but provably safe and reliable. A decade of collaboration This work wouldn't be possible without the talented team we've built and the customers who've trusted us to help secure their critical infrastructure. Across AWS, teams are continuing to expand automated-reasoning capabilities. Senior principal scientist Daniel Kroening and his team in Annapurna are advancing hardware verification. Senior applied scientist Nadia Labai is pioneering auto-formalization research, teaching AI systems to convert natural language into formal mathematical proofs. Principal applied scientist Tristan Ravitch and his team in AWS Security launched Peri to automatically track data flow across all accounts in AWS and Amazon, with zero onboarding for service teams. These efforts position AWS to make the next decade of automated reasoning just as transformative — if not more so — than the first. What we've proven Ten years later, I'm incredibly excited about what we've accomplished. We've proven that automated reasoning delivers measurable value at every layer: from preventing configuration errors to providing mathematical guarantees about AI system behavior. We've demonstrated that mathematical rigor and practical engineering can work together to solve real customer problems at scale. The journey from that first demo day in 2016 to where we are today exemplifies AWS's commitment to transforming academic advances into services customers rely on daily. We've built a foundation of trust through mathematical certainty — and that foundation will be essential as we navigate the next decade of AI innovation. Here's to the next 10 years.

Why AI-driven purchase intent so rarely becomes a completed sale

11 August 2026 at 15:00

Presented by Rezolve Ai


When an AI assistant recommends a product or brand, it generates something valuable: a purchase-ready consumer with high intent and low friction in their decision. That consumer has already compared options, asked follow-up questions, and arrived at a conclusion. They want to buy.

What they encounter next is a commerce infrastructure that was not designed for them.

The gap between recommendation and purchase

The typical enterprise commerce stack was built for a specific model: a consumer who arrives at a brand's website through search or a direct link, navigates product pages, adds to cart, and completes checkout through a multi-step form flow. That model assumed the consumer would do the work of bridging their intent to the transaction. Most commerce systems still assume exactly that.

Agentic commerce breaks that assumption. When intent is generated outside the brand's owned environment, the handoff to transaction becomes a structural problem. Context doesn't transfer. Sessions don't persist. The consumer who asked an AI assistant for a recommendation and received one now faces the same friction-laden checkout process as someone who arrived with no prior intent at all.

Cart abandonment rates have remained stubbornly high for years. Baymard Institute research puts the average at 70%. That figure predates the agentic commerce era. As more purchase intent is generated through AI interfaces, and as the gap between that intent and a brand's transaction layer widens, the abandonment problem is likely to get structurally worse before it gets better.

What the current stack wasn't built to handle

The commerce infrastructure most enterprises operate today was assembled over two decades of incremental investment. Each layer added a capability: a search tool, a recommendation engine, a personalization layer, and a checkout system. Each was built to solve a specific problem within a human-initiated shopping journey.

None of it was built to receive intent from an AI agent.

When an AI system generates a purchase recommendation, it needs to do more than surface a product page. It needs to verify real-time inventory. It needs to apply pricing logic and promotional rules. It needs to respect brand policy around which products can be recommended together, which channels apply which discounts, and what the correct fulfillment path looks like for a given consumer. And it needs to do all of that without breaking the conversational context that made the recommendation possible in the first place.

Current commerce stacks can't do this reliably. The systems that hold the relevant data, inventory, pricing, order management, fulfillment, are not exposed in ways that AI agents can safely and accurately access. The result is a journey that starts with intelligence and ends with a broken experience: a link out to a product page, a generic checkout flow, and a consumer who arrived ready to buy and left without completing the transaction.

The conversion problem is an architecture problem

The industry has treated conversion optimization as a front-end problem for most of its history: better copy, cleaner checkout UX, fewer form fields, smarter retargeting. Those interventions were appropriate for the model they were built to serve.

The agentic commerce era introduces a different kind of conversion failure, one that front-end optimization cannot fix. When intent is generated externally, conversion depends on whether the back-end infrastructure can receive that intent, act on it accurately, and complete the transaction within the guardrails the brand has established. That is not a UX problem. It is an infrastructure problem.

Brands that are investing heavily in AI-powered discovery while leaving their execution layer unchanged are widening the gap between the promise AI makes on their behalf and the experience they can actually deliver. That gap has a cost, measured not just in lost transactions but in consumer trust that erodes each time the promise and the reality don't match.

Rezolve Ai commissioned research across 1,500 US consumers in January 2025 that found consumers who encounter friction immediately after an AI recommendation are significantly less likely to complete a purchase than those who encounter friction at the top of a traditional funnel. The implication is direct: AI raises the expectation bar at the moment of intent. Brands whose infrastructure cannot clear that bar are paying a conversion penalty they may not even know they're incurring.

What closing the gap requires

Closing the gap between AI-generated intent and completed transaction requires rethinking which layer of the commerce stack carries the most strategic weight in an agentic world. For most of the past decade, that weight sat with discovery and experience. The brands that invested most in search, personalization, and content won a disproportionate share.

In the agentic era, the weight shifts to execution. The brands that can reliably take AI-generated intent and turn it into a governed, accurate, brand-safe transaction will have a structural advantage over those whose infrastructure stalls at the handoff.

That is a different investment thesis than the industry has operated on. And most enterprise commerce roadmaps have not yet caught up to it.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Robot Recycler Salvages Parts From Broken Machines

10 August 2026 at 18:01


Objects constructed by robots are ubiquitous. If you’ve used a car, household appliance, or smartphone today, you’ve used an object constructed at least in part by robots. The more products that manufacturers want to produce (and consumers want to consume) at lower costs, the more industrial robots will be needed.

There are over 4 million industrial robots in use worldwide, according to the International Federation of Robotics. And researchers predict that number will grow to over 16 million by 2030, as manufacturing rapidly increases. But what’s going to happen when they start breaking down? A new system designed by researchers at the Karlsruhe Institute of Technology (KIT), in Karlsruhe, Germany, can predict the defect in a broken product and disassemble it while protecting valuable parts from damage. To continue robotic development sustainably, the industry should prepare for the dismantling, recycling, and rebuilding of our robotic systems.


The system consists of a predictive algorithm that guesses how a product is broken, along with robotic manipulators that actually take the broken product apart. At every stage of the process, the system checks to see if the results align with its predictions, and updates its methods if necessary. For example, in the video below, the system begins by unscrewing a broken component. To simulate a stuck screw, the researcher replaces the screw. When the system observes the screw still in place, it switches to milling away material to remove the part.

Building a product with new parts is easy, says Jan Baumgärtner, one of the designers of the system. Each step is clearly outlined, and there are no expected deviations. But taking apart something that’s broken is unpredictable. “We can imagine 100 ways that something can go wrong.” And if you start taking something apart without knowing how it broke, you might have to undo part of your work when you find the problem. For example, if you have to unscrew 100 screws holding two parts together, but the last screw is stuck, you’ll have wasted time unscrewing all those screws when you should have used a different method to remove the part in the first place.

How to Take Apart a Product

KIT’s robotic disassembly system relies on a CAD model of the broken product and of each part, so it can see how the parts should behave and understand if anything is out of the ordinary. It also uses a mathematical model to predict the damage done to a broken part.

When you give the system a broken device and a CAD model, it first guesses how each part of the broken device should move. The axes each part can move along are called degrees of freedom (for example, a screw should rotate, but not move side to side). The disassembler nudges each part to see if it moves as expected. Based on how the part actually moves, it then uses the mathematical model to predict what went wrong with the part: A corroded part might move less than you think it should, a loose screw may move more, and a deformed part might have different degrees of freedom than expected.

At the beginning of disassembly, the system formulates a plan. It guesses what might be wrong with the device it’s taking apart, and then can change its guess based on observing each piece it takes apart. For example, if there was a screw loose in the part, that might be hard to guess from an initial photograph of the broken part. But when the system moves the screw, it will notice that it can move in more ways than a screw should move, and take that loose screw into account when deconstructing the device. You can also tell the disassembly system which parts are most important to salvage intact from a broken device, and it can adjust its strategy to preserve those specific parts.

The Automated Circular Economy

Baumgärtner’s motivation behind the design of the robotic disassembler is to help create a circular economy, where old devices are repaired instead of thrown away, reducing waste. “The big future is saving our planet,” he says.

Baumgärtner envisions scaling up this one system, composed of a few robotic arms, to have many robotic disassembler arms, each with different tools. These arms will specialize in a different part of the disassembly process so that an entire factory could use different robotic limbs to disassemble a wide range of products. Think of an industrial robot factory that creates cars, but instead is specialized to take them apart. Or, as he puts it, “as a giant robot with 100 arms.”

Ultimately, if this system works as intended, it would be a fully automated way of extracting a broken part from a system, replacing it, and rebuilding the device. Then the circular economy would really shine, as people replaced broken parts in old devices instead of buying new ones all the time. “That’s why we need to think about scaling this,” he says. “Because it means it becomes so cheap that it’s cheaper to repair this [electronic device] than to produce it. That’s the goal.”

This research was presented at the IEEE International Conference on Robotics and Automation (ICRA) 2026 in Vienna.

This story was updated 11 August 2026 to clarify that the disassembly system works for products in general, not only robots.

Million-Person Study Finds a Rare Gene Variant That Slashes the Risk of Diabetes and Heart Disease

11 August 2026 at 14:00

The discovery could lead to treatments and demonstrates the power of efforts to unearth rare, beneficial genes in large populations.

“Burn fat, build muscle.” It’s a familiar workout slogan, but the benefits go far beyond aesthetics. Having less belly fat and more muscle guards against heart attacks, Type 2 diabetes, and a host of other metabolic diseases.

Some people may have a genetic edge.

A massive study of over one million people across three continents discovered a rare mutation in a gene called FNIP1 is linked to a healthier metabolic profile. The gene helps cells sense nutrients and generate energy. All of us have FNIP1, but about one in 7,000 people inherit a protective version. On average, they had a 60 percent lower risk of heart disease and metabolic disorders.

Silencing FNIP1 in human liver cells switched on a genetic program that breaks down fats. In mice fed a tasty but high-fat diet, disabling the gene curbed weight gain, prevented fatty liver disease, improved insulin sensitivity, and kept their blood sugar levels steady.

The findings are great news for everyone else. Rather than relying on a naturally occurring mutation, future gene editing therapies could potentially recreate its protective effects in people against a host of cardiometabolic diseases, a leading cause of death worldwide.

Everyone has a unique metabolic profile shaped by both genes and environment. By analyzing diverse populations, the study fished out a protective variant that spans ancestries and lifestyles. The broad reach suggests targeting FNIP1 could benefit people around the world.

The study illustrates the power of efforts to find rare, beneficial genes across large populations, wrote the authors at Regeneron Pharmaceuticals, a New York biotechnology company.

Mutant Protector

Small changes in DNA can have large consequences. Some genetic variants raise the risk for health issues. The APOE4 variant, for example, increases the chances of developing Alzheimer’s disease. Others, however, are a gold mine for new treatments.

A notable example is CCR5. People who inherit a rare mutation in both copies of thegene are naturally resistant to HIV. The mutation prevents the virus from tunneling into immune cells and replicating. The discovery has led to multiple success stories in which bone marrow transplants from donors carrying the mutation kept HIV at bay, without the need for lifelong antiviral drugs.

Protective mutations could also lower the risk of heart disease. Rare variants of PCSK9, a gene involved in cholesterol metabolism, disable the gene and slash dangerously high levels of LDL, or “bad” cholesterol that clogs arteries. The discovery has already spurred a handful of therapies that block the gene or its protein with early successes.

“Identifying genetic variants associated with protection from disease is a powerful strategy,” wrote the authors. “However, protective genetic variants are often extremely rare, so finding them requires sequencing the genomes of large populations.”

Go Big

To better understand cardiometabolic diseases, the team sequenced the genomes of over a million people from 11 studies across the Americas, Europe, and Asia, including people with African ancestry. They also linked genetic data with participants’ health records.

The researchers searched for gene variants that influence a blood biomarker for cardiometabolic disease. Called TG:HDL, the biomarker is the ratio between two types of fats. The first, triglycerides, is packaged into tiny “bubbles” that circulate the bloodstream. High levels are linked to heart attacks, strokes, and other metabolic problems. In contrast, high-density lipoprotein, often called “good” cholesterol, ferries excess fat away from tissues and blood vessel walls to the liver, where it can be cleared.

Across the populations in the study, a lower TG:HDL ratio—that is less TG, more HDL, or both—tracked with better metabolic health. People with lower ratios had reduced insulin levels, lower blood pressure, and less fat buildup in the liver and muscles. The biomarker also predicted diabetes risk, heart problems, and liver scarring, making it a powerful snapshot of overall metabolic health.

The team then scanned the genome for rare gene variants linked to TG:HDL. Roughly 60 genes popped up, all involved in energy storage and active in the liver and fat tissues.

But one gene stood out: FNIP1. Rare variants essentially disable the gene by disrupting its protein-making instructions. People with one copy of these variants had lower liver fat and blood sugar and roughly 60 percent lower risk of cardiometabolic disease.

The finding “was remarkable and thought-provoking, and immediately motivated us to dig deeper into the biology of this discovery,” wrote the team. But a key question remained: Were the variants actually protecting people, or were they simply correlated with better health?

To find out, the team silenced the gene in human liver cells using a method called siRNA. Rather than snipping the gene, siRNA blocks cells from producing targeted proteins. Without functional FNIP1, liver cells ramped up genes involved in breaking down fats.

The researchers then turned to mice. Using CRISPR-Cas9, they got rid of FNIP1 and related signaling pathways specifically in mice fed a high-fat, high-sugar diet. The intervention rapidly activated mitochondria—the cell’s energy factories—and lysosomes, the acid-filled recycling centers that break down waste. Despite gorging on the unhealthy diet, mice lacking functional FNIP1 had less body and liver fat, more muscle mass, and better sensitivity to insulin.

That’s not to say FNIP1 is a “villain” gene. Normally, it acts as a metabolic brake, helping the body conserve precious energy when food is scarce. But many of us now face the opposite problem, an abundance of calories and not enough physical activity. Releasing that brake, through medication or gene editing, could rev up the body’s natural fat-burning machinery.

Turning the finding into a therapy won’t be simple. The protective effects were found in people who carried the mutation from birth. A short-term drug or gene therapy delivered later in life might not reproduce the same effects.

Safety is another major concern. Paradoxically, people who have mutations in both copies of FNIP1 develop heart disease and immune deficiency. And mice without functional FNIP1 throughout the body are more prone to liver damage and cancer. Targeting treatments specifically to the liver—for example, using lipid nanoparticles—could limit side effects, but any potential therapy will need to be thoroughly tested for safety.

The team is searching for drug candidates that inhibit FNIP1. But for now, they’ve shown the power of large-scale genetic screens across diverse populations to find rare protective variants—and potential paths towards treating diseases that affect millions of people.

“Identifying FNIP1, a previously poorly characterized gene involved in lipid metabolism, is highly novel and promising for future drug development for metabolic health,” Satoshi Koyama at the Broad Institute, who was not involved in the study, said in a research briefing. “I sincerely hope that this discovery will one day benefit patients with metabolic disorders.”

The post Million-Person Study Finds a Rare Gene Variant That Slashes the Risk of Diabetes and Heart Disease appeared first on SingularityHub.

Mistral AI wants to build 1 gigawatt of European compute by 2030 — and lock in customers now.

Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached.

The French artificial intelligence company announced Tuesday a three-part expansion of its infrastructure business: regional inference endpoints that let customers choose whether their AI workloads run in Europe or the United States, a new "Priority Tier" backed by an uptime guarantee for mission-critical deployments, and a coalition of European enterprises making multi-year compute commitments that Mistral says will underwrite 200 megawatts of infrastructure across Europe by the end of 2027 — and a full gigawatt by the end of 2030.

In a move that may raise eyebrows among sovereignty purists, the company also said it will begin hosting third-party open models on its platform, starting with GLM-5.2 from Z.ai, the Chinese AI lab formerly known as Zhipu.

Taken together, the announcements mark a decisive shift in how Mistral positions itself. The company that built its reputation training open-weight language models is now selling something closer to critical infrastructure: assured capacity, regional control, and contractual reliability for enterprises and governments that want frontier AI without surrendering control over where it runs.

"When we spoke in June, the story was around how Mistral was building a full-stack AI offering," Timothée Lacroix, Mistral's co-founder and chief technology officer, told VentureBeat in an exclusive interview ahead of the announcement. "Today, the announcement is about strengthening one part of this infrastructure, which is the inference part."

That one part, it turns out, comes with a price tag measured in the tens of billions of dollars.

Inside Mistral's plan to build 1 gigawatt of European AI compute by 2030

The headline numbers deserve scrutiny, because they imply staggering capital requirements. Mistral currently operates less than 200 megawatts of capacity, according to the company. Details shared with VentureBeat show the near-term buildout resting on three sites: a 44-megawatt facility near Paris that became operational in the second quarter of this year, a 23-megawatt facility in Sweden built in partnership with EcoDataCenter using renewable energy and advanced cooling, and a 10-megawatt site in Les Ulis, France, that came online in the third quarter.

Getting from there to one gigawatt by 2030 is a different order of magnitude. Independent estimates suggest just how different: research firm Epoch AI calculates that a typical one-gigawatt AI data center requires roughly $38 billion in upfront capital expenditure, with servers and GPUs — not buildings or land — consuming the majority of the cost. Goldman Sachs Research pegs next-generation AI facilities at $15 million to $20 million per megawatt before accounting for the chips inside them.

Lacroix did not dispute the scale of the challenge. The investment required for a gigawatt of capacity "is a large investment that requires also a lot of scaling and revenue behind it," he said.

The urgency, in his telling, comes from a supply crunch that is about to get worse. "More and more, and especially around 2027 and 2028, we see that the demand for AI compute is exceeding what the market has to offer, especially in Europe," Lacroix said. McKinsey has estimated that meeting global AI demand could require $5.2 trillion in data-center capital expenditure by 2030 — and Europe, by most analyses, is starting from behind.

A company valued at a fraction of its American rivals cannot close that gap with venture capital alone. Which explains the most consequential — and most unusual — piece of Tuesday's announcement.

European Compute Units turn AI sovereignty into a five-year contract

Mistral is assembling what it calls an anchor group of enterprises whose long-term commitments will collectively finance infrastructure none of them could justify alone. Those commitments convert into "European Compute Units," or ECUs — a claim on Mistral-built capacity over multiple years that participants can spend on inference, training, model adaptation, or other AI workloads as their needs evolve.

If that structure sounds more like a power-purchase agreement than a cloud contract, that appears to be the point. Data-center financing increasingly resembles large infrastructure projects — gigawatts, substations, energy agreements — rather than traditional technology spending, and lenders want demand locked in before capital gets deployed. Mistral raised €830 million ($962 million) in debt earlier this year to fund its data center near Paris, TechCrunch reported in March, and pre-committed enterprise demand is exactly what makes that kind of financing repeatable at ten times the scale.

Lacroix was unusually direct about the mechanics. "The entire point of compute units is to have commitment," he said. "The goal is to have customers commit for around five years, or at least a long time." Asked what happens if a customer wants out early, he didn't soften the answer: "There is no getting out."

What makes a five-year, no-exit commitment palatable, he argued, is flexibility in how the capacity gets consumed. "Typically this can be spent on raw inference that you then feed through any other AI stack. It can be spent on raw compute as managed Kubernetes, and it can be spent at the very top with our full AI offering," he said. "My hope is that they will use it with our full-stack services and will love it."

The anchor group already includes some of Europe's industrial heavyweights. Amadeus CEO Luis Maroto said in a statement that "capacity, deployment control, and operating continuity become increasingly important for all enterprises." ASML chief Christophe Fouquet — whose company led Mistral's $13.4 billion (€11.7 billion) Series C last year — called building European AI capacity one of the few industrial endeavors that "will matter more to Europe's next generation," while Capgemini's Aiman Ezzat framed it as "a question of who shapes the future of European industry." CMA CGM chairman Rodolphe Saadé said the shipping group's Mistral deployment is "already under way among thousands of employees."

Commitments of that duration only make sense, of course, if the sovereignty being purchased is real. On that question, Mistral's announcement contains an asterisk worth reading closely.

The fine print on sovereign AI: what data can still leave Europe

The centerpiece product is Mistral Regional Endpoints, now generally available, which let customers pin inference and its associated processing to Europe or the U.S. Alongside it, the new Priority Tier — in public preview — offers committed service levels, custom rate limits, and an uptime SLA for mission-critical workloads.

Mistral claims it is the only European AI lab offering both a choice of processing region and an SLA-backed service tier, and Lacroix said a third option is coming: an endpoint "that stays on Mistral-controlled infrastructure, so on Mistral compute" — for customers who want their inference not just in Europe, but off hyperscaler hardware entirely.

Then comes the fine print. Mistral's own materials note that in-region inference remains subject to "limited, safeguarded transfers" to sub-processors that may sit outside the chosen region. Pressed on what actually leaves Europe, Lacroix pointed to the connective tissue of modern AI applications: tool calls.

"There are some tool services, like some tool calls, that might be hosted in places where we don't fully control this," he said, citing web search as an example. "A few of our web-search providers might not all be in Europe, and in that case, we need to potentially gate that capability."

His answer to the compliance question — would this satisfy a European bank or a defense ministry? — was that gating is the feature, not the bug. Capabilities that cannot be sourced in-region can be switched off entirely, restricted to certain users or workspaces, or, given sufficient demand, rebuilt with European providers. "Any capabilities that we don't find a provider for in Europe — if it needs to be done in Europe, we'll find some way to implement it or find ways to address it," Lacroix said.

For enterprise buyers, that is a more honest framing than most sovereignty marketing offers: full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision rather than a default. The same pragmatism runs through the announcement's most surprising line item.

Why Europe's open source AI champion is hosting China's GLM-5.2

A French national champion — one that has partnered with the French army and positioned itself as Europe's answer to American AI dependence — hosting a Chinese lab's model invites an obvious question. Lacroix's answer was disarmingly matter-of-fact.

"It's a great model. Everyone loves it. It's open weight, so there was no good reason for us not to do it, really," he said, noting that Mistral's own stack is already built on open-source software like Kubernetes.

On security vetting, he argued that open weights fundamentally change the risk calculus. "The risks in taking a new model, at the layer of the weights, are — at least in my opinion — rather limited," Lacroix said. "We checked basically all of the safety and compliance evals that we have. We'll control that model, its outputs, and what it does the same way we do any of our models. We have the same inputs and outputs and monitoring capabilities over all of it."

The strategic logic is worth unpacking. By hosting third-party open models under European regional controls and the same SLAs as its own, Mistral is repositioning itself from model vendor to sovereign distribution layer — the trusted intermediary through which any open model, regardless of origin, can be consumed by a regulated European enterprise that could never call a Chinese API directly. It is the "model garden" playbook the hyperscalers run with Bedrock and Vertex, executed on European soil with European guarantees.

Customers appear to be reading it that way. "Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements," Matan Griberg, CEO of AI software-engineering company Factory, said in a statement.

Lacroix stressed the move is not a retreat from frontier training: the model Mistral had in training as of June "is still training, and we're still very excited about it," he said. But openness to rivals' models signals where the company now believes its moat lies — not in any single model, but in the infrastructure underneath all of them. Which makes its relationship with the world's most powerful infrastructure company all the more interesting.

How the multibillion-dollar Microsoft deal funds Mistral's independence

Hovering over every sovereignty claim is Mistral's deepening relationship with Microsoft. In July, the two companies announced a multibillion-dollar expansion of their partnership under which Microsoft will rent capacity from Mistral's European data centers to serve its own cloud and AI demand, while adding Mistral Medium 3.5 and OCR 4 to Microsoft Foundry, bringing Medium 3.5 to Copilot Studio, and enabling Mistral models on Azure Local for disconnected, customer-controlled environments. Mistral CEO Arthur Mensch told The Wall Street Journal at the time that two-thirds of Mistral's customers already work with Microsoft.

How does a company selling independence from U.S. hyperscalers square taking one on as its largest tenant? Lacroix described Microsoft not as a patron but as an anchor customer that de-risks the buildout.

"It allows us to scale different parts of the business differently by building infrastructure with Microsoft as a customer," he said. "We can scale that team, we can scale our infrastructure, and make sure that we can then, on the side of it, also build for ourselves and for our customers." He compared the arrangement to the neocloud playbook — companies that built businesses supplying capacity to the hyperscalers themselves. "As that part of our business resembles that of neoclouds, we're following the same thing."

It is a genuinely clever inversion: rather than renting American infrastructure, Mistral is renting infrastructure to one of America's largest companies, using Microsoft's demand to finance capacity that also serves European sovereignty customers. But the independence has limits no contract can engineer away — the GPUs filling Mistral's European data centers come overwhelmingly from Nvidia and other American chipmakers, as SiliconANGLE noted in its coverage of the July deal.

Asked directly why a customer should choose Mistral over an EU region on AWS or Azure, Lacroix gave two answers. "The simplest possible answer is capacity. There is more demand than supply right now, and so it adds another option," he said. The second cuts closer to the pitch: "We are a European provider, and on the region that would be Mistral compute, we are fully independent. That's a truly differentiated offering than all of the hyperscalers or pure inference companies can provide."

The economics of open models: why agentic AI is pushing inference to the cloud

There has always been a tension at the heart of Mistral's business: its best-known models are free to download, and open models have historically been difficult to monetize through APIs. Asked how free weights fund a gigawatt buildout, Lacroix offered the clearest articulation yet of the company's thesis — that the economics of self-hosting are collapsing under the weight of the models themselves.

"When the models were smaller, and we were before the explosion of agentic AI, it was doable for enterprises to host their own — up to, let's say, 100-billion-parameter dense models — on their premises," he said. "More and more, with models going into the trillion or more parameters, with the current hardware, and with the increasing amount of tokens that need to be processed, it becomes harder."

His conclusion was blunt: "I don't see how, with the current trend of model size and growth of agentic tokens, we keep the full inference on-prem. To me, that is why we think we're going to monetize our cloud inference." Inference, he noted, is particularly well suited to the cloud because it "does not need to hold any data" and can be encrypted in transit.

In other words: open weights get Mistral into the enterprise, and the physics of trillion-parameter agentic workloads brings the inference — and the revenue — back to Mistral's data centers. The thesis will get an expensive test. Mistral has raised roughly $4 billion to date, according to PitchBook data — a fraction of the war chests assembled by OpenAI and Anthropic — and Bloomberg reported in June that the company is in talks to raise about €3 billion at a roughly €20 billion valuation, nearly double its Series C mark. The revenue behind the buildout will have to come from exactly the enterprises Tuesday's announcement is courting.

And Europe, in Mistral's telling, is only the first market for what it is selling. Asked whether the framework could be replicated in the Middle East, Asia, or anywhere else anxious about AI dependence, Lacroix didn't hedge: "It's completely right. We're starting this in Europe because it's also an easier part of the world for us to scale into, especially in the infrastructure. But we definitely want to extend this, depending on customer demand." Every layer of the stack, he said, "can be controlled, changed, replaced depending on where we operate and what the requirements are — that's pretty much where we excel."

That is the wager underneath the SLAs, the compute units, and the Chinese model flying a European flag: in a world where the U.S. and China dominate frontier AI, the durable business is selling everyone else control. To fund it, Mistral is asking Europe's largest enterprises to sign five-year contracts with no exit — while making a bigger, longer commitment of its own. A gigawatt, after all, is a promise measured in decades. For Mistral, too, there is no getting out.

NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents

11 August 2026 at 13:01
Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning...

Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning model for every execution step adds cost and latency. NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed for harnesses…

Source

Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

11 August 2026 at 13:00
Decorative image.Building an AI agent does not end with choosing a single model. Each model has its own strengths, weaknesses, and cost profile, which can shift from one...Decorative image.

Building an AI agent does not end with choosing a single model. Each model has its own strengths, weaknesses, and cost profile, which can shift from one workload to another—or even within the same workload. For example, an agentic task may need classification for one step, reasoning for the next, and a smaller model for routine follow-up tasks. Sending every request to the largest model can…

Source

❌