Normal view
How mobility gives language models a deeper understanding of place
We May Be Wrong About How the Brain Stores Memory
In a new study, mice recovered their memories by regrowing brain connections lost during artificial hibernation.
Our cherished memories may be more resilient than previously thought.
Long-term memories are stored in synapses, the connections between neurons. These structures sit on tiny protrusions called dendritic spines, which dot neurons’ branching arms.
When we learn, these spines grow. Larger spines tend to form stronger synapses and are more likely to persist during learning. In Alzheimer’s and other diseases that eat away at these connections, memories can fade.
At least, that’s the traditional picture. A new study suggests the story is more complicated.
Mice in artificial hibernation rapidly lost roughly half of their synapses, both large and small. Yet once awakened, they resurfaced memories of previously learned tasks. Spines that had withered during the induced deep sleep regrew in their original spots, once again forming functional synapses. This suggests their brains had rebuilt parts of broken circuits.
A small number of stubborn synapses that survived hibernation may explain how this happened. These synapses formed clusters that preserved memories as patterns of neural activity called engrams. The more surviving clusters the mice had, the better they performed on a previously learned task after awakening.
“It was astonishing. Logically, if all our engram synapses were essential in memory retention as traditionally thought, memory should have massively deteriorated,” said study author Yu-Ju Lin at Japan’s Okinawa Institute of Science and Technology Graduate University in a press release.
The findings suggest that memories may not depend on preserving every individual synapse. Instead, they may be distributed across a higher-level architecture of connections, with some synapses acting as anchors that can reconstruct the rest.
Artificial hibernation is an extreme case, and it’s far too early to know how the findings translate to diseases like Alzheimer’s. Still, they suggest that even under extreme circumstances, the brain can bring back memories once thought lost.
Forest for the Trees
Neurons are often called the brain’s computational units. But each one is actually a sophisticated mini computer in its own right.
A neuron’s branching arms receive signals from neighbors, while a long, winding extension carries outgoing messages to other neurons. Spines dot the receiving branches. These structures can strengthen, weaken, appear, and disappear depending on the input. This allows synapses to simultaneously gather data, learn, and store memories. When neurons repeatedly activate each other, the connections between them grow stronger, mostly because of larger spines. This is the idea behind the popular neuroscience saying: “Neurons that fire together, wire together.”
For episodic memories—the when, where, what, and who of our lives—these changes begin in the hippocampus, a region central to forming and retrieving memories, and one of the first areas damaged by Alzheimer’s disease.
During the day, the hippocampus forms engrams associated with individual memories. During sleep, some of these are erased, while others are gradually incorporated elsewhere in the brain for long-term storage. The hippocampus also helps recall memories by adding context, such as where something happened or how you felt at the time.
All of this should, in theory, require relatively stable brain circuits. “Long-lasting changes in synaptic connections are widely thought to provide the structural basis of memory,” wrote the team.
But recent studies have challenged that view. The brain is anything but static. Synapses are constantly being remodeled. Even which neurons are recruited into a particular engram can change over time. Some synapses may effectively hand off information to others, freeing themselves to encode something new.
If physical traces of memories are always shifting, why don’t our memories disappear with them? That’s the question the new study explored.
Going Under
To probe the paradox, the team turned to an unorthodox method: Artificial hibernation. Like natural hibernation in bears and other animals, artificial hibernation dramatically lowers body temperature and metabolism and causes animals to enter a sleep-like state. As the brain decreases its activity to conserve energy, synapses begin to wither.
Yet hibernating animals do retain memories. Chipmunks, for example, remember where they’ve stored food, returning to their stashes when periodically awakening for “midnight” snacks. This suggests hibernation could be a useful way to study how memories survive major changes in the brain.
“Our brains are incredibly complex. If hibernation can reduce and simplify brain activity and structure, it could make studying these convoluted systems a bit easier,” said study author Kazumasa Tanaka. “That’s why I wanted to use artificial hibernation techniques to study memories.”
The team first trained mice on two standard memory tasks. In one, the critters received a mild electrical zap to their paws inside a chamber with distinctive smells and decorations, teaching them to associate that setting with danger. In the other, they learned to navigate a maze towards a sugary reward.
The researchers then activated a neural circuit that drove the mice into artificial hibernation for two days. Using fluorescent proteins, they tracked changes in the animals’ synapses throughout the process.
Spine remodeling began within minutes. Some rapidly shrank and disappeared, taking their synapses with them. Within a day, over half of the synapses were gone. Even the larger spines thought to be especially important for long-term memories were pruned.
Yet memories survived. When the mice awoke and revisited the shock chamber, they froze in fear. In the maze, they still knew how to find the reward. Previously pruned spines also returned, with roughly 80 percent growing back at their original locations along the neuron’s branches.
To test whether this recovery is unique to hibernation, the team compared the animals with a second group that underwent anesthesia and were dosed with a drug that blocks synaptic changes—a combination known to cause amnesia. These mice also lost a large number of synapses but never recovered their memories.
A core cluster of unusually resilient synapses may explain the difference. These synaptic clusters formed a unique architecture in which one neuron linked to multiple neighbors like Grand Central Station. The clusters were often located in areas where spines were tightly grouped—making them more likely to receive inputs from multiple sources at once. Somehow, they kept memories intact even as surrounding synapses disappear.
“This suggests that for long-term memory, only particular clusters of synapses matter—the rest may be dispensable,” said Tanaka.
Exactly how these clusters preserve memories remains unclear. How does the brain create and maintain them? Do they anchor multiple memories? And could the same mechanism help explain why some memories remain as synapses are lost in disease?
The team is now using genetic and molecular tools to decipher what makes the clusters so resilient. Tinkering with their formation could better reveal their role preserving memories and, in theory, inspire ideas for tackling synapse loss in the early stages of diseases.
Beyond neuroscience, demystifying how memories linger could inspire neuromorphic chips—hardware that loosely mimics the brain—or even new AI models. For now, the findings offer a twist on an old idea: A memory may not need every single synapse that helped create it. It may just need the right ones to rebuild the rest.
The post We May Be Wrong About How the Brain Stores Memory appeared first on SingularityHub.

Running Codex as a Headless Agent
Turning Codex from an interactive assistant into a programmable automation component
The post Running Codex as a Headless Agent appeared first on Towards Data Science.
Estimating from No Data: Deriving a Continuous Score from Categories
A walkthrough of and the maths behind using low-capacity networks to acquire fine-grained scoring when only categorical labelling is available for training
The post Estimating from No Data: Deriving a Continuous Score from Categories appeared first on Towards Data Science.
-
The latest research from Google
- An AI tool for prioritizing candidate biomarkers from wearable sensor data
An AI tool for prioritizing candidate biomarkers from wearable sensor data
-
AI Infrastructure Archives - The New Stack
- Anthropic’s new browser tool doesn’t actually run a browser
Anthropic’s new browser tool doesn’t actually run a browser
Anthropic launched a new Browser Use tool that gives Claude a structured view of a web page in addition to what is visually rendered. Announced Thursday, the tool uses the page’s accessibility tree to help Claude find and interact with specific elements directly rather than having to work out where they are on the screen.
Browser Use is part of a broader Anthropic release that also brings Computer Use, the Skills API and Files API into general availability. Developers can access the browser tool through the Claude API using browser_toolset_20260801.
Browser Use is part of a broader Anthropic release that also brings Computer Use, the Skills API and Files API into general availability.
The change gives Claude a more direct way to interact with a web page. Instead of working out a button’s position from a viewport image and targeting coordinates such as x: 640, y: 320, Claude can receive a reference such as ref_3 tied to that element and use it when it wants to act.
Page references replace coordinates
Computer Use can operate across an entire desktop by looking at screenshots and sending mouse coordinates and keyboard commands. Browser Use works within the browser itself, where it can use page structure that would be difficult to recover reliably from pixels alone.
When Claude calls read_page, the developer’s executor returns a text representation of the accessibility tree, in which elements such as links, buttons, and text boxes can be tagged with references. If Claude later wants to click a button represented by ref_3, it can send that reference along with the requested operation rather than trying to calculate where the button is on the screen.
That said, if the tab navigates to a new page or the page changes enough, a reference that pointed to a button a moment ago may no longer work. The API will not catch that on its own, so the executor has to recognize when the reference no longer matches the underlying element, reject the action and have Claude read the page again before continuing.
Batching cuts model calls
Playwright, for example, can represent a page as an ARIA snapshot and locate elements by role rather than coordinates. At the same time, Microsoft’s Playwright MCP server already exposes structured accessibility snapshots with references a model can use to identify elements. The concepts line up closely with Browser Use, but the protocols do not: Playwright MCP speaks MCP, while Anthropic’s tool uses its own client-toolset protocol, so developers would still need an adapter that translates Claude’s requests into Playwright actions and returns the results in the format Claude expects.
Puppeteer offers many of the same building blocks, exposing the browser’s accessibility tree via Accessibility.snapshot() and providing APIs for controlling Chrome and Firefox. A developer could use those APIs for navigation or page reads, then maintain Anthropic’s reference mappings on top.
A developer could use those APIs for navigation or page reads, then maintain Anthropic’s reference mappings on top.
Slightly confusing, an unrelated open-source project also called Browser Use runs AI browser agents against Chromium through the Chrome DevTools Protocol. Despite the shared name, it has no connection to Anthropic’s tool and comes with its own agent loop and browser abstractions, so connecting the two would still require integration work.
Several browser actions can happen in one turn
Anthropic is also reducing the back-and-forth between Claude and the browser by allowing multiple actions to be requested in a single model turn. Now actions can arrive together as several tool_use blocks. The application executes them in order and sends the results back together, avoiding another model call between every click and keystroke. Anthropic says that can lower latency and costs, particularly as workflows scale from a handful of interactions to dozens or hundreds.
That matters more as browser tasks get longer. Cheaper models alone will not solve the token cost problem in agentic workflows, so cutting unnecessary model calls is another way to reduce costs.
If Claude has to return to the model after every click or keystroke, a long browser task can quickly rack up model calls. Batching cuts out some of that back-and-forth by letting Claude request several actions at once, but the browser still has to carry them out in order because each one depends on what happened before it. If Claude asks to click a button, fill in a field, and submit a form, for example, the executor cannot simply move on to the next step if that first click fails, because everything that follows is now based on a page state Claude never reached.
Batching cuts out some of that back-and-forth by letting Claude request several actions at once, but the browser still has to carry them out in order because each one depends on what happened before it.
Developers host the browser
Browser Use is currently limited to the Claude API and is not available inside Claude Managed Agents. Adding it to a Messages API request exposes 27 browser operations by default. Claude can decide which of those operations it wants to use, but Anthropic does not execute them. The application has to translate each request into an action inside its own browser environment, preserve the session between turns and return enough information for Claude to understand what happened.
Loading all of those operations has a token cost. Anthropic’s pricing documentation says the default Browser Use toolset adds roughly 6,600 input tokens to a request, before counting screenshots, accessibility trees and other results sent back to Claude. Developers can turn off operations they do not need to reduce that overhead.
It also creates a different hosting split from some of the other tools Anthropic announced Thursday. Skills uploaded through the Skills API can run inside Anthropic’s code execution sandbox, while the Files API stores documents that can be reused by ID. Browser sessions, along with their downloads and uploaded files, stay in the developer’s environment.
Approval gates need rethinking
Claude can still encounter a prompt injection in web content or be redirected to an unexpected location, which is why Anthropic recommends running the browser in an isolated container or virtual machine with minimal access. JavaScript and file uploads should remain disabled unless needed, since code generated by Claude runs with the page’s privileges and can reach data or make requests available to that page.
Batching makes approval a little trickier because several actions can arrive at once, and a routine click at the beginning of a sequence could eventually lead to something that requires the user’s permission. That means the executor has to check actions as they happen and stop for approval when needed.
The post Anthropic’s new browser tool doesn’t actually run a browser appeared first on The New Stack.
Nvidia finds that simple linear math can replace costly AI model handoffs
When an agentic AI system hands a task from a small model to a larger one — or back down again — it pays a steep tax: the receiving model has to recompute the entire conversation from scratch, driving up compute costs and latency. This is a major bottleneck for enterprises building long-horizon, multi-LLM workflows.
To solve this challenge, researchers at Nvidia have introduced a cross-model KV cache transfer technique that directly maps the prefilled KV cache from a source model into the target model. This technique aligns with real-world agentic applications where large contexts accumulate across many turns.
For real-world AI applications, cross-model KV cache transfer can reduce compute costs and latency on long-running, multi-LLM workflows — and it does so with simple linear math, not an expensive deep learning model.
Experiments show that, on compatible model pairs, this linear mapping process runs 2.7 to 25 times faster than recomputing the conversation while retaining up to 98% of the target model's standalone accuracy.
Why swapping models mid-session is so expensive
Examining how LLMs handle memory helps understand why multi-model workflows hit a performance wall in production. When an LLM receives a prompt, it must first execute the “prefill” stage, which is the initial forward pass that computes the keys and values for all input tokens and populates the Key-Value (KV) cache.
After that, it enters the “decode” phase, where it computes and generates the next tokens in the sequence. During this phase, the model reads from this KV cache to predict new tokens one by one, bypassing the need to re-evaluate the entire history of the conversation for each new token.
In multi-turn conversations or long-horizon agentic sessions, the context gradually becomes longer. Because the computational cost of the prefill stage scales directly with both model size and input length, processing these long sessions becomes increasingly expensive and introduces significant latency if the KV cache is invalidated.
This invalidation happens whenever the AI system tries to swap models mid-session, such as routing a complex reasoning step to a larger model or dropping to a smaller model to save costs. Because different LLMs have different architectures, they expect their cache inputs in different formats.
As a result, any model switch forces the receiving model to repay the entire prefill cost from scratch to recompute the KV cache for the accumulated context.
Mapping memory between models without starting over
The Nvidia researchers studied cross-model KV cache transfer to see how developers can transform the KV cache of one model into the expected format of another without running the prefill phase again.
If solved, cross-model KV cache transfer has benefits in both directions. Small-to-large model transfer upgrades the quality of the output. For example, a cheap, small model handles the routine parts of an agentic workflow but struggles with a complex reasoning problem, and you map the KV cache to a larger model and continue the process seamlessly.
On the other hand, large-to-small model transfer reduces compute costs. A highly capable, large model might be used to unpack a massive, complex system prompt or synthesize a dense PDF at the start of a session. Once the heavy lifting is done, the session's KV cache is mapped down to a smaller, more economical model to handle the rapid-fire, conversational turns that follow.
There have been previous efforts to solve the KV cache transfer problem, but they suffer from a few key limitations. These include the need for expensive gradient-based training or very strict architectural constraints.
For this initial study, the authors restricted their focus to within-family transfers, such as transitioning between different-sized models in the Qwen, Llama, or Ministral families. These models share tokenizers, training data DNA, and core architectural styles but differ in size and depth. However, this framework leaves plenty of room for future experiments. The researchers note the technique could eventually be expanded to cross-family transfers, mismatched KV head counts, or hybrid architectures that blend standard attention with other memory mechanisms.
The key finding of the Nvidia study is that cross-model KV cache is a significantly linear structure. This means you can do the mapping with simple algebra tricks and without the need for heavy neural network training. For example, when experimenting on KV cache transfer from a 14-billion parameter Qwen3 model to a 32-billion parameter version, the authors discovered that a simple linear regression mapping from one source layer to a target layer can recover 56% of the variance in the target’s keys and 32% of the variance in its values. When combining multiple source layers, those numbers climbed to 79% and 65% respectively.
To translate this linear relationship into a practical system, the researchers designed a closed-form per-head ridge mapper with three key components:
Per-head ridge regression: Instead of using complex deep learning to train the system, they fit a simple linear regression using a tiny calibration set of a few hundred text sequences. This technique solves a classic line-of-best-fit problem independently for every attention head.
Cross-layer source selection: Because the source and target models have different numbers of layers, the mapper evaluates and selects the most predictive source layers to feed into each specific target layer. This way, the system picks only the most helpful pieces of memory from the old model to construct the new model's memory.
Content-space mapping: Before translating the data, the mapper strips away the RoPE encodings. RoPE, or Rotary Position Embedding, is a standard mechanism that applies a mathematical, position-dependent rotation to the data so the model understands the order of the tokens in a sequence. Stripping the RoPE values makes it possible for the mapper to generalize to sequences of lengths larger than its training data.
Putting the linear mapper to the test
To test whether the technique works, the researchers evaluated the transfer pipeline across six “matched-KV” model families. Matched-KV means the source and target models share the same KV head count and per-head dimensions, which is typical for different-sized models within the same family.
The model families included Qwen3, Llama 3.1, and Ministral 3, with tests for KV cache transfer across different sizes ranging from 3 billion to 70 billion parameters. Their experiments included a massive 8.8x parameter leap from Llama 3.1 8B to 70B.
To cover a wide range of tasks, they evaluated the models on five core accuracy benchmarks (ARC-Challenge, HellaSwag, WinoGrande, MMLU, and GSM8K) as well as language modeling perplexity on WikiText-2 and a multi-turn conversation task called CoQA. To fit the linear translation mapper, they used a tiny calibration dataset of just 500 text sequences of 1,024 tokens each.
The researchers compared the framework against the baseline ceiling accuracy where the target model does a full, traditional prefill. They also compared their full system against ablated configurations, such as reducing the number of selected layers or deactivating different components. Additionally, they compared their simple method against a deep neural network trained with backpropagation to see if heavier deep learning could recover accuracy on pairs where the linear method struggled.
For four of the six tested pairs, the fast, closed-form linear ridge mapper retained 73% to 98% of the target's standalone prefill accuracy — including the massive leap from Llama 3.1 8B to 70B, which retained 72.8% of target accuracy.
The mapper also runs between 2.7 and 25 times faster than re-prefilling. For example, when translating a 32,768-token KV cache from a Qwen3 14B to a 32B model, the transfer took just 278 milliseconds, compared to nearly 7 seconds for a standard re-prefill.
The system also demonstrated high stability on tasks that run across many steps. When tested on multi-turn conversations, the drift, or accuracy loss, between the target baseline and the transferred cache remained incredibly small across 10 turns, proving it will not cascade into failure during long agentic sessions.
However, the straightforward linear approach did run into limitations on specific model pairs. For two of the Ministral configurations, the linear mapper degraded sharply because the simple linear fit failed to extrapolate outside calibration data. To fix this, the researchers swapped the linear mapper for a nonlinear multi-layer perceptron (MLP) with two 1,024-unit hidden layers trained on the same data. This added a complexity and training tax to the setup, but it recovered their accuracy to above 90%.
A bigger industry problem than one paper can solve
The introduction of cross-model transfer is part of a broader, industry-wide push to solve the KV cache bottleneck, which has emerged as one of the key hurdles for scaling enterprise AI. As developers push LLMs to process massive documents or code bases and execute long-running reasoning tasks, managing this memory layer is becoming as important as the models themselves.
Over the past year, researchers have attacked this compute and memory problem from multiple angles. For instance, Nvidia recently introduced dynamic memory sparsification (DMS), a technique that intelligently evicts less important tokens from the KV cache to cut reasoning costs by up to 8x.
Other approaches focus on aggressive data compression. MIT researchers developed an algebraic compaction technique called Attention Matching that compresses the KV cache by 50x without degrading quality. Similarly, Nvidia introduced KV Cache Transform Coding (KVTC), which borrows media compression concepts to shrink memory by 20x without altering the underlying model weights.
Beyond compression, researchers are also attacking the computational overhead of memory retrieval. Optimizers like IndexCache strip away redundant layer calculations to deliver significantly faster time-to-first-token in long-context applications. And models like DeepSeek and the GLM series are optimizing the KV cache through architecture innovations.
As AI systems take on longer-horizon tasks and more complex architectures, the underlying memory infrastructure is becoming as important as the models themselves. Cross-model KV cache transfer gives developers one more tool for keeping inference costs down as they scale multi-model agentic systems.

Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG
Enterprise Document Intelligence [Vol.1 #7sexies] - The unit of retrieval doesn’t have to be a page or a paragraph. When the corpus carries tables, each body row with its column headers is a chunk in its own right, and it’s often the one row the reader asked about
The post Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG appeared first on Towards Data Science.
GPU-Accelerated Clustering for Financial Instruments at Scale
Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor loadings, and structural-break signals at single-GPU and multi-node scale Quant strategies routinely group instruments for portfolio construction, risk aggregation, statistical arbitrage, and trade surveillance. Incorrect groupings can make…
-
Amazon Science homepage
- SOP-Bench: A new benchmark for evaluating AI agents on real business procedures
SOP-Bench: A new benchmark for evaluating AI agents on real business procedures
The Types of Dimensions in a Star Schema, and How to Use Them
Dimensions are one of the two main object types in dimensional modelling. But what are the different types of dimensions? And how can you use them?
The post The Types of Dimensions in a Star Schema, and How to Use Them appeared first on Towards Data Science.
Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS
AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency. Not every megawatt translates to revenue-generating compute. Power distribution, cooling…
Waymo doubles spending on lobbying in robotaxi battle with Uber
Waymo has sharply increased its lobbying spending as it seeks to persuade US regulators to clear a path for fully autonomous taxi services, intensifying its battle with rival Uber over the future of robotaxis.
The Alphabet-owned company spent more than $1 million between April and June on lobbying the federal government, more than double its outlay a year earlier, according to filings. That put Waymo’s spending close to Uber’s and well ahead of rivals including Amazon’s Zoox and Tesla.
The rise in lobbying comes as Waymo and Uber push competing visions for the future of ride-hailing. Waymo wants a faster route to fully driverless commercial services, while Uber is advocating a staggered rollout in which robotaxis operate alongside human drivers.


© Alex Kent, Bloomberg
-
AI Infrastructure Archives - The New Stack
- Forget the model wars, Stripe and Ramp just started the router wars
Forget the model wars, Stripe and Ramp just started the router wars
I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.
Model triage is becoming one of the most important skills for the AI-native developer. I’ve argued all summer that the people getting the most out of frontier models are the ones disciplined enough not to run the best model by default. On Wednesday, Stripe and Ramp validated that idea 70 minutes apart: Stripe bought OpenRouter, and Ramp released its internal router.
Bloomberg puts the OpenRouter price tag above $7 billion, and Axios says it’s more than $8 billion in cash and stock. Stripe has not released the terms, so the details remain fuzzy.
While the acquisition made headlines, the architecture is the story. For the last couple of years, picking a model was something written into an application, a string in a config file, and swapping models took some work. A router changes that workflow. Stripe bought the layer, and Ramp built it. Both are betting their existing relationships give them a unique wedge to own this critical layer in the new AI stack.
Picking a model is becoming a runtime decision
OpenRouter is an endpoint serving more than 400 models from over 80 providers, processing more than 10 trillion tokens a day. Andrej Karpathy calls it the transfer switch for AI. Ramp’s Router does the same job on a smaller catalog and claims roughly 40% lower cost for the same output.
Both are betting the model name in a codebase is a liability. They’re mostly right. Back in June, I pointed to Mitchell Hashimoto, who found a standard coding task that cost about $1.50 on GPT-5.5 and roughly $9 on Claude Fable, with both producing equally acceptable results. A router automates that triage, making decisions on every request rather than only on those a developer explicitly configures.
This is becoming a large problem and a large opportunity. Our own Amanda Caswell reported this week that Anthropic’s /claude-api skill was burning about 200,000 tokens before answering a single question, and that loading its reference docs on demand instead of up front cut that to roughly 25,000. Hafiz Hassan wrote for us last week about why AI pipelines cost 10x more than the demo, and every culprit on his list is an engineering decision: system prompts resent every turn, whole conversation histories appended, oversized RAG chunks, raw JSON dumped into context. It’s a great practical guide, and none of the items Hassan identifies are procurement problems.
The token bill is generated by your code, which is why the tools to control it are arriving there as well.
Stripe and Ramp want the same layer for opposite reasons
Stripe is attacking the problem from the bottom up, through developers. The company’s investor letter, leaked Wednesday by Eric Newcomer, makes the argument directly: “Up until now, every developer has needed a straightforward and reliable way to manage their revenue pipeline, and serving this need gave rise to Stripe. Going forward, however, every developer will also need a straightforward and reliable way to manage their intelligence pipeline.”
Stripe wants to control AI spending through the long tail of developers. Ramp wants to control it through its existing relationship with finance.
Stripe has built this product before. Its payments business hides dozens of local payment methods behind a single API, routing each transaction to the payment method most likely to convert. The AI version is the same idea applied to models instead of payment networks.
Ramp is attacking it from the top down, through finance. The company bought the router.com domain and says its customers already buy quadrillions of tokens a month through Ramp. Founder Veeral Patel’s launch post pitches the service simply: “Monitor and control your AI bill across every provider.” Adam Wazzan sums up Ramp’s strategy better than I can: “when a CFO ships a product for CTOs.”
Stripe wants to control AI spending through the long tail of developers. Ramp wants to control it through its existing relationship with finance. Both are chasing what is rapidly becoming one of the largest line items in corporate technology budgets: tokens.
On X, Kabir Goel pushes back on Stripe’s framing. Routing tokens is a way to spend less, while Stripe’s other products are designed to help businesses make more.
“Stripe is just not where teams go to understand how much they’re spending,” he writes. “That’s pretty squarely Ramp territory.”
He has a point about where teams look today. Whether that’s still true three years from now is exactly what Stripe just spent billions betting against.
The router worth pointing at is the one with no model to sell
Whichever router you point at decides which model writes your code, and not every router is disinterested. Our own Paul Sawers flagged the problem in July when he covered the first wave of Cursor’s, Ramp’s, and Meta’s routers. Cursor backs Grok and Composer. Meta is building Muse Spark. Both have reasons to send work to their own models, and Paul quoted developer Elvis Saravia asking whether routing logic ought to be open source rather than a vendor’s private judgment call.
Stripe and Ramp do not sell models. OpenRouter CEO Alex Atallah says as much in Stripe’s own announcement: Developers “need a neutral layer to orchestrate and manage them all.” Investor Gavin Baker frames the opportunity the same way, arguing that Stripe can become the neutral infrastructure layer for AI, just as it became the neutral infrastructure layer for payments.
Neutrality isn’t free, though. Stripe takes a percentage of token spend, and Ramp wants your spending relationship, so “free through 2026” is a customer acquisition strategy with an expiration date.
The obvious objection is that vendor motives are the wrong thing to worry about, and routing quality is what really matters. That’s fair, and Towards Data Science published one of the best practical examples I’ve read. Pratik Rupareliya describes a routing layer that cut a support agent’s inference bill by 40% but also broke the product. A classifier sent “simple” queries to a cheaper model, but some of those “simple” queries were actually fraud investigations.
The cheaper model answered them confidently and incorrectly. Customers stopped using the agent, churn rose above baseline in month four, and retention costs were four to five times higher than the savings. It took three months to surface and another month to identify the cause. His fix was per-tier quality monitoring combined with an uncertainty-routed cascade, which ultimately settled at 35% savings without sacrificing quality. It’s an excellent article to read before diving into model routing.
So instrument the routing. Log what the router picks on every request and break out your quality metrics by the model that served them. Ramp Router reportedly records the model, provider, tier, tokens, latency, cost, and fallback attempts for every call. Stripe OpenRouter rankings have been a public version of that telemetry for years. You can get similar visibility with either approach.
Right now, the model is becoming an implementation detail. The competition is shifting to the layer that decides which model gets the job. Stripe and Ramp are betting that developers won’t care what sits behind the endpoint, so long as the bill is lower and the results are good enough.
The post Forget the model wars, Stripe and Ramp just started the router wars appeared first on The New Stack.
Raise Robotics unveils plans for shipyard coating robots
-
Towards Data Science
- Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions
Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions
AI systems should not automate a decision simply because they can provide a prediction. A decision system should consider how uncertain the prediction is and defer if a mistake would be costly.
The post Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions appeared first on Towards Data Science.
-
NVIDIA Technical Blog
- NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks. The challenge is how to build the agent architecture that makes frontier language models work reliably on extended…
Where Security Fits in an AI Agent Stack
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell, agent developers, open-source projects, and partners across the ecosystem, AI safety and security teams at NVIDIA offer their perspective on the emerging agent stack—including the role of each layer…