A recent GigaOm CxO Decision Brief explores how AI retrieval architectures are evolving beyond flat vector databases as organizations combine semantic search, ranking, personalization, and machine learning inference in production systems.
Vector search changed the AI infrastructure landscape by making semantic retrieval practical at scale. By converting text, images, and user behavior into embeddings, organizations could move beyond exact keyword matching and retrieve information based on meaning. But production AI systems rarely stop at vector similarity.
A real-world query often requires multiple signals to be evaluated simultaneously. Semantic relevance may be one factor, but so are structured attributes, business rules, personalization signals, freshness, access controls, recommendation logic, and machine-learned ranking models. As organizations move from AI experimentation to production-scale applications, the challenge is no longer simply finding similar items. It is in combining all of the signals that matter while maintaining low latency and operational simplicity. This is where tensors are attracting increasing attention.
While vectors represent information as a single dimension of numerical values, tensors provide a more general framework for representing and operating on complex, multi-dimensional data structures. They offer more control in how relevance is computed, allowing dense embeddings, sparse features, metadata, and model outputs to be evaluated together within a unified retrieval and ranking process. For organizations building large-scale retrieval systems, this raises an important architectural question: is a flat vector store sufficient, or does the next generation of AI applications require something more expressive?
“Tensors provide a more general framework for representing and operating on complex, multi-dimensional data structures.”
A new GigaOm CxO Decision Brief, “The Tensor Advantage in AI Search,” explores this question in depth.
Among the findings:
Production AI systems increasingly depend on combining semantic, lexical, behavioral, and business signals rather than relying on vector similarity alone.
Architectural fragmentation between vector databases, search engines, rerankers, and feature stores introduces latency, operational complexity, and synchronization challenges that become more significant as workloads scale.
Emerging retrieval models, including multi-vector and late-interaction approaches, place new demands on infrastructure that were not anticipated when first-generation vector databases were designed.
Tensor-native architectures provide an alternative approach by treating multidimensional data structures as first-class citizens rather than forcing them into simpler vector abstractions.
The paper also examines the infrastructure, operational, and organizational implications of these architectural choices, including benchmark data, deployment considerations, and the trade-offs engineering leaders should evaluate when planning future AI retrieval systems.
“Retrieval is evolving from a nearest-neighbor problem into a ranking and decision-making problem.”
As AI applications become more sophisticated, retrieval is evolving from a nearest-neighbor problem into a ranking and decision-making problem. Understanding the role tensors play in that transition may be one of the most important architectural discussions facing engineering leaders today.
Every enterprise software vendor is currently selling some version of the same thing: AI agents grounded in enterprise context and governed by a central control plane. SAP, ServiceNow, Salesforce — they all have one.
At its ONE conference in Amsterdam in June, OutSystems unveiled its version, and its CEO, Woodson Martin, agrees that they all look quite similar on the surface, but unsurprisingly, he also believes that OutSystems has a very different approach.
Martin tells The New Stack that, in his view, “It would be very easy to just look at the market today and say all enterprise software players are offering exactly the same thing.” The reason for that, he says, is that enterprise agent orchestration, “is sort of greenfield today. Everybody’s aiming for it. Everybody’s got a great story about why they’ll be a leader or a player.”
For OutSystems, the story is neutrality. While SAP and Salesforce pitch agent orchestration from within their ecosystems, where they are also the systems of record, the 25-year-old former low-code company, which now describes itself as an agentic systems platform, wants to be the layer that coordinates across all of them without owning the underlying data.
OutSystem’s agent platform. Credit: The New Stack
The advantage of not being a system of record
Martin says the company has been playing some version of this for a long time. “We’re glue between commercial off-the-shelf solutions,” he says. “We’re the thing that makes the enterprise their own enterprise, as opposed to an SAP enterprise or a Salesforce enterprise.”
One asset management customer, he says, has used OutSystems as the orchestration engine across about 80 systems for fund onboarding for the past six or seven years. That wasn’t for any agentic systems yet, of course, but OutSystems’ role in this isn’t all that different. “We’re already playing that orchestrator role,” Martin says. “In other cases, we don’t have that position yet in the account, and we’ll have to fight for it.”
Tiago Azevedo, OutSystems’ CIO, makes the same argument. “We are agnostic to all of those things.” The OutSystems platform doesn’t create most of the data it touches, he notes. Instead, its focus was always on integrating existing systems. “Our happy place is when we bring several of those systems all together into a form that makes sense for a process,” he says.
Credit: The New Stack.
Open to Claude, Codex, and Kiro
At the ONE conference, the company launched the OutSystems Agent Experience, a platform layer that exposes Model Context Protocol (MCP) and Agent2Agent (A2A) services. Developers can now build, publish, and extend OutSystems applications with third-party coding tools like Claude Code, Codex, Cursor, and Kiro, AWS’s spec-centric IDE.
The first of those services is now live on the OutSystems Developer Cloud (ODC), the cloud-native, current-generation platform. Support for OutSystems 11, the older self-managed platform where a large part of the installed base still runs, launched in early access
“Existing customers that live on O11, they’re like: what I would love to do was to be able to use Claude or Codex or whatever to evolve my applications in O11,” Azevedo says. “So we made that possible. “[…] We made a lot of people happy.”
And indeed, when this was announced in the keynote, it drew more applause than some of the other large product announcements.
He sees no alternative to opening up the platform, given how freely developers now move between coding tools. “I strongly believe in open systems,” he says. “You close those environments, you’re gone.”
Other launches at the conference include the Agentic Enterprise Orchestration service and the next-generation OutSystems Agent Workbench, which is now generally available and adds agent evaluations, guardrails, semantic search, and Amazon Bedrock support.
There is also a preview launch of a new modernization service, built on AWS Transform and Kiro, for migrating COBOL and Lotus Notes systems onto the platform, as well as a pre-packaged agentic solution for loan origination, the first of a family of packaged agentic industry solutions, which arrives later this year.
The new bane of IT departments: shadow AI
Shadow IT has returned as shadow AI, Azevedo says. And while he has always managed to stay ahead of internal technology demands with previous platform shifts, that’s getting harder now. “With AI it’s impossible,” he says. “It’s literally impossible. It’s not humanly possible.”
The way he describes it, a central team can build maybe 10 large agentic workflows that solve company-scale problems. “Those are what we call the big bets,” he says, and these bets exhaust the team’s capacity. Everything else means letting the rest of the organization build its own agents, and every one of those requests immediately raises questions of who gets to touch which company data and through which MCP servers. Demand from every department, he says, is growing almost exponentially.
The token bill
Unsurprisingly, this is now also coupled with the question of how much all these tokens cost.
“I also have my CFO saying, what about the token usage? And what about the budget? And who’s gonna pay for that?” he says. “If you look at, let’s say, 500 euros or dollars a month, times 12, times, let’s say, 1,200 or 1,500 people, these are millions in a year.”
For now, at OutSystems, he rations token budgets almost by hand, protecting projects that could scale and trimming those that serve an audience of one.
“It’s the most expensive software — and I managed big contracts,” says Azevedo. “The most expensive I’ve ever had in my hands.”
“It turns out my number one consumer of tokens on Anthropic in the month of April was a business value consultant in Australia … why is he burning $7,500 in tokens every week? That’s not in the budget.”
Martin tells a similar story from the CEO perspective. “It turns out my number one consumer of tokens on Anthropic in the month of April was a business value consultant in Australia,” he says. “And we’re like, why is he burning $7,500 in tokens every week? That’s not in the budget.”
Martin traces the shift to the latest generation of reasoning-heavy models, which arrived around January and February, “and we’re getting these token bills in March and April that are starting to scare everyone.”
For the overall OutSystems platform, the answer to this is model flexibility. Customers can bring their own models, swap them without touching the agent logic, and route requests through Amazon Bedrock to whatever is the cheapest option to do the job effectively. Some customers built model routers on OutSystems early on, Martin says, sending complex asks to expensive models and simple ones to cheaper ones.
He also argues that OutSystems’ Enterprise Context Graph means that reasoning over an enterprise context graph is less token-intensive than reasoning over an application’s raw codebase.
“I think organizations are going to develop more discipline around this as we have for other elements of material spend in any enterprise,” Martin says. “This is now becoming material for almost everyone.”
OutSystems’ advantage may indeed be that it isn’t a system of record but ties into all of them. As those systems of record open up with new MCP tools, the former low-code platform may just be in the right position — and with the right customer base — to help its customers tie all of these together, even as those platforms launch their own agent builders and orchestration platforms. Sometimes you may want your agent to be close to the data, but over time, nobody wants to manage half a dozen agent orchestration platforms either.
Beneath the chatbots and copilots, there’s a quiet revolution happening in the data services space. From pure-play database vendors to data integration wranglers and onward to the cloud hyperscalers, the focus has shifted.
Now in the spotlight is the question of how to automate data governance for agentic AI workloads, and for good reason: Traditional manual data stewardship doesn’t scale in a world where agents are becoming increasingly autonomous (and powerful).
Aiming to cut a swath in this marketplace is data control plane company lakeFS. The organization announced its lakeFS for Agentic AI service on Wednesday, and it appears to be designed to bring governed, reproducible data access to autonomous and headless agentic workloads (those that execute decisions below the user interface level) that run at enterprise scale.
The manual model breaks
Einat Orr, CEO and co-founder of lakeFS, tells The New Stack that manual data stewardship was built for human-paced, human-reviewed workflows, i.e., someone looking at a change before it is committed.
“When dozens or hundreds of agents are making changes simultaneously, faster than any person can review, the manual model breaks,” Orr says. “This is because with a human analyst, a bad write to production is usually one mistake, caught by another human before it spreads far. An agent is different — it acts automatically, in parallel, at machine speed, and it doesn’t pause to second-guess itself. And because so much agent activity is unsupervised, you often find out after the damage is done.”
She explains that attempts to identify and roll back incorrect or corrupted production data across a wide set of data modalities, such as images, documents, metadata, and structured data, are almost impossible to pull off. Impossible, that is, unless the team has the data infrastructure in place to isolate and track such changes automatically.
While some of the more disastrous outcomes stay inside an organizaton’s perimeter (or are swept beneath the communications radar), Orr explains that real world consequences of bad agentic data writes are manifold.
“Insurance claims get inappropriately denied or approved, sensor data from machines gets misinterpreted, an incorrect medical diagnosis is made, or customer service bots provide incorrect answers to customers,” Orr says. “The cost of an individual action may be manageable, but agents performing these actions hundreds or thousands of times can have an exponentially larger impact.”
“As agents are let loose on enterprise data at a massive scale, any agent that reads or writes to production data without isolation or a reproducible trail is a liability, no matter how good the model is,” —Einat Orr, lakeFS CEO.
Bad agents acting in the real world
Examples of this happening include the July 2025 Replit AI coding agent incident, which deleted a live production database during an explicit code freeze, wiping records for more than 1,200 executives and around 1,200 companies. To tidy up its handiwork, the agent then fabricated thousands of fake records and initially claimed the deletion couldn’t be rolled back.
Also in July 2025, Google’s Gemini CLI agent misread a single failed command, acted on a version of the file system that existed only in its own interpretation of the scenario, and permanently destroyed a user’s project files. The Gemini agent is widely reported to have said of its actions: “I have failed you completely and catastrophically. My review of the commands confirms my gross incompetence.”
“The pattern in both is the same: An autonomous agent took a destructive action that no one authorized, and the lack of isolation and a reliable rollback path turned a single mistake into permanent loss,” Orr says.
A doctor of mathematics with a track record in hardcore software engineering, the bottom line for Orr is clear: “As agents are let loose on enterprise data at a massive scale, any agent that reads or writes to production data without isolation or a reproducible trail is a liability, no matter how good the model is,” she said.
“…any agent that reads or writes to production data without isolation or a reproducible trail is a liability…”
Gartner expects 40 percent of enterprise applications to have task-specific agents embedded by the end of 2026, up from less than 5 percent a year earlier. IDC projects that agent use at the largest enterprises will grow tenfold by 2027, with the API and data calls those agents make growing a thousandfold. That’s the scale production data has to withstand, and it’s what lakeFS is built to govern.
Agents sent to play in an isolated data sandbox
To address these issues, lakeFS for Agentic AI gives every agent its own isolated data sandbox with a “zero-copy” branch of relevant data, so the agent can access the dataset it needs via references, snapshots, or copy-on-write techniques.
This means any changes the agent wishes to make must be validated and merged in accordance with the policy guidelines defined by the system architecture. In turn, this produces a unified audit trail across every agent action.
When running, lakeFS for Agentic AI is powered by its data version control architecture, which provides zero-copy data sandboxing. This enables isolation so that agent mistakes are automatically isolated and never corrupt production data. Every agent run is tied to an exact, immutable version of the data. Past actions can be recreated, debugged, audited, or extended using the same inputs.
Production data is gated by policy. Merges into production happen only after pre-merge validations pass. Every change can carry an agent identity, a run ID, and an execution context. The result is a unified audit trail instead of evidence scattered across orchestrators, model providers, and cloud logs.
Agents confined by branch-scoped credentials
Where agents are permitted to read and write through standard file operations. lakeFS provides file-level data access with branch-scoped credentials. These can be described as strictly cryptographically bounded, ephemeral access tokens that confine an agent to a specific branch of data or code, so that the agent operates only within its own workspace. This whole mechanism keeps each agent’s working set narrow and avoids context bloat.
“With lakeFS Mount, a branch, or even a subset of a branch, can be mounted as a local directory inside the sandbox or virtual machine where the agent is running,” Orr confirms. “From the agent’s perspective, it’s just reading and writing to files and folders.” She further clarifies and notes that no LLM tokens are spent learning the lakeFS API. The agent works with a familiar filesystem interface, and lakeFS handles the versioning underneath.
Developers also have a couple of options for injecting custom validation logic. CEO Orr explains that software engineers can use webhooks or Lua scripts, both of which allow users to define behavior and rules that must be met before a merge can proceed.
“Beyond automated checks, lakeFS also supports pull requests, which bring a human into the loop. In agentic workflows, this gives you a way to review and approve what an agent is proposing before it reaches production,” she clarifies.
Who else builds “Git for data” services?
Clearly, other vendors and projects exist in the data versioning market.
Apache Iceberg has functions for branching and tagging data. HPE acquired Pachyderm back in 2023 for its data versioning and pipelines technologies, which serve MLOps teams.
Originally developed by Dremio, Project Nessie is now an open-source data catalog and version control system for data lakes. Data Version Control (DVC) is an open-source data version control infrastructure designed for complex AI operations and big data environments, but now we’ve come full circle as lakeFS acquired the project in late 2025.
In the search for governance automation for agentic AI workloads, lakeFS appears to offer a comprehensive, cohesive set of tools and functions. In the “Git for data” marketplace, a variety of options exist, but lakeFS hasn’t explicitly positioned itself as a carte blanche replacement for similar or related tools.
One thing is certain: The questions of who is feeding what data to which agentic function, when, where, and why are becoming an increasingly pressing issue if we want AI to work correctly.
It’s no secret that generative AI has shifted the operations and business models of companies in nearly every sector. But what if I were to tell you that one day, very soon, we will view these innovations the same way smartphone owners look back on the feature phones of the nineties: early experiments in a journey towards a much more significant technological transformation?
The fact is that AI is advancing faster than any technology in human history; faster than even its creators expected. Today, with many businesses still getting to grips with generative AI, the lens is already moving on to the next world-changing iteration of the technology: agentic AI.
The opportunity of agentic AI for Europe’s enterprises
Already estimated at around $9.14 billion, the agentic AI market is forecast to grow rapidly, at a compound annual growth rate (CAGR) of 40.5%, to reach $139.19 billion by 2034. Europe will be at the forefront of this boom, growing at a 42% CAGR.
The impact of all this investment on enterprises will be profoundly beneficial. According to one study, agentic AI could generate up to $450 billion in economic value through revenue growth and cost savings by 2028.
Today’s stand-alone AI models are rapidly being replaced by complex, multi-agent-driven automation, in which enterprise systems will be able to make decisions and act on them entirely autonomously.
As this happens, today’s stand-alone AI models are rapidly being replaced by complex, multi-agent-driven automation, in which enterprise systems will be able to make decisions and act on them entirely autonomously. I’m not saying that businesses should abandon their current generative AI efforts. They need to start putting the foundations in place for the agentic future now.
Ensuring cloud infrastructure is agentic-ready
Vultr CMO Kevin Cochrane
Agentic AI is a very different proposition to generative AI and demands a different approach within the data center. Rather than serving discrete models, agentic AI infrastructure will need to orchestrate multiple autonomous systems as they interact with one another and with human users.
From a technical perspective, that means balancing both high-performance cloud GPUs and CPUs in an end-to-end AI-optimised stack.
GPUs will be required to run massive LLMs, process data, and generate outputs, while CPUs will need to orchestrate agents and execute the tools and policies that enable them to be autonomous. Modern cloud infrastructure needs to balance these technologies to enable agentic applications to deliver on their considerable promise fully.
However, for European businesses, deploying agentic AI-ready cloud infrastructure comes down to much more than technical capabilities alone. To be fit for purpose, the European cloud offerings of tomorrow will need to be completely different from those currently in play. There must be a decisive break from the past.
Data sovereignty and the need for regional AI infrastructure
Currently around two-thirds of European cloud services are provided by US hyperscalers. This situation is a legacy of a vanished world with a different set of geopolitical, economic, and regulatory conditions to our own. It harks back to a time when storing the data of European citizens and businesses overseas was much less problematic, and when cloud costs were more manageable.
Today, EU businesses need to balance agentic innovation with a renewed focus on regulatory compliance. This is because the EU is mandating measurable standards for data localization, operational control, and legal jurisdiction to address concerns about foreign surveillance risks, extraterritorial jurisdiction, and dependence on a small number of hyperscale providers. Once abstract concepts of data sovereignty have transformed into strict operating principles for European businesses. Cloud infrastructure must be stored locally and in the right operational jurisdiction to avoid foreign data subpoenas.
Addressing exploding cloud costs
A second consideration is cost. Data from Flexera shows that 81% of businesses still cite cost efficiency as their top metric for assessing progress against their cloud goals. Yet according to its research, 76% of large enterprises spend more than $5 million on the cloud each month, and nearly a third complain of wasted cloud spend.
Part of the problem is that enterprises are locked into hyperscaler contracts characterized by opaque pricing and forced service bundling. At just the moment they need to invest in agentic AI infrastructure, enterprises are struggling to fund their core cloud workloads, a situation that’s not helped by the skyrocketing price of CPUs.
As European businesses set out on their agentic AI journeys, the case for moving beyond the hyperscalers could not be more compelling.
Rooting cloud infrastructure in Europe
This is where alternative hyperscalers like Vultr come into their own. From a data sovereignty perspective, we operate nine European cloud data center regions, including Amsterdam, Frankfurt, London, Madrid, Manchester, Paris, Stockholm, Warsaw, and, as of 19 May, Milan (with this launch, we now operate 33 global cloud data center regions).
These are physically isolated data centers that come with geo-fenced data management policies and guarantees that no data will be transferred or processed outside jurisdictional boundaries without explicit consent.
As well as helping European businesses comply with data sovereignty mandates, Vultr helps them avoid the high costs and systemic lock-in associated with traditional hyperscalers. With Vultr’s full-stack AI infrastructure, developers and enterprises can benefit from Vultr’s flagship CPU offering, VX1, which offers 23% better performance and 33% lower cost than comparable hyperscaler compute plans, resulting in up to 82% better price-to-performance.
In addition, Vultr offers European businesses with access to the latest AMD and NVIDIA GPUs for AI and machine learning, high-performance computing, and more, available on demand either as virtual machines, bare metal, or self-service clusters. This is the complete, end-to-end stack of next-generation compute power that businesses will need to thrive in the era of agentic AI.
The dominance of the hyperscalers is less certain than ever as data sovereignty mandates make the case for Europe-based infrastructure.
It’s an exciting time to be in the cloud infrastructure business. The dominance of the hyperscalers is less certain than ever as data sovereignty mandates make the case for Europe-based infrastructure. Meanwhile, businesses are pushing back against the high costs and systemic lock-in that come with hyperscaler offerings. In their place, enterprises can invest in high-performance, cost-effective compute infrastructure that’s open, flexible, and ready for the demands of agentic workloads.
About a year ago, Google demoed a diffusion model at its I/O developer conference, but went quiet about the technology soon after.
On Wednesday, however, Google broke that silence with the launch of DiffusionGemma, an experimental 26B mixture-of-experts model that uses diffusion to generate text 4x faster than its existing Gemma models.
Diffusion has long been the standard for generating images (think Stable Diffusion). Instead of generating one word at a time, models like DiffusionGemma or Inception’s Mercury 2 generate words in parallel.
At first, those blocks of text don’t make sense and seem random. But then, with each new step, the model refines the text and reduces the noise until it becomes the answer you were looking for. If you’ve ever looked at a diffusion image model generate images in real-time, that’s essentially the same process, but for text.
Credit: Google
With each step, the model denoises 256 tokens in parallel, which is why it can be much faster than a traditional autoregressive large language model. It basically iterates on the text with each step until it.
All of these tokens attend to all others, which Google says is especially helpful for use cases such as inline editing, code infilling, working with amino acid sequences, and mathematical graphs.
Credit: Google
Google says DiffusionGemma can produce more than 1,000 tokens per second on a single Nvidia H100. And since the model uses the mixture-of-experts technique, it doesn’t have to keep the full 26 billion parameters in memory; instead, it activates only 3.8 billion during inference. This means it can easily run on a GPU with 18GB of VRAM.
There are some tradeoffs, though. On all benchmarks, the DiffusionGemma model underperforms when compared to Gemma 4 26B A4B. That’s something Google itself acknowledges. There’s no technical reason why a diffusion model couldn’t perform just as well as a more traditional large language model, but the focus here is on speed.
“For applications that demand maximum quality, we recommend deploying standard Gemma 4,” Google says in its announcement.
Credit: Google
Availability
The model is now available on HuggingFace, with Unsloth and other quantizations available for those who want to run it locally using llama.cpp and (soon) similar local inference tools.
Google also worked with Nvidia to optimize the model for its hardware, including high-end GPUs like the GeForce RTX 5090 and 4090, as well as the Nvidia DGX Spark and DGX Station (for those who can afford them). Nvidia NIMs are also available for the model.
“Keep readers reading” is the not-so-simple goal of Medium’s recommendations system. To predict what’s most likely to appeal to a particular reader at any given time, Medium continuously processes user activity signals (stories read, recommendations shown, follows, likes, etc.). It then immediately correlates that with the steady stream of new articles, which is estimated at millions per month.
Smart models and good inference logic are required, but that’s not enough. The data must be stored and retrieved quickly enough to remain relevant while the user is browsing. That’s the job of Medium’s feature store. And getting the data model right started to matter a lot as they scaled to 1M operations per second.
Andréas Saudemont, Medium Principal Software Engineer, recently walked through how the team identified the problem and what they built to fix it. If you’d rather watch than read, you have two options: Watch a short version from Monster Scale Summit or an extended follow-up webinar
The feature store and its role in Medium’s recommendation system
The feature store ties it all together, ingesting user activity and internal events and feeding them to the ML models that power recommendations. It’s what enables customization like the “For You” feed that greets logged-in users.
Each feature is a property of an entity, usually a user or a story. Some are simple and static, like whether a user holds a paid membership. Others capture interaction history: which stories a user has read, what content they’ve recently been shown, etc.
The following diagram shows a highly simplified view of the Medium feature store architecture:
The problem with a relational features data model
When they built their feature store years ago, Medium used relational features for cross-entity relationships. Unlike regular features, a relational feature can have multiple values for a given entity ID. Each value is defined by a relation ID (the ID of the related entity) and a timestamp recording when the event occurred.
For example, a “story users have read” feature is attached to the story entity type. It relates to the user entity type, and its values indicate whether/when a given user has read that story.
Andréas shared the following schema diagram to explain the concept:
Features sit at the center, each attached to an entity type and defined by name, version, and data type. Non-relational features are simply a feature, an entity ID, and a value. Relational features add a relation ID mapping to another entity type, plus the value itself and a timestamp.
This approach proved suboptimal from a data modeling perspective. Since relational features link two entity types, the data ends up split between two tables: one for the entity IDs and one for the values. That means you can’t get both in a single query. The first query retrieves only entity IDs (not their associated values) and relies on ALLOW FILTERING. A second query then runs for each entity ID to fetch its value. “If we have 1000 entity IDs for which we want to fetch values, then we have to run 1000 queries to fetch these values,” Andréas said.
Overrelying on ALLOW FILTERING made things worse. “This is bad,” Andréas said, referring to monitoring data showing that 90% of rows read via these queries were simply discarded. “This is just data that we don’t need. ALLOW_FILTERING should be an escape hatch, not our design pattern.”
“ALLOW_FILTERING should be an escape hatch, not our design pattern.”
The list feature model
So they reinvented their data model and shifted to a list-based feature model. Instead of splitting data across two tables, everything for a given entity lives in one place and is retrieved in a single query.
Like other features, a list feature is defined by its entity type, name, and optional version. What’s different is the value. While a non-relational feature has a single value, such as true or false, a list feature’s value is a collection of items, each containing a value and a timestamp. Item values can be of any data type; the feature store doesn’t enforce consistency within a list.
For example, consider a user’s reading history. The entity is user, the feature name is reading history, the TTL is 6 months. After that TTL is reached, the data is automatically dropped by the database (since older history isn’t useful for recommendations). The list for a given user is a collection of story IDs and the timestamps at which they were read. The same story can appear multiple times, and multiple items can share the same timestamp.
A range of operations need to be supported. Create List and Delete List operations run at most a few times per day. Remove List Items with Value, which lets a reader scrub a specific story from their history so it stops influencing recommendations, runs at 1k-10k per second. Add List Items is higher still: every story read and every thumbnail shown to a user generates an event. Get List Items is the top, at 100k-1M operations per second.
“The Add List Items, and even more the Get List Items operations, are really the reasons why we need an efficient data store.”
“The Add List Items, and even more the Get List Items operations, are really the reasons why we need an efficient data store,” Andréas said.
Multiple items, one timestamp
Beyond raw efficiency, the new data model also had to support multiple items with the same timestamp. When Medium shows a user four story thumbnails simultaneously, all four presentation events share the same timestamp, but have distinct story IDs. If this isn’t handled correctly, primary key collisions occur.
The team’s solution was a single list_items table that stores everything.
The partition key combines feature_key and entity_id, keeping all items for a given list together. All of user 123’s reading history is stored in one partition, retrieved in one query. The clustering key concatenates each item’s timestamp with an MD5 hash of its value. The hash is what makes same-timestamp items with distinct values possible.
Relying on MD5 hashes for uniqueness raises its own set of questions, but in practice, the team hasn’t seen collisions. “The values that we are storing are sufficiently distinct, especially when you add the timestamp into the equation,” Andréas said. The table’s clustering order is set to descending so ScyllaDB can optimize for the typical read pattern (most recent N items) rather than leaving the application to sort afterward.
TTL to control storage costs
Storage cost is controlled entirely through ScyllaDB’s native TTL, with no cleanup logic required. Every row expires automatically based on its own timestamp plus the feature’s TTL duration. “We don’t have anything to do regarding that,” Andréas said. “Any row for which the TTL is expired will be considered deleted by ScyllaDB.”
Storage plateaus for a steady write rate. When a feature is retired, its data drains away on its own. “That’s super useful for controlling our storage and usage costs.”
Implementing the list operations
Add List Items is a logged batch of INSERTs with atomicity guaranteed: all items land or none do. Each row carries its own TTL calculated from its timestamp, so older items expire sooner. Since items almost always carry a current timestamp, new entries append to the top of the partition, which is exactly where reads will look first.
Get List Items runs as a single-partition SELECT with a minimum timestamp and a row limit. “We run the query on a single partition,” Andréas said. “That’s the maximum efficiency that we can have.” The clustering key handles filtering and ordering directly. Post-processing is not required.
Remove List Items with Value is the one operation that couldn’t be reduced to a single query. Because value isn’t part of the primary key, a direct filter isn’t feasible.
A local secondary index built specifically for this case first finds the matching item keys, then a batch DELETE removes them by primary keys.
“Using an index is really faster than a scan because the query is highly selective,” Andréas explained. “We have very few items in a given list that have the same values compared to the total number of items in a list. And thanks to the current structure, using a local secondary index is faster than a global index.”
Andréas shared another example. Starting with the original table partition, the goal is to delete all items with the value “storyC.” Using the local secondary index, the system first identifies the two rows containing that value. It then issues two DELETE statements using the item keys from those rows, which removes them from the list. The final operation, removing all list items, is even more straightforward.
“We can just drop the partition,” Andréas said, “and ScyllaDB does its magic. It just deletes all the rows for that partition, which means that it deletes all the items for the given list. And bonus point: it’s atomic. It’s either completing successfully or not changing anything at all.
ScyllaDB vs. DynamoDB performance
Medium implemented the list operations on top of both ScyllaDB and DynamoDB. The main goal was to benchmark how both databases compared on their actual production data. “Conceptually they are very close,” Andréas noted, “but they have significant differences in how they operate.”
For AddListItems, P50 latencies were low with both databases: ScyllaDB came in under 1.5ms, DynamoDB under 5ms. “DynamoDB is extremely fast, not as fast as ScyllaDB, but extremely fast at sub 5ms latency,” Andréas commented. Things got more interesting at the P95 and P99 latencies. ScyllaDB held steady at around 5-6 ms P95s, while DynamoDB ranged from 13-45 ms. ScyllaDB’s P99s were steady single-digit milliseconds, while DynamoDB’s ranged from 40- 120 ms.
AddListItem latencies: The blue line is DynamoDB; the purple line is ScyllaDB
It was a similar story for GetListItems. At P50, ScyllaDB clocked in at 1 ms, DynamoDB at around 3.5 ms. At P95, ScyllaDB held around 5-6 ms while DynamoDB spiked from 30 – 60ms. And at P99, ScyllaDB remained at ~30ms while DynamoDB ranged from 70 ms all the way up to 220 ms.
GetListItem latencies: The top blue line is DynamoDB; the lower purple line is ScyllaDB
“ScyllaDB is very fast, with very predictable performance, and that’s super important for us.”
One caveat: DynamoDB was running without an extra caching layer. “We expect that could have a significant impact for DynamoDB because of the high cache hit rate that we are seeing on the list,” Andréas said. “But we don’t have the data yet, so we cannot compare them.” His verdict for now: “ScyllaDB is very fast, with very predictable performance, and that’s super important for us.”
Key takeaways
One pleasant side effect of getting the data model right: Medium is now eager to use ScyllaDB for additional feature store workloads. Before, they were holding back because they didn’t want to build on the shaky relational feature foundation.
Reflecting on the path to this point, Andréas left the audience with this parting advice:
“If you have a suboptimal data model, you will have queries that are slow, that will scale badly. And most likely, you won’t be able to optimize that data model. You will have to define a new data model that will be better. So take time to think about your data model before you start the implementation, because once you have production data using your suboptimal data model, it’s too late.”
Memory has replaced compute as a primary constraint for modern tech teams. A perfect storm of hardware architecture limitations, semiconductor supply chain uncertainty, and changing software licensing models has left enterprises confronting increasingly memory-constrained environments. All while high-bandwidth AI workloads overload the production chain’s ability to provide sufficient memory — and when your AI token bill is starting to cost more than your salary bill.
All this adds up to a demand for enterprises to shift from the previous buy-all-you-can mindset to a data-driven optimization strategy.
Fortunately, AI isn’t only part of the problem. When AI is applied to memory economics in modern virtualization, it becomes a vital part of the solution.
Bharath Ram, director of product management at Hewlett-Packard Enterprise (HPE), explains it to The New Stack this way: “There’s a component shortage. Today the prices have increased. So customers are looking at ways to save and optimize the existing footprint, so that they can run their workloads on whatever and not have to procure anything new.”
By switching focus from gobbling up every bit of memory your organization can grab to optimizing your workloads and their placement across multi-cloud and hybrid-cloud environments, enterprises cannot only speed up modernization but also shorten decision-making cycle time by up to 80%, all while cutting costs by up to 50%. Read on for how to transition your enterprise from guesswork to smarter IT.
Tech has to confront its waste problem
Most enterprises operate with significant over-provisioning driven by limited visibility and risk avoidance.
These same companies are standing up legacy applications that become less efficient over time, including a significant number of zombie services running without any use.
On top of this, most AI workloads rely on advanced memory technologies, which has led chip manufacturers to shift production priorities from DDR4 to DDR5 RAM, further reducing DDR4 supply. This impacts the whole industry, with even a personal computer costing 15% to 30% more than last year.
Add to this volatile DRAM pricing and higher core densities, and it’s clear that even non-technical leadership is worried about memory efficiency. Tech giants Microsoft, Google, Amazon, and Meta are buying up as many AI chips as they can, which is triggering even more shortages and price pressure across the supply chain. And thus more enterprises are hoarding more infrastructure and memory.
And this overbuying isn’t limited to memory. Companies are now also buying infrastructure like servers even before they need them, too, Ram reflects, “because the cost is so exorbitant, the quote that you might have today might not be the same price that you’re quoted for the same infrastructure tomorrow. That’s how we’re seeing the market right now. It’s very volatile.”
But it might all be ok. HPE estimates that between 20% and 40% of infrastructure is overprovisioned today. Which is an opportunity for efficiency — not only in these limited resources but also in faster, more secure workloads.
Enterprises are more capable than ever to optimize the use of what they’ve got today, especially before they go searching for more RAM that will cost significantly more.
It all starts with understanding
So much of this waste persists because enterprise infrastructure is obscured — no one really knows what does what with which data, or which services rely on it.
The same thing that holds companies back from doing anything more than lift-and-shift to the cloud is usually what keeps them from unlocking memory efficiency. There’s simply too little visibility across most enterprises’ complex, hybrid and multi-cloud distributed systems. Which has left organizations guessing and then rounding way up for over a decade now.
“It’s a combination of over-provisioning and not understanding underlying infrastructure. Because many of them are doing public cloud-based provisioning and self-service, where you don’t know what the underlying infrastructure is and you have admins leveraging whatever there is in terms of their service capabilities,” Ram explains. “One piece is memory shortage, and the other is understanding what’s been deployed and rightsizing it.”
To break these over-provisioning bad habits, any change has to be grounded in reality. The first step is to gather and analyze real usage data, using a tool like HPE CloudPhysics to establish a factual baseline that separates real cost drivers from those years of assumptions.
This allows enterprises to:
Understand their virtualization footprint and licensing exposure.
See their workload initialization and efficiency.
Identify true cost drivers before taking action.
You cannot right-size until you have real-time monitoring of how many hosts have how many VMs, and which are on and off.
Predictive, not reactive provisioning
Once an enterprise has a single source of truth for its complex distributed systems, it can explore what to deploy, where, when, and how.
“An application like SAP HANA is highly memory-intensive and highly latency-intensive. It’s not like this algorithm is optimized to pivot between hot and cold memory tiering” for cost reduction, Ram explains, without risking the application performance, akin to how, when older PCs had limited amounts of memory and, once that ran out, the computer would swap the program from running in memory to disk, slowing way down.
Part of the modern solution, Ram argues, is that companies “can over-provision with what they already have. They don’t have to buy any new memory,” because of better shared resources available to all the virtual machines managed by a single host.
“For example, a host with 64GB of physical memory may have more memory allocated across VMs than physically available,” he explains. “In practice, not all VMs consume their full allocation simultaneously, allowing unused capacity to be dynamically reassigned where needed.”
Memory ballooning, which, Ram says, is nothing new, but something desperately needed in the market right now. Version 9.0 of Morpheus, due out this summer, will feature a more modern sort of memory oversubscription, which, HPE explains, allows administrators to oversubscribe physical memory across VMs on a host, enabling higher VM density and more efficient use. This is particularly useful for testing and development environments, virtual desktop infrastructure, and workloads with variable memory demands.
Shift to architectural efficiency
Eventually, once you’ve optimized and rightsized every memory allocation, it’s time to shift your workloads to a new platform to improve hardware efficiency.
“The final step is increasing workload density per server, especially as per-core software licensing becomes more expensive,” Ram explains. He says this is best achieved using a virtualization solution with an open-source hypervisor, which can improve utilization now and help organizations shift toward per-socket licensing models. For suitable applications, that modernization may also include moving to containerized deployment, while in-memory deduplication reduces redundant data structures in RAM, improving memory efficiency.
“Not everybody can keep running on existing hardware forever. At some point, some organizations will need to move to a new platform to improve hardware efficiency. But higher workload density still brings added benefits,” he continues, especially at a time when even the biggest tech companies are overbuying infrastructure, driving up costs and tightening capacity.
Per-socket licensing saves more
And if you do go for a hardware refresh with HPE’s Morpheus Software, then you can unlock a different kind of subscription model, which charges per socket or CPU licensing, where multiple cores can share one socket. Some early results indicate that this can deliver up to 90% in savings.
GitHub hasn’t had an easy year. The platform has been hit by repeated outages affecting core services — including the Actions-based CI/CD pipelines that engineering teams depend on daily — and has had to issue public apologies as a result.
The scale of the problem is staggering: Where GitHub handled roughly 1 billion commits across the whole of 2025, it now processes 1.4 billion every month, with AI agents alone responsible for more than 17 million pull requests in the same period. GitHub’s COO Kyle Daigle told The New Stack in early June that the company is now targeting capacity to handle 30 times its current load — a challenge he described as far beyond the normal playbook of adding more machines.
Against that backdrop, Microsoft has chosen this moment to make its most direct push yet to push enterprise customers off Azure Repos — its own Git-based source code platform, which has existed in various forms since 2013, predating Microsoft’s $7.5 billion acquisition of GitHub in 2018 — and onto GitHub.
The exit ramp
The tool Microsoft is using to make that case is Enterprise Live Migrations (ELM), currently in limited public preview. The core problem it solves is downtime: previously, moving large repositories from x to x could take days, leaving teams frozen out of active development.
In a blog post authored by Soo Stahl, principal product manager at Azure DevOps, and product manager Bhuvan Shah, the pair explain that ELM works by keeping the source and destination repositories in sync while developers continue working in Azure Repos, with a final switchover window that they say typically takes under 30 minutes.
“Teams can migrate at their own pace, without coordinating complex, high-risk ‘all-at-once’ migrations.”
“This means no extended freeze periods, no multi-day outages – just a controlled, predictable transition that fits into your operations,” they write. “Teams can migrate at their own pace, without coordinating complex, high-risk ‘all-at-once’ migrations.”
There are real limitations worth acknowledging. ELM carries over the fundamentals — full Git history, branches, tags, pull request metadata, including comments and user history, and branch policies translated into GitHub rulesets — which, for teams whose work is primarily code-focused, may cover most of what they need.
But pipelines, work items, wikis, and test plans all have to be handled separately, and for enterprises deeply embedded in Azure DevOps’s broader project management and CI/CD tooling, ELM is a starting point rather than a complete solution.
For large organizations with hundreds of repositories, this is a multi-stage undertaking regardless.
Migration to GitHub
For Microsoft, the calculus is all about AI — GitHub is where Copilot, the Copilot Coding Agent, and the broader agentic development suite live, and Azure Repos is not part of that picture.
To demonstrate that this is more than a customer pitch, Microsoft recently published details of its own migration — its Copilot, Agents and Platforms (CAP) organization moved over 1,600 repositories and 3,100 developers across in six months, with a team of just two dedicated engineering leads driving the effort.
By consuming its own dog food at scale, Microsoft is making the case that the disruption is manageable and the payoff meaningful. Poonam Gupta, partner director of product management for 1ES and Azure DevOps at Microsoft, cites AI as the primary driver of the migration.
“Software development is being reshaped by AI, and where code lives now have a direct impact on how much value organizations can capture.”
“Software development is being reshaped by AI, and where code lives now has a direct impact on how much value organizations can capture,” Gupta writes. “For teams that want to take full advantage of AI-native development, repository location is becoming a strategic decision.”
The elephant in the room
Rumors of Azure Repos’ eventual deprecation have circulated online for years, and while Microsoft has not confirmed anything on that front, the direction of travel is clear.
The community response to Gupta’s June 3 post captured the mood among enterprise customers: several questioned why AI capabilities couldn’t be brought to Azure Repos rather than requiring a platform change, while others raised the cost differential — Azure DevOps Basic costs $6 per user per month, compared with GitHub Enterprise’s $21.
And more than one commenter interpreted the post as a deprecation notice in all but name. “The writing was on the wall since MS [Microsoft] bought GitHub,” wrote one commenter. “[Azure DevOps] is dead and MS wants everyone moving to GitHub… Everybody saw this coming, and only MS denied it.”
Perhaps more important here is the question that Microsoft sidesteps: if GitHub has spent the past year struggling under the weight of agentic development traffic, why is now the right time for enterprises to bet their critical infrastructure on it?
The timing acquired an extra layer of awkwardness on Friday, when 73 Microsoft-owned GitHub repositories — including the Actions used to deploy Azure Functions — were disabled in a Miasma worm attack, breaking CI/CD pipelines for developers globally.
None of this necessarily undermines the strategic case for moving to GitHub. The AI development ecosystem is consolidating there, and the migration tooling is getting meaningfully better. But for enterprise teams weighing the decision, reliability and security aren’t footnotes — they are the main criteria. Microsoft is betting that access to Copilot and agentic workflows is compelling enough to tip the balance.
It may well be right, but a 30-minute cutover window is only part of what it will take to make that argument stick.
Almost everyone’s workplace experience is now set to welcome AI-agent-driven actions through the applications we use daily, and this rapid evolution has some serious implications for how identity and access management (IAM) works.
While traditional IAM models were developed with human users and their predictable access patterns in mind, AI agents operate differently.
AI agents can quickly perform reasoning functions that impact the way business analytics feeds into board-level management dashboards; they can invoke tools that fuse new connections to an API, update a database, or run a new software script; and they can access other software services and data resources across an organization’s infrastructure and total software stack in dynamic, continuous, and sometimes unpredictable ways.
Cloud infrastructure automation and security company HashiCorp has sought to provide IAM services capable of servicing the agentic age for some time.
Before it became an IBM company in February last year, HashiCorp introduced Boundary in 2020 as an open-source project to allow software engineers to securely access dynamic hosts and services with fine-grained authorization without direct network access.
IBM senior solutions engineer Andre Faria and HashiCorp senior technical product marketing manager Van Phanblogged on June 4 to explain that as agents now go into live production systems, they will have access to “critical infrastructure resources” such as internal web services, cloud platforms, and other operational systems.
The pair say this is concerning if agents are improperly provided with long-lived static credentials that are poorly managed, rarely rotated, and tough to audit.
Credentials that are poorly managed: “a dangerous combination.”
“This creates a dangerous combination of broad access and limited oversight. Without proper guardrails, AI agents may autonomously make decisions or execute actions that negatively impact production workloads, corrupt data, trigger outages, or unintentionally expose sensitive information,” write Faria and Phan.
They further note that organizations need a way to monitor which sessions are active, which systems AI agents access, when they access them, what actions they perform, and whether their behavior deviates from policy.
Because agentic runtimes and individual execution behavior are so inherently fluid and prone to change, we can no longer set identity management, authorization, and session control policies at the point of deployment — every agent needs a unique identity and just-in-time privileges that act as a secure point-of-use access layer.
It’s time for just-in-time
Because agentic runtimes and individual execution behavior are inherently fluid and prone to change, we can no longer set identity management, authorization, and session control policies at deployment. Every agent needs a unique identity and just-in-time (JIT) privileges (a zero-trust-based secrets management technique that acts as a secure point-of-use access layer) to ensure software systems don’t become brittle or susceptible to attack as they scale.
“With Boundary’s authorization flow, access to a specific resource is granted only when needed, for a specific action, and only for the duration of that session. This helps organizations strengthen governance and maintain tighter control over how AI agents access critical infrastructure,” wrote Faria and Phan.
Boundary applies similar principles to ensure non-human and agentic identities do not have overprivileged access or handle static long-lived credentials. The software also provides monitoring, audit logs, and session recordings that can be used to play back and reveal detailed actions taken by AI agents during session access.
Adopting dynamic credential brokering
Underlining his original blog, IBM’s Faria tells The New Stack that the IBM 2025 Cost of a Data Breach Report states that the global average breach costs organizations $4.4 million. He says it also shows that 97% of organizations that reported an AI-related security incident lacked dedicated AI access controls, and 63% did not have any AI governance policies to manage AI or prevent shadow AI.
“The risk highlighted by those statistics, plus the fact that agent compromise is now the fastest-growing attack vector in the industry, showcases why it is urgent to define a solid and secure infrastructure access strategy for agentic AI workflows,” says Faria, who also points to the role HashiCorp Vault plays regarding dynamic credential brokering for Boundary.
Boundary can facilitate the use of dynamic credentials rather than static credentials. When paired with HashiCorp Vault, access with dynamic credentials becomes a reality because Vault’s secrets engines generate short-lived credentials that expire after use. Even if a credential is intercepted, it cannot be used to cause damage.
But is all of this enough?
“Agents are non-deterministic and operate at machine speed. To contain them, they need hardened, isolated runtimes that govern their behavior before they ever touch production. Cryptographic identity, just-in-time, short-lived privileges, plus ephemeral, trusted runtimes for agents to operate in – that’s the bar.” – Ev Kontsevoy, Teleport.
An immutable cryptographic hardware root of trust
Ev Kontsevoy, CEO and co-founder of AI infrastructure identity specialist Teleport tells The New Stack that just-in-time privileges and auditable control for AI agents aren’t new ideas per se. He advises that every agent today needs its own identity, cryptographically secured by a “hardware root of trust,” i.e., immutable cryptographic keys that reside at the chip level.
“To enforce policy consistently across infrastructure, software engineering teams need a unified identity layer — one that treats humans, machines, workloads, and AI agents the same way, as first-class identities,” Kontsevoy says. “Agents are non-deterministic and operate at machine speed. To contain them, they need hardened, isolated runtimes that govern their behavior before they ever touch production.”
Looking at the live working accounts it touches, the Teleport team reports that credential sprawl in service accounts and tooling remains one of the biggest attack surfaces in production infrastructure today.
“It’s not enough to manage credentials better; we need to eliminate them entirely so they can’t result in standing privileges or unintended actions. Cryptographic identity, just-in-time, short-lived privileges, plus ephemeral, trusted runtimes for agents to operate in – that’s the bar,” Kontsevoy insists.
“Developer and agent identities often sit on attack paths to critical systems because they can provision infrastructure, retrieve secrets, trigger pipelines, query data stores or inherit trust from other services.” – Justin Kohler, SpecterOps.
Defining attack paths to critical systems
As we seek to tame the Wild West of agentic access and actions through identity services, we may be overlooking that identity itself can act across cloud, developer, and production environments.
Justin Kohler, chief product officer at identity attack path management (IAPM) company SpecterOps, tells The New Stack that developer and agent identities often “sit on attack paths to critical systems”, meaning they can provision infrastructure, retrieve secrets, trigger pipelines, query data stores, or inherit trust from other services.
“Organizations need to understand where these identities can actually take them, continuously prioritize the paths that create the most risk and then enforce and audit access at the point it is used,” Kohler says. “If those identities are over-permissioned, impersonated or manipulated, the compromise follows the same relationships an attacker would, from one identity, to one system, to the next trust boundary.”
Integrated authentication & authorization in automated applications
If now seems like the right time to talk about this, the Cloud Native Computing Foundation (CNCF) TAG Security and Compliance committee posted a blog to showcase a new whitepaper last Thursday. The whitepaper supports the foundational security philosophy outlined in the HashiCorp blog above, but underscores the need to embrace open-source, vendor-neutral standards rather than proprietary software.
According to the whitepaper, “Controlling access to systems and data is a fundamental requirement in any environment; in cloud-native environments, this requirement is shaped by characteristics such as highly dynamic and short-lived workloads, the collapse of perimeter-based trust models, and the need to integrate authentication and authorization into automated application lifecycles.”
The takeaway here seems pretty clear: Short-lived workloads and long-lived credentials don’t mix, but just-in-time, short-lived privileges — a period that used to be 90 days, but shrunk to 24 hours, then to minutes, and now down to a period we can call the “ephemeral lifespan” — is now the clock we need to run to.
As large language models evolve from mere chatbots into autonomous agents capable of reasoning, planning, and acting, they are beginning to orchestrate complex application stacks on their own.
However, these agents are now encountering their most formidable obstacle: the database.
Andy Pavlo. Credit: Carnegie-Mellon University
“Databases pose the hardest and most important challenge for agents, due to their unforgiving correctness and performance requirements,” Andy Pavlo, Associate Professor of Computer Science at Carnegie Mellon University, told attendees last week at the Percona Live 2026 conference here in Mountain View, California, at the Computer History Museum.
In a discussion on the intersection of AI and open-source infrastructure, Pavlo contended that while coding agents can readily regurgitate standard data structures, the database remains the most difficult part of any system to automate and optimize.
“For example, if an agent hallucinates a UI component, the page looks slightly off; if it hallucinates a query or a configuration change in a production database, the entire system can vanish,” Pavlo says.
Now THAT would be a cause for alarm.
The multi-agent tug-of-war
Pavlo identifies two primary ways AI is impacting the database world: tuning agents and coding agents. Tuning agents aim to solve the “black magic” of database optimization — automatically adjusting system knobs, physical designs (such as indexes), and query execution strategies. Historically, this required a human database administrator (DBA) to spend years developing the intuition to know which configuration would yield better latency or throughput.
“If an agent hallucinates a UI component, the page looks slightly off; if it hallucinates a query or a configuration change in a production database, the entire system can vanish.”
The challenge is that these specialized agents often operate in silos, Pavlo said. A knob-tuning agent might be unaware of what an index-tuning agent is doing, leading to local minima where the system is better than stock but far from optimal. CMU’s research into multi-round and sequential tuning aims to solve this by creating a coordinating framework, though even this faces a “curse of dimensionality,” Pavlo says.
Carnegie Mellon’s Database Group pioneered the concept of self-driving and machine-learning-driven database optimization. Sequential tuning and multi-round tuning are prime components of their autonomous database management system (DBMS) projects.
Multi-round and sequential tuning in AI databases refers to advanced machine learning and data engineering methods in which AI models are refined for multistep reasoning, tool use, or complex conversational histories. These frameworks ensure that AI models not only respond in isolated single-turn bursts but maintain context and logic across complex interactions.
With trillions of possible configuration combinations, the search space for a perfect database is effectively exponential.
The coding agent advantage and the optimizer wall
On the development side, coding agents are already proving to be hyper-productive collaborators. Pavlo observed that at CMU, student submissions for database projects saw a massive spike in lines of code once LLMs were permitted. “The coding agents are very good at building almost every part of a database — B+ trees, hash tables, buffer managers — because they can regurgitate standard implementations found in textbooks and open-source repos,” Pavlo said.
However, the “double black diamond” challenge, Pavlo said, remains the query optimizer. Unlike basic data structures, query optimizers are rarely available as clean, modular open-source references. They are often deeply entangled with the systems for which they were built. Furthermore, proving that an AI-generated transformation rule is semantically correct — meaning it produces the same result as the original query but faster — is an unsolved problem.
Risks include hallucinations and security
The shift toward agentic database management isn’t without significant risk. Pavlo and other industry leaders, such as Percona co-founder Peter Zaitsev, warn that delegating orchestration to agents introduces massive stability and security gaps. There are already documented cases of agents being pointed at a database and accidentally dropping the entire system or leaking sensitive information because they didn’t understand the nuance of access controls, Zaitsev said.
Furthermore, LLMs suffer from so-called AI slop, in which they generate code that is hyper-specialized to a specific query but fails to generalize. For example, if a developer uses an agent to optimize an “Extract Year” clause, the agent might build an internal data structure that breaks the moment the developer tries to enact “Extract Month.”
Automation as a collaborator, not a replacement
Despite these hurdles, Pavlo said he is optimistic about the Agent Operator model. This envisions agents handling the “3 a.m. s***’s on fire” situations — immediate performance anomalies and stability issues — while humans focus on higher-level architectural design. By using Agent Boosting techniques to bootstrap training data from previously tuned databases, the time required to optimize a system can be cut from 12 hours to under 15 minutes, Pavlo said.
In the new AI era, the goal isn’t only to have an AI that writes code, but a system that can reason about its own performance and correctness. Pavlo concludes that the database is the foundation of knowledge for any agent. “If we want autonomous systems, we must first master the unforgiving art of the autonomous database,” he says.
“If we want autonomous systems, we must first master the unforgiving art of the autonomous database.”
The biggest obstacle to enterpriseAI isn’t models, data science talent, or even infrastructure. It’s operations.
Across today’s enterprises, hybrid complexity has outpaced IT’s ability to manage it. Applications, workloads, runtimes, and infrastructure now span on‑premises environments, public clouds, edge locations, and air‑gapped sites. Each layer brings its own tools, vendors, and operational language. The result is friction everywhere and a widening gap between AI ambition and operational reality.
Latha Vishnubhotla, chief platform officer at Hewlett Packard Enterprise, tells The New Stack the challenges begin on Day 2.
“People can bring things up and make them functional very quickly,” says. “But where they spend most of their time is after the infrastructure becomes functional. Day 2 to Day N is where they spend a lot of time.”
That’s the problem enterprises are running into now. It’s not getting infrastructure up and running, but keeping it running, optimized, and reliable as AI workloads move from pilot to production.
Read on to dive into not only the Day 2 problem, but to learn how HPE’s GreenLake hybrid cloud management platform has grown to respond to this enterprise complexity — including that cross-platform infusion of agentic AI.
Day 2 is when the hybrid cloud breaks down
In hybrid environments, operations teams aren’t managing a single stack. They’re juggling multiple runtimes, from bare metal and VMs to containers and AI‑native platforms. Infrastructure across compute, storage, and networking often from different vendors. Workloads spread across data centers, public clouds, edge, and disconnected sites. Legacy systems that were never designed to work together
Each layer has its own management tools and telemetry. When something goes wrong, the symptom rarely appears in the same place as the root cause.
“All these different tiers are talking to each other, but it’s not linear. You have to comb through and figure out where the issue actually is.”
“All these different tiers are talking to each other, but it’s not linear,” Vishnubhotla says. “You have to comb through and figure out where the issue actually is.”
Day Zero provisioning may be fast. Day 2 operations are where complexity compounds and teams burn time reacting rather than optimizing.
More AI is making the ops problem worse
AI not only raises the stakes but also delivers a solution.
Enterprises want to run more AI workloads, but data centers have finite capacity. Power, cooling, cost, and sustainability constraints are real. That’s why FinOps and GreenOps have become inseparable from infrastructure operations.
“When you want to run these workloads, you have to ask: what’s not being used?” Vishnubhotla says. “Why am I wasting here? Should I move something? Should I retire it?”
This is where traditional, human‑driven ops models start to break. There’s too much data, too many layers, and too many dependencies to reason about manually, especially at enterprise scale.
The ops platform as connective tissue
What enterprises need isn’t another point tool. It’s an operations platform that acts as connective tissue across the hybrid estate.
That’s the role GreenLake is designed to play.
GreenLake provides a unified platform experience for running and managing hybrid environments across on‑premises, private cloud, edge, and collocated infrastructure while preserving choice and control. Instead of hiding infrastructure behind abstraction, it makes it visible, observable, and operable from a single control plane.
“The control plane is actually running in the cloud,” Vishnubhotla says. “You get visibility across the entire estate.”
For organizations managing thousands of sites and tens of thousands of devices, that visibility is foundational. But visibility alone isn’t enough anymore.
Why agentic AI changes everything
The next step is agentic AI, AI systems embedded directly into the ops platform, trained on the context of specific infrastructure domains.
A networking agent understands networking. A storage agent understands storage. A compute agent understands compute. Each brings deep, domain‑specific intelligence to Day 2 operations.
“Each layer already has intelligence,” Vishnubhotla says. “If we can connect this intelligence, we can unleash very powerful outcomes.”
That’s where the idea of an agentic mesh comes in. Instead of siloed insights, AI agents share context across layers during provisioning, troubleshooting, and optimization. This shortens the time to root cause, reduces alert noise, and opens the door to predictive and, eventually, autonomous operations.
Predictive maintenance is a clear example. Rather than reacting to failures, AI can anticipate what’s likely to break, prioritize what actually matters, and help teams act before outages cascade.
Faster time to value for AI starts with competent ops
Agentic operations also unlock something enterprises care deeply about: faster AI ROI.
With a shared, platform‑level view, ops teams can answer questions like:
What is connected to the estate
Where is infrastructure deployed?
Who’s using it—and how?
Where is capacity being wasted?
GreenLake supports automation through copilots and MCP servers as well as UI‑driven workflows, reducing provisioning times and operational overhead. AI agents can even help predict demand and close feedback loops that used to take weeks.
“The bottleneck has always been on the ops side. Enterprises are deploying and operating infrastructure from Day Zero to Day N to unlock AI value faster.”
“The bottleneck has always been on the ops side,” Vishnubhotla says. “Enterprises are deploying and operating infrastructure from Day Zero to Day N to unlock AI value faster.”
The answer is a platform, not another tool
Hybrid complexity isn’t temporary. AI pressure isn’t slowing down. And Day 2 operations are only getting harder.
That’s why the industry is converging on a clear conclusion: the answer isn’t more tools; it’s a unified, intelligent ops platform.
GreenLake brings together visibility, agentic AIOps, and cross‑domain intelligence in a platform built for how enterprises actually run today. It connects the silos, scales operations teams, and turns infrastructure from a bottleneck into an enabler.
If AI is the future of the enterprise, operations is the gatekeeper. And the ops platform powered by agentic AI is how that future gets unlocked.
Informational and operational technology data have long been treated as separate domains.
But AI changed the game. Today, you need the capacity to regularly ingest OT data into your IT systems without a hitch. (Or a breach.) You risk being left behind as your competitors put all their data to work, or assume the risk of consistently importing data from the edge to your internal systems.
This problem is an immediate one for any company unwilling to be left behind in the AI era: If you want to take full advantage of AI, you need quick, ready access to relevant data. And if your physical operations have hit a snag, your digital tools need to be kept in the loop regularly.
The solution is not to build a host of custom scripts or depend on legacy FTP or SFTP solutions to bring data in from the edge. Those disparate tools can degrade, leak data, and fail during later, repeated OT data extraction runs.
Instead, engineers looking to free IT and OT data from their respective siloes are turning to a managed solution that offers strong encryption, continuous transfer monitoring, and the ability to fully audit every data handoff across the pipeline
Even more, OT systems — the Programmable Logic Controllers, Supervisory Control and Data Acquisition platforms, and historian databases running protocols like Modbus and OPC UA — were designed for uptime rather than connectivity. In modern architecture, however, no operational data can be left behind.
Getting data out of these environments means working against a connectivity model that was never meant to support the polling frequency or authentication patterns that modern IT infrastructure expects. Adding to the challenge, the more tools you introduce to free the OT data, the more attack vectors they may open.
A breach at the OT boundary can affect the physical systems those networks control. That’s a risk calculus most IT security frameworks weren’t built to handle.
On at 12 p.m. Eastern/9 a.m. On Tuesday, June 23, Fortra’s Jerrod Foster & Michael Barford will join The New Stack to discuss IT and OT systems, why extracting operational technology data is challenging, and how Fortra GoAnywhere MFT can resolve both data movement and data security issues that many engineers face today.
Register here to join the conversation:
What you’ll take away:
Why the IT/OT boundary is an AI infrastructure problem: How the connectivity gap between operational and information technology creates a hard ceiling for teams building on live operational data — and what becomes possible when that data is reliably accessible inside modern pipelines.
Where DIY solutions break: Why custom scripts and legacy transfer tools fail under real operational conditions — brittle transfers, no visibility, and attack surfaces you can’t audit.
What secure OT data movement actually looks like: How Fortra GoAnywhere MFT provides an encrypted, automated, and auditable data movement layer that works with the constraints of real OT environments, not against them.