❌

Normal view

Why most AI projects fail: It’s infrastructure and people 

Astronaut standing on floating platform amid abstract digital data blocks

AI trash-talkers love to rip on the technology for failing to produce meaningful business results, often pointing to studies like that from MIT NANDA, which reveals a 95% failure rate for enterprise AI solutions, or that from IDC, which states “only 9% of [Europe, Middle East, and Africa] organizations have been able to deliver measurable business outcomes from most of their AI-related projects over the past two years.” 

What many AI skeptics fail to account for is the experiential nature of AI prototypes; not all these projects are actually meant to move beyond the testing phase. Still, a 5% success rate is embarrassing. 

What’s the holdup? 

Two things. First, most organizations build AI prototypes on sand; that is, the data infrastructure on which they build early applications can’t support later moves to production. Meanwhile, the operational teams responsible for managing those applications in production often lack the human power to keep up with engineering’s growing output. 

4 reasons prototyping infrastructure  ≠ production infrastructure

When asked why so many AI prototypes don’t make it to production, Phillip Merrick, co-founder, CPO, and chairman, pgEdge, tells The New Stack that data infrastructure is largely to blame. 

Specifically, he explains that prototype environments don’t meet the requirements of large enterprises for production, naming four main ways they fall flat.

First, Merrick says prototyping environments lack the deployment flexibility organizations need to move from prototype to production. 

Vendor-managed cloud platforms, he acknowledges, may seem like an obvious choice for prototyping, as they allow teams to get up and running quickly. Still, he warns they lack the technical chops to support AI applications in production, especially in security, compliance, and governance. Particularly for organizations in healthcare, finance, or other regulated industries, vendor-managed cloud platforms often lack the stringent controls found in self-managed cloud or on-prem environments. 

“You’ve got to be able to choose where that AI prototype is ultimately going to be put into production.”

In this way, flexibility and security go hand in hand. Merrick asserts. “You’ve got to be able to choose where that AI prototype is ultimately going to be put into production.” 

Similarly, when it comes time to shift to production, Merrick says vendor-managed cloud platforms can introduce data sovereignty challenges at both the enterprise and regional levels. 

“Your data layer is obviously where you enforce this,” notes Merrick. But he says the environments most teams use for prototyping muddy the waters: “If it’s on a vendor-managed platform in who knows what cloud, what region, then you’ve lost data sovereignty.” 

Lastly, Merrick brings attention to reliability, explaining that AI prototypes can’t move into production without assurance of high availability. For example, when it comes time to upgrade the database or swap hardware, can it be done without downtime? 

“In the vendor-managed cloud world, the answer to that is almost always no,” says Merrick, reiterating his point that moving AI apps from prototype to production requires enterprise-grade infrastructure. 

So why are developers prototyping where they can’t productionize?

If data infrastructure selection is what’s holding developers back from moving AI prototypes into production, then why do they keep starting on the wrong foot? 

As Merrick explains, many are attracted to the ease of use of vendor-managed cloud platforms. “These prototyping environments admittedly make it very easy to get started,” he says. But prototyping shortcuts, it seems, don’t pay off in the long run, as someone else is ultimately on the hook for making those prototypes production-ready. 

Still, Merrick doesn’t blame developers for looking for the easy way out. Rather, he says there’s a disconnect between the prototyping playground and the production battleground that prevents developers from understanding what it will take to productionize prototypes down the pike. 


More pgEdge articles in The New Stack


Years ago, he says, tooling decisions were primarily made top-down without developer input: “Then, starting 15–20 years ago, the developers won back, quite rightly, the power in being able to choose their own tools.” 

The problem now, Merrick claims, is that developers make tooling decisions exclusively for prototyping environments, without anticipating production needs. Internal divisions mean developers are often only responsible for building prototypes before passing the baton to an entirely separate operations team for production: 

“The upshot is you really don’t have, in some organizations, this throughline of understanding [of] what the production requirements are all the way back to the developer making the initial choices.” 

For Merrick, this disconnect is where AI projects start to fall apart, as teams are left trying to move AI prototypes from accessible-but-inadequate vendor-managed cloud platforms to enterprise-grade data infrastructure that meets requirements for deployment flexibility, security, data sovereignty, and high availability. 

“But if you make the right data infrastructure choice, you won’t have that disconnect,” he says, “because you’ll have this throughline from prototype to production.” 

He names Postgres as the data infrastructure that best helps developers bridge this divide, calling it “the Swiss army knife of databases” due to its extensibility, fully open-source nature, and ability to address diverse data management problems, from unstructured data to vector embeddings to geospatial data. 

Where and how Postgres is run matters too, Merrick points out, again drawing attention to the limits of many vendor-managed cloud environments that often lack the governance controls to meet data requirements and/or the deployment flexibility to shift to compliant on-premises or BYO cloud environments. 

But picking the right data infrastructure only solves half the problem

Merrick says there’s another part of the equation most organizations are overlooking: people, or more precisely, database administrators (DBAs) and their growing workloads. 

Per Stack Overflow’s 2025 Developer Survey, 84% of respondents use AI tools, up from 76% the year prior. Meanwhile, Supabase says over 60% of databases on its platform have been launched “by some sort of AI tool.” As Merrick points out, this explosion of productivity comes with a catch: there aren’t enough DBAs to keep up.

“You’ve had this massive, massive step shift in developer productivity,” he explains. “But you have to have some way of managing that on the production side; these databases can’t go unmonitored.” Operations and administration teams were already struggling to keep track of existing Postgres databases before agentic engineering added even more, he says: “Who’s going to manage them?” 

He says it’s time for agentic operations to catch up with agentic engineering. 

AI DBA agents can give humans “superpowers” 

If it seems Merrick is proposing organizations look to fully autonomous DBA agents to take over, he says the industry isn’t there yet: 

“The world is not ready for fully autonomous databases administered by AI DBA agents. But there is a massive resource shortage and productivity problem, and DBAs can only manage so many databases,” he explains. Meanwhile, new “AI applications require so many more databases to put in production.” 

“The world is not ready for fully autonomous databases administered by AI DBA agents.”

So how can organizations increase their operational capacity? 

Merrick says DBAs should look to new AI DBA agents, not to take over but to give them “superpowers” to monitor and manage more databases with less manual slog. 

pgEdge’s Ellie is one example. Part of the pgEdge AI DBA Workbench, Ellie is an AI agent that has 21 MCP tools and can run EXPLAIN ANALYZE, inspect schemas, query historical metrics, and walk through multi-step diagnostic workflows. When a database falters, Ellie finds the problem, diagnoses it, and provides a solution in the form of working SQL code for the human DBA to review. “When you’ve reviewed it and agree that it’s the right course of action, you literally press the play button, and the agent plays that SQL code into the database, and you solve your problem,” explains Merrick.

In this way, Ellie should bring more capacity to operations teams, where Merrick insists organizations are starved for DBA expertise. To his point, some industry predictions say 41% of today’s database professionals intend to leave the industry in the next decade, half moving into retirement and the rest seeking other work. 

“An agent … can actually respond to those alerts far more quickly and productively than a human can.”

Without AI agents, Merrick argues, DBA work is tedious, laborious, and time-consuming. As he explains it, a database may have been humming along just fine, but when there’s a snag, trouble can manifest across multiple applications; it’s then up to the DBA to comb through monitoring data to observe and diagnose the problem, essentially scouring for a needle in a haystack. 

“An agent,” he says, “can actually respond to those alerts far more quickly and productively than a human can.” 

Better infrastructure AND people: It takes two to improve AI prototype success rates 

In nearly any context, AI raises questions about quality over quantity, and enterprise AI projects are no exception. Agentic engineering means developers can now produce more, but all those prototypes don’t just fly directly into production. Limitations in both infrastructure and operational human power are creating obstacles that cause many AI prototypes to fail. 

For Merrick, easing the transition from prototype to production requires not only great AI tooling but production-ready data infrastructure, paired with agentic operations that can keep up with the agentic engineering boom. 

The post Why most AI projects fail: It’s infrastructure and people  appeared first on The New Stack.

Palantir’s Alex Karp and Mistral’s Arthur Mensch agree: AI lock-in is coming for enterprises

Bundle of colorful electrical wires hanging in a tangled mass

Palantir CEO Alex Karp went on CNBC’s Squawk Box last week to discuss a new partnership with Nvidia to deploy open-weight AI models in sovereign government environments. But viewers got a nearly 20-minute broadside against the entire frontier AI model industry, calling it “effing insane” and accusing companies like OpenAI and Anthropic of overcharging enterprises while harvesting their proprietary data.

Days later, Mistral CEO Arthur Mensch made a strikingly similar case on LinkedIn, warning that closed AI providers are gaining “immense leverage” over enterprise customers as organizations connect proprietary workflows to hosted models. He suggests open-weight models, open data systems, and enterprises building their own training flywheels.

The two executives are approaching this from opposite ends of the market, yet their convergence on the same message within the same week underscores architectural control.

Two pitches, one argument

Karp runs a company that sells an application and ontology layer designed to sit between enterprises and the models. The Palantir-Nvidia deal pairs Nvidia’s open Nemotron models with Palantir’s Sovereign AI Operating System, built on AIP, Foundry, Ontology, and Apollo, enabling government agencies and critical infrastructure operators to deploy, fine-tune, and audit AI models within their own air-gapped environments.

When CNBC’s Becky Quick told Karp he sounded angry, he pushed back, saying, “This is the voice of American business that is being channeled through me,” and urged the panelists to call any CEO privately to verify.

“This is the voice of American business that is being channeled through me.”

Mensch’s company sells open-weight models, and he has a custom training platform called Forge, which frames the problem differently but reached the same conclusion. He argues in the post that closed providers have a track record of going after their most successful customers once they learn what those customers are building. His program runs from open models to open data stores, strict access controls, and a continuous training flywheel that improves systems on internal interactions.

Lock-in gets an upgrade 

If you’ve been building software at scale for any length of time, you already know that when new technology arrives, enterprises can’t help but rush to adopt it. But the dependency problem becomes impossible to ignore.

We saw the same problem with cloud computing when companies went all-in on a single hyperscaler’s proprietary services, only to later discover that the cost of switching providers could exceed the cost of staying, even when staying meant overpaying. It’s one of the reasons the industry spent years building abstraction layers, portability tooling, and multi-cloud strategies in response.

Foundation models are raising the same questions, but there’s a twist. When an enterprise connects a model to its internal data, including customer records, proprietary processes, and domain-specific knowledge, the dependency becomes informational. Karp’s argument is that model quality is converging across providers, but the operational leverage accrues to whoever controls the deployment layer and the data flowing through it.

Mensch’s claim has a concrete referent that enterprise architects will recognize. In 2025, Anthropic cut off model access to coding startup Windsurf while building its competing product, Claude Code. The Brookings Institution has separately warned that model providers increasingly compete with their own customers as they chase application-layer revenue.

When access disappears overnight 

When the U.S. government ordered Anthropic to suspend access to its most advanced models for foreign nationals, the company cut access across the board, including to enterprise customers in Europe who had built workflows on top of those models.

Access has since been restored, but for CIOs and enterprise architects who had treated model APIs as stable infrastructure, it was the same as if a cloud provider pulled compute resources without warning. It’s probably why the incident sent European policymakers into overdrive. Mensch, whose company had open-weight alternatives ready, seized the moment.  

Enterprise teams need a plan for when a critical dependency is modified, repriced, or revoked by a provider whose incentives may not always align with their own.

Enterprise teams need a plan for when a critical dependency is modified, repriced, or revoked by a provider whose incentives may not always align with their own.

Architecture shifts toward portability 

Enterprise AI architecture is now essentially this: don’t marry a single provider, build for portability, keep your most sensitive data and logic under your own control.

In practice, this is showing up in three ways.

Firstly, we’re seeing a portfolio approach. A powerful closed model is maintained for complex reasoning and customer-facing work, while an open-weight model is used for repetitive, high-volume tasks. For businesses handling sensitive information, the appeal is that an open model can be run entirely on their own infrastructure, so data never has to leave the organization.

The second is the rise of the model-routing layer, abstraction frameworks that let organizations swap models without rewriting their applications. Palantir’s ontology pitch sits here as does the emerging crop of agent orchestration tools that treat the LLM as a pluggable component behind a standardized interface.

The third is the open-weight movement itself. Nvidia shipped Nemotron 3 Ultra in June under a permissive Linux Foundation license. Meta’s Llama continues to expand. Mistral’s Forge platform lets enterprises train custom models on their own data. And Mistral is teasing an upcoming open-weight model this summer, with early access opening in July.

Follow the commercial incentives

It would be naive to ignore the commercial interests at play. Karp’s Palantir sells the deployment and governance layer; it benefits directly if enterprises treat models as interchangeable commodities. Mensch’s Mistral sells open-weight models and a training platform, and it benefits directly if enterprises distrust closed providers. Zoho’s Sridhar Vembu, who endorsed Karp’s position publicly last week, has his own reasons for wanting enterprises to own their AI infrastructure rather than rent it from Silicon Valley.

But the fact that multiple executives across different market segments are saying companies need to own their data, maintain deployment flexibility, and not hand their competitive advantage to a provider who might become their competitor suggests the argument is resonating.

What developers should watch

If your most sensitive data is flowing through a third-party API with terms of service that can change, you’ve made a governance decision that your compliance team may not have fully evaluated.

For the engineering teams actually building on foundation models, the takeaway is that if your most sensitive data is flowing through a third-party API with terms of service that can change, you’ve made a governance decision that your compliance team may not have fully evaluated.

The engineering choice, to abstract the model layer, to evaluate open-weight options alongside closed APIs, to think about deployment portability the same way you think about cloud portability, is increasingly strategic.

The post Palantir’s Alex Karp and Mistral’s Arthur Mensch agree: AI lock-in is coming for enterprises appeared first on The New Stack.

Andrej Karpathy, Google and Garry Tan agree Markdown is the answer, but they’re not solving the same problem

Illustration of businessman jumping across lightbulbs toward a glowing bright idea

In April, Andrej Karpathy published a GitHub gist file called “LLM Wiki,” a brief text document designed to help one build a personal knowledge base using LLMs. It’s based on the premise that an AI agent will keep what it knows as linked Markdown files it can read and rewrite, because a language model does not get bored maintaining cross-references and can touch fifteen files in a single pass. It was only a few thousand words with no product attached.

Two months later, Google turned that instinct into a published standard called the Open Knowledge Format. The OKF packages organizational knowledge, metrics, tables, and runbooks as plain Markdown that any agent can read without a proprietary account. Google is careful to call it v0.1 — a starting point rather than a finished standard.

Garry Tan, the Y Combinator president, got there first in a different lane. His gstack, an MIT-licensed Claude Code setup that crossed 66,000 GitHub stars within weeks, comprises 23 specialist roles, each a Markdown file. No runtime; no code; just prose that runs across ten different coding agents.

Markdown has become the substrate agents read and write

Three approaches, three different needs, one common solution. Karpathy sought agent memory, Google aimed for enterprise context in BigQuery agents, and Tan wanted a way to summon an engineering team from a terminal. All three turned to the same basic resource: a folder of Markdown files versioned in git.

Developers had already established this practice. CLAUDE.md and AGENTS.md are present in millions of repositories as the initial files an agent loads. OKF and gstack are the evolved forms of this convention – one focused on what the agent knows, the other on how it behaves.

This is the Git and JSON playbook tied to the agent’s knowledge. The formats that survived are the ones you could start using without changing anything. You can simply cat the file, clone the repo, and any tool you already use can parse it. MCP remains important as the interface an agent connects to. Markdown is becoming the format that carries the content.

The lock-in moved from the model to the files

The significant factor to observe here is the competitive advantage, not technical specifics. For two years, the belief was that owning the best model meant controlling the developer.

This perspective is now shifting. Replacing Claude with GLM or Codex, gstack continues to operate because the core intelligence evolved, but the documentation did not.

The moat is shifting from the model to the Markdown a team owns and accumulates over time.

The moat is shifting from the model to the Markdown a team owns and accumulates over time. A company’s OKF bundle, including its runbooks, metric definitions, and architecture decisions, is, by design, portable across clouds, models, and frameworks.

That kind of portability is the reason vendor-neutral formats exist and why Google’s OKF deserves a closer look.

If no one develops consumers for it, it remains just a good idea that Google released on a slow Friday.

The area where I am most likely mistaken is durability. Declaring Markdown standards is easy, but making them reliable is difficult. OKF is merely a 0.1 draft with a reference implementation, not a full ecosystem. If no one develops consumers for it, it remains just a good idea that Google released on a slow Friday.

The direction remains determined by three separate bets targeting the same file format within a single quarter. Your next agent is likely to interpret its context from a Markdown folder, and the creator of that folder now possesses an advantage that the model vendor cannot easily replicate.

The post Andrej Karpathy, Google and Garry Tan agree Markdown is the answer, but they’re not solving the same problem appeared first on The New Stack.

❌