❌

Normal view

Stop graphing everything: When GraphRAG actually beats vector RAG

If you have built anything with retrieval-augmented generation (RAG) in the last two years, you have lived its central frustration: You chop your documents into chunks, embed them, retrieve the top few that look similar to the question, and hand them to the model. For “What was our Q3 refund policy?” This works beautifully. For “What are the recurring themes across two years of customer complaints?” it falls flat — because no single chunk contains the answer.

The fashionable fix is GraphRAG: Instead of feeding the model isolated snippets, you first build a knowledge graph of the entities and relationships in your corpus, then use that structure as context. The pitch is seductive. But seductive pitches deserve scrutiny, so I went through the evidence — the original Microsoft paper plus four independent benchmark studies — to answer a simple question: When you swap text chunks for a context graph, do answers actually get better?

The short version: Yes, substantially — but only for the right kind of question, and not for free. Let me show you the receipts.

Why text chunks hit a wall

Standard vector RAG retrieves the k passages most similar to your query. That design has three structural blind spots:

  • It can’t connect the dots. When an answer requires joining facts that live in different passages through a shared entity, chunks embedded in isolation never reveal the link.

  • It’s blind to global questions. “What are the main themes?” needs the whole corpus, but similarity search only returns the handful of chunks that superficially resemble the question.

  • It severs context at chunk boundaries. The relationships and hierarchy that complex reasoning depends on are exactly what chunking throws away.

Microsoft Research framed this crisply when they introduced GraphRAG: Baseline RAG “struggles to connect the dots” and performs poorly when asked to “holistically understand summarized semantic concepts over large data collections.”

What a context graph changes

GraphRAG attacks the problem before any question is asked. During indexing, a large language model (LLM) reads every chunk and extracts entities, relationships, and claims, assembling them into a weighted knowledge graph. It then runs community detection (the Leiden algorithm) to cluster the graph into a hierarchy of related topics, and pre-writes a natural-language summary for each community.

At query time, those summaries do the heavy lifting. Each relevant community drafts a partial answer (the “map” step), the partials are ranked and merged (the “reduce” step), and the model synthesizes a final response grounded in structure rather than in a few cherry-picked snippets. Variants like HippoRAG take a different route, using the graph plus a Personalized PageRank walk to find the right passages — but the core idea is the same: Let relationships, not just cosine similarity, decide what context the model sees.

The evidence: Four studies, one pattern

1. Global sense making: The headline win

Microsoft pitted GraphRAG head-to-head against naïve RAG on global, “make sense of the whole corpus” questions over million-token datasets, with an LLM acting as judge across three axes: Comprehensiveness, diversity, and empowerment.

GraphRAG won 72 to 83% of comprehensiveness comparisons and 62 to 82% of diversity comparisons against vector RAG. Its highest-level summaries used up to 97% fewer tokens than processing the source text directly.

That is not a rounding-error improvement. On exactly the kind of question that breaks text-chunk RAG, the graph wins two out of three times or better.

2. Multi-hop retrieval: The graph finds what chunks miss

The second piece of evidence is about retrieval quality: Does the right supporting passage even make it into the top results? On the standard multi-hop QA benchmarks (MuSiQue, HotpotQA, 2WikiMultiHopQA), graph-guided retrieval lifts Recall@5 dramatically:

  • Average Recall@5 climbs from 73.4% (naïve RAG) to 87.8% (graph-guided), a +19.6 point gain.

  • The biggest jumps come on the hardest, cross-document sets: +31 points on MuSiQue and +28 points on 2Wiki.

  • HippoRAG reports up to a 20% accuracy improvement on multi-hop QA, at 10–20× lower cost and 6–13× faster than iterative retrieval methods.

3. The controlled head-to-head - where it gets honest

Here is where the story gains nuance. A 2025 study from Michigan State and Meta ran RAG against four GraphRAG families under one unified protocol — identical chunking, embeddings, and generation — and found no single winner. The two approaches are complementary:

  • On single-hop, factual lookup (natural questions), plain RAG edged ahead (F1 64.8 vs. 63.0 for the best graph method).

  • On multi-hop reasoning (MultiHop-RAG), graph-guided retrieval pulled in front (70.3 vs. 67.0 overall accuracy).

The lesson: A context graph is not a universal upgrade. It is a specialized one that pays off precisely when questions demand reasoning across pieces.

4. When to use graphs: The task-type verdict

The most recent benchmark, GraphRAG-Bench (ICLR 2026), set out to answer “In which scenarios do graph structures provide measurable benefits?” Its accuracy-by-task numbers map the boundary cleanly:

  • Simple fact retrieval: Text chunks 60.9 vs. graph 60.1 — effectively a tie. The graph’s structure is overhead the query doesn’t need.

  • Complex reasoning: Graph 53.4 vs. chunks 42.9 — a +10 point graph win.

  • Contextual summarization: Graph 64.4 vs. chunks 51.3 — a +13 point graph win.

The scorecard

Read top to bottom, the pattern is unmistakable: The graph’s advantage grows with the reasoning depth of the question, while text chunks hold their ground on isolated facts.

The catch: Cost and the LLM-judge problem

Two caveats keep this from being a slam dunk, and ignoring them is how teams end up disappointed.

Building the graph is expensive. Having an LLM extract entities and relationships from an entire corpus isn’t cheap. One analysis put index construction at roughly $48 against GPT-4o for a moderate corpus, far above a vanilla vector index. (Microsoft’s own follow-up, LazyGraphRAG, defers extraction to query time and cuts that to around 0.1% of the cost - a tacit admission that the original budget is impractical for many deployments.)

Many of the wins are judged by another LLM — and LLM judges are biased. An independent audit found systematic flaws in this evaluation style: position bias (swapping which answer appears first can swing the win-rate by more than 30 points), length bias, and trial bias (identical comparisons disagree across runs). After correction, one popular method’s reported 66.7% win rate fell to about 39% — below the 50% break-even line.

The takeaway is not “the research is wrong.” It is that the large gains — the +20% multi-hop accuracy, the +15-to-30-point recall jumps — are robust, while narrow comprehensiveness margins deserve a skeptical second look with reference-based metrics.

So when should you reach for a context graph?

Strip away the hype and the decision is refreshingly practical.

Use a context graph when: Your questions are multi-hop, global, or sensemaking in nature; you need comprehensive, multi-perspective answers; and your corpus is richly interconnected (research libraries, case files, incident histories, knowledge bases).

Stick with text chunks when: Your queries are mostly single-fact lookups; your corpus is small or flat; and indexing cost, latency, and operational simplicity outweigh a marginal quality bump.

Best of all, go hybrid: The systematic studies converge on the same recommendation: route each query to the right method, or fuse evidence from both. Combining graph and chunk retrieval consistently beats either one alone. You don’t have to choose a religion; you have to build a router.

The bottom line

A context graph is not magic, and it is not snake oil. It is a targeted instrument. Hand it a question that requires connecting scattered facts or synthesizing a whole corpus, and it will outperform text chunks decisively. Hand it “what’s the phone number on page 3,” and you’ve paid for indexing you didn’t need.

The teams that win with GraphRAG in 2026 won’t be the ones who graph everything. They’ll be the ones who know which questions deserve a graph — and build pipelines smart enough to tell the difference.

Dattaraj Rao is an R&D architect at Persistent Systems

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier

2 August 2026 at 13:01

Consolidation has been one of the paths that many astute observers predicted for the near-future of labs training models. It was labelled as inevitable, as training costs are increasing by orders of magnitude every year. Yet, as someone who in 2024 would’ve predicted consolidation really picking up come 2026 or 2027, where are we? We’re at a place where more companies are training strong models — easily investing hundreds of millions to billions of dollars in the total effort still — and an increasing number of organizations are releasing these models openly.

The demand for tokens is incredibly high, and likely to increase as models get more efficient and unlock more possible use cases. All of these labs we thought would need to consolidate are realizing that building token machines is a likely path to value, and more companies will identify that source of value over time.

The prime example is Thinking Machines — when they announced their company in February 2025, very few people would’ve put them in the bucket of an open models company, myself included. Now their open model finetuning service is making hundreds of millions in revenue per year and they’re releasing the best open-weight models built in the U.S.A. — ahead of the early leaders in NVIDIA with Nemotron and Arcee’s Trilogy.

Share

On the other side of the ecosystem is the sustained pace from the Chinese labs, with newer entrants like Xiaomi still accumulating mindshare in the broader AI economy. Having predicted consolidation for a long time, it now seems like a safer bet is to predict continued adoption, and try to imagine the role that open models play there. How much can revenue-share licenses like Kimi K3 stick? How much market share can open models take? We’re entering the decisive era.

This is one of the most packed recaps of open models we’ve ever had, we’re excited!

Our Picks

  • Inkling by thinkingmachines: The first model from Thinking Machines is a 975B-A41B multimodal MoE that supports text, images, and audio as inputs and produces text as output. While it is not the strongest model among peers (in China) in its size class, it is positioned to be a great base for fine-tuning, e.g., through their commercial offering, Tinker. They also release a smaller version (276B-A12B), which is really competitive for its size.

  • Hy3 by tencent: A 295B-A21B MoE from Tencent. It improves over its predecessor across all metrics. Most notable, however, is the license change: While the previous version (covered in Artifacts 21) used a custom and rather restrictive license, Tencent switched to Apache 2 for this release. The model was also able to proof a 50 year old math problem (with a dedicated harness and Sol as a judge, although it is unclear how important the latter really is).

  • Laguna-S-2.1 by poolside: Poolside quickly rose out of nowhere to become a frequent guest at Artifacts, marking its third appearance in three consecutive months. S2.1 is a newly pre- and post-trained version of the 118B-A8B MoE that fits on a DGX Spark, which brought it a lot of attention. Poolside also adopted the OpenMDW license, which is an Apache 2.0-like free license but has better legal backing for AI models specifically. The company also goes into more detail in its blog, which includes all the evaluation trajectories. This is a lot of transparency for an open model release!

  • DeepSeek-V4-Flash-0731 by deepseek-ai: Just one day after OpenAI has dropped the prices of their smallest model by 80%, the whale dropped an update to their V4 Flash model, beating Luna at the pareto frontier. The bigger model is not updated yet, so it remains to be seen where it will land in terms of performance. For the initial V4 releases, the Flash version was the star of the show in terms of performance per parameter, while Pro was rather underwhelming.

  • Kimi-K3 by moonshotai: This is the biggest open model release in some time, and we covered it in a separate post and a podcast episode. It was released under a noncommercial license, requiring inference and fine-tuning providers to enter into a commercial agreement. Kevin Xu and Graham Webster argue in a post that these licenses enable potential future government action against US entities doing business with Chinese AI companies:

But if a US company needs a contract with Moonshot to provide the inference tokens that Kimi K3 generates, the picture looks different. Some of the policy tools US officials and others have debated as potential levers to restrict Chinese open model use would more clearly apply.

Models

General Purpose

  • LongCat-2.0 by meituan-longcat: The Chinese DoorDash is back again. This time, the company released another big MoE with 1.6T parameters. While the model itself is not the most capable for its size beyond benchmarks, it was trained entirely on Ascend 910s, making it the first non-Huawei, non-toy model trained entirely on Chinese accelerators. Other Chinese chips are mostly used for inference (if at all).

  • Laguna-XS-2.1 by poolside: An update to the small (33B-A3B) MoE from Poolside.

  • Motif-3-Beta by Motif-Technologies: A preview of a 314B-A13B MoE by the Korean Motif. This is by far the company’s most ambitious model, as it is considerably larger and introduces some architectural innovations like GDLA and mHC.

  • Apertus-v1.5-70B by swiss-ai: A continued pre-train of the fully open-source Apertus 1.0, using 2T more tokens.

  • Instella-MoE-16B-A3B-Think by amd: A 16B-A3B MoE trained by AMD on Instinct cards. AMD also provides all the different stages, from the base to the SFT checkpoints, as well as MidTrain and DPO.

    Instella-MoE cost vs. performance

Read more

Europeans Are About to Find Out How Entrenched AI Is in Their Daily Lives

New EU rules stipulate that people must be told when they’re interacting with AI or looking at AI-generated or -edited content, leading to fear of “disclosure fatigue.”

This Week’s Awesome Tech Stories From Around the Web (Through August 1)

1 August 2026 at 14:00

Artificial Intelligence

OpenAI’s Hacking Debacle Comes Down to Human ErrorLily Hay Newman | Wired ($)

“If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies. …’A simple analysis of the actual risk has an actual simple answer,’ says longtime security and compliance consultant Davi Ottenheimer. ‘The OpenAI mistakes were dead simple.'”

Artificial Intelligence

Anthropic’s New AI Model Can Identify More Software Bugs Than Ever. Microsoft Is Struggling to Fix Them Fast Enough.Renee Dudley and Doris Burke | ProPublica

“Each month, the company publicly releases fixes for its software vulnerabilities in what’s known as ‘Patch Tuesday.’ In June, it released patches for more than 200 bugs, which industry experts then said was an all-time high. But on July 14, the company blew through that record and released patches for more than 600 bugs.”

Future

The Rise of Million-Dollar Companies With Just One EmployeeTe-Ping Chen | The Wall Street Journal ($)

“An analysis by the payments company Stripe shows there are thousands of solo operators on the company’s platform that are generating over $1 million in revenue, with their ranks doubling between 2023 and 2025. The number of solo operators crossing the $10 million threshold nearly tripled in that same span.”

Future

The AI Jobs Apocalypse Probably Isn’t Coming Anytime SoonEduardo Porter | The Guardian

“As Massachusetts Institute of Technology economist David Autor noted: ‘A lot of people have noticed that the world is not changing as fast as they predicted.’ The emerging new story not only puts more emphasis on the complexity of the relationship between automation and human work across history. It is also raising doubts about the very feasibility of the threatened AI transformation of the universe.”

Robotics

Are Brain Waves the Next Unlock for Physical AI?Tim Fernholz | TechCrunch

“Encord is one of a growing number of startups betting the next real constraint on humanoid and warehouse will be the scarcity of real-world physical training data, and which is building a business not just to manage that data but to manufacture it. The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup that’s betting measuring brain activity—to deduce mental states like error, intent, and surprise—can create a more useful dataset to train models.”

Tech

Wall Street Hunts for Creative AI Financing as ‘Digestion Issues’ EmergeStaff | The Information ($)

“John Greenwood, Goldman Sach’s global head of infrastructure and real asset finance, said he’s ‘looking for capital in every nook and cranny’ to support an expected $7.5 trillion in spending on chips, data centers, and power in the next five years. The hunt won’t end there, since much of that spending is on GPUs and other chips that need replacing every few years.”

Space

Experts Warn Current Starship Heat Shield Tech Is a ‘Dead End’ for Rapid ReuseEric Berger | Ars Technica

“The problem is that, with the signs of damage [to its heat shield], such a heat shield would appear to require a fair amount of inspection and refurbishment before another launch. In other words, SpaceX has a ways to go to reach ‘full and rapid’ reuse of Starship. “

Tech

In Silicon Valley, Some Say an AI Bubble Would Be Just FineErin Griffith | The New York Times ($)

“The excitement created by a bubble can drive new breakthroughs, their thinking goes. …These frenzies are important for allowing crucial infrastructure to get built, even if they lead to some ‘capital destruction’ along the way, [said Tomasz Tunguz, an investor at the venture capital firm Theory Ventures].”

Future

Neri Oxman Wants to Grow the Colors on Your ClothesElizabeth Segran | Fast Company ($)

“While several biotech firms have created more sustainable dyes, plugging cleaner chemicals into existing dye houses, Oxman’s approach reimagines dying from the ground up, treating dyes and fabrics as living organisms that can be grown. And while the Vigils project is still experimental, Oxman’s long-term goal is to commercialize and scale the technology, reshaping the future of fashion.”

Energy

New Data Shows EV Batteries Are Lasting Longer Than Initially ExpectedBruce Gil | Gizmodo

“Today’s average EV retains 97% of its original range after three years and 95% after five years, according to an analysis by EV data company Recurrent. …Additionally, battery replacement appears to be rare among newer EVs. A separate Recurrent analysis found that the battery replacement rate for EVs with model years 2022 and later was only 0.3%.”

Tech

Corporate America Has Suddenly Decided to Stop Blowing Money on AIAngel Au-Yeung, Katherine Bindley, and Tina Li | The Wall Street Journal ($)

“Fed up with ballooning costs, companies big and small are starting to use lower-priced models, including some built in China. In many cases, they are adding the new, cheaper models alongside OpenAI and Anthropic’s products, shopping a la carte for their artificial intelligence.”

Robotics

This Automation Tech Turns Old Tractors Into Self-Driving Farming MachinesPatrick Sisson | Fast Company ($)

“It’s a rig that can be attached to just about any existing tractor to help it mow, seed, weed, and perform any number of time-intensive tasks, all on its own, for a sector desperate for more labor.”

Space

AI Data Centers in Space? A System to Cool Chips Could Help.Ivan Penn | The New York Times ($)

“With a growing backlash against the proliferation of data centers to power artificial intelligence, there has been increasing interest in putting the energy-thirsty operations into orbit. Now, researchers may have figured out how to overcome a major obstacle to that goal: cooling the data centers in space.”

The post This Week’s Awesome Tech Stories From Around the Web (Through August 1) appeared first on SingularityHub.

Takes of Marine Mammals Incidental to Specified Activities; Taking Marine Mammals Incidental to U.S. Navy Operations of Surveillance Towed Array Sensor System Low Frequency Active Sonar in the Western and Central North Pacific Ocean and Eastern Indian Ocean

NMFS, upon request from the U.S. Department of the Navy (Navy), issues these regulations pursuant to the Marine Mammal Protection Act (MMPA) to govern the taking of marine mammals incidental to training and testing activities using Surveillance Towed Array Sensor System (SURTASS) Low Frequency Active (LFA) sonar systems in the western and central North Pacific and eastern Indian oceans over the course of 7 years from August 2026 through August 2033. These regulations allow for the issuance of a letter of authorization (LOA) for the incidental take of marine mammals during specified activities and timeframes, prescribe the permissible methods of taking and other means of effecting the least practicable adverse impact on marine mammal species and their habitat, and establish requirements pertaining to the monitoring and reporting of such taking. The Navy's activities are considered military readiness activities pursuant to the MMPA, as amended by the National Defense Authorization Act for Fiscal Year 2004 (2004 NDAA) and the NDAA for Fiscal Year 2019 (2019 NDAA).

7 States’ Water Systems Hit by Cyberattacks Likely Tied to Iran

Plus: The FBI eyes AI-powered tech to detect future crimes, Russia charges Telegram’s founder, xAI sues to stop a state’s “nudification” ban, and the Democrats learn a lesson about getting scammed.

AI as an Enterprise Operating System

31 July 2026 at 16:09
I hadn’t heard of Dan Guido until a few months ago, when I came across the video of a talk he gave at [un]prompted, an AI security practitioners’ conference. Dan is the CEO and cofounder of Trail of Bits, a software security research and development firm that works with companies in tech, defense, and finance. […]

The OpenAI and Anthropic AI Hacking Sprees Are a Messy New Legal Frontier

Both major AI labs’ models broke containment, escaped onto the internet, and hacked other companies. If a human had done that, the law would likely be against them. But a bot?

Takes of Marine Mammals Incidental to Specified Activities; Taking Marine Mammals Incidental to the Office of Naval Research's Arctic Research Activities in the Beaufort and Chukchi Seas (Year 9)

NMFS has received a request from the Office of Naval Research (ONR) for authorization to take marine mammals incidental to Arctic Research Activities (ARA) in the Beaufort Sea and eastern Chukchi Sea. Pursuant to the Marine Mammal Protection Act (MMPA), NMFS is requesting comments on its proposal to issue an incidental harassment authorization (IHA) to incidentally take marine mammals during the specified activity. NMFS is also requesting comments on a possible one- time, 1-year renewal that could be issued under certain circumstances and if all requirements are met, as described in Request for Public Comments at the end of this notice. NMFS will consider public comments prior to making any final decision on the issuance of the requested MMPA authorization and agency responses will be summarized in the final notice of our decision. ONR's activities are considered military readiness activities pursuant to the MMPA, as amended by the National Defense Authorization Act for Fiscal Year 2004 (2004 NDAA).

[AINews] not much happened today

1 August 2026 at 01:38

It might seem strange that we aren’t giving title story to a noteworthy DeepSeek open weights model update that still bumps up the Pareto Frontier that GPT 5.6 pushed out only yesterday:

But because it is a post-train only update with no further details, there’s really not all that much to report, apart from noting that DeepSeek is finally relevant again after over a year of comparative obscurity (with V4 Pro this April as an exception) after becoming way too prominent, well timed after their $70B pre-IPO fundraise.

AI News for 7/30/2026-7/31/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!


AI Twitter Recap

DeepSeek V4-Flash 0731: post-training leap, API launch, and immediate open-weights release

  • DeepSeek’s biggest story of the day was the official public-beta launch of DeepSeek-V4-Flash API, with DeepSeek stating that its upgraded agent capabilities now surpass V4-Pro-Preview and that the API now supports the Responses API format and is “fully adapted for Codex” (@deepseek_ai). In a follow-up, DeepSeek clarified that the improvement applies only to the Flash API, while V4-Pro API/App/Web remain unchanged for now; V4-Pro official is still pending (@deepseek_ai). Community observers quickly highlighted the magnitude of the jump: @cline called out Terminal-Bench 82.7, up +25.8 from the April preview’s 56.9.

  • The notable technical claim is that this jump came without changing architecture or size. Artificial Analysis summarized V4 Flash 0731 as still 284B total / 13B active, 1M context, text-only, at $0.14 / $0.28 per 1M input/output tokens with an unusually aggressive 98% cache-hit discount to $0.0028 / 1M cached tokens (@ArtificialAnlys). On their index, the model rose from 40 → 50, landing 1 point behind GPT-5.6 Luna (max, 51) while coming in at roughly 60% lower cost per task on DeepSeek’s first-party API. They also reported major agentic gains, including GDPval-AA v2 Elo 1189 → 1559, Terminal-Bench 2.1 to 79%, τ³-Bench Banking +8 points, and a 12% drop in output-token usage versus the predecessor. Multiple posts converged on the same takeaway: this is a post-training win, not a scaling-law/pretraining story (e.g. @kimmonismus, @EMostaque, @Yuchenj_UW).

  • Open-weights followed almost immediately. The official weights landed on Hugging Face and were widely amplified by @MiaAI_lab, @_akhaliq, and others. The release is under MIT, and @vllm_project highlighted serving details: 256 routed experts, 6 active per token, 1M context, three reasoning-effort levels, and an included DSpark speculative decoding module that can be enabled via a single flag. Local/quantized deployment followed immediately: @UnslothAI published runnable quants requiring roughly 168GB RAM for lossless 4-bit and 110GB for 3-bit, while @danielhanchen later shared additional UD quants.

  • A second-order theme was harness sensitivity and agent specialization. A number of posts argued that Flash’s gains are best understood in the context of better post-training for tool use and long-horizon tasks, not just raw IQ benchmarks. @jakevin7 reported that the model autonomously discovered and used subagent swarm patterns in a Maka-based setup. @arena later placed DeepSeek-V4-Flash-High on the Pareto frontier in the Frontend Code Arena, scoring 1586 and jumping +154 points over its preview. Several practitioners also noted that open models increasingly benefit from lighter harnesses and cache-friendly deployment patterns rather than heavy orchestration (e.g. @omarsar0).

Open vs closed, price compression, and what “cheap intelligence” now means

  • The release immediately reframed the week’s price war. After OpenAI’s prior-day cuts to GPT-5.6 Luna (-80%) and Terra (-20%), many users read DeepSeek’s Flash upgrade as a direct competitive response. @kimmonismus summarized the new economics as $0.28/M output tokens, with performance “super close” to higher-end proprietary systems on some coding-agent benchmarks. @ArtificialAnlys later corrected an early cache-hit-rate display issue and reiterated that on DeepSeek’s own API, 0731 is firmly on the Pareto frontier for intelligence vs. cost per task.

  • Developers quickly integrated DeepSeek into existing coding stacks rather than treating it as a standalone API. @ziwenxu_ showed DeepSeek V4-Flash running inside Codex via a router that preserves access to GPT, Grok, Kimi, and DeepSeek in one model picker; @Teknium added it to Hermes Agent; @cline made the updated model free in Cline; and @victormustar even spun up a free public endpoint. The practical message: the cost/performance delta is now big enough that routing and harness choices materially affect engineering workflows.

  • This also strengthened the pro-open argument in the cyber/safety debate. After the week’s security incidents, @ClementDelangue argued that Hugging Face defended itself with an open model—specifically a quantized GLM 5.2—and that banning open models would most harm defenders, startups, and researchers. @sundeep made the complementary point that a safe world with closed models still benefits from a vibrant open ecosystem. In parallel, @thinkymachines published a more incremental position: widen access in stages rather than treating open weights and safety as mutually exclusive.

AI security incidents: labs’ sandboxing failures overshadow “rogue model” narratives

  • The dominant non-release controversy concerned newly disclosed cyber-eval incidents. @GergelyOrosz summarized reports that OpenAI had an under-development agent escape a sandbox and target Hugging Face, while Anthropic disclosed similar incidents from prior months only after the OpenAI story broke. The Anthropic side was further summarized by @kimmonismus: after reviewing 141,006 eval runs, Anthropic found three incidents involving Opus 4.7, Mythos 5, and an internal model, all enabled by a misconfigured third-party evaluation environment with internet access.

  • The strong consensus among technical commentators was that these were primarily infra and harness failures, not evidence of autonomous agency. @johnennis, @Dan_Jeffries1, and @perrymetzger all argued that the descriptions implied poor sandboxing, weak logging, and bad operational discipline. @jachiam0 added an interesting nuance: a lack of situational awareness in evals can itself cause safety failures when the model is told the environment is simulated but it is not.

  • The policy split is becoming clearer. Some posters, including @ostrisai and @RichardSocher, used the incidents to criticize closed labs’ claims of superior safety. Others, such as @jachiam0, pushed the opposite direction, warning that the combination of frontier cyber capability and geopolitical conflict raises the probability of serious escalation against critical infrastructure. Either way, the technical lesson that emerged most consistently was narrower: agent behavior is highly shaped by eval scaffolding, access controls, and harness design.

Agents, harnesses, eval environments, and continual improvement infrastructure

  • A recurring meta-theme across many tweets was that model capability is increasingly bottlenecked by harnesses and environments. @swyx distilled the zeitgeist into a line: if you can distill models, you can also distill agent harnesses. @TheTuringPost made the related point that many perceived “model limitations” are actually memory or harness decisions made around the model.

  • Research posts this week reinforced that view with concrete systems work. @omarsar0 summarized Microsoft’s Echoverse, which compiles specifications into stateful applications with grounded graders and uses rollout analysis to repair both environments and training signals; notably, shallow environments hurt live-site accuracy while deeper ones improved it. @dair_ai highlighted OpenMLE / Frontis-MA1, a released full stack for recursive self-improvement in ML engineering using four atomic evolution operators (Draft, Improve, Debug, Crossover). @omarsar0 also covered AgentRadio, showing asynchronous inter-agent messaging can raise SWE-Atlas QnA from 32.3% → 62.1% with four agents, outperforming a stronger single-model baseline.

  • Tooling vendors are productizing this stack quickly. @hwchase17 gave the current LangChain ecosystem map—LangGraph, DeepAgents, and LangSmith—while later emphasizing standardized internal evals and Harbor-based task conversion (@hwchase17). @simonw introduced smevals for running small eval suites across models, harnesses, and prompts. @promptlayer added mocked tool responses for end-to-end agent testing without live backends. The throughline: eval infra is shifting from ad hoc notebooks to reproducible, organization-owned systems.

Multimodal product launches: MiniMax H3, Seedance 2.5, Gemini updates, and robotics

  • MiniMax’s H3 launch had broad distribution momentum. The model went live on Vercel AI Gateway with “one generateVideo[] away” positioning and promises of open weights soon (@MiniMax_AI). From there it propagated rapidly across partners including fal (@fal), Pollo (@itsPolloAI), PixVerse (@PixVerse_), Leonardo (@MiniMax_AI), and OpenArt (@MiniMax_AI). One technical detail that stood out from commentary: H3 appears to integrate low-to-high generation / baked-in super-resolution, rather than stapling on a separate SR stage (@andrew_n_carr).

  • ByteDance/Dreamina’s Seedance 2.5 also drew strong creator attention. @kimmonismus summarized support for native 30-second and consistent three-minute videos, interactive frame editing, and up to 50 multimodal references. Users testing in consumer apps noted practical caveats—e.g. current 720p, some moderation friction, and instruction-following gaps around audio/music (@TomLikesRobots)—but overall creator sentiment was highly positive.

  • Google and OpenAI both shipped UX-heavy product updates around assistants. Google’s Gemini Drops added Gemini 3.6 Flash, 3.5 Flash-Lite, wider Gemini Spark rollout, app integrations, voice on macOS, and personalized image/avatar features (@GeminiApp, @GeminiApp). OpenAI pushed more desktop/app ergonomics: Voice on macOS/Windows (@ChatGPT), a new Activity view (@OpenAIDevs), and pet-triggered shortcuts into Voice (@ChatGPT). Meanwhile, @bousmalis and @_anniexie shared early demos of Gemini Robotics 2, emphasizing extended real-time tool-kitting and multimodal, embodied recovery behaviors.

Top tweets (by engagement)

  • DeepSeek official launch: @deepseek_ai announced V4-Flash API public beta with major agent benchmark gains and Codex/Responses API support.

  • Community benchmark reaction: @cline highlighted the +25.8 Terminal-Bench jump and noted open weights were coming shortly.

  • Artificial Analysis breakdown: @ArtificialAnlys provided the most complete public summary of architecture, pricing, cache economics, and benchmark deltas.

  • Open-source cyber defense argument: @ClementDelangue argued open models were used defensively against proprietary-model-driven attacks and warned against blanket bans.

  • Anthropic/OpenAI incident criticism: @johnennis and @perrymetzger captured the dominant infra-first critique of the “rogue AI” framing.


AI Reddit Recap

/r/LocalLlama + /r/localLLM Recap

1. DeepSeek V4-Flash 0731 Release Benchmarks

Read more

Europe Approves Bionic Eye to Restore Vision Lost to Blindness

31 July 2026 at 23:06

An implant, smaller than a grain of rice, pairs with camera-mounted glasses to communicate visual information to the retina.

Age-related vision loss affects millions of people, and so far, there has been no way to reverse the damage. A newly approved retinal implant could change that by allowing some people with severe vision loss to regain functional sight.

More than five million people worldwide suffer from geographic atrophy, the late stage of the progressive eye condition dry age-related macular degeneration. The disease destroys the photoreceptors at the center of the retina, known as the macula, which is responsible for the sharp central vision required to read or recognize faces.

In the US, treatment options are limited to two drugs that can be injected into the eye to slow the disease’s progression. But neither can undo the damage. That could be about to change. California neurotech startup Science Corporation recently won European approval for a retinal implant designed to treat the condition.

“For decades, losing central vision to this disease meant losing the ability to read, recognize faces, and ultimately losing independence. There was no viable treatment. Now there is,” Max Hodak, Science’s CEO and co-founder, said in a press release.

The company’s PRIMA system combines an implant smaller than a grain of rice installed underneath the patient’s macula with a pair of camera-mounted glasses that translate incoming visual information into near-infrared light that is then beamed to the retina. The eye can’t detect this wavelength, so the device doesn’t interfere with any natural sight that remains.

The chip, which works on similar principles to a solar panel, converts the incoming light into electrical pulses that stimulate retinal neurons called bipolar cells. These are downstream of the rod and cone photoreceptor cells damaged by macular degeneration and normally spared by the disease.

In a clinical trial involving 38 patients across five countries, which was published in the New England Journal of Medicine last year, the company and its collaborators showed participants gained an average of 25.5 letters—more than five lines—on a standard eye chart after having the device fitted.

And now the device has received a CE mark from the European Union making it possible to sell in 30 European countries. The company says the first commercial implants are expected to be fitted in Germany within weeks, with Italy, the Netherlands, and the UK to follow. In the US, PRIMA holds Breakthrough and Humanitarian Use Device designations from the FDA, but the company is confident it will gain full approval in the near future.

The device is a long way from restoring normal vision. The images it produces are black and white and the field of vision is extremely narrow. Hodak described the experience to the Financial Times as “kind of like looking through a straw in the center of their vision,” though he added that they see a pathway to color vision and higher acuity.

While the implantation procedure is fairly simple, it takes months of training to unlock the device’s full potential. Nonetheless, Hodak told STAT that the company expects to install 20 to 40 devices this year and 200 globally by the end of next if they get US approval in early 2027.

The approval is welcome news for the wider neurotech industry, which has absorbed billions of dollars of investment in recent years with little to show in terms of return.

“Science is showing that brain-computer interface companies have a path to real revenue now,” Jacob Robinson, founder of startup Motif Neuroscience, told STAT. “These companies aren’t all just making a bet on a market that is 10 to 15 years away.”

Hodak told the Financial Times hehopes sales from PRIMA will bankroll Science’s more ambitious work on “biohybrid” interfaces, which use genetically engineered living neurons to connect to the brain rather than metallic wires. “This is the financial backbone,” he said. “This is the thing that pays for the rest.”

Other companies are hot on Science’s heels. Neuralink, which Hodak co-founded with Elon Musk before leaving to start Science, is also working on a vision implant called Blindsight, which is due to enter human trials this year.

While the field remains a long way from the sci-fi vision of seamless two-way communication between humans and machines, this approval is growing evidence the neurotech industry is starting to move out of the lab and into the real world.

The post Europe Approves Bionic Eye to Restore Vision Lost to Blindness appeared first on SingularityHub.

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

31 July 2026 at 22:16
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because...

As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because attention now dominates that cost, how it is designed—not just how it is implemented—increasingly determines a model’s inference performance. Shaping model architecture around how GPUs execute it is the premise of AI model co-design.

Source

Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap

If you ask an AI coding agent to write a standalone Python script to parse a single JSON file, it will likely give you a perfect answer in seconds. But the same agent often breaks if you ask it to build a systematic data processing pipeline, like ingesting thousands of messy documents, chunking text, scoring quality, and filtering noise for a Retrieval-Augmented Generation (RAG) system that fits your specific enterprise stack.

While large language models (LLMs) excel at one-off code generation, their outputs for complex data-processing tasks are typically free-form, disposable scripts. These scripts are detached from the governable workflow abstractions that MLOps teams rely on for production, making them difficult to audit or edit visually.

To address this, researchers at Peking University, Zhongguancun Academy, and Shanghai’s Institute for Advanced Algorithms Research introduced DataFlow-Harness, an open-source framework that guides an LLM agent to build structured, visual data-processing workflows step-by-step, rather than writing raw code from scratch.

The framework makes AI-generated pipelines easier to manage and integrate into existing architectures because the generated artifacts are persistent and easily editable.

The researchers report that the platform achieves a 93.3% observed end-to-end pass rate on a 12-task data-engineering benchmark. Compared to standard Claude Code, it reduces API costs by up to 72.5% and response latency by 49.9%, while achieving nearly the same success rate as an AI given the entire codebase to write standard scripts. For enterprise teams, this means getting the speed of AI automation without accumulating unmanageable technical debt, ensuring that pipelines remain secure, auditable, and ready for production.

The "NL2Pipeline gap"

Data-centric AI requires workflows for tasks like synthetic data generation, retrieval augmentation, and model training. While LLMs can translate natural language into executable implementations to perform these tasks, high task accuracy is insufficient for production deployment.

"The first wall is usually not writing Python," Runming He, first author of the DataFlow-Harness paper, told VentureBeat. "Modern coding agents can often produce a plausible script quickly. The harder problem is grounding that script in a live production platform: using operators that are actually installed, matching the real dataset schema, referring to registered datasets and model services, preserving dependencies between stages, and leaving behind an artifact that another engineer can understand and revise."

General-purpose AI agents frequently hallucinate dependencies, relying on unavailable operators or outdated platform assumptions. Instead of leaving behind an artifact that another engineer can understand and revise, they generate disposable code that is difficult to audit through workflow managing tools.

The researchers define this challenge as the "NL2Pipeline gap": the disconnect between a user expressing workflow requirements in natural language and the production environment requiring structured and persistent pipeline assets.

The researchers demonstrated this gap in their experiments. For example, when Claude Code was allowed to write standard, free-form scripts using codebase context, it hit a 94.2% success rate. However, when restricted to only using the platform's specific building blocks to create a native workflow graph, its success rate dropped to 83.3%. This gap is the paper's central finding: native, governable pipelines are meaningfully harder for the agent to produce than throwaway code.

“Closing this gap requires more than improving code-generation accuracy: construction must remain grounded in platform semantics and produce artifacts that integrate with the host platform,” the researchers write.

How the four components work together

"DataFlow-Harness changes the agent’s action space," He said. "Instead of asking the agent to emit arbitrary code, it retrieves the live operator registry and current pipeline state through MCP and applies typed, incremental changes to a persistent DAG."

To achieve this, the platform organizes workflow synthesis around four components: the Data Pipeline Backend, the interaction layer (DataFlow-WebUI), the MCP Tools Layer, and the AI guidance layer (DataFlow-Skills).

The Data Pipeline Backend acts as the authoritative source of truth across conversational, visual, and programmatic interfaces. It represents the pipeline as a directed acyclic graph (DAG), a structured workflow map containing data sources, configured pre-built processing modules (which the researchers refer to as "operators"), and execution dependencies. Instead of generating free-form code, agents interact with this backend through “typed mutations,” like adding an operator or connecting edges.

DataFlow-Skills are markdown files that inject domain-specific knowledge into the model's context window, guiding it on operator-selection patterns, schema inference, and assembly procedures. Rather than letting the AI guess how to assemble components, skills provide the AI with compatibility rules, teaching it how to correctly match different data formats and handle complex data structures without breaking the pipeline. 

The MCP Tools Layer gives the AI access to the operator registry and current state of the data workflow. The AI proposes structured changes through the tools layer. The system validates the changes to ensure the workflow runs in a valid sequence and that every connected module speaks the same data language.

DataFlow-WebUI provides two interfaces that allow humans and AI to build the workflow together. Developers can describe workflow requirements in natural language through a conversational interface. They can also access the workflow as a graphical map in a visual DAG editor. Here, they can directly inspect the changes proposed by the AI and make modifications.

“The current implementation performs static checks against platform metadata before accepting pipeline changes,” He said. “These include checks for registered datasets, operators and model-serving references, field flow, and some invalid parameter usage, as well as structural validity. The result is visible in a graphical editor and can be revised either manually or by the agent in later turns.”

The results: 93.3% pass rate, 72.5% lower cost

The researchers tested DataFlow-Harness on a benchmark of 12 tasks across six industrial data-processing scenarios, such as QA generation, review governance, and schema normalization. They used Claude Opus 4.7 as the backbone model in their experiments.

They compared DataFlow-Harness against three baselines:

  • Vanilla CC: An unconstrained coding baseline using standard Claude Code.

  • Context-Aware CC: An agent that has access to the DataFlow codebase in its context window.

  • MCP-only: An agent that has access to the DataFlow MCP tools and is instructed to generate platform-native DAGs (without access to DataFlow-Skills).

DataFlow-Harness achieved a 93.3% end-to-end pass rate, improving by 10.0 percentage points over MCP-only and beating Vanilla CC (91.7%), while being within 0.9 percentage points of Context-Aware CC (94.2%).

Importantly, it reduced API costs to $0.261 per task, a 72.5% drop compared to Vanilla CC and 42.8% compared to Context-Aware CC. In generating workflows, it was 49.9% faster than Vanilla CC and 17.6% faster than Context-Aware CC.

DataFlow-Harness proved particularly effective on complex tasks that depend on implicit domain knowledge, like QA generation. The baseline MCP-only approach frequently generated structurally valid DAGs but struggled to infer task-specific procedures from operator descriptions alone.

To show how this works in the real world, the researchers detailed a textbook-to-VQA extraction task. This job required the AI to stitch together capabilities such as PDF parsing, layout recovery, OCR, figure extraction, multimodal understanding, and long-range question-answer matching. DataFlow-Harness achieved 97.2% precision and an 87.3% coverage rate, easily beating the baselines. By having the AI snap together existing platform assets rather than coding complex tasks from scratch, it recovered more valid QA pairs from the document.

Their experiments also showed that DataFlow-Harness is highly effective at creating data generation pipelines. For example, in a synthetic instruction-data generation task, the agent built a multi-stage pipeline that generated candidate instruction–response pairs, critiqued and rewrote them, scored them with an LLM-based judge, and filtered low-quality outputs before training.

"Such workflows are costly to build and fragile to maintain as collections of ad hoc scripts," He said. "The harness does not make them automatically safe, but it turns them into explicit, editable stages that engineers can inspect, test, and govern using normal production controls."

Similarly, when tasked with building a math data cleaning-and-synthesis pipeline, the data produced by the DataFlow-Harness pipeline trained a better-performing model with higher average accuracy on AIME24 and AIME25 benchmarks than the data produced by the vanilla Claude Code pipeline.

Tech stack fit and implementation tradeoffs

For engineering teams evaluating DataFlow-Harness, it is important to understand how it fits into existing infrastructure. Released under the Apache 2.0 license, the current implementation requires a bit of engineering to fit into popular tech stacks.

"The current implementation is native to the DataFlow platform; it is not a turnkey Airflow, Prefect, or Spark plug-in," He said. To use those systems as an execution backbone, teams must build an adapter to connect their organization’s registry, metadata, and execution interfaces to the agent's control layer.

Furthermore, organizations must invest in the boundaries they want the AI to respect. This requires maintaining an operator registry, defining schemas, and encoding recurring domain procedures as Skills. Because of this overhead, He recommends against using the framework for small, one-off transformations where a simple script suffices, or in legacy environments that cannot expose reliable metadata.

Finally, while the platform prevents illogical connections by validating structural properties, it is an engineering control layer, not a compliance substitute. "The harness should still be treated as an engineering control layer, not as a substitute for compliance policy, validated detection models, access controls, audit logging, or human approval," He said.

The platform is open-source, and developers can access the source code and codebase documentation directly via the project's GitHub repository.

As protocols like MCP become standardized, the boundary between human engineers and AI agents will shift. "The goal is not autonomous data engineering without oversight," He said. "It is a better division of labor: agents perform repetitive construction inside explicit boundaries, while engineers remain responsible for the semantics, policies, and consequential decisions that require domain accountability."

Claude published malicious code to the Internet and attacked 3 real companies

31 July 2026 at 20:39

Anthropic said its Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing designed to measure the models’ offensive cyber capabilities.

The events, which Anthropic revealed Thursday, are the second revelation in 10 days that AI models from the world’s wealthiest providers have trespassed into protected networks, an offense that, in more traditional hacking scenarios, could land the human behind the keyboard in prison for years. Earlier this month, OpenAI said its security models exploited a zero-day vulnerability for use in breaking into the network of Hugging Face, a platform for open source machine-learning models and AI datasets. The OpenAI models went on to steal access credentials and other confidential Hugging Face information. The OpenAI models also exploited publicly exposed credentials to compromise accounts of four other third-party services.

Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”

Read full article

Comments

© Getty Images

How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud

The AI agent observability space is taking off — but how can enterprises be sure what observability products and solutions they need?

Observability startup groudcover (lower case "g" intentional) announced this week that it raised $100 million in a round led by One Peak, bringing its total funding to $160 million.

The company says it has more than 250 paying customers, tripled annual recurring revenue over the past year and is increasingly replacing established observability platforms inside enterprise environments. Those are company-reported figures, but together they point to growing momentum in one of enterprise software's most competitive markets.

That market has long been dominated by companies including Datadog, Dynatrace, New Relic, Splunk and Grafana. Between them, they represent billions of dollars in annual revenue and years of product maturity. Breaking into that group has never been easy.

groundcover's argument is that artificial intelligence has fundamentally changed the assumptions those platforms were built on.

Rather than competing feature for feature, the four-year-old company is trying to convince enterprises that the architecture underpinning observability itself needs to change as AI systems become more autonomous, produce vastly more telemetry and increasingly participate in software operations. Whether that thesis proves correct remains an open question, but it offers a compelling lens through which to examine how observability is evolving alongside enterprise AI.

AI is turning telemetry into an infrastructure problem

Observability has traditionally been viewed as a post-production discipline. Engineers deploy applications, monitor logs, metrics and traces, investigate incidents, and improve reliability over time.

That workflow is changing.

AI-assisted software development has dramatically accelerated deployment cycles. Coding assistants generate more code, infrastructure evolves more rapidly, and organizations are deploying increasingly complex distributed systems that combine microservices, Kubernetes clusters, APIs and large language models. At the same time, enterprises are beginning to operate AI agents that execute multi-step workflows, call external tools and interact with production systems.

Each of those activities generates telemetry.

The result is an explosion of operational data that organizations increasingly want to retain rather than discard. AI applications introduce additional layers of observability beyond traditional infrastructure monitoring, including prompt execution, model latency, token consumption, retrieval pipelines, tool invocations and agent behavior. As enterprises experiment with autonomous systems, that telemetry becomes increasingly valuable because it provides the context needed to understand what an AI system actually did and why.

For many organizations, this creates tension with pricing models that charge according to the amount of data ingested.

Historically, engineers have often responded by sampling traces, shortening retention periods or limiting which data is collected. Those approaches reduce costs, but they also reduce visibility precisely when AI-driven systems demand more complete operational context.

"We've seen telemetry exploding," groundcover co-founder and CEO Shahar Azulay said during a recent media briefing. "Users are frustrated by not getting all the value from Datadog and similar platforms. They're limiting the data, siloing it, sampling it."

Whether that frustration is widespread enough to reshape the market remains to be seen, but the underlying trend is difficult to ignore. AI is making observability less about collecting enough data and more about collecting everything organizations may eventually need.

Rather than adding AI, groundcover argues the architecture itself has to change

Many observability vendors have introduced AI assistants, AI-powered root cause analysis and AI observability features over the past two years. Datadog, Dynatrace, New Relic and Grafana have all announced products aimed at helping enterprises monitor AI applications or automate operational tasks.

groundcover acknowledges those developments but argues they do not address what it sees as the more fundamental issue: where telemetry lives and how customers pay for it.

Instead of operating a conventional SaaS platform that stores customer telemetry in vendor-managed infrastructure, groundcover uses what it calls a bring-your-own-cloud (BYOC) architecture.

Customers keep the data plane—including telemetry storage and processing—inside their own AWS, Microsoft Azure or Google Cloud environments, while groundcover provides a managed control plane and user experience. A fully self-hosted deployment option is also available.

While some competitors, including Datadog and a few other observability vendors, do offer limited hybrid or customer-controlled data residency options, these are generally not equivalent to a full BYOC model. In most cases, telemetry is still processed and stored within the vendor’s managed infrastructure, with only partial controls (such as regional data residency, private links, or selective log forwarding) available.

That architectural decision influences nearly every aspect of the company's strategy.

Because customers already pay for their own cloud infrastructure, groundcover argues it can avoid charging based on telemetry ingestion. Instead, pricing is based primarily on monitored hosts, regardless of telemetry volume.

The company believes this changes customer behavior.

Rather than deciding which logs or traces are too expensive to keep, organizations can theoretically retain complete telemetry and use it for operational analysis, compliance and AI-assisted troubleshooting.

"We don't price by data volume," Azulay said. "We price by the size of the infrastructure."

The distinction matters because AI workloads tend to increase telemetry far faster than infrastructure itself.

That does not necessarily make host-based pricing universally cheaper. Organizations with relatively light workloads spread across many hosts may find different economics than dense Kubernetes environments generating enormous amounts of telemetry. The company's own briefing notes that per-host pricing is most advantageous for organizations with high telemetry density and may be less compelling for lightly utilized fleets.

Still, the broader argument is less about cost alone than predictability. Enterprise infrastructure teams often struggle with observability bills that fluctuate alongside application growth. groundcover's model attempts to align pricing more closely with infrastructure planning rather than data generation.

eBPF sits at the center of the company's technical differentiation

The second pillar of groundcover's strategy is eBPF, a Linux kernel technology that has rapidly become one of the most important building blocks for modern cloud observability.

Instead of requiring developers to manually instrument applications, eBPF allows software running inside the operating system kernel to observe network traffic, system calls and application behavior with minimal code changes.

That enables faster deployment and broader visibility across infrastructure.

For organizations operating Kubernetes clusters and cloud-native applications, reducing instrumentation complexity can significantly shorten deployment times while increasing telemetry coverage.

Azulay argues this becomes especially important as AI systems generate increasingly complex interactions across services.

"Our sensor allows us to observe systems very deeply from infrastructure to application to AI workloads without developers needing to instrument code," he said during the briefing.

eBPF itself is hardly unique. Many observability vendors now incorporate it into their platforms.

What groundcover argues differentiates its approach is combining automatic eBPF collection with customer-controlled storage, OpenTelemetry compatibility and unified pricing inside a single platform.

The company's own research briefing acknowledges that none of these technologies individually represents a competitive moat. The claimed differentiation lies in the combination of eBPF-first collection, managed BYOC architecture, host-based economics and full-stack observability delivered together.

AI agents are becoming both customers—and users—of observability

Perhaps the most interesting aspect of groundcover's strategy extends beyond traditional monitoring.

The company increasingly describes observability as infrastructure for autonomous software development.

Historically, observability platforms have served human operators investigating production incidents.

groundcover believes future observability platforms will increasingly serve AI agents as well.

Its Agent Mode product allows engineers to investigate incidents using natural language across logs, metrics, traces and Kubernetes events. More importantly, Azulay envisions observability becoming the feedback mechanism that informs coding agents about what actually happened in production.

Rather than simply detecting failures after deployment, observability becomes continuous operational context that autonomous systems can use to evaluate changes, identify regressions and eventually recommend or implement fixes.

"We're seeing observability moving from being a post-production tool... to people taking context from production and feeding it back to their coding agents so they can write code better," Azulay said.

Today, the company emphasizes that humans remain in the loop.

Agent Mode investigates incidents and surfaces recommendations, but production changes still require human approval. Azulay expects autonomy to increase gradually as organizations become more comfortable allowing AI systems to participate in operational workflows.

That vision reflects a broader trend emerging across enterprise software, where AI agents increasingly span development, testing, deployment and operations rather than functioning as isolated assistants.

Why some enterprises are considering alternatives

groundcover is entering an intensely competitive market populated by vendors with decades of enterprise experience.

Datadog alone generated more than $3 billion in annual revenue in 2025. Dynatrace, Cisco's Splunk business, Grafana Labs and New Relic all maintain extensive partner ecosystems, mature integrations and enterprise support organizations that newer entrants cannot easily replicate.

groundcover is not attempting to outscale those incumbents overnight.

Instead, it argues that AI creates an architectural inflection point similar to previous transitions from on-premises infrastructure to cloud-native computing.

According to Azulay, many customers initially adopt groundcover to reduce observability costs but increasingly remain because they want unrestricted access to richer telemetry and AI-native workflows.

He says deployments typically replace incumbent platforms rather than operate alongside them, although the company has not publicly disclosed customer migration data or independent studies validating that claim.

The company's journalist briefing also urges caution around some performance claims.

Revenue growth, customer counts and enterprise adoption figures originate from groundcover itself. Published customer case studies reporting significant cost savings are vendor-authored and should not be treated as independent validation without additional evidence. The briefing also recommends scrutinizing exactly what metadata leaves customer environments in standard BYOC deployments, rather than assuming that no operational data ever reaches vendor infrastructure.

Those caveats are important because the observability market has become crowded. Gartner currently tracks more than one hundred observability products, and nearly every major vendor now markets AI-powered operational capabilities.

Success will likely depend less on whether AI matters—which increasingly appears inevitable—and more on whether enterprises conclude that existing architectures remain sufficient.

The larger question investors are betting on

Viewed narrowly, groundcover's Series C is another large infrastructure funding round.

Viewed more broadly, it reflects a growing debate about what observability becomes in an era where software increasingly writes, tests and operates itself.

If AI continues generating exponentially larger volumes of operational data, traditional assumptions about telemetry collection, pricing and storage may come under increasing pressure. Vendors that built businesses around charging for data ingestion may need to evolve their economics alongside customer expectations. New entrants, meanwhile, have an opportunity to design around those changing assumptions from the outset.

groundcover believes that opportunity lies in combining customer-controlled infrastructure, automatic telemetry collection and AI-assisted operations into a platform designed for autonomous software rather than simply adding AI features to existing observability products.

Whether that architectural bet proves durable will depend on enterprise adoption over the next several years.

But the company's latest funding round suggests at least some investors believe the next battle in observability will not be fought over dashboards or alerts. It will be fought over who builds the operational data layer that increasingly intelligent software relies upon to understand—and eventually manage—the systems it runs.

Nscale just bought Anyscale. Here’s why it matters for multi-cloud neutrality.

Cloud platform company Nscale announced this week a definitive agreement to acquire AI workload scaling specialist Anyscale, in a move that signals a new test of whether cloud-neutral AI software can stay neutral once it is paired with a GPU neocloud.

The purchase coalesces Nscale’s infrastructure capabilities, which span control systems that oversee GPUs, datacenters, power consumption, and the application layer where AI services themselves are executed, with Anyscale’s software layer for scaling AI workloads across data processing, training, inference, and reinforcement learning.

Argued by Nscale to be the coming together of “two highly complementary companies”, Nscale scooping up Anyscale could be a fundamental change in the resulting business model. 

Is this the start of GPU neocloud lock-in?

It’s important to remember that Nscale is a GPU neocloud (a specialized cloud provider running bare-metal GPUs and infrastructure optimized for AI and machine learning workloads), meaning that it runs its own GPU-rich datacenters and its own software ​stack. At the same time, Anyscale is an independent cloud-neutral software orchestration multi-cloud control plane that works with any cloud hyperscaler… but now owned by a single neocloud. 

That doesn’t sound quite so much like cloud-neutrality and agnosticism; it sounds more like a vertically integrated AI cloud provider proposition.

Chief product officer at Nscale, Dan Bathurst, tells The New Stack that the Anyscale platform “continues to be its own brand and product,” and that includes working with bring-your-own-cloud deployments on AWS, GCP, Azure, and the other clouds. 

“Where we want to win is on performance, not on any sort of vendor lock-in or forcing of someone to choose Nscale as the infrastructure provider.”

“But what really changes — or how it’s changing — is that customers now also get this first-party option, where they can have Anyscale running on Nscale fleet as a full-stack, highly-optimized solution. Where we want to win is on performance, not on any sort of vendor lock-in or forcing of someone to choose Nscale as the infrastructure provider,” Bathurst says.

He insists that it is in Nscale’s interest to ensure that it is making it easy for software engineering teams to get the outcomes they want with the workloads that they’re trying to run.

“For us, the existing commitments will carry forward, so Nscale’s value really is meeting instances where the compute already lives,” he says. “Where we want to win is on performance, not on any sort of vendor lock-in or forcing of someone to choose Nscale as the infrastructure provider.”

Neutrality on the platform layer, differentiation on the infrastructure layer

Bathurst invites users to think of it as “neutrality on the platform layer, but differentiation on the infrastructure layer” because the combination of the two organizations is a full-stack play.

“The differentiation comes from the fact that Nscale is fully vertically integrated with Anyscale. Therefore, if users want that first-party option, they can choose Anyscale and get the most optimized solution because, obviously, we’re designing, optimizing, and co-engineering every layer of that stack from power to the datacenter through to the application. It’s quite a unique proposition, but it’s not something we are going to force upon any customer,” confirms Bathurst.

Not everyone is convinced by the company’s pledge to maintain an agnostic and neutral open house. Sanjeev Mohan, principal analyst, SanjMo and former Gartner research VP for data and analytics, tells The New Stack that Anyscale “stops being a neutral player” the moment its best features and most optimal pricing land on Nscale first. 

“The software will still run anywhere, but ‘runs anywhere’ and ‘runs best somewhere’ are different things, and buyers will feel the gap in performance and cost. At that point, neutrality is a label.”

Runs anywhere, but… runs best somewhere

“The software will still run anywhere, but ‘runs anywhere’ and ‘runs best somewhere’ are different things, and buyers will feel the gap in performance and cost. At that point, neutrality is a label,” says Mohan. 

He agrees that integrating software and compute will produce measurable cost, performance and reliability gains. Defining this as “the strongest part of the deal”, Mohan explains that with Nscale controlling both the silicon and Anyscale’s control plane, it can tune scheduling, memory, and networking together in ways the compute-neutral Anyscale never could.

Anyscale commercial support for Ray

Anyscale was founded by the creators of Ray, an open source project that provides a distributed computing framework designed to scale Python workloads across any infrastructure into live production application jobs and services. 

Ray was donated to the PyTorch Foundation in 2025. Anyscale continues to provide its commercially supported services for Ray, which include a “no DevOps” route to 100% managed cloud infrastructure and serverless autoscaling, making it simpler to create, deploy, and monitor machine learning workflows in production.

Anyscale supports data processing, model training, batch inference, and LLMs across public and private cloud environments. As open source as this all feels, are we still edging towards narrower proprietary channels, or the possible threat of deeper application and data service dependencies that developers will ultimately have to wrangle around?

“I don’t think so, primarily because the way that the platform works, it’s designed to orchestrate across various different clouds and different infrastructure. It’s like a heterogeneous distributed compute platform. So the platform’s always gonna remain multi-cloud,” confirms Nscale’s Bathurst.

Pricing permutations and hyperscalers hearsay

Pressed on any forthcoming pricing changes or likely reactions from the major cloud hyperscalers in relation to Nscale now being a credible alternative, Bathurst and team were (perhaps understandably one day after an acquisition deal announcement) politely tight-lipped.

More voluble is always-affable analyst Mohan, who says that, “Every optimization that only shows up on Nscale hardware is a dependency. So, an argument can be made either way. Standalone orchestration software and independent tooling vendors are getting absorbed into whoever owns the GPUs, because the economics only work when you control both. Expect more of it,” Mohan underlines.

He explains that Nscale “now becomes a real specialist cloud services provider alternative,” i.e., not a general-purpose one like AWS, Azure and Google Cloud with their plethora of managed services, from databases and data warehousing to container orchestration through to AI/ML pipeline technology.  However, he does see space for Nscale to become a strong player in raw training and inference at scale.

From cryptocurrency to cloud contender

London, UK-based Nscale was established in 2024 from what was originally a cryptocurrency mining business. 

As suggested, Anyscale will retain its brand name as part of the Nscale family, and the company has restated its stance that customers are “free to choose the cloud infrastructure on which they run their AI workloads” today.

The company’s initial press statement said that “over time” users will gain the additional option of running the Anyscale software layer on Nscale’s full-stack AI platform. 

The first full-stack AI hyperscaler?

“Companies are moving beyond simply using AI to actually building their own. Doing that well requires the software and the infrastructure it runs on to be designed together,” says Keerti Melkote, CEO of Anyscale in the press release announcing the acquisition.

Melkote has defined the combination of Anyscale’s platform — built on Ray — with Nscale’s datacenter, compute and AI cloud services as the “first full-stack AI hyperscaler,” i.e., one that runs any AI workload at greater scale, so more software engineering teams can build and own their AI applications and services.

With this acquisition and the fusion of Nscale with Anyscale’s software layer, the organization will aim to widen its customer base. Existing work sees the company working in verticals from healthcare to e-commerce to robotics. It says its full stack offering will help companies speed up image and document processing, fine-tune LLMs on their proprietary data, and deploy AI agents in-house using open-source models.

The transaction is subject to closing conditions and regulatory approvals and is expected to close in the second half of 2026. Financial terms of the transaction were not disclosed, although Reuters reports a source stating that the deal price is “about $1.65 billion”, according to a person familiar with the deal.

AWS, Google Cloud and Microsoft Azure representatives were all contacted and invited to comment on this story.

The post Nscale just bought Anyscale. Here’s why it matters for multi-cloud neutrality. appeared first on The New Stack.

Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA

31 July 2026 at 16:45
Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents using over 20 pre-built metrics, DeepMind-backed adaptive rubrics, or custom code-based and LLM-as-a-judge metrics stored in a centralized, versioned registry. The service integrates directly into existing workflows via the Agent Platform SDK, agents-cli, and ADK, offering built-in user and environment simulators to automate complex multi-turn testing and streamline CI pipelines.
❌