❌

Normal view

AI Automation’s PR Problem: Businesses Keep Selling What it Can Do, Not Why it Should Be Trusted

28 August 2026 at 19:11
By James Hilditch, co-founder and executive creative director, BearJam AI automation has a messaging problem, but it’s not about capability. It’s about trust. Most businesses selling AI-powered services talk about speed, scale and output. Far fewer talk about the human oversight behind those results, the limitations that still exist, or what any of it means […]

Alibaba just released Qwen3.8-Flash: β€œAn early preview of the architecture in Qwen4”

Alibaba this week unveiled Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model.Β 

Hot on the heels of Qwen 3.8 Max, which arrived at the start of the month, this 125-billion-parameter AI model is offered as a prelude to Qwen 4. It is positioned as both a performance and value-for-money play. As such, it is claimed to have β€œsuperior capabilities in coding and office tasks” and an optimal balance among capability, latency, and cost.

β€œIn this release, we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4,” confirmed Alibaba in its release blog.

How and why does Alibaba offer early access precursor models?

By releasing the architectural changes in Qwen3.8-Flash early, the organization hopes the community will examine the mechanics, constructs, and components within and start road-testing them before the full Qwen4 model family is built on top of them.

By stating that Qwen3.8-Flash paves the way for Qwen4, Alibaba is showcasing (and, importantly, openly sharing) design forms that it will carry into its subsequent models, so that developers can start building (or at least planning) their next codebases early.Β 

Specifically then, Alibaba has stated that Qwen3.8-Flash β€œplays the same role” that Qwen3-Next played for Qwen3.5 i.e. by which the company means that Qwen3.8-Flash introduced developers to the company’s hybrid Gated DeltaNet + Gated Attention design (computational components that enable AI models to balance long-context efficiency with contextual focus and attention during inference and training), which was then used across the Qwen3.5, Qwen3.6, Qwen3.7 and Qwen3.8 series.

β€œQwen3.8-Flash-Next upgrades the model systematically along four aspects – attention, residual, embedding and optimization –  improving model capability while further optimizing computational efficiency, model capacity and training stability,” stated Alibaba.

Benchmark scores against rival models

When benchmarked on agentic coding, long-horizon agent tasks and multimodal intelligence, Qwen3.8-Flash appears to perform respectably against models including DeepSeek-V4-Flash and Claude-Opus-4.6Β across SWE-bench Pro (agentic coding), CoWorkBench (long-horizon office work), Toolathlon Verified (real-world tool use), MathVision (visual maths problem-solving), AndroidWorld (agentic mobile use) and ERQA (embodied intelligence).

β€œQwen3.8-Flash-Next upgrades the model systematically along four aspects – attention, residual, embedding and optimization – improving model capability while further optimizing computational efficiency, model capacity and training stability.”

Tested on agentic coding using SWE-bench Pro, Qwen3.8-Flash-Next scores 62.5, compared to 61.7 on Qwen3.8-27B, 55 on Qwen3.7-Plus, 56.0 on DeepSeek-V4-Flash-0731, and 53.4 on Claude-Opus-4.6 (Max).

Importantly, Qwen3.8-Flash requires only what Alibaba details as β€œaround one-ninth of the training resources,” while delivering superior performance. The company says that this means Qwen3.8-Flash β€œsignificantly reduces” both training and inference costs compared with Qwen3.7-Plus, a model three times its size.Β 

What’s the difference between Qwen3.8-Flash-Next and Qwen3.8-Flash?

For the sake of nomenclature, Qwen3.8-Flash-Next is the open-weight research-frontier model available to developers on both the Hugging Face AI developer hub and Alibaba’s ModelScope community portal. Built on the same underlying architecture, Qwen3.8-Flash is the production version of the model, offered via the QwenCloud API with 1 million tokens by default and official built-in tools.Β 

Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It natively supports 262,144 tokens of context and can be extended to 1,000,000 tokens with YaRN.

What architectural updates have happened?

As noted, Qwen3.8-Flash introduces architectural extensions across attention mechanisms, residual connections, embeddings, and optimization. Its hybrid attention architecture combines Gated DeltaNet (GDN), which compresses historical information, with Qwen Sparse Attention (QSA), a novel design that uses a lightweight compressed indexer to select relevant context, substantially reducing attention costs for long sequences.

Alibaba has explained that the Gated Residual (GR) mechanism expands data pathways between layers while strengthening cross-layer information flow and training stability. The N-gram Embedding technique scales model capacity with minimal additional computation, while the Muon Optimizer enhances the efficiency of large-scale model training.

β€œI’d love it if local LLMs actually could replace remote ones today, and I have no reason to lie about my experience either…it’s just not there (yet) for local professional software development.”

What do developers think of Qwen3.8-Flash?

In terms of developer reaction, it’s been a mixed bag so far.Β 

Mechatronics engineer Alok posts on X, saying he thinks the Video RAM barrier (i.e., the need for physical, high-speed GPU-based memory for LLM memory management in the face of model quantization that aims to enable better long-context inference) is now officially dead.Β 

β€œI just ran Qwen3.8-Flash-Next (MoE) 125B A6B with aΒ  250,000 context window on a single 24GB RTX 4090 – 21 tokens/sec decode. 364 t/s prefill – no mtp. No dflash. No KV cache quantization! We are running datacenter models on consumer hardware,” enthused Alok.

Multi-disciplined software developer Embedding Shapes is less happy.

They post on Hacker News using some colorful language to describe how models keep [insert expletive]-ing up very basic things before saying, β€œI’d love it if local LLMs actually could replace remote ones today, and I have no reason to lie about my experience either. But I too got hopeful reading the sentiment on the Internet about Qwen 3.8, but it’s just not there (yet) for local professional software development.”

Pricing and access

Qwen3.8-Flash can be accessed via API Model Studio and Qwen Cloud, Alibaba’s AI-native cloud platform. Pricing per 1 million tokens is US$0.16 for input and US$0.47 (or 3 RMB for developers inside China) for output.

The model is also available on QwenWork, Alibaba’s workplace AI agent platform, where it runs a redesigned β€œstandard mode” that cuts token consumption per task by 75% and β€œroughly doubles generation speed” compared with the current mode.

Alibaba has claimed that this brings flagship-level capabilities within reach of everyday workloads and that developers can run this model on hardware that they likely already own.

When is Alibaba’s Qwen4 scheduled for launch?

Alibaba has not confirmed a firm launch date for Qwen4. Still, a casual web search on the topic yields a range of commentators and market watchers who broadly agree it will arrive before the end of the year, possibly as early as September.

Conjecture in this space explores whether the next model family could be optimized for complex 3D coding and design tasks in advanced spatial modeling. Others think that the Mixture-of-Experts (MoE) architecture will be extended and that deeper native multimodal processing capabilities might be featured.

Alibaba was contacted for broader comment on this story but declined to engage.

Qwen, in traditional Chinese, ι€šηΎ©εƒε• (pronounced TōngyΓ¬ qiān wΓ¨n), translates to β€œa thousand questions on general meaning” in English. Use that in your local pub quiz this weekend.

The post Alibaba just released Qwen3.8-Flash: β€œAn early preview of the architecture in Qwen4” appeared first on The New Stack.

Best Manufacturing Payroll Software for Mid-Sized Companies

28 August 2026 at 18:09
Manufacturing payroll involves serious challenges. Between piece-rate workers on the floor, salaried supervisors, seasonal contractors, and rotating shift schedules, the best manufacturing payroll software has to handle the demands that generic platforms weren’t built around. Keeping up with FLSA rules, state tax codes, and union agreements across multiple facilities adds another layer entirely. After reviewing […]

The three layers of agentic AI security: A defense-in-depth architecture for autonomous agents

28 August 2026 at 17:52

Presented by Nutanix


Autonomous systems that can reason, make their own decisions, and execute actions across an environment introduce a category of risk that application-level controls were never built to contain. Treating that risk as a single problem produces incomplete architectures, says Oscar Wahlberg, senior director of product management at Nutanix.

"The guardrails to catch a malicious prompt won't stop an agent from hallucinating and doing something it never should have done, like accidentally deleting databases or leaking sensitive data with a credential it was granted but then uses for something entirely different," Wahlberg says. "That's the central problem as enterprises move autonomous agents out of experimentation and into production."

Once an agentic system is granted execution privileges across the data center, the security posture has to scale into a defense-in-depth architecture spanning infrastructure, storage, compute, networking, and a governing control plane. Each layer addresses a distinct category of risk, rather than duplicating the same controls across the stack. No single security control or vendor can provide that protection on its own. Defense-in-depth depends on those layers working together.

By dividing the responsibilities across the stack and adhering to zero trust segmentation, organizations can create a secure framework that improves their overall posture. Understanding which risks belong in each layer is what turns the principle of defense-in-depth into a practical security framework, with three layers that each have a distinct responsibility.

Infrastructure layer: Establishing trust where AI agents run

The infrastructure layer’s foundational responsibility is establishing a root of trust that answers a simple question: who is operating in the environment? That trusted identity becomes the prerequisite for every security control above it. Before an organization can trust what an agent does, it first has to trust the integrity of the environment where the agent runs. When an agent requests permission to execute an operation, the system must be able to verify that the request came from the legitimate agent β€” not something impersonating it.

Delivering that kind of assurance depends on technologies that root trust in the hardware itself, including platform attestation, confidential computing, and secure boot, alongside controls that prevent unauthorized access both within a server and beyond it. For regulated industries such as financial services, this layer provides the ability to isolate AI production workloads so that neither the agent nor the environment can operate outside its assigned scope. That mitigates risks including model and runtime tampering, supply chain compromise, and unauthorized access to sensitive AI workloads.

Network layer: Governing how AI agents communicate

Once agents begin communicating with other agents, APIs, applications, and enterprise systems, they generate a level of concurrency and dynamic communication that traditional static network configurations were never designed to handle. An agent configured to call APIs, query data sources, and spin up additional agents without constraint creates a sprawling web of east-west traffic that becomes very difficult to reason about, and that complexity can easily mask lateral movement or data exfiltration when the right network security layers are not in place.

"We should treat AI agents as a new class of network identity, and make sure that an agent can only talk to other agents or data sources where it's explicitly allowed to do so," Wahlberg says. "That means moving away from rigid static rules toward dynamic policy enforcement."

Nutanix's solution is Agent Gateway, part of the Nutanix Agentic AI solution. It's a unified, governed layer that is designed to provide cost control and governance capabilities to help manage autonomous agent users. Coupled with agents grounded in zero trust segmentation and using capabilities like Nutanix Flow for micro segmentation and integrating with networking vendors, including its integration into the Cisco Secure AI Factory, Agent Gateway helps enterprises govern interactions across agents, models, data sources, and enterprise applications.

The network layer governs lateral movement, data exfiltration, and gates the agent's network interactions. A zero trust framework with access blocked by default and scalable interaction monitoring is important for agents since they can exhibit unreliable behavior. The Nutanix software integration with Cisco UCS servers and Cisco AI PODs delivers the turnkey physical infrastructure (compute, storage, and networking) that the AI factory runs on.

Control plane layer: Governing what AI agents are permitted to do

The control plane is the brains of the operation, providing a central point for managing agent permissions, tool access, resource consumption, and runtime visibility. What matters most is having a single place where policies can be enforced consistently rather than reinvented for every agent, Wahlberg says.

"Agent Gateway acts as a universal endpoint for different models and tools, so an IT team can configure their agents to talk to this single control point," he explains.

The centralized AI gateway enables the admin to observe, audit, and control access to models as well as MCP tools protecting data and gating privileged access. This layer is designed to help mitigate risks such as privilege misuse, runaway agents, unauthorized tool usage, data leakage, and the excessive model consumption that can lead to increased token consumption when agents get stuck in runtime loops. And it depends on treating governance as a runtime control system rather than a compliance afterthought.

Why one-size-fits-all security fails agentic AI environments

The biggest architectural mistake enterprises make is assuming a single security model can be stretched across every layer of an AI stack. When an organization tries to solve for hardware-level trust with application-level software, or leans on static legacy network rules to manage dynamic agents, it builds an architecture that either blocks the agentic system from doing its job or leaves critical doors wide open. One-size-fits-all thinking tends to produce significant performance penalties and operational friction.

"By failing to assign specific responsibilities to the appropriate layers, enterprises end up with blind spots in governance," Wahlberg says. "They might secure the model output but miss that there's data leakage between agents, or they might secure the network but lack the control plane visibility to understand that they're wildly burning tokens because the agents are stuck in some kind of runtime loop."

Focusing exclusively on the model leaves the largest gaps of all, because a guardrail that catches a malicious prompt does nothing to stop a hallucinating agent from misusing a legitimate credential. Embedding security across the full stack helps ensure that even when a model level threat slips past the initial filters, the agent remains constrained by hardware rooted trust, network isolation, and access controls at the agent layer.

How Intel, Cisco, and Nutanix build defense-in-depth together

The three-way partnership from the three companies demonstrates how the layered architecture comes together in practice as a well-governed, enterprise-grade AI Cloud. Intel supplies the computer to run agentic workloads and secures the execution environment through hardware-rooted trust and confidential computing, while also driving costs down through their accelerators. Intel Xeon 6 processors with built-in AMX accelerate AI inference efficiently without relying exclusively on expensive GPUs.

Cisco wraps the environment in a secure fabric that governs communication between agents and enterprise tools, while Nutanix provides the software platform, minimizing architectural silos, and the central control plane that enforces permissions, delivers visibility and cost governance, and ties the architecture together into a defense-in-depth solution that lets enterprises scale agentic AI.

Of the three layers, enterprises currently underestimate the control plane the most, Wahlberg says. A true control plane extends far beyond initial deployment to simplify Day 2 operations, he explains, giving IT teams the continuous observability, and strict token governance required to keep autonomous agents secure and cost-effective in production.

"Apart from model and tool selection, governing the agent deployments and their access to models and business tools in a tightly integrated full stack platform will be important for the success of AI projects," he says, pointing to a near future in which organizations move from a handful of AI use cases to thousands of agents working autonomously to drive the business.

Technology leaders should prioritize building a centralized governance layer today that can manage agent identities, tool permissions, and token budgets in real time, because that control point is what builds the operational muscle to scale safely.

"You can't build an AI system without getting into a lot of complex decisions," he explains. "And you need a control plane that talks across multiple vendors and infrastructures to help you solve for those defense-in-depth strategies."

Learn more about the Nutanix Agentic AI solution here.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Video Friday: Meet Microduck

28 August 2026 at 16:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

Humanoids Summit Seoul: 22–23 September 2026, SEOUL
IROS 2026: 27 September–1 October 2026, PITTSBURGH
CoRL 2026: 9–12 November 2026, AUSTIN, TEXAS

Enjoy today’s videos!

Nvidia just paid US $12.9 billion for the company that acquired Pollen Robotics, and this must be why.

Meet Microduck. πŸ¦† The 25-centimeter, 780-gram robot that waddles, falls, gets back up, and learns new tricks.

Packed inside: 15 degrees of freedom, a front camera, an 8x8 lidar, two IMUs, mics, a speaker, NFC, Wi-Fi, and Bluetooth.

Out of the box, Microduck already walks, sits, crouches, roller skates, picks up objects with its articulated beak, and recovers from falls on its own. Drive it with a game controller, plug-in accessories, and NFC tagged objects, run autonomous behaviors, or gather several Microducks for races and football.

Software fully open source. Ready for whatever you throw at it.

On pre-order for an astonishingly low $399, and ships before Christmas.

[ Microduck ]

Thanks, Matthieu!

If you’ve chosen to ignore all the earlier DARPA Lift Challenge videos that we’ve posted, now you can get all caught up in about five minutes.

[ DARPA ]

You had me at β€œ54-gram robot that out-jumps a kangaroo.”

[ IEEE Transactions on Robotics ]

Sometimes, you just need a video like this.

Most fish-inspired robots are built for one size and one job, so scaling them up or down usually means starting from scratch. A team of engineers says it has found a way to solve that problem. They’ve unveiled ScaFi, a robot modeled on fish like cod and mackerel.

[ New York University ]

Thanks, Leah!

Martin writes, β€œWe’re a small robotics team in Czechia, Europe, building practical hardware around the Unitree G1. Here’s a short demo of our lightweight gripper picking up a strawberry; the gripper weighs under 200 grams and is designed for simple, sensitive manipulation without adding a complex multifinger hand.

[ Sentio Robotix ]

Thanks, Martin!

Hybrid visual markers that are useful for both cameras and lidar is a neat idea.

[ Hello Robot ]

Thanks, Binit!

EmoLo brings emotion-inspired expressive locomotion to Open Duck Mini V2, a low-cost, open-source bipedal robot inspired by Disney’s BDX droids. With a single reinforcement learning policy, the robot can generate distinct walking styles associated with different emotional expressions, showing how characterful and expressive whole-body motion can be achieved on an accessible robotic platform.

[ EmoLo ]

Thanks, Masato!

If it’s possible for a robot with a completely immobile face to look frustrated, this robot absolutely does, starting at three minutes into this video.

[ DLR RM ]

Noble Machines deployed its first general-purpose robots to a Fortune Global 500 industrial customer within 18 months of the company’s launch and met its first delivery milestone, made possible by its AI-driven whole-body control and industry-leading end-to-end autonomy.

[ Noble Machines ]

We’ve reduced the time it takes to go from physical prompt β†’ robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work.

[ Generalist ]

I know this video is mostly a gimmick, but I would totally rent a moderately heavy lift quadruped for a couple of days to help with a move.

[ DEEP Robotics ]

Is taking two minutes to excellently fold a shirt too long, or do we even care how long it takes, as long as it’s a robot doing it?

[ Tokyo Robotics ]

TRON 2 Γ— Wuji Hand 2 handles TCM pharmacy work: picking, weighing, grinding, and packaging. The omnidirectional base frees the hands, while precise gripping and dual-arm force control enable midair operations.

[ LimX Dynamics ]

Person who genuinely knows things about robots, Christian Hubicki, explains everything about robots smashing into walls.

[ Christian Hubicki ]

Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect

28 August 2026 at 17:06
Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing,...

Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing, post-processing, and runtime code. NVIDIA TensorRT Model Connect open collection of reference implementations helps to address this challenge. TensorRT Model Connect shows you how to run supported models with NVIDIA TensorRT in native C++…

Source

Are We on the Verge of an Intelligence Explosion? Maybe Not.

28 August 2026 at 16:45

Recursive self-improvement, where AI continuously builds better versions of itself, might be harder than some hope.

There’s growing excitement in the AI industry about the idea that today’s leading models could build the next generation of the technology. But a new study recently found top AI agents struggle on the kind of genuinely open-ended research problems required to push the field forward.

Large language models have made rapid progress in many of the day-to-day jobs involved in machine learning research, such as writing code, generating and curating data, and running experiments. Last year, startup Sakana AI’s AI Scientist-v2 even managed to write a paper that cleared peer review for the prestigious International Conference on Learning Representations.

These advances have led to speculation that models are close to being able to build better versions of themselves with little human oversightβ€”a process called recursive self-improvement. The idea underpins predictions that we may be on the verge of an intelligence explosion that could quickly lead to AI superintelligence.

In a recent paper, researchers put the idea to the test using a new approach they call shadow evaluations. This involves taking the research question from a high-quality, unpublished machine learning paper and asking AI agents to solve the problem. The original paper’s authors then grade the results. When the team tested Claude Opus 4.8 on two papers submitted to the prestigious machine-learning conference NeurIPS 2026, the authors rejected both.

β€œThe papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” Sayash Kapoor from Princton University, who co-led the study, told MIT Technology Review.

Previous efforts to get AI agents to do machine learning research have often targeted problems focused on engineering, such as reproducing previous research or training smaller models against a benchmark.

In the new experiments, the researchers challenged models with more open-ended tasks that required them to devise hypotheses, decide what evidence is needed to validate them, judge when a research direction was fruitless, and go back to the drawing board.

One research question was whether the personality traits a language model displays can be measured and adjusted by observing and editing its weights; the other attempted to detect when a model that works with tabular data has quietly stopped being reliable.

In each case, the AI researchers were given $3,000 of API credits, a budget for time on GPUs to run machine learning experiments, a dedicated Linux virtual machine, and unrestricted internet access. They were then given six days to produce a paper that could pass NeurIPS’ stringent peer-review criteria.

In both cases, the models got a good start. The agents surveyed the literature effectively, came up with opening hypotheses that mirrored those of the authors, and successfully ran hundreds of experiments.

But they quickly went off the rails. Although they could monitor their own use of time and their API and GPU budgets, they rushed through the process. One left 110 hours of unused time on the clock, and both failed to spend even 50 percent of their API budget.

Both agents also settled on a research direction within just 10 hours and failed to change approaches despite repeated negative feedback from another AI designed to review drafts of their papers. The reviewer identified problems the human authors would also flag in the final paper, but the models simply added caveats to their findings and ploughed on. Ultimately the papers received a β€œstrong reject” and a β€œreject” decision from the human reviewers based on NeurIPS grading protocol.

The authors admit their approach has limitations. The reviewers knew AI had written the submissions, and some of the team are on record as doubting an imminent intelligence explosion. The original human-authored papers also took far longer than six days to produce and used many more GPU hours to reach their conclusions (though, as the researchers note, the models did not use their allocated budget in any case).

Nonetheless, the results suggest that today’s models still have some way to go before they can tackle the most challenging problems in machine learning research. Until that happens, the dream of recursive self-improvement is likely to remain a distant prospect.

The post Are We on the Verge of an Intelligence Explosion? Maybe Not. appeared first on SingularityHub.

Meta researchers taught an 8B AI model to match Claude Opus 4.5 β€” without the frontier price tag

28 August 2026 at 16:33

Consider an AI agent tasked with a complex enterprise workflow like migrating massive batches of customer records from a legacy CRM to a cloud database. The agent cannot rely solely on its internal context window for a job spanning hours and depends on the runtime layer, aka the harness.

This harness provides execution feedback, like server logs, to help the agent maintain an accurate understanding of dynamic API connections. It also provides state trackers and control-flow mechanisms to manage completed and pending subgoals, ensuring the agent doesn't skip or duplicate data batches. When unexpected errors occur, such as a database rejecting a batch due to strict API rate limits, the harness provides tools and instructions to help the agent recover.

The main way to tell an agent how and when to use its tools is to have a human developer write a set of rules and instructions telling it what to do step-by-step. For example, a developer might instruct the agent to always search the company wiki before writing an email. Because the agent is just following a rigid script, it lacks true autonomy. It hasn't been trained to independently weigh the costs and benefits of its actions.

To solve this, researchers at Meta AI and University of Illinois Urbana–Champaign introduce EvoHarness-RL, a framework that adds a layer of abstraction to the agent's harness and teaches the underlying model when to read, update, or consolidate the information it obtains from its environment.

In long-horizon tasks, how AI agents read and process the information they obtain from their environment is pivotal to their success. The agent must update its understanding of its environment, track completed and pending subgoals, recover from failed actions, and reuse procedures from previous experience. This execution depends on the harness.

A series of self-evolving agentic frameworks like Harness-1 solve part of the problem by accumulating past trajectories and distilling them into structured procedural memory, like reusable skills, workflows, or code libraries for future tasks. However, they generally separate this long-term skill curation from real-time, within-episode state tracking. They aren't actively training the agent on how to manage its immediate environmental reality or track its active task steps while working.

Xuying Ning, co-author of the EvoHarness-RL paper, told VentureBeat that manual logic and rigid memory structures are primary culprits draining engineering resources.

"The optimal harness often changes with the model," Ning explained. "Different models may need different prompts, memory designs, permissions, or sandbox configurations. If all of this logic is manually coded, every model upgrade can lead to another long cycle of tuning and debugging."

Furthermore, existing memory systems that simply accumulate experience can actively degrade an agent's reasoning. "Append-only memory assumes that more context is always helpful, which is not necessarily true," Ning said. "Over a long task, the memory may contain outdated conclusions, failed attempts, or information that is no longer relevant." As a result, long-horizon agents need a dynamic memory capable of updating, compressing, and replacing information to avoid repeating past mistakes.

EvoHarness-RL: A unified belief, progress, and experience workspace

To overcome the limitations of rigid, manual prompts, the researchers introduce EvoHarness-RL, a training technique that teaches the agent to make optimal use of its harness. Instead of blindly following hardcoded instructions, the agent learns how to construct a structured workspace from messy execution data and decide when and how to consult that external state during complex workflows.

To simplify the management of different components of the harness, EvoHarness-RL consolidates the agent’s support systems into a single, unified interface. This interface, known as the Belief, Progress, and Experience (BPE), categorizes the agent's external needs into three functional areas:

  • Belief: Maintain an accurate read on the current environment.

  • Progress: Manage completed and pending subgoals.

  • Experience: Reuse historical knowledge across tasks.

Instead of using complex, domain-specific APIs, the AI interacts with this clean dashboard using four compact meta-actions: track, commit, recall, and note. It issues commands to track the live environment, commit to workflow updates, recall past strategies before acting, and write notes to save newly discovered insights for future runs.

These states map directly to high-value enterprise verticals. "In software engineering, Belief can represent the agent’s current understanding of the repository," Ning said, detailing how the agent monitors component interactions and workspace changes. "Progress tracks what has already been completed, what still needs to be done, and which steps depend on others." Meanwhile, Experience captures lessons, like user feedback on a mistake, to guide future actions.

The same idea applies to finance, Ning said. During a compliance audit, Belief might describe the applicable rules and available evidence. Progress tracks which checks have been completed and which exceptions remain open. Experience helps the agent recognize recurring discrepancies or know when an issue should be escalated.

"Together, these states help prevent the agent from losing track of its work or repeating the same failed approach," Ning said.

To teach the agent both the mechanics and the strategy of managing its external workspace, the researchers designed a two-stage training recipe. In the first stage, supervised harness fine-tuning, the base model learns how to extract and structure useful facts from messy interaction logs into the BPE framework.

However, querying memory or updating trackers consumes time and compute tokens, meaning the agent cannot afford to blindly check its tools at every step. To solve this, the second stage uses β€œcost-aware” reinforcement learning to teach the agent efficiency. This phase trains the agent to calculate when accessing its external state is worth the budget cost. This two-step process transforms tool-use from a rigid, hardcoded prompt into a learned runtime behavior.

EvoHarness-RL in action

To validate EvoHarness-RL, the researchers evaluated the system using the ALFWorld benchmark, a text-based environment featuring multi-step tasks that test sequential logic and state tracking.

They used Qwen3-8B as the base model to train. The team pitted the trained 8B model against three large frontier models (Claude Opus 4.5, GPT-4.1, and GPT-5), frozen agent frameworks with static tools (such as ReAct, ExpeL, and ReasoningBank), and advanced trainable methods (e.g., standard GRPO, SkillOS, and SkillRL).

The results show a significant jump in performance for smaller, cost-effective models. With EvoHarness-RL, the Qwen3-8B model achieved a 96.9% average success rate, a 49.0 percentage point improvement over its baseline ReAct counterpart.

Furthermore, the trained model outperformed advanced trainable frameworks like SkillRL (89.9%) and SkillOS (80.2%). Most impressively for enterprise developers looking to optimize compute costs, the 8B model effectively matched the performance ceiling of expensive closed models like Claude Opus 4.5, which scored 96.4% out-of-the-box.

Beyond empowering smaller models, the experiments show that the BPE framework has universal benefits across all model scales, even without the extensive reinforcement learning phase. When researchers equipped frozen, out-of-the-box frontier models with the BPE prompt-time harness, their execution improved significantly. GPT-4.1's success rate improved by 22.1 points and GPT-5 by 25.7 points.

Aside from the results, the researchers recorded effects during the experiments that demonstrate the dynamic behavior the LLMs acquire as they go through the EvoHarness-RL training. During the reinforcement learning phase, they observed a behavioral shift as the agent internalized knowledge over time, which they called "harness annealing".Β 

Early in training, the AI relied heavily on querying its Experience and Progress trackers for almost every step. However, as it mastered routine actions, it actively reduced its reliance on external tools, embedding the successful patterns directly into its parameters. In a real-world enterprise setting, this translates directly to lower latency and reduced compute costs. By annealing its tool usage, the AI stops wasting tokens and time querying databases for standard workflows it has already mastered.

Simultaneously, the agent demonstrated "harness evolution," where it dynamically adapted its strategy based on the complexity of the situation at hand. While it bypassed its tools for simple, familiar tasks, it actively chose to scale up its use of the Belief and Experience modules the moment it encountered novel environments or unexpected roadblocks. For example, if an AI agent is migrating standard database records, it moves fast. When it encounters a strange legacy API endpoint or a complex validation error, it slows down, pulls up the live server logs, and queries its historical tickets to safely resolve the edge case rather than hallucinating a guess.

Bringing EvoHarness-RL into existing systems

Despite these massive gains, adopting a new framework often introduces friction for enterprise engineering teams. However, EvoHarness-RL utilizes an environment adapter that allows internal implementations to remain domain-specific to an organization's existing tools while sharing the trainable layer.

"I think there is significant potential to integrate BPE into existing orchestration systems," Ning said. "It does not necessarily require teams to replace their current tools or agent frameworks. BPE can work as an additional state-management layer that continuously organizes what the agent currently believes, how far it has progressed, and what it has learned."

For enterprise builders worried about inference costs, the framework addresses the hidden engineering cost of consolidation. Because consolidation requires strong reasoning, teams can adopt a hybrid, asynchronous architecture to optimize budgets.

"One possible compromise is to use a frontier model to generate high-quality consolidation data, then fine-tune a capable open-weight model to handle routine state management," Ning said. Furthermore, "because consolidation can happen asynchronously, it does not always need to slow down the agent’s main execution loop."

Teams must also carefully evaluate when a trainable BPE harness is necessary versus when it is overkill.

"For a short and stable task, ReAct or standard RAG may already be sufficient," Ning said. "BPE becomes much more valuable when an agent works for many hours, days, or even weeks." In those complex scenarios, an agent needs a compressed understanding of its decisions to avoid getting lost, relying on Experience to iteratively improve from previous failures and human feedback.

Ultimately, this approach signals a shift for AI orchestration engineers. "It is not a complete replacement of workflow engineering," Ning said, "but a transition from directly scripting agent behavior to creating systems in which better behavior can be learned."

Cohere Parse 5 loses the benchmark on points. It wins on cost per page.

28 August 2026 at 15:30

Enterprises trying to feed PDFs, slides and scanned documents into AI pipelines keep running into the same wall: the tools either miss the structure β€” tables, charts, layout β€” or cost too much to run at scale.

Cohere released Parse 5 on Thursday, positioning it on price-to-performance, not raw accuracy β€” the right cost-capability mix for enterprise scale. Parse 5 is a 2.3-billion-parameter vision language model built to convert PDFs, slides and images into structured Markdown at enterprise scale.Β 

Cohere's own published benchmark comparison puts Parse 5 behind three larger, general-purpose frontier models on accuracy. GPT-5.5, Opus 4.8 and Gemini 3.5 Flash all score higher than Parse on the three ParseBench dimensions Cohere reports. Cohere is not claiming the top score. It is claiming the best price for a score close to the top.The company priced the model at $1.50 per 1,000 pages through its API, with Model Vault, Cohere's secure, single-tenant platform for managed inference, available for higher-volume deployment.

"Document parsing isn't solved because the hard part isn't reading text, it's preserving structure and meaning," Nils Reimers, VP of AI Search at Cohere, told VentureBeat. "Enterprise documents mix tables, diagrams, charts, and formatting that change the interpretation of the data. Most tools still drop structure or hallucinate content, and even frontier models break on layout‑heavy pages."

Inside the single-pass architecture

Parse 5 takes a page as an image, runs it through a single vision-language model pass and returns structured Markdown, collapsing the OCR-plus-model pipeline most tools run as separate steps.

Architecture. It is a 2.3-billion-parameter vision language model built on Cohere Labs' North-Micro-Vision-Instruct architecture, with an 8,192-token context window and roughly a 4.6-gigabyte footprint. It accepts a PDF, PowerPoint or JPEG page as a base64-encoded image and returns Markdown in reading order, with tables rendered as HTML, image descriptions and bounding box coordinates for tables and images.

Language coverage. Arabic, English, French, German, Italian, Japanese, Korean, Portuguese and Spanish get stable accuracy, with lower-accuracy zero-shot support elsewhere.

Output modes. The default output returns a Markdown string per page. A blocks mode returns typed elements, where each table carries its own HTML, bounding box and description, the format Cohere positions as what makes citation-level traceability possible for agents.

Availability. Parse 5 is generally available now through the Cohere API, Model Vault, Microsoft Foundry and AWS SageMaker.

The benchmark shows a trade-off, not a win

ParseBench is a benchmark that scores document-parsing tools against human-verified enterprise pages. Cohere reports Parse 5 scoring 79.2 across three dimensions: tables, content faithfulness and semantic formatting. That puts Parse 5 behind GPT-5.5 (84.4), Opus 4.8 (84.3) and Gemini 3.5 Flash (81.8), and ahead of LlamaParse's Cost Effective tier (78.3), Mistral OCR 4 (74.5), Databricks AI Parse (72.4) and Azure Document Intelligence (69.3).

Cohere's table notes two excluded dimensions, Layout and Chart, and attributes both to product scope rather than a performance gap. Parse 5 returns reading-order Markdown instead of per-element bounding boxes for text, and describes charts rather than extracting their underlying data, with chart-data extraction planned for a future version.

Reimers said that design choice reflects where agentic workflows actually break.

"For charts, for example, we provide a general description of the chart together with an indicator, how Agentic AI can visually inspect the chart," Reimers explained. "Other solutions try to extract the data from the chart, but then miss out critical information (for example, the color or the pattern of a line) that leads to hallucinations in Chat and Agentic AI applications."

Cost is where Cohere makes its real case. Reimers pointed to a workflow the company modeled for a large financial services firm.

"We ran the numbers for a large financial services workflow that processes 750 million documents a year and showed that choosing Parse 5 over a large general‑purpose model like GPT‑5.5 would reduce costs by more than 98 percent."

That figure is Cohere's own estimate for a single modeled workflow, not an audited deployment.

Where Parse 5 sits against the field

There is no shortage of options for enterprises looking at parsing solutions.

General-purpose frontier models, GPT-5.5, Opus 4.8 and Gemini 3.5 Flash, top the accuracy comparison but carry the cost and latency of a large model on every page.Β 

Then there are specialized parsers, including Mistral OCR 4, LlamaParse and open-weight options like Chandra OCR 2 and RedNote's dots.mocr.

Hyperscaler document intelligence services, AWS Textract, Google Document AI, Azure Document Intelligence and Databricks AI Parse, compete more on ecosystem convenience than on raw parsing quality, and score lowest in Cohere's own comparison.

Kevin Petrie, VP of Research at BARC US, said document parsing sits at the center of enterprise AI adoption right now.

"We're completing a survey now that shows document analysis is by far the #1 use case for AI, with 62% adoption rates among organizations we polled," Petrie told VentureBeat. "Documents and other unstructured objects, including images and so on, hold the proprietary context that organizations need to differentiate their agentic AI initiatives."

Petrie added that only time will tell how Cohere's cost-performance stacks up against frontier models, but strategically his view is that Cohere has the right focus.Β 

Stephanie Walter, Practice Leader for AI Stack at HyperFRAME Research, sees Cohere Parse 5 as sitting in a good spot between legacy OCR and using an expensive frontier model on every page.Β 

"Its potential advantage is delivering structure, spatial provenance and private deployment at a price suitable for high-volume ingestion," Walter told VentureBeat. "It does not need to win every benchmark. It needs to make reliable enterprise-scale parsing economical."

The real test is downstream, not on the benchmark

"Parsing is the first quality gate in the enterprise AI stack," Walter said. "If tables, headings, images, or reading order are lost during ingestion, better embeddings and larger models cannot recover that missing structure."

A benchmark score isn't the only input that matters here. "Enterprises should test parsers against their own most difficult documents and measure downstream retrieval and task accuracy, not how clean the extracted text looks," Walter said. "The right question is not 'Did it read the PDF?' but 'Can the agent now use the information correctly?'"

From Manual Loading Toward Automated Machining: A Practical CNC Production Upgrade Roadmap

28 August 2026 at 11:30
Manufacturers are under growing pressure to produce more parts with shorter lead times, consistent quality, and fewer interruptions. For many CNC shops, this naturally leads to one question: when should production move from manual loading toward automation? The answer is rarely as simple as replacing an operator with a robot. The productivity of CNC machining […]

Authorities arrest 2 alleged members of prolific hacking group TeamPCP

28 August 2026 at 11:15

Authorities in Australia said Wednesday that they arrested two men accused of participating in cybercrimes for TeamPCP, a prolific group of hackers that, over nine months, has carried out a relentless series of supply-chain attacks that infected more than 1,000 organizations worldwide.

In a statement, the Australian Federal Police said the two men were arrested and charged with 14 offenses. The statement said the men were members of TeamPCP, which by the authorities’ count, compromised more than 1,000 organizations worldwide. The statement didn’t identify the men, except to say they lived in the Western Australian towns of Cottesloe and Mandurah. KrebsOnSecurity, citing a lengthy investigation, provided what it reports to be both defendants' names, along with an extensive background of their lives and the mistakes that led to their downfall.

The hacks that keep on hacking

TeamPCP has vexed law enforcement officials and security personnel around the world since it emerged in December. The group is best known for a sustained series of supply-chain attacks that laced open source software with malware that self-propagated from one package to another. The viral infections worked by targeting organizations’ CI/CD pipelines, which are used to rapidly develop, update, and deploy software.

Read full article

Comments

Β© Australian Federal Police

❌