Normal view

Building a Foundation Stack for General-Purpose Robots

13 July 2026 at 10:19


This article is brought to you by X Square Robot.

Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the field does not yet agree on what it is.

X Square Robot, a Chinese embodied-AI company, has made an unusually explicit bet. It argues that the recipe is an integrated stack, spanning the data a robot learns from, a world model for predicting changes in the physical world, and an action model that brings together perception, planning, reasoning, and decision-making to generate executable robot behavior. The company also believes that the stack should be built and released in the open.

X Square Robot shares its vision of bringing robots into real homes.X Square Robot

X Square Robot’s embodied AI stack

What holds the stack together is a small set of principles rather than a single overarching model.

  • The first is that the basic unit of robot data is an interaction, not a trajectory; a demonstration is successful only if it changes the world as intended, not simply because the joints moved.
  • The second is that pretraining should yield usable capability, not just an initialization for later fine-tuning.
  • The third is that behavior should be modeled around physical events rather than fixed slices of time.

These principles make the layers interdependent, since the same robot-free data that trains the action model is also structured to feed the world model. It is worth being precise, though. The company describes the world model and the action model as complementary but independent model families that share a code base. Both sit within its broader World Unified Model, which it has presented as an architecture for training vision, language, action, and physical prediction together.

Robot learning data: Engineering for quality and cost, not scale

For the X Square Robot team, one of the biggest constraints on general-purpose robots is the cost and quality of interaction data, not the number of parameters. To address that, the company built its Universal Manipulation Interface (UMI) data collection system, QUANXTA Zero Series. It works by collecting demonstrations from people wearing a rig with dual grippers rather than teleoperating a robot. This approach is not itself new, and builds on established methods for robot-free data capture. What sets it apart are two engineering choices.

Person using VR headset and handheld controllers to teleoperate a dishwashing robot system X Square Robot emphasizes data quality control, recording trajectories and replaying them on a real robot, with only those that actually complete the task counted as valid.X Square Robot

The first is quality control, and it is the most distinctive part. Rather than accepting recorded trajectories as they are, the system runs a closed inspection loop, and its notable step is physical playback. A sample of trajectories is replayed on the real robot, and only those that actually complete the task count as valid. That makes the validity rate a measured quantity rather than an assumption. For example, a gripper that closes a fraction of a second too early still looks like a grasp in the data, yet it has pushed the object away, so it shouldn’t be classified as valid. A smaller clean dataset can be worth more than a larger noisy one.

The second choice is how lower-cost human data and scarce robot data are combined. The company pretrains on a large volume of robot-free demonstrations to build general representations, then adds a small amount of real-robot data as an anchor to the specific machine’s dynamics. It reports that this reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection, driven mainly by how much cheaper the wearable rig is than a teleoperation setup.

The resulting dataset is deliberately model-agnostic, formatted to feed both action models and world models. The caveat is that the strongest results are measured on the company’s own robots and data-collection pipelines. Broader independent testing will help confirm and extend these promising results across a wider range of settings.

A world model organized around events

In developing its world model, called WALL-WM, X Square Robot took a differentiated approach. Most action models predict a fixed-length chunk of motion from the current image and instruction. That is convenient, but it segments behavior into fixed-duration windows, so the boundaries fall where elapsed time dictates rather than where one action ends and the next begins. WALL-WM instead treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.

Collage of robot arms manipulating kitchen objects with charts of multimodal AI performance X Square Robot’s world model, called WALL-WM, treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.X Square Robot

WALL-WM’s design reflects a specific concern about not discarding what large video models already know. To achieve that, a text-to-video model is coupled to a freshly initialized action network that reads from the video features without overwriting them, which preserves the visual prior. From that one process, it offers two modes. An event mode runs in variable-length segments and suits reasoning over long horizons, while a fixed-length mode produces the steady, real-time output a controller needs. That places WALL-WM between mainstream chunk-based action models and pure video world models, keeping the predictive character of a world model while still yielding executable control.

In a series of experiments, the company relied on a generalization test that is more specific than most. A model trained on a limited dataset was evaluated on long-horizon tasks in unseen settings and, on the company’s real-robot benchmark, reportedly outscored baselines that had been fine-tuned on related data. That is a meaningful result if it holds. For now, it is measured on the company’s own benchmark. With the code now being released, the broader community will have the opportunity to test, reproduce, and build on them across more settings.

A policy that runs before fine-tuning, and action tokens with meaning

The action layer carries two connected ideas. The first is a requirement the company sets for itself with Wall-OSS-0.5, its vision-language-action model: The pretrained model should run on a real robot before any task-specific fine-tuning.

The interest is less in the scores than in the design behind them. The model trains three objectives together, namely discrete action tokens, language grounding, and continuous action generation. And it keeps gradients flowing through all of them rather than freezing parts of the network as some rival designs do. It’s also a more strict method, since it reports untuned behavior such as approaching, grasping, and recovering, including on a deformable task held out of training.

Dashboard of robot training metrics with charts and photos of a robot sorting objects As part of X Square Robot’s Wall-OSS-0.5 vision-language-action model design, the pretrained model should run on a real robot before any task-specific fine-tuning. X Square Robot

The second idea is the action interface itself, called X-Tokenizer. Most systems that turn continuous motion into discrete tokens produce codes that the language model cannot interpret. X-Tokenizer reframes tokenization as learning a semantic interface, so that the top-level code stands for the intent of a motion while lower-level codes carry finer detail, all aligned with the language model’s own features.

A useful consequence is stability. Adding noise to an action barely moves the intent code, which is what lets one tokenizer to be reused across robots without re-tuning. The tokenizer inside the production action model is a related variant of this approach. Together, the two ideas give the action layer something rather powerful: capability that transfers.

The future of embodied AI stacks

X Square Robot is betting that its unique approach combining three layers, each specialized in solving a key part of the problem, will stand out from other embodied AI stacks. The physical-playback step that grounds data quality is uncommon and sensible. The reframing of world modeling around events, with one backbone serving both reasoning and control, is a genuinely distinct approach. And the pairing of a deployable pretraining standard with a tokenizer designed as a semantic interface gives the action layer unusual coherence.

X Square Robot’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

The next phase will bring broader validation. Much of the current evidence comes from X Square’s own robots and benchmarks. With the world model code now being made public, and as the community begins to test, reproduce, and build on the work, the reported capabilities will be tested across more robots, tasks, and settings.

X Square Robot’s recent funding rounds reflect similar confidence. The company’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

What’s next for X Square Robot

To learn more about its future plans, the following Q&A with the X Square Robot team further explores the company’s technology, strategy, and vision.

What made now the right moment, technically, to commit to this stack? What recently became possible that wasn’t possible a couple of years ago?

It is not one breakthrough but several trends maturing together. Foundation models gave us a shared representation across vision, language, and action, so we can model what a robot sees, what it is asked to do, and how its actions change the world in one framework, rather than as separate perception, planning, and control modules.

Compute and infrastructure are finally sufficient for large-scale pretraining over long-horizon, multi-embodiment data. Just as importantly, we realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical. The useful question is no longer how to predict a few seconds of video, but how to understand the ways actions change objects, contacts, and task states. Two years ago these ingredients existed separately. Today they are mature enough to work as one system.

“We realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical.”

Your data system captures demonstrations with a wearable VR rig and custom grippers rather than teleoperating robots. What was wrong with standard teleoperation?

Teleoperation is built around controlling the robot. It forces the operator to work within the machine’s kinematics, latency, and viewpoint, and the resulting demonstrations are slower, stiffer, and less diverse. We built our system around capturing human skill instead. Manipulation is really about contact, timing, finger coordination, and recovery, not just the path the hand takes, and a wearable rig records those before the behavior is compressed onto one particular robot. It also breaks teleoperation’s expensive scaling law, in which every demonstration needs a robot.

People can generate rich data independently of any robot, and the crucial property is that those demonstrations can still be replayed and executed on a physical robot through the model. Mobility is convenient, but that replay is the real point, because it is what lets the same data be reused across different platforms.

Robot and person loading a washing machine together in a modern laundry room. In X Square Robot’s approach, demonstrations can be replayed and executed on a physical robot through the AI model, allowing the same data to be reused across different platforms.X Square Robot

X Square Robot reports that its pipeline has roughly an 85 percent data-validity rate. Why is quality control such an underrated bottleneck?

Because errors in robot data are far more expensive than in language data. A small timing or contact error can change what a demonstration means. If a gripper closes a fraction of a second too early, the motion still looks like a grasp, but physically it has pushed the object away. A dataset that mixes failures and accidental successes teaches ambiguity, not skill, because the real unit is the interaction, not the trajectory.

So we run automated inspection, kinematic checks, and physical replay, where we play a sample of trajectories back on the real robot and count only the ones that actually complete the task. Data quality sets the ceiling on how good a policy can be. In our experience a smaller, cleaner dataset often beats a much larger, noisier one, which is why we treat quality control as part of the model, not a preprocessing afterthought.

The model runs in both “event mode” and “chunk mode.” When does each matter?

Both matter, for different reasons. The physical world changes through events—when contact occurs, a grasp forms, or an object slips—not in fixed-frame windows. Event mode concentrates the model’s attention on those moments, and it matters most for long-horizon tasks, like clearing a table, where progress is a sequence of semantic events rather than a smooth stream. It runs in variable-length segments that follow the task rather than a clock. Chunk mode matters for deployment. Real controllers need a stable, real-time interface, and fixed-length chunks integrate cleanly with existing control systems.

We organize learning around events in the first place because a fixed window can split one motion in half or merge two together, which turns training into short-horizon pattern matching and weakens the model on long tasks. So the world model’s job is to connect event-level understanding, which is where the reasoning happens, with a fixed-length output a real robot can actually run.

Why make “deployable before fine-tuning” the criterion?

Pretraining should produce capability, not just a good starting point. If a model is only useful after heavy fine-tuning, then most of the intelligence still lives in the downstream supervision, not in the foundation model. Deployable before fine-tuning is a more honest test of what pretraining actually learned. A well-pretrained robot should already know how to approach, grasp, move, avoid obstacles, and correct itself. Fine-tuning should adapt it to a specific task or robot, not create the ability from nothing. It is also a practical requirement. A robot in a home or a workplace shouldn’t need a brand-new dataset and a new policy every time the task changes, so a foundation model that already carries general skill, and some ability to recover, is the minimum bar for something genuinely useful in the real world.

What is the most challenging part of cross-embodiment learning?

Robots differ in control frequency, delay, compliance, sensing precision, and contact dynamics, so the same instruction can require different action decompositions and recovery strategies, and a behavior that works on one arm cannot simply be copied to another. Cross-embodiment learning needs an intermediate abstraction, lower than language but higher than joint angles: how you approach an object, how you make contact, how you apply force, and how you recover from a mistake.

When we say cross-embodiment, the main capability we mean is multi-embodiment generalization: transferring across robots, training on many embodiments at once, and adapting to different kinematics. Human-to-robot transfer and other techniques are specific approaches to that goal.

“A robot in a home or workplace shouldn’t need a new dataset and policy every time the task changes. A useful foundation model should already carry general skills and the ability to recover.”

What would you most like to see other researchers attempt to reproduce or stress-test?

Three things, above all. Whether event-level representations really generalize beyond our own datasets, across more tasks, scenes, objects, embodiments, and failure conditions. Whether pretraining stays effective on robots the model never saw during training, or whether its capability is still too tightly coupled to what it has already seen. And whether real-robot evaluation can become a shared language for the field, so that we compare not just success rates but the reasons systems fail, where an instruction was misread, where perception broke down, or where recovery fell short. Robotics has been driven too often by impressive demonstrations, and real progress comes from results that are reproducible and diagnosable.

What capability is still missing before robots become dependable in homes?

Benchmarks measure competence, like whether a model can finish a task. Homes demand reliability, safe and consistent operation over time in a place that changes every day, with objects moving, instructions that are vague, and people interrupting. The missing piece is not a higher one-time success rate: it is robust recovery. A dependable home robot has to know when it is uncertain, when to slow down, when to ask for help, and how to bring the world back to a safe state after it drops something or misunderstands a request.

In a real home, failure recovery matters more than raw success, because the home does not reset itself. Homes also demand careful personalization, learning a household’s routines and preferences over time, with safety and trust as first principles. That combination, not any single skill, separates a capable demonstration from a robot people can live with.

Humanoid service robot stands by a table in a modern living room. X Square Robot’s approach is that, in a real home, failure recovery matters more than raw success, because the home does not reset itself and it demands careful personalization, with safety and trust as first principles. X Square Robot

How do the open-source components fit into X Square Robot’s World Unified Model direction?

We see these releases as layers of the World Unified Model direction rather than isolated projects. Wall-OSS-0.5, the action model, asks whether an open vision-language-action model can gain directly measurable capability from large-scale pretraining, so it is the capability layer. WALL-WM, the world model, asks how a robot should understand change in the world, shifting from fixed windows to event-level modeling, so it is the representation layer. The data system supplies the interaction data that both of them learn from.

Together they form a loop in which models produce capability, world models organize understanding, and the open-source community drives reproduction and improvement. World Unified Model is the broader architecture those layers support, bringing vision, language, action, and physical prediction together.

We are releasing these pieces openly because embodied intelligence cannot be solved by one organization; it needs many embodiments, many real tasks, and broad feedback, and the long-term goal is a stack that keeps learning and ultimately moves robots from laboratory demonstrations toward reliable everyday use.

Video Friday: A World Cup for Robots

10 July 2026 at 16:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH
Humanoids Summit Seoul: 22–23 September 2026, SEOUL

Enjoy today’s videos!

For the first time, two full teams of humanoid robots played an 11-vs-11 soccer match on hardware, bringing one of robotics’ most ambitious long-term visions closer to reality. Never before have two full-sized humanoid robot teams played a soccer game against each other.

[ RoboCup ]

Engineers at MIT and EPFL in Lausanne, Switzerland, have designed a robot that can swim underwater and flap out of the water to continue flying through the air, much like a diving bird. The robot can help scientists study the mechanics that enable these actions in aquatic aviators and may help launch a new class of aerial-aquatic drones and vehicles.

[ MIT ]

We’re excited to announce our breakthrough robotic hands for the NEO platform: hands that match or exceed human-level dexterity, strength, safety, and reliability. Designed from the ground up, these 25-DoF hands combine 25 fully actuated degrees of freedom with a tendon-driven system, rich tactile sensing, and built-in compliance. The result is a hand capable of true in-hand manipulation, precision tool use, and delicate interaction.

[ 1X ]

This match, Tech United played against IRIS at the midsize league at RoboCup 2026 in Incheon, South Korea.

[ Tech United ]

Atlas arrived pitchside at NYNJ Stadium in front of 80,000 people gathered to see Brazil vs. Norway. After performing some of the sport’s most memorable player celebrations, Atlas helped kick off the second half by delivering the match ball!

[ Boston Dynamics ]

Navigating discrete terrain such as stepping stones remains a major challenge for legged robots. Conventional approaches often rely on dense environment reconstruction from cameras or lidar, which can be affected by latency, occlusions, and significant computational overhead. We show that proximity sensors integrated into the bottom of a quadruped’s feet enable safe, terrain-seeking autonomous locomotion.

[ Paper ]

On this holiday, Digit is on grill duty. It turns out precise force control is good for more than payload handling. Happy 4th of July from all of us at Agility.

[ Agility ]

We’ve created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99 percent on tasks where previous models achieve 64 percent, completes tasks roughly 3x faster than state-of-the-art, and requires only one hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step toward our mission of creating generalist intelligence for the physical world.

[ Generalist ]

Four years at Figure.

[ Figure ]

Reachy Mini is becoming your real AI companion. The Conversation App makes it able to talk fluently with you, help you with your to-do list, remind you of important tasks, and even chat about music. Long-term memory, voice interaction, always ready to help.

[ Reachy Mini ]

Is this sort of thing now a real job for humanoid robots, then?

[ Unitree ]

Quite a story, but is it a real job?

[ EngineAI ]

If you have a cute animal logo for your research, I will always share it.

[ BIEVR-LIO ]

This is very delicate work, although the real challenge would be picking those nuts out of a jumbled bin full of randomly sized nuts, which is how most of us live our lives.

[ Sanctuary ]

Not for me, thank you, although I’m not saying that most of the other humanoid robots out there are any better looking, fundamentally.

[ UBTECH ]

Robotics professor Dr. Christian Hubicki judges robot soccer skills while knowing very little about soccer himself.

[ ORL ]

In this presentation, Brendan Schulman, vice president of policy at Boston Dynamics, outlines the critical role of government engagement in driving the success of the humanoid robotics industry. He demonstrates how legged robots like the Spot quadruped and Atlas humanoid are moving beyond factory settings to deliver real-world value in infrastructure inspection, industrial manufacturing, and public safety. Schulman highlights the intersection of AI and robotics, showcasing how large behavioral models and reinforcement learning enable robots to navigate slippery floors and autonomously avoid workplace hazards. Ultimately, he calls for a proactive national robotics strategy focused on workforce training, safety standards, and ethical frameworks to support supply chain resilience and global competitiveness.

[ Humanoids Summit ]

Ground Robots Inherit the Kill Zone

10 July 2026 at 11:00


Borys Drozhak has a vision: a front line almost free of humans, patrolled by flying drones and ground robots, and continuously monitored by AI-controlled sensor networks. And it’s not a pipe dream. Ukrainian roboticists have made major strides in that direction over the past four years. Remotely controlled ground vehicles fitted with machine guns and grenade launchers now patrol the no-man’s-land straddling the front, part of a robotic legion that has stymied Russia’s territorial ambitions so far this year.

Drozhak is a co-founder and CEO of RoverTech, which manufactures the Zmiy, one of Ukraine’s most successful ground robots. Zmiy, Ukrainian for snake, is an 800-kilogram (1,700-pound) rover, 2.15 by 1.5 meters in size, with 75-centimeter diameter wheels. The Zmiy comes in various configurations—for demining, logistics, fighting fires, firing a machine gun, or launching grenades.

According to Drozhak, the uncrewed ground vehicle (UGV) is a record-breaker among Ukrainian ground robots. It’s engineered to be nearly noiseless and emit as little heat as possible, helping it to elude Russia’s intelligence, surveillance, and reconnaissance (ISR) drones. As a result, a Zmiy rover completes on average 57 missions across the kill zone before being destroyed. The kill zone is the roughly 35-kilometer-wide swath of land that straddles the front line; its width is variable and determined mainly by the growing range of the drones.

“Usually, a UGV on the battlefield lasts about seven missions,” Drozhak says. “The Zmiy is quite a bit bigger and stronger” in comparison with most other UGVs, “and can make it back even if two of its wheels get destroyed.”

Drozhak is a software engineer turned roboticist whose story is echoed everywhere in the Ukrainian defense establishment. Before the Russian invasion, he was living a quiet life in Ireland, working for an international software development firm. He returned home shortly after the war began to help defend his homeland. Together with his friend, Vasyl Korenovskyi, who had been a mining engineer, he founded RoverTech with the goal of building robots to perform some of the most dangerous tasks in the war zone. In 2023, they rolled out their first product—the Zmiy de-miner. Earlier this year, one of RoverTech’s assault UGVs was part of a widely reported operation that forced a group of Russian soldiers to surrender without the presence of any Ukrainian troops. Such feats, Drozhak insists, are not rare on Ukrainian battlefields these days.

UGVs are the latest chapter in the military-technology race spurred by the war in Ukraine. Scores of Ukrainian startups have developed dozens of different small ground robots, each with typically multiple variants, over the past three years. They’re mostly replacing human-driven tanks and other military vehicles that used to crisscross the war zone. These remotely controlled robotic vehicles cost a few tens of thousands of dollars apiece compared to millions for a traditional tank, and they can be tweaked and modified in frontline workshops to serve the most urgent needs.

Zelenskyy Orders Up 50,000 More UGVs

In April, Ukraine’s President Volodymyr Zelenskyy signed an order for the government to procure 50,000 UGVs for Ukraine’s military forces by the end of 2026. That’s more than three times as many as the government purchased in 2025 and a massive increase from the 2,000 procured in 2024, according to defense analyst Marc C. Lange.

The rise of UGVs, Lange explains, is a direct response to the warfighting revolution ushered in by the speedy evolution of uncrewed aerial vehicles that came to define the war in Ukraine.

As the number of drones zooming above the front line rose and their range increased, the battlefield became completely transparent. Today, anything that enters the kill zone gets hit by a first-person view (FPV) kamikaze drone within minutes.

“Any armored formation, any resupply and logistics vehicle, and any manned formation anywhere near the edge of the battle area has between seconds to a low amount of minutes before it gets turned to dust,” Lange says. “The Ukrainians were losing drivers. Traditional methods of evacuating injured soldiers became impossible. That space is basically unsurvivable.”

Ukraine, suffering from a shortage of infantry, has taken that problem more seriously than Russia, which has a larger pool of fresh recruits to draw from. UGVs began ferrying supplies to troops at frontline positions in 2024. Gradually, they took over the complex and risky evacuations of the wounded, using special enclosures to protect the soldier being transported. But this year, Lange says, is “the year of the assault UGV.”

Emerging Ukrainian tactics combine UGVs with real-time reconnaissance and surveillance from aerial drones, which discover enemy troops, often under cover of night. The reconnaissance data are then used by remote operators who guide UGVs as they stalk, corner, and shoot to kill. Oleg Fedoryshyn, the head of research and design at DevDroid, another prominent Ukrainian UGV developer, said the ground robots can be controlled from as far as 100 kilometers away using Starlink connectivity, LTE networks, or mesh-networked military radio systems. The UGVs can also carry strike UAVs (uncrewed aerial vehicles), serve as communication relays for drones, or carry and launch communication relay drones that further extend the range of the attack vehicles. The UGV can lurk in position for up to one week without needing a battery charge, Fedoryshyn said, and wait for the enemy to move closer.

“It’s better than to put people there,” he notes. “A guy with a machine gun is always the first target for the enemy.”

An Ukrainian soldier adjusting an unmanned ground vehicle\u2019s machine gun. The Droid TW 12.7, by DevDroid, is shown here outfitted with a 0.50-caliber M2 Browning machine gun that can be aimed and fired by a remote operator using a tablet and an encrypted communications link.DevDroid

Fedoryshyn estimates that UGVs could eventually help cut the number of soldiers needed along the front line by 30 to 40 percent. Drozhak is even more ambitious. He envisions a future front line that’s entirely automated, relying on sensors and other systems that are only occasionally serviced by humans.

A guy with a machine gun is always the first target for the enemy.

“Right now, we need a lot of UGVs because there are people on the front line and we need to deliver supplies to them,” he says. “But we can substitute many of them with sensor systems, servicing robots, and UGVs, and then we will not need that many for logistics. At some point, we could have only robots in the kill zone.”

Ukraine, with a prewar population of around 41 million, has lost over 150,000 fighters in the war since 2022, according to estimates by the Center for Strategic and International Studies and others. Many thousands of others have been mutilated or permanently disabled. Even those who return without physical injuries suffer lasting psychological trauma. Drozhak dreams that a future robot army would put an end to the ability of autocratic regimes worldwide to brutalize their neighbors.

“There will be no need to push people on the battlefield anymore,” says Drozhak, the RoverTech CEO. “Once we achieve that in Ukraine, any country with a decent economy would be able to defend themselves just with technology.”

RoverTech’s Tarantula active-protection system, which uses acoustic and visual sensors combined with AI algorithms to detect approaching killer drones, is the first step in that direction, he declares.

“The future battlefield will rely on networks of robotic sensors and autonomous systems that can continuously monitor dangerous areas, provide early warning, and reduce the need for soldiers to expose themselves to direct threats,” he says. “Human operators will remain responsible for critical decisions, but increasingly advanced sensing technologies will help move people away from the most dangerous positions on the battlefield.”

Why UGVs Are Vulnerable

Militaries around the world were looking at UGVs prior to Russia’s 2022 invasion of Ukraine. But those were quite different, explains Samuel Bendett, a defense analyst at the consultancy CNA. They were larger, more complex, and conceived to operate in smaller numbers. The more compact forms now seen in Ukraine are the result of an evolution that paralleled that of the first-person view (FPV) attack drones. Both needed to be cheap as they don’t last long and small to be less conspicuous. Now, the West is trying to understand the overall role of UGVs in future warfare. So far, in Bendett’s view, the impact of UGVs on warfare isn’t as profound as that of the FPVs and other aerial drones.

“Not every terrain would be applicable to using a UGV,” Bendett explains. “So far, a lot fewer countries are seeking to integrate them into their combat operations than UAVs, which very much democratized the way of enabling short-range to mid-range strikes against adversaries.”

UGVs, he points out, are much more susceptible to communication disruptions than UAVs, while being less suitable for autonomous operations and swarming due to the complexity of ground terrain.

“With UAVs, communication is much easier,” according to Bendett. “There are no interferences between the ground station and the UAV save the distance, Earth’s curvature, and the radio horizon. But on Earth, there’s lots of different obstacles that interfere with radio signals.”

Most UGVs rely on Starlink as the first choice for operator control, but even that comes with problems. Starlink signals are easily disrupted by trees and buildings. And Russia, having been cut off from Starlink, is working hard to find ways to jam the system.

On top of that, Lange says, as UAV autonomy progresses, UGVs could be left behind. The reason is that UGVs are likely to remain dependent on operator communication links for some time and will therefore be vulnerable to enemy UAVs that can’t be stopped by jamming systems that still provide some protection today.

“The low production cost of strike drones will mean that UGVs will have to endure a barrage of strikes that might be too much,” Lange says. “The question is whether you can make UGVs more survivable on the front line both in terms of command and control and the actual survivability of that many strikes.”

Still, he thinks there’s “no path back from UGVs.” The idea of distributing a whole range of tasks, performed in the past by a single large and expensive tank, to a fleet of small, cheap UGVs provides more resilience against the omnipresent drones. Moreover, although many international commentators now say that Russia appears to be losing, the war grinds on—and so does the cat-and-mouse game of lethal innovation.

IEEE Honors Robotics Pioneer Toshio Fukuda

7 July 2026 at 19:02


Toshio Fukuda has been blazing trails for most of his career. He is considered to be one of the most prolific scholars in robotics, writing more than 2,000 research papers and authoring several books on the field. He’s an influential figure thanks to his pioneering work developing biomedical robotic systems, industrial robots, micro-nano robotics, mechatronics, and AI-driven automation.

Fukuda launched one of the first robotics conferences, the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). It is still popular almost 40 years later.

Toshio Fukuda


Employer

Egypt-Japan University of Science and Technology, in Alexandria

Title

Professor and vice president of research

Member grade

Life Fellow

Alma maters

Waseda University, in Tokyo; University of Tokyo

An IEEE Life Fellow, he is a professor emeritus in the department of micro-nano systems engineering and a visiting professor at Nagoya University, in Japan, where he taught for nearly 25 years. Currently, he is a vice president of research at the Egypt-Japan University of Science and Technology, in Alexandria, Egypt.

Within IEEE, Fukuda has held top volunteer positions including the organization’s highest office: He served as IEEE president in 2020, becoming the first person of Asian descent to hold the role.

He’s a former program director of Japan’s Moonshot program, which by 2050 intends to develop advanced AI robots.

Born in Japan, Fukuda has been recognized by the country for his contributions to science with two of its highest awards: the Medal of Honor with a purple ribbon in 2015 and the Order of the Sacred Treasure in 2022.

IEEE honored him with this year’s Richard M. Emberson Award for “distinguished service advancing the technical objectives of IEEE, especially in the area of robotics.” The IEEE Board-level award is sponsored by the IEEE Technical Activities Board. Fukuda received the award on 24 April at a ceremony in New York City.

As a former IEEE president who has served as a master of ceremonies at several of the organization’s major award events, Fukuda noted that he is more accustomed to bestowing awards than receiving them.

“It’s very interesting to be on the receiving end,” he says.

The journey into robotics research

As a teenager, Fukuda spent his summer breaks teaching himself how to build things including transistor radios and steam engines.

“It was very nice to have a hands-on hobby and make these kinds of things myself,” he says. His experimentation led him to study engineering.

He earned a bachelor’s degree in engineering in 1971 from Waseda University, in Tokyo. He says one of his professors there—Ichiro Kato, regarded as the father of Japanese robotics research—was a good mentor who made a positive impact.

Fukuda’s research interests were robotics and mechatronics, a field that combines robotics, electronics, computer science, and control systems.

He went on to earn a master’s degree and a doctorate in science from the University of Tokyo, in 1971 and 1977. During those years, he also attended Yale, where he conducted research on advanced control theory in 1973.

He reflects fondly on his time at Yale: “It was a very nice environment and a kind of free-thinking atmosphere. It motivated me to study more.”

“IEEE doesn’t care who you are, what you do, what country you are from, or whether you are male or female. IEEE accepts people who have energy and passion.”

While at Yale, Fukuda served as an assistant to his advisor—which led him to consider a career in academia, he says, because he enjoyed the freedom that research work afforded him.

But he realized that such freedom comes with a price. University researchers are expected to raise the money that funds their work. He compares researchers to small-business owners who have to bring in money to keep their enterprise afloat.

That realization led him to select robotics as his field because he intended to develop technologies useful to industry, he says.

After earning his doctorate, he returned to Japan in 1977 to work as a research scientist at the government’s Mechanical Engineering Laboratory, later renamed the National Institute of Advanced Industrial Science and Technology, in Tsukuba.

“There was a lot of research going on at the lab, including practical robotics and theory,” he says.

He left Japan in 1979 to become a visiting research fellow at the University of Stuttgart, in Germany. During his year there, he studied systems, software problems, and related topics.

He returned to Japan and was hired as an associate professor of mechanical engineering at the Tokyo University of Science. He conducted research into practical uses for robots by visiting industrial plants. He decided to develop robots that inspect industrial equipment such as those used in assembly plants, oil refineries, and power stations—places that “can be hostile environments for humans,” he says.

His work drew interest from chemical, oil, and utility companies.

“I got a lot of money from them for this very practical application, which funded my research,” he says, laughing.

Developing popular robotic systems

Fukuda grew tired of making those robots, he says, so he switched to creating ones for scientific applications. He developed many techniques, but he probably is best known for his modular, cellular robotic systems (CEBOTs), which he introduced in 1985.

He has described how CEBOTs work in numerous papers published in the IEEE Xplore Digital Library.

The CEBOT system is composed of a number of autonomous robotic cells that stick together like interlocking Lego plastic bricks, he says.

Each cell is a fundamental modular unit that has a function. When a simple task is given, the system can analyze it and generate the structure of the cellular manipulator. The cells connect to and detach from each other through connection mechanisms and cooperate mutually, creating complex structures and configurations.

“You start developing from the component-wise to the cell-wise to a small functional unit—and then you come up with clusters that make bigger systems. We can make a society of robot beings like that,” he explained in his oral history published on the Engineering and Technology History Wiki. “It’s a distributed robotic system, a self-organized robotic system, and also an evolutionary robotic system.

“It’s also a fault-tolerant robot system because if something is wrong, you just remove those things and make a new one. You keep the system working. That’s a great thing.”

Today CEBOTs are used for a variety of tasks such as delivering medication in hospitals, assisting with planting crops, and transporting products in distribution centers. Check out IEEE Spectrum’s Robots Guide for news from the world of robotics.

In 1989 Fukuda joined Nagoya University as a professor of mechanical engineering and micro-nano systems engineering. During his 24-year career there, he was director of the university’s Center for Micro-Nano Mechatronics. He developed a long list of technologies at the university, including many for medical applications. He also conducted groundbreaking research into intelligent robotic systems and micro- and nano-robotics.

Another technology he is known for is brachiation robots, which he helped develop in 1988. He calls them monkey robots because they’re based on the pendulum-like movement of monkeys swinging from tree to tree. The gravity-based locomotion enables continuous movement.

Brachiation robots now are inspecting high-voltage transmission towers and bridges, searching damaged buildings for survivors, and performing maintenance on pipelines and cables.

Fukuda retired from the university in 2013 and was named professor emeritus.

He didn’t stay retired for long, though. He next held a teaching appointment at Meijo University, in Nagoya, until he left in 2022 to join the Egypt-Japan University.

A prominent volunteer

He joined IEEE in 1980 at the encouragement of one of his research advisors, Professor Fumio Harashima, now an IEEE Life Fellow. After attending conferences and reading the organization’s publications, Fukuda says, he looked forward to becoming more involved.

“I wanted to know how to organize a conference and how to edit a paper for one of its Transactions,” he says. “I wanted to know what was going on from inside the organization, not just the outside.”

In 1988 he was the founding chair and organizer of IROS, in Tokyo. The conference had 330 attendees that year, and was supported by Harashima. Today it is one of the largest and most prestigious conferences on the topic, attracting more than 9,000 people annually. Out of 120,000 conferences, it was the only conference in the Nature Index database for this year, Fukuda says.

In 1996 he and other members launched IEEE Transactions on Mechatronics.

He was the founding president of the IEEE Nanotechnology Council, which was established in 2002. He is considered a pioneer in nanotechnology research, particularly regarding how it relates to robotics.

Over the years, he has held numerous volunteer positions on IEEE editorial boards and committees.

He was the 1998–1999 president of the IEEE Robotics and Automation Society, becoming the first non-U.S. member to hold the title.

He was director of IEEE Division X (2001–2002 and 2017–2018), which covers intelligent systems, biological engineering, robotics, control systems, and photonic technologies. He served as the 2013–2014 director of IEEE Region 10 (Asia-Pacific).

As the 2020 IEEE president, Fukuda saw the organization through the early part of the COVID-19 pandemic. Because of travel restrictions, he realized IEEE should change how it offered its in-person services, specifically educational programs. He encouraged IEEE Educational Activities to develop an online learning platform. The IEEE Learning Network started with just three courses and now offers nearly 2,000 courses, webinars, and learning materials.

An award-winning member

The Emberson Award joins a slew of other recognitions Fukuda has received from IEEE. They include several from the IEEE Robotics and Automation Society: a 2004 Pioneer Award, a 2009 Saridis Leadership Award, and the 2011 Harashima Award for Innovative Technologies. He is also a recipient of the Board-level 2010 IEEE Robotics and Automation Technical Field Award.

He says he feels strongly that IEEE should be a diverse organization that is welcoming to all. As IEEE president, he led efforts to devise a diversity, equity, and inclusion program. Several policies, procedures, and bylaws were revised to give members a safe, inclusive place for discourse.

“It’s important for IEEE to make everyone feel comfortable,” he says. “DEI programs are important. All people should be equal. IEEE doesn’t care who you are, what you do, what country you are from, or whether you are male or female. IEEE accepts people who have energy and passion.

“It accepted me, from the Far East. That’s why I like it.”

You can learn more about Fukuda and his career from the oral history conducted by the IEEE History Center.

Japan Pioneered Humanoid Robots—Can It Now Catch China?

4 July 2026 at 11:00


“In the future, the relationship between humans and robots will deepen, and the distinction between them will probably disappear.” This prediction, from one of the attendees at the recent Humanoids Summit in Tokyo, might have been unremarkable had it not come directly from an android that was first introduced to the world 20 years ago.

Geminoid HI-6 is the sixth-generation of a robot originally designed in 2006. The mechanical twin of Osaka University professor Hiroshi Ishiguro, Geminoid HI-6 is now equipped with a large language model trained on Ishiguro’s own writings and interviews. It has advanced conversational skills and can even have a chat with its creator, an eerie spectacle. But at the Humanoids Summit, Geminoid was one of the few humanoid robots from Japan, the country that pioneered the form factor.

While the event in Tokyo had only about 40 robots on display, Chinese systems outnumbered Japanese by roughly three to one. Some Japanese robotics firms were even using Chinese robots in their own technology demonstrations, something that would have been unthinkable in the recent past—one Japanese engineer described the situation as “sad.” The conference was a stark reminder of how Japan has ceded its early lead in humanoid robot development to overseas competitors, and the challenge it now faces to secure a place in an ecosystem increasingly dominated by general-purpose robots powered by AI.

Twenty-five years ago, Japan was turning out groundbreaking humanoids that were showstopping in their abilities, but they were not commercialized as practical machines in any meaningful way. Heavily influenced by science fiction and lacking practical applications, they were mostly expensive technology demonstrations that were eventually mothballed. What Japan retains, however, is robotics design and know-how, which it must leverage to be a key player in the rapidly evolving humanoid ecosystem.

Learning to Walk—Then Standing Still

To anyone who has seen recent videos of Chinese humanoids doing kung-fu and synchronized acrobatics, as well as half-marathon races, China’s remarkable progress in the field is nothing new. At the Humanoids Summit, Toyota showed a video of its latest basketball-playing robot, and Honda exhibited its latest robot hand, but the full-scale humanoids on the floor were mostly Chinese–the kid-size K1 machines from Booster Robotics of Beijing were dancing to Michael Jackson tunes. The full-scale G1 humanoid from Unitree Robotics of Hangzhou was also doing demos.

“You cannot sell these bipedal systems in Japan for safety and compliance reasons,” says Shuichi Nagao, a frequent visitor to China as CTO of Omakase Robotics, a division of Zeals, a Japanese humanoid robot developer. Omakase was exhibiting a G1 modified with an external PC controller, a dextrous hand, a suction-cup manipulator and a sensor “hat” with an extra speaker, mic, and camera.

“In China, the government is pushing humanoid development. They didn’t have an industry 20 years ago. The people pushing it are young, in their 20s and 30s. It’s a really different mentality out there,” says Nagao. “Big players in Japan are still looking for use cases for humanoids. In China, they’re already doing mass production and reducing the cost, so other countries can’t compete with them anymore.”

Another Japanese company showing off G1 bots was summit sponsor GMO AI & Robotics, a subsidiary of Japanese internet company GMO. It’s using the robots in partnership with Japan Airlines to load and unload cargo containers at Tokyo’s Haneda airport. The cargo project is a trial—like many other humanoid experiments—but the fact that Chinese machines have penetrated so far into Japan’s ecosystem upends a long history.

In 1973, scientists at Waseda University in Tokyo built WABOT-1, considered the first full-scale humanoid robot, which was capable of slow bipedal locomotion, grasping objects, and simple communication. It inspired Honda’s groundbreaking Asimo humanoid, but Asimo was never commercialized. It was eventually retired in 2022, the year ChatGPT was released. Two years later, Unitree’s G1 went on sale for US $16,000.

A 65 centimeter tall bipedal robot, its design features the head of an anime-style girl with legs directly underneath. China’s High Torque Technology Co. showed off its Mini Pi biped, customized with an anime-inspired head, at Humanoids Summit in Tokyo. The regular version is priced at $3,500. Tim Hornyak

Supply and Demand

Japan’s development of humanoids happened before practical applications or widespread demand were in place, but bad timing is only part of the story—Japan also has a history of developing technologies that might appeal to domestic consumers but not necessarily those overseas. For example, decades after its highly engineered multifunction toilets first appeared, they have only recently found a following abroad.

Japan’s humanoid prowess was partly built on the back of its legendary industrial automation, yet even that stronghold has eroded. Ani Kelkar, a partner from McKinsey & Company in Boston who produces analytical reports about the robotics industry, told the summit audience that while Japan occupied the top spot in the world in manufacturing robot density (the number of multipurpose industrial robots in operation per 10,000 employees) from at least 1994 to 2009, it then slipped to second in 2014, third in 2019, and fifth in 2024. In that year, South Korea was at the top of the leaderboard with a robot density of 1,220 compared to Japan’s 446.

The International Federation of Robotics estimates China now has the most operational industrial robots in the world, with around 2 million total units, approximately 4.5 times more than Japan. “The annual installation numbers are impressive too: 54 percent of all robots installed worldwide in 2024 were deployed in China,” the IFR said in a release in April 2026.

“I think the loss of Japanese leadership is more to do with the rise of China as a manufacturing powerhouse including for sectors that Japan had high export levels,” Kelkar said in an email interview. “The recovery has not yet happened as Japan “missed” the rapid acceleration in AI for robotics and is now playing catch-up.”

How Japan Can Adapt

Kelkar believes Japan has a $100 billion opportunity in general-purpose robotics, which are machines that can perform a wide variety of tasks, and it cannot rely on the slower-growing industrial robot market, which is centered on factory machines that do one simple and predictable task like welding car parts. He points to a McKinsey white paper suggesting that while Japan has much of the hardware and technology experience needed to support general-purpose robot development, it must change its strategy to capture a larger share in AI, software, data collection, and robotics platforms.

Tetsuya Ogata is a professor of engineering and director of the Institute for AI and Robotics at Waseda University, the birthplace of humanoids in Japan. He briefed the summit on how a nonprofit he chairs, the AI Robot Association (AIRoA), is working with Toyota and other members to develop foundational technologies for collaborative use.

For instance, AIRoA has collected some 80,000 hours of data on remote operation of mobile manipulators, which Ogata believes is the largest dataset of its kind. Using the data, it built and verified vision-language-action (VLA) models, and it has also started data collection for dual-arm mobile manipulation. In an interview, Ogata acknowledged Japan’s struggle to find its place in the changing landscape.

“The world of AI is inherently a game of scale,” says Ogata. “Therefore, Japan’s absolute prerequisite is to secure a competitive baseline of scale—in data, computing resources, and talent. Beyond that, what I consider most critical is a mind-set shift: Rather than trying to hoard scale within a single nation or company, we must grow stronger by collaborating with a diverse ecosystem of domestic and international players.”

Specifically, this means creating a “collaborative domain” to address data—the single biggest bottleneck—through industry-wide cooperation rather than data siloing. By collectively nurturing a precompetitive, shared data infrastructure and foundation model, individual companies can then compete on top of it with their own applications. “By offering this open ‘data ecosystem’ to the world, we can engage global players and establish a ‘third pole’ alongside the U.S. and China,” says Ogata. “I believe this is how Japan can reclaim its global presence.”

In 1999, Japan introduced the world’s first mobile internet services platform. But being first didn’t turn Japan into a smartphone manufacturing or design center—it’s now merely a supplier of parts to other countries that are leading the smartphone industry. If Japan can avoid a repeat of that experience and successfully deregulate, diversity, and commercialize its original humanoid dreams, it stands a better chance of influencing the direction of the industry and reaping billions in value. As automobiles and electronics were pillars of Japan’s industrial strategy in the last century, Japan could make humanoid robots one of its key value generators in the 21st century, an approach that would not only deliver economic benefits but give Japan greater clout in how the industry will evolve. Just like Japanese cars, electronics, and even toilets, Japanese humanoids could stand for craftsmanship and reliability. It’s a legacy that Japan can’t afford to give up.

This article appears in the September 2026 print issue as “Japan Seeks a Humanoid Robot Comeback.”

Video Friday: An Earthbound Mars Rover for the Moon

3 July 2026 at 15:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH

Enjoy today’s videos!

NASA is considering a mission concept for an advanced, nuclear-powered rover to be deployed to the Moon’s South Pole as part of the agency’s Moon Base plans. The PROMISE (Polar Rover for Observation, Mapping, and In-Situ Exploration) mission concept relies on the Curiosity Mars rover mission’s testbed rover. Some elements of the Perseverance Mars testbed rover shown in this video could be used as well. As exact duplicates of Curiosity and Perseverance, the testbed rovers are equipped with flight-proven engineering systems capable of carrying technology as well as science instruments that would advance Moon Base efforts.

A Mars rover for the Moon? That’s some OPTIMISM right there.

[ JPL ]

This is the absolute best thing since Festo’s AirPenguin.

The project explores soft, lightweight robots that can gently float around people in indoor environments and invite playful, affectionate, and everyday interactions. Unlike conventional drones, our robot is designed to be quiet, soft, touch-safe, and socially approachable. Through this work, we ask what future indoor companion robots might feel like if they were not rigid machines, but gentle floating beings that share space with us.

[ Paper ]

Thanks, Mingyang!

Today, we’re launching our home robot, Isaac 1. Deliveries will begin this fall.

US $500 per month, with some basic task autonomy, plus teleoperation.

[ Weave Robotics ]

A couple of things from this new Figure video: Thing one is that the cart-pulling is a good illustration of how clumsy humanoid robots still are at basic tasks relative to humans. Thing two is that there are absolutely no humans anywhere near these robots. You can see one guy at 0:19, which I can only assume is an accident, because these robots are not safe to be around from an industrial safety perspective.

[ Figure ]

Our very own Kohava Mendelsohn met some robots at ICRA in Vienna, and only one of them was murderous.

[ ICRA 2026 ]

Welcome to Robot Park, where we’re building the future with Apollo 2. Robot Park is where Apollo learns today, getting the experience needed to make a difference tomorrow. Today we’re announcing Robot Park, our nearly 90,000-square-foot facility where Apollo 2 is collecting real-world training data needed to advance autonomous humanoid robots.

[ Apptronik ]

UBTech Robotics, the world’s first publicly traded humanoid robot-maker, has launched a humanlike robot that features lifelike silicone skin and “emotional AI,” as Chinese tech firms increasingly transition robots from the factory floor to the family living room.

[ SCMP ]

Spherephones are redefining how we experience sound. Created at Georgia Tech, this wearable uses spatial audio to alert users to movement from every direction—including behind and below. Built for safer human-robot collaboration, the technology is expanding into gaming and accessibility applications. See how music is becoming a new language for awareness and interaction.

[ Georgia Tech ]

Humanoid robots are meant to carry out long-horizon autonomous missions in a world built for humans. This is hard. These missions consist of many steps, each of which requires them to perceive, navigate, and interact with the environment. This is exactly Flexion’s goal: building the general-purpose intelligence that turns any robot into a useful helper.

[ Flexion ]

We’re introducing KinetIQ Ascend—our reinforcement-learning approach designed to reach 99.9 percent manipulation reliability at human speed and beyond.

[ Humanoid ]

Dr. Sebastian “Basti” Scherer has worked in field robotics since the first DARPA Grand Challenge in 2004. He runs the AirLab at Carnegie Mellon’s Robotics Institute and is the director of safe embodied AI at FieldAI. While much of the industry is focused on local skills like tabletop manipulation, Dr. Scherer sees the greatest value in solving dirty, dull, and dangerous tasks that require operating in uncertain environments where the robot needs to “just work.” When robots “just work,” they become less like robots and more like tools. “That’s the big challenge that we have to overcome,” he says. “And that’s the challenge that FieldAI is really primed to solve.”

[ FieldAI ]

Look, I really appreciate how valuable robots like ElliQ can be, and robots that do good work and offer a financial benefit are incredibly important, especially in the context of family care. But in my opinion, you really shouldn’t suggest that a robot with FaceTime or whatever is an equal replacement for in-person human companionship, nor should you suggest that AI can replace a human wellness coach. If you can’t afford those things, then sure, ElliQ can offer some of those capabilities in a very limited way, but that’s all.

[ ElliQ ]

Very cool moves! Now get a job!

[ DEEP Robotics ]

Drawing inspiration from restaurant waiters in Morocco and Turkey, among other places, we equip a robot with a hanging tray to transport objects from one location to another without dropping them or spilling their contents. We incorporate this approach into an interactive robot waiter demonstration, which uses computer vision and visual servoing to steer toward a person with a raised hand to serve them.

[ Paper ]

If you’re going to make robots wear skirts or shorts or pants, you have to give them butts, or it’s just not going to work. That is all.

[ TechShare ] via [ Kazumichi Moriyama ]

It’s Los Alamos, so of course we have robots. Some work inside gloveboxes, while others probe unexploded ordnance in the field and aid with repetitive lifting, Doc Ock–style. Legend has it there’s a fro-yo robot in the cafeteria.

[ LANL ]

Here are a couple of talks from the recent Humanoids Summit in Japan, from Ali Agha of FieldAI as well as Hiroshi Ishiguro.

[ Humanoids Summit ]

Video Friday: Give Robots a Hand

26 June 2026 at 16:30


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH

Enjoy today’s videos!

The best way of introducing a new robot hand is to have a disembodied one crawling across a table.

[ Tangent Robotics ]

MIT CSAIL’s Improbable AI Lab Director Pulkit Agrawal explains his “SoftMimic” approach to making robots safer around humans.

[ SoftMimic ]

I now have absolutely no interest in a humanoid robot for my home unless it can do this.

[ PNDbotics ]

The DARPA Lift Challenge is open to the public August 6-9, 2026, at the National Museum of the US Air Force.

[ DARPA ]

Getting Digit to step and shuffle around an obstacle on the floor is a real test of reactive footstep planning. Digit has to spot something small and moving, recalculate where to place each foot, and keep working—all without breaking stride or losing balance. That’s the same dynamic footwork Digit uses to navigate clutter and foot traffic on a real warehouse floor.

[ Agility Robotics ]

This is the most aggressive firefighting robot I’ve ever seen.

[ DEEP Robotics ]

Wait a sec, Dusty can print things on floors besides construction layouts? How is this not in every city, making sidewalks exciting and fun everywhere?!

[ Dusty ]

I am the first to admit that for US $4,900, the performance of the Unitree R1 is very impressive. But what is it going to do out in the world such that it will give you some sort of return on that investment?

[ Unitree R1 ]

Event cameras are extraordinarily powerful because they can see motion, but what if everything is moving because your camera is moving? Oh no!

[ University of Zurich Robotics & Perception Group ]

Can we understand whale behavior and language? Harvard SEAS Professor Stephanie Gil explains the possibility of understanding animal language and behavior using AI-driven robots and machine learning. With ongoing whale research and advancements in artificial intelligence, the potential for animal communication with whales could become a tangible reality.

[ Harvard SEAS ]

Rodney Brooks, founder and chief technology officer of Robust.AI, sits down with Forbes Assistant Managing Editor Kerry Dolan to discuss how he came up with the idea of the Roomba vacuum cleaner and the future of robotics.

[ LinkedIn ]

Here are a couple of interesting presentations from UIST 2025, including everyday objects that move around your home with a mind of their own and a project featuring teamwork between helium balloons and ground robots called Buoyancé.

[ UIST 2025 ]

Video Friday: Do Robots Even Need Legs?

19 June 2026 at 15:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH

Enjoy today’s videos!

Eno is our first agentic robot: an AI agent and a general-purpose robot working as one system. It reasons, plans, and acts in the real world. Human in capability, not in form. Every detail with a purpose, reduced to what matters. Designed not to resemble us, but to extend us. Eno is built end to end at Genesis.

[ Genesis ]

Engineers from NASA’s Jet Propulsion Laboratory are field-testing advanced capabilities for potential future Moon and Mars rovers. In the Colorado Desert near Plaster City, California, teams used a prototype rover called ERNEST (Exploration Rover for Navigating Extreme Sloped Terrain) to test software for a potential future long-range lunar mission. The software enables the rover, developed at JPL, to operate autonomously and travel extreme distances with minimal intervention from human operators.

ERNEST is a lot more capable than it may look; here’s some recent research showing the kinds of terrain it can handle:

[ NASA's Jet Propulsion Lab ]

Table tennis can produce moments that are difficult even for experienced players to anticipate…like when the ball clips the net and suddenly changes direction. For the Ace research project at Sony AI, these events were a key test of the system’s ability to operate reliably in unpredictable real-world conditions. Ace addresses this uncertainty by simulating counterfactual ball trajectories in real time. In the video, the green overlays show these alternative paths the system considers while planning its response.

And check out some of these rallies that the robot has with Miyuu Khiara.

[ Sony AI ]

This video of an ANYmal deployment in a concrete plant is worth watching because it makes explicit how quadrupeds make money in inspection contexts: Among other things, “a cracked crusher foundation [was] caught before a week-long shutdown, avoiding roughly $630,000 in lost production.” That pays for a lot of robots.

[ ANYbotics ]

A lot of interesting footage here from GITAI’s prep for a robotic satellite servicing demo mission. The thruster test-firing isn’t a robot, exactly, but it may be the coolest part.

[ GITAI ]

Anyone who’s tried to take a half decent photo underwater knows that it’s basically impossible, so let’s try and teach robots to cope.

[ Bi-AQUA ]

Thanks, Masato!

Handling delicate, irregular or unpredictable objects is one of the hardest problems left in automation, and one of the most important. It’s what’s holding back the next wave of robots from doing more in the real world. That’s why we’re working with PSYONIC on a new approach. Their Ability Hand, worn by hundreds of people every day, captures real-world data on touch, pressure and grip. Our GoFa cobot brings the industrial-grade accuracy and repeatability to turn that human data into reliable robotic performance.

[ ABB Robotics ]

Sanctuary AI has achieved world-class performance on a complex wire-plugging production task with a global Tier 1 automotive supplier. In this demonstration, Sanctuary AI’s Physical AI successfully performs a high-speed wire-plug insertion task, achieving a validated task success rate of over 99.5% with a cycle time of just 2.54 seconds, meeting live production benchmarks established by the customer.

WHY IS THIS STRESSING ME OUT SO MUCH?

[ Sanctuary ]

This video is quite obviously fake, but I suppose maybe there’s a market for extra beefy quadrupeds? Maybe?

[ Kepler ]

I cannot overstate how much I do not want any robot to look at what I’m wearing and then attempt to sell me things based on what it thinks it can guess about my personality or interests.

[ MagicLab ]

I am here for fed-up robots learning how to move boxes by just kicking them.

[ ATARI Lab ]

Ah, yes, very useful and very important robots that make me very uncomfortable.

[ Paper ]

I built GrowBot ( a ~6”, two-servo bipedal robot) that runs entirely on a $15 Raspberry Pi Zero 2 W, ~$100 in parts. An LLM drives it directly: it reads the raw IMU stream with no translation layer and narrates its own motion (“rocked side to side like a baby”), riding on a 50-Hz reinforcement-learning walk policy trained in sim and transferred to the real body.

The idea here is to build an open course around this project, Brit says, “so everyone can experience physical AI right now in a low-risk way.”

[ GrowBot ]

Thanks, Brit!

What Amazon’s Astro Taught Me About Giving Robots a Soul

19 June 2026 at 10:00


In 2018, Amazon brought me in as the lead UX Sound Designer for Astro, its first consumer home robot. Astro used cameras and other sensors to map and navigate your home and workplace, and could proactively patrol, check up on loved ones, and transport small items using its built-in cargo bin. While there was a well-defined feature set and form factor, initially there was no character direction. In fact, even before Astro had a name, there were two main questions—was it simply Alexa on wheels, or was it a robot with its own character?

The Astro team was divided. One option was to focus on Alexa, and treat the mobile robot simply as an added utility. Along with the majority of the UX team, I argued for Astro to not focus on Alexa. Our belief was that a thing that moves through your home and turns toward you with intent can never be just an appliance. People would ascribe character to it whether we wanted them to or not, and so the only question was whether we shaped that character or let it happen by accident.

Ultimately, Astro became Astro rather than Alexa, and user testing backed up our decision. People didn’t see the robot as Alexa. They saw it as its own character, and that’s what they wanted it to be. Alexa on the device felt somewhat strange and creepy, but building Astro its own voice was too slow and expensive in 2018. So, we settled on Alexa as a supporting character that handled any actual talking, while Astro was the main character, communicating as much as it could without words, through sound, motion, and facial expressions.

I had been brought on to the Astro team to define the robot’s sound design language and voice. But there was no one to flesh out the robot’s actual character. You cannot make a single real decision about a character without defining it first. Every choice about how Astro moved, sounded, paused, or reacted was a character choice, and those choices required all disciplines working together. As sound lead, I was weaving together sound, motion, and character, and how they played together inside each story moment. The animators, who programmed Astro’s motion and facial expressions, were extraordinary at what they did, but the emotional arc they were animating came from the sound (and therefore character) work first. So I stepped into that role, which is where my real work started. What I learned about building character for robots applies to nearly everything being built in embodied AI right now.

Character Is a Design System

Developing a character for Astro meant answering questions that had never been asked about a product at Amazon: What is the emotional range of this robot’s baseline state? How does this robot communicate uncertainty without eroding trust? Where is the line between being expressive and annoying? What are the vulnerabilities of this device’s character?

These are design questions. They have real answers, and every team working on the product has to build from them. For example, Astro’s emotional range was designed to be relatively small at first. We never wanted Astro to get too sad or too angry. It could play sad, but would snap out of it quickly and end the reaction on a high note to keep things positive.

Character leaks out of every seam and can create a disjointed experience if not defined correctly. Even if it’s just animation timing that’s slightly off, or a response that’s technically correct but contextually tone-deaf, users feel every one of these inconsistencies, even if they can’t name them. Watch what happens at the beginning and end of this Sing sequence:

Astro goes from nothing, into the emotional moment, and then lands back on nothing. No buildup, no cooldown, no sense that the feeling came from somewhere or had anywhere to go. I pushed hard for better character stitching, the transitions in and out of expressive moments that make a performance feel continuous rather than assembled, but it never got implemented. The moment itself works. But without the stitching, it reads as a clip playing on a robot rather than coming from within the robot character itself.

Story and Sound at the Beginning

We had decided that Astro would have no spoken dialogue, but it had something that functioned the same way: a vocabulary of sounds, tones, and rhythms that acted as its voice. This vocabulary became the leading output of the character’s personality. The robot’s motion and facial expressions were built around it.

Astro’s wake-up sequence is a great example. Waking wasn’t just a boot animation on the screen; it was an entire performance. Slow and humble at first, the robot oriented itself quietly, then stretched its screen, checked its wheels, and finally, with an upward gesture toward its telescoping mast, it popped it up slightly, and did a little dance of joy. Sound, motion, and eyes hit every beat together in full choreography.

The character’s output in that sequence was first written as a story. Astro is waking up in its new home for the first time. Its main aspiration is to be part of a family, so this is the moment it has been waiting for, this is its purpose. Being the responsible character that it is, it wants to make sure everything is good to go before it introduces itself and starts learning its new home.

This narrative came first because it drove every other decision that we made. After the story was written, sound gave that story a metaphorical voice: the excited tones, the pacing as it checked its wheels, and the bright melodic phrase as Astro looked up at its new family for the first time and introduced itself. Once the sound was laid down, the animation team did their thing with motion and facial expressions, taking cues from the emotional arc the sound had established. Motion didn’t lead—it followed the feeling of the story and the sounds, the same way an animator follows a recorded vocal take.

That wake-up sequence became one of the most-discussed moments in early user testing. People described it as “alive.” What they were responding to wasn’t any single element. It was all three channels (sound, motion, and facial expressions) expressing the same defined character in harmony.

Context Is Where Character Becomes Real

The most compelling characters are defined not by a fixed disposition but by how they respond to their environments and the people in them. They’re still recognizably themselves even as they adapt. This is what I call contextual character. A robot living in a home doesn’t occupy a single emotional state. It moves through rooms with different energy, encounters people in different moods, operates at different times of day, and responds to an endless range of social situations it was never explicitly designed for.

We got close to a contextual character output with Astro’s sound. When a specific piece of environmental context was fed in, the system adapted beautifully, and Astro felt completely alive. But every state like this was still a prediction we made by hand—a situation we had to imagine in advance and design a response for. A random home throws more situations at a robot than anyone can possibly predict, so there was always a longer tail of moments the system was never prepared for.

The difference between a product people describe as “smart” and one they describe as “aware” often comes down to this. Smartness is capability. Awareness is context. Presence is character. And character is always in reaction to the people around it, to its environment, to its own evolving state. That’s what makes it feel like something is emotionally present with you.

This is where AI changes the game for character design in ways that go well beyond what was possible with Astro. AI-driven adaptation doesn’t require the contextual predictions that we relied on. It learns the specific rhythms, preferences, and emotional context of the people it lives and works with. The character doesn’t just respond to context. It grows into it.

What Industry Is Missing

The character and soul of the impending wave of embodied AI products appears to almost always be an afterthought. And character defined late is character defined by default. It becomes the sum of a thousand small decisions made by different people thinking about anything but character. People project character onto devices whether you plan for it or not, especially if those devices move—a robot that moves is already a character. If nobody has designed this character, the result will be products that feel like nothing, or worse, feel confusing and not trustworthy. Technically impressive, but lifeless.

We did not get this fully right with Astro. So many things were moving in parallel that character was rarely treated as a utility, and it made sense why. When you are building a first-of-its-kind product, the things that are the loudest are the ones that break, the deadlines, the costs, the features a customer can point to on a box. Character is quieter than all of that. It’s easy to assume it can come later. On a team as large as the Amazon Astro team, it’s lucky to get any idea onto the road map when it is competing with a hundred others that all feel more urgent in the moment. None of this came from people not caring. It came from character being the kind of thing that is hard to prioritize until you see what its absence costs you.

My Asks to Product Leaders

If you are building a product that will share physical or conversational space with people, three things are worth considering:

Define character before you define interactions. You need a defensible character with enough emotional logic to answer hard questions consistently. Find answers to character questions early, and have every discipline build from the same foundation.

Build story and sound into the character pipeline, not the production pipeline. Story and sound developed alongside character definition has the chance to inform motion, expression, and interaction logic. This requires a different kind of collaboration, and a different kind of hire.

Design for adaptation, not just consistency. A consistent character is necessary, but the products that will matter most in people’s lives are the ones that deepen through use. The infrastructure to support that is more and more accessible, but the design thinking to take advantage of it is still rare.

An expanded version of this story is available on Medium.

The Secret to Marathon-Winning Humanoid Robots

17 June 2026 at 12:19


On 19 April 2026, the Honor Lightning humanoid robot ran a half-marathon in 50 minutes and 26 seconds, beating the human world record by 7 minutes and the best robot time from 2025 by almost 2 hours.

How did Honor do it? Is there some magical technology or technique that unlocked this performance? How did the company beat the significantly better-known Unitree (which reportedly had to supply its robot with an ice backpack to try and complete the race without overheating)? My doctoral thesis involved building and controlling hopping and running robots, and since then I’ve tried to design and build efficient commercial legged robots, giving me a decent idea of the constraints involved. In this article, we take a look at the fundamental underlying constraints to try and answer these questions.

The Physics of Running

Running consists of alternating phases of a leg pushing against the ground (“stance phase”) and the body flying through the air (“aerial phase”). In the aerial phase, the body falls due to gravity, losing vertical momentum. The leg in stance phase pushes against the ground to redirect the vertical momentum upward, while the other leg swings forward to reposition for the next foothold.

Electric motors use energy to produce torque—the higher the torque, the more energy is lost as heat. Adding a gear train after the motor amplifies its torque and reduces its speed. A large reduction helps with torque production, but since the rotor of the motor itself has to spin faster, it becomes very sluggish at accelerating its output. This is obviously bad for the swing phase described above. These competing effects mean that for a particular motor, there is usually a sweet spot for the gear ratio:

A graph showing the relationship between gearing and motor efficiency, with an optimal gearing ratio in the relationship between stance and swing. The power consumed by a robot leg is minimized at an optimal gear ratio (30:1 in this example).Avik De/Datawrapper

How Honor Did It

While the Lightning’s motor specifications are not published, the hip and knee motors roughly have a 110-to-150-millimeter outer diameter. For an approximate set of motor parameters, I looked to the ILM115x25 motor due to its relevant size and detailed specifications.

We can use a simple physics model to estimate the power consumption for running at 7 meters per second (the Lightning’s average half-marathon speed) as gear ratio varies:

A graph showing that optimal gearing for a robot\u2019s motor dissipates the amount of heat that the motor generates.The light blue curve shows how to pick the optimal gearing (45:1). The dark blue curve shows how much heat will be produced in the knee motor, ~150W for the optimal gearing.Avik De/Datawrapper

We see that the drivetrain is not magical: with a gear ratio chosen for this task (we’ll return to this below), the approximate robot power consumption would be a very reasonable 400 watts.

However, the dissipated knee power ( typically the main thermal limiting factor) is approximately 150 W. This is almost an unavoidable consequence—running at human speeds with a humanoid-size robot will inevitably generate this amount of heat! Over a prolonged period, keeping the motor from overheating would be a challenge, but the Lightning has a trick up its sleeve:

According to Honor, the liquid-cooling pipes penetrate deep into the motors like capillaries. The high-power liquid pump has a heat-exchange flow rate of more than 4 liters per minute. Each of the four drive motors in the lower limbs is equipped with an independent liquid-cooling circuit.

Liquid cooling is not new, but it’s definitely not a commodity. It has shown up in research periodically, and on the commercial side Apptronik tried it for a few of its prototypes but (to my knowledge) does not use it on its main Apollo platform. Basic air-convection-based cooling would not continuously be able to extract 150 W out of the knee motor, and so the cooling technology is a key enabler of this type of performance.

Why Others Couldn’t Compete

Why did Honor’s competitors, including more established and widely shipped humanoids such as from Unitree or Agibot, not compete as well?

We can use the same model to generate an equivalent energetics plot for walking at 1.5 m/s, a much more modest but potentially more common activity for a commercial humanoid robot:

A graph showing that robots with gear ratios optimized for running or walking are inefficient when walking or running respectively. The solid and dashed light blue lines show a running-optimized design, while green lines show a walking-optimized design. The optimal ratio for walking is much lower (30:1 vs. 45:1). However, the power dissipated in the knee motor while running [dark blue] is much higher at 30:1 vs. 45:1—the price to pay for running with a walking-optimized design.Avik De/Datawrapper

The plot adds a new green curve for the walking power, and the optimal gearing is significantly different!

Let’s say you design your robot to excel at the normal walking task and choose the green design with 30:1 gearing. The knee motor power to run a half marathon is over 300 W (red arrow), more than two times what we had with the running-optimized design. It wouldn’t be so surprising to need ice packs!

Conversely, visually following the green curve shows that the running-optimized robot wastes more power for walking. Using larger motors sized for running increases the weight of the robot and wastes power when it is standing or walking. The larger motors also pose practical issues like bumping into objects while operating in homes or factories.

Closing Thoughts

Honor’s half-marathon performance was an impressive engineering effort and result. It didn’t need any magical leaps in technology, but the deployment of the capillary motor cooling solution is a notable advance without which this running pace would have been unsustainable. The cooling, weight optimization, and robustness advances may well be useful for more practical purposes like carrying heavy payloads down the line.

A comparison showing two similar humanoid robots, but one has significantly smaller motors on its hips. The Honor Lighting robot [right] has much larger motors driving its legs than the Unitree H1 robot, making it a more efficient runner but a less efficient walker.Left: Wei Zhiyang/Zhejiang Daily Press Group/VCG/Getty Images; Right: VCG/Getty Images

However, the Lightning is not as well-suited to other tasks as a robot designed for greater versatility. Engineering is always characterized by trade-offs, and making the correct ones separates good products from great ones. With consistently improving AI language models, this very human skill is becoming the most valuable one an engineer can have.

The news coverage seemed to overly focus on the fact that the human half-marathon record had been broken by a robot. Machines and humans have very different capabilities and constraints, so why should we ever have expected the half-marathon time for a robot and human to be related? As in Deep Blue’s 1997 defeat of Garry Kasparov in chess, where it couldn’t physically move the pieces, the Honor robot’s capabilities are much narrower than a human running elbow to elbow with other runners while visually navigating the course without GPS. Comparing the robot runner to a human runner is just an apples-to-oranges comparison, which only risks diminishing Honor’s engineering achievement on one hand and human athletic achievement on the other.

Visual Language Models Train Robots to Read Human Emotions

13 June 2026 at 13:00


This article is part of our exclusive IEEE Journal Watch series in partnership with IEEE Xplore.

As robots advance in terms of dexterity and other physical capabilities, it becomes more likely that humans may find themselves working alongside them. If that happens, how will robots’ emotional capabilities need to advance for them to successfully work with people?

In a recent study, researchers trained collaborative robots to read human emotions by not only accounting for facial expressions, but also contextual factors in the interactions as well. Through experiments with 40 volunteers, the researchers then evaluated how a robot’s ability to read human emotions and adjust its behavior in turn impacted a human’s perception of the robot and its capabilities as the two collaborated on tasks. The results—which show that the emotional capabilities of robots only go so far with humans—were published 18 May in IEEE Robotics and Automation Letters.

Seung Chan Hong led the study as part of his undergraduate thesis while studying at Monash University, in Melbourne, Australia. He notes that, while there has been a lot of hype in the advancing physical abilities of robots, this is only one piece of the puzzle. “We need to also innovate when it comes to them actually interacting with humans, not just their physical capabilities,” he says.

This prompted him to dig deeper into the emotional aspects of human-robot interactions. First, Hong and his co-authors decided to train a robot to read human emotions using a vision language model (VLM), which is similar to large language models (LLMs) such as ChatGPT, but which can also take visual inputs.

Training VLMs for Human Emotion Recognition

To evaluate their VLM, which used Gemini 2.5, the researchers had volunteers watch videos of robots handing over objects to humans—with varying degrees of success—and describe the emotions the humans were expressing. Importantly, the volunteers labeling these videos were able to take into account more context in these interactions, rather than reporting solely on the facial expressions of the humans in the video. For example, a person pausing to think with a furrowed brow may simply be concentrating on their task at hand and not necessarily be angry. Contextual factors such as drumming their fingers, pursing their lips, or other behaviors can point to the real cause of a person’s furrowed brow.

The researchers then compared their VLM to a conventional AI system that relies on standard facial analysis and object tracking that is used in human-robot interactions. They found that the VLM outperformed the traditional approach. On a scale from 0 (no similarity in meaning to the emotion identified by the human volunteers) to 1 (a perfect match in meaning), the conventional AI system achieved a score of 0.77. In comparison, the VLM achieved a score of 0.86.

Hong says, “I think [the VLM] was able to align with what human observers were seeing a lot better, because it wasn’t just looking at the person’s face for a brief amount of time, but seeing the whole scene—where the person was and what they were doing, and how they were interacting with the robot.”

In a second experiment, the research team asked 40 volunteers to interact with a robot using their VLM—but purposefully programmed the robot to make an error. The robot then had to offer either an emotionally adaptive apology that accounted for the human’s perceived response to the mistake or a pre-scripted spoken apology.

Participants overwhelmingly preferred the emotionally adaptive response, with 31 out of 40 people favoring this approach over a boilerplate apology.

However, their survey responses underscored how this emotional adaptivity was far less important than the robot’s functionality. After collaborating with a robot that failed in its task, many participants ranked their trust in the robot as lower, regardless of how it apologized for its mistake. “A personalized apology acts as a social lubricant, but it cannot repair the trust lost by the robot failing its physical task,” Hong says.

Interestingly, the VLM classified the emotions of its human partners similarly to human volunteers who observed an interaction from a third-party perspective. But when the VLM’s assessments were measured against humans’ self-reported emotions during the second experiment—the most accurate descriptions of their true emotions—its ability to accurately predict emotions dropped significantly.

“While the VLM is a good observer of outward social cues, it isn’t a mind reader,” Hong says. “It matched third-person human observers well, but it didn’t always align with the users‘ internal, self-reported feelings.”

Together, these results show that robots are not perfect at reading human emotions. So while people might appreciate their efforts, they still ultimately will want competent co-workers.

This story was updated on 15 June 2026 to correct where the research was conducted and clarify that the researchers evaluated the performance of a pre-trained model.

Award-Winning Researcher Trains Robots to Make Educated Guesses

12 June 2026 at 18:00


Yen-Ling Kuo always wanted to understand how things worked. When she was growing up in Taiwan, reading the story of Michael Faraday in elementary school piqued her curiosity about the natural world. During that time, she was introduced to Logo, a computer program with a turtle cursor to help children learn basic coding through hands-on experimentation.

It was Kuo’s introduction to programming logic.

Yen-Ling Kuo


Employer

University of Virginia in Charlottesville

Title

Assistant professor of computer science

Member grade

Member

Alma maters

National Taiwan University; MIT

In high school she learned the capacity computers held. She could write programs that completed tasks independently, she realized.

“Once I discovered how powerful computers could be,” she says, “I knew I wanted to focus on using them to solve real-world problems.”

Kuo, an IEEE member, never lost her interest in the “how” behind processes and tools. Her curiosity, combined with a stint working at a Silicon Valley company, led her to focus on innovations that live at the intersection of cognitive and computer sciences.

Kuo, now an assistant professor of computer science at the University of Virginia in Charlottesville, last year received the IEEE Robotics and Automation Society’s inaugural Outstanding Women in Robotics and Automation Early Career Contribution Award. The award is part of the IEEE-RAS Women in Engineering’s Outstanding Women in Robotics and Automation (WiRA) Paper Awards, which promote excellence and recognize the impact that female researchers have on robotics and automation fields at different stages in their academic careers.

Kuo’s winning paper, “Diff-DAgger: Uncertainty Estimation with Diffusion Policy for Robotic Manipulation,” demonstrates a novel method to help robots better identify and estimate uncertainty when faced with scenarios on which they’ve not been trained. The method reduces the amount of human supervision, improves a robot’s rate of successful task completion, and opens up a path to introduce more complex models with bigger data demands into interactive robot learning.

She says her research will help people working in the robotics and automation fields more efficiently collect the data needed for effective model training.

Silicon Valley’s impact

Kuo earned bachelor’s and master’s degrees in computer science at the National Taiwan University, in Taipei, in 2009 and 2012. As she was nearing completion of her master’s degree, she did what many computer science graduates do: She pursued a summer internship at a tech company.

She spent the summer of 2011 at Google’s campus in Kirkland, Wash., working on the company’s comparison ads project.

When her internship ended, she joined the MIT Media Lab as a visiting student, working on the Open Mind Common Sense project with Henry Lieberman.

As she was considering pursuing a Ph.D., a call from Google changed her plans. The company offered her a full-time role as a software engineer.

“I viewed the job offer as a positive development,” she says. “I believe it can never hurt your future research career to get some real-world experience under your belt.”

She was hired in 2012 and helped build techniques that incorporate computer vision and natural language processing to improve the customer shopping search experience. She led the company’s Shop the Look initiative, a predecessor to Google’s current AI-powered shopping experience. The project connected social media content with search results, something the company had struggled to do in the past.

Kuo and her team were tasked with building a connection between the natural language people use to describe an item and an image that matches the searcher’s intent. It was at a time when the neural network—using deep learning models to power Google products—was gaining momentum at the company. Integrating neural network tools into her work was a requirement—which raised questions for Kuo.

“I was applying the neural network tools,” she says. “But I didn’t have 100 percent certainty about how they actually worked.”

She considered how she could become more knowledgeable about deep learning models. It was a full-circle moment. She decided that after nearly four years at Google, it was time to earn a Ph.D. in computer science. She returned to MIT in 2016.

The question that changed everything

Boris Katz, one of Kuo’s Ph.D. advisors, is a principal research scientist and the head of the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL)’s InfoLab. He also led the creation of the START Natural Language System, the world’s first Web-based question-answering system.

When the two met, Katz asked Kuo why she wanted to pursue a doctorate degree. She explained her interest in understanding how neural networks work and in using that knowledge to connect the physical world with human language.

He suggested she attend a summer course at MIT’s Center for Brains, Minds, and Machines, a research initiative that ran from 2013 through 2025. CBMM’s objective was to bring together computer scientists, cognitive scientists, and neuroscientists to understand how human intelligence works. The goal was to use the resulting insights to establish an engineering practice to build artificial intelligence systems.

For Kuo, it was a chance to better understand human intelligence and identify ways it could be replicated in machines.

“It was an opportunity for me to interact with other scientists and gain insight into how people learn, understand, and figure things out in the world,” she says. “I saw it as a very useful and inspiring way to incorporate those ideas into my own research work.”

During her Ph.D. studies, she was a research assistant at CSAIL. The experience helped shape her doctoral research, which focused on building AI systems that apply past learning to new situations. She developed machine learning models to support the efforts, including language understanding and social interactions.

She completed her Ph.D. in computer science in 2022 with a minor in cognitive science.

After graduation, she continued her work and collaboration at CSAIL, particularly on projects that involved the “theory of mind” concept.

Theory of mind spurs innovation

Theory of mind isn’t new, having originated with primatologists studying chimpanzees in the late 1970s. The theory recognizes that others have their own thoughts, beliefs, and perspectives. It’s a skill that allows humans to infer someone’s mental state and predict their behavior without verbal communication.

“It’s like when college roommates are moving into their dorm. They may not talk too much, but they work together naturally to coordinate their activities and accomplish goals,” Kuo says. “They can infer and mentally interpret each other’s behaviors and signals to make decisions and complete tasks without words.”

She brought her theory of mind research to the University of Virginia when she joined as an assistant professor in 2023.

Kuo conducts her research in UVA Engineering’s multidisciplinary cyberphysical Link Lab. Her broad focus is on developing computational models that help robots interpret both direct data and silent signals, from language and movements to a person’s gaze. If successful, it could give robots the same sort of physical and theory of mind reasoning capabilities that power physical and social interactions among humans.

“There are no computational frameworks yet available that will translate this kind of understanding into a robot efficiently,” she says.

She adds that the process to get there begins with improving how robots learn to perform tasks.

The evolution of robot learning

Historically, one way robots learned was to mimic humans. A researcher would manually guide a robot through a task, like cutting an apple, and it would repeat the movements. The robot was successful until the environment changed, such as when its hand was in a different position or the apple was at a different angle. The robot was then faced with a situation for which it hadn’t been trained. Without any data available to help it correct course, the robot would start making small errors that eventually led to a full system crash.

Diagram of a robotic gripper delicately holding a potato chip. Labels describe how the gripper\u2019s visual perception and tactile sensing prevent the chip from breaking. This diagram describes how the robotic gripper’s visual perception and tactile sensing prevents a potato chip from breaking.Xuhui Kang, Yen-Ling Kuo, et al.

To solve the problem, researchers developed the dataset aggregation (DAgger) method. As a robot performed a task, a researcher was on standby to provide real-time corrections during unexpected scenarios. The correction data was continuously added to the robot’s model, teaching it how to recover from mistakes.

To reduce the human monitoring effort, robot-gated DAgger was created to enable bots to query humans when the machines became uncertain.

The most popular approach to make the query decision is to train multiple models to consider when determining a course of action. If the models all agree, the robot proceeds. If they don’t agree, the robot is likely to get stuck and ask for help.

Although the multiple model approach was widely adopted, it has limitations. Practically speaking, as models become more complex, it is hard or impossible to train multiple copies. A more fundamental issue is that disagreement among models doesn’t always imply uncertainty; it could just mean there are different ways to accomplish a task.

The Diff-DAgger solution

That is the gap Kuo’s research team closed with the novel Diff-DAgger research. The approach builds on diffusion policy, a technique that helps robots account for different ways a task can be performed.

The new method repurposes diffusion loss, the signal a robot uses to improve its model during training, as a real-time confidence check. During task execution, the robot computes the signal and compares it against values from its training data using a statistical test. The signal spikes when the robot faces an unfamiliar situation and is uncertain how to proceed. The signal stays silent when the robot’s current action is close to what it learned before.

The spike represents the robot’s ability to self-diagnose and predict an imminent failure. Human intervention is triggered only when the signal spikes. No spike means the robot can be left to complete its decision-making process on its own.

Kuo’s team achieved significant results: Failure prediction rates were improved by 39 percent. Task completion rates were increased by 20 percent, and tasks were completed nearly eight times faster.

Her research at UVA gained attention from the National Science Foundation, which honored her last year with a Career Award, the foundation’s flagship grant for early-career researchers. The five-year US $665,000 grant supports her research that builds computational models for human-robot interactions through theory of mind reasoning.

She also received the Toyota Research Institute’s Young Faculty Researcher Award to teach cars to reason about interactions on the road and with the driver.

As service robots and self-driving vehicles become more available, such works are likely to make interactions between humans and robots more intuitive and useful.

Kuo ultimately wants to build more robust robots that are able to integrate into a social space with humans by engaging with us through grounded interactions, she says.

The impact of IEEE

Like many IEEE members, Kuo was introduced to the organization as a student. In 2018 she submitted her first paper, “Deep Sequential Models for Sampling-Based Planning,” to the IEEE/Robotics Society of Japan International Conference on Intelligent Robots and Systems while pursuing her Ph.D. at MIT. Her IEEE involvement grew alongside her professional career.

“It was a natural segue to transition from student to a full IEEE member,” she says. Today she is an active volunteer with the IEEE Robotics and Automation Society, a reviewer for submitted papers, and a presenter and panelist at conferences.

She says one of the best parts of attending conferences is having the opportunity to engage with students. She also enjoys participating as a panelist at luncheons, she says, because it gives her one-on-one time with student attendees. She can share her knowledge and offer insights as they prepare to embark on their career.

Her goal in the coming years, she says, is to broaden her involvement with IEEE initiatives and branch out to other technical committees. Sharing knowledge and learning from others is essential to anyone’s career growth, she says, and “IEEE offers a great opportunity for both.”

Video Friday: Robotic Motion Discovery Reveals Unusual Behaviors

12 June 2026 at 17:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO

Enjoy today’s videos!

We present MotionDisco, a framework that discovers contact-rich, long-horizon humanoid loco-manipulation motions from scratch, without relying on teleoperation or motion retargeting from human demonstrations.

Some of the discovered behaviors are a little nutso:

[ MotionDisco ]

Not sure I’d say any of this is ‘effortless’ but those claws are pretty cute.

[ Deep Robotics ]

It turns out running a workout class is a decent way to stress-test whole-body range of motion. Coordinating fluid movement across every joint at once—timing, velocity, balance compensation—is one of the harder control problems in humanoid robotics.

[ Agility ]

Our very own Gwendolyn Rak made a robotic shoulder-friend at Computer Human Interaction in Barcelona.

Here’s a bit more about it:

[ MIT ]

At AIRoA, we’re bringing robots into real homes. Check out our exclusive first video to see how they work in our development hub and real-life household settings! The project aims to develop home robots that can assist people with everyday tasks and become long-term companions in daily life. In this video, we demonstrate Toyota’s Human Support Robot (HSR) deployed in real homes, where it assists residents with everyday tasks such as tidying rooms and fetching objects.

[ AI Robot Association ]

Thanks, Naoaki!

MIDAS Hand is a fully open-source, tactile-sensor-integrated dexterous robotic hand platform for manipulation, teleoperation, and robot learning research. MIDAS stands for Modular low-Impedance Direct-drive Anthropomorphic Sensing Hand.

[ MIDAS Hand ]

Thanks, Jun Kim!

This video presents a novel flight maneuver for a flying bipedal robot. During forward flight, the robot performs aerial braking by swinging its legs to adjust the orientation of foot-mounted thrusters.

[ Paper ]

Seems like a really good application for autonomy, to be honest.

[ Built Robotics ]

In this time-lapse, controllers on the ground are repositioning Dextre, our robotic handyman currently installed at the end of the Canadarm2. They used Dextre to unload equipment from the unpressurised Dragon trunk. Such a beautiful choreography to watch with Earth in the background!

[ European Space Agency ]

This video demonstrates how AI Sapiens learns and performs humanoid motions from video-based motion capture using only a smartphone camera, without professional motion capture equipment. ROBOTIS plans to release an open-source motion generation and learning pipeline for AI Sapiens, enabling users to generate humanoid motions from video and bring them to the real robot.

[ ROBOTIS ]

NAO LIVES!

[ Maxtronics ]

Tumblenauts are a swarm of minimalist, bacteria-inspired robots designed for collaborative inspection of pressurized microgravity habitats such as the International Space Station. Unlike current intra-vehicular robots that rely on complex actuator-dense mechanisms for precise motion, the Tumblenauts use a stochastic run-and-tumble locomotion and collective cooperation inspired by bacterial colonies.

[ Self-Organizing Swarms and Robotics Lab ]

LUMOS Robotics Founder and CEO Yu Chao officially introduces Project EDGE—inviting global builders, universities, robotics labs, and creative technologists to explore the future of humanoid robotics together. To supercharge the global developer community, we are providing 100 complimentary LUMOS NIX robots to selected global partners.

[ Lumos Robotics ]

How do you progress from early childhood computational thinking to advanced high school robotics? Sphero’s product offerings are intentionally scaffolded to scale for students by building critical skills and concepts at every grade level.

[ Sphero ]

Defining Autonomy for Wellness Robots in Senior Care

11 June 2026 at 10:00


An examination of how socially assistive wellness robots could support the seven dimensions of senior wellness, and how a framework can measure their autonomy.

What Attendees will Learn

  1. Why the senior care crisis exceeds incremental automation. Demographic pressure, workforce shortages, and a daily wellness-programming gap all strain traditional care models.
  2. What defines a wellness robot as a category. The seven ICAA wellness dimensions and eight properties separate these robots from companion and medical devices.
  3. How autonomy can be measured with CRAS. This six-level scale, modeled on the SAEJ3016 driving standard, evaluates four care dimensions.
  4. What maps the road to full autonomy. The paper examines technical capabilities, clinical evidence, and a three-phase roadmap toward the early 2030s.

Beyond Dexterity: Why Contact May Define the Next Era of Robotics

9 June 2026 at 12:51


This article is brought to you by AGILINK.

Throughout the exhibition hall at the 2026 IEEE International Conference on Robotics (ICRA), in Vienna, one demonstration seemed to attract a disproportionate amount of attention.

Two robotic hands were making a balloon dog. Slowly and deliberately, the robot twisted a long balloon into loops, bends, and joints without popping it. Visitors stopped, watched, and often returned with colleagues to watch again.

Crowd at a robotics expo watches a humanoid robot demonstrate its arm movements. AGILINK’s balloon dog demonstration draws a crowd at ICRA 2026.AGILINK

At first glance, the demonstration appeared almost playful. Among roboticists, however, balloon twisting is widely recognized as an unusually difficult manipulation task.

A balloon is lightweight, highly deformable, slippery, and extremely sensitive to force. Every twist changes its geometry and internal pressure, turning a seemingly simple activity into a continuously changing physical interaction problem.

Humans navigate those changes almost intuitively. While making a balloon animal, people rarely think consciously about force regulation, slip prevention, or contact stability. They simply adjust.

For robots, those adjustments remain remarkably difficult. The challenge is not merely moving fingers to the right positions. The harder part is maintaining stable interaction while the object itself is changing.

Highlights from AGILINK’s ICRA 2026 demonstrations, including visuotactile sensing, in-hand manipulation, balloon-animal shaping, and other contact-rich tasks enabled by the company’s latest OmniHand platform.AGILINK

That distinction helps explain why the balloon dog drew so much attention in Vienna. What appeared to be a dexterity demonstration was, in many ways, a demonstration about contact itself.

As robotic manipulation continues to advance, a growing number of researchers are arriving at a similar conclusion: many of the hardest problems in robotics begin only after contact occurs.

Motion and Contact Intelligence for Robot Manipulation

Balloon twisting combines two challenges that robotics has traditionally struggled to solve simultaneously: long-horizon task execution and contact-rich manipulation.

The first concerns motion.

A balloon dog is not created through a single grasp or twist. It emerges through a carefully ordered sequence of manipulations, each setting the conditions for what follows. A small rotational error introduced early may appear insignificant at first, yet several steps later it can prevent the final structure from forming altogether.

In that sense, balloon twisting is a long-horizon task. Success depends not only on performing individual actions correctly, but also on preserving the future feasibility of the entire manipulation process.

To address this challenge, AGILINK began by collecting demonstrations from professional balloon artists. Human actions were mapped onto robotic hands to establish an initial manipulation policy. But successful demonstrations alone were insufficient.

In practice, some of the most valuable learning occurred when execution began to drift toward failure. Whenever instability emerged, human operators intervened and corrected the manipulation in real time. Those interventions were recorded and incorporated into reinforcement-learning cycles, allowing the system to learn not only how successful demonstrations unfold, but also how experienced operators recover when things start to go wrong.

Through this process, the robot gradually acquired the capabilities required for long-horizon task execution—a collection of abilities that AGILINK groups under the term motion intelligence: the ability to generate actions, coordinate bimanual behaviors, and execute extended manipulation sequences under real-world uncertainty.

Two robotic hands, one white open palm and one black forming an OK gesture, on display. OmniHand 3 Ultra-M on display at ICRA 2026.AGILINK

Yet motion alone does not explain why balloon twisting remains difficult. The second challenge is contact.

The robot must continuously regulate force, adjust contact locations, and respond to subtle changes in the object’s state. These decisions are difficult to encode through explicit rules. Even skilled human operators often rely on tactile intuition developed through experience rather than consciously articulated strategies.

Analysis of those interventions revealed that many failures did not originate from incorrect action sequences, but from the breakdown of contact itself.

To better capture those interaction dynamics, AGILINK collected contact-centric intervention data and incorporated those interactions into reinforcement-learning training. Rather than learning only which motions to perform, the system also learned how humans maintain stability when contact conditions begin to deteriorate.

AGILINK describes this capability as contact intelligence: the ability to establish, maintain, and adapt physical interaction as force distribution, friction, deformation, and contact geometry continuously evolve.

The distinction between the two capabilities is subtle but important. Motion intelligence determines what the robot intends to do. Contact intelligence determines whether it can continue doing it. For balloon twisting, both are necessary. One provides the sequence of actions. The other keeps those actions physically viable.

Robot makes balloon animal for visitor at tech expo booth. YouTuber KhanFlicks follows OmniHand’s motions while learning to fold a balloon dog at the AGILINK booth.AGILINK

Between a balloon slipping away and a balloon bursting lies a narrow region of stability. Successful manipulation depends on finding that region—and remaining within it throughout the task.

Introducing the OmniHand 3 Ultra-M Dexterous Hand

The balloon dog demonstration showcased a manipulation capability. It also revealed a broader question. How much contact intelligence can be achieved through learning alone? A robot can only regulate what it can perceive. It can only respond as quickly as its hardware allows.

As manipulation tasks become increasingly complex, researchers are finding that progress depends not only on better policies, but also on richer sensing and faster physical response.

That realization formed the backdrop for AGILINK’s second major announcement at ICRA 2026. Alongside the balloon dog demonstration, the company introduced the OmniHand 3 Ultra-M.

Two robotic hands beside a human hand, all raised open on a display table. OmniHand 3 Ultra-M closely matches the size of an adult human hand.AGILINK

The two exhibits represented different stages of the same technological trajectory. If the balloon dog demonstrated what contact intelligence can already accomplish today, Ultra-M was designed to explore what contact intelligence may require next.

Building Hardware for Contact Intelligence

Roughly the size of an adult human hand, the OmniHand 3 Ultra-M integrates 20 active degrees of freedom within a human-scale form factor.

Its most distinctive feature is a fully direct-drive architecture. By adopting direct-drive actuation throughout the system, the hand is designed to enable faster and more transparent force regulation and higher force-control bandwidth, enabling faster response as contact conditions change. For contact-rich manipulation, responsiveness can be as important as sensing itself.

By adopting direct-drive actuation throughout the system, the OmniHand 3 Ultra-M is designed to enable faster and more transparent force regulation and higher force-control bandwidth, enabling faster response as contact conditions change.

The platform also incorporates tactile sensing across nearly the entire hand. Each fingertip contains a miniature vision-based tactile sensor, while more than 300 three-dimensional tactile sensing points are distributed throughout the palm. Together, they provide information not only about where contact occurs, but how contact is evolving.

The system is designed to estimate pressure distribution, shear forces, local deformation, slip tendencies, and other interaction dynamics that often remain invisible to conventional position-based control systems.

According to AGILINK’s tests, individual sensors achieve force resolution of approximately 0.005 N—roughly equivalent to detecting the weight of a sheet of paper resting on a fingertip. Spatial resolution reaches approximately 0.04 mm, while sensing density approaches 50,000 sensing points per square centimeter.

Robot arm delicately holds a feather, inset shows colorful dotted texture close-up. OmniHand 3 Ultra-M recognizes feather texture through vision-based tactile sensing.AGILINK

For dexterous robots, contact has traditionally been a largely hidden process. Ultra-M is designed to make that process more observable.

Rather than simply detecting that contact has occurred, the system attempts to resolve where interaction is happening, how forces are distributed, whether instability is beginning to emerge, and how manipulation strategies should adapt in response.

The balloon dog offered a glimpse of what contact intelligence can already accomplish. Ultra-M explores a different question: what capabilities may be required to push contact intelligence further?

The Physical World Remains the Hardest Benchmark

The significance of contact intelligence extends far beyond balloon animals. Many tasks that continue to resist automation involve unstable or deformable interaction: cable insertion, garment handling, flexible packaging, delicate assembly, connector mating, tool use, and household manipulation.

These tasks are difficult not because robots cannot reach the correct location, but because maintaining stable interaction after contact begins remains extraordinarily hard.

For decades, robotics achieved many of its successes by reducing uncertainty. Factories were engineered to make robotic motion predictable, repeatable, and highly structured. The physical world behaves differently.

A growing share of robotics research is shifting toward interaction itself—understanding how robots can establish, maintain, and adapt physical contact within environments that remain fundamentally unpredictable.

Objects shift. Materials deform. Friction changes. Contact evolves. Real environments rarely follow scripts. Seen through that lens, the balloon dog was never really about the balloon dog. What attracted attention at ICRA was not simply a visually impressive demonstration, but what it revealed: intelligence in the physical world is ultimately measured through interaction.

As motion generation continues to mature, a growing share of robotics research is shifting toward interaction itself—understanding how robots can establish, maintain, and adapt physical contact within environments that remain fundamentally unpredictable.

For robots moving beyond structured environments and into less predictable real-world settings, managing contact may become as important as motion itself.

How JPL Keeps the 13-Year-Old Curiosity Rover Doing Science

9 June 2026 at 12:00


Thirteen years ago last August, I was camped out in NASA’s Jet Propulsion Laboratory press room in Pasadena, Calif., waiting to see whether the Curiosity rover would survive its descent and skycrane-assisted landing on the surface of Mars. It did, and it was awesome.

Since then, Curiosity (also known as Mars Science Laboratory) has traveled nearly 37 kilometers, drilled into and sampled 42 different rocks, and as of publication has snapped nearly 763,000 photos. The fact that this robot is still hard at work, getting real science done at the age of 13, is absolutely incredible—not only is Mars an actively hostile environment for robots, but the only kind of maintenance that JPL engineers can do is to send very, very careful software updates.

Nevertheless, the clever folks at JPL have managed to keep Curiosity safe, warm, mobile, and sciencing, despite well-worn wheels and less and less power every day. One of those folks is Alexandra Holloway, the assistant team chief for engineering operations for Curiosity, who spoke to IEEE Spectrum about keeping Curiosity roving, what its future looks like, and how JPL has used that experience to make rovers like Perseverance even more capable.

How astonished should we be that after 13 years on Mars, Curiosity is not only still doing science, but actually getting more capable?

A woman with large green eyes and a shaved head Alexandra Holloway is the assistant team chief for engineering operations on the Curiosity Mars rover at the Jet Propulsion Laboratory.Alexandra Holloway

Alexandra Holloway: I’m astonished! The longevity comes from a lot of ongoing work. It’s not just that Curiosity was built robustly; it’s also because we’re continuously putting in effort to ensure it can continue to have that lifespan. I think about all the different kinds of embedded systems there are, from cars to refrigerators, and none of them have the kind of longevity that we have with the rover. It’s mind-boggling, and it’s inspiring.

Is the Perseverance rover, which is nine years younger than Curiosity, significantly different in terms of its hardware and software?

Holloway: In terms of hardware, the rovers are actually very similar. Both use a RAD 750 processor and have the same amount of memory. However, Perseverance has an extra processor specifically for visual odometry, which allows it to drive autonomously. This difference reflects their primary mission designs: Perseverance was designed for driving long distances, while Curiosity is a mission focused on sampling as it goes. So Perseverance’s onboard scheduling capabilities are there to optimize its driving. In fact, just last year, Perseverance surpassed Curiosity’s driving distance after only about three years on Mars.

Curiosity Rover Memory and Software Fixes

Do you have some examples of significant tweaks the team has made to keep Curiosity roving?

Holloway: One of my favorite examples comes from a processor anomaly that happened on Sol 2172 [Ed. note: “Sol” is the term for a Martian day—about 24 hours and 40 minutes]. Curiosity has two computers, A and B. We landed on A, swapped to B due to a NAND memory anomaly early on (Sol 200). For years, we were chugging along on B, until one day there was a problem—B booted up, but it couldn’t mount its drive partition. We’d never seen this before. To preserve B’s data, we swapped back to A, which we hadn’t trusted in two thousand Sols. A also had a degraded memory, with only two gigabytes of usable storage space instead of four. We painstakingly transferred data from B over to A and then down to Earth, and eventually we ran out of stuff we wanted to transfer, which was really good, because A then started acting funny in the same way it did on Sol 200. It was acting like its memory was coming unsoldered. That’s bad.

We quickly swapped back to B, formatted it, and got it working again. The problem then became that we couldn’t trust A’s memory at all, but we needed a second computer as a “lifeboat” for diagnostics and transfers if B failed again. We realized we had one other place of memory: where we keep our flight software. We have four copies of the flight software (two current versions and two older versions) in different banks of very small amounts of memory, just 32 megabytes each. What if we just jettisoned the old flight software copies and used that 64-megabyte NOR memory as our file system for computer A?

So that’s what we did. It was so elegant! Computer A is operating with less than 1 percent of its original memory, but we can run a mission on it. A small mission, but we haven’t had to jettison any core capabilities. We can still drive, we can manage data, we can even theoretically do science. Everything works fine, just much slower and much smaller. That flight software release was even called “R-Hope“ because we hoped it would work.

What are the constraints on Curiosity’s lifespan?

Holloway: Our biggest hardware challenge is wheel wear. It looks like we’re driving on this sandy terrain with some rocks in it, and our intuition said that we could just drive over these rocks and they’d get pushed down into the sand and it would be no big deal. But what we ended up seeing was that those little rocks are actually the tips of giant boulders buried in the sand, and they’re razor sharp. Our wheels were getting ripped apart driving over them, especially our front wheels, so we started driving backwards.

We also monitor consumables. We consider the number of times we move our actuators. That’s a consumable. Curiosity hasn’t taken a selfie in a while, and one of the reasons is that it’s really hard on the joint actuators. Our onboard memory is a consumable, but surprisingly we’re not anywhere near our life cycle for memory. Our biggest consumable is power; we have an RTG, a nuclear power source, which decreases its output as it ages.

Newer missions are flying Snapdragon [processors], but Curiosity’s RAD 750 is a power hog. One of the things that we’ve rolled out that’s going really well is a way of reducing the amount of time we spend with the computer powered on, by harvesting time when we finish activities early and going to sleep, which lets us turn off the computers and some of the heating. Another thing we’re looking at is doing stuff in parallel when we’re on, like being able to drive or use the arm while communicating with an orbiter.

So power is decreasing, and that’s causing us to do all this parallelism work and become more efficient and nuanced in the way we operate. But we are not having any degraded science output at this time. Our wheels are still going, our arm is still okay for now, knock on wood. I would say maybe the bottleneck is budget.

Curiosity Rover’s Impact on Future Mars Exploration

What have you learned from Curiosity that will improve future missions?

Holloway: As an embedded flight software person, I think about how we can change, add, or modify software capabilities during the mission. There’s definitely a sweet spot for loading and patching flight software—some of these concepts were pioneered on Spirit and Opportunity and then inherited by Curiosity and Perseverance, making it easier to understand and change the software.

Some of the things that I wish we had now on [the Mars Science Laboratory] include a better understanding of where our power is going. I want to see how much power each component is drawing every minute, so that we could architect a software system that could balance loads better. We have some of this information that was built in by the engineers who designed the rover, but as an operator, I want something slightly different. So if I were building a mission, I would have those discussions earlier and get operators into the room to say, “what do you want your data products to look like?”

The key takeaway for designing future missions is to talk to all your users early in the design process. It needs to happen upfront.

What does Curiosity’s long-term future look like?

Holloway: That’s a conversation that happens, and it’s a really delicate one. We have a lot of science instruments, and a lot of them have to do with contact science and sampling and rely on the arm. If we lose the arm, what science can we still do? Well, we have a lot of remote sensors too, like cameras, environmental sensors, and radiation sensors. All of these things are important for the future of space exploration and humans on Mars.

From a power perspective, our RTG is projected to start degrading science output in the sixth extended mission, but we’re going to be fine through 2035 and potentially even beyond that. So we have a long and exciting future ahead of us. We need to figure out the best way of operating within our constraints, but we’re still kicking.

Video Friday: Watch This Running Robot Not Fall Down Stairs

5 June 2026 at 15:30


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO

Enjoy today’s videos!

It’s been a while since a humanoid robot video actually impressed me, but the beginning of this does.

Hard to know how much of that recovery was luck, though.

[ Deep Robotics ]

When you’re very confident in your MPC-based balance controller...

I feel you, buddy. And thanks for posting this. We’ve all been there, in one way or another.

[ DARoS Lab ]

GENE01 designed from scratch and sent to batch production. Two scalable lower bodies. Physical AI deployed for motor control and world-action modeling. All in three months. Generative Bionics is running.

[ Generative Bionics ]

Alex, the newest humanoid robot built entirely by ‪IHMCRobotics‬, takes its first steps outdoors! This was a significant milestone for our team, especially because Alex is the first humanoid robot developed entirely by IHMC Robotics to venture outside the lab. These outdoor trials were conducted in preparation for a demonstration in Maryland, where Alex later successfully walked completely untethered.

[ IHMC ]

Built on the Enlight platform, Flexiv Mico is a compact dual-arm system engineered for safe, seamless collaboration in any workspace.

[ Flexiv Robotics ]

Is it weird that I’m jealous that robots can have feet that are swappable?

[ Boston Dynamics ]

Midweek at ICRA 2026. Cable-climbing robots that work as a squad: CCRobot-S is a team of robots with reconfigurable cable-driven manipulation that collaboratively inspect and maintain long-span bridge stay cables. Parallel operation for speed, morphological reconfiguration for reach.

[ IEEE Transactions on Robotics ]

I would love to know the story behind this odd choice of hat.

[ Robotis ]

How did Atlas learn football—and why? Go behind the scenes of School of Football and discover a glimpse of the Next of robotics.

What I really want to know is, what kinds of things is it possible to do in football (soccer) when your robot’s joints are not constrained by biology?

[ Boston Dynamics ]

This DIY Bipedal Robot Used Pneumatic “Air-Muscles” Instead of Motors

31 May 2026 at 13:00


In 1987, Richard Greenhill, a British photographer who was fascinated by (but had no actual training in) robotics, decided he wanted to build a life-size humanoid that could do useful things, like carrying luggage. He was working at a startup called Intergalactic Robots, but he couldn’t convince anyone there to build such a machine, so he set about building one himself, in his attic.

To help with his project, he organized a weekly get-together of a dozen or so like-minded folks. Every Wednesday night, his wife, Sally, would make a big pot of spaghetti, and the group would tinker with components scavenged from old printers and picked up from junkyards. They called themselves the Shadow Group. They eventually constructed several different robots, but their main project was the two-legged Shadow Walker.

Two color photos of a casually dressed white man in a workroom posing with a partially assembled wooden robot. In 1987, photographer Richard Greenhill organized a weekly gathering of DIY enthusiasts to work on projects in his attic, including the Shadow Walker. Richard Greenhill and David Buckley

Greenhill’s friend David Buckley, a robotics and animatronics expert he’d met at Intergalactic, sketched out a rough design based on medical textbooks of human bone structure and muscle movement. The robot’s skeleton, made of maple, was greatly simplified—only one bone in the lower leg and a single wide toe on each foot. The ankle’s double-axis design allowed for two degrees of movement. The knee had no complicating kneecap.

Greenhill didn’t want the robot to use motors, so its movement was controlled using compressed air to extend and contract 28 “air-muscles”—his version of a McKibben muscle, invented in the 1950s to mimic musculature with pneumatics. The muscles were connected to the bones across eight joints (hips, knees, ankles, toes), which provided 12 degrees of freedom.

RELATED: The Short, Strange Life of the First Friendly Robot

The robot’s headless torso held the control valves, electronics, and computer interfaces. It stood 168 centimeters tall and 46 cm wide and weighed about 38 kilograms. The group managed to get the robot to stand up reliably and balance itself; it could even regain its center if pushed a little. But walking turned out to be more of a challenge.

Rich Walker joined the group as a teenager and began writing software to get the robot to stand. He was particularly interested in using neural networks to solve balancing problems, although he ran into a number of hardware obstacles, including the unreliability of the sensors and the valves, and the robot’s overall fragility. Over time, Walker and the team developed a standard library of routines to control the robot. Walker wrote a detailed description of the Shadow Walker in 1999, which is available on David Buckley’s website.

The 1st International Robot Olympics

By the time the Shadow Group began developing Shadow Walker, engineers in academia and industry had been working on robotics for several decades. The world’s first industrial robot, the Unimate, debuted in 1961, and in 1967 Donald Michie and others began building a series of Freddy robots to investigate machine intelligence. The IEEE created its first dedicated robotics organization in 1984 when it established the IEEE Robotics and Automation Council, which became the IEEE Robotics and Automation Society in 1987. Also in 1987, the nonprofit International Federation of Robotics was established to promote research, development, use, and cooperation in the field of robotics.

As Shadow Walker pushed the limits for a DIY humanoid robot, industrial humanoids were also gaining ground. In 1986, Honda began working on its experimental (E-series) and later the prototype (P-series) humanoid robots, finally unveiling the P2 in 1996. The P2 stood 183 cm tall and weighed 210 kg. It was the first humanoid capable of stable, autonomous walking. This work eventually led to the development of the groundbreaking ASIMO.

Two color photos of a casually dressed bearded white man posing with a wooden robot leg and with a computer and other equipment. Greenhill’s friend, roboticist David Buckley, consulted medical textbooks to create Shadow Walker’s humanoid design.Richard Greenhill and David Buckley

In the late 1980s, the public was both fascinated and horrified by the potential of robots. Businesses saw robots as a way to increase productivity, while workers worried they would take their jobs. Children viewed them as wondrous toys, while people with disabilities embraced them as tools of liberation. Military experts hoped robots would fight wars without endangering human soldiers, while politicians pondered if robots might eventually get to vote. Philosophers thought robots could challenge our notions of intelligence (and stupidity), while the religious struggled with concerns about the human race in a robot-dominated future.

Photo of two articulated feet made of pieces of wood strung with wires and other components. Shadow Walker’s simplified anatomy included only one bone in the lower leg and a single wide toe on each foot.Science Museum Group

Peter Mowforth, cofounder of the Turing Institute in Glasgow, noted these disparate visions for robots when he announced the 1st International Robot Olympics, to be held in 27 and 28 September 1990 and hosted by the Turing Institute and the University of Strathclyde. The Olympics would round up the world’s best robots and showcase them head-to-head.

Mowforth himself thought all of the competing visions of robots were overblown. Steeped in machine learning research and robotics development, he knew firsthand the limitations of the state of the art: Robots rarely worked as intended, easily broke down, and glitched over seemingly trivial problems. He envisioned the Robot Olympics as a testbed to assess what the latest generation of robots could and could not do.

Photo of a headless and armless humanoid robot wearing red pants. At the 1990 Robot Olympics, held in Glasgow, Shadow Walker wore pants to conceal its pneumatic “air-muscles” from competitors.Adam Hart-Davis/Science Source

The call for participation was wide open. Instead of having predetermined categories of competition, the organizers opted to see who applied to compete and then group them based on their claimed capabilities. In addition to picking the winners of individual events, the judges would select an overall Olympic champion based on the quality of the hardware, the sophistication of behavior, and novelty. Other prizes were given for young competitors, technologies that showed commercial potential, and design. In the end, more than 50 robots were entered, from a mix of universities, industry, and hobbyist groups from Canada, France, India, Japan, Mexico, the Soviet Union, the United States, the United Kingdom, and Yugoslavia.

There were plenty of disappointments. Trolleyman, a golf-cart-like wheeled robot, suffered a power failure while carrying the opening Olympic torch through the streets of Glasgow. The pile rug in the arena tripped up many robots that had been trained only on flat, smooth floors. David Buckley later concluded that the events were too difficult, and that the Olympics didn’t push development forward.

Of course, there were winners. In a surprise triumph for vintage technology, the fully mechanical 19th-century Japanese Archer from the Museum of Automata in York, England, won gold in javelin, beating out competitors more than 100 years its junior. The overall Olympic Champion was Yamabico, Shoji Suzuki’s entry from the University of Tsukuba, in Japan, which won bronze in obstacle avoidance and gold in wall following, but was disqualified in the talking category for not speaking English.

The Shadow Group had high hopes for Shadow Walker. Unfortunately, though, it failed to take a step, and the biped race was won by the Cardiff University Biped. Shadow Walker now resides in the collections of the Science Museum in London.

The Legacy of Shadow Walker

In 1997, a paying customer in search of a robotic leg compelled the Shadow Group to get serious and become a registered company. Shadow Robot is now Britain’s oldest robotics company. Rich Walker, who had left the Shadow Group to earn a B.A. in mathematics and a diploma in computer science at the University of Cambridge, joined Shadow Robot in 1999 as technical director. Today he’s the director of the company.

Shadow Robot specializes in durable robot hands rather than walking robots. But the focus on hands is also a legacy of the Shadow Group. Walker remembers that the Shadow Group’s first humanoid hand in the late 1990s was impressive simply for being able to pick up a pint of beer (a smooth-sided, thin-walled glass). Today, Shadow Robot’s hands are testbeds for dexterity. Gone are the pneumatic muscles, replaced by actuators that move each finger with precision. The classic model contains 20 motors, allowing for abductive and adductive movement with 24 degrees of freedom.

Black and white photo of a two-legged humanoid robot with its left leg raised, next to a man with his right leg raised while another man looks on. Shadow Walker’s operator wore a data suit that captured his movements and allowed the robot to copy them.Richard Greenhill

In a recent blog post, Sejal Parsotomo, senior marketing executive at Shadow Robot, wrote that while humanoid robots are great for public relations, specialized dexterity is key for success: A robot that can walk into your factory may be impressive, but a robot that can reliably manipulate objects is transformative.

In its struggles to take more than a few steps, the Shadow Walker showed the inherent difficulty that robots had in mastering even low-level skills. In August 2025, Beijing hosted the World Humanoid Robot Games. Competing in sports such as gymnastics, soccer, and track events, as well as more “useful” tasks like hotel cleaning and sorting medicine, these robots could literally have run circles around the competitors in the first Robot Olympics 35 years earlier. And yet, there is still so much work needed in order for robots to navigate the human-built environment. Despite the astonishing progress, we’re still not all that close to actually useful humanoid robots.

Part of a continuing series looking at historical artifacts that embrace the boundless potential of technology.

An abridged version of this article appears in the June 2026 print issue as “Learning to Walk.”

References


Richard Greenhill gives an overview of his life and the founding of the Shadow Group in a post on Shadow Robot’s corporate website.

David Buckley has a compilation of resources on the Shadow Biped Walker, including specifications from the 1999 iteration and a brochure from the 1st International Robot Olympics.

There is coverage of the Robot Olympics worthy of a gossip sheet in La Repubblica and lovely footage of the competition in this TV-am interview of Peter Mowforth by Lorraine Kelly.

Video Friday: Extreme Omnidirectional Robot

29 May 2026 at 17:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

ICRA 2026: 1–5 June 2026, VIENNA
RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO

Enjoy today’s videos!

What is the right number of legs for a robot? Two? Four? No, the answer is obviously all of them. All of the legs.

[ Argus ]

Sigh, yet another skill that I as a soccer-playing human should have but a robot has instead: the rabona.

[ Boston Dynamics ]

Robots are rapidly becoming part of our everyday lives, from drones and industrial machines to home assistants and humanoid robots. As their presence continues to grow, an important question arises: How can we choose the right robot—not only in terms of performance and cost but also in terms of sustainability? This video introduces the Eco‑Score for Robots, a new approach to evaluating the environmental impact of robotic systems. Just as eco-labels help consumers make informed choices in other industries, the Robotics Eco‑Label provides a clear and transparent way to assess how sustainable a robot truly is.

[ Robotics EcoLabel ]

Thanks, Bram!

Uh oh, five-fingered hands.

[ Agility ]

Robotic manipulation has come a long way since the 1990s. We’ve gone from the two-ball paddle juggling robot to AthenaZero, who can juggle barehanded using onboard vision feedback. By moving away from task-specific passive end-effectors such as cups or paddles and using multifingered hands, it can transition between a wide range of patterns including cascade, half-shower, tennis, shower, and box.

There needs to be a robot circus show already.

[ Robotics and AI Institute ]

Zero legs. One hat. $13K.

[ Astribot ]

From its elegant design to the advanced technology powering every step, Luna is more than a machine—it’s a leap into the future.

[ LimX Dynamics ]

Thanks, Jinyan!

You got a quadrotor in my quadruped! No, you got a quadruped in my quadrotor!

[ MARS Laboratory ]

A human hand, a robot’s arm—together tracing circles of trust and precision. No missteps. No hesitation. Just pure, algorithmic grace.

[ UBTECH ]

Low-gravity planetary exploration with a quadruped just looks like fun.

[ Autonomous Robots Lab ]

Here it is, that robot Kool-Aid that everyone seems to be drinking. Including me!

[ Generalist ]

Don’t shoot Mini Pupper!

[ MangDang ]

We show here the ARISTO (Anthropomorphic, Robotic, Integrated-Sensing, Tendon-Operated) Hand. Developed in collaboration with Sony Group Corporation, this research platform is engineered to address the complex requirements of manipulating small, thin, and fragile objects.

[ University of Texas Human Centered Robotics Lab ]

Okay, but did you really have to call it the T800?

[ EngineAI ]

Moby shows what useful mobile manipulation looks like in the real world: picking up, carrying, and placing adaptable payloads. The video shows payload handling across increasing crate loads, including a 50.3-pound load, while maintaining balance, control, and mobility. This is the kind of capability that matters outside the lab—moving real objects, in real spaces, with practical reliability.

[ Noble Machines ]

What does it take to make a robot look human? Harvard SEAS students Hailey Block, Henry Tavistock, and Evan Crowley created “Hollow Minds,” a pair of animatronic heads capable of speaking, blinking, tracking movement, and displaying lifelike facial expressions.

[ Harvard University ]

The longevity here is impressive, but the obvious question here is why the heck you’d ever do this task with a bipedal humanoid robot. It also doesn’t seem to have any error recovery, which is obviously fixable, but highlights the fact that real humans are versatile and humanoid robots are not.

[ Figure ]

Kacper Nowicki, CEO and cofounder of Nomagic, recently sat down for a deep dive into the “humanoid vs. purpose-built” debate during a panel discussion at the Web Summit in Vancouver 2026.

[ Nomagic ]

Video Friday: Atlas Versus a Fridge

22 May 2026 at 16:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

ICRA 2026: 1–5 June 2026, VIENNA
RSS 2026: 13–17 July 2026, SYDNEY
Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO

Enjoy today’s videos!

Just months after its debut, Atlas is proving why it is the world’s most capable and dynamic humanoid robot, ready for real work. Lifting a mini-fridge is a feat of strength, but the true breakthrough is in the underlying reinforcement learning and controls systems. The robot is learning to navigate real world adaptability: handling heavy objects by bracing and accounting for the mass and inertia; using whole-body control, not just hands to maneuver; and demonstrating superhuman range of motion and balance. This marks a critical shift in robotics where humanoids move beyond the lab and into dynamic industrial settings.

Watching Atlas move a fridge may be less impressive than whatever the heck it does at 4:10.

[ Boston Dynamics ]

SpikerBot is a robot you teach by wiring neurons, not writing code. Drag spiking neurons in the app, connect them to sensors and motors, then press play. It moves, reacts, and changes behavior based on the brain you built.

Already funded on Kickstarter with a robot kit starting at US $219.

[ Kickstarter ] via [ Backyard Brains ]

Thanks, Greg!

Wheeled-legged robots, which have wheels at their feet and achieve high mobility by coordinating wheel drive and leg drive, have been developed. In this paper, we address the problem of how to draw out the potential task-execution capability of the legs by freeing them from the roles of locomotion through external body support.

[ WiXus ] from [ JSK Robotics Laboratory ] via [ ICRA 2026 ]

A very clever idea for electronics-free, multi-dimensional touch sensing.

[ Nature Communications ]

Using external voice commands, G1 is directly controlled to generate a wide range of actions in real time. This video was recorded in a single take, with on‑site audio recording.

[ Unitree ]

Hummingbirds are impressive flyers, and advancements in high-speed photography, instrumentation, and measurement techniques have revealed much about their aerodynamics, flight behaviors, and wing and body kinematics. However, comparatively less is known about their natural flight dynamics, which is the relationship among a bird’s flight velocities, the control actions of its wings, and the acceleration of the bird in flight. To investigate this, at the Advanced Vertical Flight Laboratory we have designed, built, and flight tested a biomimetic robotic hummingbird on which is implemented the same techniques for flight control as observed in hummingbirds.

[ Advanced Vertical Flight Laboratory ]

I guess if you’re going to make a robot dog, it’s only fair to give it the ability to frolic in the water.

[ MagicLab ]

The original automated layout robot—the one that showed up when the construction industry was pretty sure robots were lame and then proved otherwise. It has printed millions of square feet of layout across thousands of projects. It built an entire category of construction technology. The category of: Stuff That Actually Does Helpful Work on Real Jobsites. But FieldPrinter 2 is here. It’s faster, tougher, smaller, and smarter. So for FieldPrinter 1, it’s time. Time for a quiet retirement. A mug. Maybe a plaque... But nay, good knight! Thou shalt expire in a blaze of thunderous glory!!

[ Dusty Robotics ]

Here’s an interesting idea for an inflatable monocopter drone.

[ AIRLAB ]

Meet the Lynx S10—a compact all-terrain robot built to deliver industry-grade performance in a lightweight form factor under 20kg.

[ DE Robotics ]

Noble Machines builds general-purpose robots for heavy industry, supporting people with the most hazardous and physically demanding tasks. Attendees at NVIDIA GTC 2026 witnessed the power of autonomous industrial work with Noble Machines Moby.

[ Noble Machines ]

I’m sorry, but Lego bricks should be for humans only.

[ LimX Dynamics ]

Need a robot that can go places? Huskies were around way before legged humanoids, and I bet they’ll be around way after, too.

[ Clearpath Robotics ]

I know this little dude is just a research platform at Disney, but I still want one to be my friend.

[ Paper ]

In March 1982, General Motors announced a rapid and aggressive conversion to robotics. By 1990, GM wanted 14,000 robots in their factories doing everything from painting to welding to assembly. Nowadays, we dream of robots in the factories, doing everything end to end. In the dark. Lights out. Guess what? GM dreamed the same 40 years ago, and they spent an estimated US $60 billion to try to make it reality. In today’s video, we look at General Motors and their dreams of the automated, all-robot factory.

[ Asianometry ]

❌