Normal view

Video Friday: Lift Happens

14 August 2026 at 17:00


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH
Humanoids Summit Seoul: 22–23 September 2026, SEOUL

Enjoy today’s videos!

Speaking from experience, I can tell you that the best part of any DARPA challenge is when things go horribly wrong. And after you enjoy all the crashes (followed by all of the battery fires), get caught up with the DARPA Lift Challenge with video recaps of the final few days.

[ DARPA Lift Challenge ]

Drone delivery: coming soon to a moving vehicle (or perhaps even through an open window) near you.

[ HKUST Aerial Robotics Group ]

This tiny little robot called STEMbot (as in stem, not STEM) can climb up and around plant stems to check for pests. It’s not very fast, but it sure is adorable.

[ STEMbot ]

Monumental’s robots delivered the brickwork for a semi-detached home, laying around 20,000 bricks in a new community.

[ Monumental ]

Meet the world’s most “truss’t-worthy” robot.

[ Modlab University of Pennsylvania ]

Stanford BDML and Honeybee Robotics propose a payload to test gecko-inspired adhesives in spaaace!

[ NASA ]

How can a legged robot organize its own walking while maintaining a desired direction? In this work, we present a Differential Adaptive Steering (DAST) mechanism for directional adaptation in legged robots under decentralized adaptive control.

[ BRAIN VISTEC ]

I do not care even a little bit if a robot fails (safely, of course), as long as it recovers from that failure.

[ Sanctuary AI ]

Even for a robot that doesn’t drink champagne, those are some pretty light pours.

[ Kawasaki Robotics ]

If we as a society would just accept that the appropriate place to store clothing is in a pile on the floor, robots would have a much easier time of it.

[ LimX Dynamics ]

To be fair, this is also the speed at which I fold shirts.

[ Sharpa ]

Our DR02 humanoid robot takes on the stairs with stable, controlled movement—steady steps, steady progress.

[ DEEP Robotics ]

Two words: structural minifridge. Or is it mini fridge...? Whatever, THREE words.

[ AgileX ]

Video Friday: An Italian Humanoid Comes to Life

24 July 2026 at 15:30


Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.

Summer School on Multi-Robot Systems: 29 July–4 August 2026, PRAGUE
Actuate 2026: 18–19 August 2026, SAN FRANCISCO
IROS 2026: 27 September–1 October 2026, PITTSBURGH
Humanoids Summit Seoul: 22–23 September 2026, SEOUL

Enjoy today’s videos!

In just six months, our team turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact. Its full-body multimodal skin perceives touch, proximity, force, and temperature, bringing Physical AI closer to safe and natural collaboration with people. Not a render. Not a concept. This is GENE.01. The future of Physical AI is taking its first steps.

[ Generative Bionics ]

Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between. Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical common sense that transfers to new hands and new ways to grasp, push, pull, twist, and more.

And to illustrate this concept, a surprise spatula.

[ Generalist ]

This paper presents the design, fabrication, and flight validation of a flat-packable flying wing built primarily from corrugated cardboard. The aircraft is manufactured from three laser-cut sheets and assembled through a fold-and-lock architecture that forms load-bearing wing structures with minimal tooling and no permanent fasteners. The full airframe can be assembled in under 15 minutes, demonstrating strong potential for rapid deployment, low-cost logistics, and scalable field use.

[ AIR Lab ]

A $14,000 open-source data-collection system that includes beat-down capability.

[ MEVION ]

Thanks, Kento!

Together with Niantic Spatial and Nvidia, [we] can now scan a real deployment site with off-the-shelf hardware, reconstruct it into a photorealistic Gaussian splat, and run massively parallel RL [reinforcement learning] training. The policies trained in our Gym environment then transfer zero-shot to the real robot and environments they were trained for. This enables faster deployment of more capable and robust policies for the end user.

[ Flexion ]

I don’t know why, but the version of Tron 2 with the stubby little legs is just adorable.

[ LimX Dynamics ]

Uh, get a real job already...?

[ PNDbotics ]

Well, I guess we can all stop asking what humanoid robots are good for.

[ EngineAI ]

I think the right thing to do here is only post the disclaimer included with this video: “This film is a conceptual creative production, and certain scenes are presented for demonstration purposes only and do not represent the actual in-store operating process. The final store environment, robot appearance, and functionality are subject to the actual deployment. During actual operations, the robot will autonomously perform only designated preparation steps for specified ice cream products, and its hands will be fitted with protective gloves that comply with applicable food safety requirements.”

[ Sharpa ]

Drone delivery is now an essential part of the South West London Pathology (SWLP) modernization agenda. Since February 2026, our highly automated aircraft have been delivering urgent NHS samples across south west London, with service up to 85% faster than ground transport. We are thrilled to be part of this initiative, supporting clinicians in providing timely, effective care for patients and contributing to a greener, more resilient NHS.

[ Wing ]

Take a closer look at what’s next for the Aurora Driver. Designed to move freight farther, faster, and more efficiently, this next generation of the Aurora Driver delivers greater performance, built to last one million miles and cut hardware cost in half.

[ Aurora ]

How does a robot learn to recognize an object it’s never encountered? In this case, a demo can be worth a thousand words. Short human demonstrations can be used to create fully automated training datasets, sidestepping the prompting limitations that hold back vision-language models. Rather than describing objects with language, the system tracks what a person touches and manipulates during a demo, follows those objects through time, and clusters detections to handle objects merging or splitting apart in the scene. This bypasses a core weakness of VLMs, which struggle to reliably detect unusual or novel objects even with repeated, carefully engineered prompts.

[ Robotics and AI Institute ]

Beyond Dexterity: Why Contact May Define the Next Era of Robotics

9 June 2026 at 12:51


This article is brought to you by AGILINK.

Throughout the exhibition hall at the 2026 IEEE International Conference on Robotics (ICRA), in Vienna, one demonstration seemed to attract a disproportionate amount of attention.

Two robotic hands were making a balloon dog. Slowly and deliberately, the robot twisted a long balloon into loops, bends, and joints without popping it. Visitors stopped, watched, and often returned with colleagues to watch again.

Crowd at a robotics expo watches a humanoid robot demonstrate its arm movements. AGILINK’s balloon dog demonstration draws a crowd at ICRA 2026.AGILINK

At first glance, the demonstration appeared almost playful. Among roboticists, however, balloon twisting is widely recognized as an unusually difficult manipulation task.

A balloon is lightweight, highly deformable, slippery, and extremely sensitive to force. Every twist changes its geometry and internal pressure, turning a seemingly simple activity into a continuously changing physical interaction problem.

Humans navigate those changes almost intuitively. While making a balloon animal, people rarely think consciously about force regulation, slip prevention, or contact stability. They simply adjust.

For robots, those adjustments remain remarkably difficult. The challenge is not merely moving fingers to the right positions. The harder part is maintaining stable interaction while the object itself is changing.

Highlights from AGILINK’s ICRA 2026 demonstrations, including visuotactile sensing, in-hand manipulation, balloon-animal shaping, and other contact-rich tasks enabled by the company’s latest OmniHand platform.AGILINK

That distinction helps explain why the balloon dog drew so much attention in Vienna. What appeared to be a dexterity demonstration was, in many ways, a demonstration about contact itself.

As robotic manipulation continues to advance, a growing number of researchers are arriving at a similar conclusion: many of the hardest problems in robotics begin only after contact occurs.

Motion and Contact Intelligence for Robot Manipulation

Balloon twisting combines two challenges that robotics has traditionally struggled to solve simultaneously: long-horizon task execution and contact-rich manipulation.

The first concerns motion.

A balloon dog is not created through a single grasp or twist. It emerges through a carefully ordered sequence of manipulations, each setting the conditions for what follows. A small rotational error introduced early may appear insignificant at first, yet several steps later it can prevent the final structure from forming altogether.

In that sense, balloon twisting is a long-horizon task. Success depends not only on performing individual actions correctly, but also on preserving the future feasibility of the entire manipulation process.

To address this challenge, AGILINK began by collecting demonstrations from professional balloon artists. Human actions were mapped onto robotic hands to establish an initial manipulation policy. But successful demonstrations alone were insufficient.

In practice, some of the most valuable learning occurred when execution began to drift toward failure. Whenever instability emerged, human operators intervened and corrected the manipulation in real time. Those interventions were recorded and incorporated into reinforcement-learning cycles, allowing the system to learn not only how successful demonstrations unfold, but also how experienced operators recover when things start to go wrong.

Through this process, the robot gradually acquired the capabilities required for long-horizon task execution—a collection of abilities that AGILINK groups under the term motion intelligence: the ability to generate actions, coordinate bimanual behaviors, and execute extended manipulation sequences under real-world uncertainty.

Two robotic hands, one white open palm and one black forming an OK gesture, on display. OmniHand 3 Ultra-M on display at ICRA 2026.AGILINK

Yet motion alone does not explain why balloon twisting remains difficult. The second challenge is contact.

The robot must continuously regulate force, adjust contact locations, and respond to subtle changes in the object’s state. These decisions are difficult to encode through explicit rules. Even skilled human operators often rely on tactile intuition developed through experience rather than consciously articulated strategies.

Analysis of those interventions revealed that many failures did not originate from incorrect action sequences, but from the breakdown of contact itself.

To better capture those interaction dynamics, AGILINK collected contact-centric intervention data and incorporated those interactions into reinforcement-learning training. Rather than learning only which motions to perform, the system also learned how humans maintain stability when contact conditions begin to deteriorate.

AGILINK describes this capability as contact intelligence: the ability to establish, maintain, and adapt physical interaction as force distribution, friction, deformation, and contact geometry continuously evolve.

The distinction between the two capabilities is subtle but important. Motion intelligence determines what the robot intends to do. Contact intelligence determines whether it can continue doing it. For balloon twisting, both are necessary. One provides the sequence of actions. The other keeps those actions physically viable.

Robot makes balloon animal for visitor at tech expo booth. YouTuber KhanFlicks follows OmniHand’s motions while learning to fold a balloon dog at the AGILINK booth.AGILINK

Between a balloon slipping away and a balloon bursting lies a narrow region of stability. Successful manipulation depends on finding that region—and remaining within it throughout the task.

Introducing the OmniHand 3 Ultra-M Dexterous Hand

The balloon dog demonstration showcased a manipulation capability. It also revealed a broader question. How much contact intelligence can be achieved through learning alone? A robot can only regulate what it can perceive. It can only respond as quickly as its hardware allows.

As manipulation tasks become increasingly complex, researchers are finding that progress depends not only on better policies, but also on richer sensing and faster physical response.

That realization formed the backdrop for AGILINK’s second major announcement at ICRA 2026. Alongside the balloon dog demonstration, the company introduced the OmniHand 3 Ultra-M.

Two robotic hands beside a human hand, all raised open on a display table. OmniHand 3 Ultra-M closely matches the size of an adult human hand.AGILINK

The two exhibits represented different stages of the same technological trajectory. If the balloon dog demonstrated what contact intelligence can already accomplish today, Ultra-M was designed to explore what contact intelligence may require next.

Building Hardware for Contact Intelligence

Roughly the size of an adult human hand, the OmniHand 3 Ultra-M integrates 20 active degrees of freedom within a human-scale form factor.

Its most distinctive feature is a fully direct-drive architecture. By adopting direct-drive actuation throughout the system, the hand is designed to enable faster and more transparent force regulation and higher force-control bandwidth, enabling faster response as contact conditions change. For contact-rich manipulation, responsiveness can be as important as sensing itself.

By adopting direct-drive actuation throughout the system, the OmniHand 3 Ultra-M is designed to enable faster and more transparent force regulation and higher force-control bandwidth, enabling faster response as contact conditions change.

The platform also incorporates tactile sensing across nearly the entire hand. Each fingertip contains a miniature vision-based tactile sensor, while more than 300 three-dimensional tactile sensing points are distributed throughout the palm. Together, they provide information not only about where contact occurs, but how contact is evolving.

The system is designed to estimate pressure distribution, shear forces, local deformation, slip tendencies, and other interaction dynamics that often remain invisible to conventional position-based control systems.

According to AGILINK’s tests, individual sensors achieve force resolution of approximately 0.005 N—roughly equivalent to detecting the weight of a sheet of paper resting on a fingertip. Spatial resolution reaches approximately 0.04 mm, while sensing density approaches 50,000 sensing points per square centimeter.

Robot arm delicately holds a feather, inset shows colorful dotted texture close-up. OmniHand 3 Ultra-M recognizes feather texture through vision-based tactile sensing.AGILINK

For dexterous robots, contact has traditionally been a largely hidden process. Ultra-M is designed to make that process more observable.

Rather than simply detecting that contact has occurred, the system attempts to resolve where interaction is happening, how forces are distributed, whether instability is beginning to emerge, and how manipulation strategies should adapt in response.

The balloon dog offered a glimpse of what contact intelligence can already accomplish. Ultra-M explores a different question: what capabilities may be required to push contact intelligence further?

The Physical World Remains the Hardest Benchmark

The significance of contact intelligence extends far beyond balloon animals. Many tasks that continue to resist automation involve unstable or deformable interaction: cable insertion, garment handling, flexible packaging, delicate assembly, connector mating, tool use, and household manipulation.

These tasks are difficult not because robots cannot reach the correct location, but because maintaining stable interaction after contact begins remains extraordinarily hard.

For decades, robotics achieved many of its successes by reducing uncertainty. Factories were engineered to make robotic motion predictable, repeatable, and highly structured. The physical world behaves differently.

A growing share of robotics research is shifting toward interaction itself—understanding how robots can establish, maintain, and adapt physical contact within environments that remain fundamentally unpredictable.

Objects shift. Materials deform. Friction changes. Contact evolves. Real environments rarely follow scripts. Seen through that lens, the balloon dog was never really about the balloon dog. What attracted attention at ICRA was not simply a visually impressive demonstration, but what it revealed: intelligence in the physical world is ultimately measured through interaction.

As motion generation continues to mature, a growing share of robotics research is shifting toward interaction itself—understanding how robots can establish, maintain, and adapt physical contact within environments that remain fundamentally unpredictable.

For robots moving beyond structured environments and into less predictable real-world settings, managing contact may become as important as motion itself.

The Future of Physical AI Isn’t Smarter Robots, It’s Smarter Interfaces

21 May 2026 at 10:00


This sponsored article is brought to you by Wetour Robotics.

A field technician on a wind turbine, harness clipped, both hands on a wrench, needs to send a command to the diagnostic device hanging at her belt. A logistics worker on a loading dock, gloves on, eyes on the pallet, needs to redirect a connected lift. A person using an assistive mobility device on a crowded street wants to nudge it forward without taking out a phone or speaking aloud. None of these moments call for a smarter robot. They call for a smarter way to be heard by the machines that already exist.

The industry has been building from one side

The past three years of Physical AI have been a story of remarkable progress on the robot side of the loop. Companies like Boston Dynamics, Figure, and Unitree have advanced actuators, locomotion, and dexterity to a level that would have seemed implausible a decade ago. Google DeepMind’s Gemini Robotics has redefined what vision-language-action models can do in unstructured settings. The trajectory of the hardware and the foundation models is real, and it is accelerating.

But there is another side to this loop, and it has been treated as a solved problem for too long. The interface between humans and machines has defaulted, for 40 years, to three input modalities: screens, buttons, and voice. Each of those assumes the user can stop, look down, and translate intent into structured commands. That assumption breaks the moment the work moves into a real environment. On a turbine. On a dock. On a sidewalk. In any setting where hands are occupied, eyes are committed, or speaking is impractical, the conventional interface stack quietly fails.

Spatial Intent Fusion is the simultaneous processing of three streams of human-centered information, namely spatial position, visual context, and gestural intent: Your body is the interface.

The bottleneck on the human side of the loop is becoming as important as the one on the machine side. And solving it requires a different question. Not how do we make the robot more capable, but how do we let the human participate in the computing system as naturally as the robot already does.

Wetour Robotics’ bet: put the human back into the computing loop

Wetour Robotics is betting that the next architectural leap in Physical AI is not about making the robot more capable. It is about making the human a first-class node in the computing network, with the same kind of low-latency, high-fidelity participation that connected devices already enjoy.

Wetour Robotics’ engineers frame the problem this way: a wristband that recognizes a gesture is not enough. A camera that recognizes a scene is not enough. The information a human carries about what they are about to do is distributed across multiple channels, including where their body is in space, what their eyes are attending to, and what their muscles are preparing to do, and any single channel observed in isolation is ambiguous. Reconstructing intent reliably means fusing those channels at the operating system level, with latency low enough that the loop feels closed rather than mediated.

This approach has a name. Wetour Robotics calls it Spatial Intent Fusion: the simultaneous processing of three streams of human-centered information, namely spatial position, visual context, and gestural intent, fused into a single real-time command for any connected physical device. It is the technical implementation behind a simpler positioning statement the company uses externally: your body is the interface.

Sleek silver rectangular electronic device labeled \u201cORCHESTRA\u201d on a light gray background. Orchestra is a portable intelligent hub running the operating system that handles sensor fusion, intent inference, command translation, and safety arbitration. The reference compute platform is NVIDIA Jetson Orin Nano Super, which provides enough on-device inference capacity to keep the entire control loop at the edge, with no cloud dependency on the critical path. Wetour Robotics

The architecture: three layers, four engines, one loop

Orchestra is not a single device but a layered platform, designed from the start to be sensor-flexible and actuator-agnostic. The architecture decomposes into three perception layers and four coordination engines.

Orchestra itself is the local compute and orchestration core: a portable intelligent hub running the operating system that handles sensor fusion, intent inference, command translation, and safety arbitration. The reference compute platform is NVIDIA Jetson Orin Nano Super, which provides enough on-device inference capacity to keep the entire control loop at the edge, with no cloud dependency on the critical path. Edge inference is non-negotiable for this application. Full-chain latency from biosignal acquisition to actuator command is held under 100 milliseconds, the envelope inside which closed-loop control feels natural rather than laggy.

VisionLink handles visual and spatial perception. Cameras feed into vision models that identify objects, estimate distances, and track environmental context. VisionLink is designed not as a passive recognition layer but as a real-time command generator: its outputs feed directly into Orchestra OS to be fused with biosignal data.

Conductor is the biosignal pipeline. It ingests raw surface electromyographic (sEMG) data from a wrist-worn device, classifies temporal patterns into discrete gestures or continuous control signals, and outputs actuator commands. The technically interesting property of sEMG for this use case is that the signal precedes visible motion. Motor unit action potentials appear at the skin surface roughly 50 to 80 milliseconds before a finger completes the corresponding gesture. Wetour Robotics calls this property pre-motion intent sensing, and it is what allows Orchestra to anticipate user intent rather than react to it.

On top of the three perception layers, Orchestra OS runs four coordination engines. The Perception Engine ingests and normalizes raw sensor streams. The Intent Engine performs Spatial Intent Fusion across modalities, resolving what the user is trying to do given where they are, what they are looking at, and what their hand is signaling. The Orchestration Engine translates intent into device-specific command sequences for any connected actuator. The Safety Engine arbitrates conflicting commands, enforces operational envelopes, and gates execution against runtime safety conditions.

The trade-offs we’re honest about

No system that bridges the human body and the digital world is finished. Three engineering challenges remain open, and the company addresses each with a deliberate trade-off rather than a claim of having fully solved it.

Baseline stability of sEMG under motion. In a stationary user, continuous gesture recognition from sEMG is reliable. Once the user is walking, climbing, or otherwise moving, motion artifacts and electrode drift degrade the signal in ways that are difficult to fully compensate for. Rather than overpromise on continuous control in dynamic settings, Orchestra defaults to a smaller set of robust discrete gestures in complex operating environments, and reserves continuous control modes for contexts where the signal-to-noise ratio supports them.

Miniaturization of edge AI compute. Running the Orchestra control loop entirely at the edge requires real on-device inference, which has historically meant trading off between compute capacity, battery life, and form factor. Wetour Robotics’ approach has been a compact carrier board paired with a thermal design and a battery module sized for all-day wearability. The result is a hub that travels with the user rather than tethering them to a desk, and that performs the full perception-to-actuation loop without offloading to the cloud.

Heterogeneity of third-party device protocols. The actuator side of the loop is a fragmented landscape. Different manufacturers expose different command interfaces, different communication stacks, and different safety conventions, and a Physical AI operating system has to integrate with all of them. Wetour Robotics uses an AI-agent layer to negotiate connection and protocol translation adaptively, so that Orchestra OS can ingest data from a wide range of devices, run them through neural network models that infer human intent, and emit the right command on the right protocol for the device on the other end.

Why this matters, and why it helps the rest of the field

The history of computing is a history of interface revolutions. Command lines gave way to graphical user interfaces, which gave way to touch, which gave way to voice. Each transition expanded who could participate in the system and what they could do with it. The next transition is not about a new screen or a new microphone. It is about treating the human body itself as a participant in the computing network, capable of contributing intent at the same speed and fidelity that any other connected node can.

The history of computing is a history of interface revolutions. The next transition is not about a new screen or a new microphone — it is about treating the human body itself as a participant in the computing network.

This path is not a competitor to the work being done on humanoid robots, foundation models for embodied AI, and dexterous manipulation. It is the missing complement to that work. The hardest open problem for humanoid systems is the data: every natural interaction between a human and the physical world is a potential training signal, and most of those interactions are currently invisible to any computing system. As more humans become first-class nodes in the loop, those interactions become observable, structured, and ultimately useful for training the next generation of embodied AI, including the humanoid robots being developed today.

In other words: putting the human back into the computing loop is not just about better interfaces for individual users. It is about generating the kind of grounded, in-the-wild human-machine interaction data that the broader Physical AI ecosystem will need to keep advancing. The robot side and the human side of the loop are not two competing futures. They are two halves of the same one.

That is what Wetour Robotics means when it says: Your body is the interface.

Learn more at wetourrobotics.com.

❌