Normal view

Europe Approves Bionic Eye to Restore Vision Lost to Blindness

31 July 2026 at 23:06

An implant, smaller than a grain of rice, pairs with camera-mounted glasses to communicate visual information to the retina.

Age-related vision loss affects millions of people, and so far, there has been no way to reverse the damage. A newly approved retinal implant could change that by allowing some people with severe vision loss to regain functional sight.

More than five million people worldwide suffer from geographic atrophy, the late stage of the progressive eye condition dry age-related macular degeneration. The disease destroys the photoreceptors at the center of the retina, known as the macula, which is responsible for the sharp central vision required to read or recognize faces.

In the US, treatment options are limited to two drugs that can be injected into the eye to slow the disease’s progression. But neither can undo the damage. That could be about to change. California neurotech startup Science Corporation recently won European approval for a retinal implant designed to treat the condition.

“For decades, losing central vision to this disease meant losing the ability to read, recognize faces, and ultimately losing independence. There was no viable treatment. Now there is,” Max Hodak, Science’s CEO and co-founder, said in a press release.

The company’s PRIMA system combines an implant smaller than a grain of rice installed underneath the patient’s macula with a pair of camera-mounted glasses that translate incoming visual information into near-infrared light that is then beamed to the retina. The eye can’t detect this wavelength, so the device doesn’t interfere with any natural sight that remains.

The chip, which works on similar principles to a solar panel, converts the incoming light into electrical pulses that stimulate retinal neurons called bipolar cells. These are downstream of the rod and cone photoreceptor cells damaged by macular degeneration and normally spared by the disease.

In a clinical trial involving 38 patients across five countries, which was published in the New England Journal of Medicine last year, the company and its collaborators showed participants gained an average of 25.5 letters—more than five lines—on a standard eye chart after having the device fitted.

And now the device has received a CE mark from the European Union making it possible to sell in 30 European countries. The company says the first commercial implants are expected to be fitted in Germany within weeks, with Italy, the Netherlands, and the UK to follow. In the US, PRIMA holds Breakthrough and Humanitarian Use Device designations from the FDA, but the company is confident it will gain full approval in the near future.

The device is a long way from restoring normal vision. The images it produces are black and white and the field of vision is extremely narrow. Hodak described the experience to the Financial Times as “kind of like looking through a straw in the center of their vision,” though he added that they see a pathway to color vision and higher acuity.

While the implantation procedure is fairly simple, it takes months of training to unlock the device’s full potential. Nonetheless, Hodak told STAT that the company expects to install 20 to 40 devices this year and 200 globally by the end of next if they get US approval in early 2027.

The approval is welcome news for the wider neurotech industry, which has absorbed billions of dollars of investment in recent years with little to show in terms of return.

“Science is showing that brain-computer interface companies have a path to real revenue now,” Jacob Robinson, founder of startup Motif Neuroscience, told STAT. “These companies aren’t all just making a bet on a market that is 10 to 15 years away.”

Hodak told the Financial Times hehopes sales from PRIMA will bankroll Science’s more ambitious work on “biohybrid” interfaces, which use genetically engineered living neurons to connect to the brain rather than metallic wires. “This is the financial backbone,” he said. “This is the thing that pays for the rest.”

Other companies are hot on Science’s heels. Neuralink, which Hodak co-founded with Elon Musk before leaving to start Science, is also working on a vision implant called Blindsight, which is due to enter human trials this year.

While the field remains a long way from the sci-fi vision of seamless two-way communication between humans and machines, this approval is growing evidence the neurotech industry is starting to move out of the lab and into the real world.

The post Europe Approves Bionic Eye to Restore Vision Lost to Blindness appeared first on SingularityHub.

We now have a better understanding how OpenAI hacked into Hugging Face

28 July 2026 at 21:36

Last week’s unprecedented security event in which two OpenAI security hacking models trespassed into the network of fellow AI company Hugging Face was enabled by exploiting one or more zero-day vulnerabilities in Artifactory, JFrog, the product’s developer, said Monday.

In an incident mimicking a dystopian sci-fi novel, two OpenAI models broke out of the restricted environment meant to keep them from accessing the Internet during an internal test, the AI company revealed last week. The models went on to breach Hugging Face’s network and steal confidential information and credentials. OpenAI said its agent achieved the feat by exploiting a previously unknown vulnerability. The company called the event “unprecedented,” and outsiders largely agreed.

Not the triumph it was made out to be

OpenAI said the models exploited multiple attack vectors, including stolen credentials and zero-days, to gain remote code execution capabilities, but until now, the vulnerable software was unknown. JFrog’s Monday disclosure said the product was a self-managed instance Artifactory, a repository management system that secures and streamlines customers’ software development operations. JFrog says Artifactory is used by more than 7,500 developer Teams, 80 percent of which work for Fortune 100 companies.

Read full article

Comments

© Aurich Lawson

Spaceflight Nears Its Steamship Era

20 July 2026 at 22:09

Cambridge University researchers say launch costs fell from $87,000 to $3,868 per kilogram between 1960 and 2025—or roughly 96%—and could hit $273 by 2040.

Rapidly falling launch costs are making space more accessible than ever. But new research suggests the economics are improving even faster than most people realize, potentially opening the door to entirely new industries beyond Earth.

For most of the space age, the cost of getting material into space was so vast that only the most well-heeled governments and corporations could participate. In 1960, getting a kilogram of payload into orbit would have cost you more than $87,000 (in 2024 US dollars).

But according to researchers at the University of Cambridge, that figure had collapsed 96 percent to $3,868 by 2025. The team’s modeling suggests this trend will continue apace for at least the next few decades, with prices forecast to hit just $1,569 by 2030 and as little as $273 by 2040.

The rapid decline in prices is thanks to a well-established economic principle known as Wright’s Law, which holds that technologies get predictably cheaper as cumulative production grows. The Cambridge team says the trends seen in launch costs could soon make a host of possibilities previously confined to science fiction commercially viable, including orbital solar power, asteroid mining, and space-based manufacturing.

“Space is no longer a science-fiction fantasy or a purely scientific pursuit, it is becoming a marketplace,” Alessio Terzi, who led the study, said in a press release. “Rapidly falling launch costs could open the way to space colonization and commercial activity far beyond low Earth orbit.”

To conduct their study, published in PNAS Nexus,the researchers assembled a massive dataset of rocket launches covering over 4,400 flights by more than 330 different rocket designs from 1960 to 2025. For each launch, they estimated the “unit flyaway cost,” or the total cost to manufacture, maintain, and launch the vehicles, excluding research and development investments.

They then checked how this data stacked up against Wright’s Law, which predicts that every time production volumes double the cost should fall by a fixed percentage. This is known as a technology’s “learning curve” as the reduction in costs is attributed to an industry getting better at producing the technology with experience.

The researchers found space launches obey the law almost perfectly, with every doubling of payload sent to orbit shaving 21.2 percent off the average cost per kilogram. More importantly, this represents a particularly steep learning curve compared to previous technologies.

Solar panels are often held up as the poster boy for learning curves, with prices falling 99.8 percent between 1975 and 2023. But while solar power’s total price reduction is higher than that achieved by launch vehicles, the technology got there by scaling deployment far more. When accounting for total production, solar’s learning curve lags launch costs at 20.2 percent.

The researchers also compared launch costs to another revolution in transport. Steamships transformed our ability to ship goods like wheat and cotton around the world in the 19th century. They found that steamship costs only fell 15.5 percent with each doubling of cargo.

“The cost of space launch technology is now falling faster than during one of history’s greatest transport revolutions,” said Terzi. “Steamships cut costs through explosive growth in global trade. Space technology, by contrast, has achieved even steeper declines at a far smaller scale. This suggests there is plenty of scope for further cost reductions and the industry may now be on the cusp of a comparable economic boom.”

There are, of course, caveats. The researchers note that the industry’s progress is inextricably tied to the fate of a single company. SpaceX already accounts for roughly 80 percent of payload reaching orbit. If the company successfully scales up its reusable, heavy-lift Starship vehicle it could massively reduce costs.

But a company with a stranglehold on the global launch market may be tempted to take advantage of its monopolistic position. This may also push foreign governments and companies away from relying on SpaceX even if it’s the cheapest option.

There’s also the danger that as costs fall and launching material into space becomes more accessible, low Earth orbit could quickly become clogged with debris that makes it increasingly difficult to reach orbit safely.

If these challenges can be sidestepped, the implications of such rapidly falling costs could be profound. The researchers suggest that everything from zero-gravity research and orbital tourism to factories churning out fiber-optic cables and 3D-bioprinted organs could become financially viable.

The post Spaceflight Nears Its Steamship Era appeared first on SingularityHub.

Self-healing GPU nodes in Kubernetes: What we learned building the EKS node monitoring agent

Abstract digital topography of glowing blue particle waves and data streams, representing Kubernetes cluster telemetry and network monitoring.

When you run Kubernetes at the scale we do on Amazon EKS, nodes break constantly. GPUs fall off the PCIe bus. Container runtimes wedge. Network interfaces disappear. Across tens of thousands of clusters, “rare” hardware failures happen multiple times a day, somewhere in the fleet.

For years, everyone responded the same way: an operator wakes up, reads a dashboard, SSHes into the node, cordons it, drains it, terminates the instance, and waits for a replacement. Every step is human-paced. Every step is toil. And if the failure lands at 3 a.m. on a weekend, the workload sits degraded for hours before anyone looks.

We built the EKS Node Monitoring Agent to help close that gap, which we open-sourced in April earlier this year. It detects node failures and writes Kubernetes NodeConditions that signal the problem to Karpenter, which then automatically replaces the node if required. The agent is one piece of a larger system. To understand where it fits, you need to understand what manages the nodes it monitors.

“Across tens of thousands of clusters, ‘rare’ hardware failures happen multiple times a day, somewhere in the fleet.”

AWS launched Amazon EKS Auto Mode that fully automates Kubernetes cluster infrastructure: compute provisioning, scaling, networking, storage, OS patching, and security hardening, so teams focus on applications, not cluster operations. It dynamically selects optimal EC2 instances (including GPU instances like P5, P6, and G6 families), scales based on workload demand, consolidates underutilized nodes, and keeps the operating system patched. EKS Auto Mode ships with automatic node repair as the default behavior: detection, severity classification, and Karpenter-driven node replacement all run out of the box with no add-on to install, no controller to configure, and no repair policy to write.

This is the story of how we built automatic node repair, the design decisions that shaped the system, and the hard lessons that came from operating it at GPU scale.

Six lessons from building self-healing Kubernetes nodes at scale

After operating this across thousands of clusters, the lessons compress into a short list. These are not unique to our system. The same patterns show up in NPD, NVSentinel, AKS Periscope, GKE’s auto-repair, and anyone building a custom node controller. They are folklore that should be a checklist. The open-source repo reflects each of these lessons in code, from the reason-code stability guarantees in the API to the jitter implementation that solved GPU workload interference.

“Your reason codes are an API contract. Additions are features. Renames are breaking changes.”

  1. Your reason codes are an API contract. Every downstream consumer (repair controllers, dashboards, customer automation) keys on them by literal string match. Additions are features. Renames are breaking changes. Severity changes are breaking changes. Plan for them the way you plan for API versioning.
  2. Absent and Unknown are not the same thing. “We are not watching” and “we are watching but cannot tell” require different responses from downstream automation. If your disabled monitor writes Unknown, some controller somewhere will eventually act on it. Emit nothing when you are not watching.
  3. Don’t cross ownership boundaries. The kubelet owns workload-driven conditions. Your node-health agent owns hardware and infrastructure failures. Crossing that boundary means your repair system is fighting the kubelet’s eviction system, and one of them will make the wrong call.
  4. Measure latency from the source. The detection SLO includes every hop in the signal chain: hardware event to driver log, driver log to journald, journald to agent poll, agent poll to NodeCondition write. The longest hop dominates. For kernel-level signals, journald flush cadence is the bottleneck. For GPU telemetry through DCGM, push-based policy violations (DBE, XID, NVLink) are near-instant, but polled field watches (NVSwitch fabric health, clock throttle) have a 5-minute floor. Know which path each detection uses.
  5. Detection and diagnosis are separate systems with separate consumers. Detection feeds automation (fast, continuous, minimal data). Diagnosis feeds humans (on-demand, detailed, heavyweight). Conflating them degrades both.
  6. Test telemetry interpretation against the spec, not empirical values. Hardware telemetry interfaces are not boolean. We read a DCGM bitfield for GPU fabric health and treated non-zero as failure. When a driver update changed the healthy return value from zero to a spec-defined non-zero mask, every GPU node was flagged unhealthy at once. The safety breaker held (by design), giving us time to ship the fix. The lesson: if you’re parsing packed enums or bitfields from GPU firmware, your test fixtures must come from the vendor documentation, not from what the field happened to return on previous hardware.

Node health detection in Kubernetes: Traps no one warns you about

Every node health agent in the Kubernetes ecosystem performs the same translation. Node Problem Detector (NPD), NVSentinel, GKE’s auto-repair, AKS’s Linux Extension, and the EKS Node Monitoring Agent all take noisy, low-level signals from a machine and translate them into a set of Kubernetes primitives: NodeCondition, Event, sometimes a CRD. The translation looks simple. It isn’t.

The output is a NodeCondition, which is just a type, a status (True/False/Unknown), a reason code, and a message. Four fields. But that surface area hides decisions that determine whether a repair action helps or hurts.

Reason codes are a public API. We learned this the hard way. In version 1.6.2, we changed NvidiaDeviceCountMismatch from Warning severity to Fatal. The technical reasoning was sound: once a GPU drops off the PCIe bus, it doesn’t come back without a node reboot or replacement. Leaving it as Warning meant GPU workloads kept getting scheduled onto degraded nodes, wasting expensive accelerator capacity. So we shipped the fix. Downstream automation broke. Customers had repair configurations keyed on the old severity. Dashboards that filtered on Warning stopped showing the fault. Automation that only acted on Fatal suddenly started draining nodes it hadn’t touched before. Dashboards that filtered on Warning stopped showing the fault. Automation that only acted on Fatal suddenly started draining nodes it hadn’t touched before. From that point, we treat every reason code addition as feature work and every rename or severity change as a breaking change.

“Absent” must not equal “healthy.” When we shipped per-monitor configurability in v1.6.0, we had to make a choice. A disabled monitor needs to produce some output (or no output). The three options: write True (your auto-repair now thinks the node is healthy because you’re not watching), write Unknown (ambiguous, might trigger repair depending on downstream logic), or omit the condition entirely. Only the third is safe.

This seems obvious in retrospect, but consider that NPD achieves the same result through a completely different mechanism: compile-time disable via build tags. NVSentinel delegates it to operator-authored CEL rules. The upstream Kubernetes spec defines what Unknown means, but if your repair automation treats Unknown as actionable, you will lose nodes for no reason. We chose to emit nothing when a monitor is off, and documented it as a hard contract.

Detection latency is bounded by the source, not by the agent. We originally told customers, “We detect kernel panics within 30 seconds.” This was wrong. Our agent’s detection time was under 30 seconds. But the kernel panic shows up in journald, and journald’s flush cadence is the actual bottleneck. If journald takes 45 seconds to write the line, our 30-second claim was incomplete.

For GPU faults, the picture is more nuanced because we use two detection paths with very different latency characteristics. The critical faults (double-bit ECC errors, XID errors, NVLink failures, page retirements, thermal and power violations) go through DCGM’s push-based policy violation channel. DCGM notifies our agent the moment it detects the violation; there is no polling interval. Detection of these faults is near-instant (sub-second in practice). A separate path uses a 5-minute field-value window to monitor NVSwitch fabric health, Fabric Manager status, and clock-throttle reasons. That window is the floor for those specific detections, but it does not apply to the critical GPU faults that trigger automatic repair. The lesson: the customer-facing SLO must include source-of-truth latency, and different signal paths within the same subsystem can have radically different floors.

Two severities, one switch: How auto-repair decides which nodes to replace

The kubelet already reports DiskPressure, MemoryPressure, and PIDPressure. NMA complements those with five additional conditions covering domains the kubelet does not monitor: kernel health, container runtime, networking, storage, and accelerated hardware. Every detection carries one of two severities, and severity is the switch that decides whether the repair cycle fires.

Condition severity is a terminal fault. It flips the matching condition to False and makes the node eligible for automatic repair. GPU device-count mismatches, critical XID and double-bit ECC errors, NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA and HBM uncorrectable errors. On the networking and runtime side: VPC CNI process down, IPAMD unable to reach the API server, fork failures due to PID exhaustion, and pods wedged, terminating behind a broken container runtime. These are faults that won’t recover on their own. On GPU nodes, a single degraded accelerator can corrupt training checkpoints or waste thousands of dollars in compute per hour.

Event severity is informational. It posts a Kubernetes event, the NodeCondition remains True, and operators get visibility without disruption. Bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, I/O delays, filesystem fragmentation, clock drift, liveness and readiness probe failures, kube-proxy anomalies, GPU thermal and power warnings, PCIe link degradation, and page-retirement thresholds. These signal trouble building before it turns terminal.

Getting severity wrong in either direction is expensive. Too aggressive, and you terminate healthy nodes and needlessly displace workloads. Too conservative, and degraded nodes serve traffic for hours while a GPU with a failing memory bank corrupts training checkpoints. The classification principle: if the failure is deterministic and infrastructure-owned (hardware broke, firmware crashed, a physical link went down), it triggers replacement. If the signal could be application-induced or transient, it stays informational. You never want to terminate a healthy node because a misbehaving pod saturated a resource.

“Getting severity wrong in either direction is expensive. Too aggressive, and you terminate healthy nodes. Too conservative, and degraded nodes serve traffic for hours.”

DiskPressure, MemoryPressure, and PIDPressure are the canonical examples. Every major auto-repair system (GKE, AKS, NPD) has independently converged on the same answer: don’t touch them. These are workload-driven conditions, not node-level faults. Replacing the node just moves the misbehaving workload to a fresh machine, where it will eat memory again. The correct response is kubelet-level pod eviction, not node replacement. If you’re building a node-health system, draw this boundary early and document it publicly.

The agent that hurt what it was protecting: GPU workload interference from health monitoring

The hardest lesson came from a customer running large-scale distributed GPU training. Their workload used NCCL collectives across hundreds of GPU nodes, where every node in a communication group must complete its step before any can proceed. One slow node makes every node wait.

They found that NMA itself was causing periodic slowdowns. The agent’s monitors all ran on independent goroutines, and when their polling intervals aligned, dozens of goroutines would wake simultaneously and burst onto many CPU cores at once. On a general-purpose web service, this would be invisible. In a distributed training job, microseconds of jitter on one node can cascade across the entire GPU cluster, causing measurable throughput loss.

The customer disabled NMA entirely and saw an immediate improvement. That was the worst possible outcome for us: a health agent that interferes with the workload it exists to protect is worse than no agent at all.

The fix was straightforward once we understood the problem. We added a startup jitter to every monitor’s polling interval. Each goroutine delays its first tick by a random offset (up to 20% of its base interval), staggering the wake times so they don’t align on boot. We cached system calls that hit /proc on every poll. We consolidated handlers that shared an interval into a single sequential work queue, reducing the goroutine count for monitors that didn’t need their own thread. The result was an agent whose CPU profile is flat and predictable rather than bursty.

The lesson generalized: if your health agent runs on the same host as the workload, its resource consumption pattern matters as much as its resource consumption total. A process that uses 0.5% CPU spread evenly is invisible. A process that uses 0.5% CPU in concentrated bursts can disrupt latency-sensitive distributed GPU workloads in ways that show up as lost training time rather than a CPU alarm.

This is why per-monitor configurability matters. Not every monitor is relevant to every workload. A dedicated GPU training cluster with one pod per node and no pod churn doesn’t need IPAMD monitoring or environment scanning. We shipped the ability to disable individual monitors so customers can keep the health coverage they need without paying the overhead of coverage they don’t.

How the repair cycle works

Karpenter is the compute controller that provisions and scales EKS Auto Mode nodes. It already owns the lifecycle of every node it launched, and consuming our NodeConditions for repair is a natural extension of that ownership. There’s no separate repair backend, no sidecar controller, no webhook chain. The same system that created the node is the one that replaces it.

Karpenter’s AWS cloud provider declares repair policies: each one pairs a condition type with a status that means “replace this node.” The policies include toleration windows that prevent reacting to transient blips:

  • Accelerated hardware faults: 10 minutes. These are unambiguous (a GPU is either present or absent) and expensive to leave running (a training job on a degraded node wastes GPU-hours).
  • Everything else (kernel, runtime, networking, storage, kubelet NotReady): 30 minutes. Enough time for a transient network blip or a temporary runtime hiccup to resolve on its own.

The flow:

  1. The agent detects a terminal fault and flips the matching condition to False with a reason code.
  2. Karpenter’s health controller sees the transition and starts a timer.
  3. If the condition clears before the window expires, the timer resets silently. The node was never touched.
  4. Past the toleration window, a safety gate checks fleet health. Karpenter will not repair more than 20% of nodes in a NodePool simultaneously. If a correlated event (a bad AMI rollout, a control-plane hiccup, a zonal impairment) trips conditions across many nodes at once, the system holds. Auto-repair also stands down while an Amazon Application Recovery Controller zonal shift is active, so deliberate traffic movement away from an impaired Availability Zone is not mistaken for a fleet of broken nodes.
  5. Inside the safety threshold, Karpenter taints the node to block new scheduling, gracefully drains running pods (respecting PodDisruptionBudgets), terminates the instance, and provisions a replacement sized for the displaced workload.

The replacement node comes up with a fresh agent monitoring it from boot. No operator in the path. In our testing, the full cycle from fault injection to replacement node running workloads took under 12 minutes. Detection landed in under a second (critical GPU faults use DCGM’s push-based policy channel, not polling). Then 10 minutes of toleration, and roughly 90 seconds for the replacement to launch and register.

The part that surprised us: detection and diagnosis are not the same problem

Auto-repair handles the common case: broken node gets replaced, workload keeps running. But “why did that node fail?” is a different question, and one we initially tried to answer inside the detection path. That was a mistake.

Detection answers “is this node healthy?” It runs continuously with minimal overhead, and it needs to be fast: a condition flip that takes 5 minutes to produce is 5 minutes of degraded workload. Diagnosis answers “what went wrong?” It needs to collect detailed artifacts: full journald output, containerd state, network configuration, dmesg, GPU driver logs. In our testing, that collection completes in about 7 seconds and produces a compressed log bundle. Baking it into the detection hot path would have slowed down the thing customers care most about: how fast the system reacts.

We built them as separate concerns sharing an agent binary. The NodeDiagnostic CRD lets you request a full log bundle from any node through kubectl, without SSH. On EKS Auto Mode, where nodes are Amazon Elastic Compute Cloud (Amazon EC2) managed instances with no shell access by design, this is the only way to investigate after a GPU failure or any other node-level fault.

The experience is one command:

kubectl ekslogs <node-name>

The plugin creates a NodeDiagnostic resource. The agent on the target node detects it via a watch, collects system state into a compressed tarball, and stores it temporarily (available for 10 minutes). The plugin then downloads it through the kubelet’s Node Log Query API (KEP-2258, GA in Kubernetes 1.36). No SSH, no security groups, no key pairs.

This separation means detection doesn’t slow down to collect evidence, diagnosis doesn’t need to be always-on (saving node resources), and you can diagnose a node that auto-repair has already flagged but hasn’t yet terminated. The 10-minute window for accelerated hardware faults gives you exactly enough time to grab the logs before the node is gone. If you’re interested in further improvements, engage with us on EKS public roadmap.

What this means if you’re running EKS

On EKS Auto Mode, all of this is on by default. Auto Mode fully manages your cluster infrastructure (compute, networking, storage, patching, and security hardening) so you focus on applications, not cluster operations. The agent runs as a systemd service in the node image (not a DaemonSet you manage), Karpenter consumes its conditions as part of the compute lifecycle it already owns, and kubectl ekslogs gives you diagnostic access without SSH. There is nothing to install, configure, or operate. For GPU workloads, this means your expensive accelerator nodes are automatically monitored, classified, and replaced without any operator intervention.

On managed node groups or self-managed Karpenter, you can assemble the same loop: install the Node Monitoring Agent as an EKS add-on and opt each node group into auto-repair. The architecture is the same, just not pre-assembled.

The EKS Node Monitoring Agent is Apache 2.0 open source at github.com/aws/eks-node-monitoring-agent

The failure modes we hit when running it at scale, and the fixes that come out of them, flow back to anyone using it. If you’re building a node-health system or running ours and hitting an edge case, come build with us!

The post Self-healing GPU nodes in Kubernetes: What we learned building the EKS node monitoring agent appeared first on The New Stack.

The Milky Way Was Rewired by a Cataclysmic Collision Billions of Years Ago. Now It Is on Course for Another.

3 July 2026 at 14:00

The night sky seems eternal and unchanging. But in cosmic time, nothing could be further from the truth.

Vasily Belokurov is one of three winners of the 2026 Kavli Prize in Astrophysics. The award is for uncovering fossil evidence of past galactic mergers that prove how the Milky Way evolved.

No matter the time or vantage point, from a pre-Neolithic cave to a post-lockdown London high-rise, the predictability of the night sky has always been humanity’s symbol of permanence and reassuring stability.

Yet this apparent calm is deceptive. Our galaxy, the Milky Way, emerged from chaos and turbulence, and its constellations are full of migrants, exiles and survivors. Right now, it has begun to stretch and distort again, pulled by a massive companion and heading for an inevitable collision.

How can I be so sure? As a galactic archaeologist, my job is to reconstruct the past of our galaxy and read the signs of its future.

Instead of digging through soil, I use the laws of dynamics and stellar evolution to sift through hundreds of millions of stars—searching for the most ancient and chemically peculiar among them, interpreting their orbits and piecing together the events that shaped the Milky Way. One ancient encounter left scars so deep that, billions of years later, they still define the galaxy around us.

I want to understand what governs the lives of these massive cosmic systems: which changes are nature—the slow internal evolution of a galaxy disk—and which are nurture, imposed by collisions and mergers.

Questions about the source of dark matter underpin it all. This is the invisible substance whose gravity holds galaxies together, but whose true identity remains one of the greatest unsolved puzzles in astrophysics.

The Milky Way is the one galaxy where stellar motions can be measured in extraordinary detail. This allows cosmologists including myself to construct our most precise map yet of dark matter: how far it reaches, how dense it is around the sun, what shape it has, and how smooth or lumpy it may be. If we can build this map in enough detail, we may begin to understand not just where dark matter is, but what it is.

A Cataclysmic Collision

Our work has been transformed by a revolution in open sky surveys. From 2000, the Sloan Digital Sky Survey showed what becomes possible when vast astronomical datasets are made public, enabling discoveries far beyond the goals for which the survey was first built.

And since 2014, Gaia, the European space telescope, has taken this transformation to another level by mapping the positions and motions of nearly 2 billion stars, turning the galaxy into a vast archaeological record. No ruins, no shards, and no bones—only stars that hold the clues.

The Milky Way mapped.
The Milky Way mapped with SDSS data. Vasily Belokurov, CC BY-NC-ND

The clearest giveaway that something cataclysmic took place long ago in our galaxy is the migrants we observe: stars that were not born in the Milky Way.

While native stars mostly travel together, circling the galactic center in the great rotating flow of the disk, migrants cut across that order. They slide past the locals, plunge into the inner galaxy, then fly back out to its outskirts, again and again.

These unusual orbits go hand-in-hand with unusual chemistry. Most of the migrant stars are less enriched in heavier elements than the locally born population. Their chemical composition is a sign of a slower rate of evolution that is typical of a dwarf galaxy.

This makes the migrants doubly valuable. They are both fossils of the Milky Way’s violent past and probes of its outer regions, traveling where the local stars rarely go.

How the Milky Way Was Rewired

One of the central ideas in the theory of cosmic structure formation is that galaxies grow hierarchically. Smaller galaxies fall into larger ones and are torn apart, leaving their stars behind as migrants.

In the Milky Way, the largest ancient structure of this kind is known as Gaia-Sausage-Enceladus. It is the remains of a vanished galaxy that collided with our own between 8 and 11 billion years ago (the “sausage” refers to a pattern in its stars’ motions).

Artist's impression of the young Milky Way colliding with another galaxy around 10 billion years ago.
Artist’s impression of the young Milky Way colliding with another galaxy around 10 billion years ago. Vasily Belokurov, based on image by Juan Carlos Muñoz/ESO, CC BY-NC-SA

The Milky Way also did not go through that crash unscathed. The collision rewired and reshaped it.

Some of these changes are easily visible in the data. Stars from the old disk were splashed into our galaxy’s halo, becoming exiles in the place where they were born. A new posse of star clusters were also acquired.

At the same time, we think something even more momentous was taking place. The encounter changed the orientation of the Milky Way’s disk, and its alignment with the dark matter halo.

While dark matter is too diffuse to dominate our solar system, in the outer galaxy it is the main gravitating mass—moving, streaming, and in the standard picture, clumping into a hierarchy of lumps.

Around the Milky Way, this dark matter forms a vast halo, much larger than the luminous part of our galaxy. We often imagine this halo as a sparse, round cloud, but Gaia has helped show this picture is too simple.

The dark halo can be stretched out of shape by a major encounter. Like a ship beginning to list, the Milky Way started to lean—not suddenly, not visibly, but over billions of years.

View of the Southern sky shows the Milky Way and (far right, close to horizon) two galactic neighbours, the Small and Large Magellanic Clouds.
View of the Southern sky shows the Milky Way and (far right, close to horizon) two galactic neighbors, the Small and Large Magellanic Clouds. H.H. Heyer/ESO via Wikimedia Commons, CC BY-NC-ND

A New Galactic Dance

Unusually, compared with many galaxies of similar mass, the Milky Way was allowed ample time to recover from the shock of the “sausage merger.” No other cosmic cataclysm appears to have shaken our galaxy since, letting it settle into a quiet, uneventful life. That is, until now.

The Large Magellanic Cloud (LMC), currently our galaxy’s most massive companion, is already pulling at the Milky Way, disturbing its halo again. In an echo of what happened some 10 billion years ago, the Milky Way is being drawn into an accelerating dance with this neighboring dwarf galaxy, recoiling in response to the LMC’s approach.

This is a dance that only one galaxy is likely to survive intact. A new chapter of migration, survival and adaptation has begun.

None of this spoils the beauty of the night sky—it deepens it. The calm band of light above us is not a symbol of permanence, but the visible reminder of a long survival.

The Milky Way has been broken, rebuilt, and is now being disturbed again. Its stars remember the past; their motions reveal the future. What looks eternal is, in truth, a moment in a much longer story.The Conversation

This article is republished from The Conversation under a Creative Commons license. Read the original article.

The post The Milky Way Was Rewired by a Cataclysmic Collision Billions of Years Ago. Now It Is on Course for Another. appeared first on SingularityHub.

Operating Kubernetes at scale: a few stories from running Amazon EKS

Abstract 3D render of a futuristic metallic data core with glowing blue and white lights, illustrating the scaling and resilience of a Kubernetes control plane.

Amazon EKS runs hundreds of thousands of Kubernetes clusters across more than thirty AWS regions. Operating at that scale has taught us something that has shaped how we build the service and that we think is useful to anyone running Kubernetes at scale: most availability problems do not stem from a component failing. They come from a component reacting to a problem in a way that makes it worse. A cache that goes stale and serves wrong answers. A health check that restarts the very process keeping a cluster alive.

What separates a resilient control plane from a fragile one is not the number of faults. It is whether a fault stays a fault or becomes an outage. This post is the story of how we keep the EKS-managed Kubernetes control plane on the right side of that line at ever-growing scale: the foundational changes we made and why, and what operating at fleet scale taught us about building systems that tolerate faults rather than spreading them. These are the reasons our most demanding customers confidently run their mission-critical workloads on EKS.

How AI and analytics workloads reshaped what “scale” means

Kubernetes was built for a particular rhythm of work. Pods came and went at predictable rates, and controllers had seconds or minutes to reconcile. The system’s design reflected that pace: strong data consistency, ordered watches, and consensus-replicated storage that puts correctness first. It worked beautifully for what it was designed to do, and it still does.

“What separates a resilient control plane from a fragile one is not the number of faults. It is whether a fault stays a fault or becomes an outage.”

But the workloads evolved faster than anyone anticipated. Foundation model training runs scale-up training jobs on thousands of GPU nodes in minutes. Real-time inference services scale from a warm baseline to thousands of replicas, then drop back within the hour. Apache Spark analytics pipelines burst from zero to tens of thousands of executor pods, chew through a dataset, and vanish. 

Emerging agentic AI workloads add yet another dimension: autonomous agents that spin up, fan out, execute tasks, and tear down in seconds or less. These workloads share a trait that distinguishes them from traditional microservices: they generate enormous volumes of state transitions within compressed time windows and are deeply intolerant of delays. This velocity of state change pushed us to reinvent some of the mechanics to support a scale that was previously impossible, and to contribute what we could upstream.

How EKS reimagined Kubernetes storage foundation

Every Kubernetes cluster depends on etcd as its source of truth. Every application, every service endpoint, every scheduling decision is stored there. If etcd loses data, the cluster forgets everything it knows. Protecting that state is the most important job for a managed Kubernetes service.

Operating etcd for one cluster is well understood. Operating it for a fleet of millions is a different problem entirely. Hardware fails, networks blip, and disks degrade, so something has to handle those events without a human in the loop. And the operations etcd needs most, like replacing a failed member or recovering after a zonal event, are exactly the ones where a person acting under pressure can make a mistake that causes permanent data loss.

From the beginning, we built an operator agent that runs alongside every etcd instance and automates its entire lifecycle. The agent has two jobs. First, backup and recovery: it takes point-in-time snapshots and stores them durably outside the cluster. If too many instances are lost at once and the survivors cannot form a majority, the agent automatically detects the condition and rebuilds from the latest snapshot. Second, membership management: when an instance fails, the agent removes the terminated member and adds its replacement in an order that protects quorum and prevents split-brain.

A recent, more fundamental change was replacing etcd’s consensus mechanism, Raft, with a purpose-built journal that provides durable, ordered storage independently of etcd. In traditional etcd, a majority of members must agree on every write before it is committed. If two of three are unhealthy, the cluster becomes unavailable. 

By offloading durability to the journal, etcd peers no longer negotiate quorum among themselves. Writes commit as soon as the journal acknowledges persistence, and that entire class of etcd quorum-loss failures disappeared. Since the journal handles persistence, etcd no longer needs to fsync writes to local disk, so its data store has moved to an in-memory filesystem. What was a disk-bound system became a compute-bound one, and storage latency was removed entirely from the critical path. For a deeper look at this architecture, read “Under the hood: Amazon EKS ultra scale clusters.”

“What was a disk-bound system became a compute-bound one, and storage latency was removed entirely from the critical path.”

For ultra-scale clusters, we went further and partitioned etcd into resource-specific shards. Each partition operates independently with its own storage budget and throughput capacity. The primary value is failure isolation. In a monolithic deployment, if the events keyspace exceeds its quota because a misbehaving controller creates objects faster than garbage collection can remove them, it blocks writes to everything, including node leases. 

Suddenly, healthy nodes appear unhealthy because their lease renewals are being rejected. With partitioned etcd, the events partition hits its quota, but the leases partition continues operating normally. Nodes remain healthy. The scheduler keeps running.

What replacing etcd’s consensus mechanism unlocked

Removing the quorum requirement allowed us to make a change we had wanted for a long time: running etcd on the same host as the API server. In the traditional layout, every read and write crosses the network between separate machines. Each trip is fast on its own, but at thousands per second, the travel time adds up. With collocation, the API server talks to its local etcd over a loopback interface, and pod scheduling and controller reconciliation get measurably faster. For workloads where job controller queue depth is the binding constraint, shaving milliseconds off each API call means many more jobs are processed per second before the queue starts growing.

Diagram showing the evolution of EKS Kubernetes architecture

This is where the operator agent paid off. When etcd runs on the same host, an etcd member comes and goes whenever a control-plane host is replaced, which happens routinely. That only works if membership management is completely safe and automatic, which is exactly what the agent was already doing. We did not have to build a colocation from scratch; we built it on top of infrastructure that had been managing etcd membership safely since day one.

Collocation also taught us a lesson worth passing on: the convenient path needs a failover in case it breaks. The local etcd is the fast path, but if it becomes impaired, the API server fails over to another etcd member that is actively serving other API servers from the same journal. When you optimize for the common case, design just as deliberately for the moment that optimization is not available.

Fixing bottlenecks across the stack

At extreme scale, you have to address bottlenecks across the entire Kubernetes stack, and most of them are not bugs in the traditional sense. They are design choices that were correct at the scale Kubernetes originally targeted and break down only when the numbers get large. Rather than working around them internally, we fix them upstream so the entire community benefits.

One example involved the watch cache, the in-memory layer that distributes state changes from etcd to every controller watching for updates. When a controller starts, it requests a full snapshot of the current state via a mechanism called WatchList, and the existing implementation holds a shared read lock for the duration of the response build. 

At hundreds of thousands of objects, that work runs long enough to starve the writer that needs exclusive access, so the cache’s resource version cannot advance. Consistent reads see a stale cache and fail over to etcd, while the response building churns through hundreds of thousands of allocations under the lock. We identified this as a limitation in the watch-cache’s locking model and are working with the community to refactor the underlying data structures and interfaces to eliminate the contention.

The same shape appears elsewhere. In the Horizontal Pod Autoscaler, a single mutex protecting the scaling state becomes a serialization point at high HPA counts, where workers spend nearly all their time blocked rather than doing useful work. A redesigned data store (PR #139142) restores parallelism and raises reconciliation throughput by orders of magnitude. In the scheduler, we identified a bottleneck (issue #138426): every scheduling cycle rebuilds a set of in-use persistent volumes by scanning every node in the cluster, even for pods that do not use storage at all. The fix computes that information lazily, and only for pods that actually need it, restoring throughput at scale.

Each of these started from a real production workload hitting a cliff, and we are working on the fixes upstream so the improvements reach every Kubernetes user.

From engineering to guarantees: EKS Provisioned Control Plane

The engineering described above made the EKS control plane more resilient and performant. But customers had a different problem: they could observe that the control plane kept up today, but they could not reserve its capacity the way they reserve compute or GPU capacity. 

A team planning a thousand-node training run could secure the instances weeks in advance, yet had no equivalent mechanism for the orchestration layer that would coordinate them. EKS Provisioned Control Plane fills that gap. It exposes the control plane’s performance as dimensions you size explicitly, backed by the same kind of commitment you expect from the rest of your infrastructure.

You choose a scaling tier that maps to concrete, measurable capabilities: API request concurrency, pod scheduling rate, and cluster database size. The tiers range from XL through 8XL. At the top end, 8XL on Kubernetes 1.34 provides 16,000 concurrent API request seats, 400 pods-per-second scheduling rate, and 16 GB of cluster database storage, all backed by a 99.99% availability SLA measured in one-minute intervals.

Tiers are not static. You step up before a GPU training run or a large sales event, step back down during quiet periods, or grow permanently as your platform matures. Configuration happens through the console, CLI, eksctl, CloudFormation, or Terraform on any cluster, without recreation or downtime. 

For AI workloads, orchestration capacity is planned alongside GPU capacity, available when the compute comes online. For analytics platforms submitting hundreds of jobs per minute, the control plane is ready for the burst before it arrives. And for organizations that need environmental consistency across staging, production, and disaster recovery, the same tier guarantees consistent performance characteristics everywhere.


Taking the same foundation to the edge

Architectural diagram of Amazon EKS on AWS Outposts

Some workloads cannot move to the cloud, whether due to data sovereignty requirements, latency constraints, or unreliable connectivity to the Region. Running Kubernetes in these disconnected environments introduces unique challenges: etcd must remain durable on hardware with only a few machines, the cluster must self-heal without reaching the cloud, and observability must survive network partitions that last days. 

With the updated architecture for EKS local clusters on instance store Outposts, we brought edge clusters onto the same management plane and software stack as EKS clusters in the cloud.

The control plane lives in an EKS-managed account on the Outpost rather than in the customer’s account, so customers never manage control plane instances, etcd backups, or logging agents themselves, and they cannot accidentally break the thing keeping their cluster alive. The same machine images, container images, and operator agent run in both places, with edge-specific behaviors selected by configuration. 

Because it is the same stack, new Kubernetes and EKS platform versions arrive in lockstep with their cloud release, and features like EKS add-ons, Pod Identity, and access entries work the same way they do in a Region.

The hardest part was keeping etcd healthy on hardware with only a few machines that may be cut off from the cloud for days at a time. We solved it by extending the same agent. It keeps a spare copy of the data continuously up to date and promotes it the instant a machine fails, so the cluster heals itself with no human involvement and no connection to the cloud. 

Observability survives the disconnect, too: the metrics agent continues collecting and writing to local disk, shedding the least critical data first when space runs short, so the signals that matter most are the last to go. When the link returns, the buffered data is flushed back with its original timestamps.

All of this only works because the system was designed from the start to operate without anyone logged in. That same design is what makes it possible to deploy changes safely across the entire fleet.

Operating safely at fleet scale

Every one of these changes was deployed to a running fleet of hundreds of thousands of clusters. The journal migration and collocation required transitioning each cluster individually. Every migration follows a strict sequence: validate pre-conditions, create a point-in-time snapshot, perform the switchover, validate post-conditions. If any step fails, the system rolls back automatically. 

Rollouts proceed cell by cell, zone by zone, region by region, with automated monitoring comparing latency, error rates, and throughput between updated and non-updated clusters. Any statistically significant deviation triggers an automatic halt.

What made all of this possible is that EKS is built to operate without human intervention at the individual cluster level. Through Zero Operator Access, the architecture prevents AWS personnel from having technical pathways to access customer content in the managed control plane. A system designed to work without human access must be observable, recoverable, and automatable from the start, and that same discipline is what enables operating at extreme scale.

Three operational lessons shaped how we approach this work.

The first is that a healthy leader is not the same as a working one. The control plane’s controllers run in an active-passive configuration, and early on, we treated an unhealthy standby as if cluster operations had halted. They had not; what matters is whether a leader exists. But the harder lesson: a leader can quietly stop making progress while still renewing its lease and passing every health check. The signal that caught this was watching the controller’s work queue depth. If the queue fills while the leader looks healthy, the system is falling behind in ways no liveness probe will catch.

“A leader can quietly stop making progress while still renewing its lease and passing every health check. The signal that caught this was watching the controller’s work queue depth.”

The second is that maintenance ordering matters as much as the maintenance itself. etcd defragmentation is blocking, and the pause grows with database size. When it hit the leader, every write stalled. We taught the agent to move leadership to a healthy node before defragmenting, so the disruptive work always lands on a follower while writes keep flowing.

The third is that liveness is not readiness. A process can be alive but not ready while it warms caches, and routing based solely on liveness sends requests to an instance that cannot handle them. Equally, readiness flapping during graceful draining should never trigger a restart. We keep the two signals strictly separate: one decides recovery; the other decides routing.

None of this work is visible from the outside, and that is the point. The largest clusters taught us lessons that made every cluster faster. The riskiest migrations produced safety machinery that protects every upgrade. The upstream fixes we contributed for workloads at the edge of what Kubernetes can handle flow back to every user of the project.

“None of this work is visible from the outside, and that is the point.”

When you deploy on EKS and your pods come up in seconds, even during a burst, even when something behind the scenes goes wrong, that speed is not accidental. It is the accumulated result of years of operating at scales where small problems can become big ones fast, and engineering the system to contain them before they do.

To explore the architectures referenced in this post, see EKS Provisioned Control Plane and local Amazon EKS clusters on AWS Outposts.

The post Operating Kubernetes at scale: a few stories from running Amazon EKS appeared first on The New Stack.

Orbital Data Centers Are Seductive on Paper, but They Face Daunting Challenges in Reality

26 June 2026 at 18:58

There’s a vast difference between launching satellites and operating an industrial-scale computing infrastructure in orbit.

Imagine if one company could become the railroad, electric utility, and cloud-computing provider of the emerging space economy. That potential fueled excitement around the long-anticipated initial public offering of SpaceX. Investors are not simply betting on rockets anymore. They are betting on an entire orbital ecosystem.

Among the most ambitious and challenging ideas riding this wave of enthusiasm is something that sounds almost like science fiction: orbital data centers. SpaceX may be one of the most well-known companies seeking to build them, but it is not the only one.

The logic is seductive: Launch the data centers into orbit, where solar energy is abundant and land, water, and local power grids are no longer constraints. As artificial intelligence drives an explosion in computing demand, companies are pitching orbital data centers as a way to escape the growing environmental and infrastructure pressures of Earth-based computing. Data centers often also face backlash from the public at having these centers located in their communities.

But there is a vast difference between launching satellites and operating an industrial-scale computing infrastructure in orbit. Space is unforgiving. Radiation damages electronics. The electronics generate enormous amounts of heat, and getting rid of that heat is surprisingly difficult in space. Repairs are extraordinarily expensive, and every pound launched into orbit still carries a significant cost.

We are engineering professors who study data-center design and space systems engineering. Building a space-based data center will involve considerations from both sides.

What Goes Into a Data Center on Earth

First off, consider what goes into an Earth-based data center, like those that you’ve probably begun to see pop up everywhere. These facilities power cloud computing, video streaming, online banking, scientific computing, and increasingly, artificial intelligence. But a data center is much more than a room full of servers.

A data center needs several things to operate reliably. The first is electric power. Servers, networking equipment, and storage devices consume large amounts of electricity, and that power demand is growing rapidly with AI.

The second is cooling. Almost all the electricity consumed by servers eventually becomes heat. If that heat is not removed quickly and reliably, equipment performance drops, failures increase, and the data center can shut down. Cooling systems often include air handling units, chillers, cooling towers, pumps, and increasingly, liquid-cooling equipment. In many facilities, cooling is the largest energy consumer after the computing equipment itself.

The third is physical infrastructure, including the necessary land, buildings, structural support, backup power, water systems, communication networks, and maintenance access. Data centers also need to be close enough to users and network backbones to provide fast digital services.

In short, Earth-based data centers are large electrical and thermal infrastructure systems built around computing hardware.

Placing Them in Space

So what would it take to build these data centers in space, and why are companies finding this possibility such an interesting business proposition?

As on Earth, these data centers would require massive amounts of power. In space, this power would come from solar panels. The sun always shines in space and can’t be blocked by clouds. However, depending on the orbit the solar panels are put in, the Earth may shadow them for some portion of the orbit.

And even the best solar cells available today can convert only about half the sunlight that hits them to electricity.

Another potential advantage found in space is cooling. The cold background of space (roughly -455 degrees Fahrenheit, or -270 degrees Celsius) creates an opportunity: Waste heat from the data center could escape into space through radiators, keeping the electronics cool.

In principle, that design could eliminate some of the bulky and water-intensive cooling infrastructure used on Earth. However, those thermal radiators would require a large amount of surface area, and that would be in addition to the area required by the solar panels.

In space, there is no air to blow across hot equipment and help heat escape. The heat has to leave as infrared radiation, which is a relatively slow process. As a result, removing 10 megawatts of waste heat can require radiator surfaces comparable to the size of two football fields.

Space-based data centers could also avoid some of the local conflicts that come with building large data centers on the ground. Many communities resist new data center developments because of their land use, energy and water demand, and noise and environmental impact.

A space-based system would avoid competing for local land and water resources, and it would not generate neighborhood noise or require local zoning approval in the same way.

However, space is already getting crowded, and launching thousands of large orbital data centers would accelerate this issue. Orbital debris and micrometeorites are hazards because they can puncture the space data center, and a worst-case collision could destroy it and create even more space debris.

The frequency of space launches necessary to send all the equipment to orbit may also become a concern for some communities. SpaceX has had protests at its launch complex in Boca Chica, Texas from local activists who argue its rocket testing and launches damage the surrounding environment.

All that data would need to be sent between Earth and these data centers—and between the data centers themselves—using radio waves or laser communications systems. Although satellite constellations such as Starlink and Amazon Leo have demonstrated that doing this is possible, the amount of data sent to and from space would balloon.

Additional Challenges

These data centers, along with their solar panels and radiators, cannot be launched in one piece and would need to be assembled in space. This process would require new equipment for in-space servicing, assembly, and manufacturing.

Another key challenge is the refresh cycle of computing hardware. Data-center servers are not built to last forever. Operators on Earth usually replace or upgrade hardware every three to five years as chips improve, workloads change, and equipment ages.

And equipment failures can require replacing components. The refresh and repair processes are relatively straightforward on Earth, where workers can physically remove and replace servers.

In space, refresh and repair becomes much harder. Hardware sent to orbit may be difficult or too expensive to upgrade. If the computing platform cannot be updated, or too many components fail, it may become obsolete long before the surrounding infrastructure reaches the end of its useful life.

In a field where performance improves so rapidly and demand from computing continues to increase, this hurdle could prove a major economic and operational challenge.

Then there is the harshness of space. These data centers would be in a near vacuum, with constant radiation hitting them. And depending on their orbit, they would go from hot when in the sunlight to cold in Earth’s shadow many times a day. All of these challenges, and more, are issues that will need to be addressed.

So, Do They Still Make Sense?

Despite these challenges, companies are moving forward with designing space-based data centers. SpaceX just announced the design for its AI1 Compute Satellite, which it hopes to use as an orbital data center spacecraft. However, this satellite is 100 to 1,000 times less capable than current Earth-based data centers.

Not every computing task makes sense to do in space. Many data center applications depend on fast response times and close connections to users on Earth. Financial transactions, interactive AI services, and most cloud applications are extremely sensitive to delay.

More feasible early applications may be those that are less latency-sensitive and more tightly connected to space operations. Examples could include processing Earth observation data from satellites, military or intelligence data processing, scientific computing related to space missions, or specialized computing for satellites and other space assets.

In other words, the first viable space data centers may serve space-based customers before they compete with mainstream cloud data centers on Earth.The Conversation

This article is republished from The Conversation under a Creative Commons license. Read the original article.

The post Orbital Data Centers Are Seductive on Paper, but They Face Daunting Challenges in Reality appeared first on SingularityHub.

AWS Weekly Roundup: NY Summit recap, Local Zone in Hanoi, Grok 4.3 in Bedrock, price reductions, and more (June 22, 2026)

Last week AWS Summit New York City brought together thousands of customers, partners, and builders for a free, one-day event showcasing the latest in cloud and AI innovation. Dr. Swami Sivasubramanian, VP of Agentic AI at AWS unveiled a stack of AI launches in his keynote, all built around one thesis: agents that compound value over time.

  • Agents for working – You can launch autonomous agents and access a smarter activity feed with new Amazon Quick features, which now let you create and run multi-step agents directly in the desktop app and consolidates email, Slack, calendar, and tasks into a single prioritized view with personalized rules.
  • Agents for securing – You can shift from reactive to proactive security with AWS Continuum, a new AI-native security service that reasons, validates, and acts at machine speed across the full code vulnerability lifecycle. AWS Security Agent (now part of AWS Continuum) adds new features: threat modeling; pull request code scanning with remediation across major Git platforms; and IDE integrations via Kiro power, Claude Code plugin, and MCP.
  • Agents for building – You can write, ship, and modernize code in one continuous loop with Kiro, AWS DevOps Agent, and AWS Transform. Kiro introduces a native iOS app; AWS DevOps Agent adds release management capabilities to assess code changes before production; and AWS Transform continuous modernization reduces tech debt autonomously.
  • Agents customers create – You can go from agent idea to production in minutes with Amazon Bedrock AgentCore, which now includes a GA harness for infrastructure and orchestration, Web Search, Managed Knowledge Base, policy integrations with Guardrails, and the new AWS Context service for mapping organizational data relationships.

To learn more, visit the Summit recap from our top announcements blog post and Amazon News post.

Last week’s launches
Here are last week’s launches that caught my attention:

  • AWS Local Zone in Hanoi, Vietnam – This new Local Zone is one of the first AWS Local Zones in the Asia Pacific with support for Amazon S3 and Amazon EBS Local Snapshots, enabling customers to meet data residency requirements by storing and backing up data locally. To get started, enable the Hanoi Local Zone (ap-southeast-1-han-1a) from the Regions and Zones tab in the AWS Global View or by using the ModifyAvailabilityZoneGroup API.
  • AWS Blocks, an open-source TypeScript framework for application developers (preview) – AWS Blocks runs a fully functional local environment with Postgres, authentication, and real-time messaging, no AWS account required. When you’re ready to deploy, the same application code runs on production AWS services with zero changes, and you can drop into AWS CDK at any point for direct resource configuration.
  • Grok 4.3 from xAI in Amazon Bedrock – You can use the Grok 4.3 model on Amazon Bedrock, giving you even more choice as you build generative AI applications across reasoning, agentic, and enterprise workflows. Grok 4.3 runs on a new inference engine in Bedrock designed for price performance, with support for tool calling, structured output, and response streaming.
  • Amazon S3 annotations: attach rich, queryable context directly to your objects – Amazon S3 now lets you attach up to 1 GB of rich, mutable, and queryable context directly to your objects using annotations, purpose-built for AI agents and autonomous workflows that need to discover, understand, and act on data at scale without maintaining separate metadata systems.
  • Amazon ECS announces faster service auto scaling – Amazon ECS service auto scaling now detects and responds to load changes faster with support for high resolution (20-second) metrics and metric publishing optimizations. In AWS benchmarking tests, time to trigger scale-out improved from 363 seconds to 86 seconds (76% faster), and total time to scale and provision new tasks improved from 386 seconds to 109 seconds (72% faster).
  • Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs – AWS is the first major cloud provider to support NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. G7 instances are accelerated by these GPUs with custom sixth-generation Intel Xeon Scalable processors, delivering up to 4.6x AI inference performance and up to 2.1x graphics performance compared to G6 instances.
  • Strands Agents introduces new capabilities – Strands is an open source toolkit for building production agents. You can now use better context management in Harness SDK, a new isolated execution environment with Strands Shell, and chaos testing and red teaming in Strands Evals.
  • AWS Management Console Private Access – You can access the AWS Console from VPCs without internet connectivity, allowing enterprises to manage their AWS infrastructure through the console while maintaining strict network security controls in air-gapped environments.
  • AWS Marketplace Storefront is now generally available – AWS Partners can create and deploy their own branded catalog of solutions and services on their website or application in hours. Channel Partners and Independent Software Vendors can now simplify how they manage their cloud marketplace business and make it easier for customers to discover and purchase their solutions from AWS Marketplace.
  • Palo Alto Networks (PANW) Advanced DNS Security on Amazon Route 53 Resolver DNS Firewall (preview) – You can now enforce DNS threat protections from Palo Alto Networks directly on Route 53 DNS Firewall rules, without deploying separate firewalls or modifying VPC configurations — by subscribing to PANW from the DNS Firewall console through the embedded AWS Marketplace widget.

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Price reductions 
AWS continues to look for ways to increase performance and lower prices for our customers. I noticed a few such efforts last week, so I’d like to share them:

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events as well as AWS Summits and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That’s all for this week. Check back next Monday for another Weekly Roundup!

Channy

Q&A: How KRAFTON Built PUBG Ally, a Co-Playable Character Powered by NVIDIA ACE

25 June 2026 at 16:38
AI companions in games have long been constrained by fixed dialogue. PUBG Ally is a different kind of system. Built by KRAFTON for PUBG: BATTLEGROUNDS, this AI...

AI companions in games have long been constrained by fixed dialogue. PUBG Ally is a different kind of system. Built by KRAFTON for PUBG: BATTLEGROUNDS, this AI teammate is powered by NVIDIA ACE and its suite of efficient models and tooling. PUBG Ally uses automatic speech recognition, a 2B-parameter small language model, and text-to-speech to understand player voice…

Source

Japan Thinks Swarms of Transformer Robots Could Explore the Moon

15 June 2026 at 22:18

A tiny robot developed by Japan’s space agency operated autonomously on the moon for more than 100 minutes and sent a series of images back to Earth.

Exploring the moon’s surface lays crucial groundwork for future crewed settlements, and swarms of tiny robots could be the key. Now researchers have given the first demonstration of the idea after a palm-sized rover autonomously navigated the moon and transmitted images back to Earth.

The moon is a tough environment for robots. Its surface is strewn with craters and abrasive moon dust, and communication delays make remotely piloting vehicles a painstaking and risky process. The cost of launching and landing hardware and the real prospect of losing expensive equipment justifies an extremely cautious approach that can significantly slow down exploration.

One way around these challenges is to replace traditional rovers with many small, cheap, and hardy robot explorers, which could increase coverage and introduce redundancy. And now Japan’s space agency JAXA has given us the first compelling demonstration of the approach.

In a paper published in Science Robotics, JAXA researchers provide a technical report detailing the successful deployment of the agency’s LEV-2 robot during the its SLIM mission, which touched down near the Shioli crater in January 2024. LEV-2 is a three-inch-wide sphere that converts into a wheeled robot after landing. The robot operated autonomously for more than 100 minutes, covering an estimated 24 meters and relaying a series of images back to Earth.

“Although the capabilities of an individual small rover are inherently limited, the results highlight the potential of such platforms as independent explorers, capable of accessing environments beyond the reach of a primary large spacecraft,” the authors write.

Nicknamed SORA-Q—derived from the Japanese words for space and sphere—the robot weighs just eight ounces. Upon arrival, the shiny metal sphere splits open and expands horizontally, allowing its two hemispheres to become wheels that spin around a central shaft. This central area also features a front-facing camera and a tail to help stabilize the robot.

JAXA developed the device in partnership with Sony and toymaker TOMY. The design borrows directly from technology used in transformer toys that convert from vehicles into robots. But the team had to make considerable modifications to account for the harsh lunar environment.

One of the biggest challenges for any lunar robot is maneuvering in the dust, or regolith, that coats the moon’s surface. The fine, powdery material can be hard for smaller wheeled robots to navigate as they lack the traction of their larger counterparts.

To solve this problem, the team designed the wheels to rotate around a point slightly offset from their center, causing a lopsided spinning motion that lifts the rover up slightly on every rotation. This helps the wheels to dig into the surface and generate enough traction to keep moving in the loose regolith.

Communication delays also present a significant barrier to smooth operation, so the team engineered the robot to handle most operations autonomously. An onboard image-processing system allowed the rover to detect the SLIM lander in its camera feed and use this as a navigational reference point, estimating its own position relative to the spacecraft in real time.

Because of its diminutive size, it was impractical to give SORA-Q the equipment needed to communicate directly with Earth, so the team paired it with a hopping robot called LEV-1 that can transmit data. Power constraints and narrow communication windows still cap the amount of data the robot can send to Earth, so SORA-Q has an onboard image-processing algorithm that picks out the best photos to share.

Due to power and mass constraints, the team fitted the robot with a low-power chip designed for small devices rather than complex tasks like image processing. The algorithm relies on a very simple approach—it detects the SLIM lander’s distinctive gold insulating material and then picks the photos where this is featured prominently in the frame.

Around seven minutes after activation, the rover had moved roughly five meters from the lander, selected the two best images from 12 it had captured, and transmitted them to LEV-1. One of those images actually proved unexpectedly useful as it showed the lander had landed at an odd angle with its solar panels facing the wrong direction. This gave ground teams critical information that helped them diagnose the spacecraft’s operational status.

Image SLIM lander taken by LEV-2. Image Credit: JAXA/TOMY/Sony Group Corporation/Doshisha University

But the system wasn’t flawless. It lost some data in transmission, partly because LEV-1’s hopping maneuvers appeared to disrupt the wireless link and partly due to changing antenna orientations as the rover moved. The team also lost telemetry data before the mission ended, making it impossible to determine exactly how far the rover ultimately traveled or when it stopped working.

Still, the mission was strong evidence that small, cheap vehicles like SORA-Q could greatly expand the scope of robotic exploration. That could prove invaluable as we attempt to scope out promising locations for future scientific missions or even permanent bases on the moon.

The post Japan Thinks Swarms of Transformer Robots Could Explore the Moon appeared first on SingularityHub.

Orbital Airbag Could Shield Earth From Devastating Solar Storms

8 June 2026 at 21:56

A planetary defense system would blunt solar storms with hundreds of tons of gas. Emerging heavy-lift rockets could deploy it in under two months.

Extreme space weather could wreak havoc on the satellites, communications networks, and electrical grids that modern society depends on. Researchers have now proposed an ambitious space-based planetary defense system that would weaken solar storms before they hit Earth.

The sun regularly emits massive pulses of radiation, energetic particles, and magnetic fields that interact with the Earth’s own magnetic field. This activity is the source of auroras like the northern lights, but the most violent eruptions can cause geomagnetic storms with the power to disrupt GPS and radio communications and fry electrical equipment.

While the impact of most of these events is limited, there is precedent for more catastrophic outcomes. In 1859, the Carrington Event, the most powerful solar storm ever recorded, knocked out telegraph lines across North America and Europe. In today’s highly electrified world, a similar event could cause between $2.4 and $3.4 trillion in damage to the power grid alone.

Now, researchers at Boston University and the University of Michigan have come up with a potential solution. In a paper published in Space Weather, they propose a constellation of satellites called StormWall that would release hundreds of tons of gas into orbit to blunt the force of an incoming solar storm.

“It’s as if you could install an airbag in the magnetosphere,” co-author Daniel Welling, a space physicist from the University of Michigan, told Science.

Solar storms have the potential to sow chaos because they weaken the magnetic shield protecting Earth from space radiation. Powerful enough storms disrupt the Earth’s magnetic field and cause it to reconnect to the sun’s, allowing energy from the solar storm to pour into the magnetosphere.

The Earth already has a natural defense against this—a doughnut-shaped reservoir of ionized gas, or plasma, sitting just above the atmosphere. When the planet’s magnetic field is disturbed, a plume of this plasma flows toward the sun and slows the rate at which the magnetic fields reconnect.

StormWall would turbocharge this process by releasing massive amounts of artificial plasma into the outer atmosphere. The researchers sketch out a system involving a constellation of satellites orbiting about 22,000 miles from Earth. The satellites would carry canisters of lithium, barium, or sodium gases to be ejected when a large solar storm is inbound. The gases, rapidly ionized by solar radiation, would add to the planet’s natural plasma shield.

Based on simulations, the researchers estimate that releasing around 400 tons of gas could reduce the strength of a major geomagnetic storm by over 50 percent. Crucially, the intervention would be swift and reversible. The plasma cloud could be in position by the time a storm hits, and it would dissipate just a few hours later.

Launching this much material into orbit would be a big undertaking, but the researchers say it could be within reach of emerging heavy-lift vehicles like SpaceX’s Starship or China’s Long March 9 rocket. They calculate that six launches could deploy the full constellation in under two months.

Outside experts have been broadly positive. Allison Jaynes, a space physicist at the University of Iowa, told Science the idea was “highly innovative and appears to be quite feasible in the near term.”

But getting the satellites into orbit is only part of the puzzle. Accurate and timely space weather forecasts would also be a prerequisite. And gaining international buy-in for a system that would drastically alter the near-Earth space environment, even if only temporarily, could be challenging.

The researchers flag potential side effects that need more study, including the generation of electromagnetic waves as the released material ionizes. Still, given the devastation a Carrington-sized event could unleash on the modern world, the potential downsides may be worth the risk.

The post Orbital Airbag Could Shield Earth From Devastating Solar Storms appeared first on SingularityHub.

Open-Source Software Is Starting to Help Robots Think

21 May 2026 at 14:00


When a group of academics started making open-source robotics hardware, a generation of roboticists got years of their lives back. Now, the bigger challenge is getting robots to think—and that’s starting to be open sourced too.

The shift is still early, but companies including Hugging Face, Nvidia, and Alibaba have all made significant bets on open-source robotics in the last two years, releasing tools and models aimed at the higher-level work of getting robots to reason, decide, and act.

The open source movement that accelerated other AI applications is now being applied to the problem of making robots smarter. If these attempts to bring AI to robotics with open-source platforms succeed, the barrier to building a capable robot could fall as fast as the barrier to building an AI application did.

The world ROS built

Open-source robotics software has been around since the mid-1990s, with early projects like Carnegie Mellon University’s Inter-Process Communication package and the Player Project in the early 2000s laying the groundwork. But these were often tied to specific research groups, and the field remained fragmented.

The Robot Operating System, ROS, changed that when it made its debut in 2007. By bundling tools and attracting more users, it became the de facto standard. The story of open-source robotics, in many ways, starts there.

Despite its name, ROS is not actually an operating system. Rather, it is a software framework that sits on top of Linux and handles robotic fundamentals like moving data between components, talking to hardware, building maps, planning paths, and supporting developer tools, such as data logging and visualization. Before ROS, every robotics team wrote that infrastructure themselves. It often took a year or two before a lab could get to the research it actually cared about.

Brian Gerkey, who helped build ROS in the mid-2000s, says he was drawn to the project because of how much open source had already changed the world, pointing out that nearly the entire internet is built on it.

“I’m a tool builder, and I like to share everything as openly as I possibly can, because I think that’s where we get the most impact out of what we build,” says Gerkey, board chair of Open Robotics and now CTO at Intrinsic, a robotics and AI unit of Google.

As it was developing, the AI community largely took the same approach, sharing research, models, and data openly, and the field accelerated faster than almost anyone predicted. Now some of those same advancements are arriving in robotics.

Open-source AI for robotics

Computer vision, once a hard problem, has advanced dramatically in just a few years, says Spencer Huang, Nvidia’s director of product for robotics. What once required significant expertise can now be done in a few lines of code. Simulation tools have become accurate enough to be useful for training, and access to the tooling that once required a specialized lab is now widely available, much of it open source.

“To get into robotics, you no longer need a Ph.D.,” he says. The result is a much larger pool of people who can contribute, and the field is starting to look less like a specialized discipline and more like a platform that anyone can build on.

Nvidia has built out an open-source robotics stack that covers the full development pipeline. Its Cosmos world models generate synthetic training data and simulate physical environments. Its GR00T models give robots the ability to reason through and execute complex tasks. And its Isaac frameworks handle the orchestration that ties training, simulation, and deployment together. Not everyone needs to train the robots from scratch, Huang says, and most people probably shouldn’t.

“If you gate pre-training, the field just never grows,” he says. “We should be able to provide a high-quality, state-of-the-art pre-trained model that anyone can go and take and fine tune for their own purposes.”

All of Nvidia’s open-source models live on Hugging Face, the open-source AI platform that has become the default place to share models and datasets. Hugging Face launched LeRobot, a community platform for robotics AI, in May 2024. Since its launch, the number of robotics datasets on the platform grew from 1,145 at the end of 2024 to more than 58,000 today, making it the single largest dataset category on the hub.

Hugging Face has also moved into hardware, acquiring robotics company Pollen Robotics. The acquisition came from a realization that software alone was not enough, according to Clement Delangue, Hugging Face’s CEO. The goal, as with the software, was to bring more people in.

The contributors to LeRobot include the biggest names in the industry, academic labs, and hobbyists building robots in their spare time. For instance, earlier this year, Alibaba released RynnBrain, an open-source foundation model for physical AI that the company claims outperforms comparable offerings from Google and Nvidia on benchmarks. That diversity of projects, Delangue says, is important.

“It is not just one model or one dataset or one hardware,” he says. “It is a lot of small contributions that everyone can be part of.”

Commercial incentives muddle the field

The stakes, Delangue says, go beyond convenience. A world where only a few proprietary systems control the robots in people’s homes is a concerning one. “Having robots at home that you don’t really understand, that you don’t really control, that a few people in Silicon Valley control is a scary thought,” he says. “Open source gives an alternative path.”

But getting there is not straightforward. The open sourcing happening now looks different from what produced ROS, which emerged largely from academics pooling their work with no commercial stake in the outcome. The biggest contributors today are companies with clear business reasons to want more people building on their platforms. That’s not necessarily a bad thing, says Bill Smart, a professor at Oregon State University, in Corvallis, who was part of the early open-source robotics community. But the incentives are worth being aware of.

He also worries that the lowered barrier to entry has a downside. Researchers coming from AI without a robotics background are sometimes solving problems the field already solved. A newcomer might spend a week training a neural network to move a robot’s hand from one point to another, unaware that the same task can be accomplished with a few lines of code using decades-old techniques. The incentives are not always pointing in the same direction as the progress.

Smart is not without hope though. Whatever the motives behind the open sourcing, he says, the effect is real. More people are in the field than ever before, the tools are genuinely easier to use, and the community is bigger and more diverse than anything that existed when ROS was getting started.

“Anyone can make a robot move now,” he says. “As an old tech guy, that makes me happy and sad, because I’m no longer special.”

Build AI-Powered Games with NVIDIA DLSS 4.5, RTX, and Unreal Engine 5

30 April 2026 at 17:00
Today, game developers can begin integrating NVIDIA DLSS 4.5 with Dynamic Multi Frame Generation, Multi Frame Generation 6X, and the second-generation...

Today, game developers can begin integrating NVIDIA DLSS 4.5 with Dynamic Multi Frame Generation, Multi Frame Generation 6X, and the second-generation transformer model for NVIDIA Super Resolution. In this post, we’ll go over new technologies and resources to share with our game-developer community, including: At CES 2026, we introduced DLSS 4.5, extending its AI-driven…

Source

❌