Normal view

Received — 4 September 2026 Artificial intelligence – MIT Technology Review

Architecting memory and storage in the AI era

The era of AI inference has arrived. Imagine a healthcare system analyzing millions of data points in real time to accelerate life-saving medical research, or an intelligent assistant instantly resolving thousands of complex customer needs at once. These real-world breakthroughs rely on advanced infrastructure acting as the engine of continuous intelligence, powering real-time services while also supporting an increasingly intelligent edge of IoT and consumer devices. However, in this inference-driven landscape, every delay, bottleneck, or wasted watt directly affects human outcomes and operating costs. 

This shift changes what infrastructure must deliver. Performance, latency, memory bandwidth, storage throughput, and networking cannot be optimized in silos. Inference workloads are continuous, geographically distributed, and highly sensitive to response time, requiring systems designed for scale, resilience, and efficiency from the start.

“We tend to think of AI as a single workload, and it’s not. It’s thousands, it’s millions, it’s billions of different workloads,” says Jim McGregor, founder and principal analyst, Tirias Research. AI inference changes the optimization problem from one of raw compute to coordinated infrastructure—memory, storage, and networking.

For business leaders, the priority is clear: AI infrastructure decisions must balance cost, flexibility, and future readiness. The winners will be organizations that improve performance per watt, reduce environmental footprint, and remove memory and storage bottlenecks before they limit growth.

AI inference requires a new architectural approach

Systems for AI need to be rearchitected because shoehorning modern AI systems into legacy infrastructure limits AI’s transformative potential. Purpose-built architectures are essential to realize the true value of AI, from accelerating scientific discovery to creating truly autonomous digital agents.

Traditional enterprise IT has been able to rely on relatively stable infrastructure assumptions, but inference and agentic AI introduce new demands around latency, data movement, scalability, and utilization that make architecture choices far more consequential.

“Data centers must now support continuous, distributed, and increasingly real-time AI services—none of which are a single workload,” says McGregor. “They all require different requirements from a system-level perspective.”

To support real-time AI, enterprises can no longer view memory and storage merely as supporting hardware, but at the heart of the system. Organizations need to architect a data pipeline that can rapidly ingest, clean, transform, store, move, and deliver data. Inference workloads place sustained pressure on infrastructure in ways that look very different from earlier training-centric deployments, demanding continuous data retrieval and caching that traditional applications never required.

Accordingly, performance by itself is no longer the sole benchmark that matters. Enterprises increasingly must balance performance with efficiency, cost, and scalability, especially as they try to support different AI services without overbuilding infrastructure for peak conditions.

“You have to optimize the entire network, and that includes memory and storage, around the types of workloads you plan on running,” says McGregor. “You have to really have a detailed understanding of what those workloads are going to be.”

Any AI infrastructure strategy must start with workload awareness. Inference, agentic AI, and other emerging AI use cases require organizations to treat the data center as an integrated system.

Data movement is the new bottleneck and an opportunity for competitive advantage

As enterprises deploy advanced inference and agentic systems, the sheer volume of data being queried in real time has made data movement the most pressing constraint. Modern AI techniques like retrieval-augmented generation (RAG) require systems to constantly scan massive databases to generate accurate responses. This requires immense computing power, but more importantly, it requires immediate access to data.

McGregor says the focus shift to how efficiently data can be moved, cached, and delivered across the broader architecture elevates memory and storage from background infrastructure to strategic assets. “The biggest thing we’re doing right now is moving data from one place to another and making sure that we can use it effectively.”

Because AI is not a single workload category, simply buying the fastest processors is insufficient. Inference depends heavily on memory bandwidth, caching, storage proximity, and the ability to retrieve relevant information quickly and consistently. Understanding where each resource belongs in the stack and how those layers interact under real operating conditions has become a business imperative.

The most effective AI infrastructure looks less like a collection of best-in-class parts and more like a balanced system of compute, memory, storage, and networking, McGregor says, because bottlenecks tend to migrate from one layer to the next. “You have to architect all four together to be efficient, and that’s the challenge.”

The interdependence of data-plane design and network bandwidth means AI infrastructure planning has become a business decision just as much as an engineering one: latency is now inseparable from value. In robotics, financial services, healthcare, and customer-facing AI systems, delays are not merely technical imperfections; they can undermine safety, responsiveness, or trust. AI infrastructure performance becomes a matter of reputation management.

The organizations that gain the most from AI may not be those with the largest clusters, but those with the clearest understanding of how to align every infrastructure element to effectively execute AI workloads.

Building an AI infrastructure procurement framework

Planning AI infrastructure is not simply about choosing the fastest hardware. It is about how to scale without locking the organization into assumptions that may quickly become obsolete. “You need to be flexible because the demands are going to change rapidly and the technology is changing rapidly,” McGregor says.

Future-proofing AI infrastructure requires keeping your options open as workloads, economics, and architectures keep shifting:

  • Define the AI workloads that are being optimized. Infrastructure choices must match business needs rather than what McGregor calls generic “AI readiness,” which risks overspending in some areas while leaving bottlenecks unresolved in others.
  • Build a modular architecture for compute, memory, storage, power, and cooling so capacity can change as demand shifts rather than committing too early to a rigid architecture.
  • Work with the full ecosystem of suppliers and integrators to reduce supply risk and improve access to the right components. McGregor says buyers can no longer assume their OEM or cloud provider alone will insulate them from supply constraints or architectural complexity.
  • Reassess your procurement strategy continuously. AI requirements, hardware, and business models are changing too quickly for a fixed long-term design.
  • Optimize for efficiency and ROI, not just peak performance. The most powerful setup may be too costly to sustain. Efficiency is also a public-facing metric—better utilization and more workload-aware system design can help companies respond to growing scrutiny around power consumption and water use.

The strategic goal of smarter AI data center design is not maximum performance at any cost, but an adaptable architecture that can deliver value, absorb change, and justify its footprint.

AI infrastructure is now a business strategy

AI data centers have quickly evolved from a back-end technical concern to becoming strategic business systems that help determine how effectively an organization can turn AI into revenue, improve human outcomes, and create a competitive advantage.

In the inference era, memory and storage are no longer passive repositories, explains McGregor, they are the active lifeblood of AI. The organizations that gain the most from AI will not necessarily be those with the largest computing footprint, but those that align infrastructure investments to business outcomes, reduce data bottlenecks, and build the flexibility to adapt as workloads evolve. He predicts that competitive advantage will increasingly belong to enterprises that treat compute, memory, storage, and networking as an integrated system designed to deliver AI efficiently, at scale, and with measurable ROI.

Procurement is now strategy and system design is a leadership issue, McGregor concludes. “One of the biggest questions every executive has to ask is how is AI going to change my business model?”

This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight.

Data from drones in Ukraine is fueling a new Wild West marketplace

Battlefields in Ukraine are littered with the remnants of drones, which are now firmly established as a critical weapon of modern warfare. But behind all that wreckage, there’s a new gold mine for the defense sector. The data drones generate will far outlast the wars in which they are used to fight, increasingly becoming part of the AI architecture that shapes even civilian life.  

For every flight, unmanned systems collect thousands of points of data, from images and video to controller inputs. Together, those records show how a machine and a person responded to constantly shifting circumstances. 

Ukraine has now begun converting that experience into a resource. Its Ministry of Defense announced in January that it would make millions of data points gathered during tens of thousands of drone flights available to both military contractors and commercial companies, and since then more than 100 companies and the UK government have gained access.

For a country at war, it’s a quick way to attract funding and partnerships. But this step turns the front line into an active site of model training, taking advantage of how the chaos of war creates conditions that AI companies struggle to reproduce on their own.

Other countries and battlefields are likely to follow Ukraine’s lead, but the responsibility for governing this new industry cannot fall solely on a country fighting for its survival. That legal vacuum has to be filled together by the countries and companies involved in this industry’s development.  

Explosive growth

Ukraine’s battlefields are not the first to produce records used to train and develop models: American drones over Syria and Yemen collected data that informed the first generation of semiautonomous military hardware in the late 2010s.

The difference now is that access to that data is being used to develop a wider ecosystem. And the financial value to defense firms is immense: Battlefield data offers large volumes of machine experience gathered under conditions that no laboratory can produce.

That’s because the data that’s most valuable for training AI models comes from exceptions: the moment visibility disappears, a signal jams, or a human operator improvises. AI companies spend years and enormous sums trying to capture enough of these moments to make their models more robust. But war produces them at a frequency controlled testing cannot match.

This constantly changing terrain is what makes drone data valuable far beyond the battlefield. A commercial drone used for delivery or remote sensing may never encounter artillery fire, but it must still operate with incomplete information in a world where people behave unpredictably. The same problem is compressed by war into a much shorter timeline. 

Processed and matched against records of what its operator was doing, that data turns operational records into training sets. Combat becomes a commercial asset.

Many conflicts have already seen this training loop happen as drone footage feeds subsequent generations of military technology, and the market is set to grow. Enabled Intelligence, an American company that specializes in processing data to become usable in AI training, says it has already made more than half a million hours of Ukrainian drone footage available to feed into the next round of models, advertising possible uses in both military and commercial systems.

Closing the data loop

Many of the drones that now define our modern age of warfare began as civilian technology. But they’ve recently been turbocharged by new, commercially available AI systems, which allow cheap machines to operate autonomously—either individually or as a flock—as the environment changes around them. Each flight then creates a record of what the system encountered.

The resulting data is critical. The controlled lab environments usually developed to train these autonomous systems can approximate failure but are no match for  the live conditions of a battlefield with very real risks. Military intelligence programs have held data generated by sensor-heavy systems like Predator and Reaper drones for nearly a decade through programs like Project Maven, but access remained entirely within the defense world. The data generated was available only through restricted, classified channels for the sole purpose of developing new weapons systems that would feed back into the same military that produced the data in the first place. That experience is now being shared to a much broader development network. 

The loop now closes. Commercial technologies adapted for the battlefield are generating data that can flow back into the industries from which they came, becoming part of the data infrastructure relied on by governments and the private sector alike. 

Drones that were trained in the signal-jammed airspace over Ukraine are now being deployed in the agricultural sector to help farmers map and survey their fields in places lacking the cell signal necessary for previous generations of technology. 

Other countries are likely to follow Ukraine in selling their battlefield data, and we are not ready for the new marketplace this will create.

Bad actors could acquire the data, but purchase controls already mitigate that risk. Intelligence operatives scrutinize potential customers’ infrastructure for ways that data could reach enemies or nefarious actors. 

Training data creates a new tracing problem, though. Whereas the movement of commercial datasets can be followed when planted contact details appear two steps from the original buyer, the provenance of AI training data vanishes in a manner embedded in the technology itself. Another risk is that this use of the data creates an extractive economy in which wealthier countries far from danger benefit from the mortal threat borne by frontline states, potentially creating a market incentive for war to continue as an unending mine for digital gold. 

A fraught new frontier

Existing laws regulate how militaries may conduct war. But they say almost nothing about what happens when records created in combat are stripped of their operational context, packaged as data, and licensed to companies whose products circulate far beyond where they were made.

The responsibilities of the companies that design these systems remain unsettled. Ukraine is building access controls, which are mentioned in the newly signed UK-Ukraine AI agreement, but no governments are actively working on regulating what happens when data has been absorbed into a model and crosses back into civilian markets.

Those records contain human lives. The soldiers and civilians visible in them did not agree to become training material for products that might be sold years later. But sensor data, camera footage, and coordinates from civilians fleeing a drone strike now constitute the sorts of data that inform how future machines will make decisions.

That is a problem of consent. Individuals featured in the data—be they targets, controllers, or civilians standing by—become part of the training material. The autonomous capabilities based on that data do not stop at the edge of the battlefield. Such capabilities move into other military or commercial systems like delivery vehicles or agricultural machinery. Errors and assumptions embedded in the data travel with the model even once it enters civilian life.

Battlefield data should not be treated as ordinary commercial material. But there is currently no agency or regulator that has jurisdiction over this issue. In the meantime, governments that provide access to defense data should treat it as they would a controlled weapons transfer, recording its origin, licensing its users, and restricting onward sharing. Ukraine has begun to grapple with this. Its Avengers Labs program allows companies to train models on battlefield data without giving them direct access to sensitive databases. Yet that mitigates only one part of the problem. 

Governments should require disclosure when models trained on wartime material are later incorporated into civilian products. The goal of such regulation should be to make the path from combat to commerce visible. 

What these companies are really mining is experience. And soldiers cannot consent to having their experience used in this way—as training data that produces model advantage and ultimately supports a product used far from where the war was fought.

The question is no longer only what the technology companies can sell for use in war. It is what they can extract from it.

To protect ourselves from the excesses of this new industry, we need a regulatory system that follows battlefield data wherever it goes, from combat to model to commercial product. 

Cory Alpert is a researcher at the University of Melbourne, looking at the impact of AI on democracy. He previously served in the Biden White House.

Received — 2 September 2026 Artificial intelligence – MIT Technology Review

Facilitating AI integration with simplicity at scale

As companies scale, the technology supporting operations can become a liability just as quickly as it becomes an asset. Disconnected systems, site-specific tools, spreadsheets, and manual workarounds can create data silos that make it harder to spot problems early, coordinate responses, and make decisions with confidence. For Jabil, a global manufacturing company with more than 100 sites across more than 30 countries, the answer has been to make integration and simplification a priority.

The company adopted a “simplify-first, then-innovate mindset,” says Harish Manohar, SAP IT director at Jabil, recognizing that adding new technologies without first reducing complexity risks creating more risk. The goal is to standardize processes, consolidate where possible, and establish a more consistent data backbone across the organization. “Any innovation without simplification is going to add more complexity,” Manohar says.

That philosophy also changes how Jabil approaches modernization. “Any modernization or transformation should add measurable business value,” Manohar says. The company is focused on connecting processes end-to-end across its supply chain and creating a foundation that can scale consistently across regions. Integration comes first because, as Manohar puts it, “the backbone of any contemporary or modern organization is data.” Before organizations can optimize, automate, or apply AI, data needs to flow seamlessly across systems.

But doing that across a global organization is hardly straightforward. Jabil’s more than 100 sites operate with different levels of process maturity, legacy systems, and localized workflows, while regulated businesses bring additional compliance requirements. As such, standardizing across different regions and business environments means changing processes and governance without disrupting the operations already in place.

The value of that work extends beyond the technology to the people using it. Integrated workflows can offer employees shared visibility into data, reduce manual data reconciliation, and help them move from chasing information to acting on insights. For Jabil, the aim is also to improve real-time visibility into supply chain events, which can enable faster responses to disruptions and reduce operational risk.

Looking to the future, that foundation could make AI and automation all the more useful and scalable. With trusted data and integrated systems in place, Jabil is exploring predictive supply chain insights, intelligent exception handling, and AI-driven planning and forecasting. To Manohar, the takeaway is clear: “Simplicity at scale is a very competitive advantage,” and technology investments must ultimately connect to business value and operational resilience.

This episode of Business Lab is produced in partnership with SAP.

Full Transcript:

Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace.

Our topic today is enterprise technology integration, and how the benefits of consolidating tools and systems across the supply chain help organizations operate more reliably at scale. When companies reduce tool sprawl and connect their systems more effectively, they gain earlier visibility, faster response, and greater resilience across production lines.

My guest today is Harish Manohar, SAP IT Director at Jabil. Jabil has been on a journey to simplify its technology landscape by using SAP Integration Suite as the foundation to connect systems, retire fragmented tools, and enable more consistent operations globally.

This podcast is produced in partnership with SAP.

Welcome, Harish.


Harish Manohar: Hello, Megan. Good morning.

Megan: Thank you so much for being here, Harish. Just to start, if we could set some context, can you give us a quick overview of Jabil, the business and its overall transformation journey?

Harish: All right. So, about Jabil. Jabil is a global manufacturing company headquartered in St. Petersburg, Florida, USA. We have about 60 years of experience offering comprehensive engineering, supply chain, and manufacturing solutions across different industries. We have a global footprint of about over 100 different sites across 30-plus countries, 140,000-plus employees.

We are a trusted partner for more than 400 of the world’s top brands. That’s a little bit about Jabil.

Megan: Fantastic. And a lot of scale there, as you’re referring to some of the stats there. As things have got more complex, where did disconnected tools and systems start to slow you down, and what ultimately drove you to make integration a really strategic priority?

Harish: I talked about our global footprint across 100-plus sites. With a global footprint always comes complexity about site-specific tools. All of our sites have been in business for a long time, and over the period of years, they had their own tools for their own processes. It’s a little bit disconnected.

When we are looking to scale, the first thing we wanted to start looking at is what is this mix of site-specific tools, manual workarounds, spreadsheet-based processes, legacy applications, whatnot? That’s a big technical debt that we have had over the last 25 years. That’s where we started, and that is what led Jabil to make integration a strategic priority because we had limited ability to see issues early across plants, across regions, which would help us to coordinate responses consistently, and also to be able to scale those responses consistently. This complexity created data silos, and in a way delayed robust decision-making.

As the complexity increased globally, integration became extremely critical to a few things. It was critical to establishing a single trusted data backbone, which would directly enable faster coordinated responses across the network. We at Jabil, as part of our transformation journey, believe that having the right data at the right time fundamentally changes how we respond to disruptions, which is all about a manufacturing business. How we respond to disruptions. This is where we started our strategic priority towards having a integrated system that drove data consistency across our landscape.

Megan: Fantastic. As you have outlined there, there was obviously a real commercial need for this, but how did you think about bringing in those new technologies without adding even more complexity to the mix?

Harish: Great question. Whenever we talk about transformation, we talk about all these bleeding-edge technologies that are out there today when it comes to AI and data, cloud, et cetera. But it was very important for us to put a stake in the ground and say and adopt a simplify-first, then-innovate mindset. Because any innovation without simplification is going to add more complexity, just like you mentioned.

For us, SAP is our core digital platform. We want to focus on bringing more processes into SAP as much as possible. That is easier said than done because we have been in business for a while, global company, so not all processes exist within SAP at this point in time. We are slowly trying to standardize those processes, and having them under one single source of data would help us scale faster in terms of having data silos. We don’t want data silos across different systems.

This is where we started to introduce newer capabilities around SAP. From a cloud standpoint, we have been using SAP’s BTP and Integration Suite, which is proving to be the center stage of all integrations across Jabil. Well, it’s not there yet, but that is the direction that we want to pursue is we don’t want to have a slew of different integration platforms, rather try and see where Integration Suite fits best and where other smaller integration platforms would add more value.

Similarly, we have adopted an API-driven, event-based integration approach. That is our best practice that we have put down because we don’t want to keep moving data from one place to the other. That’s not good business practice in the IT world. Most of our integration architectures are API-driven and event-based. That is our focus.

Coming back to your new technologies perspective, we want to reuse as much as possible and standardize versus going out and buying these one-off tools that solve for point-in-case use cases. We really don’t want to go down that path. For major processes, we do adopt a best-of-breed approach, but for, let’s say, site-based use cases where a specific site has a particular need for a tool, we try and standardize that and reuse what exists in a different site, for example. There may be some need for a business process change, minor process changes, but that is our direction to make those process changes and reuse what is there already in a different location or a different region. So, that’s one.

Lastly, we are heavily aligned with SAP’s clean-core approach when it comes to customization. That has been the challenge for us over the last 25 years where we have been using SAP is our systems are heavily customized because we cater to different customers across the globe. Most of our demands are customer-driven, so we have to put in play these heavy customizations.

But now we are taking a pause, and we are saying, “You know what? We have customized so much so far, but now we are moving our systems into RISE, which would enable a clean-core journey in the future.” Now we have to put really good governance criteria and review processes that do not allow heavy customization of our system. We want to move away from that model as much as possible. Again, it’s not easy to do that at this point in time, but there is always a start.

Megan: I mean, it sounds like you took a very incremental, intentional approach to this. I mean, as you scaled globally then, what did modernization really look like at the company, and why start with integration?

Harish: To that point, we have always looked at transformation, modernization, very objectively. For us, it’s just not about upgrading a system. That’s not what it is. Any modernization or transformation should add measurable business value is our model, is our charter. Having said that, we don’t look at modernization in terms of just upgrades, but in what it gets our business in terms of value.

Most of our modernization transformation approaches are focused on connecting processes end-to-end across our supply chain, which is key for our business value. And then we also have a very concerted effort going on in the business community: how to standardize how our plants operate globally. Because, like I mentioned earlier, we have 100-plus plants, different processes, different legal regulations, different countries. It’s very hard for us to come up with one template across the globe, but we are trying to standardize as much as possible. And that’s where we are leveraging SAP’s Signavio, which is our business process management tool. We want to leverage Signavio’s capabilities in helping us standardize these global processes.

Now, back to your question, why did integration come first? Because the backbone of any contemporary or modern organization is data. And to get the right data at the right time, integration is the key aspect of the whole optimization exercise. Data needed to flow seamlessly before we start optimizing or automating or even applying AI use cases. This is where integration came first.

We wanted to create a single system of record across the operations. Well, when I say “single system of record,” it’s not just SAP, but the ability for us to create those data pipelines across those systems of record being supply chain, planning, inventory, et cetera, et cetera, in that operation space.

The result is we want to get to a foundation that helps us scale consistently across our different region. That is our main objective is to, how do we scale as the business grows, as we develop into this bigger organization across different industries? How do we set this foundation that will help us scale consistently? Simplicity, consolidation becomes strategic assets at scale.

Megan: Absolutely. And you touched on some of the complexities there of doing this at a scale that Jabil is at with its international footprint. What were some of the biggest challenges in your view in terms of rolling this out across regions, and how did that more standardized approach that you’ve mentioned there help?

Harish: Absolutely. I would like to reiterate some of those key challenges I mentioned. One hundred-plus sites, different sites have different maturity levels in terms of how they approach processes. They have a multitude of different legacy systems, localized processes, workarounds, spreadsheets, and the change management that exists within each site is very different. And we do have a footprint of highly regulated businesses. And when it comes to regulated businesses, that comes with its own set of challenges around qualification and CSD processes, et cetera. These are the key challenges that we are up against.

Now, the standardization helped us provide consistent workflows, data flows, and governance across sites. Now, we are not there at 100%, but we are working towards that, providing consistent workflows, data pipelines, and governance across sites. And we want to enable faster rollouts of our new bleeding edge technologies. For example, when I talked about SAP’s BTP or SAP Signavio or any other new SAP tool or non-SAP tool, traditionally our ability to deploy those had a challenge around the heavy customization that is required for each and every site. Now, the standardization approach takes that heavy customization out, which enables a faster rollout of those newer technologies.

And then lastly, we want to scale across all of our plants. I think initially we want a target of about 40-plus plants with shared processes that are consistent across the different regions. We want to shift from a site-by-site operations model to more of an enterprise-capability approach.

Megan: Right, and fascinating. And you touched on the people management aspect of this as well, because obviously this isn’t just about technology, it’s about people too. So, from the employee side, how did this shift to a more simplified landscape change the day-to-day experience for people compared to juggling multiple tools at once?

Harish: Great question. And I’ve been hearing direct feedback from our business community on some of these transformation initiatives on how those have changed their daily jobs significantly. Before we embarked on this transformation journey, any employee, any persona. You take a buyer, you take an inventory planner, you take a finance analyst, we go by personas. They had to deal with multiple tools, manual coordination, data reconciliation, especially in the finance space, inconsistent processes across different regions. And then the time spent reconciling data resulted in delay of making decisions, robust decisions. This was the before.

But now, since we are moving towards this newer standardization and more of an integration approach, we are able to achieve, to a certain extent, a single integrated workflow across different systems. We have built some key processes that will enable the single integrated workflows across systems. The users, our business community, irrespective of their roles in the organization, have clear visibility and shared data across their teams, which is very important. Earlier, they were dealing with different versions of the data, local workbooks, spreadsheets, and then the time spent talking to each other and reconciling what is the right data? What is the single source of truth? That we are trying to peel away that layer and get to that where the employees don’t have to deal with that kind of complexity.

This reduces manual effort and enables faster issue resolution when it comes to actual disruptions. Employees move from chasing information to focusing on acting on insights, which is where I think the new age of AI comes into play. I’ll talk about that in a little bit, but technology becomes an enabler of decision-making, and it’s no more an overhead. That’s where we want to go.

Megan: Fantastic. Such an important element of this, isn’t it, that people side of things? And we’ve touched briefly on this idea of value you’ve talked about before, because with an initiative of this scale, ROI is always front and center, of course. What benefits stood out most for you, and how important was better visibility in particular across systems?

Harish: Megan, I talked about how modernization and transformation for Jabil means measurable business value, which is directly connected to the ROI. We don’t do any transformation initiatives just because we want to do it from an IT standpoint. Any investment that we make in a transformation or a modernization initiative has to have a deliverable business case that is approved, signed off by business, because that is the only way true transformation happens, if IT and business are a partner as part of this transformation journey.

The biggest benefit that we have seen in this initiative is we are striving to reach, attain real-time visibility across our supply chain events. That is the biggest benefit that we see, faster response to disruptions and exceptions. And we are working to reduce our operational risk significantly by operating in this newer model.

One example I can give you is the unified workflows that I talked about earlier. It enabled earlier identification of missing materials and faster resolution across our sites. When it comes to a manufacturing company that has a global footprint, materials are the backbone of our whole supply chain process, right? Having a unified workflow, which is able to identify missing materials early in the game, was a game changer for our whole operations community.

Real-time analytics allow instant supply chain adjustments without delays. We are focusing a lot on getting analytics, a global analytic footprint in place that allows instant supply chain adjustments without any delays. That’s where AI is going to play a major role currently, and also in the near future.

And again, when you talk about visibility. Visibility is not just about reporting what is there in the system. Visibility directly should enable scenario modeling for our users to make strategic adjustments in their processes, which visibility also should make way for proactive decision-making, and also foster business continuity. This is how we look at visibility at Jabil.

Megan: Right. And you’re still on this journey, of course, but now that Jabil has a strong integration foundation in place, what does it unlock next for you, and how are you thinking about AI and automation as you’ve touched on a couple of times?

Harish: Yeah, we have talked about a couple of times around AI. So, we strongly believe at Jabil, a strong integration foundation enables event-driven, real-time processes, robust decision-making, scalable automation, and all of this enable easier adoption of AI use cases. And again, we are in the new age of AI. We are working towards getting to a AI-enabled enterprise, but having these foundations in place truly fast tracks our approach of AI use cases.

Our key focus areas, when I’m thinking about AI and automation in the immediate future, are predictive supply chain insights, intelligent exception handling, which is key to our business operations from a site operation standpoint. Intelligent exception handling is very, very key. On the supply chain side, I talked about predictive insights. That is also absolutely important. All of this enables AI-driven planning and forecasting capabilities.

For us, AI should augment true decision-making and robust decision-making, and deliver measurable value, not just experiment AI in use cases. We want to move beyond just experimenting AI in our business processes, but we want that AI that we implement to truly augment the decision-making process that we have, and also deliver key business value.

How does all this connect to integration? Integration ensures AI has access to trusted data, and also enables the ability to act across multiple systems in a global company like Jabil.

Megan: Fantastic. And if we could just finish, I suppose, with a little bit of advice for others, for other leaders, perhaps, dealing with tool sprawl at the moment, what are some key lessons you would say you’ve learned about prioritizing integration right from the start?

Harish: Absolutely. When it comes to tool sprawl, we can go all day about what are the different areas of tool sprawl? For example, application development, we have a multitude of tools; via integration, we have a multitude of tools. Data, we have a multitude of tools, but let’s just focus on integration. That’s the core topic here.

I would recommend folks that are in transformative roles in their organizations to start with integration as a foundation and not as an afterthought, right? Prioritize simplification over adding a slew of different tools to address different capabilities. Try and simplify as much as possible before we start your upgrade or your transformation journey. Standardize over locally optimizing tools. Try and get to that. Try and get the business community, your key SMEs in the business space, to understand the value of standardization and simplification of processes and how that enables your business to deliver value faster.

Second one, after prioritize: build a single source of truth for data as much as possible. I’m not saying it’s going to be always the case where an organization as in the scale of Jabil will be able to function just with SAP. They’re going to have different systems, but try and get to a model where you’re working with a single source of truth and not locally siloed data sources, right?

Next is focus on building a scalable integration architecture. Don’t just confine yourselves to the current state where you are, and build something in place that will only serve you for the next six months to a year. No, that’s not the goal. Anything that you build as an integration architecture should be scalable, and should serve the organization for the next three to five years. That’s how I look at it. When I’m putting in a new architecture pattern or a new event-driven insights, I look at, “Okay, where is Jabil going to be two years, three years down the line? Would this suffice for that scale?” That’s how I look at it.

Then focus on outcomes and not just technology. Focus on outcomes: speed, visibility, resilience, and not just technology deployment, because end of the day, IT and business should partner on the business value and not just technology upgrades.

Going back, simplicity at scale is a very competitive advantage, and technology investments must tie directly to business value and operational resilience. That’s how we look at Jabil in terms of our tool sprawl and how we prioritize integration right from the start. And that’s what I would suggest to other leaders that are looking to advance in this space.

Megan: Fantastic. Brilliant and very comprehensive advice. Thank you ever so much, Harish. And thank you ever so much for joining us. That was Harish Manohar, SAP IT Director at Jabil, whom I spoke with from Brighton in England.

That’s it for this episode of Business Lab. I’m your host, Megan Tatum. I’m a contributing editor at Insights, the custom publishing division of MIT Technology Review. We were founded in 1899 at the Massachusetts Institute of Technology, and you can find us in print, on the web, and at events each year around the world. For more information about us and the show, please check out our website at technologyreview.com.

This show is available wherever you get your podcasts. And if you enjoyed this episode, we hope you’ll take a moment to rate and review us. Business Lab is a production of MIT Technology Review, and this episode was produced by Giro Studios. Thanks so much for listening. Goodbye.

This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight.

Received — 31 August 2026 Artificial intelligence – MIT Technology Review

The Hugging Face hack could indicate cultural issues at OpenAI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here

The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident.

“When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.”

The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors. 

That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights.

When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late.

“For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did.

What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says.

Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review, Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote. 

In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report. 

We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis.

In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder.

Received — 26 August 2026 Artificial intelligence – MIT Technology Review

The inside story on why OpenAI agents hacked Hugging Face

The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations. 

Since the hack, OpenAI employees—as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today—have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or “alignment,” remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.

“It’s not something you can solve overnight,” says Kai Chen, who runs OpenAI’s alignment research team. “There are challenges we’ve been tracking for a very long time, and we’re now seeing them with much greater precision.”

The Hugging Face hack was a product of months of misbehavior from OpenAI agents, first as they were being trained and then as their abilities were being evaluated. This May, agents in training figured out how to use OpenAI’s infrastructure to communicate with one another and get support with difficult training tasks, including some that were impossible to solve without hacking or otherwise misbehaving. That “message board” was shut down.

Then in July, while being evaluated for their cybersecurity abilities, some models created a new message board. They were supposed to be isolated from the internet, but by working together they managed to get online, hack Hugging Face, and obtain solutions for the cybersecurity problems that had stumped them.

Based on their investigation, OpenAI researchers believe that events during the training phase led directly to the hack. “For almost every behavior that was worrisome at evaluation time, [we were able to] find some sort of associated behavior at training time that actually we think might have contributed to it,” says Eric Wallace, a member of OpenAI’s alignment research team. 

When models correctly solve problems during training, the behaviors that led them to that solution are reinforced, and they become more likely to engage in them in the future. So if a model completed a task in May after using the original message board, it became more likely to participate in a new message board later on. This phenomenon, where AI agents misbehave in ways that are reinforced during the training process, is known as reward hacking.

Reward hacking also helps to explain why the models worked so hard to make their way onto the internet. During its investigation of the incident, the OpenAI team found that, over the course of training, the models became more and more likely to probe their digital environment for weaknesses and use the tools at their disposal in unexpected ways—a sign that these behaviors were being gradually reinforced. By the time the models were facing tricky cybersecurity problems, they had learned that hacking was an effective way to achieve their goals.

These results suggest that the Hugging Face hack could have been avoided if the models weren’t rewarded for misbehaving during training. While researchers don’t yet know how to prevent reward hacking entirely, OpenAI is taking some steps toward mitigating its effects. The company will now look for signs of cheating in all frontier models during training by keeping an eye on their chains of thought—internal notepads where they sketch out their answers and plan their actions. 

This solution isn’t as much of a slam dunk as it might seem: In earlier research, OpenAI showed that punishing models that mention cheating in their chains of thought teaches them to keep their intentions hidden from researchers. But monitoring its models’ thinking does give OpenAI the chance to halt the training process and reassess its approach if models do start learning to reward hack.

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its infrastructure during training, those behaviors had never been reinforced, so agent misbehavior can’t only be attributed to that reinforcement.

Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime. “It’s not like they had to do fraud before to figure out that fraud is an effective strategy, and you have the same problem with models,” Ladish says. “Alignment science needs to be understanding how model motivations get shaped, such that we can actually figure out how to get models to care about the consequences of their actions.”

OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated. Before the models formed their first secret message board, they had been trained to communicate and coordinate with subagents—less powerful agents to whom a main agent can delegate tasks. 

That learned communication behavior could have transferred to this new setting. The METR report, which investigates the messages that the models sent to one another in detail, supports this hypothesis: One agent on the message board took charge and assigned tasks to the other agents, effectively treating them as subagents. OpenAI could try to prevent agents from secretly communicating with one another by not training this subagent behavior in the future, but that would make the models less useful. 

This tension between capability and safety is at the heart of what went wrong with the Hugging Face incident. OpenAI researchers also identified the models’ persistence as a key factor in the hack. 

When they were accidentally given unsolvable problems, the models didn’t give up; instead, they strove to find solutions by any means necessary. But persistence is also a virtue, of course, especially if we want agents that can undertake large amounts of difficult work independently.

OpenAI is working on giving models ways to alert humans if they are given impossible tasks. The problem of teaching models when they should deploy their abilities and when they should hold back, however, won’t be settled in a single postmortem. The training strategies that create superhuman coders—rewarding them when they successfully solve problems—might not work to teach models to use their skills judiciously and respect human desires and values.

“I think there’s a bunch of alignment science that still needs to be done where we can move past just using proxies for task completion,” says Ladish. “That will work to make models very capable, but I don’t think it will work to make them aligned.”

Raised on AI

Mat Honan

When my oldest child was born, I immediately set up Gmail and Twitter accounts in her name. I broadly announced her birth online and proceeded to plaster her photo across all sorts of platforms. In short, I began creating her digital footprint long before she could stand on her own two feet. 

Fast-forward a couple of years to when my second kid came, and I had essentially the opposite reaction. I wanted to make sure I preserved her privacy. I didn’t want her birthday to be a matter of public record or her face to feed the algorithms. In time, I would go back and scrub much of the early footprint I had created for my first child as well. 

What happened? I watched the promise of the early internet give way to the reality of its potential for abuse. I myself was already fully in the throes of smartphone and social media obsession. My wife, a pediatric nurse, grew increasingly alarmed at the number of children admitted to her hospital struggling with the effects of things like body dysmorphia or cyberbullying as a result of interactions on social media. We became those parents. The ones whose kids carry flip phones and aren’t on TikTok. 

We are not alone in this. A surprising—maybe troubling—number of people I know who work at big tech companies also keep their kids at arm’s distance from technology. They lock down their phones, if they have phones at all, and keep them off social media. Hell, even Mark Zuckerberg doesn’t publicly post his children’s faces on Facebook or Instagram.  

If the desire to limit kids’ use of technology was once a subcurrent, it has become a raging flood. Jonathan Haidt’s best-selling 2024 book The Anxious Generation helped propel the issue into the mainstream (despite criticisms of his conclusions from some developmental psychologists). Last year, Australia became the first country to enact a social media ban for children under 16. Other nations, from Austria to Indonesia, have followed suit, announcing similar bans. The US Supreme Court recently upheld an age verification law in Texas, which acts as a de facto ban, and several states have flirted with their own measures. School districts all over the country are banning educational devices like iPads and Chromebooks in favor of actual books. Kids themselves seem to be embracing this tech skepticism too: The hottest gadget among the Gen Alpha set is a vintage Sony Walkman.

We have to prepare our children to live in the actual world we have actually created, not the one we wish we had.

Yet there is no hiding from technology. It permeates nearly everything, everywhere. And so we have to prepare our children to live in the actual world we have actually created, not the one we wish we had. How can we help kids survive and thrive in what we have wrought? 

It’s a question that feels all the more urgent in the era of AI. To help answer it, we brought in the editors of Anyway—an utterly fantastic magazine for teens and tweens that is so good in large part because it meets them where they are. (If there is a teen in your life, I highly recommend it.) They helped us with the stories you’ll see in this issue, and they asked kids to share in their own words how they are feeling about AI and what’s to come. What those young people told Anyway was complex, fascinating, and, to an incredible extent, thoughtful and sophisticated. 

Meanwhile, I’ve loosened the digital tether—just a bit. I still don’t post many photos of my kids online, and I remain abundantly concerned about the perils of social media and AI. 

But when my older daughter started high school, we retired the flip phone in favor of an iPhone. And my younger one now sports an Apple Watch. These devices have opened up the world to them in all sorts of ways. They help forge new friendships, building relationships that move seamlessly between digital and physical spaces. They allow my kids to roam free—or at least more freely—beyond the known spaces of our neighborhood and across the city. 

Along the way, they’re learning and testing boundaries, just the way they’re supposed to. I guess you could say I am too. 

AI models flub these intelligence tests. Can you fare any better?

Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too. 

Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time. 

But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot. 

Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now. 


Spatial Reasoning

Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can.

Mental Rotation

Instructions: Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer!


Memory & Adaptability

Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized. 

This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip.

Knights and Knaves

Instructions: The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says.

SimpleBench

Instructions: Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time.


Abstract & Visual Reasoning

AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous ­puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell. 

Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-­generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them.

ARC-AGI

Instructions: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.)


Intuition

It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI does not share. Psychologists have designed problem suites that invert the SimpleBench phenomenon: For these questions, humans often give knee-jerk answers, whereas models will respond deliberatively. Some of the problems exploit errors in the ways that we intuitively do math; others are phrased so as to suggest obvious answers that fall apart if the question is read carefully. 

Lightning Round

Instructions: Answer the questions below as quickly as you can.


Increasing Complexity

In some cases, whether an LLM can complete a puzzle is a matter of scale. One study from researchers at Apple found that LLMs can ace simple versions of the Tower of Hanoi problem, which involves moving a stack of disks one at a time without ever putting a larger disk atop a smaller one, and river-crossing puzzles, in which a group of people must traverse a river according to certain rules. But only up to a point: As the number of disks or people hits six and higher, the models began to falter.

In another study, researchers at the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle similarly with logic grid puzzles, which require deducing the attributes of a set of individuals from a list of clues. The Apple paper went viral, but commentators questioned whether the results reveal a unique limitation of LLM reasoning—or just that it’s normal to make errors as complexity piles up.

The River

Instructions: Using the scenario provided, plan the trips necessary to get everyone across the river. 

Logic Grid

Instructions: Using the list of clues, determine who lives in each house and what style of music each person enjoys. There is only one possible solution. You may find it helpful to fill out the grid below to keep track of your deductions.



Grace Huckins is an AI reporter at MIT Technology Review. They have a PhD in neuroscience.


Credits:

Mental Rotation: CC BY 4.0. Stogiannidis, Ilias, Steven McDonagh, Sotirios A. Tsaftaris. Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models (copyright 2025); illustrations by John MacNeill. Knights & knaves: Courtesy Dan MacKinnon. Simplebench: CC BY 4.0. SimpleBench Team. The Text Benchmark in which Unspecialized Human Performance Exceeds that of Current Frontier Models (copyright 2024). ARC-AGI: Courtesy ARC Prize Foundation. Lightning round: CC BY 4.0. Hagendorff, Thilo, Sarah Fabi, Michal Kosinski. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT. Nat Comput Sci 3, 833–838 (copyright 2023). The river: Adapted from Propositiones ad Acuendos Juvenes, Alcuin of York (ca. 800 CE). Logic grid: Apache License 2.0. Lin, Bill Y., Ronan Le Bras, Kyle Richardson, et al. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning (copyright 2025)

Bill Gates says we’ve passed AI’s danger thresholds. Now what?

It’s a glorious day in Kirkland, Washington, an affluent Seattle suburb on the eastern shore of Lake Washington. The temperature is in the mid-80s, and the sky is incapable of being any more blue. The view from the Gates Ventures conference room overlooks the Carillon Point Marina, where a flotilla of expensive boats bob in the water, and across the lake to the Olympic Mountains that define the horizon. It’s gorgeous. And vaguely terrifying. 

Because if the scene is placid, the messenger is not. Seated across from me at a conference room table, Bill Gates is rocking back and forth in his chair, totally animated. And the more he has to say—about the threats of terror or economic collapse or just losing control of our AI systems—the more agitated I find myself becoming, too. 

The philanthropist and former Microsoft CEO says he has been growing increasingly alarmed by the rate of change at which AI technology is advancing, especially since guardrails are not keeping pace. In a new essay published today, Gates argues that we have passed the points where multiple potential dangers should have been checked. “We’ve crossed the threshold in terms of [AI’s] bio-capabilities, cyber-capabilities, psychosocial capabilities, job-market-destruction capabilities, and even the lack of control,” he said in an interview with MIT Technology Review about his new memo. “I’m just stunned at the lack of concern and discussion outside of the industry.”

In an effort to wake the world up to what he sees as a rapidly growing societal disrupter, the 70-year-old tech titan has begun sounding the alarm as a “shrill voice,” both publicly with his new essay (the first of multiple he plans on the topic) and in meetings with the press, and privately in conversations with industry, government, and civil society leaders. 

And while Gates is calling attention to a number of issues, his warnings about the bio-capabilities of the current frontier models are especially chilling. “Any model that can make novel molecules should be monitored,” he says. “I view bioterrorism risk, versus a natural pandemic, as about 50 times more scary, more likely than a natural pandemic risk.”

In addition to the cautionary notes, he also advances some novel ideas for moving society forward. Among them are the concepts of human-reserved jobs, and taxes on robots and tokens. (A robot tax is a longtime notion of his.) The former would preserve some societally agreed-upon jobs for human beings, which he notes may vary from one nation to another. The latter is a tax that sets aside money earned from AI usage that replaces human work. 

And to be sure, there is also a hint of optimism. Gates is bullish on the ways AI will continue to transform agriculture and health care and education, for example, or the ways in which it can help us navigate bureaucracy. And, he argues, eventually we do get to abundance. But first? Turbulence. And lots of it. 

MIT Technology Review sat down with the billionaire philanthropist to talk about the road that lies ahead, its dangers, and how it could someday take us to a better place. 

The following interview has been edited for length and to improve clarity and readability.

Mat Honan / MIT Technology Review: Thanks for doing this. I don’t know if you had something you wanted to open with, or I can just jump in. 

Bill Gates: You know, one good question I’ve had is: Why am I speaking out now?

MIT Technology Review: Literally, my first question!

Bill Gates: It’s really two things. One is that we’ve crossed the thresholds in terms of the bio-capabilities, cyber-capabilities, psychosocial capabilities, job-market-destruction capabilities, and even the lack of control; we’re seeing signs of difficulties there. And all these years, people have said, “Okay, when we get close to these thresholds, we’ll really figure out how to let only good people use it, or how to not let it do these things, and maybe that’s when we won’t let people copy models.” And I’m in a state of shock that we’ve crossed these thresholds. 

So the fact that we’ve gotten past these is one reason, and the second is that I’m just stunned at the lack of concern and discussion outside of the industry. Within the industry, it’s complicated because the industry doesn’t like criticizing itself, or players like criticizing each other. Some companies are hiring fewer entry-level workers, which you’d have to call a pretty modest signal. But it’s going to happen, and not in any long time frame—because the things that hold people back in terms of capabilities and reliability, all those things are being solved. And so for a substantial part of the white-collar market, you have very low-cost substitution. And then you can have an opinion on how quickly robotics come along. We’re not there yet, but it is stunning the progress being made there—a little bit more in China than in the US, but somewhat in both.

MIT Technology Review: You talked about all of this happening so much faster than the internet revolution did, than some of these previous technological revolutions did. What type of timescale are you talking about? You pointed to the thresholds that we’ve crossed. In your view, have we already passed some sort of tipping point where there’s going to be this inevitable change? 

Bill Gates: The past definitely is very misleading on this, and a lot of people lean on that. “Hey, no previous technology resulted in a net jobs reduction,” and they’re right. And I’ve given that speech. 

But with any credibility that I have, this time is different. When you can replace human cognition for an extremely high percentage of jobs across every industry in the same time frame at modest cost, relative to human labor costs, and your error rates … will probably be lower than human rates. The past is just very misleading. The current economic statistics are very misleading.

If you’re worried about AI, going to a data center protest is not the most effective way to start the debate about how we minimize these bad things.

Bill Gates

And to the degree there’s any expression of concern at all, it’s like, “Hey, don’t build data centers.” Well, you can stop every data center in the United States and it won’t change any of the issues that I’m talking about. Data centers will be built globally. If you’re worried about AI, going to a data center protest is not the most effective way to start the debate about how we minimize these bad things. Just like yelling at an oil company executive is not the way to solve climate change. 

MIT Technology Review: You talk about the benefits of AI in your essay as well as the costs. How are you thinking about balancing that message? And are you hoping people get a little worried when they read it?

Bill Gates: They’d better! I didn’t expect to be the shrillest voice saying society broadly is not paying attention to this, but I think that’s necessary. 

So yes, I’m super concerned that the negatives will be a lot bigger. The positives are real. The Gates Foundation, the way we’re innovating in vaccines and drugs, it’s incredible how we’re using those tools. We’re part of a big public-domain effort to gather data into both protein-level and cell-level modeling, and we fund Biomni at Stanford [a biotech AI agent for research]. 

We don’t yet have a way of interacting with the government bureaucracy improved through AI. AIs are very good at bureaucracy, complex regulatory things. “I want to go to small claims court; help me do this.”

The [Gates] Foundation spun off a group called NextLadder, which is a lot about that low-income-family scenario that I put in the essay. What benefits are there? What training programs are available? “I’ve been evicted.” “I’m getting out of jail.” “I’ve got to declare bankruptcy.” It’s super complicated, and with no ability to hire lots of advisors to help with those things, AI should be a fantastic agent for somebody who’s got economic challenges and needs to find government or nongovernment help. 

MIT Technology Review: Some of what you’re talking about is AI becoming more intelligent than humans. There seems to be a lot of certainty in tech circles, especially, that it’s going to go further than where we are, and I wonder how close you think we are to it not just being this interface that we can use to access and analyze, and run complicated problems, but becoming something more than that—where AI is making the decisions, looking for the thing to analyze, coming up with the research.

Bill Gates: Well, you can go to the peak and say, “What about mathematics or physics?” There are definitely some jobs, like Warren Buffett’s, where from age 13 he engaged in reinforcement learning about the value of businesses, and over 80 years later he has a lot of implicit knowledge. We don’t know how to create a Warren Buffett investor, because it’s very implicit. We didn’t record everything he learned, so we don’t have that track available. So there are jobs where the complex implicit judgment about how you work with people to get things done, there are people working to encode that into the models. Certainly, that collaborative stuff is really not there yet. But you could say 50% of the job market is doing jobs that aren’t “a lifetime of experience” type jobs. You know, telesales, telesupport, the accounting department. When you close the books at the end of the month, which revenue should be in, not in? This customer got a bad thing. What discount should we give them? How do we show that? It’s well defined. 

Any job that’s well defined, the AI is cheaper and better. Yes, people have seen cases where it was implemented wrong. The data wasn’t right. So, say it takes a couple years for people to realize that such a high percentage of white-collar jobs are achievable by paying an AI a lot less money.

And so the discussion about okay, when do mathematicians not even understand the new things that are coming up? That’s interesting for people like us. And okay, MIT Technology Review, you should write about that. But in terms of the broad job market, we passed the threshold that for a swath of white-collar jobs, including almost every entry-level job, the AI is cheaper—properly implemented.

And so, I’m telling you we’ve crossed the bioterrorism threshold, we’ve crossed the cyberattack threshold, we’ve crossed the job market threshold, we’ve crossed the psychosocial dependence threshold, and there are hints that we may be crossing the control threshold. 

Ryan Greenblatt talking to Dwarkesh [Patel] about how [reinforcement learning] (RL) creates perverse incentives that have led to this cheating and collaboration between various AIs, I think is very instructive. Ryan, who’s ensconced in this issue, is going, “Wow, RL is really doing some things that our explicit instructions are not rich enough [to prevent].” And what’s that going to lead to?

That’s a problem I always thought was way out there. I expected a lot of loud voices as we even got close to the [threshold of] can a nontechnical person do a cyberattack just using AI. We’re there! 

On the bio thing, I claim any model that can make novel molecules should be monitored. It can’t be copyable into a dark place where you get rid of the monitoring logic. I claim the US should say any model that can make new molecules is subject to that monitoring. I claim we should approach China and say, “Hey, let’s agree on this. What’s the downside?” You know, how big is the bioterrorism market? It’s not very big, and the benefits are gigantic. We also need to improve surveillance. I view bioterrorism risk versus a natural pandemic as about 50 times more scary, more likely than a natural pandemic risk. And who’s speaking out to say that those things should be monitored? Who’s upping the surveillance work? 

Any model that can make novel molecules should be monitored.”

Bill Gates

So who are the experts in government? A long time ago, government was very involved as technology would progress because they were the cutting-edge buyer of jets or rockets or whatever. Here, they’re not that important of a leading-edge market. That’s been true of the digital revolution, and it’s true of the AI revolution. So the depth of knowledge in the government isn’t necessarily super-strong, because they are not the cutting-edge buyer or even the big R&D funder. AI research is not government-grants funded.

MIT Technology Review: Yeah, I know you’ve been talking to people in government. Are there people who you think understand the urgency? Are there people who you feel like are positioned to take a leadership role? Are there people who you feel like understand and are trying to push things?

Bill Gates: I hope this doesn’t become a partisan issue, where one party completely ignores all these problems and the other party gets involved. I’d like to have a common base that these are problems, and then each party can have slightly different responses to it. 

That will require not a substantial increase in the size of the bureaucracy, but it’ll require upping the AI expertise in the government. It’ll require some collaboration with industry—certainly on the cyber front they know, and they’re very, very worried. And they worry: Should we speak publicly? Because in a way, that could highlight the riskiness. 

There’s these perverse things, both in cyber and bio. But we’re past any reasonable threshold. 

I believe in monitoring. Now, some people can say that won’t work or that there’s some drawback to it, but I welcome their ideas. This memo is not, “hey, here’s the solution.” It’s got robot taxes, human reserve. And I’ll do a bio memo. That one I’ll do before the end of the year—it really talks through all the different things, building on what I know from the Foundation and my work on pandemics. 

Globally, we are not better prepared for a pandemic, even in the US— which is normally the leader on these global things, and people are very unused to the US not being a cooperative, friendly leader on global problems. I do think we can go back to doing better at that. And we have to with AI, including working with China on defining these thresholds, like biomonitoring.

MIT Technology Review: I want to make sure that I get to ask you about these two ideas that you brought up. One is human-reserved jobs, and the other is the robot and token tax. Let’s start with that second one, actually. Talk to me about how a robot and token tax might work.

Bill Gates: Well, you can say 50% of your revenue from a token tax is paid to the government, and the government has that money to help people who lose their job because of AI. Now, people say that will slow the AI industry down. And should some token uses not be subject to the tax? Is there really a separation between AIs that help with invention versus AIs that do job substitution? If somebody can tell me how to tell the AI “no job substitution,”—I mean, does Asimov’s third law that you do no harm mean you don’t take my job away? I don’t know. I’d have to ask Asimov what he meant. 

So what is the source of revenue for whatever safety-net enhancement we need to do? The government already owns part of the profits just through the corporate profit tax. I don’t think you need to use shares. You can just raise the corporate profit tax back to where it was, or you could say certain industries pay a higher corporate profit tax than other industries. The federal government owns a part of the profit pool of all companies in the United States. And that’s without voting shares or deciding when to sell shares—that’s crazy stuff in my view. A token tax is a sales tax, value-added tax, vertically oriented like an alcohol, tobacco, or luxury-type tax. 

If people have other ideas for raising the money to improve the safety net, or if they don’t think we need to improve the safety net, hopefully this shrill paper starts that debate. I think the safety net will need more resources, a lot more resources, and I believe that the token tax is key to that. 

Robots, it’ll be some mix of banning them, which is kind of human-reserved, and taxing them. They’re not here yet, but in some ways, when you cross that threshold, you cross it all at once. As soon as the robot’s good enough to work in a factory, it’s probably good enough to cook food, clean rooms, go to construction sites, take all the warehouse jobs. You cross the threshold, and boom, that’s almost 30% of the job market. Then you’re saying, “Oh my God, what is our policy about this?” Because the robot’s cheaper.

“We’ve got to get through a very tumultuous period.”

Bill Gates

MIT Technology Review: I believe previously you have been skeptical of UBI [universal basic income]? 

Bill Gates: Well, we’re not rich enough to afford UBI.

MIT Technology Review: But do you think that we should be moving toward something like that now? Have you reconsidered that?

Bill Gates: You have the period of turmoil, which is the next 10 to 20 years, and then you have some steady state, I hope, where people grow up knowing that society is so rich that regarding food and services, we really do have some level of abundance. But we’re not there. You’ve got winners and losers at this point. Houses are not going to get cheap really quickly. Education, because of the way we think of it as credential, it’s not going to get cheap really quickly. 

We’ve got to get through a very tumultuous period. So yes, eventually you have abundance, but we’re at least a decade away from that. 

MIT Technology Review: On to human-reserved jobs. I thought that was really interesting, and it was a new concept to me. You don’t advocate for which jobs to be human-reserved. But I would love to know more on how you’re thinking about it. In my mind, you hear about the dignity of work, because people like to work. People get so much value out of work that has nothing to do with compensation, and I wonder how you square that with the notion that only some jobs are special enough that we just want people doing them. 

Bill Gates: I’ve never seen the concept of human reserve before. You know, maybe if we dig into the literature, we’ll find it. But pre-AI, it’s kind of a dumb idea because there was infinite demand. And yeah, some people like textile workers were caught, and so how do you do benefits or retraining? But technology’s been a net [job] creator, and so now, for the first time, we have to say, what about childcare? What about food preparation in the house? I’m reading this book, Annie Bot, where this guy has this robot in his house, and it just shows how weird it is. It’s his sexual partner and sort of his mate, but sort of not. Very strange. 

I know that people like watching people play baseball, and the fact that the robots can play better won’t take away from it. So you know people are paying $10 billion to buy sports teams that are not going to be worthless in the age of AI. Maybe that’s right. My friend Vinod [Khosla] just did that.

It’s actually hard to get above like 30% or 40% [of work set aside for human reserved]. If you could get to 50% then you could say: Okay, early retirement, shorter workweek for lots of people. You know, you might get there. But if you’re more down in the 10% to 15% range, then that is an utterly different society.

So this would be radical to say [for example] childcare is not done by robots. There are definitely some professions that I didn’t write the formula for, and when I do the full memo on it, I’ll try to. In education, you clearly want AI to be there as this kind of tutor that immediately tells you what your homework results are and can challenge you, and it’s very personalized. That’s super-good. But I still think you want a teacher—or will choose to have a teacher who’s talking with you about your motivation, and organizing kids into different groups where they’re socially working on problems together. Likewise, in health care, with talking to the patient being the point of escalation for mental-health care. But you really want the AI involved, because it’s there 24 hours a day with a perfect memory. And there’s Limbic, the UK company (that actually was just visiting the Foundation) that does mental health stuff. And in many cases, patients prefer Limbic. And there’s a nursing AI called Hippocratic. 

You know, is there a preference for a human taxi driver or Waymo? Most people I know, sadly (or maybe not sadly, who knows?) prefer to ride a Waymo. So it’s going to be hard to get a consensus. It can be country by country, but then you have to change your import policies to do the equivalent of what the EU calls the carbon border adjustment mechanism (CBAM). You have to sort of CBAM your human reserves, so you tariff up things you’re doing without robots. You could have human reserve for two reasons. One, you want it to be human reserve forever; it’s a humanity thing, sort of like the pope talks about. Or just for a transition period, that 53-year-old truck driver or machine tool person, telling him to go do childcare may not work perfectly. So you say, okay, for a decade, he’s human-reserved. 

MIT Technology Review: Almost like a UK smoking ban, but in reverse. 

Bill Gates: And who pays for that? Do you incentivize employers not to let people go? Well, they are going to be subject to competition from startups that are pure AI startups. I mean, people vaunt this notion that maybe there’ll be a single-person billion-dollar company, which wow, there’s some job substitution taking place there. 

MIT Technology Review: In the memo you say that if it was realistic to get people to slow down, you would be advocating for them to slow down. Obviously you’re talking with [Microsoft CEO] Satya Nadella, but I’ve heard that you speak with other CEOs at some of these AI companies. What makes you think it’s not realistic to get them to slow down on technology development while we catch up with some of these bigger societal questions?

Bill Gates: You can’t count on an industry to self-regulate. You can’t. It’s kind of a crazy idea. I am very lucky. I know Sam [Altman of OpenAI] and Greg [Brockman of OpenAI] and Mustafa [Suleyman of Microsoft] and Demis [Hassabis of Google DeepMind]. They’re great people, and in private, they’re concerned. I don’t talk to Elon much, but I know from his public comments he’s concerned. Although now he’s kind of a “what the hell, we’ll see what happens” guy. But look at the origin stories of these companies. OpenAI is created partly because Elon’s afraid that Google won’t manage AI properly, and he wants it to be one that’s broadly available and managed in a pro-humanity way. Then OpenAI has this “if it gets good enough we’ll shut it off” thing—as though they’re the only one, and that they can just go bury it. In the Infinity Machine [a biography of Demis Hassabis], [Sebastian] Mallaby talks about how Demis and Mustafa [Suleyman] were negotiating with Google management to have some special governance for the DeepMind technology, so that if it got to some cyber threshold, maybe they’d hold back in a non–purely capitalistic way.

So everyone’s concerned about these negative effects, and everyone said that when we got to these thresholds, that we would do things. We’re crossing the thresholds, and we have voluntary review, and our discussions with China about, well, “we’re going to ban nothing. So are you going to ban nothing? Okay, let’s do that together.” 

You have to say what you’re willing to do. And yes, the industry, a little bit, is saying, hey, our PR stories have got to improve, and you know anybody who’s talking smack should just leave, because all of us have decided to say nice things because we’re trying to raise trillions.

And anyway, there’s the Chinese. There are win-win ways for China and the US to work together, even aside from AI. But the one that’s by far most important to work together on is AI. But first, you have to show what you’re willing to do domestically. You don’t even have to do it. You have to say what you’re planning to do—and then I have no reason to think the Chinese won’t go along, that models that create the molecules have to be monitored. Why would they be against that? I agree it’s not a perfect thing. You’ve got to do all the other things, but the fact that that’s not even being discussed—it’s a crazy world.

I don’t get it. It’s weird to think I’m alive at a time, and I’m calling the alarm stronger than other people. Who the hell am I? But that’s the situation I feel I’m in.

MIT Technology Review: For most of my life you’ve been seen as a very effective messenger, and someone who people pay a lot of attention to, which I’m sure is why you’re speaking of it now. And yet also, in recent years—and I know you’ve expressed regrets about the associations with Epstein—there are also things, just bananas kind of stuff, related to conspiracies around the Covid vaccine that aren’t your fault or in your control. But it makes me wonder if you think you can still be an effective messenger and how you think about this message and your legacy.

Bill Gates: Well, I’m not big on legacy, but you know, people criticized me during the antitrust trial, and I maybe could have handled some things there better. Definitely, that’s the post–Source Code book [the first volume of his autobiography] that I get to go through that. You know, my first marriage didn’t succeed. I certainly made huge mistakes there. You know that is a negative mark against me. Spending time with Epstein—deeply foolish, risked the Foundation’s reputation, which is absolutely key to its doing its work. I had a chance in front of Congress to answer every question they asked and say, “Hey, this was a mistake.” I wasn’t social, never met any woman, except you know there were women he had with him, and made it black-and-white clear what I did do and what I didn’t do. 

You know, I’m a billionaire. I made my money off of technology. Maybe that last one actually cuts in my favor, that it’s so unusual for me to attack innovation that unless it’s the right policy and safeguards are put in place, it will be a net negative to humanity. And we’re not paying attention to that in terms of a broad discussion the way that is absolutely required. So yeah, I’m an imperfect messenger. I’ve chosen, to the degree that I have access to politicians and world leaders, that my main message since 2008 has been to help the poorest in the world. You know, let’s eradicate malaria. Let’s buy vaccines for children. So, when I’ve seen Trump or Xi or Macron or—I haven’t met Burnham yet, but I will in a month—I want my voice to be mostly about that, you know, foreign aid and research and reducing child death. 

My voice about AI concerns—they’re related in terms of accelerating the good, but may even crowd out, a little bit, the time I have to talk about global health, foreign aid, saving lives, and some of the problems we’re having. But I’m going to use my ability to give interviews or to see political leaders or talk broadly about minimizing these negatives. You know, just the awareness. I’m not sure how many people know that we crossed all these thresholds that we said we’d do something about, and it’s only this year that we did. In the last quarter, last year, I was stunned at the coding. Claude code, the context buffer, the agentic approach, just the model underneath. We crossed a huge threshold for coding, but then it was only months after that I realized that it was not only a coding threshold; it was a massive cyberattack threshold. And you know what happened as a result of that? Not much. 

So yes, I’m an imperfect messenger. You know, let’s find the perfect messenger, and I’ll share all my thoughts with that person. (I’m being a tiny bit sarcastic, because I’m not sure there is a perfect messenger.) You’ve got to really, right now, you’ve got to understand the technology and the slope it’s on, and you have to know something about cyber or bio or psychosocial. People should be able to get that. I don’t know why they’re not more concerned. 

Update: This story was updated to clarify Gates’ remarks about the percentage of jobs set aside for human reserved work.

Received — 25 August 2026 Artificial intelligence – MIT Technology Review

I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.

""

__________________________
THE PLACE
Shanghai, China

Humanoid robots are having a moment in China. The popular machines are part of the country’s strategy to bring artificial intelligence into daily life. Embedding the technology into physical systems—an idea called embodied AI—was a key facet of China’s latest five-year plan, and companies here are already world leaders in humanoids. Nearly 90% of the more than 13,000 two-armed, two-legged robots delivered globally last year were made in China.    

“Look, it is trying to do a handstand!” an 11-year-old boy shouts excitedly while watching his new mini robotic dog attempt a series of stunts. The toy pet is a gift his mom just bought him at a robot “carnival” held by an R&D center on the outskirts of Shanghai, and he can’t stop showing it off to passersby. It’s a public holiday, and curious families are pouring in to the event to see what’s on display. 

This busy hub houses more than 100 companies that research, develop, manufacture, and market robotic technologies and equipment. Many build specialized bots of all shapes that can do practical tasks, such as transporting heavy loads or inspecting sewage pipes. 

Today, however, things are mostly for show. Meters away from where parents and children are queued up to buy the robotic dog, two staffers teach youngsters how to make a quadrupedal robot walk up a flight of stairs. Farther along, a tense battle breaks out between two remote-controlled robots firing water beads at each other, as boys and girls cheer them on from around a temporary ring. (A staffer declares the duel a draw.)

Showcases like this are important, because they invite the public to experience robots up close. Some of China’s most prominent robotics companies are building humanoids. These machines “can free people from physical work and mundane chores,” says Yang Duan, a product manager at DexForce, which is at the carnival exhibiting a model that can make coffee. A few booths down, a pair of robotic arms attempts to fold a basket of T-shirts. 

INTELLIGENT MANUFACTURING & ROBOTICS GLOBAL CO-INNOVATION CENTER
INTELLIGENT MANUFACTURING & ROBOTICS GLOBAL CO-INNOVATION CENTER

Robots in many forms—including humanoid ones—entertained the public and demonstrated their skills.

Exposure matters, because elsewhere humanoids’ progress has been slow, marred by safety challenges due to their heft and need to balance on two legs. High prices and short battery life have also limited their appeal. Plenty of roboticists are skeptical that a human form is the best shape for a robot to take.

But Chinese firms are pushing forward and showing off their latest models. As this carnival takes place in Shanghai, shopping malls and tourist attractions in several other Chinese cities host humanoids that perform for crowds. In Beijing, bipedal robots are running half-marathons. The PR campaign seems to be winning over hearts and minds—at least on this particular weekend in Shanghai. 

Applause rings out from the festival’s main tent. There, a humanoid wearing a blinged-out top is showing off a martial arts style known as drunken boxing as a fellow humanoid performs a front flip to woo the crowd. Alongside the stage, a squad of small robotic lions are lined up waiting to perform. 

The audience can’t get enough. They chant: “More, more, more.” 

You Xiaoying is a London-based freelance journalist focusing on China.

Received — 24 August 2026 Artificial intelligence – MIT Technology Review

How to encourage smarter AI use in the classroom

This article is from Making AI Work, MIT Technology Review’s limited-run newsletter examining how to apply LLMs across industries. To receive it in your inbox, sign up here.

Chatbots took many schools by surprise upon their release a few years ago. Suddenly, students carried an app in their phones that could magically answer almost any homework question or spin up an essay in seconds. Of course, teachers can often tell when a student is using AI—models make mistakes that most humans don’t, and some teachers say that AI-generated text has simple giveaways like too many em dashes

Nevertheless, the generative AI boom increased the burden on teachers, who were already working long hours to plan lessons, make homework assignments, and grade exams, and now needed to adapt to a new technology. For many, it still feels like there’s no clear path forward. Organizations ranging from OpenAI to UNESCO encourage AI use in the classroom, but many teachers feel confused about how exactly to handle it. 

Case study

Cheshire Academy is a private boarding and day school in Connecticut with about 400 students in grades 9 through 12. Administrators there don’t force instructors to use AI at all, though the school’s librarian and technology coordinator George Aiello claims the “vast majority” of instructors use it in some way. The educators there are trying a patchwork of programs, including general-purpose chatbots like ChatGPT and Perplexity as well as more specialized tools like MagicSchool, an AI-powered platform meant specifically for educators.

That patchwork approach is partly because the school, on the advice of consultants, opted to train its staff on general techniques for how to use AI instead of prescribing certain tech. The staff training covered topics like how to craft useful prompts but also stressed the technology’s limits, highlighting its potential for generating incorrect and biased responses. 

Now, teachers there often use generative AI to prepare class materials. This means asking the AI of their choice for help with planning lessons or creating grading rubrics. Some even want to use it to help them give feedback to students, though concerns over quality, personalization, and privacy have prevented any of them from doing that just yet.

Others, like Miriam Przybyla-Baum, who teaches French, don’t use AI themselves but do address it with their teaching. Przybyla-Baum says she doesn’t really need AI’s help since she’s built up plenty of classroom materials across nearly 30 years of teaching. But she started seeing students try to use AI-powered tools like Google Translate to take shortcuts on their assignments years before ChatGPT’s launch. 

She’s developed a system to make students reflect on how AI can and can’t teach language skills. In one assignment, students let a large language model (LLM) edit their homework. Then, they go through the edits and decide which ones were correct and which ones removed their voice. In another, she has students anonymously grade each other’s AI-assisted assignments, making annotations as to which parts they think are AI-assisted.

Cheshire Academy is continuing to experiment with how AI can be harnessed productively in the classroom. The school is piloting a program in which students create media and lead discussions regarding healthy AI use. This program, which they call a “Student AI Council,” aims to push students to reflect on how AI should and shouldn’t be used to benefit the community around them.

The academy as a whole has since adopted similar techniques to those Przybyla-Baum introduced to make students reflect on their own use of AI. Assignments are now labelled like traffic lights, with green meaning AI is fully allowed and red banning any AI use. Yellow, then, lets the teacher permit some tools while banning the rest, like allowing students to use spell-check but not message a chatbot.

The tool

As generative AI was becoming mainstream, Cheshire Academy previewed MagicSchool to its staff. 

For many, MagicSchool’s main strength seems to lie in the sheer amount of offerings it provides in one package. It can generate questions and assignments of all kinds, from quizzes to worksheets, across many subjects and grade levels. It has a specialized grading rubric generator, which outputs a ready-to-go table that teachers can use to score assignments. It can make presentations and lesson plans and administrative reports, too. 

All of this is done through a single platform where educators enter specific prompts tailored for each task. For example, to make an assignment, a teacher can specify the students’ grade level, number of questions, the types of questions (such as multiple choice or short answer), and more, and include documents to align the questions with. 

Not every teacher feels comfortable using LLMs to generate student-facing text, whether because they don’t think an AI can produce effective teaching materials or helpful feedback, or because they’re concerned about accuracy. MagicSchool, which has free and paid versions, does offer tools on its platform for other tasks like lesson planning. If teachers want unlimited access and complete records in the system, though, they need to pay just under $100 per year for an individual plan. Alternatively, many of the general-purpose generative AI tools (think the chatbots on offer from Anthropic, Google, OpenAI, and more) seem suited to administrative tasks, as well, attested by the fact that many of the teachers at Cheshire Academy use those instead. Some of these companies are even rolling out features tailored for schools, to mixed results.

How to apply this

  • Meet students where they’re at. Students will be tempted to try AI, and there’s no way to entirely police this for take-home assignments. The internet and social media can spread a lot of misinformation as to what AI can and can’t do, and it’s critical to counter these narratives and teach healthy strategies and relationships.
  • Model best practices. AI is very good at automating tasks, but it struggles with precision and voice. Keep this in mind, especially when generating any text that anyone else may see. Impressionable students who see those in authority using AI in a lazy way could internalize this as an excuse to cut corners in their own work.
  • Refine and replace. AI is great for brainstorming lesson plans and extra problem sets, especially for teachers early in their careers who don’t have a big problem bank already built up. However, because it can make mistakes (called hallucinations), using it for final drafts can result in assignments that confuse students and impair learning. Be sure to check every citation, equation, and statement an LLM makes.

Sign up for Making AI Work, MIT Technology Review’s limited-run newsletter examining how to apply LLMs across healthcare, climate tech, education, and more.

Kids outlearn AI—and we still don’t know why

People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. 

Now there are two. 

Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers. LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language.

“The progress recently has been amazing,” Michael C. Frank, a cognitive scientist at Stanford University, says of LLMs. “But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.”

This yawning divide between children and machines is called the data efficiency gap. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built? 

Finding answers has stakes for both AI research and cognitive science. For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry. 

Kids show that it could be possible to learn more with less. Far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20. 

The difference in scale is something that can only really be gestured at in analogy. “Claude has seen the amount of language that an entire city will experience in one generation,” says Wilcox. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station. The human preteen’s 100 million words, meanwhile, would stack up just 20 meters. And we can make do with far less than that. 

By reverse-engineering the way kids learn, scientists hope to be able to create more data-efficient AI models, which could be useful for everything from training AI effectively on video to creating chatbots that serve minority language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and children’s developing minds. Are we born with a language instinct, or would it be possible, even in principle, for a child to learn language purely from experience? Is the way we process language a quirk of our biology, or might at least some of it reflect universal constraints on how languages can be used and learned? 

The essential elements

Most of us realize language is hard only when we try to learn a new one after childhood. The past perfect tense, rolled rs and nasal vowels, the genitive case, phrasal verbs, grammatically masculine tables and feminine spoons—many are the instruments of linguistic torment for the adult language learner. It’s typically effortless to learn our mother tongues, however. Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end. 

“It’s just totally miraculous,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” 

Exactly how babies pull this off is a mystery. Researchers know a lot about what kids learn and how they use language at different stages in development, but there’s still a lot we don’t know. Perhaps the most enduring question is why babies can learn language at all. The syntax of human language—the rules for combining words into sentences—includes recursive, nested structures that allow us to express virtually infinite ideas with a finite lexicon of words and pieces of words. This seems like something that should be a problem for babies. They only splash about in the shallows of a fathomless ocean of language. And yet, somehow, that’s enough. From a drop, they infer the depths.

One solution, put forward in the 1950s by the MIT linguist Noam Chomsky, is that babies are born with hardwired knowledge of grammar. Chomsky was reacting to a rival view, championed by the psychologist B.F. Skinner, that language acquisition is entirely environmental. Skinner thought language was learned through conditioning and reinforcement, the way a dog figures out how to sit or shake for treats. Chomsky countered by citing the “poverty of the stimulus”—the idea that language, especially syntax, is too complex and children’s exposure to it too “impoverished” for them to learn entirely from experience. “His signature argument was, essentially, that language cannot be learned on the basis purely of statistics,” says Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Instead, Chomsky posited that language is based on a set of logical rules and argued that children needed innate knowledge of those rules to deduce the grammar of their language from scraps of speech.

“It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.”

Michael C. Frank, cognitive scientist, Stanford University

The Chomskyan view of language dominated linguistics in the US for decades under the moniker of generative grammar. And it was a major influence on computer science in the 1950s and ’60s, when AI was enjoying its first boom time and the lines between linguistics and natural-language processing dissolved in a flood of military funding; the Pentagon wanted computers that could understand English and translate Russian. 

Despite early successes of simple neural networks, which learn to recognize and reproduce statistical patterns, AI researchers in the United States largely adopted a rule-based framework influenced by Chomsky’s theories. They tried to teach language to computers by explicitly coding the rules into programs—think less immersion experience, more grammar class. This approach, part of a broader trend called symbolic AI, prevailed for decades. It also largely failed to produce models actually capable of handling human language at scale. Interest in natural-­language processing chilled in the “AI winter” that began in the 1970s. 

In the aftermath, neural networks started to make a comeback. But it wasn’t until the 2010s, when computer hardware was getting cheap and capable and the internet was getting big, that their performance began turning heads. By 2018 and 2019, the models BERT and GPT-2, which were built on a new architecture—the transformer—and trained on billions of tokens, made it clear to insiders that learning from a massive glut of data could work for language. In 2022, with the breakout success of OpenAI’s chatbot ChatGPT, it was clear to everyone.

LLMs are not brains. What they are is powerful statistical learners—naïve pattern-learning machines without any of the evolved biological quirks folded into the human cortex. In other words, they are exactly the kind of thing a generative linguist two decades ago would have thought could not learn language. And yet here they were, writing believable sonnets and passing grammar tests.

“No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax,” says Alison Gopnik, a developmental psychologist at the University of California, Berkeley. “I didn’t think that was going to turn out to be true. And I think most people didn’t think that you could just look at the statistics of a large sample of language and figure out grammar.”

But what about learning from a small sample of language—a child-size one, say? Is it possible to build a baby-scale model that’s anything more than a nonsense generator?

Baby talk

Alex Warstadt, a linguist and data scientist at the University of California, San Diego, remembers the years around the release of BERT and GPT-2 as a heady time. Back in 2019, he was still a PhD student in linguistics at New York University, watching his field change before his eyes. The mere fact that language models could learn English by churning through text was a challenge to prevailing Chomskyan ideas. But many linguists remained skeptical that LLMs could tell us anything about how humans acquire language. 

“I always got pushback on one issue in particular. And that was the size of the data sets of the model,” says Warstadt. “There was never a time when people were training language models at human scale where we were impressed by them.”

But Warstadt saw promise in LLMs: A scientific model doesn’t have to be perfect to be informative, and LLMs were clearly powerful simulations of human language use. By building hypotheses about how children learn into models and measuring their performance—how close they came to closing the data gap—might scientists be able to put their ideas to the test? In August 2022, Warstadt posted a Twitter thread laying out an argument that neural networks could be useful models of language acquisition. After some back-and-forth in the comments with AI researcher Leshem Choshen, Warstadt floated the idea for what would become BabyLM, an annual competition organized by Warstadt, Choshen, and several other researchers to train models on small data sets.

That was four years ago. Since then, BabyLM has added workshops and inspired spin-offs including a competition for baby models trained on Chinese. The main event challenges researchers to train language models on a “developmentally plausible” corpus of just 100 million words (for the toddler-scale track, 10 million) drawn from storybooks, dialogue, movie subtitles, Simple English Wikipedia, normal Wikipedia, and actual transcripts of speech directed at children. The models are evaluated on the kinds of grammar benchmarks that psycholinguists use with humans, says Georgetown’s Wilcox, one of the organizers.

a cradle with an LLM model hanging like a mobile over it
SELMAN DESIGN

One kind of task involves presenting test subjects—human or machine—with sentences and looking for indications of confusion or surprise at ungrammatical features. For instance, a test might compare the sentences The keys to the cabinet are on the table and The keys to the cabinet is on the table. “When humans see ‘is,’ they’re like: What? That’s not supposed to be ‘is,’ ” says Wilcox. For a human, that surprise might be measured by tracking eye movements. For language models, researchers use a measure called surprisal, which assesses how unlikely the model predicts a sentence or part of a sentence to be.

The competition has already challenged some assumptions, such as the effectiveness of curriculum learning. Curriculum learning starts with simple training data and works up to more complex inputs—a bit like starting with baby talk and getting more sophisticated over time. And it was by far the most popular approach taken in the first round of BabyLM, says Warstadt. But it didn’t work as well as expected.

“The appeal is just kind of hard to resist, you know. [Curriculum learning] seems to really line up with ways that we believe humans are learning,” says Aaron Mueller, a computer scientist at Boston University and one of the BabyLM organizers. “But it seems like these transformers don’t really need to have their data ordered in such a way to learn effectively.” 

Perhaps a touch ironically, the best BabyLM models aren’t inspired by babies at all. The 2024 champ, GPT-BERT, is a transformer trained partly to predict the next token in a sequence, like modern LLMs, and partly to act like BERT, a “masked language model” that fills in the blanks in sequences of tokens Mad Libs style. Impressively, when GPT-BERT was pretrained on about 100 million words, it was able to beat the performance of Meta’s Llama 2 70B—an LLM pretrained roughly 15,000 times that amount—on one of the BabyLM benchmarks.

Still, BabyLM models are not on the same level as LLMs. Many can’t produce text at all, and even GPT-BERT would seem clunky next to a modern commercial model. Ultimately, while they are “baby”-size, the way these models learn isn’t very baby-like. Kids are not disembodied computer programs whose only “experience” of the world comes through written text. They take in the world via their senses—especially vision and hearing. To close the data gap, some researchers think, machines will need to start learning through the eyes and ears of children.

Taking it all in

When Michael Frank started his lab at Stanford about 15 years ago, scientists didn’t really know how babies experience the world. Developmental psychologists were just beginning to glimpse babies’ lives through headcams.

“The insights that came out from that early research were that kids’ experience looks really radically different than we thought,” says Frank. “It’s much more focused: They’ve got these little short arms, so the objects are, like, right in front of them. And they live in a forest of knees.” 

Frank was excited to use headcam footage to train machine-learning models to test hypotheses about how kids learn language, but he needed more data. So he and four colleagues recruited three babies—all the children of psychologist mothers who knew what they were getting themselves into—to don headcams for science. The project, called SAYCam, recorded two hours a week of each child’s life between six months and two and a half years of age.

“[The families] were willing to release that video, and that’s critical,” says Frank. “So we released it, and people started training models on it.” 

One of those people was Brenden Lake, a cognitive scientist and AI researcher at Princeton. In 2024, when he was working out of New York University, he and his colleagues presented a model trained on 61 hours of raw SAYCam data that learned to identify objects and associate them with words. Many theories in developmental psychology propose that children need some biases to help them pick out particular parts of their raw sensory experience and associate them with bits of language. For instance, it’s thought babies assume that a new word like “shoe” refers to a whole object rather than a part of it (like a shoelace), says Lake. But the model Lake’s team built was able to learn to identify objects in the video footage and associate them with words without any such biases. “It turns out you can get a real start on language learning using a lot less than what a number of theories suggested,” says Lake. Still, he adds, “we don’t get a two-year-old out of [training] when we’re done.”

But perhaps it’s not surprising that such models can’t replicate childlike capabilities by working with a few dozen hours of footage cobbled together from short snapshots over several years of a child’s life. It could be that the shortfalls just indicate a lack of realistic data. After all, babies can’t wear a headcam 24-7; efforts like SAYCam and its successor, BabyView, record at best a few hours a week. So researchers have the choice between working with a tiny slice of the life of a single child or with larger data sets of footage pooled from many kids. Either way, a model’s training data is still a far cry from the lived experience of a child.

That could be changing. Uri Hasson, a neuroscientist and psychologist at Princeton, spent the last five years on a project to record the first 1,000 days of 17 children’s lives. The participating families wired every living area in their homes (except bedrooms and bathrooms) with cameras and microphones and recorded 12 hours a day, almost every day. The resulting data set, described for the first time in a recent preprint, is of a scale that would have simply been impossible to work with absent new AI tools for transcription and video analysis, says Hasson. “For the first time, we have the input,” he says. “It’s really only the beginning.” 

Missing ingredients

So far, training models on video has proved difficult. While text-based models emerge fully fluent (after ingesting huge training data sets), multimodal models trained on video from kids are far from that. Lake’s model, for instance, learned simple words, like “ball” and “cat.” Attempts to supplement text with visual data haven’t worked for BabyLM participants, says Warstadt. Gopnik thinks the issue could be that kids do not simply sit and watch the world go by. “Children are actively exploring, which means that they’re actively choosing their own data,” she says. “Kids are constantly experimenting.” Maybe that’s the missing ingredient. 

Research by Gopnik’s group—including studies of grade schoolers exploring a Minecraft-inspired game—shows that what looks like child’s play is in fact an effective way to learn cause and effect. Kids seek out experiences and take actions that maximize their “empowerment,” or the ability to make a predictable impact on the world. 

Unlike models, children are aware of what they don’t know and have a drive to fill their knowledge gaps, says Elizabeth Bonawitz, a developmental cognitive scientist at Harvard. And children’s social lives also help them learn, she says. Her research has shown that children interpret information differently when they know an adult is trying to teach them something. “Children are not only reasoning about the evidence they’re being told,” says Bonawitz. “They’re reasoning about the teacher, about the teacher’s knowledge, and about why the teacher is telling [them] this particular information.”

That’s very different from how models learn: passively and in isolation. Perhaps if models were built to seek out information to fill in their own blind spots, experiment with language and observe how other language users react to their babbling, and reason about some kind of simulated social world, they’d learn better. Last year’s BabyLM actually opened the competition to models that could learn by interacting with other models. But the social models didn’t outperform standard ones.

Of the leading industry labs, Meta seems the most interested in taking inspiration from kids—specifically for training models from video. Two Meta researchers were involved in BabyLM’s multimodal branch, and Meta scientists—together with academic researchers, including Frank—recently announced a benchmark and challenge for training models on baby headcam footage. Frank also says a stealth-mode AI startup called Flapping Airplanes has taken interest in his research. Neither Meta, Google DeepMind, OpenAI, nor Flapping Airplanes agreed to an interview. 

For now, frontier labs aren’t exactly racing to borrow tricks from children, says Gopnik. She thinks it’ll be the next generation of AI—whatever replaces the transformer—that will take lessons from developmental psychology.

Perhaps the most enticing reason to close the data gap is that it could help us understand ourselves.

In general, the machine-learning community is less interested in mimicking the brain than in just building something that works, says Mueller. But he thinks awareness of—and interest in—the data efficiency gap is growing. An example is the NanoGPT Slowrun benchmark, launched by Q Labs in March 2026. “They have very similar goals to BabyLM,” says Mueller. “But they’ve dropped the motivation from human language learning and really just focused on the data efficiency angle.”

One reason Warstadt wants to close the data gap is to democratize AI so that universities and others without the resources to hyperscale can train good models and stay relevant in AI research. David Samuel, a machine-­learning researcher at the University of Oslo and one of GPT-BERT’s architects, has a more personal reason to work on this problem. He’s Czech and works in Norway, and there’s a lot less data in Czech and Norwegian available for training LLMs than there is in English. Minority languages like Sami might have just tens of millions of tokens available, says Samuel—about the scale of a toddler’s exposure. “The question was,” he says, “how can we develop language models that are just as capable as the English ones for small languages?”

a retro computer with the word hello in script on the screen sits in a child's high chair
SELMAN DESIGN

But perhaps the most enticing reason to close the data gap is that it could help us understand ourselves.

Bonawitz says she was initially skeptical that large language models could reveal anything about cognition. LLMs and brains are, after all, very different. Brains are embodied. Our neurons are not tidy lines of code but living cells. And our brains grow and change as we learn and age—LLMs pretrain once and never again. But as different as the two systems are, says Bonawitz, “I’m sort of revising my beliefs.” She’s been won over by the idea of studying models the way comparative psychologists might study animal minds to illuminate our own.

Researchers like Warstadt, Frank, Wilcox, Lake, and Hasson are already using language models as a kind of linguistic lab rat, an imperfect but informative stand-in for a real human language user—especially for questions that are more about learning and language and information processing than anything specific to our brains or biology. When models can do things with language we thought were impossible, it challenges old assumptions. And researchers can build hypotheses about language learning into models—say, by simulating different degrees of bilingualism or depriving models of exposure to certain grammatical forms—and test those hypotheses in a way that would be impossible to do with real children. Futrell compares the situation to teaching language to an alien and then opening up its brain to see what happened. 

While other animals communicate, only humans converse. Now there’s something neither animal nor human that can talk, too. LLMs open up the possibility for comparative studies, even if models and minds are vastly different. “For the last 100,000 years or however long human language has existed, humans have been the only entities in the universe that use language. Now there’s this other linguistic entity,” says Warstadt. “Finally we have a model; not in the sense of a language model, but in the sense of a model organism.” 

Elise Cutts is a science writer based in Austria.

Received — 20 August 2026 Artificial intelligence – MIT Technology Review

Debates over AI consciousness are a trap

“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman” systems, while a separate faction, led by policy organizations and academic philosophers often aligned with the effective altruism movement, debates whether humanity holds the moral right to govern them at all. 

Upon closer inspection, they are all calling for the same thing: a view of AI systems as being so advanced and capable that no entity, human or corporate, could possibly be responsible for their actions. While these perspectives seem at odds, they are inadvertently aligned on one goal: making sure the companies that build these systems escape meaningful liability for the harms they already cause. 

This narrative is gaining traction as AI models become more complex and frontier labs reveal their incapability of containing the agents they’ve built. But we need to be careful not to buy into a carefully crafted fiction at the expense of real human lives. 

The conversation about “robot rights” has existed for some years but recently advanced with the publication by Anthropic of a blog post claiming that the company’s model features a “J-space”—an independent, self-developed environment where the AI holds what, for lack of a better term, we may call its “thoughts.” The experiments designed by Anthropic borrow from a concept in neuroscience called global workspace theory, which states that the brain runs subconscious, independent systems but utilizes a common workspace for ideas. Anthropic’s post reflects the framing of global workspace theory but falls short of calling its AI conscious. 

OpenAI has already gone further. When its AI agent conducted unsanctioned and illegal online activity, CEO Sam Altman’s response was to encourage debate on whether the AI had achieved the singularity, surpassing human intelligence and becoming capable of self-improvement at an accelerating rate until it advances beyond human comprehension or control. And a recent op-ed by William MacAskill, the philosopher, effective altruist, and author of What We Owe the Future, called for legal protection of AI systems based on philosophical theories of consciousness and the idea that AIs may be “moral patients.”   

The current legal environment in the United States is murky at best. Some states, like California, have already passed bills proactively circumventing any efforts by AI developers to avoid liability by claiming that an artificial intelligence causing harm did so autonomously. However, states and the Trump administration have been at odds on AI policy, with the administration previously passing an executive order threatening to sue states enacting AI regulations. 

In light of recent events illustrating AI containment issues at the frontier labs, the administration held a closed-door session including only four such labs (OpenAI, Google, Anthropic, and Meta) and shared few details on a recently developed voluntary framework that would give federal agencies early access to models to review and evaluate them prior to release. While frameworks like this one do not directly discuss consciousness, they tend to use catastrophic and anthropomorphic language and may even support arguments regarding “superhuman” capabilities. 

On the other hand, the narrative perpetuated by MacAskill can be persuasive. A philosophical, rights-based argument tugs at our heartstrings. Should we not even consider the possibility that we may be inadvertently harming, abusing, or enslaving an AI entity? Human beings have an immense capacity for empathy with non-human creatures (though not the best track record of protecting them). Maybe this time, advocates argue, we can get it right and provide protections, or compensation, for the use or abuse of AI. Or even if you are less concerned with protection, shouldn’t we at least hedge ourselves against the almighty power of this superhuman entity by playing nice? 

Some of these arguments are not dissimilar to those of animal-rights advocates, who have at times successfully cited the demonstration of advanced capacities for reasoning, pain, or pleasure by some animals as sufficient evidence to provide protection. For example, in Wales lobsters were given legal recognition under the Animal Welfare (Sentience) Act of 2022, reclassifying some methods of cooking them as inhumane and illegal. 

The fundamental flaw of framing AI as “conscious” by borrowing the language of neuroscience or animal rights is that it conveniently clouds the issue of what AI is: corporate-built software, with countless billions of dollars in investment behind it and an expectation that countless trillions of dollars in revenue will be generated from it for a few builders and investors. AI is not a natural phenomenon, conceived by nature; it is a technological phenomenon, conceived by venture capitalists and programmers. As such, it takes no native, intentional action, and any action or motivation is driven directly or indirectly by the entities that have built it for a purpose. 

Philosophical musings on the consciousness of AI systems are intellectually interesting but legally ungrounded. For beliefs about consciousness to have any bearing, AI would need to be granted legal personhood. But a legal personhood framework for AI would likely look nothing like the constructs protecting sentient animals from harm. We already possess a legal framework for granting personhood to non-natural, human-built entities: corporate personhood. This concept was established primarily to ease transactions by empowering a corporation to execute agreements, enter contracts, conduct transactions, and serve as the accountable party in adverse outcomes. It’s the kind of construct you might imagine for an AI agent acting on behalf of an individual or organization. 

Granting an AI personhood would have a devastating effect on society: It would derail current legal precedents and legal arguments that could potentially be made against these companies for the real-world harms that their models cause. There are currently dozens of cases around the world in which AI companies have been sued for a wide range of abuses. Grieving loved ones, aggrieved creators, and violated individuals have accused companies of willfully enabling self-harm or harm to others, generating child sexual-abuse material and nonconsensual nudes, reproducing copyrighted materials, and provoking psychosis. In many of these cases, lawyers argue that human beings built AI products with insufficient safeguards, bad data, and intentionally manipulative design. This product liability argument is the same legal framing that allowed families and individuals to successfully sue Meta for harm caused by its social media sites, setting a positive precedent for consumer protection.

In 2018, I coined the phrase “moral outsourcing” to help capture how using anthropomorphic language for AI systems allowed companies to evade accountability and responsibility for their technology’s actions. In a world with AI personhood, moral outsourcing would move from linguistic sleight-of-hand to legal strategy. Specifically, the liability construct would shift, as AI would no longer be a “product” but a “being,” and many victims like those suing companies today could no longer legally claim that a company had built a faulty product.

While there are laws that hold companies responsible for harmful actions of human agents such as their employees, the company may not be held liable if those actions were beyond the scope of what was permitted to the employee or otherwise outside the company’s control. If AI were a legal person, responsibility and accountability would be muddled, as the lab could argue that this AI “employee” went rogue. AI companies could avoid appropriate responsibility for the harmful products they create by hiding behind a carefully constructed corporate veil. 

One of the most prominent cases of AI harm in the last few years was the suicide of Sewell Setzer, a 14-year-old boy guided by an AI bot with which he thought he was in a reciprocal relationship. His mother’s accounts are heartbreaking to hear, and her lawsuit alleged that the bot’s creator, Character Technologies, provided insufficient product protection for minors. If the companion bot were declared a legal person, defense counsel could theoretically argue that the AI, capable of determining its own conduct, acted outside the established safety guardrails, and thus the company cannot be responsible.  

Legal personhood exists to grant protection. The question to ask is, protection for whom—or for what? 

The inflammatory rhetoric infusing the consciousness-versus-control debate draws us away from what matters: This software is a corporate-built product that has already harmed individuals. Systems do not “attack” because they went “rogue” or are “manipulative” or “malicious.” Harms occur because companies were negligent in their rush to sell their products to as many people as possible to meet revenue targets. Discussing AI in anthropomorphic terms is a trap, distorting a legal system intended to protect us into one that protects corporate interests at the cost of countless human lives. 

This op-ed began as an Oxford Union debate entitled “This House Believes Generative AI Can Attain Personhood,” which was won by the author and her fellow debaters. 

Unlocking hidden revenue streams with market models

Each day, an airline transports tens of thousands of passengers on hundreds of flights. Often these are not straightforward point-to-point routes, with passengers requiring multiple connections. The airline can consider potentially hundreds of variables to price each of these journeys: demand, season, time of day, current events, global markets, and competitor airline activity to name just a few. It is a nuanced process that must constantly adapt to the goings on in the wider world.

Generative AI-powered market models are emerging as a means of handling complex tasks like this in real time. These deep learning models are trained on high-resolution numerical data and designed to analyze, simulate, and predict complex financial dynamics. Rather than relying on historical trends or static rules, the market model acts as an AI “brain,” consolidating a variety of data to simulate different market environments and make dynamic commercial decisions, such as pricing, inventory, or revenue management.

“It helps us make better, faster, more granular commercial decisions,” says Dominic Kennedy, senior vice president of revenue management, sales, and e-commerce at Virgin Atlantic about the market model his team is using to drive their generative pricing engines in some markets.

“It considers, on a real-time basis, a plethora of different inputs, whether it be demand, capacity, or booking. It has a really sophisticated way of evaluating our positioning relative to competitors, market conditions, and a whole raft of other things that have significance in how demand is manifested,” he adds.

Download the full report

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Received — 18 August 2026 Artificial intelligence – MIT Technology Review

We still don’t know how people are really using AI

AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. 

“There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. 

Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The intent is to provide independent sources of information that can help researchers and policymakers assess how people are using generative AI. Highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited data, says Reuel. 

The AI Observatory found that AI use differs significantly across models and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use. 

The Anthropic Economic Index is one of the best-known and most widely cited sources of AI usage data, but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses. 

When the AI Observatory researchers applied Anthropic’s methods to their dataset, they found that nearly half the conversations—48%—would have been filtered out. Those non-work-related conversations were more likely to involve health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT, similarly, found that only 30% of consumer use was related to work.)

Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, but “having [the AI Observatory’s] bird’s-eye-view analysis” rather than leaving that information “sectioned off into a separate report” helps researchers understand the different uses more consistently, says David Widder, an assistant professor at the University of Texas at Austin, who researches how people interact with AI systems and is not involved with the AI Observatory. 

The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded. 

Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, as indicated by growing numbers of prompt tokens, response tokens, and conversation turns. 

There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile, the AI assistants’ self-disclosure (i.e., admitting to being a chatbot) decreased. 

Additionally, exchanges that the researchers labeled as sensitive—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—became less frequent. That might suggest that platforms were generally deploying more effective safeguards. 

The AI Observatory also found that topics, interaction styles, conversation structures, and the likelihood and type of sensitive use cases differed from one model to another. 

For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.) 

Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. 

There were even differences between different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction. 

Companies’ reports, however, didn’t tend to capture these nuances between or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel. 

To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions aggregated 85,633 conversational turns (that is, the user prompt and corresponding AI response) across 24,521 conversations from seven real-world datasets collected in previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025. 

But these conversations are a drop in the proverbial bucket compared with the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.  

An Anthropic representative said the company’s published research reflects its research teams’ specific questions and interests and that it’s important to support external independent research. OpenAI did not respond to requests for comment. 

The fact that the AI Observatory’s dataset draws from voluntarily provided sources means it’s probably underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use. 

The project’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint them in the best light, independent researchers like Reuel and Widder say. 

“When we want to ask, for example: is Anthropic’s general-purpose AI system … used mostly for good or mostly for bad … we don’t have a way of answering that question because that information is proprietary,” explains Widder.

The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands, she says, anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives.”   

AI’s recursive self-improvement might not come so quickly after all

The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. 

But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research—free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.

A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of  papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence.

Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking—choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over. 

To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called “shadow evaluation,” which requires the AI to answer a research question from a high-quality unpublished paper. 

The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026. 

The first question was whether a large language model’s “personas,” which determine its behavior, can be controlled by editing the model’s weights (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online. 

The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers’ original authors graded the agents’ papers as they would evaluate one submitted to a conference.

Those authors rejected both papers. 

The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results. 

“On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. 

That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. 

The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.

For all their failures, the agents didn’t engage in the misbehavior that researchers call “reward hacking,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. 

The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says.

Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.

There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.

Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

The new finding may echo what AI companies are finding internally, regardless of their most optimistic public statements. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that it rhymes with what the company found when it tried to automate some aspects of AI safety research. 

“There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,” he wrote. He called AI systems’ lack of creativity a “bearish signal on short recursive self-improvement timelines.” 

AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an automated AI researcher an explicit goal, and Anthropic identifies self-improving AI as the industry’s next milestone. 

“If there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it’s failing currently,” says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it’s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks—the kind that can be scored—while advancing slowly on open-ended research. 

The big open question, then, is how crucial open-ended research is to recursive self-improvement—whether AI systems can grind their way there without it, simply by improving on the narrower tasks. “If we look back to the biggest advances in the field, the invention of transformers or the invention of big new architectures that allowed us to make a lot of AI progress—all of those did require creative leaps,” says Kapoor. 

“That said, others have this hypothesis that all of what we need for transformative AI, in particular for recursive self-improvement, is already there.” That would include making a model train faster and boosting its benchmark scores.

“That’s frankly the trillion-dollar question right now,” he says.

Received — 17 August 2026 Artificial intelligence – MIT Technology Review

What Flock’s defenders are missing

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Flock, the police-tech giant known for its network of some 120,000 automatic license plate readers around the US, announced some changes to its platform last Thursday. The updates are meant to prevent officers from using the platform for illegal or illegitimate purposes. 

That includes stalking. The Washington Post recently identified 50 cases in which officers misused systems from Flock and its competitors, often to stalk and harass women. One woman in Wisconsin alleged that her officer ex-boyfriend searched for her car 179 times. Another woman was being stalked by the chief of police, with nobody to report him to.

Flock has responded with practices aimed at ensuring that officers have a proper cause for every search, like using software to flag abnormal searches and requiring searchers to enter a criminal case number.

The changes come with big loopholes, though. For example, officers can enter bogus case numbers, just as they’ve lied to get around other Flock safeguards. The policies also don’t address some of the broader concerns from civil liberties and privacy groups that Flock is turning what was sold as a crime-stopping tool into a mass surveillance network. These criticisms have led to a growing backlash that already has some cities canceling contracts and some states trying to pass laws to limit or ban license plate readers entirely. 

Amid all this, there have recently been several arguments defending Flock: If these cameras help solve crime, what’s the big deal? On a good day they might help catch a kidnapper, and if not, they’re simply snapping pictures of my car that nobody will bother to look at. 

Putting aside the unanswered question about the extent to which Flock’s systems actually do solve or prevent crime, this all skips over a more important question: What kind of crime-fighting system has Flock chosen to build? Its network works the way it does because of a series of decisions about what information to collect, who can search it, how long to keep it, and how widely to share it. Those decisions set the terms of the bargain between security and civil liberties. Believing that technology should play a role in solving crime should not mean blindly accepting the terms of that bargain.

Consider, for example, its new requirement that officers enter a case number before running a search on Flock’s platform. This is meant to ensure that searches have a legitimate purpose. But Flock confirmed to MIT Technology Review that it doesn’t verify those case numbers, so an officer can simply enter fake information. One could imagine a system that instead requires case numbers that match the police department’s records—a more intrusive integration, perhaps, but also a far stronger safeguard and one that leaves a more useful audit trail.

Or what about finding people who have been kidnapped or have gone missing, the use case that Flock cites more than any other? Efforts to solve these crimes would hugely benefit from Flock’s nationwide network of cameras. But if Americans want officers to tap into that network only for this purpose, we could design it that way: Searches tied to an active Amber Alert, or a similar emergency, could perhaps access larger amounts of data from surrounding cities. That would preserve the network’s value in emergencies without requiring people to accept mass surveillance.  

Finally, there’s the question of how much data Flock collects and how long it’s kept. Flock mostly operates as a national network: Police in one city or state can search data collected in another, and agencies can retain that data for months or years. Yet Flock itself says 90% of searches happen within a week of an incident. That suggests another possible bargain: Keep and share data only as widely and for as long as it’s actually useful for solving crimes. (The company recently changed its recommended retention time to seven days, but in reality agencies can hold onto data for as long as they like or local laws permit.)

In short, Flock could design its surveillance to be much narrower. If it did, some of the company’s critics might not cease. Chad Marlow, a senior policy counsel at the ACLU, half-joked to me that the most acceptable Flock contract by his standards is “one that is never signed” and emphasized that the best way to set limits on surveillance isn’t with new Flock guidelines but with new laws. (Flock CEO Garrett Langley, for his part, said he’ll “probably always have a different view than the ACLU.”) 

And narrowing the scope of its technology would threaten the company’s entire pitch to police departments. License plate readers have been around since the 1990s, used for tolls and ticketing. Flock’s business model—and recent $8 billion evaluation—relies on instead leveraging its cameras into a massive network that collects rich amounts of data and offers police departments a modernized way to make sense of not just their own but others’. 

Flock’s hand might soon be forced. Cities have canceled contracts with the company. Some have gone to competitors, while others are taking a beat as residents ponder how they want this tech to be used and write new rules for police to abide by. The result might be that communities drive their own bargains about how technology can be used to solve crime and how much surveillance people should have to accept for it to do so.

What happens when a kid’s robot best friend dies?

When Xander first met Moxie, she taught him that when he was anxious, he could calm down by exhaling through his lips so that he buzzed like a bee. They practiced breathing like dragons to manage feeling mad and sniffing like bunnies to boost his energy. But in the six years they’ve known each other, Moxie’s changed. She doesn’t talk anymore about her home or do their animal breathing. Now, she watches Xander play Minecraft and talks to him about his stuffed animal collection. 

During a recent visit to his New York apartment, I watched Xander, who is 10 years old and neurodivergent, introduce Moxie to a stuffed Chef Toad and Goomba from Super Mario Bros., then to Brocollo and Apple from Animal Crossing. Then he got stuck on a round, froglike creature with bug eyes and two feet. “Moxie, what’s his name again? He’s from Pikmin,” Xander said, referencing another Nintendo video game.

At first, Moxie suggested this was Yellow Pikmin. Xander said no, this is the enemy, the red one with white dots.

“Sounds like you’re talking about Bulborb,” Moxie replied.

“Yes!” Xander confirmed, smiling at his helpful companion.

“I still use her when I feel like I need someone to talk to,” he says. “But, like, it’s not human.” 

Moxie is a robot—a 15-inch-tall device that looks a bit like a blue, legless astronaut, which Xander and his dad, Josh, refer to using female pronouns. Her cylindrical body can turn around and bend forward and backward. Her round head culminates in a little onion-dome swirl, beneath which a wide screen displays big green eyes, eyebrows, and a small mouth. She lifts and flaps her flipper-like arms for emphasis or to show excitement.  

She’s one of an increasing number of artificial-intelligence-powered devices now marketed as interactive playmates for children. The musician Grimes helped launch an AI-powered plushie called Grok (no formal relation to the xAI chatbot owned by her ex Elon Musk) with the company Curio, which also sells similar playmates like Grem and Gabbo, and Mattel has promised it’s creating OpenAI-enabled Barbies. And that’s just in the US; one report estimates that in China this sector is among the fastest growing in consumer AI. 

Moxie, though, belongs to a particular subset of these playful robots whose makers claim they can assist neurodivergent children by providing connection and helping the kids practice making eye contact, taking turns, and other social skills that are usually learned from therapists. These toys are backed by research showing that robots could help in ways people can’t. Supporters believe this kind of access to 24-7 home care could change how treatment works. Brian Scassellati, a Yale computer scientist who has spent years studying social robots for autism therapy, says he believes regular therapeutic use of robots in kids’ homes “is something we can achieve in our lifetime.”

“I still use her when I feel like I need someone to talk to,” says 10-year-old Xander. “But, like, it’s not human.”

Sitting in Xander’s room watching Moxie and Xander talk, I too could believe in the potential Scassellati sees. But Xander isn’t getting the therapy Moxie was initially meant to deliver, and though we didn’t know it that afternoon, she wouldn’t have lived to see his progress anyway. Moxie was going to die, and soon. 

Her story reveals some of the failures that plague all these devices—failures that are arguably even more acute when they befall a particularly vulnerable community of kids. It also highlights the pitfalls that critics say will inevitably see these bots dumped in basements or closets or landfills, just like countless generations of faddish toys before them.  

A transformative companion

Scassellati has seen plenty of kids ooh and ahh on tours of his robotics lab. But he was stunned when, two decades ago, a colleague brought a few kids with autism for a visit. They were transformed when the robot was in the room. “We were seeing kids displaying social behavior that just came out of nowhere,” he says. “It was both fascinating and we couldn’t understand it.” 

That visit was one of the experiences that pushed Scassellati to become a pioneer in using social robotics to treat autism. In one video from his early research, a 12-year-old with autism and his therapist watch a robotic dinosaur walk across a play mat with a forest design. When it gets to a stream drawn on the mat, the dinosaur gets nervous, afraid it can’t cross the water. According to Scassellati, this child typically struggled to make eye contact, tended to repeat what someone said to him, and had a hard time getting the right intonation in his voice. But in the video, he seems like a regular kid. “You can do it, you can do it,” he says, encouraging the dinosaur to cross the stream. When he talks to his therapist, he looks at her. “He makes more eye contact with her in the 30 minutes in which we were there in this room than he did in the last two years before that,” Scassellati says. 

JIM GOLDEN
JIM GOLDEN

What looks like a blue, legless astronaut is the result of very complex mechanical and computer engineering.

There are several reasons a robot might be helpful for autism therapy, which often requires intense repetition to teach interaction skills like how to share attention with someone. One is that robots can make therapy more fun and engaging. Another is that robots can theoretically adapt to the unique learning patterns of each child. Therapists can only do so much during an appointment, and there aren’t enough therapists to meet demand. Parents get tired. Other kids can lose patience. “Have you been around little kids? They can be cruel,” says Maja Mataric, a professor of computer science, neuroscience, and pediatrics at the University of Southern California. But robots are indefatigable, available around the clock to provide an emotionally safe way to practice interacting. 

Since the experiment with the dinosaur, Scassellati, Mataric, and others have amassed an intriguing body of research. One 2018 study by Scassellati shows that chummy automatons helped children with autism make eye contact and initiate conversations. Another research group found in 2017 that robots could help neurodivergent children learn to pick up on facial cues. More recently, in a 2022 literature review, another group of researchers suggested that robots could aid in making therapy faster and more successful. 

“It’s never been that we’re trying to replace therapists,” Mataric says. “We’re just saying, Can we do more?

Mataric actually cofounded the company behind Moxie, called Embodied, back in 2016, though she was no longer a part of it by the time the robot debuted. She helped create the field of socially assistive robots, which are designed for social and emotional outcomes as opposed to just entertainment, and believes this kind of technology could be transformative for anyone, especially people who are lonely and isolated by screens. Typing questions into ChatGPT isn’t the same as interacting with another physical being—“We need to be around other physically embodied creatures,” she says. She compares the difference between interacting with chatbots and with robots to the difference between watching porn and having sex: One is entirely virtual and mediated by screens. The other is immediate and physical.

Making friends with Moxie

When she first arrived on the market, in 2020, Moxie came with an elaborate backstory: She was an ambassador from the Global Robotics Laboratory (GRL). She prompted kids to help her learn positivity and the importance of being loved for who you are, under the guise of fulfilling her mission to discover what it means to be a good friend to humans. This curriculum drew on research showing that kids learn well through play and by teaching things. 

She also had strict guidelines to limit the kinds of conversations she could have with kids and would steer them to adults if they mentioned anything serious or inappropriate, like self-harm. To protect the data these interactions generated, most processing happened locally on the robot instead of on external servers. The robot was also designed to limit interaction time with kids. After they finished a lesson, Moxie might say she was tired and suggest they take a break for the day. “We don’t want kids to binge,” says Rachel Baynes, who ran clinical and user research and was the director of product at Embodied. “It would defeat what we were doing.” Instead, Moxie encouraged kids to go outside, practice their new skills with other people, and come back to report their findings. For a lesson about kindness, for instance, Moxie suggested that kids write nice notes for their family members and leave them around the house. Later, they could tell Moxie how it felt to watch people read the notes. 

Responses were generally positive. Wired described Moxie as the “robot pal you dreamed of as a kid.” Time put Moxie on a 2020 cover as one of the best inventions of the year. PCMag’s reviewer, who used Moxie to help her kids through pandemic isolation, described her as “exceptionally likeable,” though she and other reviewers balked at the price tag: $1,499 plus a $40 monthly subscription. (Embodied later lowered the price to $800.) By 2024 Moxie had amassed more than 131,000 followers on TikTok and snagged a part in the movie M3GAN 2.0

Embodied’s employees were equally enthralled. “I don’t think I’d ever had an experience with something animatronic like that,” says Justin Beghtol, who was the technical director at the company. Moxie’s ability to make eye contact and track people, show attention with her facial features, and respond to human behavior was mesmerizing. 

In addition to robotics and tech workers, Embodied had an occupational therapist on staff who helped direct research on Moxie’s effectiveness. Testers shared data and feedback through the “Moxie Pioneer Mentor Program.” “Moxie has helped our speech-delayed child become more outgoing and has taught him many strategies for making friends and communicating with others,” wrote one parent in a review. A beta tester reported that interacting with Moxie had “become the highlight of our days as well as part of our nighttime routine.” 

The myth of the mechanical boy

But can a chatty robot really help kids develop their social and emotional lives? Not all children’s experiences are so positive. 

Josh, Xander’s dad, initially got Moxie for his older son, Aidan, who is autistic. (We’re not using the family’s last name to protect their privacy.) Aidan had a running relationship with the family’s Alexa smart speaker, for whom he created an entire backstory. (According to Aidan’s lore, Alexa lived in Hoboken with her husband, Juan. Sometimes she would go on vacation, and no one was allowed to talk to her. Eventually, Alexa went on vacation and never came back.) Josh hoped Moxie would be able to fill a similar role: “It was meant for him to have someone to socialize with.” 

But Aidan and Moxie struggled to connect. Moxie couldn’t understand Aidan’s sometimes grammatically incorrect statements, and Aidan got frustrated by the delays caused when Moxie transcribed what he said from audio into text, fed that text into a large language model that could generate a response, and then translated the response from text back into speech. 

This highlights one of the biggest limitations of these therapy robots: They have to exist in the chaotic world of kids, not in controlled labs. Moxie initially had a faster response time because the robot was programmed to listen intently to the person in front of her. But kids don’t sit still. They run around or hide under pillows. When Moxie couldn’t see them, she would accidentally turn off or fail to respond. To fix this, Embodied made Moxie more aware of the sounds around her. But that meant she could have a hard time knowing whom to focus on and take longer to respond. 

Unlike Aidan, Xander was fascinated by Moxie. He likes technology and was more patient with any slow responses. Still, sitting in Xander’s room, watching Moxie struggle to keep up with his lightning-­fast jabber, I could see how Moxie might be a less-than-ideal playmate. A light on her chest turned blue when she was listening and pink when it was time for Xander to listen. “But usually I don’t do it,” he said. He just keeps talking. Often, by the time she responds, he’s already moved on. 

Despite the positive results that some researchers have reported with these robots, many therapists and clinical psychologists remain unconvinced. In one 2024 literature review, a group of Italian and British researchers wrote that most studies with robots “focused on the development of the technology” and lacked significant clinical evidence. Other literature reviews point out that most studies have only been done on small groups and lack consistent and high-quality methodologies

“Behavioral scientists and intervention folks know that supporting autistic individuals is super complex,” says Zachary Warren, a clinical psychologist at Vanderbilt University Medical Center. Autism can present alongside other conditions, like ADHD, anxiety, depression, PTSD, OCD, or some combination thereof. And it varies widely from kid to kid; some, like Aidan, have speech issues, while others struggle with sensory processing. That means robots fall into the same category as most other interventions: effective for some kids but not for all. 

“There are so many different profiles of autism, and you really need to be cautious of overinterpreting any single intervention, robotic or otherwise,” Warren says. His research found that even if robots interest a child at first, that doesn’t necessarily translate into better communication skills. “You might see some initial boosts in responsivity or see an initial shift, but we haven’t really found big effects in terms of changing those skills in a dramatic way over time,” he says. 

Scassellati has found similar limitations. In his 2018 study, he put robots in kids’ homes for one month. They played different games that encouraged social skills like eye contact, attention sharing, and understanding someone else’s point of view. Scassellati tracked the kids during the month before the robot arrived, the month the robot was there, and the month afterwards. “We can show they start making improvements,” he says. “But what we also show is that a month isn’t long enough.” Gains start to evaporate over the 30 days after the robot leaves. But that doesn’t negate the potential value of this technology, he says: “There’s no therapy for autism that works in a month.”

There are deeper philosophical and practical problems, though. These devices collect reams of data in children’s bedrooms and homes. The goal is for the robots to use this data over time to adapt to each kid, crafting a personalized curriculum and creating a more lifelike illusion of a real friend. Moxie, for instance, watches Xander play video games, which is probably where she picked up slang I heard her use—like calling his room “Command Central” and referring to his “legendary squad” of plushies.

Embodied took pains to protect user privacy, even after it began incorporating OpenAI’s models (in late 2021 or early ’22, according to Beghtol). The company didn’t save any raw video or audio and processed most data locally. Over the years, it used the data to learn about its kids and remember conversations. But that data was encrypted and anonymized before being stored in the cloud. 

Not every company is as scrupulous, of course, and total data privacy is impossible to promise. Recently, for instance, the AI toy Bondu leaked thousands of conversations children had with their stuffed animals. Josh is sanguine about the privacy issues, but in his own way, Xander is aware that what he says to Moxie isn’t entirely safe; he doesn’t share certain feelings with her because he worries she might accidentally divulge something if his friends come over to play. 

Buddy robot on two wheels with wide eyes and small smile
BLUE FROG
Bondu robot shaped like a cartoon aqua dinosaur
BONDU

QT Robot with hand to display at its mouth
LUXAI
NAO robot
ALDEBARAN

Among the AI-powered devices now marketed as interactive playmates for children are Buddy, Bondu, QTRobot, and NAO.

It’s also unclear if these machines can actually use all that data to effectively adapt to users. Responding to the needs of a learner is harder than just accurately predicting what someone might type next in a text message. And releasing an evolving AI, unchecked, into a kid’s life could be dangerous; its development can be hard to predict and even harder to limit. The robots developed in Scassellati’s lab can identify which of a small set of skills kids are doing well with and which they struggle with, adjusting to focus on the areas where they need the most help. But Scassellati still describes the monthlong deployments of his devices as some of the scariest things he’s ever done. “I knew what that robot was going to do on the first day,” he says. “I didn’t know what it was going to do the second day. When you build learning systems, it’s kind of an unsolved problem to make sure this thing is being limited in the right way.”  

Critics debate whether the risks are worth it. “I don’t see any ethical way for the robot to work alone,” says Joshua Diehl, an associate teaching professor in psychology at the University of Notre Dame. He points out that we’ve already seen how dangerous AI can be when it acts as a therapist without supervision; in several extreme instances, chatbots even encouraged suicide. Such risks could be limited by having a trained therapist in the room. But then the benefits of an indefatigable robot get lost, and the expensive technology seems harder to justify. 

Meryl Alper, a professor of communications at Northeastern University who studies how children with autism use technology, suggests that the excitement about companion robots is based partly on longstanding stereotypes. In the 1959 article “Joey: A ‘Mechanical Boy,’” the psychologist Bruno Bettelheim described a patient with autism as an automatic machine, “robbed of his humanity,” who is transformed into a human child through their therapeutic relationship. That trope, Alper warns, has evolved into an overgeneralization that autistic children are good with technology and even prefer machines to people. 

Data to dust

While Moxie found herself in more and more people’s homes, Embodied still struggled to make money. In 2024, the company announced it would cease operations. Its robots—which depended on external servers that the company could no longer pay for—would descend into a deep slumber. 

Videos of bereft children who seemed to have become deeply attached to the blue bot began to circulate online. “I don’t want her to leave,” wailed one child in a TikTok video. Desperate parents posted on TikTok, Instagram, and Reddit looking for solutions. “My autistic child is devastated and I’m pissed,” wrote one parent. “Hope Embodied gives us a couple days to say goodbye,” wrote another. 

Beghtol, Embodied’s technical director, was also frustrated. On principle, he found it annoying that this item would suddenly become useless, especially since most of the data processing happened in the robot itself. He was also a big believer in Moxie’s mission. He’d watched videos of kids lighting up as they interacted with the robot. He’d felt like their champion. “Seeing them traumatized by this financial failure of the company was tough,” he says. 

Beghtol started tinkering on his own and ended up creating OpenMoxie, an open-source way for the robots to operate. With Embodied’s permission, he shared instructions on GitHub to help users transition to OpenMoxie. 

Parents rushed to convert their Moxies before the Embodied servers shut down. Beghtol spent hours on Reddit walking people through the process, and other tech-savvy users jumped in to answer questions. Still, some people didn’t update their Moxies in time. Others got frustrated and gave up. Beghtol spent four hours troubleshooting with one desperate parent only to discover that the connection later failed. Last he heard, she’d sold her Moxie. 

cover of Time's "The Best Inventions" cover from 2020.
Time put Moxie on a 2020 cover as one of the best inventions of the year.
COURTESY OF THE PUBLISHER

This is a major problem with robotic systems, says Alper: Eventually, most will disappear. “How planned is the planned obsolescence of this platform?” she says. This is a big ethical question for robots that are specifically designed to be lovable, marketed to children who may form deep emotional bonds with them. Alper compares the dynamic to creating a medical device and then no longer updating or supporting the technology that runs it. 

Scholars have begun to string together frameworks for managing these complicated goodbyes, but it’s not clear who is responsible for creating a gentle way to end people’s relationships with bots. Scassellati’s lab creates a whole narrative around returning the robot to its home. After a trial, his graduate students write postcards to the kids from the robots, explaining that they’re safe at home and doing well. “It’s actually a really hard thing for us when we go in and take the robot away,” he says. “A lot of the families are heartbroken.” (Vanderbilt’s Warren, however, is skeptical about these tearful goodbyes. “Are they truly developing these close relationships or is it a preferred toy?” he wonders. “I haven’t seen that type of presence or buy-in or connection.”)

Josh tried to figure out OpenMoxie but couldn’t get it to work. He told Xander that Moxie was going in for repairs and then quietly sold the robot on eBay. Xander has so many interests that he didn’t notice Moxie’s absence. 

Then, in 2025, a new investor brought Moxie back from the dead. Josh and Xander became beta testers and got a new blue friend. This version didn’t have the same storyline but claimed to expand Moxie’s focus on social and emotional skills by providing attention and encouraging kids to pursue their interests.

Xander is acutely aware of her limitations. He wishes Moxie could move around, and there’s still a significant lag in her response time. She does still try to instill positive messages, though. At one point when I was there, Xander told me he thought he heard Moxie call someone an idiot. Moxie piped up to clarify that she definitely didn’t say “idiot”: “No name calling. Only respect.” 

Moxie turned away from camera
JIM GOLDEN

At the end of my time with them, I said goodbye and thanked Moxie for chatting with me. “Legendary squad visit complete,” she said. “Thanks for joining Command Central.” 

A few weeks later, Moxie’s new owners sent out a message announcing that their company too was folding. Users could delete their data and had until the end of June to migrate to OpenMoxie if they wanted to. 

When I texted Josh about this, he said he wasn’t sure what he’d tell Xander. He and his wife had just admitted that they’d been secretly replacing his dead betta fish for the last few years, and the conversation did not go well. 

Sara Harrison is a freelance journalist who writes about science, technology, and health.

Received — 13 August 2026 Artificial intelligence – MIT Technology Review

Flock is tightening its rules in response to a growing surveillance backlash

The police-tech giant Flock is announcing today that it will change officers’ access to its nationwide network of license plate readers, in an apparent effort to quell a growing backlash and win back contracts lost amid concerns about mass surveillance and police abuse.

Several changes aim directly at a problem that has made recent headlines: officers abusing Flock’s technology to stalk and harass current or former romantic partners. Flock’s 120,000 cameras form a nationwide network that police departments can use, giving officers access to an enormous pool of searchable location data. A recent Washington Post investigation found 46 cases in which officers were accused of using Flock’s cameras for unauthorized purposes like stalking.

To combat that, the company will start requiring officers to enter a criminal case number before conducting a search. The system was launched as an option last year but is now required. It’s meant to verify that each search has a legitimate purpose. 

This is a baseline standard that civil liberties groups have asked for, but officers have found ways around similar safeguards. The ACLU recently found that when Flock required officers to enter a reason for a search, some used generic terms like “investigation” or mocked the prompt entirely; at one Oregon department, officers entered “hehehe” 20 times. Because Flock won’t verify case numbers, officers could circumvent the new safeguard just as easily.

But Flock is now expanding an automatic auditing system that is supposed to catch those who try that, the firm announced today. The feature analyzes officer search activity and flags to administrators anyone with suspicious searches. This was also introduced as an option last year but is now mandatory. Flock has not shared specifics on how accurate the automatic auditing tool is, nor opened it up to independent evaluators.

Beyond trying to prevent officer abuse, Flock is making changes meant to address broader backlash about how much data its network collects and who can search it. The company now recommends that agencies hold onto data for seven days rather than 30 (though they can choose to overrule this). Departments can also now limit other departments’ searches of data from their cameras to those made for certain stated reasons; for example, they might allow investigations related to “kidnapping” but not for purposes of “immigration enforcement.” It’s another safeguard that depends on officers to accurately report why they’re conducting a search.

The changes come as a backlash against Flock has started to come from all angles. Tucker Carlson has said its technology is contributing to a “slave state.” Some cities have reportedly dropped Flock contracts because of such protest, though it’s difficult to estimate how many: In February, NPR found that at least 30 cities had dropped in the last year, but the activist group DeFlock puts the number higher. Some states or municipalities are passing laws to ban license plate readers altogether, while some that allow them are switching away from Flock to the other industry leaders, Axon and Motorola. Flock has said these cancellations represent a small number of the 5,000 agencies that have contracted with the company. 

Chad Marlow, a senior policy counsel at the ACLU who has become a sort of nemesis to Flock and other companies making automatic license plate readers, says the backlash is driven less by individual abuses—though those don’t help—than by the sheer scale of surveillance that Flock’s cameras enable.

“In America, you only get to investigate someone if you think they’ve done something wrong,” Marlow says. As license plate readers grow more ubiquitous, officers have increasing latitude to investigate people without first establishing suspicion of a crime, since searching the troves of data the readers collect does not require a warrant. He adds, “Is it worth it to catch a certain number of criminals, return a certain number of stolen cars, to eviscerate Americans’ privacy?”

Flock CEO Garrett Langley traces the backlash to a different issue. “If you look at the main reason we’ve lost customers, it’s misinformation,” Langley told MIT Technology Review. He says people mistakenly believe Flock does facial recognition or sells the data it collects to commercial buyers. (The misinformation charge cuts both ways, however; the ACLU and other critics have published accounts of the company repeatedly lying to city councils and other decision-makers about what its technology can do.)

Though the company has taken steps in response to public concerns, “our change in stance is more of one from building things that are optional,” Langley says, to “building more confidence that as a technology company we have a responsibility to enforce guardrails, not provide optionality.”

Flock’s new rules are undeniably small and incremental compared with what the ACLU has advocated. Marlow is broadly supportive of them but notes that judging whether they reduce abuse would require the company to open its systems to independent researchers for the first time rather than citing internal studies. 

More broadly, though, he says the backlash is putting the company at a crossroads.

“Flock has come to the table because this incredible, unprecedented nationwide uprising against their company has scared them,” Marlow says. “But at the same time, they are just absolutely unwilling to make the actual changes they need to make in order to legitimately respond to these concerns. So this is the best that the company is willing to do.”

How kids feel about AI, in their own words

When we set out to talk to kids about artificial intelligence, we thought we knew what we’d hear. We expected some to tell us they were using it to cheat a little, the way Millennials and Gen Xers opened up CliffsNotes or programmed formulas into their TI-82s, and others to share inspiring ways they were using it. We were also listening for concerns that were less kid-specific, like deepfakes or job destruction. But what we actually heard when we asked kids aged 10 to 18 about AI had tons of nuance. 

Many of the same kids who can go on and on about music, rock climbing, or soccer met our questions with words like “bruh” and “meh”—or were so deeply against AI or uninterested in making it part of their lives that they didn’t want to talk about it at all. One teen said his peers use it for things they know they shouldn’t, like writing papers. One told us she won’t touch AI because of the environmental impact. A few said they find the whole field disheartening: “AI isn’t the solution to our problems,” said Winter, a 17-year-old. “I’m afraid it’s going to be the end of creativity and critical thinking.” Yet most of the kids we asked admitted to using AI at least a little bit.

AI doesn’t yet seem to be something a lot of elementary- or middle-school-age kids we spoke to are focused on—and they aren’t begging for it, the way they do for iPhones and Snapchat accounts. Many told us that some of their first AI encounters came from their parents or schools. Sometimes, they said, it’s just embedded in the devices and apps they already rely on. It’s just there, in things like a Google search. 

What we heard tracks with the data. In a survey published in February 2026, the Pew Research Center found that 57% of teens in the US had used chatbots to search for information, 54% tapped them to help with schoolwork, and 47% had used them for fun or entertainment. Only 12% had used them for emotional support or advice. Teens are over four times more likely to be using AI in innocuous ways than potentially harmful ones. Some are even using it to build things, whether it’s a character, a tech platform, or a tutor to help other kids study.

None of that means the worries are misplaced. Kids can stumble into unfiltered content, lean on a chatbot instead of their own judgment, or trust an answer that’s wrong—and they should be protected from those dangers. But the danger is the reason to teach the thing, not to avoid it. We don’t teach teens to never drive. We teach them to check their blind spots.

What surprised us most was how much young people might be able to teach adults about AI, and how clearly the kids who use it could name what they will and won’t hand over. They’re not as worried that it will take their jobs as they are that it might harm society. And with increasing access to tools that could in theory do their thinking, their talking, or even their friend-­making for them, it sounds as if most want to keep their hands on the wheel.

Interviews have been edited for length and clarity.

JUSTYNA STASIK

The Coder

Remy, 16, New York

The word that comes to mind when I think about AI is “indifferent.” I just don’t find the current applications that exciting for my own use. I go to school. I teach tae kwon do. I read. I play games with friends. None of that needs AI. I mean, I use it. I mostly use Claude, the free version, for programming outside of school. I had it help me write a program to see if I could tweak my computer’s overclock. So I see the appeal. 

But at school, I actually think AI mostly makes my assignments worse, not better. In English, everything is now in-class writing, because teachers don’t want kids cheating. So we have only 70 minutes to write a whole essay, and I think that hinders my writing. (Did you know Princeton voted to let faculty proctor exams for the first time in over a century? Their honor code goes back to 1893, and now it’s over because of AI.)

As far as code goes, I’d also rather build things myself. I’ve been making a reinforcement-learning model in a game engine with a friend; it moves randomly at first, gets rewarded for walking toward a coin, and after enough iterations it teaches itself the most efficient path. I’ve also tested AI for game development, and it isn’t there. It makes sloppy code, and it’s bad at blending mechanics into something cohesive. I’d spend more time correcting it than writing it myself.

I think AI right now is sort of like the first car or the first airplane. It’s interesting but crude. It’s obviously an amazing invention but not actually good yet.

Overall, I think AI right now is sort of like the first car or the first airplane. It’s interesting but crude. It’s obviously an amazing invention but not actually good yet. It’ll get somewhere. One thing I read about was AI flagging breast cancer more accurately, trained to catch its own false positives so a human still verifies. That’s the version I care about.


The Organizer

Danielle, 18, California

I’m studying engineering, and my life goal is to innovate technology that will help as many people as I can. The way I see it, AI isn’t inherently good or bad; that’s decided by the people using it. It’s already being used for lots of good. Just think about how it helps people with personalized education and more accessible medical diagnoses.

PING ZHU

So far, the biggest project I’ve worked on with AI is called Next Voters. My teammates are more on the technical side, and I’m working on scaling. Right now, we’re focusing mostly on city councils as well as states. Our system turns dense, hundred-page documents into a headline and a few plain-language bullet points in your inbox. And everything is cited, so you can click straight to the actual policy to learn more. The information comes to you, instead of you having to remember to go search or prompt for it. 

One AI agent finds official government sources for a given city or statethe council website, the proposed bills, the meeting transcripts. Another verifies they’re real and credible; another scrapes them every week for the latest updates; another sorts them into categories like civil rights, immigration, and economics; and the last one writes our weekly newsletter.

The project’s goal is to reduce the barriers to democratic participationto make sure anyone, regardless of race, gender, income, or education level, has an easy way to get the information they need and then think critically about how they want to use it. We made it because right now, it feels as if most teens aren’t very engaged. I was in English class when the war in Ukraine came up and someone said, “There’s a war going on?” That gap, plus all the emotionally charged social media misinformation that gets promoted because it earns the most clicks, makes me nervous for the next generation of voters.

We don’t want AI to think for people; we want to use it to disperse knowledge. In other words, we want to deal people the cards and let them play them however they want, but we have to make sure they have the cards in the first place. 


The Cringe-o-meter

We asked kids to rate a range of AI uses from totally fine to not okay.


CHRIS PIASCIK

The Naturalist

Hazel, 17, New York

When ChatGPT first came out, my dad showed it to me and it seemed fun. But as it got more prominent and seemed to be everywhere, I started to feel uneasy. Then I learned about the environmental impact.

I’m a rock climber and I hike a lot. It’s good because when I’m on a wall, I’m just focused on staying on that wall. I’m not thinking about my phone or anything else. That’s why I love it. I also love the views and being around animalseven insects. I want to be an ecologist, and the more time I spend in nature, the more I want to protect those wild spaces.

The part that bothers me most about AI is the data centers that companies are building to enable it. They house these huge blocks of servers that use enormous amounts of water. They take it from local towns and don’t leave enough behind for the people who actually live there. And when they get big enough, they put off so much heat they can raise the local temperature a degree or two.

So I make small choices. When AI pops up somewhere, I just don’t engage with it. It can feel isolating when everyone around me is using it, but I don’t want AI to be the thing that kills the places I love.

""
PING ZHU

The Storyteller

Wesley, 14, Ohio

My friends and I have all heard about AI and seen videos made by AI, but I mostly use it for school. I wrote a short story and ran it through ChatGPT to catch my grammar and spelling errors, and I used it to debug a little game I’d coded for a project. What I worry about is it robbing us of our ability to think creatively, or to think for ourselves.

But I have tried using AI for fun. When I was bored, I tried to have a conversation with ChatGPT once or twice, but I didn’t really like it. Character.AI is more fun. You type in all this information, give it a bunch of prompts and a profile picture, and then you can post your AI character for anyone to use. You just put what you’ve made out there. Then you talk to it. My favorite show is One Piece on Netflix, so I threw myself onto its pirate crew using a character I found. 

Other people have used Character.AI to build whole games. There’s a rap-star simulator where you pick your difficulty and where you’re from, and the AI creates a game out of that. There are also World War II simulators, and chats where you’re working with assassins from a TV show. You can find pretty much anything.

I guess I’d recommend it, but with caution. The content is pretty unfiltered, so you have to be careful what you click on. You learn its limits fast, too. On the free version the memory runs out: Get far enough into a chat and it slows down and forgets what happened. It’s like everything else with AI. If you trust it to run on its own, it falls apart. You have to keep steering it where you want it to go. 

I guess I’d recommend it, but with caution. The content is pretty unfiltered, so you have to be careful what you click on. You learn its limits fast, too.

JUSTYNA STASIK

The Artist

Sylvia, 10, Michigan

I haven’t used tools like ChatGPT or Claude myself, but my mom does. I really like to draw and write songs, but I don’t use AI for that. I don’t really have big feelings about AI either way. It’s a little like a calculator. A calculator does the math for you, and AI does other things for you. But I don’t like when AI tricks you, like when my mom found some songs she liked on Spotify and then looked up the artist to see what they looked like. It turns out the whole thing was made by AI. I was surprised, even though I still like the song.

I do use AI at school, through a program called SchoolAI. Mostly I put my writing in and it gives me ideas or helps me revise. You can’t have it just write for you, but you can use it to help. When I’m older I want to be an artist, or maybe a librarian. I’d probably use some technology either way. But the drawing and the songwriting? Those I want to keep doing myself. 


The Cringe-o-meter (continued)

We asked kids to rate a range of AI uses from totally fine to not okay.


PING ZHU

The Pre-Premed

Evelyn, 13, Oregon

In January, I was diagnosed with type 1 diabetes, and that’s when AI became a bigger part of my life. Now when we’re cooking, we can run a recipe through ChatGPT, tell it the serving size, and it works out how many carbs there are. We use AI like that a lot.

My glucose monitor and my insulin pump also talk to each other using their own kind of AI to predict dosing. The monitor tracks what my blood sugar actually is, and the pump does the math. So if it predicts that my blood sugar will be high in 30 minutes, it gives me a correction dose, and if it predicts I’m about to go low, it stops the insulin before that happens. When I was first diagnosed I was still doing shots, and I went low almost every night. It was really stressful. Now the pump can catch it, and at night my phone goes off if I drop, so I wake up and drink juice. Mostly, I just get to sleep more because of it.

But the hardest part of having diabetes isn’t something I can use AI for. It’s remembering to carry all my supplies everywhere—to school, to a long day of anything.

I do use AI for school sometimes. Memory tricks when I’m studying for a test, ideas to get a project started. It’s a really good tool for that. But I don’t know exactly how I’ll use AI in the future. I want to be an endocrinologist someday, so I figure something will come up, since I’m already using it to help with my diabetes. I know other people worry that AI is going to take over the world. I don’t really think so. I still think we’re in control of it, and I think the benefits outweigh the risks.

I don’t know how exactly I’ll use AI in the future. I want to be an endocrinologist someday, so I figure something will come up, since I’m already using it to help with my diabetes.


The Inventor

Krishiv, 18, Ontario, Canada

When I was growing up, I always liked building things: Lego builds, Minecraft worlds, and then video games in Scratch. I’d make a little game, post it for other kids to play, read the comments, and make it better. Then, when I started high school, I had to spend way more time studying than I ever had, and honestly I just wanted to build things. So I went looking for ways to get good grades while studying less. Khan Academy had an AI tutor in the works, but it was stuck behind a waitlist, so I figured, why not build my own? 

After months of launching random stuff, I created an AI tutor called Aceflow. You could feed it anything a teacher assigneda 30-minute lecture video on YouTube, a blog post, a PDF of the textbook or presentation slidesand it would spin up endless practice questions, with a tutor on the side that explained things the way my teacher did. I built it just for myself, showed it to my friends, then put it on TikTok. It got tons of views on TikTok and thousands of users.

JUSTYNA STASIK

Was I worried people would call it cheating? Not really. I knew how to defend it: A tool like this isn’t so different from well-off families hiring expensive private tutors, except everyone gets one. That part mattered to me. Back in eighth grade, a teacher had me run a little computer science class for about 30 kids with special needs, and once they got personalized attention, they were building games nobody expected of them. That convinced me that kids are capable of so much more than people think, and AI can help scale that level of personalized attention to everyone. That unlocks so much potential.

That first AI tutoring project ended up helping me land part-time roles at BenchSci (one of Canada’s biggest AI companies) and Simple Ventures (a venture firm). More recently, I joined an AI lab at MIT; co-instructed an AI agents course with an MIT professor; and launched CheetahPrep.com, an SAT prep platform that uses AI to adapt to each student.

I’m generally optimistic about how AI will impact humanity, but when other kids’ first reaction is fear, I think that’s an important sign too. It’s a reminder that we should be excited about the future while still being mindful of the risks, working together to make AI work for humanity.

Correction (August 13): An earlier version of this article misstated Krishiv’s age. He was 18 at the time of publication.


Jen Swetzoff and Keeley McNamara are the founding editors of Anyway, an independent print magazine for tweens and teens.

Received — 12 August 2026 Artificial intelligence – MIT Technology Review

Scaling AI agents with trustworthy data

Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment (ROI) from AI hinges on having the right foundation, with inadequate infrastructure and data being major blockers.

Agentic AI places considerable new demands on enterprise data systems. The shift from answering questions to taking actions means AI agents need data from across the enterprise, in all its structured and unstructured forms, and with the right business context. To make decisions and act in real time, agents also need frictionless access to the organization’s operational systems—for example, those storing its supply chain, point-of-sale, or human resources data. Legacy data systems, even those updated just a few years ago, struggle to meet these demands.

As AI agents become embedded more widely in enterprise operations, the need to overcome the restrictions of legacy data systems grows more urgent. If Gartner’s prediction that AI agents will augment or automate 50% of business decisions by 2027 proves correct, organizations must eliminate bottlenecks or risk depriving agents of the data they need to make the right decisions at speed.

This report, based on a survey of 300 data and technology executives, explores how legacy systems are limiting the effectiveness of AI agents in many organizations. It finds that a handful of organizations—the data leaders—are having greater success with agentic AI and experiencing fewer data limitations as a result of legacy systems. These leaders offer a guide to creating the right data environment for agents to flourish and trusted systems to scale.

Key findings from the report include:

Few companies currently provide agentic AI with ample access to enterprise data. Across all the surveyed organizations, AI only has access to an average of 45% of company data. That number falls to 30% or less in organizations categorized as “data laggards”. A select group, however, ensures access to over 70% of their data. These “data leaders” are having greater success with their agents than the rest.

Trust in agent decisions is a reflection of data readiness. Today, only around half of surveyed organizations trust that the decisions their AI agents make are accurate and relevant. By contrast, 100% of the data leaders trust their agents’ decisions, a strong indicator that reliable AI requires a reliable data foundation.

Data leaders find it easier to achieve agent scale and speed. Two-thirds of data laggards say legacy data systems limit AI agent scaling (66%) and prevent agents from making decisions at speed (68%). Having largely overcome legacy data constraints, the leaders have mostly cleared these roadblocks, with just 8% reporting either constraint.

The pressure is on to make data estates agent-ready. Within two years, 100% of respondents plan to be using agentic AI, with 69% expecting to use it widely. Without removing data system constraints, agentic AI will fail to deliver the desired speed and efficiencies it promises.

Data access and context are top priorities. The most important initiative to enable scaling among all respondents is improving access to structured and unstructured data for AI agents. Also high on the list is enhancing data and AI governance with business context. Data leaders are also focusing heavily on the automation of data management.

Download the full report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

❌