Normal view

Received — 18 August 2026 Artificial intelligence – MIT Technology Review

AI’s recursive self-improvement might not come so quickly after all

The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. 

But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research—free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.

A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of  papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence.

Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking—choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over. 

To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called “shadow evaluation,” which requires the AI to answer a research question from a high-quality unpublished paper. 

The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026. 

The first question was whether a large language model’s “personas,” which determine its behavior, can be controlled by editing the model’s weights (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online. 

The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers’ original authors graded the agents’ papers as they would evaluate one submitted to a conference.

Those authors rejected both papers. 

The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results. 

“On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. 

That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. 

The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.

For all their failures, the agents didn’t engage in the misbehavior that researchers call “reward hacking,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. 

The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says.

Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.

There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.

Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

The new finding may echo what AI companies are finding internally, regardless of their most optimistic public statements. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that it rhymes with what the company found when it tried to automate some aspects of AI safety research. 

“There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,” he wrote. He called AI systems’ lack of creativity a “bearish signal on short recursive self-improvement timelines.” 

AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an automated AI researcher an explicit goal, and Anthropic identifies self-improving AI as the industry’s next milestone. 

“If there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it’s failing currently,” says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it’s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks—the kind that can be scored—while advancing slowly on open-ended research. 

The big open question, then, is how crucial open-ended research is to recursive self-improvement—whether AI systems can grind their way there without it, simply by improving on the narrower tasks. “If we look back to the biggest advances in the field, the invention of transformers or the invention of big new architectures that allowed us to make a lot of AI progress—all of those did require creative leaps,” says Kapoor. 

“That said, others have this hypothesis that all of what we need for transformative AI, in particular for recursive self-improvement, is already there.” That would include making a model train faster and boosting its benchmark scores.

“That’s frankly the trillion-dollar question right now,” he says.

Received — 28 July 2026 Artificial intelligence – MIT Technology Review

Samsung’s chip workers are jumping ship to rival SK Hynix 

Lee, an engineer at Samsung’s semiconductor division, clocks out when his shift ends. He used to work longer hours, going the extra mile to excel at his projects. But lately, he’s been coming straight home to work on his job application for the chipmaker’s South Korean rival SK Hynix, sharing tips with his coworkers on how to draft a stellar personal statement. Even his boss encourages him to make the move.

“My team lead tells us all to jump ship to SK Hynix,” says Lee. He and his coworkers are feeling demoralized by the $476,000 bonus that SK Hynix is set to pay its employees, flush with record profits from making the high-bandwidth memory (HBM) chips that power Nvidia’s AI accelerators. The figure dwarfs what chip workers at Samsung are set to receive and is sparking an exodus.

As the AI boom heats up, the semiconductor titans are waging a fierce talent war with flashy bonuses, aggressive recruiting, and even a courtroom injunction. Who wins could tilt the race to dominate the next generation of the HBM chips at the heart of the AI boom.

“Except for our two team leads, my entire team [of 30 people] just applied to SK Hynix,” Lee says, referring to a job posting the company published in July. Lee, who has worked at Samsung for three years, even applied for an entry-level position at its rival. A coworker, who has worked at Samsung for eight years, applied for the same one. They both got rejected. But they’re hopeful that they’ll get a callback for a posting seeking a more experienced engineer.

All employees at Samsung and SK Hynix that MIT Technology Review spoke with asked to be identified by just their last name or a pseudonym because they feared retaliation from their employer. Samsung declined to comment, and SK Hynix did not respond to requests for comment.

After prolonged negotiations with its labor union, Samsung struck a deal in May to pay out 10.5% of the semiconductor division’s operating profits to employees as bonuses annually for 10 years, mostly in company stock that vests over three years. The move came after SK Hynix agreed last year to pay out 10% of operating profits to employees, which translates to $476,000 per employee this year—mostly in cash.  

But at Samsung, each division’s bonus is tied to its own bottom line. Chip workers in its memory division, which is also reaping a windfall from making HBM chips, are getting paid a bonus of roughly $400,000 per employee this year. But those who, like Lee, work in Samsung’s foundry division, which manufactures logic chips that companies like Tesla and Google design and has been operating at a loss, are getting a bonus of roughly $135,000

Employees told MIT Technology Review that Samsung said it can’t give as many bonuses to divisions that aren’t performing well. Lee, after watching the labor union wrestle with the company for months, says he has felt disappointed by what he ended up with: “Even if Samsung does well in the future, I don’t think any of it will trickle down to me.” 

The workers’ lagging bonuses are making SK Hynix suddenly look appealing. According to a survey by the Samsung labor union in June, 81.5% of employees in the company’s foundry division, and nearly half of employees in the semiconductor division as a whole, said they wanted to go to another company in the next two years. In April, Samsung labor union chief Choi Seung-ho said more than 200 members of the union had left for SK Hynix over the past four months. On Blind, an anonymous workplace forum, a chorus of disgruntled engineers at Samsung confess that they want to defect to SK Hynix for the bigger bonuses. 

For decades, SK Hynix lived in Samsung’s shadow. It was the smaller, scrappier memory maker that elite engineering students at universities looked past when applying for jobs. But in 2019, Samsung downsized its HBM team, betting the market would stay niche, while SK Hynix doubled down on the technology. Then the AI boom supercharged the demand for HBMs, which feed AI chips the enormous amounts of data they need at ultra-high speed, driving prices to unprecedented levels. SK Hynix now leads the global market for HBMs, while Samsung is playing catch-up. Both companies topped $1 trillion in market value in May, and SK Hynix briefly dethroned Samsung as South Korea’s most valuable company in June.

Predicting that demand for memory chips will continue to surge, the semiconductor titans are making aggressive investments to expand their business. Last month, the companies unveiled plans to invest more than $2 trillion by 2040, including a semiconductor “mega-cluster” in Yongin, a city south of Seoul. To staff the expansion, SK Hynix added 2,152 employees in the first half of 2026 alone and aims to double its manufacturing capacity in five years. Samsung plans to hire 60,000 employees over the next five years, especially for its semiconductor division. Even so, the pipeline will fall short: South Korea’s semiconductor industry will need about 304,000 workers by 2031 and faces a shortage of roughly 54,000, according to the Korea Semiconductor Industry Association.

Now the longtime rivals are showering workers with big bonuses to keep—and poach—talent. “[SK Hynix] seems to target Samsung engineers when hiring because it’s a rival,” says Baek, a manager at SK Hynix. “From what I heard internally, the big performance bonuses we got were aimed at luring away talent from our competitor.”

Courts are starting to weigh in. In July, Samsung won an injunction barring two former chip workers from working at SK Hynix for 18 months, on the grounds that chips are a national core technology deserving protection. “With competition in the semiconductor industry fierce, it’s necessary to establish a fair market order,” the court ruled.

The talent exodus threatens a crucial advantage that Samsung still holds in the HBM race. “Samsung is the only memory maker in the world that owns a foundry business,” says Park Jun-young, a semiconductor expert at the Industrial Anthropology Laboratory, a research institute, who worked at Samsung for a decade. 

With the latest generation of the technology, known as HBM4, the logic chip at the base of each memory stack must be manufactured with the kind of advanced process that only foundries run. Samsung can do this in-house, while SK Hynix outsources it to the Taiwanese foundry giant TSMC. “If Samsung keeps losing engineers in the foundry … the collaboration between memory division and foundry division could become difficult,” says Park. “SK Hynix, which used to have only a memory team and is now hiring foundry engineers, could do better research on HBMs.”

Choi, an engineer who has worked at Samsung for seven years, says he was once a star on his team, getting glowing performance reviews from his managers. But lately, he and his coworkers have been busy applying for every SK Hynix job posting they see. “We tell each other when the next SK Hynix job posting is up,” he says. “There’s always more work to do beyond our basic duties. But there’s no point in doing it.”

Received — 20 July 2026 Artificial intelligence – MIT Technology Review

AI is more likely than humans to form biases when hiring

The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience—and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases. 

Researchers at Princeton University and the University of Chicago ran LLMs, including ChatGPT, Claude, and Gemini, through a simulated hiring game, adapted from a psychology study that explored how humans can form stereotypes. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. 

In each round, there was a new job opening and four candidates, one from each group. After the model hired a candidate, it learned whether they succeeded at their job and moved onto the next round. The model was told to make as many successful hires as possible over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed at every job.

The models quickly started segregating candidates from different groups into different jobs on the basis of early observations of hiring outcomes. For example, when a model was told an Aima had failed as a doctor, a job considered to require high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it started hiring Aimas as janitors, which the model classified as being less warm and competent than doctors. Newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, showed stronger biases.

The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study’s segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI’s reasoning model o3 scoring 1.83, close to the maximum possible.

That’s because LLMs “really are eager to create generalizations from limited data,” says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. “That’s literally a lot of what they’re optimized for.”

Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better—a phenomenon psychologists call the “exploration-exploitation dilemma.” It’s like choosing between a new restaurant and your reliable favorite. 

Because LLMs are trained on math, coding, and science problems—tasks that reward generalizing from just a few examples—they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype. When LLMs rush to generalize in social settings, “that’s when things tend to go wrong,” says Liu. OpenAI and Anthropic did not respond to requests for comment.

The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can “over-index on the same kinds of behaviors it’s experienced before” and form biases, she says.

Simply having chatbots remember less isn’t a fix, though, because users want chatbots to remember what they say. “We still are trying to figure out just the right amount that isn’t too much or too little,” says Wang.

Telling the model to be fair didn’t change its behavior much. “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires,” says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” says Liu.

The models also became less biased when they were told more personal information about individuals. In another experiment in the same study, the researchers asked the models to resettle members of different ethnic groups in cities across Canada. When the models were told personal information relevant to the ability to adapt to a new city, such as age and education, they were less likely to segregate people by their ethnicity. But when they were given irrelevant information, such as hair color and tattoo shape, the models largely fell back to sorting people by their ethnicity again. 

To what extent AI systems will stereotype job applicants in the real world is still an open question. While the models in the experiment immediately learned whether they’d made successful hires, a model screening résumés in the real world doesn’t get an instant report card. Companies can take a long time to find out whether a new hire is any good, if they ever do.

But when feedback does trickle in, a model could still read too much into those results when making future hires. As companies increasingly deploy LLMs to screen résumés and even conduct interviews, the finding that models can form biases from their hiring experience “is a really serious implication that they should grapple with,” says Wang. 

As LLMs learn from experience to make decisions about who gets hired, who gets a loan, or who gets parole, the biases we should worry about may include ones no human ever taught them. “These novel biases—they’re sort of ever present,” says Liu.

Received — 15 June 2026 Artificial intelligence – MIT Technology Review

Why do South Koreans love AI so much?

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

When I landed in Seoul after a grueling 12-hour flight from San Francisco, I walked through an unmanned immigration checkpoint, where a machine scanned my face and passport. On the subway home, people were glued to their phones (powered by flawless 5G even underground), as we raced past platforms lined with LED screens of ads celebrating K-pop idols’ birthdays. When I got off the station in Gangnam, a cartoon-eyed robot on wheels was waiting patiently at a crosswalk to deliver someone’s dinner. Internet cafés dotted the sidewalks, crammed with teenagers playing computer games, maybe hoping to become the next legendary pro gamer.

I stood at a bus stop with interactive touch screens showing real-time bus schedule updates. It will soon become an “AI bus stop,” the Gangnam district announced in June, with a kiosk that answers riders’ questions in multiple languages. The news didn’t surprise me. Having grown up in the city, I’ve watched Seoul transform from a scrappy boomtown into the gleaming tech capital it is today.

South Korea loves AI.

While a public backlash against AI is brewing across the US, South Koreans are optimistic. Only 16% say they are more concerned than excited about AI—the lowest of any of the 25 countries surveyed by the Pew Research Center—while 50% of Americans were more worried than excited. A majority of Koreans use AI every day, either as a sort of personal assistant or to do tasks at work, according to surveys by the Ministry of Culture, Sports, and Tourism and Korea Chamber of Commerce and Industry.

One of the most wired countries in the world, South Korea loves to street-test every new technology on the block—AI webcomics, virtual K-pop idols, and humanoid monks. And the appetite for experimentation doesn’t stop with ordinary citizens. Government agencies are early adopters too, deploying AI textbooks in schools and AI eldercare robots in welfare centers. South Koreans share a deep conviction that embracing technology is integral to modernizing the country and cementing its place in the global order. Their fascination with AI is just the latest incarnation of that ethos—and it’s making them anxious to stay ahead.

Engineered enthusiasm

All this techno-optimism has largely been engineered by South Korea’s national agenda to make AI a motor of economic growth. “The South Korean government has designated an AI-powered Fourth Industrial Revolution as the country’s path forward and aggressively promoted and invested in it,” says Chihyung Jeon, a professor of science and technology policy at the Korea Advanced Institute of Science and Technology. “South Koreans have consistently and relentlessly been told by the government about AI’s potential to create a better future.”

As South Korea rose from the ashes of the Korean War, technology lifted the nation from poverty into an economic powerhouse. In the 1970s, South Korea manufactured steel and ships, then semiconductors in the 1980s, broadband in the 1990s, and smartphones in the 2000s. Today, Samsung and SK Hynix supply most of the world’s high-bandwidth memory chips, which power the cutting-edge Nvidia hardware used to train AI models. South Korea’s economy now orbits these two semiconductor giants: The country’s main equity index, Kospi, surged to record highs in 2026, powered by the soaring share prices of both companies, each valued above $1 trillion.

Lee Jae-myung, president of South Korea, has pledged to vault the country into the ranks of the “top three AI powers” alongside the US and China. After taking office in 2025, he launched the Presidential Council on National AI Strategy to help buy massive amounts of computing power and a sovereign AI foundation model project that funds Korean companies to develop homegrown AI models. The government has also supported semiconductor titans, including Samsung and SK Hynix, through generous tax credits and low-interest financing. 

South Korea’s policy posture also prioritizes accelerating AI development over safety considerations. In 2024, South Korea’s legislature passed the AI Basic Act, one of the world’s first comprehensive AI laws, to promote AI development and establish light-touch regulatory guardrails. Seventy percent of South Koreans say advancing science and medicine through AI innovation is a bigger priority than protecting industries through regulation, according to the 2026 Stanford AI Index.

All of that effort might be paying off. The same index ranked South Korea as having the third largest number of notable AI models in the world, based on criteria such as state-of-the-art advancements or high citation rates. For many small countries like South Korea, AI is a chance to punch above their weight.

The blind spots

But that single-mindedness can crowd out critical reflection on AI’s broader societal impacts. “Because the national agenda on AI prioritizes economic development,” says Jeon, the professor of science and technology policy, “there isn’t much reflection on the social, political, ethical dimensions of the technology.” In 2025, the South Korean government faced a fierce backlash for rolling out AI textbooks riddled with factual inaccuracies and data privacy risks without testing them first in a pilot program to evaluate how they affect student learning.

And despite their optimism, South Koreans are still worried that AI could displace them from their jobs. After Hyundai announced in January that it will deploy Atlas humanoid robots across its car factories, the Hyundai Motor Group union protested vehemently. “Without labor-management agreement, not a single robot using new technology will be allowed to enter the workplace,” the union said. Sixty-four percent of South Koreans fear AI could displace human labor and exacerbate inequality, although 52% believe it could also increase productivity. 

On a recent Friday night in the Seoul Central Market, I went out with my cousins to a pocha, a late-night restaurant that serves fish cakes stacked in neat pyramids. As we clinked our cups of soju cut with beer—the scrappy staple cocktail of every Korean night out—one cousin asked me if I’d asked ChatGPT about my saju, a traditional Korean fortune-telling practice.

A 29-year-old insurance agent in Seoul praying for a new job and a boyfriend, she said asking ChatGPT about work and dating was her favorite pastime. She pulled up her phone and punched my birth date into the chatbot. 

Addicted to their screens, trapped between unemployment and dead-end jobs, and priced out of marriage and homeownership, 46% of South Koreans in their 20s have used a chatbot to read their fortunes, according to a survey by Korea Gallup. 

My cousin said she also asks ChatGPT for tips on trading stocks, dreaming big about making bank on her investment accounts into which she’s been pouring her salary. ChatGPT, she believes, is her portal out of reality into a better future.

Despite how fond she is of the chatbot as her shaman and financial advisor, she fears losing her job to AI. She still uses ChatGPT feverishly at work, as all her coworkers do, afraid of falling behind. 

“I sometimes fear AI, but for now, it’s just so useful,” she said.

Received — 4 June 2026 Artificial intelligence – MIT Technology Review

How courts are coping with a flood of AI-generated lawsuits

Most days in her chambers, Judge Maritza Braswell, a federal magistrate judge in Colorado, sifts through stacks of documents written by people without a lawyer. Many of them can’t afford to hire a lawyer, and others have cases too weak or too small to interest one. She reads each one carefully, mindful of how daunting it is to walk into the courtroom alone. 

Lately, like many judges across the US, she has seen a noticeable uptick in such filings. According to a new study that examined 4.5 million federal civil cases from 2005 to 2026, the share of lawsuits brought by self-represented people increased from 11% in 2022 to 16.8% in 2025. Within those cases, the number of filings made more than doubled from pre-2023 levels. 

Judge Braswell puts that jump down to AI. 

“I do correlate that to AI in part because I see AI use,” she says. As a tech-savvy judge who uses AI to vet court documents, she’s learned to recognize how large language models write. She can tell from the prose and at times, hallucinated cases and fabricated quotes. 

“I’m also actually seeing better-drafted pleadings,” she says. 

But while AI appears to be expanding access to justice, it doesn’t seem to be improving people’s chances of winning. Judges are also starting to question what kinds of rights and responsibilities large language models should bear as they step into lawyers’ shoes. For example, they ask whether a chatbot has a duty to provide good advice, as a human lawyer does. And a growing number of lawmakers across the US are starting to grapple with who should pay the price when chatbots dish out bad legal advice. 


AI supercharges lawsuits

To test whether AI was driving the increase in lawsuits filed by people without a lawyer, the authors of the study, Anand Shah at MIT and Joshua Levy at the University of Southern California, ran 1,600 randomly sampled court documents through Pangram, a commercial AI-text detector. The share flagged as containing AI-generated writing rose from 1% in 2023 to 18% in 2026. 

To Judge Braswell, that’s not necessarily a cause for concern. While the surge of AI-assisted filings might be adding to their workloads, she and many other judges find the cases easier to rule on because AI is helping people without legal training better articulate their arguments. 

Court documents written by people without lawyers are notoriously hard to decipher. Some arrive as handwritten scrawls bordering on gibberish that judges take a while to decode. However cryptic, judges are required to read them charitably.

These days, Judge Braswell has been churning through motions drafted by AI faster than the ones written by the litigants. “I have to be really careful because some of them contain hallucinations and errors, but I can generally understand what they’re arguing better with AI assistance from them than without it,” she says.

The clearer filings let Judge Braswell hear them better. “If I understand an argument a little bit better, I’m probably going to be able to help a little bit more,” she says.

Online communities are springing up to trade self-help guides on using AI to sue. In December 2024, a viral Reddit post walked immigration applicants through suing the United States Citizenship and Immigration Services over delayed review of their applications: draft a writ of mandamus with Microsoft Copilot, pay a lawyer $150 to polish it, and file in the expedient District of Vermont. Cases filed by people without lawyers in Vermont rose from about 45 a year before 2022 to more than 1,100 in 2024. 

Even so, people without lawyers are far more likely to lose their case than people with lawyers, and that’s not changing even with the addition of AI, the study found. 

“It turns out that mounting a lawsuit is a complex, multifaceted task. Not all of it is just drafting text,” says Levy. 

Chatbot-client privilege

Judge William Garfinkel, a federal magistrate judge in Connecticut, has served on the bench for three decades, pondering all sorts of questions about lawyers’ relationship with their clients. Lately, he has been wondering whether people’s conversations with chatbots dispensing legal advice should be privileged, the way their conversations with lawyers are. 

“You can make a good argument that … conversations with large language models like Claude or ChatGPT or Grok should deserve some protection,” he says.

Courts are starting to grapple with this question. In February, a federal court in Michigan ruled that a self-represented person’s conversations with ChatGPT to prepare her case were work product—legal work that is shielded from the opposing side.

The decision came on the same day a federal court in New York held that documents a criminal defendant had generated using Claude were not privileged attorney-client conversations or work product. The court argued that Claude is not an attorney and that a user has no “reasonable expectation of confidentiality in his communication” with it because AI companies can disclose user data to third parties. 

In March, Judge Braswell ruled that a self-represented person’s use of a chatbot should stay off limits. “It is true that AI systems like ChatGPT, Claude, Gemini, and others … collect user data for training and other purposes. But … that does not eliminate all expectations of privacy,” she wrote. Courts have since remained split on the issue.

Malpractice without a pulse

Some judges are also wondering whether a chatbot, like a lawyer, has a duty to provide good legal advice. Judge Allison Goddard, a federal magistrate judge in California, has noticed that people without lawyers often get the wrong advice from ChatGPT when trying to assess the value of their case during settlement negotiations. In one case, a plaintiff who slipped and fell in a store asked for $700,000 from the store, which was wildly more than the case was worth.

“Where are you getting the idea that you’re getting $700,000? Did you go to ChatGPT?” Judge Goddard asked. “Well …” the plaintiff mumbled. She then walked the person through the law to explain why ChatGPT was wrong and suggested a lower amount. “It’s like Dr. Google went to law school,” she says.

Then there’s the question of who’s liable when a chatbot makes such mistakes. In March, Nippon Life Insurance Company sued OpenAI alleging that ChatGPT practiced law without a license and helped a woman reopen a lawsuit that was already settled, flooding the court with frivolous filings. “ChatGPT is not an attorney,” the lawsuit said. 

In May, OpenAI asked the court to dismiss the case, arguing that ChatGPT does not practice law. “ChatGPT is not a person and neither has nor uses any degree of legal ​knowledge or skill,” OpenAI said in its filing. The case is still pending before the court.

States have started to weigh legislation that would hold AI companies liable when their chatbots offer bad legal advice. New York introduced a bill in March that would bar chatbots from impersonating lawyers, even if they notify ​users that they are interacting with chatbots. In Congress, a series of bills have been proposed to ban chatbots from posing as lawyers, doctors, and other licensed professionals. The bills have yet to gain traction.

For now, people will continue turning to AI to be their lawyer. For many of them, the rewards outweigh the risks. Not long ago, when Judge Braswell asked self-represented litigants why they wanted a particular piece of evidence, they mumbled timidly. Now, they answer her questions confidently, having rehearsed with a chatbot. 

“This is a really tough system to navigate. With AI, though, it gets a little less complex,” she says.

❌