Skip to content
SATURDAY, OCTOBER 10, 2026

Independently reported.

Culture

Ten Minutes of AI Help Hurt Solve Rates in All Three Trials. Giving Up Showed Up in Two.

A UC Berkeley News story says ten minutes with a chatbot erodes persistence. The paper behind it found lower solve rates every time, and a skip-rate gap that was not statistically significant in its best-controlled experiment.

By Sana Whitfield, Culture, Science & Life

· 5 min read · Updated

A yellow pencil resting on blank graph paper beside a closed laptop on a worn wooden desk in soft morning light, no people.
Illustration: Trestlewire

Key Takeaways

  • •Three randomized trials (1,222 recruited, 1,060 analyzed) found lower unassisted solve rates after GPT-5 help was removed: 57 percent versus 73 percent, 71 versus 77, and 76 versus 89.
  • •Skipping rose significantly in the first fraction trial (20 percent versus 11 percent) and the SAT reading trial (8 percent versus 1 percent), but not in the better-controlled fraction trial (10 percent versus 7 percent, p = 0.239).
  • •In Experiment 2, 61 percent of the AI group (189 of 308) said they mostly asked for answers, and that group skipped 13 percent of test problems versus 5 percent for hint users.
  • •Participants were paid online recruits on low-stakes tasks tested right after the AI was removed, and the authors say lasting effects have not been measured.

A fraction problem fills the screen. The easiest ones in this study look like 5/6 minus 1/3, and the hardest chain three steps together. Beside the problem sits a chat sidebar where GPT-5 has already been handed the solution, so a participant who wants the answer can type a single word: "answer?" A tutor would wince at that prompt. The sidebar obliges. Twelve problems into the session, it vanishes without warning, and three more fractions are waiting.

That setup sits at the center of a paper presented this week at the Conference on Language Modeling, and of a UC Berkeley News story published October 9 under the headline "Using AI for just 10 minutes erodes your ability to persist at hard things." The first author is Grace Liu of Carnegie Mellon, with Brian Christian of UC Berkeley and colleagues at Oxford, MIT and UCLA. The paper is more hedged than the headline.

The short answer

Three randomized trials with 1,222 recruited adults tested what happens when a GPT-5 helper is taken away after about ten minutes. Solve rates on unassisted problems fell in all three trials. Skipping rose significantly in two. In the best-controlled fraction trial, the AI group skipped 10 percent of test problems versus 7 percent for controls, a gap the authors report as not statistically significant.

Three trials, one clear pattern and one shakier one

Experiment 1 recruited 354 people on the Prolific research platform. After exclusions, 185 AI-group participants and 122 controls remained. On the three unassisted test problems, the AI group solved 57 percent and the control group 73 percent. The AI group also skipped 20 percent of the test problems against 11 percent for controls, the widest raw skipping gap in the paper.

57% vs 73%

Share of unassisted test problems solved, AI group vs control, Experiment 1

307 analyzed participants, three test problems each. The authors flag a possible attrition confound in this trial.

The authors flag a problem with that result themselves. They dropped anyone who solved fewer than 3 of the 12 practice problems, but AI users could clear that bar by copying answers, and more controls were dropped. The AI group may have kept weaker solvers, which could inflate the gap.

Experiment 2 fixed this. It added a three-problem pretest for exclusions and gave controls a matched sidebar, so both groups saw an interface change when the test began. Of 667 people recruited, 585 were analyzed. The AI group solved 71 percent of test problems against 77 percent for controls, a small effect (Cohen's d of 0.19). They skipped 10 percent against 7 percent. The authors' own test put that skip difference at p = 0.239 and describe it as not significant.

Experiment 3 swapped fractions for SAT reading questions. Of 201 people recruited, 168 were analyzed. The AI group solved 76 percent of test questions versus 89 percent, and skipped 8 percent versus 1 percent (p = 0.008).

Who actually gave up

In Experiment 2, the researchers also asked the AI group how they had used the assistant. Of 308 participants, 189 (61 percent) said they mainly asked for answers. Another 82 (27 percent) used it for hints or clarification, and 37 (12 percent) barely touched it.

The answer-seekers solved 65 percent of test problems and skipped 13 percent. Hint users solved 76 percent and skipped 5 percent. Controls solved 77 percent and skipped 7 percent. So the extra giving up sits in one subgroup, which is also where the authors urge the most caution. People chose their own style of AI use, so this analysis is cross-sectional and not causal by the authors' own account.

What ten minutes can and cannot tell you

The stakes were low. Pay was flat ($2.60 in the first trial, $3.40 in the others), and in the first trial people were told that wrong answers carried no penalty and that payment did not depend on how many they got right. A skip there is closer to a shrug than a failure. The authors tested people immediately after the AI left and say they cannot speak to whether the effect lasts hours or days. Their assistant was also maximally helpful by design, handing over complete answers on request. Whether months of daily use would compound the effect, fade it, or do neither is, in their words, an open question.

Christian, a research fellow at UC Berkeley's Center for Human-Compatible AI and author of the 2020 book The Alignment Problem, keeps a notebook on his desk and writes in it every day, according to Berkeley News, partly to ward off a suspicion that too much time with AI could dull his own scholarly edge. Of the three experiments, he said, "You almost couldn't ask for a clearer story."

That reads strongest for solve rates, which fell in all three trials. For persistence, the cleanest fraction trial gave a smaller signal that did not clear the usual statistical bar.

The obvious objection is the calculator, which also does arithmetic for us. The authors raise it themselves and argue that fractions and reading comprehension are developmental prerequisites, the footing for algebra and critical reasoning. Their proposed mechanism is a hypothesis, not a tested result: when answers arrive in seconds, the expected length of a task shrinks, and unaided work starts to feel slower than it is.

The finding is not alone. The paper's literature review says earlier evidence on AI deskilling was mostly correlational or based on small samples, and cites a 2025 randomized trial of GPT-4 in high school math (Bastani and colleagues, in PNAS) that also found worse unaided performance afterward, partly offset when students used a scaffolded tutor. The direction matches. The size, in this paper's cleanest test, is modest: six points on solve rate.

The narrower version of the headline is still real. About ten minutes of answer-on-demand help left people solving fewer problems in three out of three trials. Giving up showed up in two, and in the one trial that tracked how people used the AI, it clustered among those who asked for answers rather than hints. What nobody has measured yet is the distance between a lab session and a working life.

Frequently Asked Questions

Does ten minutes of AI help erode persistence?
In the Berkeley-linked paper, solve rates on unassisted problems were lower after about ten minutes of GPT-5 help in all three trials. Skipping, the study's measure of giving up, rose significantly in two of the three and not in the best-controlled fraction trial.
Does the study show lasting damage?
No. People were tested immediately after the AI was removed, and the authors say they cannot speak to whether the effect lasts hours or days, or whether longer daily use compounds it.
Who took part?
US-based adults recruited on the Prolific platform and paid a flat fee. 1,222 were recruited across three experiments, and 1,060 were included in the analysis after exclusions.
  • AI and learning
  • persistence
  • Brian Christian
  • UC Berkeley
  • Conference on Language Modeling
  • GPT-5
  • cognitive offloading

Sources

  1. 01AI Assistance Reduces Persistence and Hurts Independent Performance (arXiv:2604.04721, COLM 2026), arXivarxiv.org
  2. 02Using AI for just 10 minutes erodes your ability to persist at hard things, UC Berkeley Newsnews.berkeley.edu
  3. 03AI Assistance Reduces Persistence project page, Study authorsai-project-website.github.io

Corrections

No corrections have been made to this article.

About the reporter

Sana Whitfield

Culture, Science & Life Reporter, Trestlewire

I started out writing about science — real lab-coat, peer-reviewed, methods-section science — and I moved into culture coverage because I kept noticing that the internet, not the lab, is often where science actually plays out in people's daily lives now. A study on sleep or attention or habit formation used to sit in a journal for a decade before anyone outside the field read it. Now it shows up in a group chat by Thursday, usually stripped of every caveat that made it interesting in the first place.

Read full bio and all stories →