The most interesting thing an AI tutor did in a recent experiment with nearly 800 students had nothing to do with how it explained anything. It simply chose which problem to serve up next. That small design decision produced a gap in final exam results large enough to surprise the researchers who built it, and it cuts against most of what the edtech industry has been selling for the past three years.
When chatbots backfire
The uncomfortable backdrop, laid out in new reporting from The Hechinger Report that has been circulating widely this week, is that the evidence on AI tutoring so far leans negative. Several studies have found that chatbot tutors can actually hurt learning because students lean on them too heavily, get spoonfed solutions, and never absorb the material. Even tutors deliberately designed to withhold answers have not consistently beaten old-fashioned studying.
That track record has not stopped schools from adopting the tools, or vendors from marketing them. It has, however, pushed some researchers to ask a sharper question: if conversational explanations are not the ingredient that matters, what is?
The Penn experiment
A team at the University of Pennsylvania, which included some confirmed AI sceptics, tested one answer with close to 800 Taiwanese high school students learning Python in a five-month after-school course. Every student used the same AI tutor, built not to give away answers. The only difference was sequencing. Half the students worked through a fixed ladder of practice problems, easy to hard. The other half got a personalised sequence, with the system continuously adjusting difficulty based on how each student was performing.
The personalised group did better on the final exam, a difference the researchers characterised as the equivalent of six to nine months of additional schooling. That is an eye-catching claim for a five-month online course, and the tutor’s inventor, Wharton doctoral student Angel Chung, acknowledged the conversion was “not a perfect estimate.” The draft paper, posted in March, has also not yet been through peer review. Even read conservatively, though, the direction of the result is hard to ignore.
The design rests on an old idea from learning science called the zone of proximal development. Problems that are too easy bore students. Problems that are too hard defeat them. The system pairs a large language model with a separate machine-learning algorithm that watches how students answer questions, how often they revise their code, and how they talk to the chatbot, then uses all of it to pick the next problem that keeps them in the productive middle.
“Students usually don’t know what they don’t know,” Chung told Hechinger. “The student doesn’t have the ability to ask the right questions to get the best tutoring.”
Engagement was always the problem
None of this is entirely new. Long before ChatGPT, researchers built intelligent tutoring systems that estimated what a student knew and delivered the right next problem, and rigorous studies showed well-designed versions worked. Their fatal weakness was that students found them dull and would not use them. A conversational AI may finally solve the engagement half of the equation: students in the personalised group spent roughly three extra minutes per problem, adding up to about an hour of practice per module, double what the comparison group put in.
The benefits were not evenly spread. Students new to Python gained the most, while experienced coders did just as well with the fixed sequence. Students from less elite high schools also appeared to benefit more, a hopeful sign for a technology often accused of widening gaps rather than closing them.
What it does not prove yet
The study’s participants were volunteers taking an optional course to strengthen college applications, many with educated parents and prior coding experience. Whether the same approach works for a disengaged 14-year-old who is behind at school, the student who most needs the help, remains unanswered. Carnegie Mellon’s Ken Koedinger, a pioneer of the original tutoring systems, is testing one hybrid: AI models that alert remote human tutors when a struggling student starts drifting. “We are having more success,” he told Hechinger.
The lesson for schools shopping for AI this autumn is worth sitting with. The feature that moved the needle was not a friendlier chatbot or a smarter explanation. It was an unglamorous recommendation engine deciding what a student should practise next. The tutoring wars may end up being won not by whoever talks best, but by whoever asks better questions.
For more coverage of AI in education, visit Mylistingo.







