Why students who use ChatGPT on homework assignments do worse on the test
Students who use ChatGPT to finish their homework get it done faster. But they also miss the point of the assignment: Learning the material. (Photo by Jessica Lewis 🦋 thepaintedsquare on Unsplash)
Aug. 31, 2026 — ChatGPT and similar AI products can quickly answer basic questions in math, science, and history, sometimes accurately and sometimes not. Parents and educators wonder, though: Are students who use AI actually learning or merely outsourcing their educational development?
Researchers who’ve studied the issue say their concern is justified.
One of the most cited works in this area is a 2024 study conducted by a team led by Hamsa Bastani, associate professor at the University of Pennsylvania's Wharton School.
Bastani’s study found that students who had access to ChatGPT while doing math homework problems did better on the homework completion but worse on a math test compared to students who didn’t have access to the chatbot.
In other words, the chatbot acted as a cheat code to give them the answers and complete the assignment faster. But in doing so they missed the point of the assignment: to actually learn the concepts and material.
Chatbot assist: zip through the homework, fail the test
In the study, students who used ChatGPT on their homework correctly solved 48% more of the practice problems than their peers who weren’t allowed to use AI. But the chatbot-aided students, when challenged to take a test without AI assistance on the same material, scored 17% wore than the students who didn’t use ChatGPT.
Bastani and her co-authors concluded that “unfettered access to ChatGPT-4 can harm educational outcomes.”
“Without guardrails, students attempt to use ChatGPT-4 as a ‘crutch’ during practice problem sessions, and subsequently perform worse on their own.”
Key question: How does gen ai affect how humans learn?
Bastani’s team wanted to understand “how generative AI affects how humans learn novel skills, both in educational settings and through the course of performing their jobs.”
“When technology automates a task, humans can miss out on valuable experience.”
— Hamsa Basmati, lead study author and associate professor at the University of Pennsylvania’s Wharton School
This has become an increasingly concerning question not just in education but in business as well. A number of thought leaders have raised worries about the reliance on AI hollowing out the rising generation of managers and company leaders. Junior staff members who perform a given set of tasks don’t just perform them—they learn from that performance, from failure and success and collaboration. If those junior staffers are replaced by AI, or by staffers who perform those tasks using AI, will they still gain the hard-learned experience they need to succeed at higher leadership levels?
“When technology automates a task,” the Wharton authors write, “humans can miss out on valuable experience, inducing a tradeoff where the technology improves performance on average but introduces new failure cases due to reduced human skill.”
Like the risk of pilots overusing autopilot
The study uses the example of airline pilots: “Overreliance on autopilot led the Federal Aviation Administration to recommend that pilots minimize their use of the technology; their precautionary guidance ensures that pilots have the necessary skills to maintain safety in situations where autopilot fails to function correctly.”
The risk of skill loss then doubles back on itself: Generative AI continues to suffer from hallucinations, instances where its output (answers) are presented as confidently truthful but are factually incorrect. A student, staffer, or leader who has gradually de-skilled, or failed to learn due to an overreliance on AI, will not have sufficient knowledge to spot the error and correct the hallucination.
What about using it in ‘study mode’?
AI developers aren’t unaware of these issues. About a year ago OpenAI released ChatGPT Study Mode, which the company said would "guide students towards using AI in ways that encourage true, deeper learning."
Chase DiBenedetto, an experienced tech reporter for Mashable, put Study Mode to the test shortly after its release. In her article, “I tried learning from AI tutors,” she found mixed results.
DiBenedetto reported that Study Mode’s brief interactions and minimalist user experience “make it easier to process what you are learning.” The product offered helpful practice tests, quick overviews, and seemed “built for learners seeking clarification on rubrics and grading standards.”
The downside: DiBenedetto found that ChatGPT Study Mode “would frequently give the answers, unprompted, and failed to let users fix mistakes before moving them on to the next step.” She found the experience frustrating, in part because “the chatbot is obsessed with getting users to practice and perfect what they just ‘learned.’”
The Mashable writer reported her experience to Hamsa Bastani, the lead author of the Wharton study. The researcher seemed unsurprised. DiBenedetto wrote:
“Across the board, Bastani says she has yet to come across an ‘actually good’ generative AI chatbot built for learning. Of the studies that have been done, most are negative or negligible as far as learning improvement.”
Follow-up research: can a better tutor be designed?
Two years after releasing the initial Wharton study, Bastani followed up with further intriguing work. Working with PdD candidate Angel Tsai-Hsuan Chung, Bastani explored small design changes might make AI tutoring more effective by emulating the most effective practices of human instructors.
In an interview about the 2026 work, Bastani said:
“Our previous research showed that generative AI can harm learning by encouraging over-reliance. Students end up asking for help too early and undermining their own ‘productive struggle.’ So one of the questions we began exploring was whether we could preserve that productive struggle through personalization.
In other words, AI could offer more than just 24/7 access to a teaching assistant; it could also enable personalized learning. In a traditional classroom, you have to teach to the lower middle, something like the 33rd percentile student. You want to make sure the strongest students are not bored, but you also need to revisit material for those who are struggling.
The idea was to personalize practice problems in a way that both creates more productive struggle and better matches each student’s current level of mastery.”
Their initial results were positive. The students who used the more personalized question sequence actually performed better on the final exam. “What’s important is that the effects are meaningful, and they’re essentially free,” Bastani said. “We didn’t increase instruction time or teacher workload.”