The Promise and Perils of AI Tutors: New Research Questions the Effectiveness of Khanmigo in Classroom Settings

When OpenAI launched ChatGPT in late 2022, the educational landscape underwent a seismic shift. Almost overnight, students discovered a tool capable of generating essays, solving complex equations, and summarizing lengthy texts with unprecedented speed. While proponents saw an opportunity for personalized learning, educators quickly identified a darker trend: students were leveraging the technology as a shortcut to bypass critical thinking, leading to widespread concerns regarding academic integrity and the potential atrophy of foundational math and literacy skills.
In response to these challenges, Silicon Valley innovators and educational organizations sought to pivot the technology from a shortcut into a scaffold. The goal was to build an AI that functioned not as a ghostwriter, but as a Socratic tutor—an entity that would withhold answers, provide subtle hints, and guide students through the struggle of learning. Among the most prominent responses to this call was Khanmigo, an AI-powered assistant developed by the nonprofit Khan Academy and released in 2023. However, a comprehensive new study covering two years of classroom implementation suggests that the gap between a well-designed pedagogical tool and actual student engagement remains vast.
The Tennessee Study: A Longitudinal Analysis
Researchers from the University of Toronto, including Philip Oreopoulos and Nina Low, conducted an extensive randomized controlled trial (RCT) in 18 Tennessee middle schools between 2024 and 2026. The study focused on students who were identified as needing academic support, specifically those performing at least one grade level behind their peers in mathematics. These students were already enrolled in mandatory remedial math periods, providing a controlled environment to measure the impact of AI-assisted instruction against traditional digital interventions.
The participants were divided into two groups. The first group continued with their standard remedial curriculum, which utilized established online platforms such as Zearn, IXL, Waggle, and DeltaMath. The second group was introduced to Khan Academy’s ecosystem, augmented by the Khanmigo AI tutor.
The findings, circulated in a draft paper by the National Bureau of Economic Research (NBER) in August 2026, revealed a sobering reality: while the platform itself showed some efficacy, the AI tutor component struggled to find its footing. Most students initially experimented with Khanmigo, but their primary objective was often to use the tool as an answer key. When the software’s programming refused to provide direct solutions—opting instead to offer hints—student interest plummeted.
"We observe how students actually used the AI tutor, and the answer is: not much," noted Oreopoulos and Low. The data indicated that seeking help while navigating personal academic confusion remained a choice, and the majority of students opted to bypass the tool rather than engage with the productive struggle of guided inquiry.
Data Trends and the "Value Add" Question
The quantitative results of the study painted a complex picture. Students who utilized the Khan Academy intervention did show performance improvements compared to those in standard remedial settings. However, the study found that these gains were statistically indistinguishable from the improvements recorded in previous years where Khan Academy was used without the AI assistant.

In essence, the digital practice platform provided a measurable benefit, but the inclusion of the "Socratic" AI tutor failed to move the needle further. This suggests that for many middle school students, the barrier to learning is not necessarily the absence of a guide, but the lack of intrinsic motivation to engage with complex, effortful problem-solving when an easier path is desired.
Chronology of AI in the Classroom: 2022–2026
To understand the current state of AI in education, one must look at the rapid evolution of the technology over the last four years:
- November 2022: The public release of ChatGPT triggers a global debate on AI’s role in education, characterized by a rapid rise in student misuse for homework completion.
- 2023: Khan Academy releases Khanmigo, aiming to provide a safe, AI-driven, Socratic tutoring experience that prevents cheating while supporting learning.
- 2024: Researchers launch a two-year longitudinal study in Tennessee, tracking the integration of AI tools within remedial middle school math programs.
- 2025: Initial feedback from schools highlights a "disengagement gap," where students find the Socratic method frustrating when they are struggling to meet immediate assignment deadlines.
- 2026: The NBER study is released, confirming that while Khan Academy’s software is effective, the AI tutor component did not significantly improve outcomes compared to baseline digital learning tools.
Institutional Responses and Adaptations
Sal Khan, the founder and CEO of Khan Academy, has been transparent about the study’s findings. Rather than dismissing the research, Khan embraced the data, utilizing his organization’s blog to analyze the trial’s implications with the same methodical approach he applies to his instructional videos.
In interviews, Khan acknowledged that the internal data collected by his team mirrored the findings from Tennessee. He emphasized that the decision to release the AI tutor in 2023 was a calculated risk, prioritized by the need to ensure data privacy, student safety, and the mitigation of "hallucinations"—the instances where AI provides factually incorrect or nonsensical information. "It’s allowed us to learn and hopefully make the new version even more helpful," Khan noted, defending the decision to deploy the technology in a real-world setting early in its development cycle.
Since the Tennessee study, Khan Academy has overhauled the user experience. The AI is no longer a peripheral tab that requires a conscious, active choice by the student. Instead, it is now integrated into the core workflow. The logic has been shifted: if a student answers a problem incorrectly, the AI proactively intervenes to offer support, rather than waiting for the student to seek it out. Furthermore, the company is experimenting with "credit-based" incentives, where a student’s engagement with the AI during a remedial exercise can count toward the mastery requirements needed to progress through the curriculum.
Implications for the Future of EdTech
The findings from the Tennessee trial raise broader questions about the future of Artificial Intelligence in K-12 education. The primary takeaway is that the existence of sophisticated technology does not automatically equate to pedagogical success. If the software requires a high degree of student agency and self-regulation—qualities often missing in students who are already struggling—the technology may fail to reach its intended audience.
The challenge now shifts from "Can we build a tutor?" to "How do we design a system that students will actually choose to use?" Educators and developers are looking toward a model where AI is not an optional add-on but a seamless, low-friction partner in the classroom. However, the requirement for evidence-based success remains high. As policymakers and school districts look to invest limited budgets in AI tools, they will require more than just promises of "personalized learning"; they will need rigorous, long-term data proving that these tools can actually improve student achievement in ways that traditional, non-AI digital tools cannot.
As the educational community moves into the 2026–2027 school year, the focus will be on these "second-generation" integrations. Whether shifting the AI from a passive responder to an active participant will increase student engagement remains the critical question. For now, the Tennessee study serves as a necessary reality check: in the race to automate education, the human element—the student’s willingness to engage with the difficult, often tedious process of learning—remains the most important variable in the equation.







