Skip to content

Signals · AI Adoption

A two-year Khanmigo trial shows that access to AI is not the same as using it

In 18 Tennessee schools, students given an AI tutor gained about as much as students on plain practice. The authors point to how rarely they used it.

by ·3 min read·

New here? Start with the free AI Survival Kit →

A two-year Khanmigo trial shows that access to AI is not the same as using it

0:00 / 4:57

This voice is generated by AI.

Part of the guide Your employer is grading your AI use. What to do before the next review

A two-year Khanmigo trial shows that access to AI is not the same as using it

A working paper posted to EdWorkingPapers and dated August 2026 reports a randomised trial of AI tutoring in real classrooms. Philip Oreopoulos and Nina Low ran a two-year cluster randomised trial in 18 Tennessee middle schools. Students chosen at random used Khan Academy with its AI tutor, Khanmigo, during their existing daily remedial maths sessions. The tutor was set up to coach rather than hand over answers. Being assigned to it raised maths achievement by 1.3 national percentile ranks per term, about 0.06 to 0.08 standard deviations over a school year. The authors report that these gains resemble what Khan Academy practice delivers without AI assistance.

The obvious headline is that AI tutoring does not work. The paper points somewhere more specific.

The tool was there. The conversation rarely was.

Almost every student tried it. 96 percent of students used Khanmigo at least once. After that, usage fell away. The median student messaged the tutor on only a third of the days they practised. When the median student made a mistake, the moment a coach is worth most, they messaged the tutor in only 17 percent of those exercise sessions.

When they did write, the authors found the messages were mostly bare answers or clicks on suggested prompts. That is a student pushing a button and waiting for the machine to do something.

The paper also estimates that a full year of active participation would reach 0.14 standard deviations, roughly double the effect of being assigned. The authors put their reading plainly: "The binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access."

The same gap runs through the office

This is a study of middle-school maths. The pattern it measures is not limited to classrooms.

Employers already count usage. Amazon set a target of more than 80 percent of its developers using AI tools each week, the Financial Times reported in May. In this trial, the usage number was 96 percent. It looked like success, and the gains matched Khan Academy practice without AI. The figure that mattered was the 17 percent: how often the median student used the tool at the moment it could have changed their thinking.

A student typing a bare answer into a tutor and an analyst pasting a prompt and copying the output are making the same mistake. Both treat a reasoning partner as a dispenser. The machine does the routine work either way. The return comes from the person who uses it to think harder, questioning it, pushing back and asking why an answer was wrong.

If the pattern holds outside the classroom, access will not be what separates people. The habit of real dialogue with the tool will, and in this trial most students did not build it.

Your Next Move

Audit a week of your own use. Go through your last twenty AI interactions and split them into two piles. In one, you extracted an output. In the other, you questioned, challenged or asked the model to explain its reasoning. If the first pile dominates, you are getting the plain-practice result from a tool that can do more.

Take your mistakes to it. The trial's weakest number was engagement after an error. Do the opposite. When a draft comes back with corrections, a forecast misses or a reconciliation fails, ask the model to walk you through where your reasoning broke before you ask it to fix anything.

Keep an impact record. If your employer tracks AI adoption, expect it to count logins first. Keep a running record of decisions you made better or faster because you worked something through with the tool. Our briefing on what employers now measure in AI-use reviews and how to evidence impact sets out how to present it before your next review.

I build systems like the one publishing this site. → Work with me

About the author

Jo

Jo runs The War Room: one signal a day on how AI is changing work, and what to do about it.

Sources

Get the Briefing

Want more intelligence like this?

The weekly briefing, free — and the AI Survival Kit with it.

Takes 30 seconds. No spam.

More in Signals

See all