How do you measure whether an AI can actually teach?

Jul 16, 2026

Charts and data analysis on paper

It is easy to make an AI that sounds like a teacher. It is much harder to know whether it actually teaches. Fluent explanations are cheap; what matters is whether a student who did not understand something understands it afterwards. So before we let any version of Chalk-1 near a student, we make it prove itself.

Accuracy is not teaching

Getting the maths right is table stakes — and it is checked separately, live, with every step verified before it reaches the board and fewer than 2% ever corrected. But a model can be perfectly accurate and a terrible teacher: lecturing instead of asking, racing ahead of the student, doing all the work itself. Teaching quality has to be measured as its own thing.

How we score it

Every candidate version of Chalk-1 teaches complete lessons — not cherry-picked exchanges — across five GCSE maths modules, against simulated students who behave like real ones: confidently wrong, quick, quiet, distracted. Those lessons are scored blind on a 5-point rubric that asks teacherly questions. Did it notice the mistake? Did it let the student do the work? Did it check understanding rather than assume it? Blind means the judge does not know which model taught the lesson — so newer versions get no benefit of the doubt.

The bar is the frontier

We score a frontier model on the same rubric — currently Opus 5, which scores 4.35 — and treat it as the bar. Chalk-1 has climbed from 3.15, to 3.60, to 4.30 in its current version: within 0.05 of the frontier benchmark, at a small fraction of the running cost. The full chart is on our technology page, and we will keep publishing it with every release — whichever direction it moves.

Why publish this at all?

Because every AI education product claims to be good at teaching, and almost none of them show their working. We think the bar for teaching children should be the same one we set for our students: show the method, not just the answer. Next up is the harder benchmark — Chalk-2, measured head-to-head against human tutors on learning outcomes.