Mastery, Unmeasured

Initially published September 29, 2026

When a dashboard says a student has mastered a skill, what has anyone actually checked? Usually, whether the student’s next answer inside the same program would be right, because that’s what the model behind the word was trained and scored on. “Mastered” claims far more than that, and researchers who build the newest of these models have written that, as far as they know, none of the dozens of papers on them checks the gap.

I build this kind of software, so the criticism applies to my own company’s claims, and the question I think districts should ask vendors is one I have to answer first.

The Next Answer

Bayesian Knowledge TracingBayesian Knowledge Tracing is a model from the mid-1990s that keeps a running probability that a student has learned each skill, updating it after every right or wrong answer. Once that probability crosses a threshold, the software marks the skill mastered. and its neural successor, Deep Knowledge TracingDeep Knowledge Tracing replaces the Bayesian model with a neural network that reads a student’s whole history of answers and predicts whether the next one will be right. It keeps no explicit estimate for each skill. (DKT), are the standard ways software estimates what a student knows. Both are fit and scored on the next answer inside the system, and the paper that introduced the deep version defined its task in exactly those terms. That target is the right one for choosing the next exercise. On its own, it says nothing about what a learner can do later, unaided, somewhere else. A later analysis found the deep model tends to learn a general ability signal rather than tracking each skill through time, so even its next-answer accuracy says little about any one skill.

Outside the Tutor

Researchers at the University of Pennsylvania and Carnegie MellonRichard Scruggs and Ryan Baker of the University of Pennsylvania’s Graduate School of Education and Bruce McLaren of Carnegie Mellon’s Human-Computer Interaction Institute, writing at the 2020 International Conference on Computers in Education. wrote in 2020 that, to the best of their knowledge, “none of the dozens of papers of DKT and its successors have explicitly attempted to measure how well these approaches perform at inferring the knowledge that is carried outside the learning system”. They set that against the early Bayesian work, which did look, and they named what replaced it: “the now-dominant paradigm of predicting immediate correctness.”

The early Bayesian checks found over-prediction. In their summary, estimates for students driven to mastery ran ahead of their external post-testA post-test is an assessment given after instruction ends, separate from the practice items inside the software. It measures what the student carries out of the system rather than what they do inside it. scores, and the gap was larger for students who had needed more remedial practice.

Change the Target

The same paper shows the check is tractable. Its authors took each model’s probability of a correct answer on every item a student attempted, averaged it per skill, and compared that estimate with a 43-item post-test outside the tutor. On ordering decimals, the averaged deep model correlated with the post-test at 0.71, against 0.44 for the standard Bayesian mastery estimate. Most of that gain came from the averaging, not the architecture: the same step lifted the Bayesian model to 0.65. That’s one skill in one study, with no delayed retention test, and the gap was smaller or absent on the other three skills. They changed how the estimate was summarized because they were trying to predict the post-test, and no new model was needed to do it. A valid mastery signal has to predict performance with the help gone, and that can be tested.

Bloom’s Split

The same mismatch between claim and measurement runs through the number the personalized-learning category leans on. Benjamin BloomBenjamin Bloom was a University of Chicago educational psychologist, known for Bloom’s taxonomy of learning objectives and for developing mastery learning, in which students move on only after demonstrating a skill. reported in 1984 that one-to-one tutoring put the average student about two standard deviations above a conventional class. A 2024 peer-reviewed meta-analysis of randomized tutoring trials pools at 0.288.

A recent reanalysisPaul von Hippel, an associate professor at the University of Texas at Austin, reworked Bloom’s underlying studies in Education Next in 2024. of Bloom’s own studies accounts for much of that distance: the tutored students also got extra testing and feedback, which by its estimate explains about half of the two sigma. So about half the number invoked for individual attention came from repeated testing and correction, a quieter mechanism that software can reproduce, and not what “two sigma” is cited to mean.

A 1982 meta-analysis of tutoring studies had already shown the measurement problem. It found tutoring effects of 0.84 standard deviations on narrow or author-made testsAuthor-made tests are written by the researchers running the study and usually track closely what was taught. Broad standardized tests are written independently and cover more ground. and 0.27 on broad standardized ones, the same dependence on who wrote the test that a district should probe in a mastery score.

Stealth Assessment

Stealth assessmentStealth assessment, developed largely by Valerie Shute at Florida State University, builds measurement into a game or simulation and infers a student’s competence from how they play, without a separate test. leans on correlation more than prediction. In a 2023 review of the field, the most common validation correlated the game’s estimate with an external measure. Fewer studies tried to predict a post-test or scored classification accuracy, and some reported no validation at all. A correlation with another measure shows the estimate tracks something real, without showing whether the competence lasts.

What to Ask

A mastery estimate trained on the next answer is a good tool for choosing the next exercise, and a district can value it for that. “Mastered” is a larger claim: it describes the learner after the software has stopped choosing items and supplying help. The Bayesian over-prediction shows why that claim needs checking, and the post-test comparison shows it can be checked.

Ask the vendor to compare its mastery estimate with a delayed, unaided assessment it didn’t design, and to report the correlation. If the only validation is next-item prediction, the score should be described that way. If nobody measured it, the signal is a usage metric wearing a competence label.

That’s the line between personalized learning and precision learning. Personalized learning adapts to what a student does inside the system: which item comes next, when to offer a hint, when to move on. Precision learning makes a claim about the student outside it, and is held to the measurement that claim requires, the way precision medicine relies on a biomarker validated against outcomes rather than one that merely moves with treatment. A product can do the first well without ever being checked as the second, and the correlation a district asks for is what tells them apart.

References

Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge

Corbett and Anderson (1995), the original Bayesian Knowledge Tracing paper; its post-test over-prediction finding is reported here through Scruggs et al.

Student Modeling in the ACT Programming Tutor

Corbett and Bhatnagar (1997) on adjusting the knowledge-tracing model when mastery estimates over-predict post-test performance.

Deep Knowledge Tracing

Piech et al. (2015) define the task as predicting whether a student answers the next exercise correctly.

Why Deep Knowledge Tracing Has Less Depth than Anticipated

Ding and Larson (2019) find the deep model is more likely to learn an ability model than to track each skill.

Extending Deep Knowledge Tracing: Predicting Post-System Performance

Scruggs, Baker and McLaren (2020): averaged estimates correlate with a 43-item post-test at 0.71 (deep) and 0.65 (Bayesian) versus 0.44 for the final Bayesian estimate, on one skill.

The 2 Sigma Problem

Bloom (1984) reports one-to-one tutoring about two standard deviations above a conventional class.

Educational Outcomes of Tutoring: A Meta-analysis of Findings

Cohen, Kulik and Kulik (1982), 65 studies; the 0.33, 0.84 and 0.27 figures are reported through von Hippel (2024).

The Promise of Tutoring for PreK–12 Learning

Nickow, Oreopoulos and Quan (2024), peer-reviewed meta-analysis of randomized tutoring trials, pooled effect 0.288 standard deviations.

Two-Sigma Tutoring: Separating Science Fiction from Science Fact

Von Hippel (2024) attributes about 1.1 of Bloom’s two sigma to extra testing and feedback, leaving about 0.9 for tutoring.

Stealth Assessment: A Systematic Review of the Literature

Rahimi, Shute et al. (2023): of 60 coded studies, 24 validated by convergent correlation, 18 by post-test prediction or classification, 18 unspecified.

The Supervision Gap

The previous essay: juniors become seniors through supervised, corrected work, and AI puts that supervision at risk.

Learning on Credit

The companion argument: AI makes student work better and learning worse, and the gap comes due later.