
The Supervision Gap
Initially published September 23, 2026
The entry-level job is disappearing, and the explanation everyone reaches for is that AI does that work now. Junior associates used to spend their first years on document review, junior engineers on boilerplate, junior analysts on the first pass of a model. Those are exactly the tasks a model now finishes in seconds, so that is changing the equation for some companies.
The best data on that claim supports half of it. Payroll records analyzed by Brynjolfsson, Chandar and Chen at Stanford show employment of 22-to-25-year-olds falling in the occupations most exposed to AI while the same age group grew everywhere else, and experienced workers in those exposed jobs show no comparable gap. The decline runs through fewer hires rather than more layoffs: nobody is being pushed off the rung, fewer people are being let onto it. It also concentrates where AI substitutes for the work; where AI assists the person doing it, employment is flat or rising. The authors present all of this as descriptive and stop short of claiming AI caused it.

The task was never what turned a junior into a senior. What did was having that task handed back with corrections by someone whose judgment was already formed and who would answer for the result. That loop is how people build taste, the sense of what good looks like and which output is quietly wrong, and it’s how they build opinion, the willingness to defend a call when someone senior disagrees. AI can take the task. The risk is that the supervision goes with it.
Education is the field arguing loudest about what AI does to learners, and it also trains its own staff through supervised practice written into licensure. Special education makes a good test case: the supervision requirement there comes with a number attached, so the constraint can be measured, while the same constraint runs through law and software without anyone counting it.
If AI took the work, you protect the work. If supervision is the real constraint, protecting beginner tasks accomplishes nothing, and the useful move is to point AI at the supervisor’s overhead until there’s room to take on a junior again.
Taste Is Taught by Correction
Every field that produces senior judgment runs some version of the same loop. A junior does real work, a senior marks it up, and the junior absorbs the fix along with the reason for it. Repeat that a few hundred times and the junior starts catching the problems before the senior does.
In law, the first-year associate’s memo comes back covered in a partner’s edits. The document review that preceded it was never the education; it was what got you access to the markup. In software the equivalent is code review. A junior engineer learns more from a senior’s comment on why an abstraction is premature than from writing the code, and that comment only exists because someone senior took the time to read it.
Medicine formalized the loop. A resident sees real patients and makes real calls, and an attending physician stays responsible for every one of them while the resident’s judgment forms.
Kitchens run it without paperwork: a line cook’s plate gets tasted before it leaves the pass, and a cook corrected on seasoning often enough eventually stops needing to be.
What transfers in all of these is calibration more than skill. The junior learns which errors matter and which don’t, where the senior draws the line, and eventually where they would draw it differently. That last part is opinion, and it only forms once someone has disagreed with you and you’ve had to hold your ground or concede.
The same payroll data hints at why. Young workers lost ground in occupations built on codified knowledge, the kind that can be written down and handed to a model, while occupations built on tacit knowledge saw experienced workers’ employment grow faster. The authors flag that finding as suggestive. Tacit knowledge is the part a junior can only get from someone who already has it.
In special education, the loop is a regulatory requirement with a number attached.
What Actually Caps the Pipeline
A speech-language pathology student can’t bank clinical hours unless a licensed clinician watches them work. ASHA’s standard says the supervision must be in real time and must never be less than 25 percent of the student’s total contact with each client. That single rule turns the size of every graduating class into a function of how many practicing clinicians are willing to host a student, and nobody budgets for willingness. A supervisor takes on a trainee on top of a caseload that is already past manageable.
Asked what limits their enrollment, speech-language pathology master’s programs put insufficient student funding first and insufficient clinical placements second. Money is the larger of the two, and I’d rather concede that than pretend otherwise. Neither one is a shortage of people who want in.
School psychology has the same shape. Programs are held small by how many faculty they have and how many internship sites they can get approved, and the one study that tested whether people leaving the field explains the shortage found it doesn’t, though it covered a single state.
The result is a special-education shortage reported in forty-five states, with school psychologists carrying roughly twice the student load their professional body recommends. Every one of those missing clinicians is also a person who isn’t hosting an intern: a district that can’t hire a school psychologist can’t offer the internship site that would have produced one.
Where the Hours Go
In special education, working directly with students is a smaller share of the job than the title suggests. Speech-language pathologists spend under 60 percent of the week on intervention, school psychologists spend the bulk of the year on eligibility evaluations, and special-education teachers work ten to twelve uncompensated hours a week beyond contracted time, a large share of it documentation.
The closest thing to a real time-use study of special-education teachers followed thirty-one of them in one region, sixteen years ago. We are five decades into a federal mandate that governs how these professionals spend their days, and nobody has measured what it costs them.
None of it is waste in the ordinary sense. IDEA mandates the documentation. The hours exist because federal law says a child’s eligibility and services have to be recorded in a specific, defensible, auditable way, and a district that cuts corners there loses in due process. You cannot lean on people to be faster at a legal requirement.
The automatable part of a special-education job is the part nobody chose and nobody exercises judgment inside. Drafting, formatting, assembling, cross-referencing, transcribing: none of it is the clinical decision, and all of it is what’s eating the hours that used to hold a student teacher.
Same Tool, Opposite Effect
Almost every argument about AI in education is an argument about students. Whether it helps them learn, whether it lets them submit work they didn’t do, whether a district should ban it or require it. That argument matters and the evidence in it is uncomfortable: when a model writes the essay, the model does the work that was supposed to build the student, and the paper improves while the person doesn’t.
A student’s essay isn’t valuable because it exists. It’s valuable because producing it forced retrieval, sequencing and self-correction, and those are the things meant to stay behind once the essay is handed in. Automate the writing and you keep the artifact while deleting the lesson.
An eligibility report is valuable because it exists. It’s a legal instrument recording a decision a licensed person already made, and nobody’s judgment is formed by formatting it, reconciling it against last year’s version, or retyping a score table. The clinician learned to make that decision through supervised practice, which is the activity the paperwork displaced.
A first-year associate’s draft brief is their student essay: the value is in writing it and having it taken apart. Reconciling redlines, renumbering exhibits and conforming citations are the eligibility report, necessary work inside which nobody’s judgment forms.
The test for any AI use, in a school or anywhere juniors are trained, is whether the work being displaced was the work that made someone better. For a student building an argument, or an associate drafting a first brief, it almost always was. For a clinician assembling a compliance document, it almost never is.
Age changes what a student’s effort is building. In K-12 the effort is building cognition itself: attention, working memory, the ability to hold an argument in your head long enough to write it down. Offload that and there may be nothing underneath to build on later. In higher education the student usually has that foundation, and what’s being built is expertise, knowledge of a domain and a judgment about what good looks like inside it. That’s the same thing a junior builds at work, and it forms the same way, through work that gets corrected by someone who already knows. A tool that is reckless near a ten-year-old can be reasonable near a graduate student, provided someone is still checking the output.
Students and staff also sit on top of the same shortage. A student who needs a speech evaluation and a graduate student who needs a supervisor are waiting on the same missing person.
What Teachers Actually Reach For
Almost everything published about teachers and AI is survey data, which measures what people say they do. Stanford’s SCALE initiative went at it through platform logs instead, and its largest look covers roughly 87,000 educators in the top five percent of MagicSchool users, meaning the ones who showed up on at least fifty-two separate days over two and a half years.
Two measures in that data pull against each other. On adoption, teachers consolidate: the general-purpose assistant is the single most-used tool by a wide margin, and SCALE’s companion work on a different platform found 84 percent of teachers running more than one thread stayed with one assistant, with no specialized tool clearing 12 percent. On intensity, the specialized tools win: at every grade level, use-case categories other than the chatbot category drew more concentrated use. Teachers reach for the general tool and work the purpose-built ones harder.
Elementary teachers point AI at administration, communication and student support, and they are the only group whose use of a text-rewriting tool exceeds their use of the general assistant. Middle and high school teachers use it on the academic core: feedback on student work, quizzes and worksheets, instructional materials, lesson plans. One group of teachers, without being told to, already aims AI at the overhead rather than at the teaching.
Elementary teachers are also underrepresented among heavy users relative to their share of the workforce, which is roughly half of all US teachers against a third of this cohort. This is one vendor’s user base, and “top five percent” counts days visited and says nothing about depth of use, which is a narrower thing than it sounds. Nobody has done the equivalent study on a supervising clinician’s documentation load.
What the Evidence Doesn’t Support Yet
The case for pointing AI at staff workload is promising and not yet settled. Gallup, in work funded by the Walton Family Foundation, found that teachers who use AI weekly report saving 5.9 hours a week. Gallup’s own methodology note says teachers produced that figure by estimating in half-hours, task by task. It’s a self-report, from people who had already chosen the tool.
On the clinical side the evidence is thinner. The largest peer-reviewed study of AI-assisted IEP goals compared goals written with and without ChatGPT across fifty-six special educators, and found no statistically significant difference in quality. The finding is that AI-assisted goals were not worse, which is a smaller claim than the one usually made from it, and the handful of other studies on the same question have been smaller still.
The bar was set this month by the American Psychological Association’s expert report on children’s and adolescents’ learning with educational technology. It tells policymakers in as many words not to accept usage reports or engagement metrics as proof of efficacy, and to ask vendors for independent data showing knowledge retention or skill transfer that persists at least a week after use. That is aimed squarely at the companies selling into schools.
The Rung We Already Built
The apprenticeship everyone is now trying to invent already exists in the licensed professions, and each writes supervision into the path to practice. Physicians train under attending supervision through residencies of three years or more. Architects, psychologists, clinical social workers and speech-language pathologists all log supervised hours before they practice alone, and every jurisdiction in the country requires it of clinical social workers.

None of them hands the junior easier work. A resident sees real patients and a first officer flies the real aircraft, while someone licensed carries the accountability as the junior’s judgment forms.
Industries without licensure ran the same loop informally, through the partner’s markup and the senior engineer’s review, and those versions are the most exposed now because nothing written down protects them. Hand the junior the AI and the simplified work at the same time and you remove the difficulty and the correction together. They produce acceptable output from the first week without ever learning what was wrong with it or having to defend it.
The Gap Was Never the Work
AI took the entry-level work, and that’s why there’s nowhere for a beginner to start.
In any field that turns juniors into seniors, the rung is a person with judgment, some capacity, and a reason to spend it on someone who can’t yet return the favor. That capacity is under pressure everywhere at once: in special education it’s consumed by documentation federal law requires, and in the occupations AI handles best it’s shrinking because fewer juniors are being hired to correct. Protect the beginner tasks from automation and you have defended the wrong thing.
There’s a version of this where AI closes the gap instead of getting blamed for it. That version leaves the senior’s judgment where it is and gives the senior back time, with no model drafting something the senior then rubber-stamps. It only counts if some of that time goes to correcting someone who is still learning. In a school, that means giving a school psychologist back a few hours out of each of the fifty-five evaluations they run a year, and spending part of what comes back on an intern.
References
Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence
Brynjolfsson, Chandar and Chen, August 2026. ADP payroll data through June 2026 on young workers in AI-exposed occupations.
ASHA Certification Standards: SLP Clinical Practicum
Standard V-E: supervision must be in real time and never less than 25 percent of the student’s total contact with each client.
CSD Education Survey National Aggregate Data Report, 2023-2024
ASHA and CAPCSD. 52,179 applications against 23,655 offers, and the enrollment-limiting factors programs report for themselves.
An Analysis of the Workforce Pipeline in School Psychology
Sabnis et al. Single-state study of 236 graduates from eight Ohio programs, concluding attrition is not a significant contributor to the shortage.
State Teacher Shortages: 2026 Factsheet
Learning Policy Institute. Special education reported as a shortage area in 45 states, more than any other subject.
NASP Shortages Dashboard and Workforce Information
A national ratio of 1,071 students to one school psychologist for 2024-25, against a recommended 1:500.
ASHA 2022 Schools Survey: Caseload Characteristics Trends
Twenty-two hours a week on direct intervention, six on documentation and four on diagnostic evaluations.
NASP 2020 Membership Survey
School psychologists complete an average of 55 initial evaluations and reevaluations a year.
Special Education Teacher Time Use in Four Types of Programs
Vannest and Hagan-Burke, 2010. Thirty-one teachers across 24 schools in the Southwest, and still the closest thing to a time-use study.
Teachers Are Using AI to Help Write IEPs. Advocates Have Concerns
Education Week, October 2025. Elizabeth Bettini on the hours special educators work beyond the school day.
How Highly Active K-12 Educators Are Using AI Tools Like MagicSchool
Stanford SCALE Initiative. Platform behaviour from roughly 87,000 educators in the top 5 percent of MagicSchool users, May 2023 to November 2025.
What K-12 Educators Are Actually Prompting to AI
Stanford SCALE Initiative. More than 150,000 prompts on a different platform, where teachers consolidated onto a single general assistant.
Three in 10 Teachers Use AI Weekly, Saving Six Weeks a Year
Gallup with the Walton Family Foundation, fielded spring 2025. The 5.9 hours is self-reported by weekly users.
IEPs in the Age of AI: Examining IEP Goals Written with and Without ChatGPT
Waterfield et al. Fifty-six special educators across 22 states, finding no statistically significant difference in goal quality.
APA Expert Report on Children’s and Adolescents’ Learning With Educational Technology
American Psychological Association, September 2026. Usage and engagement metrics are not evidence of efficacy.
ASHA Clinical Fellowship
The 1,260-hour supervised clinical fellowship required for certification as a speech-language pathologist.
NCARB Architectural Experience Program: Experience Requirements
3,740 experience hours, of which 1,860 must be under the supervision of a licensed architect.
ASPPB: Supervised Experience Requirements for Psychologist Licensure
Jurisdiction-by-jurisdiction supervised hours, most US states between 1,500 and 3,600.
ASWB: Clinical Social Work Supervision, Comparison of Requirements
Post-MSW supervised experience required in all 56 jurisdictions; 3,000 hours is the most common requirement.
Learning on Credit
The companion argument: AI makes student work better and learning worse, and the market eventually collects on the gap.
