A notification can break attention even when the phone is never touched. In a 2015 experiment, people pressed a key for every digit in a stream and had to hold that press back for one rare target, while a phone they never handled buzzed with a call, buzzed with a text, or stayed silent. Both interrupted groups failed to hold back more often than the silent group, by a margin the authors set alongside the cost of texting while driving. The phones stayed out of their hands. Receiving the request was enough. A computer handles the same problem with an [[interrupt controller::Hardware that sits between peripheral devices and the processor. It receives requests, assigns priorities, ignores masked lines, and delivers one interrupt at a time.]]. A keyboard, network card, timer, and drive can each raise a request on its own line. The controller decides which one reaches the processor next. An interrupt forces the processor to stop, save its current state, run a handler, and restore the state it had been using. The handler may take almost no time. Repeating the save and restore can still consume the machine. Under a heavy enough arrival rate, the processor enters receive livelock: throughput falls towards zero because it spends the whole period servicing requests and completes no useful work. Attention is the eighth function on this panel of twelve, and the controller is its piece of hardware. It decides which signal gets processed, which one waits, and which one is suppressed. An interrupt storm is what happens when those decisions never settle long enough for the selected work to move. Attention split into three networks in the model Michael Posner and Steven Petersen introduced in 1990. They separated alerting from orienting and executive control. Alerting gets the system ready, orienting aims it at a location or object, and executive control resolves conflict between competing responses. The three networks draw on different anatomy. Alerting runs through the locus coeruleus and its supply of noradrenaline. Orienting depends on the temporoparietal junction and frontal eye fields, while executive control recruits the anterior cingulate and lateral prefrontal cortex. The panel names the frontoparietal attention network because attention is distributed across this whole control system, with no single point that holds it. Brain imaging also separates the system you aim from the one that interrupts it. A dorsal network, including the intraparietal sulcus and superior frontal cortex, directs attention where you intend to put it. A right-sided ventral network detects salient, unexpected events and breaks into that plan. The 2002 imaging paper called the second network a circuit breaker, which is unusually literal language for the hardware analogy. That split also bounds it. Only the ventral half behaves like an interrupt controller, sitting outside the work and raising a line when something unexpected arrives. The dorsal half is closer to the scheduler that decided what the work was, and no interrupt controller does that. The locus coeruleus changes its firing pattern with the value of staying on task. Phasic firing accompanies engagement with the current work, while tonic firing accompanies disengagement and a search for something better. It receives direct input from regions that track whether the current task is still worth doing. Attention therefore includes a running decision about value, not only a spotlight pointed at sensory input. The original Attention Network Test made the three scores look independent in forty healthy adults. Alerting, orienting, and executive-control scores were essentially uncorrelated. A later appraisal pooled 1,129 people across fifteen studies and scored how well each network measure agreed with itself: .20 for alerting, .32 for orienting, and .65 for executive control. The same analysis found the three networks correlate with each other after all. A clean three-part diagram survived better than three clean personal measurements. Larger task batteries haven’t settled whether attention is one capacity or several. One study put 257 people through nine visual-attention paradigms and found a single general factor across roughly 1.3 million trials. Another gave 222 people eleven tasks and recovered four clusters: spatiotemporal, global, transient, and sustained attention. The number of attentions remains an open question because the answer changes with the tasks used to ask it. The Cattell-Horn-Carroll framework contains no broad attention ability. Attention appears inside other constructs instead, most plainly in working memory: the fourth Woodcock-Johnson already built a subtest called Verbal Attention into that score. Every cognitive test requires attention before it can measure its named function. Psychometrics has a name for the general version of this, task impurity, and every function on the panel suffers some of it; attention is the acute case because it is the one ingredient no task can be built without. A memory item has to be noticed, a reasoning rule has to stay selected, and a speeded response has to survive competing signals. The score belongs to the target function and to the gate that admitted the task. A processor can be benchmarked with its interrupt controller masked or removed from the path. That separation works because the controller is a distinct piece of hardware and the benchmark can feed the processor directly. A person offers no corresponding test condition. Factoring attention out would also factor out the instruction, the stimulus, and the response. Every function on this panel so far has broken the analogy on behaviour. This one breaks it on access. Human attention isn’t a detachable controller sitting beside an otherwise testable processor. It is part of the act of presenting any test to the processor at all, so the function being measured can’t be observed in isolation from it. Working memory has attention written into its definition. The current Woodcock-Johnson wording adds attentional control to short-term storage, making a low working-memory result partly an attention result by construction. A task that asks somebody to hold and transform material can’t distinguish a small workspace from a gate that keeps dropping the material. Randall Engle’s lab, which has spent two decades on this measurement problem, makes the strongest version of the claim from its own battery of 396 adults: measure attention control well enough and the link between working memory and reasoning is largely attention. Their own numbers don’t quite separate that account from the alternative, and the result is concentrated in one lab and its close replications. Restrict sleep and simple attention moves most, working memory moves less, and reasoning barely moves. The processor can still solve the problem. It receives fewer clean stretches in which to do so. Clinicians treat a low score on a timed test as a question about attention before treating it as a clean finding about the named function. Reading, processing speed, and memory can all look impaired when a person misses the instruction, loses the target, or trades speed for caution. That practice doesn’t yield a correction factor. It marks uncertainty that the test score can’t remove. A wrong answer doesn’t reveal which stage failed. The item may never have been selected, may have been dropped mid-operation, or may have reached an intact process too late for a timed response. Tests observe output without exposing the path that produced it. The famous attention tasks were designed to produce effects in almost everyone. The Stroop task slows the naming of an ink colour when the printed word names a competing colour. A flanker task slows a response when surrounding arrows point the wrong way. Stop-signal, go/no-go, and cueing tasks create their own conflicts, then turn the extra delay or error into an attention score. A good experiment and a good personal test need opposite properties. An experiment wants the manipulation to affect participants in the same direction, leaving little variation between people. A personal measure needs stable, sizeable differences between people so it can recognise who is who later. Craig Hedge, a psychologist who named this conflict the reliability paradox, showed why tasks that replicate beautifully in groups can rank individuals badly. [[Test-retest reliability::The degree to which a test gives people a similar rank when they take it again. A high value supports comparisons over time; a low value means the measure barely recognises the same people.]] is the property that matters here, and the classic tasks don’t have it. Three-week retests of seven of them returned correlations running from zero to .76, and the widest set reached .82. None of the reaction-time interference measures cleared .8, the level a test has to reach before it can reliably tell one person from another. The flanker cost managed .40 in one study and .57 in another. A larger battery splits the two questions apart. Several tasks agreed with themselves inside one sitting and then lost much of that ordering months later. The interval averaged 194 days, and elapsed time explained only about one per cent of the variation. The flanker effect reached .74 internal consistency and only .23 on retest. The trials cohered within the sitting, then barely recognised the same people months later. Difference scores discard much of the stable signal. Congruent and incongruent reaction times correlate strongly because both are driven by a person’s general speed. Subtract one from the other and general speed disappears, leaving a smaller residue with much more noise. Individual speed-accuracy choices add another source of variation that the subtraction can’t repair. The Stroop result may be narrower than its label suggests. In the same 396-adult battery its reaction-time effect grouped with essentially nothing but itself, and the three flanker measures formed a separate cluster of their own. Psychology’s best-known attention tasks measure stable experimental conflicts, and those conflicts don’t add up to an attention capacity. The [[psychomotor vigilance task::A simple reaction-time test in which a signal appears after unpredictable waits and the participant responds as quickly as possible. Responses slower than 500 milliseconds are conventionally counted as lapses.]] avoids subtraction. The standard version lasts ten minutes, with signals arriving after random waits of two to ten seconds. It reached .95 internal consistency in the same battery, the highest value in the table, and has almost no learning curve. Its six-month retest correlation was still only .55, so it reads today’s state and not a permanent fingerprint. The sustained-attention-to-response task asks people to respond to every digit except one rare target. Pressing on that target counts as a lapse, but caution can make the score look better. A person who slows every response may commit fewer errors without having stronger attention. The task confounds the gate with the policy used to protect it. Repetition speeds congruent and incongruent trials by roughly the same amount, so the interference score can sit still. An unchanged score on retest doesn’t show that attention held steady. It may show that both ingredients of a noisy subtraction improved together. Continuous-performance tests can’t diagnose attention deficit hyperactivity disorder on their own. A 2024 systematic review found the third Conners test was a weak predictor in five studies and adequate in only two. The evidence supports using these tests alongside symptoms and history because the score isn’t accurate enough to separate people who have the condition from people who don’t. Laboratory attention and clinical impairment overlap, but neither can substitute for the other. A useful attention measure has to match the question being asked. Conflict tasks are strong tools for studying a manipulation across a group. A ten-minute vigilance task is better for tracking whether one person is alert today. Put the two halves together and the subtraction becomes impossible in principle rather than merely hard: attention sits inside every score, and attention’s own measures can’t rank the same person twice, so the contamination can’t be estimated, let alone removed. Sustained attention degrades inside a single sitting. The 1948 clock test found detection falling within the first half hour of a two-hour watch, a pattern now called the [[vigilance decrement::The decline in detecting infrequent targets as time on a sustained-attention task increases. It appears within tens of minutes and changes with event rate and the kind of judgement required.]]. Watching for a rare event is active, stressful work, not an idle state between signals. Sleep restriction moves attention on a slower and more consequential clock. Participants assigned to four, six, or eight hours in bed for fourteen nights accumulated vigilance lapses steadily in the four-hour and six-hour groups, while the eight-hour group didn’t show the same rise. By the end, four hours a night produced lapse and working-memory scores comparable with two nights without sleep. Six hours a night matched one full night without sleep. Subjective sleepiness stopped tracking the decline. The six-hour group felt sleepier at first and then levelled off, while objective performance kept worsening, so they couldn’t tell day three from day fourteen. Sleep debt damages the instrument used to judge whether the debt is affecting you. Three recovery nights returned nobody to baseline. A separate seven-day study ran four sleep doses and then gave everyone three nights at eight hours. The three-hour group, the most damaged, recovered fastest and mostly after the first night, then stopped short. The five- and seven-hour groups recovered nothing at all: they had already settled at a reduced level during the restriction week and stayed there, finishing level with the group that had been sleeping three hours a night. Only the nine-hour group was unchanged throughout. The authors read that as adaptation, performance stabilising at a lower setting instead of climbing back to the old one. Attention is the first function sleep loss hits and the one it hits hardest. A meta-analysis covering seventy articles, 147 tests, and 209 effects found the largest group differences in simple-attention lapses and reaction time. Working-memory accuracy showed a moderate drop, processing speed a smaller one, and reasoning accuracy barely moved. An [[effect size::A standardised measure of how far apart two groups are, expressed in units of their shared variation. Larger absolute values mean a clearer separation, but they don’t say whether an effect matters in daily life.]] lets those tests sit on one scale. The near-null reasoning result fits a gate failure better than a damaged processor. Sleep-deprived people can still reason when the task reaches them cleanly. Simple lapses interrupt that delivery more often, so functions needing continuous contact with the task show more damage than reasoning accuracy itself. The popular numbers for this are mostly bad. Gloria Mark’s field programme measured attention on one screen falling from about two and a half minutes in 2004 to 47 seconds across studies from 2016 to 2021, but her early work used observation and the later work computer logging, so the series tracks switching under changing methods and not a capacity shrinking on one instrument. The famous 23-minute recovery penalty comes from a 2006 interview rather than from a paper, and the study usually cited for it found interrupted work finished faster, with more stress. The peer-reviewed figure from the same author’s field study is close but measures something else: twenty-four information workers took about 25 minutes and two intervening tasks to get back to an interrupted piece of work, which describes how a working day is organised and not how long a brain needs to recover. Heavy media multitasking, the claim with the most real literature behind it, pools to almost nothing across two meta-analyses. The eight-second attention span is not a research finding. It reached the press through a 2015 Microsoft Canada consumer report, where it sits in a figure credited to a statistics-aggregation website rather than to any of the work Microsoft actually did. A BBC journalist chased that website’s own cited sources, the National Library of Medicine and the Associated Press, and neither could find the research. Microsoft later took the report down. No laboratory measure of human attention produced the headline. Notifications have firmer evidence than a phone’s mere presence. A phone sitting on the desk appeared to cost working-memory performance in one influential study, where powering it off changed nothing even in the original result, and a preregistered direct replication of that experiment found no effect of location either. The notification experiment did produce more lapses even though nobody touched the phone. The interrupt has to arrive. Ageing changes the response policy as well as the speed. Across twelve sustained-attention studies, older adults were slower on ordinary responses and after errors, yet more accurate when withholding a response. That pattern is a trade of speed for caution. The ten-minute psychomotor vigilance task is the closest cognition gets to a blood-pressure cuff, and the only test on the panel with nothing else inside it to confuse the reading. It uses a simple response, has essentially no practice ceiling, and moves with sleep debt on a dose-response curve. The task doesn’t explain why attention changed, but it can show that the same person’s lapse rate changed before they reliably feel it. Daily use is what gives the task practical value. Repetition doesn’t create the steady climb seen on many cognitive tests, so a worse result is harder to dismiss as forgotten rules or an unfamiliar interface. A home version needs the standard structure to remain comparable with itself. Use a simple visual signal, random waits between two and ten seconds, and a ten-minute run. Count responses slower than 500 milliseconds as lapses, and keep the device, posture, input method, and time of day fixed. A short web game with constant timing measures something else. The useful record contains the date, sleep opportunity, time of day, and lapse count. The lapse count is the established sleep-sensitive outcome. Three-minute versions trade convenience for validity. A 2022 comparison found the short form had inadequate agreement with the ten-minute task, so a home shortcut is approximate. The full version also reached only .55 test-retest correlation over roughly six months in one sample. Sensitivity to today’s state is part of its value and the reason it can’t be treated as a fixed ranking. A low home score can’t diagnose attention deficit hyperactivity disorder, ageing, or a sleep disorder. A reaction-time task collects none of the symptoms and history a clinician would need. Sleep is the largest honest lever on attention, and it restores performance rather than adding a new capacity. Protecting enough time in bed removes a known impairment before any enhancement claim gets considered. Sleep extension has promising field evidence and weak causal isolation. Eleven Stanford basketball players added an average of 110.9 minutes a night for five to seven weeks, improved vigilance reaction time, cut sprint time from 16.2 to 15.5 seconds, and raised free-throw and three-point accuracy by about nine per cent. The study followed one team, included no parallel control group, and can’t separate extra sleep from the rest of the programme. Stimulants work differently in diagnosed attention deficit hyperactivity disorder and in healthy people. Across 133 double-blind randomised trials, clinician-rated symptom reductions were large for amphetamines and methylphenidate at the endpoint nearest twelve weeks, with different first choices for children and adults. A healthy-adult crossover trial found no objective gain across thirteen measures even while participants believed they had improved. On a complex optimisation task, methylphenidate, dextroamphetamine, and modafinil increased effort and persistence while reducing solution quality and efficiency. These intervention numbers sit on one scale while measuring different things. The stimulant values come from clinical symptom ratings, caffeine and meditation from objective attention tasks, exercise from broad cognition, and training from transfer beyond the practised task. Large treatment effects occur in a diagnosed group, small state changes appear elsewhere, and almost nothing travels far from training. The rest of the ladder is small. Caffeine moves accuracy and reaction time by a little over a quarter of a standard deviation across thirty-one trials, with the gain growing past 200 milligrams and reversing into anxiety and restlessness past 400. Meditation pools to about the same size on objective attention tasks, and the largest preregistered school trial, which taught mindfulness to thousands of pupils, moved no student outcome at all. Training still only teaches the trained task, as it did for working memory: reviews of working-memory training find reliable gains on untrained tasks of the same kind and nothing convincing further out once an active control is used. The prescription game EndeavorRx made the split visible in 348 children with diagnosed attention deficit hyperactivity disorder: its computerised attention endpoint moved, while every parent- and clinician-rated secondary endpoint failed to reach significance. The Food and Drug Administration cleared a changed test score without evidence of a change the adults around the child could see. Exercise shows how a small effect shrinks under better controls, falling by nearly half once active controls and baseline differences are accounted for and almost to nothing after correcting for publication bias, though a published rebuttal argues the review understates it. Everything on that list tries to make the gate hold harder. The best of it barely moves, and the one lever that clearly works repairs the gate instead of reinforcing it. Human suppression spends the resource it is trying to protect. A hardware [[mask register::A small hardware record listing interrupt lines the controller should ignore. Setting it is cheap and reliable, so a masked device can’t interrupt the processor.]] blocks a line with a single instruction and doesn’t ask the processor to resist each arriving request. A person has to notice a distractor in order to suppress it, and that act already used attention. There is no mask register in a person, nothing you set once and then stop paying for. The mask has to live in the machine, which means changing when requests are delivered instead of demanding stronger resistance once they have. Notifications, feeds, messages, and agent completions are all devices raising lines on the same controller. [[Interrupt coalescing::A computer technique that batches many device events into one delivery instead of interrupting the processor for each event. It trades a little latency for far less save-and-restore overhead.]] has been tested as a human notification policy. In a trial of 237 people, delivering smartphone notifications in three fixed batches a day improved how attentive, productive and in control they reported feeling. Hourly batches did little, while turning notifications off entirely increased anxiety and fear of missing out. Everything in that study is self-reported across two weeks, with no objective measure of attention in it anywhere. A separate three-week abstinence trial in adolescents, a pilot, points the same way: its gains in sleep, mood and working memory were gone by the two-month follow-up. Scheduled delivery beat both constant interruption and silence. My workstation creates many of its own interrupt sources. Background jobs, subagents, and review bots report when they finish, on schedules I configured rather than schedules a platform chose for me. I build with these systems every day. Their completions can collect until I read them in a batch, which keeps the automation without handing each completion the right to stop the current task. Notifications do interrupt, sleep debt does accumulate, and the machine can schedule both human messages and artificial-intelligence completions before they reach the person. There is no score to chase here. Every number in this piece was read through the gate it was trying to measure, which is why the lever is the delivery schedule and never the result. Set the mask in the machine, because there isn’t one in you.