You know the name. You have known it for twenty years. The face is right there, the first sound might even be there, and the word will not come. Four hours later it arrives on its own, in the shower. Nothing was missing. The name sat in storage the whole time and the route to it was slow. Having something and being able to produce it on demand are separate capacities, and only the second one failed. Processors split the same way, which is why this function gets the [[L3 cache::The last and largest level of cache sitting on a processor, shared across its cores. It keeps a small amount of recently and frequently used data close to hand so the processor can avoid the slower trip out to main memory. Missing it costs time, not data.]] analogy on the panel. A hit there comes back in ten or twenty nanoseconds. Main memory takes several times longer and the drive takes thousands of times longer, for exactly the same data. Nothing moved. The path lengthened. The field calls this retrieval fluency, the seventh function on this panel of twelve. The metaphor earns its place because what stays quick to reach is decided by what has been used lately and often, in silicon and in people. It breaks in one spot though: a processor reading its cache leaves the cache untouched, while a person retrieving a memory changes which route works next time. Pull one word and a neighbouring one can get harder to reach. No cache does that. The failure is not that the memory is gone. On the premise the word arrives the moment you are told it is a kind of fruit, the material was there and the search was what broke. What stays reachable is decided by recency and frequency, not by importance. A cache holds what has been touched lately and often, and human access has the same indifference: a colleague’s name stalls while an advertising jingle arrives whole, because one route has had more traffic. Eviction is not deletion. It means a slower path to the same material. Robert and Elizabeth Bjork gave that split its cleanest theoretical form in 1992. Their New Theory of Disuse gives every memory two quantities: storage strength, how well learned it is, and retrieval strength, how reachable it is right now. Retrieval strength rises with use and decays without it, while storage strength, in their account, never falls. Nobody has measured storage strength on its own, though. It gets inferred from how quickly something is relearned, so the theory is a frame that fits the evidence rather than a reading off two dials. Psychometrics took until recently to separate the two. Broad retrieval was a single ability for decades. Schneider and McGrew recommended splitting it in 2018 after two decades of factor-analytic work, and the fifth edition of the Woodcock-Johnson acted on it in 2025: long-term storage covers how efficiently material gets in, and [[Gr::Retrieval fluency, the speed at which a person can load information from long-term memory into working memory for further cognitive processing.]] covers how fast it comes back out. The publisher’s own definition of the second is the speed at which a person can load information from long-term memory into working memory, which is a cache description with the hardware filed off. The two halves behave differently against general intelligence. Storage loads on it about as heavily as reasoning does. Retrieval fluency loads near the bottom of the panel, down beside processing speed, and further from storage than storage is from fluid reasoning. Two things the field called one construct until recently sit a long way apart. The tests that measure it ask for production against a clock rather than recognition at leisure: name as many as you can, now. A person can know every item on the list and still lose time getting each one out, which is the only thing being scored. The anatomy divides the same way. An anterior temporal hub holds the meanings, integrating what things look like, sound like and do, while a separate left-sided control network centred on the inferior frontal gyrus steers the search through it. Magnetic stimulation disrupts the second without touching the first, so the division is causal rather than inferred from scans. The clinical version is a matched comparison of eight people with stroke aphasia and six with semantic dementia: identical overall scores, different errors underneath. The dementia group, whose store had degraded, was heavily sensitive to how common and familiar each word was. The stroke group, with an intact store and broken control over it, was not. No single famous patient anchors this the way H.M. anchors memory formation, which makes it less cinematic and more typical of how the field actually settles things. Brown and McNeill caught the failure in progress in 1966 by reading out definitions of rare words: [[sextant::A navigational instrument for measuring the angle between a celestial body and the horizon, used to fix a ship’s position at sea. Brown and McNeill’s own prompt for it ran: a navigational instrument used for measuring angular distances, especially the altitude of sun, moon, and stars at sea.]], [[zither::A flat stringed instrument laid on a table or across the lap, its many strings plucked or strummed rather than bowed.]], [[caduceus::The winged staff with two snakes twined around it, carried by Hermes. It is routinely mistaken for the single-snake rod of Asclepius, which is the one that actually signifies medicine.]]. Their participants, mid-struggle, appeared to be “in mild torment, something like on the brink of a sneeze.” People who could not produce the word named its first letter more than half the time, and its syllable count about half the time. The control that makes those figures mean anything came later. Asked to guess the first letter of a word they simply did not know, people managed it one time in ten. Partial form without the whole word is the evidence: the first sound, the rhythm, the syllable count all reaching working memory while the word itself stays blocked. A drive doesn’t hand back the first byte and the file’s shape while withholding the file. Which words fail is not random. Salthouse and Mandell put seven hundred people between eighteen and ninety-nine through several kinds of retrieval, and age barely predicted failure on ordinary word definitions at all. It predicted failure on proper names strongly: a politician’s name from their face, a place from its description. What ages is access to arbitrary labels, not access to language. That effect largely survives controlling for episodic memory. Someone can store events perfectly well while names get less dependable, which is why “my memory is going” is not yet a description of anything. Vocabulary makes the pattern stranger. Pooling a couple of hundred studies, older adults beat younger adults on vocabulary by a wide margin, and the advantage is larger when they only have to recognise the word than when they have to produce it. Control for education and the recognition advantage disappears while the production advantage survives. The store grows. The route into speech gets relatively slower. The growth is real and measurable. Crowdsourced across roughly a million test takers, receptive vocabulary climbs from about forty-two thousand words at twenty to about forty-eight thousand at sixty, near enough one new word every two days. That is recognition rather than production, so it doesn’t show those words were all available in conversation. A bigger store is also more to search, with more neighbours competing. Bilingual speakers provide a natural experiment in frequency. They name pictures more slowly, produce fewer category items, and report more tip-of-the-tongue states in both languages, despite intact vocabulary and general verbal ability. Each word gets used less often because use is divided across two languages, which fits the frequency-lag account. Very rare words reverse the pattern, because bilingual speakers are sometimes less likely to know them at all, and an unlearned word can’t get stuck on the tongue. Hardware eviction ignores content; human access doesn’t. Names fail differently from nouns, related answers compete, and retrieving one item can push a neighbour further out of reach. Anderson and the Bjorks measured that last effect directly, and the meta-analysis pooling five hundred samples finds it real, small, and smaller still once the obvious artefact is controlled. The mechanism is still argued over. The direction isn’t. This is the rare function on the panel whose main research instrument is public. Letter fluency gives you sixty seconds for words starting with F, then A, then S. Category fluency gives you sixty seconds for animals. A letter, a stopwatch, a count. What is gated is everything built on top: the licensed batteries with their norms and derived scores, the Boston Naming Test, the rapid-naming subtests. The National Institutes of Health Toolbox contains no fluency measure at all, which surprises people who assume it does. The two tasks look alike and aren’t. Across thirteen hundred adults they correlate only moderately, which is the case for never averaging them into one score. Education drives the letter task more than age does. Age drives animals more than education does. Sex explains essentially nothing in either. Healthy older adults tend to land around nineteen or twenty animals in a minute, and that ballpark is not a personal norm. The published tables stratify by age and education for a reason, and testing in a second language, hearing loss, depression and anxiety all move the number by amounts nobody has pinned down well enough to correct for. Splitting the list into clusters and switches makes an appealing story: staying inside pets, then jumping to sea animals, with the staying attributed to the semantic store and the jumping to frontal control. The tidy anatomical version is contested, processing speed confounds it, and the decomposition is measurably less reliable than the raw count it was carved out of. Repeated testing creates a trap peculiar to retrieval fluency. Give healthy adults animal naming twice in a week and they produce about three more animals the second time. Give it to people with amnestic [[mild cognitive impairment::Measurable cognitive decline centred on memory that hasn’t yet removed a person’s independence.]] and the gain doesn’t appear. A flat retest is therefore the absence of a gain that should have been there, which is the reverse of how memory tests mislead, where the practice gain itself hides the decline. A total word count is still the strongest thing you can produce at home, and its meaning depends on the prompt, the time limit, the language, your age, your education, and whether you have sat it before. Keep the letter and category scores apart. Rapid automatized naming sits at the edge of this construct rather than its centre, predicting reading about as well as phonological awareness does, which earns it a clause rather than a section. Category fluency declines faster than letter fluency, in ordinary aging and much faster in disease. Letter fluency leans on frontal search. Category fluency draws on the semantic store itself, and the store is what degrades. Following Spanish-speaking older adults in northern Manhattan over years, letter fluency didn’t decline in the healthy group at all, while it fell steadily in people who went on to develop Alzheimer’s disease. Semantic fluency fell in both groups, and fell about three times faster in the future patients, accelerating as diagnosis approached. The ratio between the two scores, popular as a shortcut, worked worse than reading them separately. The French PAQUID cohort produced two different famous numbers from the same population. One analysis put the first detectable change twelve years before diagnosis. Another, using a different fluency test over a different window, put it at nine. Both describe the average gap between a group already sorted by future diagnosis and a matched group that stayed well, and neither tells an individual when anything started. A version of this figure also circulates attached to a large study of anticholinergic drugs that contains no fluency measure whatsoever, which is worth knowing before repeating the attractive sentence. Animal fluency does separate healthy adults from those with amnestic mild cognitive impairment, by about four animals, which is large as group differences go. It is also useless as a personal threshold. The impaired group’s average still sits inside the normal range, the studies disagree sharply with one another, and the authors’ own check says positive results reached print more easily than null ones. Verbal fluency has the shape of a longevity measure. Grip strength and gait speed are cheap functional tests whose reach exceeds the movement they record, and a sixty-second word count has the same appeal. Gait speed has already earned its own construct: paired with a memory complaint it becomes motoric cognitive risk syndrome, which roughly triples dementia risk across several thousand people. Fluency is not part of it. The Berlin Aging Study followed five hundred adults, aged seventy to a hundred and three at the start, for as long as eighteen years. Every one of them had died by the time of analysis, which is what let cognition and survival be modelled together rather than guessed at from a short follow-up. Of nine tasks across four domains, plus a composite intelligence score, only the two verbal fluency measures predicted mortality once age, sex, social position and suspected dementia were in the model. Not perceptual speed. Not episodic memory. Not vocabulary. Not the composite. The hazard attached to each extra word is small, around five percent, and the widely quoted nine-year gap between the top and bottom quarters of the sample is that small per-word effect stretched across the full spread. Terminal decline keeps that result from becoming a prescription. Cognition falls in the years before death because underlying health is deteriorating, so poorer fluency may be a readout of the same process that shortens the life. Ulman Lindenberger, one of the study’s authors, put the limitation plainly: training the test would change the readout without changing what it reads. A kitchen-table protocol can preserve the public research format exactly. Set a sixty-second timer and produce words beginning with F, then repeat separately for A and S. On another sitting, give yourself sixty seconds for animals. Record the prompt, the total, the date, the language and anything that altered hearing, mood or attention. Letter and category trials stay separate in the record, because they load the system differently and combining them buys a smoother number with less meaning. The first sitting is a baseline rather than a verdict. Repeat exposure can improve a healthy score, and the absence of that improvement can carry information in clinical groups. That makes tight retesting a bad way to reassure yourself with a familiar prompt. Population ballparks can’t supply a private percentile. Age and education, native language, hearing, mood and the chosen category all sit inside the result. A clinic can use proper norms and watch the error patterns a stopwatch won’t capture. Home testing suits a within-person trace, not a claim about where you rank. A sustained change across comparable sittings deserves context before interpretation: letter decline and semantic decline don’t mean the same thing, and a low count can’t diagnose anything. No confound-controlled longitudinal study has shown that habitual language use preserves retrieval speed with age. The cleanest natural experiment is the bilingual frequency lag, and that evidence is cross-sectional. “Use it or lose it” therefore remains a mechanism-shaped hypothesis rather than an established prevention result. The missing study would track language use and later unaided retrieval while separating education and cognitive reserve. Current access still responds to use, though. The Bjorks’ framework predicts that retrieving a word raises its retrieval strength even when storage strength was already high, and bilingual naming supplies the matching population pattern. Neither establishes that a daily word exercise slows anything. Giebl, Mena, Bjork, Storm and Bjork ran the most direct test of lookup order before generative artificial intelligence arrived. Programming novices either attempted problems before searching or searched immediately. The answer-first group performed better on a later transfer test, and related experiments find the benefit survives even when the initial attempt is wrong. A follow-up by the same group found people prefer to search first anyway, which makes this a habit to install rather than an insight to have. Generating an answer beats reading one, across dozens of studies. Applied to access rather than to learning, that means producing the word, the name, the explanation before you reveal it, so the internal route runs at least once before an external answer replaces it. Nobody has shown this protects against age-related decline. Retrieval also costs its neighbours. Producing the same few answers over and over can make the related ones you skipped harder to reach, an effect that is real and modest. Variety in what you retrieve is more defensible than drilling one fluent list. Everything sold for this fails in the same way. Strategy instruction lifts the score without evidence it reaches ordinary word finding, and the only trial of it ran on schoolchildren. Ginkgo comes back flat on memory, executive function and attention across every trial pooled. Brain stimulation over the relevant region produces gains in small clinical samples that reviews of healthy adults can’t find. Exercise earns its place on other grounds, and the fluency-specific syntheses disagree with each other. Cognitive training reliably improves the trained task and loses most of its distant benefit against an active control, with far transfer close to zero. That makes score chasing a poor goal, especially when the derived scores are less reliable than the total they came from. Cognitive offloading starts with a cost calculation. Risko and Gilbert’s model has a person estimate the effort of doing a task internally, choose an external aid, and then update future judgments from that choice. Each successful lookup can make the next lookup feel like the obvious route. The model describes a feedback loop in metacognition without claiming that every offloaded fact weakens memory. The original Google-effect paper split into one claim that failed and one that held. The eye-catching half, that hard questions leave people primed to think about computers, did not survive preregistered replication: a two-stage attempt returned strong evidence for no effect, and the wider battery it belonged to reproduced well under two-thirds of what it tested. The half that held is duller and more useful. People remember where to find a thing rather than the thing. The broader effect on memory is real and moderate, pooled across thirty thousand people from twelve to eighty-nine. It runs larger on a phone than a desktop, and smaller in people who already know a lot. Existing knowledge buffers the cost of offloading, which hands retrieval fluency a second job. The best-designed case of an offloaded skill actually degrading is satellite navigation, and it is thin: fifty drivers, worse unaided spatial memory with heavier lifetime use, and a three-year follow-up on thirteen of them in which heavier use predicted steeper decline. That follow-up is the part that rules out bad navigators simply reaching for the tool more often. Thirteen people in one lab can’t establish that the same thing happens to words. A sharper signal comes from a different modality. After AI polyp detection arrived in four Polish endoscopy centres, the same experienced doctors got measurably worse at finding adenomas when working without it, by six percentage points. The comparison is observational, built on three-month windows either side, and can’t separate complacency from genuine skill loss. Generative artificial intelligence has no equivalent study. A survey of a few hundred knowledge workers found most of them reporting less mental effort when using it, which measures a feeling at one point in time. A much-repeated finding that most of one group couldn’t quote their own essay comes from an unpublished study of fifty-four people, on an unvalidated electroencephalography pipeline, with a formal critique attached. Neither measured later unaided word finding. The experiment that would settle the practical concern hasn’t been done. It would randomize how often people receive an answer before attempting retrieval, then measure unaided recall and word finding months later under controlled conditions. No longitudinal or controlled study currently links generative-AI use to later retrieval decline. Hesper originally treated long-term memory as one component. I split it into storage and retrieval fluency after reviewing the commercial assessment catalogue and finding the field had already made the same move. The operational choice that follows is smaller than the taxonomy: decide which words, names, facts and phrasings should stay available without the model, then make one attempt before asking it. Language models can take the slow path in milliseconds and return better phrasing than the first human attempt. That makes them unusually complete substitutes for this function, since the output is exactly the word, name, fact or sentence that retrieval was trying to supply. It also makes the intervention cheap. The model can stay open while the answer waits for one internal attempt. Use the model everywhere, but answer first when the answer is one you want to stay able to produce.