Desirable difficulty is the deliberate use of learning conditions that feel harder in the moment but produce stronger, more durable retention over time. The term comes from UCLA cognitive psychologist Robert Bjork, and it names something L&D leaders already half-suspect but rarely build into their programs: when a training experience feels smooth and effortless while someone is going through it, that ease is often a warning sign, not a compliment.
The brain encodes information most durably when it has to work to retrieve it, not when the answer is handed over cleanly. For teams under pressure to raise satisfaction scores and shorten time-to-complete, that finding is inconvenient. It may also be the most useful idea in the psychology of learning that most corporate training still ignores.
Where Desirable Difficulty Comes From
Bjork introduced the concept in the 1990s, building on research he conducted with Elizabeth Bjork on what they called the New Theory of Disuse. Their central insight was a distinction between two properties of memory that most people conflate: storage strength and retrieval strength.
Storage strength is how deeply a memory is embedded over the long run. Once something is genuinely learned, storage strength rarely goes down. Retrieval strength is different: it is how accessible that memory is right now, in this moment, and it rises and falls quickly depending on recent exposure. Re-reading a slide or rewatching a demo video boosts retrieval strength immediately. It makes the material feel familiar and fluent. What it does far less for is storage strength, the property that determines whether the material is still usable a month later, under pressure, without the slide in front of you.
Conditions that create desirable difficulty, deliberately recalling information instead of re-reading it, spacing practice out instead of massing it into one session, mixing related skills instead of drilling one at a time, do the opposite of what passive review does. They lower retrieval strength in the moment, which is why they feel harder while you are in them. But they build storage strength more effectively, which is why the learning holds up long after the training ends.
How It Works in Practice
Four conditions show up again and again in the research on desirable difficulty, and each one trades short-term comfort for long-term durability.
Retrieval practice asks a learner to reconstruct an answer from memory before being shown it, rather than simply reviewing the correct answer. Spacing distributes practice across days or weeks instead of concentrating it into a single session. Interleaving mixes different but related skills or problem types within a practice session instead of blocking them into separate, tidy units. Variation practices a skill across different contexts instead of the same static scenario every time.
Consider two sales reps learning the same objection-handling technique. One watches a demonstration video twice and can repeat the framework back accurately in the room, immediately after. The other has to reconstruct the response from memory during a roleplay, under mild time pressure, without a script in front of them. In the room, the second rep looks less polished. Three months later, in a live deal, the second rep is the one who still has it. The friction was the mechanism, not an obstacle to it.
Why Smile-Sheet Scores Get This Backward
Most L&D measurement systems reward exactly the wrong signal. Kirkpatrick Level 1 reaction surveys, the "how satisfied were you with this session" forms handed out at the end of training, capture how fluent and comfortable the experience felt in the moment. Meta-analytic research on training evaluation has repeatedly found that correlation to be weak, and in some conditions, to point in the wrong direction entirely. Learners tend to rate training highly when the material feels easy and the delivery is smooth, which is precisely the condition desirable difficulty predicts will produce the least durable retention.
This creates a structural problem most L&D leaders inherit rather than choose. Budgets, renewals, and program credibility often ride on scores collected at the exact moment those scores are least informative about whether anything changed. A program that feels effortful, that draws a few grumbling comments about pacing or difficulty, may be doing its job better than the one that earns uniformly glowing reviews and evaporates by week six.
Dan Docherty, Braintrust's Chief Coaching Officer and the author of NeuroCoaching, sees this pattern constantly in leadership cohorts. The sessions managers describe as "a little uncomfortable" in the room, the ones where they had to practice a difficult coaching conversation live instead of watching one modeled, are consistently the sessions that show up later in how those managers actually behave with their teams. The polished, easy-to-sit-through sessions rarely do.
Desirable Difficulty vs. What Most L&D Programs Actually Do
The gap between what the research supports and what most training programs are built to optimize for is wide, and it shows up in nearly every design decision a program makes.
| What Most Programs Optimize For | What Desirable Difficulty Requires |
|---|---|
| Passive content delivery: slides, videos, one-way modules | Active retrieval: recall before being shown the answer |
| One concentrated session or workshop day | Practice spaced across days or weeks, with deliberate gaps |
| Blocked practice: one skill drilled at a time | Interleaved practice: related skills mixed within a session |
| Smooth delivery and high in-session comprehension | Visible, productive struggle during practice |
| Success measured by end-of-session satisfaction scores | Success measured by applied behavior at 30, 60, and 90 days |
None of this means training should be needlessly hard or badly designed. Desirable difficulty is not an argument for friction as an end in itself. It is an argument that the specific kind of friction built into retrieval, spacing, and interleaving is doing real cognitive work, and that removing it in the name of a better learner experience quietly removes the mechanism that made the learning stick in the first place.
How to Apply This to Your Programs
Start with one existing program and change the ratio of exposure to retrieval. Most programs are built almost entirely around exposure: telling, showing, demonstrating. Replace at least one review touchpoint with a retrieval-only exercise, no notes, no slides, where learners have to produce the answer or the skill from memory first, then get corrected.
Build in deliberate gaps instead of a single event. If a program currently runs as one workshop day, break part of the reinforcement into sessions spaced a week or two apart, timed so learners are practicing right as the material starts to feel less familiar, not while it is still fresh.
Mix skills instead of teaching them in isolated blocks. If a program currently spends thirty minutes on discovery questions and then thirty separate minutes on objection handling, restructure part of the practice so participants have to move between both skills in the same scenario, the way they actually have to on the job.
Stop treating the end-of-session survey as the finish line. Pair it with a delayed check, a manager observation, a follow-up scenario, a 90-day behavior review, so the program is accountable to something the smile sheet cannot measure.
If your team's programs consistently score high on day one and quietly evaporate by day thirty, the content may not be the problem. The program may simply be engineered to feel good rather than to work. Desirable difficulty is not permission to make training miserable. It is a reminder that comfort and durability are rarely the same thing, and that the sessions your people occasionally complain about might be the only ones actually building something that lasts. If that gap between how your programs score and what your people can actually do six weeks later sounds familiar, that is worth a conversation.


