Desirable difficulties are study conditions that make learning feel slower and harder in the moment yet produce stronger long-term retention and better transfer to new situations. The term is Robert Bjork's, and the word desirable is the whole point: the difficulty is not an obstacle to learning but the engine of it. The catch is that the difficulty also makes learning feel less successful while it is happening, so learners and teachers tend to abandon the very conditions that work best.
Bjork's framework, the New Theory of Disuse, splits memory into two quantities that usually get confused. Storage strength is how deeply and durably something is learned. Retrieval strength is how easily it can be brought to mind right now. The two come apart. Re-reading a page or drilling the same item many times in a row spikes retrieval strength fast, so the material feels fluent and known, but it does little for storage strength, and the fluency drains away within days. Conditions that force effort, by contrast, raise storage strength precisely because retrieval was hard. The trouble is that learners judge their progress by current retrieval strength, the feeling of fluency, which leads to a fluency illusion: easy, massed study feels productive and is not, while effortful study feels unproductive and is. Most poor study habits trace back to trusting that feeling.
Four conditions have decades of evidence behind them.
Spacing distributes study of a topic across separated sessions rather than massing it in one block. The gaps let retrieval strength drop, so each return demands real effort and pays back in durability.
Retrieval practice means recalling material from memory, through testing or self-quizzing, rather than reviewing it. The act of pulling an answer out, even when it is effortful and error-prone, strengthens later memory far more than re-reading the same answer.
Interleaving mixes different problem types or topics within a session instead of practising one to completion before the next (blocking). Studying A B C A C B rather than A A B B C C feels messier and lowers in-session accuracy, but it forces the learner to discriminate which approach fits which problem, a skill blocking never trains because the type is given away by the sequence. Interleaving deserves singling out because it is the least intuitive of the set and the one most often skipped.
Varying the conditions of practice, changing the setting, format, or examples rather than repeating fixed ones, builds knowledge that survives outside the room it was learned in, again at the cost of feeling shakier during practice.
The principle comes with a hard boundary that the slogan often loses: not every difficulty is desirable, and the same manipulation can flip from helpful to harmful depending on the learner. A difficulty is only desirable if the learner has the prior knowledge to respond to it successfully; a difficulty the learner cannot meet is just a difficulty. For genuine novices, or for material with many interacting parts that must be held in mind at once, adding interleaving or generation can push the load past what working memory can carry, producing not deeper learning but collapse. This connects to the expertise-reversal effect documented by Slava Kalyuga and colleagues in 2003, where supports that help beginners become redundant for experts and, run the other way, challenges that sharpen experts can swamp beginners. The benefits also depend on the learner persisting through the unpleasant phase, and on the difficulty being one the task actually rewards. Difficulty for its own sake, illegible handwriting, needless complexity, is undesirable and simply degrades learning.
The first job is to break the fluency illusion: tell learners outright that smooth, easy practice is a poor signal of durable learning, because otherwise their own sense of progress will steer them wrong. Build spacing and low-stakes retrieval into the default rhythm of a course rather than leaving review to cramming. Interleave problem types once the basics are secure, and be explicit that the dip in in-session accuracy is the method working, not failing. Crucially, gate the harder conditions on readiness: front-load worked examples and supported practice for beginners, then withdraw the support and raise the difficulty as competence grows, rather than imposing maximum difficulty from the first lesson where it backfires.