Navigation
What We AreThe BrainPortfolioThe Lab's LabBuilt For YouThe WhiteboardServices & Prices
Let's Talk →
← The Whiteboard

Why can you keep a 500 day Duolingo streak and still not speak the language?

Because the streak certifies attendance, not understanding. The habit engine is the best thing Duolingo has and it is pointed at the wrong measurement. The fix is not to demand harder answers, it is to build a supported progression where the scaffold fades and the tick is earned for learning rather than completion.

The shape of it
Support fades as the proof deepens. The tick is earned at every step.
Full support. The system is doing the work, and that is correct here.
SCAFFOLDTHE LEARNER0% yoursday 1day 500
held upon your own
You watch the rule work before you ever touch it.

Nobody produces language they have never seen. First contact is a worked example: the rule inside a real sentence with the moving part called out. No blank box, no guessing, no penalty. It costs almost nothing, and skipping it is precisely why the next step feels like a cliff.

And when you get it wrong, the answer is a diagnosis
todaythe same item comes back later you recognise it you have learned that one sentence
insteadname which rule broke teach the why, and what it links to retest in a sentence you have never seen you have learned the rule

You can keep a 500 day streak and still not speak the language because the streak certifies that you showed up, not that you understood anything. The habit engine is genuinely good, daily practice really is how you learn a language, and none of that is the problem. The problem is what earns the tick. The fix is not simply to demand harder answers, because a blank box in a language you cannot speak is where people quit. It is to build a supported progression, from watching a rule work, to completing a sentence, to building one, to seeing how the rules constrain each other, with the scaffold fading as evidence accumulates and the streak point earned for the learning done that day.

Ross Jones, Founder, The Hopium Lab. Last modified 8 August 2026.

What is the streak actually measuring?

Attendance. That is the whole answer, and it is why the number can climb for 500 days while your ability stays flat.

The metric being optimised is whether you came back, and the app is extremely good at making you come back. But returning is a proxy for learning, and the moment a proxy becomes the target it stops being a good proxy. That is Goodhart's law, and a learning product is close to the purest example available. The engagement graph goes up and to the right. The outcome graph does not move. Nobody is lying, the wrong thing is simply being counted.

The tell is that day 500 looks exactly like day 1. If you had genuinely built understanding across 499 days, day 500 could not look the same, because it would have more to attach to.

Why doesn't repetition alone make it stick?

Because there is nothing for it to stick to.

Picking the right option out of four is recognition, the shallowest thing memory does. Recognition needs the answer to already be on screen, which means it can be perfectly intact while the ability to produce that same answer unprompted is completely absent. This is why the experience is so disorienting: you are not imagining the progress inside the app, you really are getting better at the task the app sets. The task just is not speaking.

Repetition works when it reinforces a structure. Without one, you are laying identical bricks in a row and never building a wall. Five hundred isolated repetitions is one repetition, done five hundred times.

Is gamification the problem?

No, and this is where most of the criticism goes wrong.

The habit engine is the most valuable thing Duolingo has built. Daily contact with a language genuinely is how you acquire it, and the streak is why millions of people come back at all. Rip out the gamification and you get a worse product that nobody opens, which teaches nobody anything. Being consistently present is a real precondition for learning.

So the fix is not to remove the engine. It is to keep the engine and change what it certifies.

Why isn't the answer just "make them produce sentences"?

Because a blank slate is a cliff, and learners will not walk off it.

This is the trap in the obvious critique, and I fell into it myself first time through. If you simply replace multiple choice with an empty box, most people will not attempt it. They do not yet know the rule, they have nothing to reach for, and an empty box in a language you cannot speak produces avoidance rather than effort. Multiple choice did not win because designers were lazy. It won because it always gives the learner something to do.

Difficulty on its own is not the goal. Difficulty the learner is equipped to meet is the goal, and the gap between those two ideas is where almost every serious learning product is won or lost.

So the answer is not a harder question. It is a process with the support built in, where each step asks for slightly more than the last and the scaffold comes away as the evidence says it can.

What would the progression look like?

Five steps, with support fading across them. The diagram above walks it, and you can watch the scaffold recede as the structure forms.

Meet the rule. You watch it work before you touch it: the rule inside a real sentence with the moving part called out. No blank box, no guessing, no penalty. This costs almost nothing, and skipping it is exactly why the next step feels impossible.

Complete the sentence. The bridge multiple choice pretends to be. The sentence is built except for the part the rule governs, and you supply that part. You are producing rather than selecting, but the surrounding structure holds you up. Learners will attempt this when they would refuse an empty box, and that difference is the entire design.

Build it yourself. The scaffold comes away. You produce the whole thing, with support available but never volunteered. Asking for it is data rather than failure, because it identifies precisely which part of the rule has not landed, which is worth more than a correct answer with no visibility into how it was reached.

Relate it to what you already have. The step that is almost entirely absent today, and where a language stops being a list. Rules do not live alone, they collide and constrain each other. Which rule wins here? What does changing the tense force you to change downstream? Where does this contradict the equivalent rule in your own language? Understanding is relational, or it is just inventory.

Transfer it. A context you have never seen, or two rules combined for the first time. If it survives that, you understood it, and it is the only step that is genuinely hard to fake.

What should happen when you get it wrong?

The wrong answer is the most valuable thing in the entire session, and it is currently thrown away.

When you miss a question, the system already knows a great deal about what happened. It knows whether you chose the wrong verb ending, the wrong gender, the wrong word order, or simply the wrong word. That is a diagnosis, and it is sitting right there. What happens next, almost everywhere, is that the same item is put back into the queue to resurface later. You see it again, you recognise it, you get it right. The product records a success.

What you have learned is that sentence. Not the rule underneath it.

Re-showing the identical item measures recall of the item, which is exactly the failure mode the whole product suffers from, reproduced at the level of a single question. It is also why people describe hitting the same handful of sentences forever without the underlying thing ever clicking.

The alternative uses the diagnosis:

  • Name the rule that broke, not just the answer that was wrong. Missing a verb ending is not a vocabulary problem, and treating them identically guarantees the wrong remediation.
  • Teach the why, and what it connects to. The rule, the reason it behaves that way, and which other rules it interacts with. That is what generalises past this one sentence.
  • Retest in a sentence you have never seen. This is the part that matters most. If the retest reuses the original item, a correct answer proves nothing, because item recall and rule understanding produce identical results. Only a new context separates them.

Same wrong answer, same ten seconds of the learner's attention. One path teaches a sentence, the other repairs a rule and everything downstream of it.

Why is the relational step the one that matters?

Because a language is a system of interacting rules, and almost nobody teaches the interactions.

You can know a tense rule, an agreement rule and a word order rule perfectly well in isolation and still produce broken sentences, because the difficulty lives in how they constrain each other. Explaining a single rule in isolation is better than matching a pattern, but it is still inventory. The understanding that actually produces fluent output is knowing which rule takes precedence, what a change in one forces elsewhere, and where the language behaves differently from the one already in your head.

That is also the step that makes everything before it stick, because a rule connected to three others has three ways to be recalled and one isolated fact has none.

What should earn the streak point?

Evidence of learning that day, at whatever step you are currently on.

This is the change that does the real work. The tick stops being awarded for opening the app and completing a lap of the mechanics, and starts being awarded for a demonstration appropriate to where you are: completing the sentence, or building it, or naming which rule applied and why. The demand rises as you do, so the streak becomes a record of understanding built rather than days survived.

Crucially, the tick is still available every single day, at every step. Nobody is locked out for being early.

Would retention drop?

That is the genuine risk in this design, and it deserves a straight answer rather than a reassurance.

Harder tasks add friction, and friction reduces daily completion. If you swapped multiple choice for unaided production overnight, completion would fall, and the people who left would learn nothing at all, which is a worse outcome than the shallow version. Anyone who tells you a deeper product is automatically a stickier one is selling something.

The reason I think this design holds retention is the scaffold. The learner is never facing a task they have not been equipped for, because the support only recedes as evidence says it can. That is what separates it from simply making the app harder. And the tick remains earnable daily, so the habit loop, the reason any of this works, stays completely intact.

What you would trade is the vanity of the number. A streak that certifies understanding will climb more slowly and mean vastly more, and some users will prefer the easy version. That is a real cost. It is also the difference between a product people use for 500 days and a product that got them somewhere.

If your core metric can climb for 500 straight days while the outcome stays flat, you did not build a learning product. You built an attendance product with very good art direction.

So why doesn't anyone build it?

Because attendance is trivial to measure and understanding is not.

Counting whether someone opened the app is free and unambiguous. Assessing whether a produced sentence is correct, whether a stated rule is sound, whether understanding transferred to a new context, all of that is expensive, fuzzy and much harder to put on a dashboard. So the industry measures the easy thing and quietly redefines it as the goal. Every incentive points the same way, because the easy metric is also the one that correlates with revenue.

That is a product decision, not a technical limitation. It was a much better excuse before language models could evaluate a produced sentence, interrogate an explanation and adapt the level of support in real time, at low cost. The measurement problem that justified the shortcut is the exact problem that recently became tractable.

Where does this argument stop?

It applies to anything claiming to build capability, and it stops where recall genuinely is the goal.

Some things simply have to be known instantly and correctly, with no underlying structure to discover: vocabulary, a safety procedure, number bonds. Drilling is the right tool there, and pretending otherwise wastes everyone's time. The critique lands specifically where a product claims to develop an ability, then measures something merely correlated with showing up.

Onboarding, compliance training and internal enablement all fail in the same shape. They measure completion because completion is easy to record. Ask instead whether the person could produce the thing, say why it works, and connect it to something they already do, and the whole picture changes.

Ross Jones, Founder, The Hopium Lab.