How ConvoLive teaches, and the research behind it

Updated

ConvoLive teaches one way. You open the app, a character starts a conversation with you in the language you're learning, and you talk back out loud until you've done the three things the scene asked of you. Order the drink, ask about the allergy, get the bill. Then it ends. Tomorrow there's another one.

That is a narrow design, and people ask why it's so narrow. Where are the grammar tables, the word lists, the multiple choice? The honest answer is that we started with those, watched what happened, and then went and read what the research says about how people actually become able to speak. This page is the second half of that answer.

Each section below is one decision, what we do in the app because of it, and the studies it rests on. In short:

  1. Immersion works because it makes you speak, not because it surrounds you
  2. Speaking is retrieval, and retrieval only improves by being practised
  3. A conversation with a goal teaches more than a drill
  4. Help should arrive after your first attempt, not before it
  5. Repeating a scene builds fluency, but repeat it tomorrow, not now
  6. A day's practice has an end, and tomorrow is already planned
  7. A streak with rest days outlasts a strict one
  8. Talking to an AI lowers the fear that stops people speaking
  9. Rewards for showing up cost motivation; honest feedback doesn't

The full sources are at the bottom, and every citation on the page jumps to its entry there.

Immersion works, but not for the reason people think

The most common language learning story goes: I studied for years, nothing stuck, then I moved there and it clicked. The usual explanation is that you were surrounded by the language. Input everywhere.

The research says the surrounding is necessary and not sufficient.

In the 1980s Merrill Swain studied children in Canadian French immersion schools. They had spent years being taught in French, thousands of hours of understandable input, and they could understand it very well. Their spoken French was still nowhere near what that much exposure should have produced. Swain's conclusion, which became the output hypothesis, was that comprehension and production are different skills, and that input alone does not build the second one. You also have to be pushed to produce language, notice the gaps when you fail, and try again1.

A cleaner test came in 2004. Freed, Segalowitz and Dewey compared three groups of French learners: a regular classroom at home, a semester abroad in France, and a seven-week intensive programme at home where students pledged to speak only French. The group that improved most in oral fluency was not the one in France. It was the domestic immersion group, who also reported speaking and writing far more French outside class than either of the others2. Twenty-eight students is a small study, and the one statistical predictor of fluency gains was, oddly, hours spent writing outside class, so the honest reading is narrower than the headline: what separated the groups was how much they used the language, not where they were standing. Living somewhere isn't the same as being made to talk.

So we didn't try to simulate being abroad. We tried to simulate the part of being abroad that works: ordinary life handing you a small problem with a goal attached, and no way to solve it except by saying something.

In the app: every lesson is a spoken conversation with a job to do. There is no reading mode and no listening mode as a separate activity. Even setup is spoken; the host asks what you want to learn and you answer out loud, so the habit starts in the first thirty seconds.

Speaking is retrieval, and retrieval has to be practised

Knowing a word when you see it and finding a word when you need it are different abilities, held in different strength. Reading and flashcards train the first. Speaking needs the second. I've written a longer piece on why that gap opens, but the short version is that the route from meaning to word only gets faster by being used.

The evidence that pulling something from memory teaches more than looking at it again is about as solid as anything in learning science. Roediger and Karpicke had students either reread a passage or take a test on it; a week later the tested group remembered far more, even though the rereading group had felt more confident3. Karpicke and Roediger then showed with foreign vocabulary that repeated retrieval, not repeated study, was what produced durable learning, with recall a week later roughly doubled4.

The study that matters most for a speaking app is Kang, Gollan and Pashler from 2013. Most language programmes ask you to repeat after a native speaker. Kang's team compared that with a condition where learners had to try to say the word themselves first and only then heard the model. Trying first produced better comprehension and better production of the words, and, the part everyone worries about, no loss of pronunciation quality5. You don't need to hear it first to say it well. You need to reach for it. Kang's study used single words and pictures, not sentences in a conversation, so it is adjacent evidence for an app like ours rather than a test of one.

Skill acquisition theory gives the mechanism. DeKeyser describes language knowledge moving from declarative (you can state the rule) to procedural (you can use it without thinking) only through practice at the actual skill, under conditions close to the ones you'll use it in6. Practice at reading makes you better at reading. Practice at producing sentences in real time makes you better at that.

In the app: you speak every turn. The end-of-lesson summary counts your lines in your own words separately from lines you read off a suggestion, because those are the lines that were retrieval.

Conversation, not drills

If production is the goal, why not drills? Say this sentence. Now this one. It would be simpler to build.

Because the research on interaction says the back-and-forth is itself where a lot of the learning happens. Long's interaction hypothesis proposed that when you and a partner have to negotiate meaning (they didn't understand, you rephrase, they check, you confirm), the negotiation itself makes the relevant language salient in a way a drill can't7. Mackey and Goo's meta-analysis of that research found interaction produced large gains on both immediate and delayed tests, and, notably, the effect on vocabulary was still strong weeks later8.

Task-based teaching takes this further: give learners a real task with an outcome that can be judged, and let the language be whatever it takes to get there. Bryfonski and McKay pooled 52 studies of task-based programmes and found a strong overall effect (d = 0.93); a later re-analysis put it at g = 0.61, and a 2023 review argues the studies are too varied to pool at all9. So the size of the effect is disputed. What is not disputed, there or in the interaction research above, is that learners need to use the language for something, and a task with an outcome gives the conversation a reason to exist.

The important word there is task. Open-ended chat has nothing to succeed at and no way to tell if it went well, which is where most "just talk to the AI" products lose people. A task has an end.

In the app: every scene has three objectives shown as a checklist. They tick off when you actually say the thing, and the conversation ends when they're done. If the character doesn't understand you, it says so and asks again, the way a person would, rather than showing a red cross; you find out whether a line worked from what comes back, and you fix it in your next one. A whole conversation takes a few minutes.

Help arrives after you've tried

Every learner hits the moment where they have nothing. A speaking app has to handle that moment or people quit. The question is when to help.

The research answer is: after the attempt, not before. Kornell, Hays and Bjork showed that trying to answer and failing, then seeing the answer, produces better learning than being shown the answer straight away, even when the attempt was hopeless10. Bjork's broader work on "desirable difficulties" collects a lot of results of the same shape: conditions that make practice feel harder and slower often make it stick better, and conditions that make it feel smooth often don't11. The Kang result above is the same finding in a speaking context: reaching for the word beats being handed it.

This is uncomfortable to build, because a hint offered up front feels kinder and tests better in the first minute. But a hint before the attempt turns retrieval practice back into reading practice.

In the app: when it's your turn there are suggestions you can tap to hear a phrase said properly, and a way to ask "how do I say..." in your own language. Both exist to keep you talking rather than staring, and both are optional: the conversation doesn't wait for you to use them. The end-of-lesson summary counts lines in your own words separately from lines read off a suggestion, because the reach is where the learning is, and that is the number to try to raise.

Repeat the scene, but tomorrow

You will notice the same situations coming around again in ConvoLive. That's deliberate, and the schedule is deliberate too.

De Jong and Perfetti had ESL students give a four-minute talk, then three minutes, then two. Half repeated the same topic, half got new topics each time. Both groups got faster during training; only the repeating group kept the gain on later tests, and their speed transferred to new topics they had never practised12. Repeating a task doesn't just teach you that task. It proceduralises the language inside it. Bygate's earlier work found the same: repeating a speaking task improved both fluency and complexity13.

Then there's the question of when to repeat. Suzuki and Hanzawa compared doing the same task six times in one sitting against spreading it across a class or across a week. Massed repetition removed pauses fastest, but it also slowed articulation and produced more verbatim repeats: learners were reciting, not speaking14. Kakitani and Kormos then tested one-day gaps against seven-day gaps between fluency sessions and found the same gains a month later for both15. A day's gap is enough. Doing it again immediately is the one schedule to avoid.

Spacing in general is one of the oldest findings in memory research; Cepeda and colleagues' synthesis covers over 250 studies16. Duolingo's own production experiment, on 3.3 million learners, found a better spacing model raised daily retention by 12% for any activity17.

In the app: a course is a hundred scenes, so the same functions (ordering, asking directions, describing a problem) come around in new settings rather than as a repeat of the same script. When you finish a scene, the app records your best run and invites you to beat it tomorrow, not now.

A day's practice has an end

How much per day is a design question too, and here the evidence is our own rather than a paper's. It is a correlation, not a trial, so treat it that way.

Of 461 people who signed up between 18 and 25 August 2026, those who finished four or more conversations on their first day came back on a later day 48% of the time, against 14% for those who finished none. But that group's median first session was 24 minutes, and most of them still never came back. A big first session is enthusiasm, not a habit. Three is a product choice, not a number from a study: enough to be a real session, small enough to finish and stop.

The best-evidenced cheap tool for turning one session into a second is the implementation intention: deciding, specifically, when you'll do it next. Gollwitzer and Sheeran's meta-analysis of 94 studies puts the effect at d = 0.65, which is large for something that takes one question18. Nunes and Drèze found people are far more likely to finish a goal that already shows progress toward it than one starting from zero19.

In the app: a day's practice is three conversations. When the third finishes the app says so, shows tomorrow's scene already lined up, and asks one question: when tomorrow? Your reminder fires then, naming that scene. You can do a fourth if you want, but the screen has already told you the day is done.

A streak you can miss

Streaks work; they're also brittle. Silverman and Barasch found that showing someone a broken streak makes them less likely to continue than showing an intact one, and that the damage is much smaller when the streak can be repaired20. Sharif and Shu found that goals with a small built-in reserve ("five days of seven") are both preferred and better attained than strict ones, and that people who miss a day recover far more often under a reserve21. Duolingo's biggest documented streak improvement was letting one small activity extend it instead of requiring the full daily goal22.

In the app: the streak is a day count with two rest days built into every rolling week. Miss two days and nothing happens. One conversation keeps it going. When it does end, the app doesn't make a scene of it.

Talking to something that isn't judging you

Speaking is the most anxiety-provoking of the four skills, and that anxiety measurably suppresses performance and willingness to speak at all23. MacIntyre's willingness-to-communicate model treats that willingness as the last step before speech happens, and so the thing every other variable ultimately has to move24.

One consistent finding from the last few years is that talking to a machine lowers the bar. Bashori and colleagues ran a quasi-experiment with 232 Indonesian secondary students: the groups that practised on speech-recognition websites ended up with lower speaking anxiety than the group in regular classes, and in interviews students said they felt less anxious speaking to the website than to people25. Tai and Chen found that adolescents who practised with a voice assistant became significantly more willing to communicate in English, with lower anxiety and higher confidence, because the environment was less threatening than a classroom26.

This is the one advantage an AI partner has over a human one, and we lean on it. The character has no opinion of you. You can be terrible at the café three times in a row and nobody but you knows.

In the app: you can start over any time, nothing is graded in front of anyone, and the first attempt at any scene is allowed to be awful.

What we deliberately leave out

Some things are absent because the evidence is against them, not because we haven't got to them.

There are no hearts, no energy, no gem shop, and no badge case. Part of that is evidence and part is taste, so here is which is which.

The evidence is about expected rewards and competition. Deci, Koestner and Ryan's meta-analysis of 128 studies found that expected tangible rewards for an interesting activity reduce intrinsic motivation for it, while information about how you're doing does not27. Mekler and colleagues found points, levels and leaderboards raised how much people did without changing how well they did it or how much they wanted to28. Hanus and Fox ran a semester-long classroom trial of badges and leaderboards and found lower motivation and lower exam scores in the gamified group29.

Hearts and energy are a different thing: not a reward but a penalty for mistakes, and the section on help above is an argument for making mistakes cheap. Leaving them out is a design choice consistent with that, not a finding.

So the end-of-lesson screen tells you what happened: which objectives you met, how many lines you said, how many were in your own words, how many the character understood first time, and your best run on that scene. Information about competence, which the same research says helps. XP is there as a quiet number. That's it.

There is also no free-form chat mode, for the reason in the conversation section: it has no end and no way to fail, and a task you can't fail can't feel like progress.

Conclusion

Say things out loud, in a situation with a goal, reaching for the words yourself as often as you can. Find out whether it worked from what comes back. Do the same kind of thing again tomorrow, not now. Three a day, then stop. Keep the streak, forgive the miss. No rewards for showing up, just an honest account of what you said.

None of that is original. It's what the research has said for a long time about how people learn to speak, applied without the parts that make an app feel nicer in the first minute and emptier in the first month. If you want to see how it feels, the app is free, and the first conversation is supposed to be hard.

The ConvoLive figures are from our own analytics for signups between 18 and 25 August 2026, with test accounts removed. "Came back" means any activity on a later calendar day, observed at least five days after signup.

Sources

  1. Swain, M. (1985). Communicative competence: Some roles of comprehensible input and comprehensible output in its development. In S. Gass & C. Madden (Eds.), Input in Second Language Acquisition (pp. 235–253). Newbury House.

  2. Freed, B. F., Segalowitz, N., & Dewey, D. P. (2004). Context of learning and second language fluency in French: Comparing regular classroom, study abroad, and intensive domestic immersion programs. Studies in Second Language Acquisition, 26(2), 275–301. doi:10.1017/S0272263104262064

  3. Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. doi:10.1111/j.1467-9280.2006.01693.x

  4. Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968. doi:10.1126/science.1152408

  5. Kang, S. H. K., Gollan, T. H., & Pashler, H. (2013). Don't just repeat after me: Retrieval practice is better than imitation for foreign vocabulary learning. Psychonomic Bulletin & Review, 20(6), 1259–1265. doi:10.3758/s13423-013-0450-z

  6. DeKeyser, R. (2007). Skill acquisition theory. In B. VanPatten & J. Williams (Eds.), Theories in Second Language Acquisition: An Introduction (pp. 97–112). Lawrence Erlbaum.

  7. Long, M. H. (1996). The role of the linguistic environment in second language acquisition. In W. C. Ritchie & T. K. Bhatia (Eds.), Handbook of Second Language Acquisition (pp. 413–468). Academic Press. doi:10.1016/B978-012589042-7/50015-3

  8. Mackey, A., & Goo, J. (2007). Interaction research in SLA: A meta-analysis and research synthesis. In A. Mackey (Ed.), Conversational Interaction in Second Language Acquisition (pp. 407–452). Oxford University Press. eprints.lancs.ac.uk

  9. Bryfonski, L., & McKay, T. H. (2019). TBLT implementation and evaluation: A meta-analysis. Language Teaching Research, 23(5), 603–632. doi:10.1177/1362168817744389. Re-analysis: Xuan, Q., Cheung, A., & Liu, J. (2025). Language Teaching Research. doi:10.1177/13621688221131127. Review: Boers, F., & Faez, F. (2023). Meta-analysis to estimate the relative effectiveness of TBLT programs: Are we there yet? Language Teaching Research. doi:10.1177/13621688231167573

  10. Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989–998. doi:10.1037/a0015729

  11. Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher et al. (Eds.), Psychology and the Real World (pp. 56–64). Worth. bjorklab.psych.ucla.edu

  12. de Jong, N., & Perfetti, C. A. (2011). Fluency training in the ESL classroom: An experimental study of fluency development and proceduralization. Language Learning, 61(2), 533–568. doi:10.1111/j.1467-9922.2010.00620.x

  13. Bygate, M. (2001). Effects of task repetition on the structure and control of oral language. In M. Bygate, P. Skehan & M. Swain (Eds.), Researching Pedagogic Tasks: Second Language Learning, Teaching and Testing (pp. 23–48). Longman. routledge.com

  14. Suzuki, Y., & Hanzawa, K. (2022). Massed task repetition is a double-edged sword for fluency development: An EFL classroom study. Studies in Second Language Acquisition, 44(2), 536–561. doi:10.1017/S0272263121000358

  15. Kakitani, J., & Kormos, J. (2024). The effects of distributed practice on second language fluency development. Studies in Second Language Acquisition, 46(3), 770–794. doi:10.1017/S0272263124000251

  16. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. doi:10.1037/0033-2909.132.3.354

  17. Settles, B., & Meeder, B. (2016). A trainable spaced repetition model for language learning. Proceedings of ACL 2016, 1848–1858. research.duolingo.com

  18. Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. Advances in Experimental Social Psychology, 38, 69–119. doi:10.1016/S0065-2601(06)38002-1

  19. Nunes, J. C., & Drèze, X. (2006). The endowed progress effect: How artificial advancement increases effort. Journal of Consumer Research, 32(4), 504–512. doi:10.1086/500480

  20. Silverman, J., & Barasch, A. (2023). On or off track: How (broken) streaks affect consumer decisions. Journal of Consumer Research, 49(6), 1095–1117. doi:10.1093/jcr/ucac029

  21. Sharif, M. A., & Shu, S. B. (2017). The benefits of emergency reserves: Greater preference and persistence for goals that have slack with a cost. Journal of Marketing Research, 54(3), 495–509. doi:10.1509/jmr.15.0231

  22. Duolingo (2020). Improving the streak. blog.duolingo.com

  23. Horwitz, E. K., Horwitz, M. B., & Cope, J. (1986). Foreign language classroom anxiety. The Modern Language Journal, 70(2), 125–132. doi:10.1111/j.1540-4781.1986.tb05256.x

  24. MacIntyre, P. D., Clément, R., Dörnyei, Z., & Noels, K. A. (1998). Conceptualizing willingness to communicate in a L2: A situational model of L2 confidence and affiliation. The Modern Language Journal, 82(4), 545–562. doi:10.1111/j.1540-4781.1998.tb05543.x

  25. Bashori, M., van Hout, R., Strik, H., & Cucchiarini, C. (2021). Effects of ASR-based websites on EFL learners' vocabulary, speaking anxiety, and language enjoyment. System, 99, 102496. doi:10.1016/j.system.2021.102496

  26. Tai, T.-Y., & Chen, H. H.-J. (2023). The impact of Google Assistant on adolescent EFL learners' willingness to communicate. Interactive Learning Environments, 31(3), 1485–1502. doi:10.1080/10494820.2020.1841801

  27. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627–668. doi:10.1037/0033-2909.125.6.627

  28. Mekler, E. D., Brühlmann, F., Tuch, A. N., & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior, 71, 525–534. doi:10.1016/j.chb.2015.08.048

  29. Hanus, M. D., & Fox, J. (2015). Assessing the effects of gamification in the classroom: A longitudinal study on intrinsic motivation, social comparison, satisfaction, effort, and academic performance. Computers & Education, 80, 152–161. doi:10.1016/j.compedu.2014.08.019