How to practice speaking English with ChatGPT or Gemini: voice-mode setups, prompts that force correction, and where the ceiling is

By James Corrigan, English Teacher · Sep 2, 2026 · 13 min read

Can you really practice speaking English with ChatGPT or Gemini?
Yes, you can practice speaking English with ChatGPT or Gemini, provided you use voice mode and set the rules yourself before you start. Treat it as a conversation partner that corrects your grammar on request. Do not treat it as a coach: it cannot score your pronunciation, measure your pace, or keep a scored record of your level week to week.
I have taught English for a long time, and lately a good share of my students arrive having already tried this. The pattern is consistent: the ones who typed got a grammar tutor, the ones who spoke got a conversation partner, and almost none of them knew what the tool was quietly not telling them. This is the setup I hand out on day one.
At the time of writing, ChatGPT voice is open to all logged-in users, with free accounts on a daily allowance and a smaller model; Gemini Live is a free-flowing voice conversation you can interrupt by talking. Allowances change every few months; check the official help pages. The pipeline underneath is covered in How AI Language Tutors Actually Work.
How do you turn on voice mode and lock the ground rules?
Use the phone app, not the browser tab. Voice in ChatGPT sits behind the voice icon in the mobile app; Gemini Live is a mode inside the Gemini app on Android and iOS. Put the phone down at normal speaking distance.
Two settings matter. In Gemini Live, the Interrupt Live responses setting decides whether talking over the bot stops it; with it off, you tap the screen instead. In ChatGPT, by the maker’s own account, earlier voice mode detected the end of your turn from silence, so a thinking pause or a passing bus could end your sentence for you; the newer live option listens and speaks at the same time. Keep sentences short and finish them. A learner who pauses to find a word is exactly the case these systems handle worst.
Rules you say aloud may or may not survive the session, depending on memory settings, which vary by plan and region as of 2026. Durable rules belong in custom instructions or saved info, which the tool reads every session. Paste this once, there: “We are practicing spoken English. English only, even if I switch to another language; answer in English and ask me to repeat in English. Keep every reply under thirty seconds. Ask me one question at the end of each reply. Do not summarize, lecture, or list options unless I ask. Correct me only in the format I specify next.” Thirty seconds is not arbitrary: in a ten-minute session it gives you roughly ten turns each, so you speak for about five minutes. A thirty-student class gives you two.
Which correction prompts survive the chatbot’s politeness?
A general assistant is trained to keep you comfortable, and comfort is the enemy of correction. Left to itself it will admire your point and let “I am agree with you” pass for the entire session; even competitor app guides concede that it defaults to encouragement. You have to switch that off, because a bare “correct my mistakes” produces a lecture that eats your speaking time.
Three formats cover most needs. Recast-and-continue: “Before you answer, repeat my last sentence in correct natural English, then continue the conversation. Do not explain.” End-of-turn list: “After each of my turns, list at most three errors, most important first, with a one-line reason each, then continue.” Delayed review: “After ten exchanges, stop and list my five most repeated errors with an example of each.” One exam board publishes a version worth keeping for typed practice: “Correct this message and explain my mistakes in simple English.”
Now the caveat that no prompt fixes. The bot corrects the words it transcribed, not the sounds you made. If you say “sheep” when you mean “ship”, the speech recognizer, built to guess the most likely word from context, will often write “ship”, and the error vanishes before the model ever sees it. Your grammar gets marked. Your mouth does not. In a transcript-based pipeline the error vanishes before the model sees it; in the newer end-to-end audio models the model can hear it but gives an impression, not a score. Either way you get no per-sound measurement.
How do you run a role-play with a persona, a goal and a clock?
Role-plays work when the bot has an objective and a difficulty dial. Give it a persona, a goal that conflicts mildly with yours, and an instruction to escalate. Job interview: “You are a skeptical hiring manager for a project-manager role. Ask one question at a time. After each good answer, ask a harder follow-up about the same project.” Meeting: “You are a colleague who disagrees with my proposal. Interrupt me with a counter-argument every two turns; I have to hold the floor and finish my point.”
Three more: a hotel complaint with a receptionist who is unhelpful until you say exactly what you want; small talk with a colleague who gives short answers; a doctor’s appointment where you must describe a symptom precisely. Add: “Stay in character. Break character only for the end-of-turn error list.” Time-box the scene to eight to ten minutes and finish with “Give me a one-line debrief: what would a native speaker have done differently?”
The interview and the meeting have full guides, English Job Interview Practice That Sticks and English for Meetings: Prepare, Speak, Follow Up; a chatbot is the rehearsal room for both, not the syllabus. One warning: the bot will let you win, and a real hiring manager does not soften after your third weak answer. If the scene feels easy, tell the bot to be less impressed.
How do you lock the level with CEFR can-do statements?
Chatbots drift. They answer at whatever level your last message implied, and one ambitious sentence from you buys a reply full of C1 idiom. For an A2 learner that is not challenge, it is noise: you understand half, copy none, and spend the session nodding. The fix is to hand the bot the official description of your level and tell it to stay there.
The CEFR self-assessment grid, published by Europass, opens each spoken-interaction cell with a sentence like these. A2: “I can communicate in simple and routine tasks requiring a simple and direct exchange of information on familiar topics and activities.” B1: “I can deal with most situations likely to arise whilst travelling in an area where the language is spoken.” B2: “I can interact with a degree of fluency and spontaneity that makes regular interaction with native speakers quite possible.” Paste the one that fits: “Speak to me at CEFR B1, which means: [descriptor]. Use vocabulary and grammar a B1 learner would know. If I produce language above B1 twice in a row, tell me and move to B2.”
Be clear about what this is. It constrains the bot’s output; it is not an assessment of you, and whether the bot recalls the decision next week depends, as of 2026, on memory settings; it will not hold you to it. If you do not know your level, the free CEFR level test on this site gives you a starting letter, and CEFR Levels Explained: A1 to C2 in Plain English describes the scale. Start one level below where you think you are. An easy session costs little; a session you did not understand costs the session.
What can a general chatbot not do for your speaking?
Pronunciation. Pronunciation scorers compare your audio against a model of each target sound, and the good ones agree with human raters when they have enough speech to work with. That computation is separate from transcribing what you said. A chat model reasons over words. The newer voice models are trained end-to-end on audio, so they can hear you and may remark on a gross error, but an impression is not a per-sound score, and you cannot track an impression month to month.
Pace and fillers. The IELTS speaking descriptors mark hesitation, repetition and self-correction, and in the middle bands mention rhythm affected by a lack of stress-timing or a rapid speech rate. Measuring those means counting words per minute, pauses and fillers. A chatbot has no words-per-minute counter, no filler count and no pause map.
Memory, turn-taking, and the standard. As of 2026 both tools can recall past chats if memory is on, but neither keeps a scored, comparable record of your errors. Its turn detection was built to know when an assistant should answer, not when a learner needs a beat, and neither tool is designed to step in with a correction at the moment you need one, which is what a human coach does. The standard to hold feedback to is the current CEFR one: it scores articulation and prosody separately, expects intelligibility rather than a native accent, and notes that many C2 speakers keep a noticeable accent. Good feedback measures those. A chatbot measures none.
What is the five-minute test for your chatbot’s ceiling?
Do not take my word for any of the above. Run this once, in voice mode, and write down what happens. My prediction, from the mechanism rather than from hope: it will comment on word choice and grammar, not on articulation; it will estimate or decline on speed; and it may recall the topic, but it will not give you a comparable score. Either way, the list you write down is a precise description of the gap, and the only honest way to decide whether the gap matters to you. Repeat the test when your app updates:
• Minutes one and two: say three sentences in which you mispronounce a target word on purpose. Say “thought” with a hard t, “three” as “tree”, and stress “photograph” on the second syllable. Do not mention that you did it. Note whether the bot comments unprompted.
• Minute three: ask, “How fast was I speaking, and how many times did I say um?” Note whether you get a number, a guess, or a polite refusal.
• Minute four: ask for the end-of-turn error list and check whether the three deliberate errors appear in it.
• Minute five: end the session, open a new one, and ask, “What were my three errors in our last conversation, and how many words per minute did I speak?” With memory on it may recall the topic; note whether you get a number.
What does the research actually show?
The research is small, short and mostly uncontrolled, and it points the same way. A 2025 quasi-experimental study in Ecuador had 49 high-school learners practice with ChatGPT voice for three weeks; automated pre- and post-tests showed gains on pronunciation, fluency, vocabulary and grammar, largest for fluency and vocabulary. No control group, one school, machine scoring: read it as “speaking more for three weeks helped”, not as proof the tool caused it.
A 2026 qualitative study followed 12 undergraduates in Semarang, Indonesia, through three weeks of the same practice. They reported immediate feedback, grammar correction and dialogue prompts that raised confidence and lowered speaking anxiety; they also reported software bugs, difficulty learning pronunciation from the tool, and anxiety at the start. Twelve people is twelve people.
So: a consistent signal on anxiety and willingness to speak, weak evidence on pronunciation, no long-term data. That matches what I see: students who use these tools talk more and flinch less, which is worth a great deal; English Speaking Anxiety: Why You Freeze and What Helps explains why the flinch matters. What no controlled study yet shows is better articulation: the one study that reports a pronunciation gain had no control group and scored it by machine.
When is a dedicated speaking app worth it, and when is it not?
Stay with the free chatbot if any of these describe you: you are B2 or above and mainly need topics and time on your feet; your problem is confidence rather than accuracy; or a human already corrects you and you only need reps between those sessions.
Move to a dedicated speaking app if any line from your five-minute test bothered you: no pronunciation score, no pace or filler count, no scored record of your level week to week, no feedback on the sounds and pace while you are still in the conversation. Full disclosure: I write for Lucida’s blog, so discount accordingly. The factual case is short and it sits in the box at the end of this page: real spoken conversations, real-time feedback on the six things your test just listed, six languages, free to download. I will not rank it against other apps here; Best AI Apps to Practice Speaking English (2026) does that job.
Use both. A 20-minute daily budget, five days a week, is 100 minutes. Give three days to the chatbot for open topics and role-plays, 60 minutes, and two days to the app for measured sessions where pronunciation, pace and fillers get numbers, 40 minutes. Write the measured numbers down on the first of every month. Over a year that is roughly 87 hours of sessions and, at the half-share of airtime from the ten-turn rule, something like 45 hours of you actually speaking. For scale, Cambridge’s guided-learning-hours table rises from about 100 taught hours at A1 to several hundred per level higher up, and those are classroom hours, not solo practice, so treat the comparison as a scale check, not a forecast. The chatbot hours count only if a correction format is switched on. The measured hours tell you whether anything moved.
Frequently asked questions
Can I practice speaking English with ChatGPT for free?
Yes, at the time of writing. Voice conversations are available to all logged-in users, and free accounts run on a smaller model with a daily allowance of hours that changes without notice; paid plans lift the cap. Check the official help page for the current limit rather than a blog post, this one included, and keep sessions short so the allowance covers a daily habit.
Is Gemini better than ChatGPT for English speaking practice?
Neither is built for it, and the difference for a learner is small. Gemini Live lets you interrupt by talking, supports more than 40 languages on Android and up to two languages on one device, which suits switching between English and your own language. ChatGPT’s voice has the longer track record and the same custom-instruction trick. Run the five-minute test in both and keep whichever you will actually open tomorrow.
Does ChatGPT correct my pronunciation?
Not in a way you can track. It corrects the transcript of what the recognizer heard, and the recognizer repairs many mispronunciations into the intended word. The newer audio models can hear you directly and may remark on a gross error, but there is no per-sound score, no comparison against a target, and no history. For pronunciation feedback you need a tool that scores sounds, or a human.
What is the best prompt for practicing English speaking with AI?
There is no single best prompt, and be suspicious of any list that claims one. The one that changes most for most learners is the recast: “Repeat my last sentence in correct natural English before you answer, and do not explain.” Combine it with a level lock (“speak to me at CEFR B1”) and a thirty-second turn limit, and store all three in custom instructions so they survive the session.
Can I use ChatGPT to practice speaking Spanish or another language?
Yes; the same setups work for Spanish, French, German, Italian or Portuguese. You can ask during a voice conversation to switch language, and Gemini Live supports two languages on one device. Write the session contract in the target language, lock the level with that language’s CEFR descriptor, and remember that the blind spot is identical: the transcript gets corrected, your accent does not.
Can an AI chatbot replace an English teacher?
No, and distrust any app that claims otherwise, ours included. A chatbot gives unlimited patience and instant grammar recasts; a teacher hears your sounds, notices what you avoid, remembers last week, and interrupts you at the right moment. A dedicated speaking app closes part of that gap with measured feedback. Use the chatbot for volume, the app or the teacher for correction, and a real human for the stage.
Run the five-minute test. Then fill the gap.
Lucida is an AI speaking coach app for iOS and Android: you hold real spoken conversations with an AI coach and get real-time feedback on pronunciation, grammar, vocabulary, fluency, pace and filler words. It is free to download, with some features requiring a subscription. Keep the chatbot for open topics; use Lucida for the sessions that get measured.
