Section 4 of 4 · 11 items · 8 minutes

TOEFL 2026 Speaking: eight minutes, eleven responses, nothing to press

The old Speaking section was four prepared tasks, three of them built on a reading or a lecture, in sixteen minutes. The 2026 section is eight minutes, no preparation time and no buttons: seven sentences to hear once and say back, then four questions from an interviewer on video that are never shown in print. It is also the section with the most raw points on the test. This guide covers both tasks as ETS defines them, how the microphone and the clocks behave, and how to prepare for a section that gives you no time to prepare.

The two tasks

Counts, points and CEFR ranges are ETS's. Both tasks span the whole scale; both are scored by machine out of five. Not adaptive.

TaskItemsPointsCEFRResponse clockWhat happens
Listen and Repeat75 each · AI-scoredA1–C2About 8 seconds a responseA scenario is set over a schematic illustration; you hear one sentence, once, and say it back. Each sentence relates to a different part of the picture, which is marked as you go. Sentences lengthen and grow more complex across the seven.
Take an Interview45 each · AI-scoredA1–C2About 45 seconds a responseA scenario screen, then an interviewer on video asks four questions aloud. The question is never shown in print. The four move from the factual to opinion, explanation and prediction.
Section1155 raw8 minLinear; all AI-scored; no preparation time

Section time, counts and points are ETS's. The response clocks are those observed in the official practice interface.

Fifty-five points in eight minutes

Every Speaking item is worth five points, and there are eleven of them: 55 raw points, against 35 for Reading, 35 for Listening and 20 for Writing. No other section comes close, and the section takes eight minutes. Whatever that means for how ETS equates the sections into bands — and ETS does not say — it means this for a candidate: each spoken response carries the weight of five Reading questions, and a response lost to a closed microphone, a restart or a silence is five points that cannot be recovered anywhere else in the test.

Each task, and where it goes wrong

Listen and Repeat — 7 items, A1 to C2

A scenario is set, aloud and in print, over a schematic drawing — a place with parts to it, a plan, a diagram. Then, item by item, a region of the drawing is marked and you hear a sentence about it, once. You say it back. ETS describes the progression: the first sentences are short and have one clause; the later ones carry compound verbs, dependent clauses, relative clauses and complex noun phrases. The scoring, in ETS's words, is on accuracy and intelligibility — the words, and the stress, rhythm and intonation that carry their meaning.

This is an elicited-imitation task, and what it measures is not memory for its own sake. A sentence longer than working memory can hold verbatim has to be reconstructed from its meaning and its grammar; a candidate who understood it can do that, and one who caught only sounds cannot. That is why the task spans the whole scale and why the long sentences separate the top bands.

Where it goes wrong: restarting. A candidate who stumbles at word four and begins again has spent the eight seconds on the first three words twice. Say the sentence once, keep the rhythm, and if a word is lost, go on to the next one — a gap is marked as a gap; a restart is marked as an incomplete sentence. The other failure is flat delivery: a correct sentence spoken as a list of words loses the intonation half of the mark.

Take an Interview — 4 items, A1 to C2

A scenario screen gives the setting and your role — a survey, an application, a conversation with someone who wants to know your view. Then an interviewer appears on video and asks four questions, one at a time, aloud. Nothing is printed. After each question a response panel opens with a countdown of about 45 seconds and the microphone is live; when the countdown ends the recording stops and the next question comes. ETS describes the questions as moving from brief factual statements to opinion, explanation, prediction and narration, with topics accessible to a general audience.

The marking, in ETS's words, is on clear, coherent elaboration with accurate grammar, varied vocabulary and intelligible prosody. Elaboration is the operative word: a one-sentence answer to a why-question is an answer at the bottom of the scale.

Where it goes wrong: using the first ten seconds to think. There is no preparation time, so thinking time comes out of the response. The shape that works is fixed: answer the question in the first sentence, give a reason in the second, give an example or a consequence in the third, and if the clock still runs, say what it means for you. The other failure is not hearing the question — it is spoken once, on video, by a person; listen for the question word and the tense, because a question about what you would do is not answered by what you did.

What the screen does

  • A microphone check comes first, before the section: you record, a level meter shows the signal, and this is the one moment at which a faulty headset can be fixed at no cost.
  • Scenarios are campus and academic, and the voices vary. ETS's Speaking page states that scenarios are based on academic and campus situations, that no specialised background knowledge is required, and that you may hear native-speaker accents from North America, the U.K., New Zealand or Australia.
  • Recording starts and stops itself. A response panel shows a microphone and a countdown; the microphone is live for exactly the countdown. There is no Record, no Stop, no Play, and no second attempt.
  • The illustration is annotated per item. In Listen and Repeat, the part of the drawing each sentence is about is marked as the sentence plays. It is context, and a cue to the vocabulary you are about to hear.
  • The interviewer is a video, and the question is only spoken. Nothing to read, nothing to re-play.
  • The scenario is given both ways. Aloud and in print, before the items. Read it; it tells you who you are in the conversation, and the register follows from that.

What the CEFR ranges mean in this section

ETS's descriptors and evidence statements for Speaking, condensed:

A1–A2

Reproduce a limited range of familiar sounds and simple phrases with support; repeat short, simple sentences with generally clear pronunciation, though stress and intonation may carry a first-language accent; respond to basic questions on familiar topics when speech is slow and clear; express basic preferences, describe routines and answer in short phrases. ETS's evidence statements at this level allow for omitted or substituted words that keep the basic meaning, formulaic answers, limited vocabulary and grammar, and frequent pauses.

B1–B2

Repeat longer utterances with mostly accurate pronunciation and stress, minor errors not impeding understanding; repeat sentences with good intelligibility and appropriate rhythm and intonation; describe experiences, explain opinions and respond with some elaboration at a conversational pace with occasional hesitation; speak fluently on familiar and some unfamiliar topics with a range of vocabulary and grammar at moderate accuracy.

C1–C2

Repeat complex sentences with high accuracy, using intonation and stress to convey meaning; repeat spoken text fluently and intelligibly with clear segmentation of meaning; respond fluently and spontaneously to complex questions with nuanced ideas, precise vocabulary and advanced grammar; speak at length with ease and clarity on abstract and complex topics with full control of language and discourse. ETS's evidence statements: fully intelligible with minimal listener effort, natural rhythm and intonation.

Two things stand out. Intelligibility runs through every level — the descriptors talk about listener effort, not about accent — so a strong accent that is easy to follow is not the problem, and a mild one that swallows word endings is. And the top of the scale is defined by prosody: stress, rhythm and intonation used to carry meaning. That is trainable, and most candidates never train it.

Preparing for a section with no preparation time

  • Practise repetition in meaning-chunks, once. Hear a sentence once, say it once, no second hearing and no restart. Work from short sentences up to ones with a relative clause. The habit is holding the sentence as a meaning, not as a string of sounds.
  • Read aloud for rhythm. Stress the content words, run the function words together, let the intonation fall at a full stop and rise at a question. The marking hears prosody; reading aloud with exaggerated rhythm for ten minutes a day is the cheapest band on the test.
  • Answer interview questions to a 45-second clock, cold. No notes, no pause before the first word. Answer, reason, example, consequence. Record yourself and listen for the ten-second gap at the start; when it is gone, you are ready.
  • Learn the question types by their first word. Do you… / Why… / How would you… / What will… map to fact, reason, hypothesis, prediction, and each wants a different tense and a different shape.
  • Fix the microphone before the test, not during it. A headset you have used before, in a quiet room, checked at the microphone screen. A response recorded through a laptop fan is a response the scorer cannot hear.
  • Retire the old material. Independent-task templates, integrated note-taking and thirty-second planning are training for a section that no longer exists.

How the Speaking practice here works

A microphone check first. Then seven sentences over an illustration whose parts are marked as each plays, heard once, with a self-running eight-second recording; then a scenario and an interviewer on video who asks four questions that are never printed, each with its own 45-second recording that opens and closes itself. The recordings are transcribed and each is marked out of five on what was said and how it was delivered — the words that were kept or lost in a repeat, and for the interview the answer's directness, elaboration, grammar, vocabulary and intelligibility — our reading of ETS's descriptors, since ETS publishes no rubric. The report quotes your transcript back to you against the sentence you were given.

Questions about Speaking

How long is the Speaking section?

ETS times it at 8 minutes for 11 items — the shortest section of the test. Each response has its own countdown: in the practice interface about eight seconds for a repeated sentence and about three-quarters of a minute for an interview answer. There is no preparation time on either task, and the countdowns are not shared between items.

Is there really no preparation time?

None. The old test gave 15 to 30 seconds to prepare; the 2026 test gives the response clock and nothing before it. For Listen and Repeat there is nothing to prepare — the sentence is the task. For the interview the first sentence of your answer has to answer the question, because the thinking time you would once have used is now inside the response.

What if I miss part of the sentence?

Say what you heard, in the right order, at a natural pace, and let the clock end. The sentence plays once and cannot be replayed. ETS's descriptors explicitly allow, at the lower levels, for omitted or substituted words that keep the basic meaning; a response with a gap scores something, a silence scores nothing, and a restart usually costs more than the gap did.

Can I re-record a response?

No. There are no buttons: the microphone opens when the countdown begins and closes when it ends, and the next item follows. What was recorded is what is marked. This is also why the microphone check at the start of the test matters — it is the only moment at which a bad headset can be fixed for free.

Do I get to read the interview questions?

No. The interviewer asks each question aloud, on video, and it is not printed anywhere on the screen. You answer what you heard. The scenario before the interview is given both aloud and in print, so you know the setting and the role before the questions start; the questions themselves are heard once.

How is Speaking scored?

By ETS's automated scoring, each of the eleven items out of five, for a section total of 55 raw points — more than any other section, and more than Reading and Listening combined. ETS describes the AI scoring as assessing features such as fluency, coherence and grammar; for Listen and Repeat the blueprint names accuracy and intelligibility, and for the interview clear, coherent elaboration with accurate grammar, varied vocabulary and intelligible prosody. The raw score becomes a 1–6 band by equating; ETS does not publish the conversion.

Sources: ETS, TOEFL iBT Test: 2026 Update — Test Blueprint and Specifications (task types, item counts, points, CEFR targets, difficulty progression, stimulus description and the section time), with the CEFR descriptors and evidence statements condensed from that document; response clocks, the microphone behaviour, the illustration and the video interviewer from the official TOEFL iBT practice interface as observed. Checked 6 September 2026. ExamStep is an independent practice resource and is not affiliated with ETS. TOEFL® and TOEFL iBT® are registered trademarks of Educational Testing Service.

In this series: the whole 2026 format · what changed · Reading · Listening · Writing · The 1–6 score scale · registration, fees and dates.