TOEFL 2026 Listening: four task types, one hearing, no way back
Listening is where the 2026 test's rules bite hardest. Every recording plays once. There is no Back. The software will not move on until you have answered. And the section is adaptive, so the 18-minute first module decides how hard the rest is. This guide covers the four task types as ETS defines them, the behaviour of the screen, and how to prepare for rules rather than for questions.
The four task types
Counts and CEFR ranges are ETS's. Together they make 47 items worth 35 raw points; the difference is the unscored items ETS trials inside the live test.
| Task type | Items | CEFR | What you hear | What it tests |
|---|---|---|---|---|
| Listen and Choose a Response | 15–19 | A1–B2 | One speaker, one short utterance — ETS caps it at six stressed syllables | Whether you caught the function of what was said: a request, an offer, a complaint, a question, a suggestion |
| Listen to a Conversation | 10 | A2–C1 | Two speakers, a short everyday or campus exchange | What was said, what was meant, and what one speaker will do next |
| Listen to an Announcement | 6–10 | A2–C1 | One speaker addressing a group about a campus or classroom matter | What is changing, for whom, and what you are being asked to do |
| Listen to an Academic Talk | 8–16 | A2–C2 | A lecture-style monologue, the longest recording — up to about 250 words | Main idea, supporting points, the speaker's attitude, and at the top of the range less common or idiomatic vocabulary |
How the section runs
Listening is two-stage. Everyone sits the same first module, the router, which ETS times at 18 minutes. Your performance on it decides the second module: the lower module, 7 minutes, or the upper module, 11 minutes. The upper module is longer because it carries the harder, longer recordings. Both branches hold the same mix of task types; only the difficulty changes, and it is the difficulty — not the raw count — that the band is built from.
- 1
The section directions
A screen naming the task types and the item range. Not on the clock.
- 2
The stage
When a recording begins the screen shows the speakers and nothing else — no question, no options, no Next. One figure means a single utterance is coming; two facing each other means a conversation. The audio plays once.
- 3
The questions
After the audio, one question per screen, the speakers shrunk to the left, the question and plain options on the right. Each question runs on its own countdown.
- 4
Next, and only Next
There is no Back. If you press Next without choosing, the software stops you. Choosing and pressing Next is final.
- 5
End of module
A screen marks the end of the router; the second module then begins with its own clock. You cannot return to the router.
Each task, and where it goes wrong
Listen and Choose a Response
The most numerous task and the shortest: one speaker, one utterance no longer than six stressed syllables, then a choice of replies. Nothing here tests vocabulary breadth. It tests whether you heard what the utterance was doing — asking, offering, apologising, checking, warning — and whether you know the reply that function takes. A speaker who says they have lost their key is not asking where keys are sold.
Where it goes wrong: matching a word in the utterance to the same word in an option. The correct reply usually shares no vocabulary with the prompt at all. Listen for the intonation and the function, then pick the reply a real person would give.
Listen to a Conversation
Two speakers on a campus or everyday matter — a room change, a missed deadline, a plan for the weekend — then questions. Typically one asks what was said or agreed and another asks what a speaker means, feels, or will do next. The second kind is answered by tone and by the last thing said, not by the topic.
Where it goes wrong: deciding early who is right and stopping listening. The question about what happens next depends on how the exchange ends.
Listen to an Announcement
One speaker addressing a group: a schedule change, a facility closing, a new procedure, an instruction to a class. Questions on the essentials — what is changing, who it affects, and what listeners are meant to do. Announcements are dense with specifics: a day, a time, a room, a condition. The question will turn on one of them.
Where it goes wrong: catching the topic and missing the exception — the one group the change does not apply to, or the one condition under which it does.
Listen to an Academic Talk
The lecture. Up to about 250 words of monologue on a topic from the sciences, history, the arts or social studies, delivered at a natural pace and built on a recognisable structure — a claim and its support, a comparison, a cause and its effects, a process in steps. The questions follow that structure: the main idea, a supporting detail, why the speaker mentioned something, and at the higher levels what an idiom or a less common word meant in context.
Where it goes wrong: listening for facts and losing the shape. A question about why an example was given cannot be answered from the example alone; it needs the sentence before it.
What the CEFR ranges mean in this section
Each task type carries a CEFR range, and the harder items in the range are what separate the bands. ETS's descriptors for Listening, condensed:
A1–A2
Follow everyday expressions and simple, familiar phrases when speech is slow and clear; pick out the main idea and the important details of a short message; predict what a speaker will do next; recognise what a short communication is for.
B1–B2
Follow everyday conversations and public announcements when clearly spoken; understand detailed spoken input on concrete and abstract topics at a natural pace; infer what is not explicitly stated; recognise rhetorical structures such as compare/contrast and cause/effect; cope with idiomatic and colloquial expressions.
C1–C2
Comprehend complex spoken texts including academic talks; extract specific information even from distorted or low-quality audio; infer speaker attitude, intention and implied meaning with little effort; understand rapid, idiomatic, colloquial speech and grasp subtle sociocultural and rhetorical nuance.
The practical reading: Choose a Response tops out at B2, so it can lift a low band but cannot on its own reach the top of the scale. The academic talk runs to C2; the highest Listening bands are earned there, on inference, attitude and idiom.
Voices, accents and what is on screen
ETS states that the recordings use AI-generated voices, with human actors in some cases, in a deliberately balanced mix of U.S./Canadian, British and Australian English and of male and female speakers; its Listening page lists North America, the U.K., New Zealand and Australia. A picture of each speaker sits on screen while the audio plays and stays there, shrunk, beside the questions. Recordings range from a single utterance to a monologue of about 250 words; the intermediate ones run 35 to 100 words and may have more than one speaker.
Two consequences. First, an ear trained on one accent will lose a question to the first unfamiliar one; practice audio has to be as mixed as the test's. Second, the picture is information: it tells you the task type, and therefore what kind of question is coming, before anyone speaks.
Preparing for the rules, not the questions
- Practise with one hearing, always. Replaying a recording in practice trains a reflex — wait, hear it again, then decide — that the test will punish. If you missed it, guess and move on; that is the skill.
- Answer in order, without the question in front of you. On the test the question appears only after the audio ends. Practice that shows the question first teaches you to listen for a keyword instead of for meaning.
- Treat the router as the exam. The first 18 minutes decide which second module you sit and therefore the ceiling on your band. Warm up before it, not during it.
- Use the stage screen. One figure: a function question is coming. Two: a conversation, so expect a question about what someone will do. A lectern or a group: an announcement or a talk, so expect specifics and structure.
- Commit inside the countdown. Each question is timed on its own and the software will not proceed without an answer. Hesitating past the clock does not give you a second recording; it takes the next question's time.
- Rotate accents. Aim for practice where British and Australian voices are as routine as North American ones.
How the Listening practice here works
The practice section reproduces the rules rather than just the questions. An 18-item router with the same task mix as the test; a 16-item second module chosen on the server from how you did, so you are never sent the branch you did not take; each recording generated in one of the four accents and played once; the stage screen, the per-question countdown and the block on an unanswered Next all present. The report tells you which branch you were routed to, marks every item with the line of audio it turned on, and gives an estimated band — estimated because ETS does not publish the conversion, and the report says so.
Questions about Listening
How long is the Listening section?
ETS's blueprint times the router at 18 minutes and the second module at 7 minutes if you are routed lower or 11 minutes if you are routed upper — 25 to 29 minutes in all, excluding the direction screen. In practice the clock is per question, not per module: each question has its own countdown once the recording ends.
Can I take notes?
Yes. ETS's Listening page states that you can take notes while listening to help you answer, and its test-content page says scratch paper and pencils are provided at a test centre and that Home Edition candidates should use a small whiteboard or a sheet of paper in a transparent protector with an erasable marker. For the shortest task there is nothing to note: a single utterance of a few syllables is over before a pen moves. For the academic talk, treat notes as a backup to attention rather than a substitute for it — the questions follow immediately and the talk does not replay.
Why can't I go back?
Because the section is adaptive and each question is timed on its own. Going back would let a candidate revisit an item after seeing later ones, which breaks the routing model. The consequence for you is simple: an answer is final the moment you press Next, and the software will not let you press Next until you have chosen one.
What happens if the recording plays while I am not ready?
It plays anyway. The stage screen — the speakers' pictures with no question, no options and no Next — is the only warning. Use it: the number of speakers and the setting tell you which task type is coming before a word is spoken.
Which accents will I hear?
ETS's Listening page says you may hear native-speaker accents from North America, the U.K., New Zealand or Australia; its blueprint describes a balanced mix of U.S./Canadian, British and Australian English, produced by AI-generated voices and sometimes human actors. If all your practice audio is one accent, the first British or Australian speaker on test day costs you a question while your ear adjusts. Our practice recordings are generated across those varieties for that reason.
How many of the 47 items count?
ETS reports 47 items and 35 raw points for Listening, and says that the item total includes unscored items used to trial future questions. You are not told which items are unscored, so every item has to be answered as if it counts — and the band is derived from the raw score by equating, not by simple arithmetic, so a raw count does not translate directly into a band.
Sources: ETS, TOEFL iBT Test: 2026 Update — Test Blueprint and Specifications (task types, item counts, CEFR targets, audio and timing), with the CEFR descriptors condensed from that document; screen behaviour from the official TOEFL iBT practice interface as observed. Checked 5 September 2026. ExamStep is an independent practice resource and is not affiliated with ETS. TOEFL® and TOEFL iBT® are registered trademarks of Educational Testing Service.
In this series: the whole 2026 format · what changed · Reading · Writing · Speaking · The 1–6 score scale · registration, fees and dates.