CELPIP Task 3: Describing a Scene, and Task 4 After It
Describing a Scene has no direct IELTS counterpart, and Task 4 uses the same picture. How the pair works, why people run dry, and how to practise without real prompts.
Describing a Scene is where CELPIP preparation most often goes wrong, and it is rarely because the task is hard.
It goes wrong because people practise it alone. Task 3 hands you an image, Task 4 immediately asks what happens next in that same image, and a description built as a list of objects leaves you with nothing to say when the second question arrives.
If you want to build the habit of talking about any picture without freezing, and without stopping the moment it feels finished, SpeakShark is free with three AI conversation sessions a day and no card, each capped at five minutes and four turns.
In this guide: what the task is · the pair problem · why people run dry · describe a situation · the language that scores · a practice routine · common mistakes · what we could not verify · how we researched this guide · FAQ
Key takeaways
- Describing a Scene is task 3 of the 8 on the official CELPIP test format list, and Making Predictions is task 4 on the same image.
- Reading the two published task lists side by side, our count is that four CELPIP tasks are not set as discrete tasks by IELTS. This is one of them, so the picture itself is new even if the habit of speaking alone is not.
- Listing what is visible is not the failure by itself, since Paragon's rated sample credits accurate detail. Skipping the overview, never addressing the listener and leaving Task 4 nothing to build on are.
- Speculation such as it looks like is safer than assertion, and it hands you material for the next task.
- Paragon publishes official free sample material, which is the only reliable source of real images.
What Describing a Scene actually asks
The official CELPIP test format page lists eight speaking tasks, one question each, across a fifteen minute section. Task 3 is Describing a Scene.
You are shown an image and asked to describe it to a listener who cannot see it, and that framing carries more requirements than it looks like. Paragon's own published sample analysis marks answers down when they never address that listener, and when they never give an overview before the details.
What the instruction does not hand you is a structure to follow, a list of things to notice, or an examiner to nudge you when you stall, because CELPIP has no interviewer in the room at all.
That last point matters more here than on any other task. On a live exam the examiner takes over again the moment your turn ends, whereas on CELPIP the recording simply keeps running until the clock stops it. Nobody responds to you at all on the day: Paragon's 2024 Data Report says responses are assessed afterwards by at least three trained and certified raters. Our CELPIP Speaking guide covers the full eight task structure and what the published scores show.
We have not reproduced any real image or prompt here. Question content belongs to the exam body, and Paragon publishes official free sample material for anyone who wants to see the real thing.
Tasks 3 and 4 are one problem
This is the single most useful thing to understand about Task 3, and most guides treat the two tasks separately.
The official list puts Making Predictions immediately after Describing a Scene, and Paragon's own free Speaking Pro study pack describes Task 4 as working from the same illustration you have just described. Having said what is happening, you say what happens next.
Prediction as a language function is not unique to CELPIP, and IELTS Part 3 pushes candidates toward comparisons and predictions too. What is specific here is predicting the next state of one particular image you have just described out loud, which is why the quality of the description decides how much you have to work with.
So the two tasks share a dependency that runs one way. A description with nothing unresolved in it leaves you with nothing to predict. If you spend your description naming a bench, a dog, a red jacket and a bus stop, then Task 4 arrives and the honest answer is that nothing in particular happens next, because you never established a situation that could develop.
Prepare them as a pair. Every time you practise a description, force yourself to end it somewhere unfinished: somebody is about to do something, somebody has not noticed something, somebody looks like they are waiting. That ending is the handle Task 4 grabs.
Why people run dry after twenty seconds
The most common failure is not vocabulary. It is running out of things to say while the recording continues.
Three causes, in the order they usually appear.
You described the obvious first and left yourself nothing. Openings that name the setting and the main people burn the easiest material in ten seconds. Better to open with the situation, which is inexhaustible, and let the details arrive as evidence for it.
You are describing objects rather than relationships. A picture contains a finite number of things and an unlimited number of relationships between them. Who is with whom, who is ignoring whom, what one person seems to want from another.
You stopped because you felt you had finished. With no examiner reacting, there is no signal that more is wanted. Candidates who practise with a partner learn to stop when the listener looks satisfied, and that habit is actively harmful here. Practising English speaking alone is the closest deliberate training for the silence.
Describe a situation, not an inventory
The strongest framing we know for this task is that the image is one frame from a story you can see the middle of.
An inventory sounds like this in structure: there is a park, there are three people, one is holding a bag, the weather is sunny. Every line of that can be accurate, and accuracy does get credit in Paragon's rated sample, but the answer never says what kind of moment this is and never speaks to the person who cannot see it, which is exactly what that analysis marks down. It also cannot be extended.
A situation sounds different. Somebody appears to be waiting for somebody else. One person has clearly just arrived and the other has been there a while. Somebody is holding something they seem reluctant to hand over.
Side by side, the difference is not effort, it is what the same observation is used for.
| Inventory answer | Situation answer | |
|---|---|---|
| Opening | Dives straight into details | Gives an overview of the moment first |
| Details are | The point | Evidence for an interpretation |
| Tenses used | Usually present continuous only | Continuous, simple, perfect, conditional |
| Can it run long? | No, objects are finite | Yes, relationships are not |
| Leaves Task 4 with | Nothing to predict | Something unresolved to extend |
| Risk | Sounds like a list | Speculation, which is safe language |
The practical structure that gets there:
- Open with an overview in one sentence, addressed to your listener. What kind of moment is this, said to somebody who cannot see the picture. Paragon's rated sample marks answers down for skipping exactly this.
- Place the people relative to each other, not relative to the frame. Beside, behind, facing away from, ignoring.
- Give the evidence. The details you would have listed now arrive as reasons for your interpretation, which is a far better use of them.
- Add what is generally true of the place, in present simple, which widens your tense range at no cost.
- End on something unresolved. This is the handle for Task 4.
The language that shows range
Four things reliably widen an answer without any extra vocabulary study.
Mix your tenses deliberately. Present continuous for the action, present simple for what is normally true of the place, present perfect for what has just happened, and a conditional or two when you speculate. An answer entirely in present continuous announces its own ceiling.
Use spatial language beyond in the middle. In the foreground, off to one side, with their back to us, further along. These are cheap to learn and immediately visible in a recording.
Vary your adjectives and put them to work. Not just a big bag but a bag heavy enough that she is leaning to one side. Description that implies something is worth more than description that labels.
Speculate out loud. It looks like, they seem to have just, I would guess that, judging by. Speculation is the safest language in the whole task because nobody can contradict it, and it doubles as preparation for Task 4.
For clarity underneath all of it, six pronunciation fixes in fourteen days targets the sounds that actually block comprehension, and reducing mother tongue influence covers the patterns beneath them. If you hesitate rather than lack words, how to stop hesitating in English is the more relevant fix.
A practice routine with no real prompts
You do not need CELPIP images to practise this. You need pictures and a way to record yourself.
- Open any photo. Your camera roll, a news site, anything with people in it.
- Describe it without stopping, and keep going past the point where it feels finished. Record yourself. Stopping early is the failure you are training out, and Paragon's official sample material is where to see the real length for this task in the real interface.
- Immediately predict what happens next, working only from what you said, and again keep going until you run dry. This is the Task 4 rehearsal and it is where you discover whether your description left a handle.
- Listen back once, for one thing only. First pass: did you list or did you interpret. Second pass another day: how many tenses did you use.
- Do it daily rather than long. Five minutes a day beats an hour on Sunday, because what is being trained is retrieval under pressure rather than knowledge.
- Practise the silence deliberately. No partner, no reactions, just you and the recorder, because that is the condition on the day.
A conversation partner still helps for everything else in CELPIP, particularly Task 6, where Paragon describes the job as explaining a decision to somebody and backing it with reasons. Role play scenarios and English job interview role plays cover that register. To keep the daily habit going, free methods to raise fluency and a 30 day plan to improve speaking are the groundwork, and you can start a free session to get the reps in.
The mistakes worth fixing first
- Skipping the overview and never addressing the listener. These are the two weaknesses Paragon's own rated sample names, and naming things instead of interpreting them is usually what causes both.
- One tense throughout. Usually present continuous, and it caps your visible range.
- Stopping when you feel finished. There is no examiner to tell you more is wanted.
- Describing Task 3 with no thought for Task 4. They share an image and a fate.
- Asserting what the image does not support. Speculate instead; it is safer and it is better language.
- Practising only with real prompts you found online. Unofficial images teach nothing that an ordinary photograph does not, and Paragon publishes official samples anyway.
If you are still deciding between tests, CELPIP versus IELTS for Canada PR covers the choice, and test your English speaking level free gives you a baseline before you commit. The nearest task anywhere else is covered in our PTE Describe Image guide, though Pearson lists that stimulus as a graph, picture, map, chart or table, which in practice means a chart, graph, map or table more often than a social scene you have to read for context.
What we could not verify
We are not stating preparation or response times for individual tasks. Figures circulate for how many seconds you get to prepare and to speak on each of the eight tasks. The official test format page publishes only the fifteen minute section length rather than a task by task breakdown, so rather than restate numbers here we point you to the per task table in Paragon's own free Speaking Pro study packs on celpip.ca, which also match the real interface.
We are not stating a scoring rule for factual accuracy. Claims that you do or do not lose marks for misdescribing something are common online and we found no official statement either way.
We have not reproduced any test content. No real image, no real prompt.
How we researched this guide
The eight task names, their order, the fifteen minute section length and the absence of a separate speaking session come from the official CELPIP test format page. Paragon's own free Speaking Pro study pack describes Task 4 as working from the same illustration as Task 3, which is why we treat the two tasks as one problem rather than inferring the link from their order. The rated sample analysis and the Task 6 description come from the same free study pack material on celpip.ca, and the multi rater process comes from the 2024 Data Report.
The count of four tasks with no direct IELTS counterpart, and a fifth that overlaps only loosely, is our own reading of the two published task lists. Neither Paragon nor the IELTS partners publish a task level mapping between the exams.
Everything else here is speaking practice advice rather than exam reporting, and we have marked it as such. Where a figure would have made the guide more concrete but could not be traced to Paragon, we have left it out and said so.
Practise speaking, from SpeakShark
SpeakShark is an AI English speaking practice app, and it is our pick for this task for one specific reason: the thing that fails on Task 3 is running dry when nobody is prompting you, and the only cure is producing a lot of unscripted speech. You talk, the AI answers what you actually said, and you get speaking feedback while the conversation is still going. The free tier gives basic feedback; the detailed pronunciation and grammar breakdown is on Premium.
Start free with three sessions a day and no card, each free session capped at five minutes and four turns. Paid sessions run ten minutes with unlimited turns. Limits are on the pricing page, and how it works walks through a full session.
We are a speaking improvement tool. We are not an exam preparation provider and we are not affiliated with Paragon Testing Enterprises, Prometric, IRCC or any exam board. For test format, rules, booking and official practice material, go to the exam body directly. Use SpeakShark to make your spoken English stronger, and use official material to learn the test.
- Site: speakshark.com
- Support: speaksharksupport@gmail.com
- Free tier: 3 AI conversation sessions per day, no card
Sources
- CELPIP test format: the eight speaking tasks in order, the fifteen minute section, and no separate speaking session
- Paragon 2024 CELPIP Data Report: responses assessed afterwards by at least three trained and certified raters
- CELPIP free resources: the free Speaking Pro and General Overview study packs, which carry the per task timing table, the rated sample analysis, the Task 6 description, and the statement that Task 4 uses the same illustration as Task 3
- Pearson PTE Academic test format: the Describe Image task, the nearest equivalent anywhere else, with the stimulus listed as a graph, picture, map, chart or table
FAQ
What is CELPIP Speaking Task 3?
How is Task 4 connected to Task 3?
What do people get wrong in Describing a Scene?
Should I use present continuous for Task 3?
Can I practise with real CELPIP images?
Do I lose marks for describing something wrongly?
Curious how it works? Explore SpeakShark's features or see plans and pricing.
Keep reading
CELPIP Speaking in 2026: All 8 Tasks, and the Real Bar
CELPIP Speaking is 15 minutes and 8 tasks with no human interviewer. Add up Paragon's published band table and most candidates land below band 9. Here is each task.
CELPIP vs IELTS for Canada PR: How to Choose in 2026
IRCC accepts both. The real difference is that IELTS Speaking has a human examiner and CELPIP does not. Which one suits you, and what nobody can tell you.
7 Best Duolingo Alternatives for Speaking (2026)
Duolingo publishes no prices on its own website, and real conversation is locked to its top tier. We checked 7 alternatives. SpeakShark is our Editor's Pick.