CELPIP Speaking Task 3 Template: Describing a Scene
A five-part scan pattern for describing any CELPIP Task 3 picture, with spatial phrases, a full model answer, and a fix for running dry at 40 seconds.
The Celpip Plus team
Written and reviewed against the official CELPIP scoring criteria
Task 3 puts a single picture on screen and gives you 30 seconds to prepare and 60 seconds to speak. The picture changes every time. What you do with it doesn't — which is why a fixed pattern helps here more than in almost any other task.
Use this as a shape to fill, not a script to memorise. Raters score Content/Coherence, Vocabulary, Listenability, and Task Fulfillment, and a rehearsed paragraph that ignores the actual image loses on all four.
The five-part structure
| Part | What goes in it | Seconds |
|---|---|---|
| Orient | One sentence naming the overall setting | 0–5 |
| Foreground | The closest, most prominent people and objects | 5–30 |
| Background | Secondary details — other people, weather, distance | 30–45 |
| Inference | One grounded guess about what is happening or why | 45–55 |
| Close | A summary line that ties the scene together | 55–60 |
The order is the point. Deciding what to mention next mid-answer is what produces long pauses; a fixed sweep from front to back makes that decision for you before you open your mouth.
Step 1 — Use your 30 seconds to sort, not to script
You cannot write out sentences in 30 seconds. You can label three zones. Look at the picture and silently name: the one thing closest to the camera, the one thing furthest back, and one detail that hints at time, place, or purpose.
That is your whole plan. Three anchors, in order.
Step 2 — Orient in a single sentence
Name the setting before any detail. The rater needs a frame to hang the rest on.
This picture shows [a type of place], the kind where [what usually happens there].
Applied:
This picture shows a covered outdoor market, the kind where people buy fresh produce.
Avoid opening with "In this picture I can see a man." That is a detail, not a frame, and it forces you to build outward from a random point.
Step 3 — Describe the foreground with spatial language
This is your longest stretch — roughly 25 seconds — so it needs more than a list of nouns. Every observation should carry a location and an action.
In the foreground, [person or object] is [action], and [second detail] is [action].
To the left of them, I can see [detail], while [something else] sits closer to [reference point].
The spatial phrases are what turn a list into a description. Keep these ready:
| Function | Phrases |
|---|---|
| Distance from viewer | in the foreground, in the middle of the picture, further back, in the background |
| Side to side | on the left, over to the right, on the far side, along the edge |
| Relative position | just behind them, next to, in front of, above, underneath, between |
Present continuous throughout. A picture is a frozen moment, so the tense that matches it is is/are + -ing: "a woman is handing over a bag," not "a woman handed over a bag." Slipping into past tense is one of the fastest ways to sound like you are telling a story instead of describing an image. Reach for specific verbs too — arranging, browsing, pointing at, leaning against, unloading — rather than doing or having.
Step 4 — Move outward to the background
Around the 30-second mark, deliberately push your eyes to the back of the image. This step exists to solve the most common Task 3 failure: running dry at 40 seconds because only the obvious central figure got described.
Further back, [what continues or repeats], and I can make out [smaller detail].
The [lighting / weather / space] looks [description], which gives the whole scene a [quality] feel.
Background material is almost always there — signage, other people, structures, sky, floor surface, what the space is made of. If you genuinely cannot find more, describe colours, materials, and how crowded or empty the place is.
Step 5 — Add one grounded inference
Around 45 seconds, shift from what I see to what it suggests. One inference, tied to visible evidence.
Given [the specific visual clue], I'd guess [reasonable conclusion].
It looks like [situation], probably because [evidence from the picture].
One is enough. Inventing a backstory for the people — names, relationships, what happened yesterday — moves you away from the task, which is description.
Step 6 — Close, don't trail off
Leave five seconds for a summary line so the answer ends rather than stops.
Overall, it seems like [one-phrase characterisation of the scene].
Worked walkthrough
Here is the whole process on one prompt, from a blank screen to a finished answer.
Prompt: Describe what you see in this picture to someone who cannot see it. (The picture: an outdoor community skating rink in a snowy park at dusk, with children lacing skates on a bench and a few people on the ice.)
1. Read it for who, what, and the ask.
You are describing a single frozen image to a listener who can't see it, so they can build the scene in their head — front to back, in order.
2. Plan — notes only.
In your 30 seconds of prep, sort the picture into three anchors — nearest thing, furthest thing, one telling detail — not sentences.
| Part | Your note |
|---|---|
| Orient | an outdoor community skating rink, flooded in a park each winter |
| Foreground | two kids lacing skates on a bench; a parent tightening a boot; boots and a red toque; a hockey stick on the boards; skaters circling |
| Background | a small warming hut with a string of lights; snow-covered pines; dusk sky |
| Inference | lights already glowing + deep-blue sky → late afternoon on a cold weekend |
| Close | a calm, cheerful winter evening at a neighbourhood rink |
3. Turn each note into a line.
Orient:
This picture shows an outdoor community skating rink, the kind a town floods in a park each winter.
Foreground:
In the foreground, two children are sitting on a wooden bench right at the edge, lacing up their skates, and a parent is kneeling in front of the smaller one, tightening a boot for her. Just to the left, a pile of boots and a red toque are lying on the bench, while a hockey stick is leaning against the boards nearby. Out on the ice, several people are gliding in a loose circle, and one man in a blue jacket is skating with his hands behind his back.
Background:
Further back, I can make out a small warming hut with a string of lights along its roof, and the whole rink is ringed by snow-covered pine trees fading into the dusk.
Inference:
Given how the lights are already glowing and the sky is turning that deep evening blue, I'd guess it is late afternoon on a cold weekend, when families come out once the day warms up a little.
Close:
Overall, it looks like a calm, cheerful winter evening at a neighbourhood rink.
4. Read it back.
This picture shows an outdoor community skating rink, the kind a town floods in a park each winter. In the foreground, two children are sitting on a wooden bench right at the edge, lacing up their skates, and a parent is kneeling in front of the smaller one, tightening a boot for her. Just to the left, a pile of boots and a red toque are lying on the bench, while a hockey stick is leaning against the boards nearby. Out on the ice, several people are gliding in a loose circle, and one man in a blue jacket is skating with his hands behind his back. Further back, I can make out a small warming hut with a string of lights along its roof, and the whole rink is ringed by snow-covered pine trees fading into the dusk. Given how the lights are already glowing and the sky is turning that deep evening blue, I'd guess it is late afternoon on a cold weekend, when families come out once the day warms up a little. Overall, it looks like a calm, cheerful winter evening at a neighbourhood rink.
5. Check before you stop speaking.
- The description sweeps in one direction — bench, floor, the ice, then the hut behind — and fills the full 60 seconds.
- The first sentence names the setting, not a person.
- Every foreground observation carries a location phrase, and the verbs stay in present continuous.
- The single inference points at a visible clue (the glowing lights and the darkening sky), and the answer closes on a summary line.
Full model answer
Prompt: Describe what you see in this picture to someone who cannot see it.

This picture shows a covered outdoor produce market, the kind where sellers lay out fruit and vegetables on open stalls. In the foreground, wooden crates are stacked several rows deep, filled with apples, green grapes and what look like tomatoes, and some of the boxes have a printed label along the front. Lower down, empty crates and a white plastic basket are sitting on the floor, while a red cloth is covering one of the tables over to the left. Just behind the fruit, an older man in a white shirt is standing at his stall, apparently keeping an eye on his goods. Further back, the market carries on under long fabric awnings held up by tall green poles, though everything at that distance is out of focus, so I can only make out one or two figures moving between the stalls. Given how soft the light is and how few shoppers I can pick out, I'd guess it is early in the morning, before the place fills up. Overall, it looks like a calm, well-stocked market that is only just getting started.
Why this scores well: it names the setting before anything else, then moves in one direction — crates, floor, red cloth, the vendor, and finally the awnings behind him — with a location phrase carried on almost every observation. It stays in present continuous throughout, and it handles the blurred background honestly: rather than inventing shoppers it cannot see, it says the distance is out of focus and describes only the shapes that are really there, which reads as accurate rather than thin. The single inference is tied to visible evidence — the soft light and the near-empty aisles — and the answer closes on a summary line instead of trailing off. The background and inference together fill the last twenty seconds, which is exactly what prevents an early finish.
Before you speak
- You have named three anchors: nearest thing, furthest thing, one telling detail
- Your first sentence names the setting, not a person
- Every foreground observation has a location phrase attached
- Verbs are present continuous and specific, not doing or having
- You have something reserved for the background — you are not spending it all up front
- Your inference points at a visible clue
- You know your closing line before you start talking
Common mistakes
- The flat list. "There's a man. There's a table. There's a bag." Every item is accurate and the answer still scores low, because nothing is positioned relative to anything else. Content/Coherence rewards arrangement, not inventory.
- Running out at 40 seconds. This nearly always means the obvious central figure was described and the rest of the frame was ignored. The background step is the cure — treat it as a required stage, not an optional extra.
- Drifting into past tense. "A woman was buying apples" describes a memory. The picture is happening now; keep is/are + -ing running the whole way through.
- Over-inferring. Deciding the two people are siblings who have argued is fiction, not description, and it costs Task Fulfillment. One grounded guess, then stop.
- Repeating one spatial phrase. If "on the left" appears five times, Vocabulary suffers. Rotate through just behind, further back, next to, along the edge, in the middle.
- Narrating your own process. "Let me see… what else is there…" burns seconds that could hold a real detail. Silence while you scan is better than filler out loud.