Task 3 — Describing a Scene
The Celpip Plus team
Written and reviewed against the official CELPIP scoring criteria
What you'll learn: a repeatable visual scan pattern for describing a single picture, plus the layered vocabulary (location, action, inference) that keeps a 60-second description from running dry.
What Task 3 actually tests
Task 3 shows you a single picture — often a scene with several people or objects — and gives you 30 seconds to prepare and 60 seconds to describe it. There's no "story" required here the way there is in Task 2; the goal is a clear, organized description of what's visibly happening. The most common failure isn't lacking things to describe — it's describing them in a random order that jumps around the image, which hurts Content/Coherence even when every individual observation is accurate.
The scan pattern: foreground, background, inference
- Orient the listener — one sentence naming the overall setting.
- Foreground first — describe the most prominent people or objects and what they're doing.
- Background second — describe secondary details (weather, other people, objects in the distance).
- Add one inference — a reasonable guess about what's happening or why, based on the visual evidence.
- Close with a summary line that ties the scene together.
Scanning in a fixed order (foreground → background → inference) means you never have to decide "what do I mention next" in the moment — the structure decides for you.
Worked example
Prompt: A picture of a covered outdoor produce market with fruit and vegetable stalls.

This picture shows a covered outdoor produce market, the kind that sells fresh fruit and vegetables. In the foreground, there are several wooden crates stacked with produce — I can see grapes, apples, and what look like tomatoes — arranged in neat rows, with a few of the boxes labelled, so the vendors clearly sort their stock by type. Just behind the crates, an older man in a white shirt is standing at his stall, most likely the vendor keeping an eye on his goods, and there's a red cloth covering one of the tables over to the left. Further back, the stalls continue under long fabric awnings held up by green metal poles, and I can make out one or two people moving between them, though it doesn't look especially crowded. Given the calm atmosphere and the soft light, I'd guess it's fairly early in the morning, before the busier shopping hours. Overall, it looks like a relaxed, well-stocked local fruit and vegetable market.
Notice the description never jumps randomly — it moves systematically from the closest details outward, then adds a reasoned guess (season, based on weather and turnout) before closing.
Vocabulary layers to prepare in advance
- Location phrases: "in the foreground / background," "on the left/right," "in the middle of the picture," "just behind..."
- Action verbs: specific over generic — "handing," "arranging," "browsing," "pointing at," instead of just "doing."
- Inference language: "it looks like...", "this suggests...", "given [detail], I'd guess..."
Timing plan for 60 seconds
- 0–5s: Orient (setting)
- 5–35s: Foreground details
- 35–50s: Background details
- 50–60s: Inference + closing line
Common mistakes
- Describing the picture as a flat list ("There's a man. There's a woman. There's a stand.") without location or spatial language.
- Guessing wildly beyond what the image supports (inventing a backstory) instead of one grounded inference.
- Running out of things to say after 20 seconds because only the most obvious element was described — the scan pattern's background step exists specifically to prevent this.
Recap
Task 3 rewards an organized visual sweep, not a random list of observations. Scan foreground to background in a fixed order, layer in specific location and action vocabulary, add one grounded inference, and close with a summary line — this structure alone is usually enough to comfortably fill the full 60 seconds with relevant, well-sequenced content.