Celpip AI Try a scored task

Calibration

How our AI scoring works

A number is only worth as much as the method behind it. This page is the method: what the scoring was tested against, how many times, what counted as a hit, and the band where it stops working.

What it was measured against

Officially rated CELPIP sample responses — answers that already carry a level assigned by the exam’s own raters. Our scoring never sees that level; it is compared afterwards.

Speaking — 110 scored calls

22 officially rated samples, each scored 5 times over, from the published audio recordings rather than transcripts of them. The run covered 2 of the 8 Speaking tasks: Task 1 (Giving Advice) and Task 8 (Describing an Unusual Situation).

Writing — 108 scored calls

36 officially rated samples, each scored 3 times over, with anchors at every level in the set.

When

August 2026, on the model that was in production that month. The date travels with the numbers because a new model means a new measurement, not an inherited one.

Answers are scored several times over on purpose: the model is not deterministic, so a single call per sample would measure luck. Across five calls on the same Speaking answer the estimates differed by 0.489 of a band on average.

What counts as a hit

Within one band means the estimate is the official level, or one CLB level either side of it. Two levels away is a miss, and it is counted as one. We report this rather than exact matches because the question a candidate is actually asking is “am I roughly in the right place?” — but the looser definition is also the flattering one, so it is stated here rather than buried.

The numbers, both halves

Measured accuracy of the AI scoring, August 2026
What was measured Result Scored calls behind it
Speaking, official CLB 3–8, within one band 87% 70
Speaking, official CLB 9 and above, within one band 43% 40
Speaking, all levels together, within one band 71% 110
Writing, official CLB 3–8, within one band 95% 108
Writing, quote fidelity — a marked phrase really is in your answer 99% 108

The 43% is 17 of 40 calls, which is 42.5%. We publish it rounded up to 43% everywhere, and print the raw count here so that nobody meets the difference by surprise. Overscoring across the whole Speaking run was 7%; the rest of the error runs strict rather than generous.

Why nothing above CLB 9 is printed as a single number

Because of the 43%. At official CLB 9 and above the estimate misses the band more often than it hits it, and a bare number there would claim a precision we have measured ourselves not to have. So every report shows a range, the app’s own test suite enforces it on the screen, and this page and the score calculator do the same.

If you are already scoring 9 or 10, treat the estimate as a rough sanity check and nothing more.

The four dimensions

A report scores four dimensions separately, under the names the published CELPIP rubric uses:

  • Content & coherence
  • Vocabulary
  • Listenability on a spoken answer, Readability on a written one
  • Task fulfillment

Using the published names is not the same as reproducing the official weighting, and we make no claim about the second. What the four give you is a diagnosis you can act on: the weakest one is marked, and the two fix-first instructions attach to it.

Why quote fidelity is on the list

On a Writing answer the report marks the phrase that cost you and the phrase that would replace it. If the marker invented that phrase, the whole report would be worthless, so it is measured like everything else: 99% of marked phrases appear in the candidate’s own answer word for word.

What we have not measured

Writing above CLB 8. The 95% covers the core band. We have not run the same upper-band measurement for Writing, so there is no upper-band Writing figure on this site.

Six of the eight Speaking tasks. The Speaking run used Task 1 and Task 8. The other six are scored by the same rubric and the same model, but they were not part of this measurement and we are not going to imply otherwise.

What happens on test day. These are estimates from a practice tool, not official scores. An official CELPIP result takes 2 to 4 business days and comes from the test provider; a scored report here takes about five seconds and comes from an AI model.

The weak number travels with the strong one. That is the whole policy: publishing only the flattering half would make the exercise worthless, and there would be no reason to believe the flattering half either.