From Paragon's results page

How CELPIP is scored: raw scores to levels, the four rating categories, and equating

A CELPIP report shows four levels and no raw scores, and most explanations of how the levels are reached are guesses. Paragon publishes the process itself, in more detail than most candidates ever read: example charts from raw score to level, the rule for blank answers, the categories its human raters score against, how many raters read each response, and why the number of questions you got right is not what you are told. This page reproduces that, and nothing else.

The levels, as Paragon describes them

Each of the four components is reported as a CELPIP level. The level is the Canadian Language Benchmark; Paragon also maps each to a CEFR level and gives it a one-line descriptor.

CELPIP levelCLBCEFRParagon's descriptor
1212C2Expert proficiency in high-stakes social, educational, or workplace contexts
1111C1Advanced proficiency in high-stakes social, educational, or workplace contexts
1010C1Highly effective proficiency in some high-stakes social, educational, or workplace contexts
99B2Effective proficiency in some high-stakes social, educational, or workplace contexts
88B2Good proficiency in more demanding social, educational, or workplace contexts
77B2Adequate proficiency in somewhat demanding social, educational, or workplace contexts
66B1Developing proficiency everyday social, educational, or workplace contexts
55B1Acquiring proficiency in everyday social, educational, or workplace contexts
44A2Adequate proficiency for daily life activities
33A2Some proficiency in limited contexts
21, 2A1Limited ability in contexts related to immediate needs

Source: Paragon, CELPIP Results, “Understanding Your CELPIP Scores”, checked 6 September 2026. Paragon also lists 1 and 0 as “insufficient information to assess” and NA for a component not administered; its raw-score charts below label the lowest band M. The shaded rows are the levels the Express Entry rules turn on.

Listening and Reading: right or wrong, by computer, then equated

Paragon's statement is short. All Listening and Reading questions are multiple choice or a similar format; every answer is scored as correct or incorrect; questions left blank are scored as incorrect; all scoring is by computer. Each component has 38 scored questions, and Paragon publishes an example chart of how that raw score corresponds to a level.

Listening — approximate score out of 38 per level

CELPIP levelScore /38
10–1235–38
933–35
830–33
727–31
622–28
517–23
411–18
37–12
M0–7

Reading — approximate score out of 38 per level

CELPIP levelScore /38
10–1233–38
931–33
828–31
724–28
619–25
515–20
410–16
38–11
M0–7

Source: Paragon, “Approximate Scores and CELPIP Levels” example charts. Paragon's disclaimer, reproduced: the charts show how scores approximately correspond to levels; since questions may have different levels of difficulty and may therefore be equated differently, the raw score required for a certain level may vary slightly from one test to another.

Two things in those charts are worth reading twice. The ranges overlap, by design: 33 in Listening appears in both the 8 and the 9 row, because the same raw score maps to different levels on forms of different difficulty. And the top is wide: Paragon prints 35 to 38 as one band for levels 10 to 12 in Listening and 33 to 38 for Reading, so the difference between a 10 and a 12 in those components comes down to a handful of questions and the form's equating.

The number you see is never the raw score. Paragon's reason: forms differ slightly in difficulty, so 30 right does not mean the same thing on two forms; raw scores are transformed into scaled scores, and scaled scores into levels by transformation rules set by language experts in a standard-setting exercise. Paragon reports an average Cronbach's alpha of 0.88 for both components, against a threshold of 0.80 it describes as excellent.

Writing and Speaking: four categories, five levels each, several raters

Paragon's performance standards name four categories per component and the factors under each. Each category is divided into five performance levels with descriptors, and a rater assigns a level by finding tangible evidence in the response that matches a descriptor.

Speaking

Content/Coherence
Number of ideas · quality of ideas · organization of ideas · examples and supporting details
Vocabulary
Word choice · precision and accuracy · range of words and phrases · suitable use of words and phrases
Listenability
Rhythm, pronunciation, and intonation · pauses, interjections, and self-correction · grammar and sentence structure · variety of sentence structure
Task Fulfillment
Relevance · completeness · tone · length

Writing

Content/Coherence
Number of ideas · quality of ideas · organization of ideas · examples and supporting details
Vocabulary
Word choice · suitable use of words and phrases · range of words and phrases · precision and accuracy
Readability
Format and paragraphing · connectors and transitions · grammar and sentence structure · spelling and punctuation
Task Fulfillment
Relevance · completeness · tone · word count

Source: Paragon, CELPIP Results, “Performance Standards”, reproduced as published.

The lists repay a slow read. Three of the four categories are the same for both components; the fourth is where the medium lives, Listenability for speech and Readability for text, and it is the one that carries grammar and sentence structure in both. Task Fulfillment ends in length for Speaking and word count for Writing, which is why a fluent answer that stops early, or an essay well outside the word range, loses on a category that has nothing to do with English.

How a response becomes a level

  1. An online system assigns tests to raters at random; the candidate stays anonymous throughout.
  2. A candidate's whole performance in a component, all tasks, goes to multiple raters: three to five for Speaking, four to six for Writing. They work independently and do not see each other's ratings.
  3. Each rater assigns a level in each of the four categories.
  4. The ratings are inspected for agreement. If they disagree, a benchmark rater, an experienced rater with a record of consistent accuracy, is assigned automatically and sees none of the earlier ratings.
  5. The dimensional ratings produce a component score, which transformation rules from standard setting turn into a CELPIP level.

Paragon also publishes who the raters are: native speakers of English or non-native speakers at CLB 11 or 12, with at least an undergraduate degree, an ESL teaching certification recognised by TESL Canada or graduate training in language education or linguistics, at least three years of relevant experience, and residence in Canada at the time of scoring. They are certified after initial training, receive ongoing feedback on their agreement with other raters, and are analysed monthly; a rater who does not improve to Paragon's standard has the contract ended.

Forms, unscored questions and equating

Paragon administers many test forms, even to candidates sitting at the same time, so that no one can hold the questions in advance. Every test also carries some new items being pre-tested: they look like scored items, are not used in the score, and are not identified, because Paragon needs candidates to try equally hard on all of them to judge the new items' quality. Only items that perform well become scored items later.

Forms are built to the same content and difficulty guidelines, but Paragon says small differences remain and that it would be unfair not to correct for them. Score equating is that correction: its own example is two candidates each answering 30 questions correctly, one on a relatively easy form and one on a harder one, whose reported scores are adjusted so that they reflect proficiency rather than the questions received.

The practical reading of all this: the raw-score charts above are the shape of the scale, not a pass mark to count towards; a blank answer is a wrong answer; and the questions that felt oddly easy or hard may have been the ones that did not count.

What a re-evaluation can change

Any component can be sent for re-evaluation within six months of the test date, once per component, for a fee paid at the time of the request and refunded for a component whose level rises; the request is final once submitted and answered in about one to two weeks. Paragon adds the sentence that matters: requesting a re-evaluation of the Listening and Reading components is unlikely to result in a change, as they are computer rated. The components with human judgement in them, Writing and Speaking, are the only ones where a second look can differ from the first.

Deadlines, fees and the rest of the booking rules are on the CELPIP registration page.

Practise against the same four categories

The practice reports on this site score Writing on Content and Coherence, Vocabulary, Readability and Task Fulfillment, and Speaking on Content and Coherence, Vocabulary, Listenability and Task Fulfillment: Paragon's categories, under Paragon's names, so that a weak category on a practice report is the same weak category a rater would find. Listening and Reading are marked right or wrong and reported as a level, the way the real report does it.

Questions candidates ask

How many questions do I need right for CELPIP 7 in Listening?

Paragon's example chart puts CELPIP 7 at roughly 27 to 31 of 38 in Listening and 24 to 28 of 38 in Reading, and says in the same breath that the raw score for a level may vary slightly from one test to another because questions differ in difficulty and are equated differently. Treat the chart as a range, not a threshold, which is how Paragon labels it.

Are wrong answers penalised?

No. Paragon's page says every Reading and Listening answer is scored as correct or incorrect and that questions left blank are scored as incorrect. A guess therefore cannot cost anything a blank would not already cost.

Why does my report not show a score out of 38?

Paragon does not report raw scores. Its explanation: test forms differ slightly in difficulty, a raw score of 30 does not mean the same thing on two forms, so raw scores are transformed into scaled scores that can be compared, and the scaled score is then transformed into a CELPIP level by rules set in a standard-setting exercise.

Who marks Writing and Speaking?

Human raters. Paragon's rating procedure: tests are randomly assigned by an online system with the candidate anonymous; each Speaking performance is rated by three to five raters and each Writing performance by four to six; raters work independently and do not see each other's ratings; if the ratings disagree, a benchmark rater who has not seen them is assigned automatically.

What are raters looking for?

Four categories each, listed on this page as Paragon publishes them: for Speaking, Content/Coherence, Vocabulary, Listenability and Task Fulfillment; for Writing, Content/Coherence, Vocabulary, Readability and Task Fulfillment. Each is divided into five performance levels with descriptors, and raters assign a level by matching tangible evidence in the response to the descriptors.

Are some questions unscored?

Yes. Paragon pre-tests new items by including some in every test; they look like scored items, are not used in the score, and are not identified, so that candidates try equally hard on all of them. Its format page says Listening and Reading contain unscored items.

Is one version of the test easier?

Paragon says many forms exist, built to the same content and difficulty guidelines, and that small differences remain; score equating corrects the final score for them so that 30 right on an easier form and 30 right on a harder one are not treated alike.

Is a re-evaluation worth requesting?

Paragon's own guidance: a re-evaluation can be requested for any component within six months, once per component, for a fee refunded if the level rises; but Listening and Reading are computer-rated and a re-evaluation there is unlikely to change anything. Where a change is possible is Writing and Speaking, the human-rated components.

Sources

  • Paragon, CELPIP Results — the level table, the approximate score charts, the performance standards, the rating procedure, rater qualifications and training, the scoring FAQs, and the re-evaluation policy.
  • Paragon, CELPIP Test Format & Scoring — component timings and the presence of unscored items.

Checked 6 September 2026. The process is Paragon's and can change; the charts are its example charts, labelled approximate by Paragon itself.

Related: Listening guide · Reading guide · Writing guide · Speaking guide · CLB calculator · CELPIP vs IELTS vs PTE Core · Registration, fee and test day · CELPIP for the PGWP · CELPIP for citizenship.