Mecca Bingo visitors who have ever compared two numbers and assumed the bigger one wins will recognise the trap that IQ scores set. A figure of 148 on one scale and 130 on another describe the same performance, while 130 on two different papers can describe genuinely different people. The number on its own carries almost no information. What carries information is the scale it sits on, the norm group behind it and the error band around it.
This page unpacks the scoring machinery: why the average is fixed at 100, why the spread differs between instruments, what a percentile really states, how much a single figure can move on retest, and why scores drift upward across generations. It is the part of testing that almost every result page skips, and the part that decides whether a number means anything at all.
- Scores are relative positions in a distribution, never counts of correct answers.
- The mean is fixed at 100 by construction, on every mainstream scale.
- Standard deviation varies: 15 on Wechsler, 16 on Stanford-Binet, 24 on Cattell.
- Percentiles are comparable across scales; raw score labels are not.
- Every responsible report gives a confidence interval, usually spanning several points.
- Norms age, which is why instruments are restandardised every decade or two.
Why the Centre Sits at 100
The mean of 100 is not a discovery about human beings. It is a decision. When an instrument is standardised, it is given to a large sample chosen to mirror the population on age, sex, region and education. Whatever that sample's average raw performance turns out to be gets labelled 100. Every subsequent score describes distance from that sample average, translated into a common unit.
Two consequences follow immediately. First, a score has no meaning outside its norm group; a figure normed on British adults says nothing precise about a different population. Second, the mean cannot drift within a single standardisation, because it is nailed down by definition. If the population genuinely improves, the raw performance needed to score 100 rises, and the score itself stays put. This is exactly what makes cross-generation comparison so awkward.
Standard Deviation and Why Scales Disagree
The second decision is how wide to make the ruler. Standard deviation sets how many points correspond to one unit of spread in the population. Wechsler instruments use 15. The Stanford-Binet fifth edition uses 16. The Cattell scales used by British Mensa use 24. All three describe an identical distribution; they simply print different numbers on the side of it.
Work through the arithmetic and the confusion dissolves. Two standard deviations above the mean is the 98th percentile on any of them. On a scale with SD 15 that prints as 130. With SD 16 it prints as 132. With SD 24 it prints as 148. Someone quoting 148 has not outperformed someone quoting 130 by eighteen points; the two figures are the same position expressed in different currencies. This single point causes more misplaced pride and more unnecessary disappointment than anything else in the field, and it is central to how the Mensa IQ test qualifying thresholds are set.
| Percentile | SD 15 Scale | SD 16 Scale | SD 24 Scale | Roughly How Common |
|---|---|---|---|---|
| 2nd | 70 | 68 | 52 | 1 in 50 |
| 16th | 85 | 84 | 76 | 1 in 6 |
| 50th | 100 | 100 | 100 | 1 in 2 |
| 84th | 115 | 116 | 124 | 1 in 6 |
| 98th | 130 | 132 | 148 | 1 in 50 |
| 99.9th | 145 | 148 | 172 | 1 in 1,000 |
Ranges and Percentiles
Because the distribution is bell-shaped, points are not evenly valuable. The stretch from 95 to 105 contains roughly a quarter of the population. The stretch from 130 to 140 contains well under two per cent. Ten points near the centre moves you past a crowd; ten points out in the tail moves you past comparatively few people, because there are few people left to pass.
This is why percentile reporting is the more honest format and why any decent result screen shows it. A percentile states directly what proportion of the norm group you outperformed. It is scale-independent, so it travels between instruments without translation, and it resists the illusion that the gap between 100 and 110 is the same size as the gap between 140 and 150.
Reading a Percentile Honestly
A percentile describes a position on one occasion against one reference sample. It is not a permanent property. Take the same paper on a different day, in different conditions, and the position shifts. Take a different instrument and it shifts again, because instruments sample slightly different abilities. The correct mental model is a range you occupy rather than a coordinate you own, and this holds whether the result came from a clinic or from a free IQ test online.
Confidence Intervals and Measurement Error
No psychological measurement is exact. Professional reports therefore give a composite alongside a confidence interval, commonly at the 95 per cent level, which typically spans somewhere around eight to ten points on an SD 15 scale. A report reading 118 with an interval of 112 to 124 is stating that the best estimate is 118 and that the true value very probably sits inside that band.
Two practical rules follow. Differences smaller than the interval should not be treated as real differences; a 121 and a 117 from two sittings are the same result. And unsupervised online figures deserve wider bands than clinical ones, because they add uncontrolled conditions, self-selected norm samples and unlimited retakes to the ordinary measurement noise. A responsible free paper says so; most do not.
Anyone about to sit something that matters can reduce the noise on their side by taking one full-length timed iq test under realistic conditions first, so the eventual result reflects reasoning rather than unfamiliarity with the format.
Ratio Scores and Deviation Scores
The original quotient was literal. Mental age was divided by chronological age and multiplied by a hundred, so a ten-year-old performing like a twelve-year-old scored 120. The method worked passably for children and collapsed for adults, since mental age stops climbing while chronological age does not, producing the absurd result that everyone declines steadily after adolescence.
Modern instruments abandoned that arithmetic in favour of deviation scoring: compare the taker with same-age peers and express the distance in standard deviations. The word quotient survives out of habit, but nothing is divided any more. The shift matters most for children, where every score is anchored to a narrow age band, a point covered in detail on the page about an IQ test for kids.
Ageing Norms and the Flynn Effect
Raw performance on standardised instruments rose substantially through the twentieth century, at roughly three points per decade across many countries. Because the mean is pinned to 100 at each standardisation, that improvement is invisible until you score a modern taker against an old norm table, at which point the score comes out inflated.
Practically, this means the edition and standardisation year of an instrument are part of the result. A figure produced against 1970s norms is not comparable with one produced against current norms, and organisations that accept prior evidence check the edition for exactly this reason. The gains have slowed or reversed in some countries in recent decades, which is an active area of debate rather than a settled matter.
Index Scores Beneath the Composite
A full battery produces several index scores before it produces one composite. Typical indices cover verbal comprehension, perceptual or fluid reasoning, working memory and processing speed. The composite is a weighted summary of those, and summarising has a cost: two people with the same composite can have completely different profiles underneath it.
Where the indices scatter widely, the composite becomes less informative and psychologists say so explicitly in the report. Someone strong in verbal reasoning and weak in processing speed is described by neither the average nor either extreme. This is also why a browser-based paper that samples one narrow ability cannot honestly report a composite at all, and why understanding the IQ test questions on any given paper tells you which slice of ability the figure actually reflects.
Reading a Score Report Without Overreaching
Four questions extract most of the available meaning from any result. Which instrument and edition produced it. Which norm sample it was compared against. What the confidence interval is. And what the percentile is, since that is the only figure that travels between scales without translation.
Answer those and the number becomes usable. Skip them and you are left with a bare figure that invites exactly the comparisons it cannot support. Scores describe performance on a defined set of reasoning problems, on one day, against one group. That is genuinely useful information, particularly where it flags an unusual gap between abilities. It is not a summary of a person, and no serious practitioner has ever claimed it was.

