"Real", "official", "accurate" and "certified" are marketing words. Any website can print all four, and thousands do. What actually separates a measurement instrument from a scored quiz is four technical properties, none of which are hard to check once you know they exist: the norms, the scale, the reliability and the ceiling.
This page explains each one in plain terms, shows what a trustworthy result page looks like, and sets out a screening routine that takes about three minutes to run on any assessment you are considering. The aim is not scepticism for its own sake - good instruments exist and they are genuinely informative. The aim is being able to tell which is which.
What "Official" Can and Cannot Mean
There is no governing body that certifies intelligence assessments for the public, and no register you can consult. When a site calls itself official it is either referring to a relationship with a specific organisation - a national high-IQ society, a publisher - or it is saying nothing at all.
The instruments that professionals treat as authoritative earn that standing through publication rather than proclamation. The Wechsler scales and the Stanford-Binet have technical manuals documenting how they were built, who they were normed on, and how they perform under repeat administration. That documentation is the credential. Anything claiming authority without it has skipped the part that costs money.
This matters practically. A supervised assessment carried out by a qualified psychologist produces a report that institutions will accept. A home session, however well built, produces an estimate for your own information. Both can be accurate; only one is evidence. The boundary shows up clearly in admission requirements such as those for the Mensa IQ test, where unsupervised results are simply not considered regardless of the figure.
Norms: The Part Almost Nobody Checks
A standardised score is not a measure of how many questions you answered correctly. It is a statement about where your raw performance sits within a reference population. Change the reference population and the same performance produces a different number - with no error anywhere in the arithmetic.
So the quality of an assessment rests almost entirely on the sample its norms came from. A properly normed instrument was administered to a large group selected to reflect the wider population by age, sex, education and region, with separate conversion tables for each age band. That work is slow and expensive, which is exactly why it gets skipped.
The cheap substitute is norming on site traffic: whoever happened to take the test last, treated as though they were the country. People who seek out cognitive assessments are not a random slice of anyone, and traffic-based norms drift constantly as the audience changes. Look for these disclosures before trusting a figure:
- Sample size and date - a stated number and year, not "thousands of users".
- Sample composition - how the group was selected, and on what basis it represents anyone.
- Age banding - separate norms per age range, because reasoning performance shifts across the lifespan.
- Re-norming history - populations move over decades, and a set of tables from 1998 has aged.
- Percentile alongside the score - the position is the real information; the number is a translation of it.
An assessment that publishes none of this is asking to be trusted on the strength of its design. That is not an unreasonable request for a free product, but it should change how much weight the result carries.
Scale: Why 148 and 130 Can Be the Same Result
Standard deviation is the second thing to check and the easiest to miss, because the score arrives without it. Every mainstream instrument sets the median at 100, but they differ in spread. On a scale with a standard deviation of 15, one person in fifty reaches 130. On a scale with a standard deviation of 24, that same one-in-fifty position is 148.
| Position in population | Score on SD 15 | Score on SD 16 | Score on SD 24 |
|---|---|---|---|
| Median (50th percentile) | 100 | 100 | 100 |
| Top 16 per cent | 115 | 116 | 124 |
| Top 2 per cent | 130 | 132 | 148 |
| Top 0.1 per cent | 146 | 149 | 174 |
Every row describes one group of people. This is why comparing figures between assessments is meaningless unless both scales are named, and why a site that reports a score without stating its standard deviation has withheld the information needed to interpret it. It is also, less charitably, why some products favour wide scales: 148 reads better in a screenshot than 130.
Reliability and the Margin You Never See
No measurement is exact, and cognitive measurement is noisier than most. Reliability describes how consistently an instrument produces the same result for the same person - typically reported as a correlation between two administrations. The established clinical batteries report figures around 0.95 for full-scale scores, which is high, and still leaves a confidence band of several points either side.
That band is not a technicality. A result of 121 with a 90 per cent confidence interval of 116 to 126 means the underlying figure is very probably somewhere in that range. Someone who scores 121 and someone who scores 125 have not been distinguished by the assessment. A responsible result page says so; a promotional one prints a single number in a large font and lets the reader assume precision that was never there.
Sources of variation are mundane and additive: sleep, illness, caffeine, screen size, an unfamiliar interface, noise, nerves. None are exotic and together they easily account for several points. Anyone who wants a defensible personal figure should sit twice on separate days and treat the pair as a range rather than picking the flattering one - the practical setup for that is covered in the material on sitting a free IQ test online under controlled conditions.
Ceilings, Floors and Inflated Numbers
Every item set has a maximum. If the hardest question is not hard enough, everyone strong clusters at the top and the instrument stops discriminating between them - a ceiling effect. A short screener might top out around 130, which means it can tell you that you are in the top few per cent and cannot tell you anything beyond that. Reporting 145 from such a set is an extrapolation, not a measurement.
The same problem exists at the bottom. A set with too few easy items cannot separate anyone below average, and it will report scores in that region with false confidence. Well-documented instruments state their effective range. Products that quietly report anything from 60 to 160 from thirty questions are producing numbers rather than measurements.
Score inflation is the commercial version of the same failure. A test that returns pleasing figures to nearly everyone gets shared, recommended and revisited. A test that tells two thirds of its visitors they are average does not. The incentive is obvious, and the easiest check is external: search for reported scores. If nobody anywhere reports a below-average result, the scale has been shifted. A well-built iq test will return a distribution that looks like the population, with plenty of results between 85 and 115, because that is where most people are.
A Three-Minute Screening Routine
Before investing forty minutes in any assessment, run through this. It is quicker than the test itself.
- 1
Find the methodology page
Look for norming sample, scale, and reliability. If no such page exists, you have your answer without going further.
- 2
Check for a timer and a difficulty curve
Speeded administration and items that genuinely climb are both required. An untimed flat set cannot be normed meaningfully.
- 3
Find out when payment appears
A stated price before you start is fair dealing. A checkout page revealed after forty questions is a conversion tactic built on sunk cost.
- 4
Look at reported results elsewhere
Forum threads reveal inflation instantly. A credible instrument produces plenty of ordinary scores.
- 5
Confirm the result page shows a range
Percentile plus confidence interval signals an instrument. A bare number in a large font signals a product.
What a Careful Result Is Worth
All of the above is worth the effort because a well-measured figure does tell you something. Reasoning ability predicts performance on tasks that require holding several constraints in mind at once, and knowing roughly where you sit is useful for choosing how to study, how to structure difficult work, and where to expect friction.
What it does not predict is outcome. The research consistently shows measured ability explaining a modest share of variation in achievement, with persistence, circumstance, health and opportunity accounting for far more. A well-normed score describes one capacity on one morning against one reference group - a genuinely useful fact and a much narrower one than the number implies.
Held that way, the figure becomes a tool rather than a verdict. That is the same disciplined approach to numbers and probability that runs through the wider material published by Virgin Games, and it applies whether you are reading a percentile, an average or an odds line. For anyone assessing a child, the interpretive care matters even more, as the material on the IQ test for kids sets out - and understanding what each item family is built to measure, covered in the breakdown of IQ test questions, makes any score far easier to read honestly.
More on IQ Testing
Continue with the rest of the material on assessment, scoring and question formats:

