Methodology · norms beta-2026.09
How scoring works.
What the test measures, how your score is calculated, and where its limits are.
What's on the test
30 questions in 4 sections, with a 25-minute limit for the whole test. The mix leans toward non-verbal reasoning, which depends less on language and schooling:
| Section | Questions | What it asks |
|---|---|---|
| Pattern reasoning | 12 | Complete a 3×3 grid of shapes by finding the rules across each row. |
| Number series | 7 | Find the next number in a sequence. |
| Spatial rotation | 6 | Spot which shape is a rotation, not a mirror image. |
| Verbal reasoning | 5 | Analogies and odd-one-out word problems. |
Every puzzle is original, and most are generated
Pattern, number and spatial questions are built from rules rather than picked from a fixed list. A pattern question, for example, combines rules like "the number of shapes goes up by one" or "each row has one of each shading". Each attempt gets its own questions, so answer keys can't leak and retests stay fair.
The eight answer options for pattern questions are built so that each wrong option differs from the right one in a balanced way. That closes a shortcut some tests allow, where you can pick the option that looks most like the others without solving the puzzle.
How your score is calculated
Each question has a difficulty rating. We use a one-parameter item response model (the Rasch model): the chance you get a question right depends on the gap between your ability and the question's difficulty. Your score is the ability level that best explains your pattern of right and wrong answers.
We calculate an expected a posteriori estimate, which starts from the population average and moves as the evidence comes in. That keeps estimates sensible at the extremes, where a raw count can't tell you much. The estimate is then placed on the familiar IQ scale (mean 100, standard deviation 15).
The range next to your score
The same calculation tells us how uncertain the estimate is. We show a 90% range: your true score falls inside it nine times out of ten. With 30 questions the range is usually around ±12 points, and wider for very high or low scores, where there are fewer questions at your level.
Where the norms stand
Question difficulties are provisional (beta-2026.09). They come from how the rules combine, and we're checking them against real results. A norming study with a demographically balanced adult sample is under way. When it's finished, we'll publish a technical report with reliability, sample details and correlations with established reasoning tests. Results will say which norms version they used.
Limits you should know
- This is a self-assessment for adults, taken unsupervised. It is not a clinical or diagnostic test and can't be used for educational or employment decisions.
- Scores near the top of the range are less precise. The highest possible estimate is about 145.
- Tiredness, distraction and practice all move scores by a few points. Compare trends across attempts, not single results.
- We are not affiliated with Mensa or any test publisher.