Online cognitive benchmarks are everywhere, and almost nobody explains what the numbers mean. A reaction time of 240 milliseconds, a digit span of nine, a chimp test level of twelve — each of those is a real measurement of something specific, and each of them is also much less comparable between people and between websites than it looks. This is a guide to reading your own scores honestly: what each standard test measures, what moves it, what does not, and why the same person can score very differently on two sites running what looks like the same test.
Almost every cognitive benchmark site runs the same handful of tasks, because they are the ones with the longest history in psychology and the clearest interpretation.
Reaction time. Wait for a signal, respond as fast as you can. No decision involved, so it measures the shortest possible path from stimulus to movement. Take it on the Reaction Time Test.
Number memory. A number appears, disappears, you type it back, and each success adds a digit. This is digit span, a measure of short-term memory capacity that has been in use since the nineteenth century. Take it on Number Memory.
Verbal memory. Words arrive one at a time and you say whether you have seen each before. The list keeps growing, so the load rises continuously. Take it on Verbal Memory.
Sequence memory. A pattern flashes and you reproduce it, one step longer each round. This is visual sequential memory — order plus identity. Take it on Color Sequence or Simon.
Visual memory. Cells light up on a grid and you click the ones that were lit. Position memory rather than sequence memory, and a different capacity. Take it on Visual Grid Memory.
Chimp test. Numbers appear scattered, hide the moment you touch the first, and you tap the rest in order. Simultaneous spatial capture. Take it on the Chimp Test.
Aim trainers and typing tests usually sit alongside these. Both are real measurements, but they are motor-skill measures rather than cognitive ones — the Aim Trainer measures pointing speed and the Typing Test measures a trained motor skill, and both improve far more with practice than any of the six above.
This is the part benchmark sites rarely say out loud. Three things move these numbers enough to make cross-site and cross-device comparison close to meaningless.
Hardware, on anything timed
A 60 Hz display adds up to 16 milliseconds of latency before a signal even reaches your eye. Input lag from a mouse or a touchscreen adds more, and browser timing adds more again. Together those can move a reaction time result by a large fraction of the range that separates fast people from slow ones. Touchscreens are consistently the worst, which is why the same person records slower times on a phone than on a desktop with a wired mouse.
Presentation speed, on anything remembered
Digit span depends heavily on how fast the digits are shown and how long you get before recall. A site that shows a number for three seconds and one that shows it for one second are measuring different things, and neither is wrong. This is why a span of nine on one site and seven on another tells you nothing about either site.
Scoring rules that are never stated
Some sites end a run on the first error and some allow lives. Some average several trials and some report your best. A best-of-ten reaction time is systematically faster than an average-of-ten for the same person, and the difference is larger than most real differences between people.
The only number worth tracking is your own, measured the same way, on the same device, over time. That figure is real and it moves.
Some things move cognitive benchmarks a great deal, and some things that people expect to move them do not.
Sleep, more than anything else
Sleep loss slows reaction time reliably and substantially, and it is one of the most sensitive markers of fatigue there is. It degrades inhibition — measurable as false alarms on Go/No-Go — before people notice any subjective tiredness. A bad benchmark day is more often a sleep signal than a cognitive one.
Caffeine and time of day
Stimulants shorten reaction times in the short term; this is one of the more robust findings in the area. Reaction time also follows a daily rhythm large enough to mislead you, so comparing a morning run with an evening one is comparing two different states.
Strategy, on every memory test
This is the big one, and it is why memory benchmarks are much less a measure of raw capacity than they appear. Chunking — grouping digits into threes and fours — routinely adds several digits to a span. Verbal rehearsal, saying the material quietly to yourself, holds material meaningfully longer than a silent visual impression. Both are learnable in minutes. The Number Chunking Tool shows you how to split a number for memory.
Practice, but only on the thing practised
You will get better at any of these tests with practice, quickly and reliably. Whether that improvement extends to anything else is genuinely contested, and the larger, better-controlled studies have mostly failed to find it. We say so on every game page rather than implying otherwise.
Use one device, always. Pick the device you will keep using and never compare across.
Average, do not cherry-pick. Five clean trials averaged is worth far more than a personal best, which is dominated by luck.
Test at a similar time of day. The daily rhythm is large enough to swamp a real change.
Warm up with one throwaway run. Cold first runs are consistently the worst of a session, and treating one as your score will mislead you.
Record the conditions when they change. A run after a bad night is not a data point about your ability.
The Score Tracker does the arithmetic: it keeps a log, shows a five-run rolling average and computes a least-squares trend, so 'am I improving' becomes a calculated answer rather than an impression.
None of these tests measures intelligence, and none of them is a clinical assessment. Digit span and Stroop variants do appear in standardised batteries, but a clinical administration is controlled in ways a web version is not — the timing, the instructions, the environment and the normative sample are all standardised, and none of that applies to a browser tab.
A score here does not diagnose anything, does not rule anything out, and should never be presented as evidence for or against a diagnosis. If you are genuinely worried about your memory or attention, a clinician is the right place to go and a benchmark result is not useful input.
What these tests are good for is being interesting, being measurable, and giving you a reason to come back tomorrow. That is enough.
They accurately measure what they measure, on your device, under your conditions. They are not accurate as a comparison between people, because hardware latency, presentation speed and scoring rules differ enough to swamp real differences. Track your own trend on one device.
Around 200-250 milliseconds is the commonly quoted range for visual reaction, but display and input latency move measured values by tens of milliseconds. Your own average across five clean trials is the only number worth comparing against.
Almost always because the presentation speed, the scoring rule or the hardware path is different. A digit shown for three seconds and one shown for one second produce different spans for the same person.
No. Some of the underlying tasks appear as components of intelligence batteries, but a single task in a browser is not an IQ measure and should not be read as one.
Yes, quickly, on the tests themselves. Memory scores respond most to strategy — chunking and verbal rehearsal — and reaction scores respond most to sleep and alertness. Whether any of it transfers to everyday ability is unproven.
Treat a cognitive benchmark the way you would treat a bathroom scale: useful for watching your own trend, useless for comparing yourself to someone standing on a different scale. Pick one device, average several runs, and watch the line over weeks rather than the number on any given day.