The anatomy of a question that tests understanding

Most examination questions can be answered correctly by a candidate who has memorised a result and never understood it. The difference between an item that measures comprehension and one that measures storage lies less in the question than in the wrong answers offered alongside it.

What a question actually measures

Every question makes an implicit claim about what it measures, and many of those claims are false. An item that names a principle in its stem and then rewards whoever recalls the matching formula is measuring retrieval of a stored result. On a mark sheet that is indistinguishable from understanding. The two come apart the moment the surface of the problem changes.

Chi, Feltovich and Glaser demonstrated the mechanism in 1981. Asked to sort physics problems into groups, novices sorted by appearance, putting inclined planes with inclined planes and springs with springs. Experienced physicists sorted by the principle required, filing a spring problem beside a pulley problem because both turned on conservation of energy. That result is a design brief. A question tests understanding to the extent that its surface offers no route to the answer and the candidate must first decide which principle is in play.

The wrong answers do most of the work

Most of the diagnostic power of a multiple-choice item sits in the options the candidate does not select. A stem can be elegant and still be worthless if two of its four options are visibly absurd, because the item then behaves as a two-option question wearing four. A distractor earns its place only when it is the precise endpoint of a specific error: a sign dropped, a boundary condition ignored, a rule carried one step outside the domain where it holds.

The item-writing literature, summarised most often through the rule taxonomy of Haladyna, Downing and Rodriguez, keeps arriving at the same conclusion: options nobody chooses carry no information. The corollary is that error patterns are data. When a large group of candidates converges on one particular wrong option, that option names their misconception, and names it more precisely than any score can. A well-built distractor set turns a scoring instrument into a diagnostic one.

The arithmetic of the guess

Negative marking sets an accuracy threshold below which answering costs more than abstaining, equal to the penalty divided by the sum of penalty and reward. In UPSC Prelims it works out at exactly 25%, which is the accuracy of a blind pick from four options, so guessing there is expected-value neutral by construction. NEET and JEE Main break even at 20%, which makes a blind pick from four options mildly positive.

That arithmetic changes what question design has to achieve. Where the scoring rule already pays for a coin flip, any item whose distractors can be discarded on appearance alone becomes strongly profitable to attempt without understanding, and the instrument stops measuring the subject and starts measuring examination craft. The remedy is not a heavier penalty, which mostly taxes cautious candidates. It is distractors that survive elimination, so that narrowing four options to two demands the same reasoning as answering outright.

What a long archive can show

The JEE Main record from 2002 to 2025 runs to 8,267 authentic questions, and 117 topics within it carry enough questions for a pattern to be described at all: 48 in Physics, 41 in Chemistry, 28 in Mathematics. Concentration inside that set is pronounced. Matrices and determinants account for 8.8% of the Mathematics questions in recent papers, 275 items across the 23 years covered. Kinematics in one and two dimensions is 6.6% of recent Physics, 180 questions. Coordination compounds is 8.7% of recent Chemistry, 171 questions.

Those proportions look like an instruction to drill the heaviest topics. They say something more useful. An idea set 275 times has been dressed 275 different ways, and repetition at that scale is what makes the distinction between understanding and recognition visible. A candidate whose accuracy holds steady as the presentation of a determinant problem changes has the idea. A candidate whose accuracy tracks familiarity with particular presentations holds a catalogue of shapes, and the catalogue fails on the version that has not been seen.

Difficulty is not discrimination

Question quality is often discussed as though it meant difficulty. It does not. An item almost nobody answers correctly may be deep, or merely ambiguous or mis-keyed. The property worth engineering is separation: whether success on the item aligns with command of the subject. An item that strong candidates fail and weak candidates pass is broken however demanding it feels, and an easy item that cleanly separates the two groups does more work than a hard one that does not.

Difficulty of the right kind is nonetheless productive. Robert Bjork's research on desirable difficulties established that conditions which slow performance during practice frequently improve later retention. Roediger and Karpicke's 2006 experiments on retrieval practice found that testing after study beat re-reading, with the advantage widening as the delay before the final test grew. Andrew Butler's later work found the benefit extends to transfer: material learned through retrieval was applied more successfully to unfamiliar problems. A good question is not only a measurement. It is an act of learning.

The explanation is part of the item

Sharp distractors carry a measured risk. Roediger and Marsh's 2005 experiments found that multiple-choice testing produced both benefits and costs, one cost being that plausible lures were sometimes later recalled as facts. Feedback removed most of it. A question whose wrong answers are all positions somebody could reason their way into therefore incurs an obligation: to say why each option fails, in the terms of the error it represents rather than by restating the correct answer more loudly.

The 1989 work of Chi and colleagues on self-explanation points the same way. Students who gained most from worked examples explained each step to themselves as they went, rather than reading the solutions through. An explanation that only narrates the right method leaves a learner nothing to explain. An explanation that names the reasoning behind each wrong option leaves a claim to test, which is the condition under which the next question of the same kind becomes answerable on principle.

The data behind this

The full set, with progress tracking and five agent perspectives per question, is in the JupiteX app — browse the exam catalogue or browse the Learn library.