Measuring knowledge, not marks

An examination score records where a candidate stood on one morning, against one paper and one cohort. What follows is the argument for a knowledge measure built apart from any single examination, and the conditions it would have to satisfy before anyone should trust it.

A mark records a placement

An examination produces a number, and the number is then read as though it described a person. In most cases it describes a placement. Cut-offs move with the size of the cohort and with the difficulty of the paper set that year, and normalisation exists because raw marks from two sittings are not comparable. A candidate who lands in the same percentile two years running has demonstrated stable standing against a shifting population, which is a different claim from stable knowledge.

The difficulty appears when a mark is read backwards, as a diagnosis. A total compresses what a candidate knew, what they were never asked, how they spent a fixed number of minutes and how much risk they accepted when unsure, into one figure that cannot be taken apart again. Messick's account of validity applies here: validity belongs to the inference drawn from a score, not to the instrument, so a score can be reliable while the conclusion drawn from it is unfounded.

The scoring rule enters the score

Scoring rules sit inside the measurement rather than beside it. Where wrong answers are penalised, the break-even accuracy for guessing is the penalty divided by the sum of reward and penalty: 25% for UPSC Prelims, 20% for NEET and 20% for JEE Main. The UPSC figure is exactly the accuracy of a blind pick from four options, so guessing there is neutral in expectation, while under the NEET and JEE Main rules the same blind pick carries a mild positive expectation.

Two candidates with identical knowledge and different appetites for risk therefore separate on the mark sheet, and a knowledge measure that borrows a competitive scoring rule imports variance that has nothing to do with knowing anything.

What a corpus shows that one paper cannot

The JEE Main record makes a useful test case, long and verifiable. Across papers from 2002 to 2025, 8,267 authentic questions in all, the shape of the domain is stable enough to describe. Matrices and Determinants accounts for 8.8% of recent Mathematics, and 275 questions across the record. Kinematics in one and two dimensions runs at 6.6% of recent Physics, 180 questions. Coordination Compounds reaches 8.7% of recent Chemistry, 171 questions. Altogether 117 topics carry enough volume for a pattern to be described at all: 48 in Physics, 41 in Chemistry, 28 in Mathematics.

Those proportions say something awkward about a single sitting. A topic holding close to a tenth of a subject over two decades may still surface once or twice on a given morning, so any inference about it from one paper rests on one or two observations. That is too thin to act on, even while the total remains reliable enough to order a hall of candidates. The domain becomes visible at the resolution learning works at only when many years are pooled.

Conditions a credible measure would have to meet

The first condition is item-independence. An estimate must sit on a scale that separates the difficulty of the questions administered from the ability being estimated, in the manner of the item response models developed by Rasch and by Lord. Without that separation a percentage means only this percentage of these particular items, and no comparison across learners or across months is available.

The second is stability. Re-measurement after a short interval, with different items and no study in between, should reproduce the estimate within a stated band. A figure that swings several points on re-administration is reporting its own noise, and a published estimate with no error attached is advertising rather than measurement.

The third is that validation must not be circular. Correlating a knowledge measure against examination rank reinstates the compression the measure was built to escape, and a strong correlation would show only that the measure had learned to imitate a mark. The defensible evidence is prospective: the measure should predict performance on material the learner has not seen, and an intervention that moves the measure should move that performance too. Kane's argument-based approach sets the standard, with each inference between data and claim named and separately evidenced.

The fourth is resistance to its own use. Campbell's warning that a social indicator used for decisions comes under pressure to corruption applies in full to anything an admissions process might consult. A measure that can be raised by practising the measure rather than the material is worthless within a year of becoming consequential. Withheld and rotated items are the only defence, and the drift has to be monitored rather than assumed.

The fifth concerns what is counted. Chi, Feltovich and Glaser found in 1981 that experts sort physics problems by the governing principle while novices sort by surface features, so an identical score can sit on quite different representations of a subject. Items have to be built so that principle and surface come apart.

Measurement that returns the time it takes

The standing objection to more measurement is the time it takes from studying. Research on retrieval practice largely answers it. Roediger and Karpicke's 2006 experiments showed that retrieving material produces far more durable retention than the same time spent rereading, with the advantage widening as the delay before the final test grows. Spacing carries a similar record, reviewed by Cepeda and colleagues in 2006. Measurement built on spaced retrieval spends a learner's time once and returns both an estimate and the retention that retrieving produces.

What such a measure does not replace

None of this argues against examinations. A competitive examination rations scarce places, and must be timed, identical for everyone and defensible in public, which pushes it towards a single ordered list. A knowledge measure answers a different question and should not be asked to do the rationing. Its bounds are worth stating plainly, since a written instrument observes only what writing carries. A learner should be able to find out what they know against a description of the subject that does not shift with the cohort sitting beside them.

The data behind this

The full set, with progress tracking and five agent perspectives per question, is in the JupiteX app — browse the exam catalogue or browse the Learn library.