Why a flat list of subjects is a poor structure for learning

Nearly every exam syllabus is a flat list of subject names, arranged so that coverage can be audited rather than so that material can be learnt. Domains, categories and topics do work that such a list cannot, and the difference shows up in scores, in schedules, and in the decision to attempt a question.

The syllabus is an audit document

An examining body publishes a syllabus to settle one question: what may be asked. A list of headings settles it efficiently, since headings can be counted, compared year on year, and used to show that nothing outside the list appeared. None of that bears on how the material coheres, or on how a candidate should spend a Tuesday evening.

The shape then propagates: textbook chapters inherit it, coaching timetables inherit it, revision plans inherit it from those. Most candidates end up modelling a subject as a row of headings of roughly equal apparent size, with no relation between them beyond alphabetical or historical accident. The actual demand of a paper looks nothing like that.

Three levels, three different jobs

Three levels are enough, and each does distinct work. A domain is a discipline of reasoning: the moves that count as an argument within it, and the standard of evidence it accepts. A category is a family of problems that share a solution structure even when their surface details differ. A topic is the smallest unit to which a real question can be assigned reliably, and the smallest unit for which an accuracy figure carries information.

A collection of topics on its own is only a longer flat list. What the upper levels contribute is transfer. Chi, Feltovich and Glaser showed in 1981 that novices asked to sort physics problems grouped them by surface features, inclined planes with inclined planes and springs with springs, while experts grouped them by the principle required, a spring problem beside a pulley problem when both turn on conservation of energy. Expertise, on that finding, consists partly in holding a better category structure. A flat list of chapter names trains the novice sort and calls it a syllabus.

Beyond three levels the returns fall away: taggers disagree more as distinctions get finer, and the questions under any single node grow too few to show a pattern.

The weight is not evenly spread

The corpus behind this argument is JEE Main's past papers, 2002 to 2025: 8,267 questions, all authentic. Sorted into topics, 117 of them carry enough questions for a pattern to be described: 48 in Physics, 41 in Chemistry, 28 in Mathematics. Physics resolves into many describable units, while Mathematics concentrates into fewer and larger ones that recombine. Giving all three subjects the same number of equally weighted headings misrepresents all three.

The weights are lopsided. Matrices and determinants account for 8.8 per cent of JEE Main Mathematics in recent papers, 275 questions across the 23 years covered. Kinematics in one and two dimensions accounts for 6.6 per cent of recent Physics, 180 questions. Coordination compounds account for 8.7 per cent of recent Chemistry, 171 questions. At subject level none of this is visible, because Mathematics is one word. Only at the third level does the question of where a week should go have an answer better than a guess.

Grain decides whether a score means anything

The best-supported findings in the study of learning are scheduling findings, and scheduling needs units. Roediger and Karpicke's 2006 experiments on retrieval practice found that being tested on material produced far more durable retention than rereading it, even where rereading felt more productive at the time. The 2006 review of distributed practice by Cepeda and colleagues found spacing benefits across a very large body of experiments. Rohrer and Taylor's work on mixed practice of mathematics problems found that blocking raised performance during practice and lowered it on a delayed test.

Each result converts into an instruction only when there is something to schedule: what to retrieve, when to return to it, which items to shuffle together. A score at subject level generates none of those instructions. It averages over items of widely different frequency and difficulty, and two candidates holding the same Physics figure can need opposite work in the week ahead. A structure with a usable grain turns a diagnostic into a plan. A flat list turns it into a mood.

The decision to attempt is made at topic level

Marking schemes make the same point from the other end. Break-even guessing accuracy is the penalty divided by the sum of reward and penalty: 25 per cent in UPSC Prelims, 20 per cent in NEET, 20 per cent in JEE Main. The UPSC figure is exactly a blind pick from four options, which makes guessing there expected-value neutral. NEET and JEE break even at 20 per cent, so a blind pick from four is mildly positive, and any elimination makes it clearly so.

That arithmetic is usable only by a candidate who knows their own accuracy at a fine enough grain. The judgement made in the hall concerns one item, not the paper average, and the useful form of self-knowledge is that one class of problem is reliable and another is not. A candidate who believes they are good at Chemistry has no basis on which to attempt or skip a particular coordination-compounds item.

A maintained instrument, not a document

The structure itself stays unpublished, and not only out of commercial caution. A taxonomy has no value apart from the corpus it indexes. Its worth lies in the assignment of thousands of real questions to nodes, the judgement calls at the boundaries, and the counts that assignment produces. Detached from that, the names are decoration. A hierarchy read as a document also tends to become a checklist, and a checklist is a flat list again, with indentation.

There is a second reason. Any structure of this kind is a claim about how a body of knowledge coheres, and claims can be wrong. Boundaries between categories are argued over, papers move, and a category that fitted the questions of one decade may fit those of the next badly. Such an index is maintained equipment, revised whenever questions stop sitting comfortably in it. Publishing what the corpus shows is the honest half of the work.

The data behind this

The full set, with progress tracking and five agent perspectives per question, is in the JupiteX app — browse the exam catalogue or browse the Learn library.