Elsie Item-Writing Academy
About 20 minutes · Academy module: Reading Step 1 Item Statistics
The Elsie Item-Writing Academy is an unofficial faculty-development resource, not affiliated with or endorsed by the AAMC, NBME, USMLE, LCME, or NRMP.
Each administered item comes back with a small set of statistics. They work the same way at every exam level; the context here is 5-option Step 1-style basic-science items, where random guessing lands at 0.20. (USMLE reports Step 1 as pass/fail; these statistics still govern the quality of your practice and formative items.)
N โ the number of responses. Every statistic below is a fraction with N in the denominator, so read N first. A p-value from 14 responses is a rumor; the same p-value from 400 responses is evidence. When N is small, treat the numbers as a preview, not a verdict.
p-value โ the proportion correct. Despite the name, this has nothing to do with significance testing; it is simply the fraction of examinees who selected the keyed answer. Near 0.20, the group performed at the level of random guessing. Near 1.00, nearly everyone answered correctly; near 0.00, nearly no one did. Neither extreme is automatically a defect: a very easy item may be doing exactly the warm-up job you designed it for. Extremes deserve a look, not a reflex โ check for a mis-key or a herding stem flaw.
Discrimination index โ who got it right. The p-value tells you how many answered correctly; the discrimination index tells you which examinees. It compares the proportion correct among the highest and lowest scorers (commonly the top and bottom thirds or quarters). Near +1.0, the item cleanly separates stronger from weaker examinees; near zero, it carries almost no information about who knows the material. Example: 0.85 โ 0.30 = 0.55 โ positive and healthy.
Negative discrimination is a review flag โ never a verdict. A negative index means lower-scoring examinees chose the keyed answer more often than higher-scoring ones. That pattern is a standing faculty-review flag: a mis-keyed answer, a second defensible correct option, or a subtle flaw misleading your strongest students. The flag starts an investigation; it never ends one. No item is deleted and no key is changed on a negative number alone โ a faculty member reads the item and confirms the problem first.
Distractor pull โ where the wrong answers went. For each option, the pull is the proportion of examinees who chose it. Healthy distractors each attract a real share โ pulls like 0.08, 0.12, 0.10, and 0.10 around a key at 0.60. In basic-science items, the best distractors are plausible mechanisms โ the wrong enzyme in the right pathway โ not random vocabulary. A distractor near zero fooled nobody โ revise it next use. A distractor that pulls harder than the key is the single most useful diagnostic in the set โ it usually marks a mis-key, a second correct answer, or a stem flaw, and paired with a negative index it tells you exactly where to look first.
The small-N rule. Below a minimum number of responses, item statistics are "not enough data yet" โ interesting, not actionable. Programs set their own minimums; whatever the threshold, an item is never revised, rekeyed, or retired on statistics alone until it clears it. Small samples manufacture dramatic numbers by chance, and acting on them is how good items get broken.
In the Exam Admin builder. Pooled per-item statistics live in the builder's item review panel: one row per item showing N, p-value, discrimination index, and a small bar for each option's pull, with the keyed option marked. Items with negative discrimination carry a faculty-review flag โ the flag routes the item to a human reader, and nothing is auto-deleted or auto-rekeyed. Opening an item puts the stem, options, and key beside the numbers. Kept this way, the review workflow supports faculty-development documentation: an evidence trail for every item you revise.
Illustrative data โ invented numbers for practice, not real exam statistics.
A 6-month-old infant presents with hepatomegaly and fasting hypoglycemia. Laboratory studies show lactic acidosis and hyperlipidemia. Liver biopsy reveals excess glycogen with abnormally short outer branches. Deficiency of which enzyme best explains these findings?
A) Glucose-6-phosphatase ยท B) Debranching enzyme ยท C) Branching enzyme ยท D) Glycogen phosphorylase ยท E) Glycogen synthase
Key: B
| Option | Chose it | Pull |
| A | 30 | 0.17 |
| B (key) | 99 | 0.55 |
| C | 18 | 0.10 |
| D | 22 | 0.12 |
| E | 11 | 0.06 |
| N | 180 | โ |
p-value = 99 / 180 = 0.55. Top-third correct = 0.85, bottom-third correct = 0.30, so discrimination = 0.85 โ 0.30 = 0.55. Every distractor pulled a real share โ note option A, the plausible-but-wrong enzyme, doing the most work. Read: a healthy, mid-difficulty item โ keep it as written.
All tables below are illustrative data โ invented numbers for practice, not real exam statistics.
| Option | Chose it | Pull |
| A | 15 | 0.10 |
| B | 20 | 0.13 |
| C | 75 | 0.50 |
| D (key) | 30 | 0.20 |
| E | 10 | 0.07 |
| N | 150 | โ |
p-value = 0.20. Top-third correct = 0.25, bottom-third correct = 0.45, so discrimination = 0.25 โ 0.45 = โ0.20.
Question: What action should the faculty member take?
Model answer: Flag the item for faculty review โ do not auto-delete it and do not change the key. Half the examinees chose distractor C while the key sits at guessing level, and the discrimination index is negative: the classic mis-key signature. The next step is to read the item and verify whether C is the correct answer; only a confirmed reading justifies a rekey.
| Option | Chose it | Pull |
| A (key) | 5 | 0.63 |
| B | 1 | 0.13 |
| C | 1 | 0.13 |
| D | 1 | 0.13 |
| E | 0 | 0.00 |
| N | 8 | โ |
p-value = 0.63. Discrimination cannot be meaningfully computed.
Question: What action should the faculty member take?
Model answer: None yet โ "not enough data yet." With N = 8, every number in this table could swing wildly with the next administration. Do not revise, rekey, or retire the item until it clears the program's minimum-N threshold.
| Option | Chose it | Pull |
| A | 45 | 0.15 |
| B | 60 | 0.20 |
| C | 48 | 0.16 |
| D | 57 | 0.19 |
| E (key) | 90 | 0.30 |
| N | 300 | โ |
p-value = 0.30. Top-third correct = 0.55, bottom-third correct = 0.15, so discrimination = 0.55 โ 0.15 = 0.40.
Question: The p-value is low. What action should the faculty member take?
Model answer: Keep the item. A low p-value is not a defect when the discrimination index is strongly positive and every distractor pulls an even, meaningful share โ that is the profile of a hard but fair item that separates stronger from weaker examinees exactly as designed. Difficulty alone is never a reason to revise.
Recording it adds the module to your Academy completion record on this device โ module, track, level, and date, ready to download from your account page for faculty-development documentation.