Elsie Item-Writing Academy
About 20 minutes · Academy module: Reading Step 2 CK Item Statistics
The Elsie Item-Writing Academy is an unofficial faculty-development resource, not affiliated with or endorsed by the AAMC, NBME, USMLE, LCME, or NRMP.
Each administered item comes back with a small set of statistics. They work the same way at every exam level; the context here is 5-option Step 2 CK-style clinical vignette items โ often "next step in management" lead-ins โ where random guessing lands at 0.20.
N โ the number of responses. Every statistic below is a fraction with N in the denominator, so read N first. A p-value from 14 responses is a rumor; the same p-value from 400 responses is evidence. When N is small, treat the numbers as a preview, not a verdict.
p-value โ the proportion correct. Despite the name, this has nothing to do with significance testing; it is simply the fraction of examinees who selected the keyed answer. Near 0.20, the group performed at the level of random guessing. Near 1.00, nearly everyone answered correctly; near 0.00, almost no one did. Neither extreme is automatically a defect: a very easy item may be doing exactly the warm-up job you designed it for. Extremes deserve a look, not a reflex โ check for a mis-key or a herding stem flaw.
Discrimination index โ who got it right. The p-value tells you how many answered correctly; the discrimination index tells you which examinees. It compares the proportion correct among the highest and lowest scorers (commonly the top and bottom thirds or quarters). Near +1.0, the item cleanly separates stronger from weaker examinees; near zero, it carries almost no information about who knows the material. Example: 0.88 โ 0.40 = 0.48 โ positive and healthy.
Negative discrimination is a review flag โ never a verdict. A negative index means lower-scoring examinees chose the keyed answer more often than higher-scoring ones. That pattern is a standing faculty-review flag: a mis-keyed answer, a second defensible correct option, or a subtle flaw misleading your strongest students. The flag starts an investigation; it never ends one. No item is deleted and no key is changed on a negative number alone โ a faculty member reads the item and confirms the problem first.
Distractor pull โ where the wrong answers went. For each option, the pull is the proportion of examinees who chose it. Healthy distractors each attract a real share โ pulls like 0.08, 0.12, 0.10, and 0.10 around a key at 0.60. In clinical items, the best distractors are the plausible-but-wrong next steps โ the right drug for the wrong stage of disease. A distractor near zero fooled nobody โ revise it next use. A distractor that pulls harder than the key is the single most useful diagnostic in the set โ it usually marks a mis-key, a second correct answer, or a stem flaw, and paired with a negative index it tells you exactly where to look first.
The small-N rule. Below a minimum number of responses, item statistics are "not enough data yet" โ interesting, not actionable. Programs set their own minimums; whatever the threshold, an item is never revised, rekeyed, or retired on statistics alone until it clears it. Small samples manufacture dramatic numbers by chance, and acting on them is how good items get broken.
In the Exam Admin builder. Pooled per-item statistics live in the builder's item review panel: one row per item showing N, p-value, discrimination index, and a small bar for each option's pull, with the keyed option marked. Items with negative discrimination carry a faculty-review flag โ the flag routes the item to a human reader, and nothing is auto-deleted or auto-rekeyed. Opening an item puts the stem, options, and key beside the numbers. Kept this way, the review workflow supports faculty-development documentation: an evidence trail for every item you revise.
Illustrative data โ invented numbers for practice, not real exam statistics.
A 68-year-old man with heart failure with reduced ejection fraction remains symptomatic despite an ACE inhibitor and a beta blocker. His ejection fraction is 30% and his rhythm is normal sinus. Which of the following is the most appropriate next pharmacologic therapy?
A) Add an aldosterone antagonist ยท B) Add digoxin ยท C) Add a calcium channel blocker ยท D) Replace the ACE inhibitor with hydralazine alone ยท E) Add a thiazide diuretic
Key: A
| Option | Chose it | Pull |
| A (key) | 143 | 0.65 |
| B | 22 | 0.10 |
| C | 11 | 0.05 |
| D | 18 | 0.08 |
| E | 26 | 0.12 |
| N | 220 | โ |
p-value = 143 / 220 = 0.65. Top-third correct = 0.88, bottom-third correct = 0.40, so discrimination = 0.88 โ 0.40 = 0.48. Every distractor pulled a real share โ note option E, the plausible-but-premature diuretic, doing honest work. Read: a healthy, mid-difficulty item โ keep it as written.
All tables below are illustrative data โ invented numbers for practice, not real exam statistics.
| Option | Chose it | Pull |
| A | 56 | 0.40 |
| B (key) | 35 | 0.25 |
| C | 21 | 0.15 |
| D | 14 | 0.10 |
| E | 14 | 0.10 |
| N | 140 | โ |
p-value = 0.25. Top-third correct = 0.28, bottom-third correct = 0.50, so discrimination = 0.28 โ 0.50 = โ0.22.
Question: What action should the faculty member take?
Model answer: Flag the item for faculty review โ do not auto-delete it and do not change the key. Distractor A out-pulled the key by a wide margin and the discrimination index is negative: the classic mis-key or second-correct-answer signature. The next step is to read the item itself and check whether A is defensible; only a confirmed reading justifies a rekey or revision.
| Option | Chose it | Pull |
| A | 3 | 0.23 |
| B | 2 | 0.15 |
| C (key) | 6 | 0.46 |
| D | 1 | 0.08 |
| E | 1 | 0.08 |
| N | 13 | โ |
p-value = 0.46. Discrimination cannot be meaningfully computed.
Question: What action should the faculty member take?
Model answer: None yet โ "not enough data yet." With N = 13, the table is a sketch, not evidence. Do not revise, rekey, or retire the item until it clears the program's minimum-N threshold.
| Option | Chose it | Pull |
| A | 52 | 0.20 |
| B | 65 | 0.25 |
| C | 39 | 0.15 |
| D (key) | 104 | 0.40 |
| E | 0 | 0.00 |
| N | 260 | โ |
p-value = 0.40. Top-third correct = 0.70, bottom-third correct = 0.25, so discrimination = 0.70 โ 0.25 = 0.45.
Question: The item discriminates well, but option E pulled zero. What action should the faculty member take?
Model answer: Keep the item โ a p-value of 0.40 with a discrimination index of 0.45 is a strong profile. The one evidence-based move is to revise or replace distractor E, which fooled nobody across 260 responses; an option no one chooses is not a distractor, it is decoration. Everything else stays as written.
Recording it adds the module to your Academy completion record on this device โ module, track, level, and date, ready to download from your account page for faculty-development documentation.