Item difficulty and discrimination index

Item difficulty is the proportion of students answering correctly, with 0.40 to 0.79 the productive band, and the higher the value the easier the question. Discrimination is the gap in correct answers between the top 27 per cent and the bottom 27 per cent of scorers, where 0.40 or above is very good. A negative value almost always points to an answer key error or an ambiguous stem.

Once an exam has been sat, you hold more than a set of grades. Item analysis reduces that data to two numbers: how many students answered a question correctly, and whether the question separated those who knew the material from those who did not.

Both can be worked out with pen and paper, and both improve the next exam directly. What follows is the arithmetic, how to read the results, and the limit to watch for in small classes.

Item difficulty: how many got it right

The difficulty index is the number of students answering correctly divided by the total number of students. It is usually written as p and falls between 0 and 1.

The name misleads: as p rises the question gets EASIER. If 24 of 30 students answer correctly, p = 24 / 30 = 0.80, and that is an easy question.

Reading it: above 0.80 nearly the whole class knew it, below 0.20 nearly nobody did. Neither contributes to the spread of grades, because both give everyone the same thing.

  • p ≥ 0.80: very easy. One or two at the start of a paper can serve as a warm-up.
  • p 0.40 - 0.79: the productive band. Most questions belong here.
  • p 0.20 - 0.39: hard. A few on a paper is normal.
  • p < 0.20: either the topic was not learned or something is wrong with the question. Reread the stem to tell the two apart.

Discrimination: did the question separate anyone

Difficulty alone is not enough. A question can sit at a fine difficulty and still be useless, if the students who scored high overall and those who scored low performed the same on it.

To compute it, rank the class by total score. Take the top 27 per cent and the bottom 27 per cent. The discrimination index is the proportion correct in the upper group minus the proportion correct in the lower group, written as D.

The result runs between -1 and +1. You want it positive and large: the students who knew the material got it, and those who did not, did not.

  • D ≥ 0.40: very good. The question can be reused as it stands.
  • D 0.30 - 0.39: good. Small edits will improve it further.
  • D 0.20 - 0.29: borderline. Look at the option counts and revise.
  • D < 0.20: weak. The question is either ambiguous or measures something unrelated to the rest of the exam.
  • D negative: an alarm. See the section below.

A worked example

Take a class of 30 where 18 students answered question 7 correctly. The difficulty index is 18 / 30 = 0.60, which sits in the productive band.

Now for discrimination. 27 per cent of 30 is about 8 students. Of the top 8 by exam score, 7 answered this question correctly; of the bottom 8, only 2 did.

D = 7/8 - 2/8 = 0.875 - 0.250 = 0.625. That is very good discrimination. The question sits at a productive difficulty and cleanly separates the two groups, so it can be kept as it is.

Negative discrimination is a fault signal

A negative D means the students who scored high overall got this question right less often than those who scored low. That does not happen by chance; almost always one of three causes is behind it.

The first and most common is an answer key error. The strong students marked the correct option while the key counted a different one. Check this before anything else.

The second is ambiguity in the stem. A student who knows the topic well spots a second reading and rules out the "correct" option, while a weaker student never sees it and marks it. The third is one of the distractors being more correct than the answer.

A question with negative D is the one to reconsider before an appeal arrives.

Option by option: which distractor worked

The two indices describe the question as a whole; the spread across options tells you where it broke. Count how many students marked each option.

A distractor nobody marked is dead and quietly turns the question into a three-option one. The odds of guessing right go up and the question becomes easier than intended.

A distractor that many students from the upper group chose is a warning of its own: that option is either too close to the answer or partly correct.

In small classes the numbers move

In a class of 20, the top and bottom 27 per cent are five students each. One student in a group of five marking differently shifts D by 0.20. So do not discard a question on the numbers from a single sitting in a small group.

Two practical fixes: take the top and bottom third instead of 27 per cent, and track the same question across a few terms so the data accumulates.

Negative discrimination still deserves attention even with small numbers. The other thresholds wobble, but strong students knowing a question less often than weak ones is not natural at any class size.

Frequently asked

What should the difficulty index be?
Aim for most questions between 0.40 and 0.79. A few easy and a few hard questions on a paper are fine; all of them piled at one end is not.
Why 27 per cent for the upper and lower groups?
Under an assumption of normal distribution, that proportion gives the best balance between group size and how different the two groups are. In small classes, taking thirds is more practical.
What do I do with a question whose discrimination is negative?
Check the answer key first, since that is the most common cause. If the key is right, read the stem and options for ambiguity. If you find a problem, dropping the question from the scoring is a defensible call.
Do I need to run these numbers on every exam?
No. Doing it for questions you intend to reuse is enough. A question analysed once and found sound can be used again next term with confidence.

More from the guide

  • How to write skill-based questions

    Adding a long passage does not make a question skill-based. When context is needed, the decorative-context trap, and how to build a question from the skill up.

  • How to build a test blueprint

    A test blueprint shows which topic an exam measures and at which level. How the rows, the columns and the cell counts get decided, with a worked example.

  • How to write multiple choice distractors

    A weak distractor hands the answer to a student who does not know the topic. How to build them from real student errors, and seven patterns that give it away.