Track progress, take quizzes and save notes on this lesson.

Free forever · no card needed

Start free
Advanced

Statistical Testing: The Sign Test and Significance

4.2.3.3 Inferential testing

Aligned to the AQA 7182 specification

Level
Advanced
Reading time
11 min
Published
1 July 2026
On this page
  1. 1.Why Psychologists Use Statistical Tests
  2. 2.Probability and the 0.05 Level of Significance
  3. 3.Critical Values and Statistical Tables
  4. 4.The Sign Test: When to Use It
  5. 5.Calculating the Sign Test: The Method
  6. 6.Worked Example: Does a New Study Technique Improve Test Scores?
  7. 7.Type I and Type II Errors
  8. 8.Common Exam Mistakes

Key takeaways

  • Inferential tests decide whether a result reflects a real effect or is likely due to chance; if the probability of a chance result is low enough, the result is called statistically significant.
  • Psychology uses the 0.05 (5%) level of significance as standard: p ≤ 0.05 means there is a 5% or lower probability the results occurred by chance if the null hypothesis were true.
  • The sign test is used for a test of difference with a related (repeated measures) design and nominal data, or data reduced to the direction of change.
  • For the sign test, S is the number of times the less frequent sign occurs, N excludes participants who show no change, and the result is significant if S is equal to or less than the critical value.
  • A Type I error is a false positive (rejecting a true null hypothesis) and a Type II error is a false negative (retaining a false null hypothesis); the 0.05 level balances the risk of each.

Why Psychologists Use Statistical Tests

When a study produces a result — say, a therapy group improves after treatment — one question decides everything: is that improvement a real effect, or could it just be down to chance? Random variation alone can make scores drift up or down. An inferential (statistical) test answers this by working out the probability that the result would occur if there were no real effect at all.

If that probability is low enough, the result is judged statistically significant and the psychologist concludes a genuine effect is present. If the probability is high, the result could easily be a fluke, so it is not treated as evidence of an effect.

Question the test answersWhat it decides
Could this result be down to chance?Whether to trust the finding
Is the probability of a chance result low enough?Whether the result is significant
Do we reject the null hypothesis?Whether a real effect is present

An inferential test does not prove an effect is real. It tells you how likely the result would be if there were no effect, and lets you decide whether that likelihood is small enough to reject that idea.

Every statistical test starts from the null hypothesis — the statement that there is no difference or no relationship. The test tells you whether the data give you enough grounds to reject it.

Probability and the 0.05 Level of Significance

Psychology cannot demand certainty, so it works with probability. The agreed standard is the 0.05 (5%) level of significance, written p ≤ 0.05. This means: there is a 5% or lower probability that the results occurred by chance if the null hypothesis were true. Reach that threshold and the result is called statistically significant.

Why 0.05 and not something else? It is a compromise. A stricter level (such as p ≤ 0.01) makes you very sure any effect you claim is real, but risks dismissing genuine effects. A looser level (such as p ≤ 0.10) catches more real effects, but lets more chance results slip through as false claims.

Significance levelMeaningTrade-off
p ≤ 0.1010% chance the result is a flukeToo lenient — accepts weak evidence
p ≤ 0.055% chance the result is a flukeStandard balance used in psychology
p ≤ 0.011% chance the result is a flukeStringent — used where certainty matters

The 0.05 level does not mean "there is a 5% chance the effect is real". It means: if the null hypothesis were true, there is only a 5% (or lower) probability of getting a result this extreme by chance.

To reach a decision, the psychologist compares a calculated value (worked out from the data) against a critical value (looked up from a table). How that comparison works is the next step.

Critical Values and Statistical Tables

Every inferential test produces a calculated value (also called the observed value) from the data. On its own that number means nothing. It only becomes meaningful when compared against a critical value taken from a statistical table.

To find the correct critical value, you need three pieces of information:

  • N or degrees of freedom — usually the number of participants in the study.
  • One-tailed or two-tailed — whether the hypothesis is directional (predicts the direction of the effect) or non-directional (predicts a difference but not its direction).
  • The level of significance — normally p ≤ 0.05.
Hypothesis typeAlso calledTable column to use
Directional (predicts direction)One-tailedOne-tailed critical value
Non-directional (predicts a difference only)Two-tailedTwo-tailed critical value

A directional hypothesis states which way the result will go ("scores will increase"). A non-directional hypothesis only states that there will be a difference ("scores will change"). This choice decides which column of the table you read.

For most tests the rule is that the calculated value must be equal to or greater than the critical value to be significant. The sign test is an exception: for the sign test the calculated value must be equal to or less than the critical value. Learning which way the comparison runs for each test is essential, and it is where many marks are lost.

The Sign Test: When to Use It

The sign test is the one inferential test AQA requires you to be able to calculate in full. It is used only under a specific set of conditions, and all of them must hold.

ConditionRequirement for the sign test
What is being testedA difference (not a correlation)
Experimental designA related design — repeated measures
Level of measurementNominal data, or data reduced to the direction of change

A related design means the same participants are measured twice — for example, before and after an intervention — so each participant provides a matched pair of scores. The sign test then throws away the exact numbers and keeps only the direction of change for each person: did they get better (a plus) or worse (a minus)?

Use the sign test when you have a test of difference, a repeated measures design, and nominal data (or scores reduced to better/worse). If the design is independent groups, or you are looking at a relationship rather than a difference, the sign test is the wrong choice.

Because it only looks at direction, the sign test is quick to calculate by hand — which is exactly why it is the one the specification asks you to compute.

Worth saving these ideas?

Turn what you've read into instant revision cards. Free to get started.

Make flashcards

Calculating the Sign Test: The Method

The calculation follows six clear steps. Learn them as a fixed sequence.

  1. State the hypotheses. The null hypothesis says there is no difference; the alternative hypothesis predicts a difference (directional or non-directional).
  2. Work out the sign of the difference for each participant: a plus (+) for an increase, a minus (−) for a decrease.
  3. Ignore any participant who shows no change (a difference of zero). They are dropped completely.
  4. Count the pluses and the minuses. The calculated value S is the number of times the LESS frequent sign occurs.
  5. Work out N — the number of participants excluding those who showed no change.
  6. Look up the critical value of S for that N, at p ≤ 0.05, using the one- or two-tailed column.

The decision rule is the part students get wrong most often:

For the sign test, the result is statistically significant if the calculated value S is equal to or less than the critical value. If S is greater than the critical value, you retain the null hypothesis.

This runs the opposite way to most tests, where a bigger calculated value is what you want. For the sign test, a small S (few contradicting signs) is the evidence of a consistent, real difference.

Worked Example: Does a New Study Technique Improve Test Scores?

Ten students sit a memory test, are taught a new revision technique, then sit an equivalent test. Higher scores are better. We predict a difference but not its direction, so the hypothesis is non-directional (two-tailed).

ParticipantBeforeAfterSign (After − Before)
11218+
21514
3916+
411110 (no change)
5815+
61420+
71019+
8713+
91317+
10612+

Step 1 — drop the no-change participant. Participant 4 scored 11 both times, a difference of zero, so they are removed.

Step 2 — count the signs for the remaining nine participants:

Pluses  (+): P1, P3, P5, P6, P7, P8, P9, P10  →  8
Minuses (−): P2                                →  1
No change (0): P4                              →  excluded

Step 3 — find S and N.

  • The less frequent sign is the minus, which occurs once, so S = 1.
  • N excludes the no-change participant: N = 8 + 1 = 9.

Step 4 — compare with the critical value. For N = 9, two-tailed, at p ≤ 0.05, the critical value of S is 1.

Since the calculated S (1) is equal to the critical value (1), and the rule is significant when S is equal to or less than the critical value, the result is statistically significant (p ≤ 0.05).

Conclusion: we reject the null hypothesis. The new study technique produced a statistically significant difference in test scores (S = 1, N = 9, p ≤ 0.05, two-tailed).

Type I and Type II Errors

No statistical decision is risk-free. Because we work with probability, a test can reach the wrong conclusion in two distinct ways.

A Type I error is a false positive: rejecting the null hypothesis when it is actually true. The psychologist claims an effect that is not really there. This becomes more likely if the significance level is too lenient — for example, using p ≤ 0.10 makes it easier to declare significance on flimsy evidence.

A Type II error is a false negative: retaining (accepting) the null hypothesis when it is actually false. A real effect exists, but the test misses it. This becomes more likely if the significance level is too stringent — for example, p ≤ 0.01 sets the bar so high that genuine effects fail to reach it.

Null hypothesis is really TRUENull hypothesis is really FALSE
Reject nullType I error (false positive)Correct decision
Retain nullCorrect decisionType II error (false negative)

The 0.05 level is a deliberate balance: strict enough to keep the Type I (false positive) risk to about 5%, but lenient enough to avoid missing too many real effects through Type II (false negative) errors. Making the level stricter cuts Type I errors but raises Type II errors, and vice versa.

A memory hook: Type I = "I claim an effect that isn't real" (false alarm); Type II = "too cautious, I miss the real effect".

Common Exam Mistakes

1. Taking S as the more frequent sign

The calculated value S is the number of times the less frequent sign occurs, not the more frequent one. In the worked example there were 8 pluses and 1 minus, so S = 1 (the minuses), not 8. Counting the majority sign inverts the whole test.

2. Including no-change participants in N

Any participant whose score does not change (a difference of zero) is removed entirely. They are not counted in the pluses, the minuses or N. With 10 participants and one showing no change, N = 9, not 10.

3. Thinking a result is significant when S is greater than the critical value

The sign test runs the opposite way to most tests. The result is significant only when the calculated S is equal to or less than the critical value. If S is greater than the critical value, you retain the null hypothesis. Reading the comparison the wrong way is a frequent error.

4. Confusing Type I and Type II errors

A Type I error is a false positive (claiming an effect that is not real, from rejecting a true null hypothesis). A Type II error is a false negative (missing a real effect, from retaining a false null hypothesis). Swapping the two, or mixing up which is caused by a lenient versus a stringent significance level, loses marks.

5. Using the sign test on an independent-groups design

The sign test requires a related (repeated measures) design, where each participant is measured twice so their scores form a matched pair. It cannot be used when two separate groups of people are compared — that is an independent-groups design, which needs a different test.

6. Reading the wrong column of the critical values table

The critical value depends on whether the hypothesis is directional (one-tailed) or non-directional (two-tailed). Reading the one-tailed column for a two-tailed hypothesis (or the reverse) gives the wrong critical value and can flip the conclusion. Match the column to the hypothesis and the significance level before reading off the value.

Key terms

Statistical significance
The point at which the probability that a result is due to chance is low enough (p ≤ 0.05 in psychology) to conclude that a real effect is present.
The sign test
An inferential test of difference for a related design with nominal data, based on counting the direction (sign) of change for each participant.
Critical value
A value looked up from statistical tables, using N, the significance level and whether the hypothesis is one- or two-tailed, against which the calculated value is compared.
The 0.05 level of significance
The standard threshold in psychology, meaning a 5% or lower probability that the results occurred by chance if the null hypothesis were true.
Type I error
A false positive: rejecting the null hypothesis when it is actually true, claiming an effect that is not real.
Type II error
A false negative: retaining the null hypothesis when it is actually false, missing an effect that is really there.
Null hypothesis
A statement that there is no difference or no relationship, which the inferential test is used to reject or retain.

Frequently asked questions

Use the sign test when you are testing for a difference, using a related (repeated measures) design, with nominal data or data reduced to the direction of change (better or worse). All three conditions must hold.

No. For the sign test the result is significant only when the calculated value S is equal to or less than the critical value from the table. A calculated value greater than the critical value means you retain the null hypothesis.

A Type I error is a false positive: rejecting the null hypothesis when it is actually true. A Type II error is a false negative: retaining the null hypothesis when it is actually false. A lenient significance level raises Type I risk; a stringent one raises Type II risk.

Generate revision on any topic you study

Type any topic you're studying and Aicademy generates a complete lesson, quiz, and flashcard set, personalised to your level.

Lessons on anything

Structured, level-matched lessons on any topic you study

Practice quizzes

Find out what you actually know before the exam does

Flashcard sets

Lock in key concepts with instant revision cards

Ask Aica

Stuck on something? Get a clear explanation, any time

Prev

Data Presentation, Levels of Measurement and Correlation

Next

Choosing a Statistical Test

Related lessons

Top students don’t revise more. They revise what counts.

Start revising free

Free to start. No card needed.