Track progress, take quizzes and save notes on this lesson.

Free forever · no card needed

Start free
Intermediate

Reliability, Validity and Features of Science

4.2.3.1 Scientific processes

Aligned to the AQA 7182 specification

Level
Intermediate
Reading time
11 min
Published
1 July 2026
On this page
  1. 1.Reliability vs Validity: The Core Distinction
  2. 2.Measuring Reliability: Test-Retest and Inter-Observer
  3. 3.Improving Reliability
  4. 4.Types of Validity
  5. 5.Improving Validity
  6. 6.Features of Science
  7. 7.Reporting Psychological Investigations
  8. 8.Common Exam Mistakes

Key takeaways

  • Reliability means consistency of a measurement; validity means whether a measure actually measures what it claims to and whether findings are genuine. They are different qualities: a test can be reliable but not valid.
  • Test-retest reliability correlates the scores of the same people tested twice; inter-observer reliability correlates two observers recording the same behaviour. A correlation of about +0.8 or above is good.
  • External validity includes ecological validity (generalising to other settings and real life) and temporal validity (generalising across time); these are commonly confused.
  • Popper argued a scientific theory must be falsifiable (capable of being proven wrong). Falsifiability is a strength of a scientific theory, not a weakness.
  • A scientific report has a fixed structure: abstract, introduction, method, results, discussion and referencing, and the method contains design, participants, apparatus, procedure and ethics.

Reliability vs Validity: The Core Distinction

Two words dominate this part of the specification, and mixing them up costs marks in almost every research methods paper. Get the definitions clear first.

Reliability means the consistency of a measurement. A reliable measure produces the same result each time it is used under the same conditions. If you weigh yourself three times in a minute and get 70 kg, 70 kg, 70 kg, the scale is reliable.

Validity means the legitimacy or accuracy of a measurement: whether a measure actually measures what it claims to, and whether the findings of a study are genuine. If your scale reads 70 kg but you truly weigh 65 kg, it is reliable (consistent) but not valid (accurate).

QualityQuestion it answersEveryday word
ReliabilityDoes it give the same result each time?Consistency
ValidityDoes it measure the true thing?Accuracy

A measure can be reliable without being valid. A broken clock stopped at 3 o'clock is perfectly consistent but tells the correct time only twice a day. Reliability is necessary for validity, but it is not sufficient on its own.

Measuring Reliability: Test-Retest and Inter-Observer

The specification names exactly two ways of measuring reliability, and both work by producing a correlation coefficient.

Test-retest reliability assesses the consistency of a test, questionnaire or interview over time. You give the same test to the same people on two different occasions, then correlate the two sets of scores. There must be enough of a gap that participants do not simply recall their previous answers, but not so long that the trait itself has changed.

Inter-observer reliability assesses the consistency of an observation. Two or more observers watch and record the same behaviour independently, then their records are correlated. High agreement means the behavioural categories are being applied consistently rather than reflecting one observer's personal judgement.

For both methods, the same benchmark applies:

A correlation coefficient of about +0.8 or above indicates good reliability. Below this, the measure is treated as insufficiently consistent.

MethodWhat is correlatedUsed to check reliability of
Test-retestScores from the same people on two occasionsTests, questionnaires, interviews
Inter-observerRecords of two or more observers watching togetherObservations

Note that test-retest uses the same test twice, not two different tests. Using a different test on the second occasion would confound the measurement.

Improving Reliability

If a measure falls short of the +0.8 benchmark, there are specific, spec-relevant ways to raise it. The fix depends on the type of study.

  • Standardise the procedure. Every participant should experience the same instructions, timing and conditions so the process itself does not introduce variation.
  • Operationalise behavioural categories clearly. In observations, define categories so precisely that any trained observer would code the same behaviour the same way, with no overlap or ambiguity between categories.
  • Train observers. Observers who practise against an agreed standard before the real observation apply the categories more consistently, raising inter-observer reliability.
  • Replace ambiguous questionnaire items. In test-retest, questions that could be read two ways produce inconsistent answers. Rewriting or removing them (for example, replacing open questions with closed ones, or clarifying wording) improves consistency.

Worked example. Two observers coding playground aggression agree only 60% of the time (a low correlation). The category "aggressive" is too vague. Splitting it into operationalised categories such as "pushes", "kicks" and "shouts a threat" gives each observer an unambiguous rule to apply, and re-running the observation lifts their correlation above +0.8.

The general principle: reliability improves when you remove sources of variation that come from the procedure or the observer rather than from the participants' actual behaviour.

Types of Validity

Validity is broader than reliability and the specification names four types you must be able to define and distinguish, sitting under two headings.

Internal validity concerns whether the independent variable (not a confounding variable) actually caused the change in the dependent variable. Demand characteristics, investigator effects and poor controls all threaten it.

External validity concerns generalisability — whether findings extend beyond the specific study. It splits into:

TypeHeadingWhether findings generalise to…
Ecological validityExternalOther settings and everyday, real-life situations
Temporal validityExternalOther time periods; still true across the years
Population validity*ExternalPeople beyond the sample studied

(*Extra context — population validity is widely taught but is not one of the four validity types named by AQA 7182; only face, concurrent, ecological and temporal validity are required.)

Two further types describe how validity is assessed:

TypeWhat it checks
Face validityWhether a measure looks, on the surface, like it measures what it claims
Concurrent validityWhether results agree closely with an already-established valid measure

Worked example. A researcher writes a new 10-item stress questionnaire. To check face validity she asks whether the items obviously relate to stress. To check concurrent validity she gives it to people who also complete a well-established stress scale; if the two sets of scores correlate closely (again, around +0.8 or above), concurrent validity is demonstrated.

Worth saving these ideas?

Turn what you've read into instant revision cards. Free to get started.

Make flashcards

Improving Validity

As with reliability, the specification expects concrete techniques rather than a vague "control things better". Match the technique to the threat.

  • Use a control group. Comparing the experimental group against a control group lets the researcher attribute any difference to the independent variable rather than to chance or extraneous factors, protecting internal validity.
  • Standardise procedures. Standardisation reduces the influence of extraneous variables and investigator effects, so the IV is the only thing that differs between conditions.
  • Use single-blind and double-blind procedures. In a single-blind design the participant does not know which condition they are in, reducing the effect of demand characteristics. In a double-blind design neither the participant nor the researcher interacting with them knows, which also removes investigator effects.
  • Guarantee anonymity. In questionnaires and interviews, assuring participants their responses are anonymous reduces social desirability bias — the tendency to answer in a way that looks good rather than truthfully — so responses are more valid.

A double-blind procedure protects validity on two fronts at once: it controls demand characteristics (the participant cannot second-guess the aim) and investigator effects (the researcher cannot unconsciously cue a response).

Features of Science

For a discipline to count as a science, it must show certain features. The specification lists these explicitly, so learn them as a set.

  • Objectivity and the empirical method. Objectivity means data are gathered free of bias or personal opinion. The empirical method gathers knowledge through direct observation and measurement rather than through reasoning or belief alone.
  • Replicability and falsifiability. Findings must be replicable — repeatable by other researchers to confirm they are genuine. Karl Popper argued that a theory is only scientific if it is falsifiable, meaning it is capable of being proven wrong by evidence. A theory that can explain away every possible result is unfalsifiable and therefore unscientific.
  • Theory construction and hypothesis testing. A theory is a set of principles that explains observations. Good theories generate clear, testable hypotheses; testing those hypotheses lets the theory be supported, refined or rejected.
  • Paradigms and paradigm shifts. Thomas Kuhn argued that a paradigm is a shared set of assumptions, methods and terminology within a discipline. A paradigm shift is a scientific revolution in which the accepted paradigm is overthrown by a new one when enough contradictory evidence accumulates.

Kuhn argued that psychology may be pre-paradigmatic — it lacks a single unifying paradigm because approaches such as the behaviourist, cognitive and biological each hold different core assumptions. This is a common evaluation point when asked whether psychology is a science.

Worked example of falsifiability. "People behave the way they do because of unconscious forces we cannot observe" is hard to falsify — no result could disprove it. "Caffeine reduces reaction time" is falsifiable: a study finding no effect, or a slower reaction time, would prove it wrong. The second is scientific; the first is criticised as unscientific.

Reporting Psychological Investigations

When psychologists write up a study for a journal, they follow a fixed structure so that others can understand, evaluate and replicate the work. Learn the sections in order.

SectionWhat it contains
AbstractA short summary (around 150–200 words) of the aims, method, results and conclusions
IntroductionA review of relevant past research, narrowing to the aims and hypotheses
MethodEnough detail to replicate the study exactly (see breakdown below)
ResultsThe findings, using descriptive and inferential statistics, tables and graphs
DiscussionWhat the results mean, links to previous research, limitations and implications
ReferencingFull citations of every source mentioned, in a standard format

The method section is itself subdivided:

  • Design — the experimental design (for example repeated measures) and why it was chosen.
  • Participants / sampling — who took part, how many, and the sampling method used.
  • Apparatus / materials — the equipment, stimuli or materials used.
  • Procedure — a step-by-step account of exactly what was done, in enough detail to repeat it.
  • Ethics — how ethical issues such as consent, confidentiality and protection from harm were dealt with.

The abstract comes first but is written last, because it summarises the whole report. The discussion, not the results, is where interpretation belongs — the results section reports the numbers without explaining them.

Common Exam Mistakes

1. Confusing reliability with validity

Reliability is consistency (same result each time); validity is accuracy (measuring the true thing). A measure can be reliable but not valid. Define both correctly and keep them apart — this single confusion loses marks across the whole topic.

2. Saying test-retest uses two different tests

Test-retest gives the same test to the same people on two occasions. Using a different test the second time is not test-retest reliability and would confound the comparison.

3. Muddling ecological and temporal validity

Ecological validity is about generalising to other settings and real life. Temporal validity is about generalising across time. Both are external validity, but one concerns place and the other concerns era.

4. Thinking a falsifiable theory is a weak one

Falsifiability is a strength: a scientific theory must be capable of being proven wrong. A theory that cannot be falsified (that explains any result) is criticised as unscientific. Do not treat "falsifiable" as a criticism.

5. Putting report sections in the wrong order

The order is abstract, introduction, method, results, discussion, referencing. Interpretation belongs in the discussion, not the results. Remember the abstract is placed first even though it is written last.

6. Quoting the wrong reliability benchmark

Good reliability is a correlation of about +0.8 or above for both test-retest and inter-observer. Stating a lower cut-off, or giving no figure at all when the question asks how reliability is judged, weakens the answer.

Key terms

Reliability
The consistency of a measurement; the extent to which a procedure or measure produces the same result each time it is used under the same conditions.
Test-retest reliability
A way of assessing reliability by giving the same test or questionnaire to the same people on two occasions and correlating the two sets of scores.
Inter-observer reliability
The extent to which two or more observers recording the same behaviour independently produce records that agree, assessed by correlating their data.
Validity
Whether a measure measures what it claims to measure and whether the findings from a study are genuine and legitimate.
Ecological validity
A form of external validity concerning whether findings can be generalised to other settings and to everyday, real-life situations.
Falsifiability
The principle, associated with Popper, that a scientific theory must be capable of being proven wrong by evidence.
Paradigm
A shared set of assumptions, methods and terminology within a scientific discipline, as described by Kuhn.

Frequently asked questions

Reliability is the consistency of a measurement, so the same result is produced each time. Validity is the accuracy or legitimacy of a measurement, so it measures what it claims to and the findings are genuine. A measure can be reliable without being valid.

A correlation coefficient of about +0.8 or above indicates good reliability. In test-retest the two sets of scores from the same people are correlated; in inter-observer reliability the records of two or more observers are correlated.

Ecological validity is whether findings generalise to other settings and to real life. Temporal validity is whether findings generalise across time, remaining true in later eras. Both are forms of external validity but they concern setting versus era.

Generate revision on any topic you study

Type any topic you're studying and Aicademy generates a complete lesson, quiz, and flashcard set, personalised to your level.

Lessons on anything

Structured, level-matched lessons on any topic you study

Practice quizzes

Find out what you actually know before the exam does

Flashcard sets

Lock in key concepts with instant revision cards

Ask Aica

Stuck on something? Get a clear explanation, any time

Prev

Ethics, Peer Review and the Economy

Next

Descriptive Statistics and Distributions

Related lessons

Top students don’t revise more. They revise what counts.

Start revising free

Free to start. No card needed.