HookIn 2015, psychologists re-ran 100 studies — fewer than half survived
In 2015 a team of 270 researchers did something the field had never seriously attempted. They re-ran 100 published psychology experiments exactly as the originals described them and checked whether the same result came back. Fewer than half did. The 'replication crisis' did not mean psychologists had been lying — it meant that small, loosely-controlled studies with hand-picked participants had been treated as settled fact when they should not have been. Every idea you meet in this course — Milgram's obedient teachers, Loftus's rewritten memories, Piaget's neat developmental stages — is only ever as strong as the method that produced it.
That is why the research-methods section carries more marks than any single theory, and why it is the most predictable part of the whole course to revise. It is not asking you to remember a story; it is asking whether a study is built to answer its own question. Learn the machinery once — how to write a hypothesis, name the variables, pick a sample, choose a design, control what could go wrong, protect the participants, and turn raw scores into a defensible claim — and you can take apart any study the examiner puts in front of you, including ones you have never seen before.
ModelHypotheses and variables — turning a hunch into something you can test
An aim is the vague goal ('to investigate whether noise affects revision'). A hypothesis is the precise, falsifiable prediction you actually test. The skill the examiner rewards is operationalisation: defining your variables so exactly that someone else could measure them the same way. 'Noise affects memory' is not testable; 'participants who revise with music at 70 decibels recall fewer words from a 20-word list than participants who revise in silence' is.
Every experiment has an independent variable (IV) — the thing the researcher deliberately changes or manipulates — and a dependent variable (DV) — the thing they measure to see the effect. Music-versus-silence is the IV; words recalled is the DV. A hypothesis can be directional (it states which way the result will go: 'recall fewer words') when past research points a direction, or non-directional (it predicts a difference without saying which way) when the field is undecided. The null hypothesis is the formal 'no effect' version — 'there is no difference in recall between the two conditions' — and the whole point of running the study is to gather enough evidence to reject it.
The two variables that wreck studies are the ones you did not plan for. An extraneous variable is anything other than the IV that could affect the DV — time of day, room temperature, how clever the participants are. It becomes a far more dangerous confounding variable the moment it varies systematically with the IV: if everyone in the silence condition was tested at 9am and everyone in the music condition at 4pm, you can no longer tell whether music or tiredness caused the difference. Naming and controlling these is where evaluation marks live.
ModelDesigns and procedures — the shape of the study
There are three experimental designs, each with a trade-off. In independent groups, different people do each condition — quick, no order effects, but the two groups might simply differ (participant variables). In repeated measures, the same people do every condition, which removes participant variables but creates order effects: people get better with practice or worse with fatigue. The fix is counterbalancing — half do condition A then B, half do B then A. Matched pairs is the compromise: different people, but paired on key characteristics (age, IQ) so the groups start comparable.
Experiments themselves come in flavours. A laboratory experiment gives tight control and easy replication but low realism. A field experiment runs in a real setting (higher ecological validity, weaker control). A natural experiment uses an IV that already varies in the world (you cannot ethically make children suffer neglect, so you study children who already have). Beyond experiments sit the non-experimental methods: observations (naturalistic or controlled; overt or covert; participant or non-participant), questionnaires (open questions give rich detail, closed questions give easy numbers), interviews (structured for consistency, unstructured for depth) and case studies (deep, detailed accounts of one person or group — rich but hard to generalise).
MechanismSampling — who ends up in the study, and why it decides everything
You almost never test everyone you care about. The target population is the whole group you want to draw conclusions about; the sample is the smaller set you actually study. The question is always whether the sample is representative enough to generalise from.
AQA expects four methods. Random sampling gives everyone in the target population an equal chance (names from a hat, or a random number generator) — fair in principle, but you might still, by luck, pick an unrepresentative set. Systematic sampling takes every nth person from a list (every 10th name) — objective, but not truly random. Stratified sampling divides the population into subgroups (strata) and samples each in proportion, so a school that is 60% female yields a sample that is 60% female — the most representative, but fiddly. Opportunity sampling just uses whoever is available and willing — cheap and fast, which is exactly why so much psychology is built on undergraduates and passers-by, and exactly why it is so easily biased. A common trap: a large sample is not automatically a good one. A biased sample of 500 is still biased; representativeness beats size.
MechanismCorrelation — spotting a relationship without proving a cause
An experiment manipulates an IV to test cause and effect. A correlation does something different: it measures whether two co-variables that already exist tend to move together, without changing either. Hours of sleep and exam score; screen time and anxiety; temperature and ice-cream sales.
A relationship can be positive (both rise together — more revision, higher marks), negative (one rises as the other falls — more absences, lower marks) or zero (no consistent pattern). Its strength is captured by a correlation coefficient, a number from −1 to +1: values near ±1 mean a strong relationship, values near 0 mean a weak or non-existent one. You display it on a scatter diagram, one dot per participant. The line every examiner wants you to hold is this: correlation is not causation. Ice-cream sales and drownings correlate, but neither causes the other — hot weather (a third variable) drives both. Correlations are brilliant for studying things you cannot ethically or practically manipulate, and useless for proving what causes what.
CaseEthics, reliability and validity — getting it right before you run it
Two quality words run through the whole section. Reliability is consistency: would you get the same result if you did it again? You build it in with standardised procedures — every participant gets identical instructions and conditions — and check it with a pilot study, a small trial run that catches problems before you waste the real sample. Validity is accuracy: are you actually measuring what you claim to measure, and does it generalise beyond the lab? Controlling extraneous variables protects validity.
Then there is the non-negotiable: ethics. British psychologists work to a code built on a handful of principles. Informed consent — participants agree knowing what they are letting themselves in for. Deception — you should not mislead participants, and where a study genuinely cannot work without it, it must be justified and followed by a full debrief. Protection from harm — participants should leave in no worse a state, physically or psychologically, than they arrived. Right to withdraw — they can stop at any point and remove their data. Confidentiality and privacy — personal data is protected and participants are not identifiable. Psychologists deal with these through consent forms, debriefing, ethics committees and a cost-benefit judgement: do the potential gains to knowledge outweigh the costs to the people taking part? Milgram's obedience study is the case examiners reach for — real scientific value, but serious questions over deception and protection from harm.
DataMaking the numbers talk — data, statistics and distributions
Psychology deals in two kinds of data. Quantitative data is numbers (words recalled, reaction time in milliseconds) — easy to analyse and compare, but thin on the 'why'. Qualitative data is words and meaning (interview transcripts, open-ended answers) — rich and detailed, but harder to summarise objectively. Data is also either primary (collected first-hand by the researcher for this study) or secondary (already gathered by someone else — government statistics, another team's results), which is quicker but was not designed for your question.
To summarise quantitative data you use descriptive statistics. The three averages (measures of central tendency) are the mean (add them all, divide by how many — uses every score, but a single extreme outlier drags it), the median (the middle value when ordered — resistant to outliers) and the mode (the most frequent value — the only average that works for categories). Spread is shown by the range (highest minus lowest — crude, because it only uses two scores). You display the results with the right chart: a bar chart for separate categories (gaps between bars), a histogram for continuous data (bars touch), a scatter diagram for a correlation, and frequency tables to organise raw scores. And you should be fluent in the maths that carries roughly a tenth of the whole GCSE: percentages, fractions, ratios, decimals, significant figures, standard form and simple estimation.
One shape underlies much of this: the normal distribution. Measure something like reaction time or a memory score across hundreds of people and the results form a symmetrical bell curve — most scores cluster around the middle and the mean, median and mode all sit together at the peak, with a few very high and very low scores tailing off each side. If a distribution leans to one side it is skewed: a test that is far too easy bunches scores near the top, pulling the tail out to the left (a negative skew).
Nine students sit a 20-mark memory test. Their scores, put in order, are: 9, 11, 12, 14, 15, 15, 15, 18, 20.
Mean: add them (9+11+12+14+15+15+15+18+20 = 129) and divide by the number of scores (9): 129 ÷ 9 = 14.3 marks (to 3 significant figures). Median: with nine scores the middle one is the 5th in order = 15 marks. Mode: the most frequent score is 15 (it appears three times). Range: 20 − 9 = 11 marks.
Now turn the average into a percentage — a classic exam maths step: the mean of 14.3 out of 20 is 129 ÷ 180 = 0.717, i.e. 71.7%. Notice the mean (14.3), median (15) and mode (15) are close but not identical, which tells you the distribution is only mildly skewed. If you had scores from 900 students rather than 9, they would settle into that bell-shaped normal distribution, and the three averages would converge on the peak. Which average would you report? The mean uses all the data but is the most vulnerable to that lone score of 9; the median ignores the outlier; the mode is really only useful for the most common result. Stating which average you chose and why is the sentence that turns a correct calculation into a full-mark answer.
VocabularyKey terms the mark scheme pays for
TrapsMisconceptions that cost marks
ExamWhat examiners want
AQA GCSE Psychology (8182) marks against three objectives, and research methods leans hardest on two of them. AO2 (apply) dominates: the exam hands you an unfamiliar study and asks you to name the IV and DV, write an operationalised hypothesis, spot an extraneous variable or choose a sampling method for that scenario — so never answer in the abstract, always in the terms of the stem. AO3 (analyse and evaluate) is where the design trade-offs pay off: strengths and weaknesses of a chosen design, sample or method, tied back to the study in front of you.
Roughly 10% of the whole qualification is maths, and most of it surfaces here — practise means, medians, modes, ranges, percentages and ratios by hand, show your working, and always give the unit. On 'write a hypothesis' questions, operationalise both variables and pick directional or non-directional based on whether earlier research points a way. Watch the command words: 'identify' or 'outline' is AO1 and wants a crisp fact; 'explain how you would…' is AO2 and wants it done to the scenario; 'evaluate' or 'discuss' is AO3 and wants balanced judgement with a conclusion. The single most-dropped mark is the interpretation sentence after a calculation — state what the number means, not just what it is.