HookThe paper that proved people can see the future — and then didn't
In 2011 the Journal of Personality and Social Psychology, one of the most respected titles in the field, published a paper by the Cornell social psychologist Daryl Bem titled 'Feeling the Future'. Across nine experiments and more than a thousand participants, Bem reported statistically significant evidence for precognition: people appeared to 'remember' words more accurately if they were going to revise them after the memory test, as though the future was reaching back to help them. Eight of the nine studies cleared the p ≤ 0.05 bar. It had passed peer review at the top of the discipline.
What happened next is the reason Research methods is a whole paper of its own. When Stuart Ritchie, Richard Wiseman and Chris French ran careful replications and found nothing, the same journal initially refused to publish them — replications, editors said, were not novel. That refusal, and the wider discovery that only about a third of landmark psychology findings could be reproduced (the Open Science Collaboration replicated just 36 of 100 studies in 2015), triggered the 'replication crisis' that reshaped how the subject polices itself. Every idea in this section — falsifiability, replicability, peer review, significance testing, the Type I error — is the machinery that decided whether Bem had found the future or had simply been fooled by chance and flexible analysis. Research methods is not the boring bit of the course. It is the part that tells you which of the other sixteen topics you can actually believe.
ModelThe experiment — the only method that touches cause
An experiment manipulates an independent variable (IV) and measures a dependent variable (DV) while holding everything else constant, which is why it is the only design that can claim cause and effect. Four flavours, distinguished by how the IV arises and where it happens: a laboratory experiment manipulates the IV in a controlled setting (high control and replicability, but artificiality and demand characteristics); a field experiment manipulates the IV in a real-world setting (more realistic behaviour, less control, consent problems); a natural experiment measures a DV where the IV varies on its own and the researcher never manipulates it (the Romanian orphan adoptions are the classic case); and a quasi-experiment uses an IV that is a fixed feature of the participants — age, sex, a diagnosis — so people cannot be randomly allocated to it.
The IV/DV must be operationalised — defined in measurable terms — and the study protected from extraneous variables (nuisance variables that add noise) and, worse, confounding variables (which vary systematically with the IV and offer a rival explanation). Three experimental designs handle who does what: independent groups (different people per condition — no order effects but participant variables intrude, so use random allocation), repeated measures (same people in every condition — controls participant variables but creates order effects, countered by counterbalancing), and matched pairs (different but matched people — the compromise). Control techniques are the toolkit: random allocation, counterbalancing (the ABBA order that spreads practice and fatigue effects), randomisation of stimuli, and standardisation so every participant meets the identical procedure.
MechanismEverything that isn't an experiment
Not every question can manipulate an IV, and the non-experimental methods each buy insight at a cost. Observational techniques watch behaviour directly and come in pairs: naturalistic versus controlled (real setting versus arranged one), covert versus overt (participants unaware versus aware), and participant versus non-participant (observer joins in or stays outside). Good observations need an observational design: behavioural categories that operationalise fuzzy behaviours into countable units, plus a sampling rule — event sampling (tally each occurrence) or time sampling (record at fixed intervals).
Self-report methods ask people directly: questionnaires and interviews, structured (fixed, replicable) or unstructured (free, rich, harder to analyse). Correlations measure the strength and direction of a relationship between two co-variables and, crucially, cannot establish cause — the defining difference from an experiment, because no variable is manipulated and a third variable may drive both. Content analysis turns communication (adverts, tweets, diaries) into data by coding it into categories for quantitative counts, while thematic analysis draws out recurrent qualitative themes. And the case study — an in-depth investigation of one person or group, from HM's amnesia to Freud's Little Hans — trades generalisability for extraordinary depth on the rare and the unrepeatable.
ModelDesigning a study without wrecking it
Before data collection, a study is a series of design decisions that examiners love to test with 'design a study' scenarios. It starts with an aim (the general purpose) sharpened into a hypothesis (a precise, testable, operationalised prediction). Hypotheses are directional (one-tailed — predicts which way the effect will go, justified when prior research points a direction) or non-directional (two-tailed — predicts a difference or relationship without saying which way), and every study also carries a null hypothesis of no effect.
Next, sampling from the target population: random (everyone has an equal chance — unbiased but often unrepresentative and hard to get), systematic (every nth person), stratified (subgroups represented in proportion — the most representative), opportunity (whoever is available — convenient but biased), and volunteer (self-selected via advert — reaches the motivated but invites volunteer bias). Every technique is a trade between representativeness and practicality, and every unrepresentative sample limits generalisation. A pilot study — a small-scale trial run — catches confusing instructions and broken materials before real money is spent. Finally, two threats stalk any study: demand characteristics (participants guessing the aim and playing along or sabotaging) and investigator effects (the researcher's expectations leaking into the results), both curbed by single- and double-blind procedures and standardisation. Overarching all of it is the BPS code of ethics: informed consent, no undue deception, protection from harm, confidentiality and the right to withdraw — managed through consent forms, debriefing, ethics committees and anonymity.
MechanismCan you trust the number? Reliability, validity and science
A finding is only worth as much as its reliability and validity. Reliability is consistency: test-retest (the same measure gives the same result on two occasions) and inter-observer reliability (two observers agree — a correlation of +0.80 or above is the working threshold), improved by operationalising categories, training observers and standardising procedures. Validity is whether you measured what you claimed: face validity (does it look right), concurrent validity (does it agree with an established measure), ecological validity (do findings generalise beyond the setting) and temporal validity (do they still hold over time), improved by control groups, standardisation and blind procedures.
Behind these sit the features of science that separate psychology from common sense: objectivity and the empirical method, replicability, falsifiability (Popper's demand that a theory make itself refutable), theory construction and hypothesis testing, and Kuhn's idea of a shared paradigm that occasionally lurches in a paradigm shift. Peer review is the field's quality gate — independent experts vet work before funding or publication — though it can suppress dissent, favour positive results and protect the status quo (exactly what delayed the Bem replications). And psychological research has real economic implications: attachment work reshaping flexible-working and shared-parental-leave policy, and mental-health treatments cutting the workplace absenteeism that mental illness drives. A full write-up follows the standard report sections — abstract, introduction, method, results, discussion, references — with Harvard-style citation.
DataDescribing the data — and choosing the right average
Once you have numbers, you summarise them, and the summary you choose depends on the level of measurement. Nominal data are frequency counts in categories; ordinal data are ranked but with unequal intervals (a Likert scale); interval data sit on a scale with equal units (reaction time, a standardised IQ score). That level dictates the legitimate measure of central tendency: the mode for nominal, the median for ordinal (robust to outliers), the mean for interval (uses every score but is dragged by extremes). Dispersion is captured crudely by the range and precisely by the standard deviation, which measures the average distance of scores from the mean — a large SD signals a spread-out, variable set.
Data also divide into quantitative versus qualitative, and primary (collected first-hand) versus secondary (reusing others' data, of which a meta-analysis — pooling many studies into one combined effect — is the powerful special case). You display quantitative data honestly in tables, bar charts (for discrete categories, bars separated), histograms (for continuous data, bars touching) and scattergrams (for correlations). The shape matters too: a normal distribution is symmetrical with mean, median and mode together, while a skewed distribution has a tail — a hard test produces a positive skew (tail to the right, mean pulled above the mode), an easy one a negative skew.
A researcher records how many words 8 participants recall: 7, 8, 8, 9, 10, 11, 12, 31. Mean = (7+8+8+9+10+11+12+31) ÷ 8 = 96 ÷ 8 = 12. But that single score of 31 has dragged the mean above every value except the outlier itself — 7 of 8 people scored below the 'average'. The median is the midpoint of the ordered set, (9+10) ÷ 2 = 9.5, which describes the typical participant far better; the mode is 8. This is the whole point of the skew-and-average topic in one dataset: the outlier makes the distribution positively skewed, so quoting the mean alone would mislead. A full-mark answer states the median and justifies it — 'the median is preferable here because the extreme score of 31 distorts the mean' — because the justification, not the arithmetic, is where the marks hide.
DataThe inferential leap — significance and the sign test
Descriptive statistics summarise your sample; inferential statistics ask whether the pattern is real or the sort of thing chance would throw up anyway. You set a significance level — conventionally p ≤ 0.05, a 5% risk of being wrong — compare your calculated statistic against a critical value from a table, and either reject or retain the null. Two ways to get it wrong: a Type I error is a false positive (rejecting a true null, claiming an effect that isn't there — the risk you raise with a lax level like 0.10, and the likeliest verdict on Bem's precognition), and a Type II error is a false negative (retaining a false null, missing a real effect — the risk of an over-strict level like 0.01). The 0.05 convention is a deliberate balance between the two.
Which test you run depends on three questions — difference or correlation, related or unrelated design, and level of measurement — captured by the standard grid (Chi-Squared, Sign test, Mann-Whitney, Wilcoxon, related and unrelated t-tests, Spearman's rho, Pearson's r). The one you must be able to calculate is the sign test: for a test of difference, a related design and nominal data.
Twelve participants rate their wellbeing before and after a mindfulness course. For each, mark a plus if they improved, a minus if they worsened, and drop anyone who did not change. Suppose 2 show no change (discarded, so N = 10), 8 improve (+) and 2 worsen (−). The calculated statistic S is the count of the less frequent sign: S = 2. From a sign-test table, the critical value for a non-directional test at p ≤ 0.05 with N = 10 is 1. The decision rule for the sign test is that S must be equal to or less than the critical value to be significant. Here S = 2, which is greater than 1, so the result is not significant — we retain the null hypothesis. That surprises students: 8 of 10 improving feels convincing. But the sign test throws away the size of each change and only counts direction, which makes it conservative — had 9 of 10 improved (S = 1), it would have cleared the bar. State the test, justify it (difference, related design, nominal data), show S, compare to the critical value, and interpret: that five-step chain is the full-mark method.
VocabularyKey terms the mark scheme pays for
TrapsMisconceptions that cost marks
ExamWhat examiners want
Research methods is examined right across the specification, but its home is Paper 2, and it is unusual because it leans on AO2 (application) far more than AO1. Most questions hand you an unfamiliar study — a stem describing a made-up piece of research — and ask you to identify the design, write an operationalised hypothesis, spot a confounding variable, or suggest an improvement. Never answer these in the abstract: quote the details of the scenario back (the specific IV, the specific sample) so the examiner can see you have applied, not recited. 'Design a study' questions are marked on feasibility and detail — name the sampling method, the design, the controls and how you would measure the DV.
This section also carries the bulk of the 10% of A-level marks reserved for maths skills at GCSE Level 2 standard, so drill the calculations that recur: percentages, ratios, standard deviation, significant figures, and above all the sign test, which is the one inferential test you can be required to compute by hand. For any statistics question, show the working — state the test and justify it from the design and level of measurement, give the calculated value, compare it to the critical value, and finish with a plain-English decision about the null hypothesis; the interpretation sentence is the mark candidates most often drop. Evaluation questions (AO3) reward the reliability/validity vocabulary used precisely: not 'this lacks validity' but 'the artificial lab task lowers ecological validity because the behaviour may not generalise to everyday memory'. Precision, applied to the scenario in front of you, is what lifts this paper.