Learn · A-Level Maths · Strand Statistics
EDX-A-MATH-S5 · Statistical hypothesis testing

Hypothesis testing — proportions & means.

Written for Edexcel 9MA0 Official specification ↗ Updated 2026.07.05

HookThe lady who could taste the milk

At Rothamsted agricultural research station in Hertfordshire in the 1920s, a scientist named Muriel Bristol claimed she could taste whether the milk or the tea had been poured into the cup first. Her colleague Ronald Fisher did not argue — he designed a test. He prepared eight cups, four milk-first and four tea-first, arranged in random order, and asked her to sort them into the two groups. If she were merely guessing, the probability of getting all eight correct by luck alone is 1 in 70. Fisher's write-up in his 1935 book The Design of Experiments turned a tea-time boast into the template for every hypothesis test since: assume the sceptical position, work out how surprising the data would be if that position were true, and abandon it only when the data are surprising enough.

That is the entire logic of S5. You will learn the language — null and alternative hypotheses, significance level, test statistic, critical region and p-value — built first on the binomial model. You will run a hypothesis test for a proportion using the binomial distribution, and a test for the mean of a Normal distribution with known variance, and — the part students most often fumble — interpret the result in the actual context, in a sentence an examiner will accept. Fisher's insight was that you can never prove a claim true; you can only ask whether the evidence against the sceptic is strong enough to act on. Muriel Bristol, for the record, sorted all eight cups correctly.

ModelThe language of a hypothesis test

Every test starts with two rival claims. The null hypothesis \(H_0\) is the sceptic's default position, and it is always an equality — a specific value of a proportion or a mean, such as \(p=0.3\) or \(\mu=500\). The alternative hypothesis \(H_1\) is what you will conclude if you reject the null, and it points in a direction: one-tailed if you suspect a change one way only (\(p\lt 0.3\) or \(p\gt 0.3\)), two-tailed if any change would interest you (\(p\neq 0.3\)). You must choose the tail from the context before collecting data, never after.

The significance level — commonly 5% or 1% — is the risk of wrongly rejecting a true null that you are prepared to tolerate. The test statistic is the quantity you actually observe: a count of successes, or a sample mean. The critical region is the set of outcomes so extreme that, if they occur, you reject \(H_0\); by construction its probability under \(H_0\) does not exceed the significance level. The p-value is the probability, computed assuming \(H_0\) is true, of a result at least as extreme as the one observed — and the decision rule is simply: reject \(H_0\) if the \(p\)-value is less than or equal to the significance level. The same machinery extends to correlation: given a sample product-moment correlation coefficient \(r\), you can test \(H_0:\rho=0\) (no linear correlation in the population) against \(H_1:\rho\neq 0\) by comparing \(r\) with a critical value from tables.

Worked example

A café claims 30% of its customers prefer a new blend. The manager suspects the true figure is lower. In symbols, the sceptical claim is \(H_0:\ p=0.3\) and the manager's suspicion is the one-tailed alternative \(H_1:\ p\lt 0.3\). The test statistic will be \(X\), the number in a sample who prefer the new blend, modelled under the null as \(X\sim B(n,\,0.3)\). Fixing the direction now, before any data are seen, is what keeps the test honest — deciding the tail after glancing at the result would quietly double your chance of a false alarm.

MechanismTesting a proportion with the binomial

A binomial test asks whether an observed count is too extreme to square with a claimed proportion. The method is a fixed ritual. State \(H_0\) and \(H_1\) and the significance level. Write the distribution of the test statistic assuming \(H_0\) is true: \(X\sim B(n,p_0)\). Then measure how extreme the observation is, either by computing the \(p\)-value — the probability under \(H_0\) of a result at least as far in the direction of \(H_1\) as the one you saw — or by finding the critical region in advance. Compare, decide, and write a conclusion in context.

Two details trip people up. First, 'at least as extreme' must match the tail: for \(H_1:\ p\lt p_0\) you accumulate the lower tail \(P(X\le x)\); for \(p\gt p_0\) the upper tail \(P(X\ge x)\); for a two-tailed test you compare each tail with half the significance level. Second, because the binomial is discrete you usually cannot build a critical region of probability exactly 5% — the true probability of the region, the actual significance level, is the largest achievable value not exceeding 5%, and examiners often ask for it.

Worked example

Test the café's claim at the 5% level. A random sample of 20 customers is taken and only 3 prefer the new blend. With \(H_0:\ p=0.3\), \(H_1:\ p\lt 0.3\), and \(X\sim B(20,\,0.3)\) under the null, we need the lower-tail probability of a result at least as low as 3.

Summing the binomial probabilities, \[P(X\le 3)=P(0)+P(1)+P(2)+P(3)=0.00080+0.00684+0.02785+0.07160=0.1071.\] Since \(0.1071\) is greater than \(0.05\), the result is not in the critical region and we do not reject \(H_0\). (Checking the region directly agrees: \(P(X\le 2)=0.0355\le 0.05\) but \(P(X\le 3)=0.1071\gt 0.05\), so the critical region is \(X\le 2\), and the actual significance level is 3.55%. The observed 3 lies just outside it.) Conclusion in context: there is insufficient evidence at the 5% level to suggest that fewer than 30% of customers prefer the new blend — even though 3 out of 20 is only 15%, that is not yet surprising enough to overturn the claim.

MechanismTesting a mean with the Normal distribution

When the data are measurements from a Normal population whose variance is known, you test the mean. The key fact is how a sample mean behaves: if individual values follow \(X\sim N(\mu,\sigma^2)\), then the mean of a sample of size \(n\) follows \[\bar{X}\sim N\!\left(\mu,\ \frac{\sigma^2}{n}\right),\] a Normal distribution with the same centre but a spread narrowed by a factor of \(\sqrt{n}\). Averages are less variable than individuals, which is exactly why a sample mean is a sharp instrument for detecting a shift in \(\mu\).

The test standardises the observed sample mean against the null value: \[Z=\frac{\bar{x}-\mu_0}{\sigma/\sqrt{n}}.\] You then compare \(Z\) with a critical value — for a one-tailed test at 5% that is \(\pm 1.645\); for two-tailed at 5% it is \(\pm 1.96\); for one-tailed at 1% it is \(\pm 2.326\). If \(Z\) falls beyond the critical value, the sample mean is too far from \(\mu_0\) to be luck, and you reject \(H_0\). Equivalently, convert \(Z\) to a \(p\)-value and compare with the significance level. The conclusion, as always, is a sentence about the real quantity, not about \(Z\).

Worked example

A machine should fill bags to a mean of 500 g, and the fill weight is known to have standard deviation 4 g. A quality manager, suspecting underfilling, weighs a random sample of 16 bags and finds a mean of 498 g. Test at the 5% level whether the mean has fallen.

Hypotheses: \(H_0:\ \mu=500\), \(H_1:\ \mu\lt 500\) (one-tailed). Under \(H_0\), \(\bar{X}\sim N\!\left(500,\ \dfrac{4^2}{16}\right)=N(500,\,1)\), so the standard error is \(\sigma/\sqrt{n}=4/4=1\). The test statistic is \[Z=\frac{498-500}{1}=-2.\] The one-tailed 5% critical value is \(-1.645\). Since \(-2\lt -1.645\), the statistic lies in the critical region, so we reject \(H_0\). Equivalently the \(p\)-value is \(P(Z\lt -2)=0.0228\lt 0.05\). Conclusion in context: there is sufficient evidence at the 5% level to suggest the machine is underfilling — the mean fill weight has dropped below 500 g.

CaseReading the result honestly

The hardest marks in S5 are not the arithmetic; they are the interpretation, because the natural-language reading of a test result is riddled with traps. Rejecting \(H_0\) at the 5% level does not mean there is a 95% chance \(H_1\) is true. It means: if the null were true, data this extreme would arise less than 5% of the time, so we choose to act as though the null is false. The 5% is the Type I error rate — the long-run proportion of true nulls you would wrongly reject if you ran the test forever. It is a property of the procedure, not a probability attached to today's particular conclusion.

Two cautions follow. First, failing to reject \(H_0\) is not proof that \(H_0\) is true — the café test found 'insufficient evidence', which is a statement about the strength of the data, not a verdict that exactly 30% prefer the blend. Absence of evidence is not evidence of absence. Second, the \(p\)-value is emphatically not the probability that the null hypothesis is true; that confusion is the same prosecutor's fallacy that mislabelled the odds in the Sally Clark case — swapping \(P(\text{data}\mid H_0)\) for \(P(H_0\mid\text{data})\). Write your conclusion as 'there is (or is not) sufficient evidence at the stated level to suggest…', tie it to the context, and you will collect the marks that vaguer answers throw away.

Worked example

Return to the café result, \(P(X\le 3)=0.1071\). A weak answer says 'there is an 89% chance at least 30% prefer the blend' — wrong, because \(0.1071\) is \(P(\text{data}\mid H_0)\), not \(P(H_0\mid\text{data})\). A full-mark answer says: 'Since \(0.1071\gt 0.05\), we do not reject \(H_0\); there is insufficient evidence at the 5% level to conclude that fewer than 30% of customers prefer the new blend.' Had the manager instead tested at the 10% level, note the decision would still stand, since \(0.1071\gt 0.10\) — a reminder that the significance level must be fixed in advance, not nudged until the data cross the line.

VocabularyKey terms the mark scheme pays for

Null hypothesis \(H_0\)
The sceptical default being tested, always stated as an equality such as \(p=0.3\) or \(\mu=500\). All probabilities in the test are computed assuming \(H_0\) is true.
Alternative hypothesis \(H_1\)
The claim accepted if \(H_0\) is rejected. One-tailed (\(\lt\) or \(\gt\)) if a direction is suspected; two-tailed (\(\neq\)) if any change matters. Chosen before seeing data.
Significance level
The tolerated probability of wrongly rejecting a true \(H_0\), commonly 5% or 1%. It equals the Type I error rate and must be fixed in advance.
Test statistic
The observed quantity the decision rests on — a binomial count \(X\) or a standardised sample mean \(Z\) — whose distribution under \(H_0\) is known.
Critical region
The set of test-statistic values extreme enough to reject \(H_0\); its probability under \(H_0\) does not exceed the significance level. The boundary value is the critical value.
p-value
The probability, assuming \(H_0\) true, of a result at least as extreme as the one observed. Reject \(H_0\) when the \(p\)-value is less than or equal to the significance level.
Actual significance level
For a discrete binomial test, the true probability of the critical region — the largest achievable value not exceeding the nominal level (e.g. 3.55% for a nominal 5%).
Distribution of the sample mean
For a Normal population, \(\bar{X}\sim N\!\left(\mu,\ \sigma^2/n\right)\); the standard error \(\sigma/\sqrt{n}\) shrinks with sample size, sharpening the test for the mean.

TrapsMisconceptions that cost marks

“The p-value is the probability that the null hypothesis is true.”
Actually: It is \(P(\text{data this extreme}\mid H_0\text{ true})\), not \(P(H_0\mid\text{data})\). Swapping the two is the prosecutor's fallacy from S3; a small \(p\)-value means the data are surprising under \(H_0\), nothing more.
“Failing to reject \(H_0\) proves \(H_0\) is true.”
Actually: It only means the evidence was not strong enough to overturn the default. Absence of evidence is not evidence of absence — the café test showed 'insufficient evidence', not that exactly 30% prefer the blend.
“A 5% significance level means there is a 5% chance the conclusion is wrong.”
Actually: The 5% is the Type I error rate — the long-run rate of falsely rejecting a null that is actually true. It describes the procedure, not the probability that today's specific decision is mistaken.
“You can decide one-tailed versus two-tailed after seeing the data.”
Actually: The tail must be fixed from the context beforehand. Choosing it to suit the result quietly inflates the true error rate, and examiners penalise a tail that was clearly picked to make the test 'work'.

ExamWhat examiners want

Open every hypothesis test with the same three lines before any numbers: state \(H_0\) and \(H_1\) in symbols, name the significance level, and say whether the test is one- or two-tailed. Then write the distribution of the test statistic under \(H_0\) — \(X\sim B(n,p)\) for a proportion, \(\bar{X}\sim N(\mu,\sigma^2/n)\) for a mean — because that line carries a method mark and dictates every calculation that follows.

Make the comparison explicit: either \(p\)-value against significance level, or test statistic against critical value, with the inequality written out. For a binomial test, match the tail to \(H_1\) (lower tail for \(\lt\), upper for \(\gt\), half the level in each tail for \(\neq\)), and be ready to state the actual significance level. For a Normal mean test, compute the standard error \(\sigma/\sqrt{n}\) first and quote the right critical value (1.645 one-tailed, 1.96 two-tailed at 5%). Finish with a conclusion in context — 'there is sufficient/insufficient evidence at the 5% level to suggest…' — and never write 'this proves \(H_1\)'; a test supports a hypothesis, it does not prove it.

Retrieve

Test yourself

Question 1 of 4

Vofti has 16 questions on EDX-A-MATH-S5 — every one hook-first, every one mapped to this section of the Edexcel spec.

Last updated · 2026.08.09 Edexcel A-Level Maths · Spec EDX-A-MATH-S5