HookThe poll that surveyed two million people and still got it catastrophically wrong
In 1936 the American magazine Literary Digest ran the largest opinion poll the world had ever seen. It mailed ballots to ten million people and got back an astonishing 2.4 million replies, then declared with total confidence that Alf Landon would beat Franklin Roosevelt in that year's presidential election by 57% to 43%. Roosevelt won by a landslide — 61% to 37% — one of the most lopsided results in American history. The magazine, humiliated, folded within two years. In the same election a young pollster named George Gallup surveyed just a few thousand people and called the result correctly. The magazine had beaten Gallup on sheer numbers by a factor of nearly a thousand, and still lost. Why?
Because the Literary Digest drew its names from telephone directories, car registration lists and its own subscriber base — and in Depression-era America, the people who owned telephones and cars were disproportionately wealthy, and the wealthy leaned Republican. The sample was enormous and biased; Gallup's was tiny and representative. That is the whole lesson of this section, and it is the opposite of most people's instinct: a sample's value comes from how it is chosen, not how big it is. You will pin down the difference between a population and a sample, use a sample to make an informal inference about the whole, and learn the standard sampling techniques — simple random, systematic, stratified and opportunity — well enough to choose one and to tear a bad one apart. Edexcel gives you a real population to practise on: the Large Data Set of Met Office weather records, from which every sampling question quietly draws.
ModelPopulation and sample — and why we rarely measure everything
The population is the entire collection of items you want to know about: every voter in the UK, every light bulb from a factory's production run, every daily rainfall reading at Heathrow in 2015. Measuring all of it is a census. A census is exhaustive and, when you can do it, unbeatable — but it is often impossible, ruinously expensive, or self-defeating. You cannot census the lifetime of every light bulb by testing it, because a lifetime test destroys the bulb; you cannot realistically ask all 47 million UK voters their intention before an election.
So instead we take a sample: a manageable subset of the population, chosen so that measuring it tells us something trustworthy about the whole. The list you actually draw the sample from is the sampling frame — a register of members, each ideally identifiable and numbered. A quantity calculated from the whole population, such as its true mean, is a parameter; the same quantity calculated from a sample is a statistic, and it is our estimate of the parameter. The entire discipline of sampling is the art of making that estimate reliable while measuring as little as possible. The trade-off is permanent: a census gives accuracy at great cost, a sample gives speed and economy at the price of some uncertainty.
MechanismSimple random sampling — every member an equal chance
Simple random sampling is the gold standard the others are judged against. It means every member of the population has an equal chance of being chosen, and — just as importantly — every possible sample of the required size is equally likely. In practice you number the sampling frame \(1, 2, 3, \ldots, N\) and then use a source of randomness — a random number generator, or the random-number function on a calculator — to pick which numbers make the sample, ignoring repeats and any number beyond \(N\).
Its strength is that it is free of bias by construction: no member is favoured, so the sample is representative in the long run and the mathematics of estimation applies cleanly. Its weaknesses are practical. You need a complete, numbered sampling frame, which for large or vague populations may not exist. And random selection can, by bad luck, scatter your chosen members across a huge area or miss an important subgroup entirely — pick 20 UK schools at random and you might get none from Wales. Those drawbacks are exactly what the other techniques are designed to fix.
A company has \(500\) employees and wants a simple random sample of \(5\) for a survey. Number the staff \(001\) to \(500\). Reading three-digit groups from a random number source gives, say, \(348,\ 501,\ 072,\ 348,\ 415,\ 006,\ 289\).
Work along the list applying two rules. Discard \(501\) because it exceeds \(500\) (no such employee). Discard the second \(348\) because that person is already in the sample — you sample without replacement. That leaves the sample as employees numbered \(348,\ 072,\ 415,\ 006,\ 289\). Every one of the 500 had an equal \(\dfrac{5}{500}=\dfrac{1}{100}\) chance of selection, and no property of an employee — department, age, seniority — influenced whether they were picked, which is precisely what makes the sample unbiased.
MechanismSystematic, stratified and opportunity sampling
Three more techniques round out the toolkit. In systematic sampling you order the frame and take every \(k\)th member after a random start, where \(k=\dfrac{\text{population size}}{\text{sample size}}\); it is quick and spreads the sample evenly through the list, but it can go badly wrong if the list has a hidden repeating pattern that lines up with \(k\). In stratified sampling you split the population into non-overlapping groups (strata) — year groups, regions, age bands — and take a simple random sample from each, sized in proportion to the stratum. This guarantees every subgroup is represented in the right mix, which simple random sampling cannot promise; the cost is that you must know the strata sizes in advance.
At the other end of the rigour scale sits opportunity sampling (also called convenience sampling): you simply take whoever is available — the first 30 people who walk past, the students in your own class. It is fast, cheap and needs no sampling frame, which is its entire appeal. But it is wide open to bias, because 'who happens to be available' is almost never representative of the population — survey shoppers at 11am on a Tuesday and you systematically miss everyone at work. A close relative, quota sampling, is opportunity sampling with targets (interview 10 men and 10 women): better balanced, but the choice of which individuals is still left to the interviewer, so bias creeps back in.
A school has \(1200\) students: \(600\) in Key Stage 3, \(400\) in Key Stage 4 and \(200\) in the sixth form. A stratified sample of \(60\) is wanted. How many come from each stage?
The sampling fraction is the same for every stratum: \(\dfrac{60}{1200}=\dfrac{1}{20}\), so take one in twenty from each. Key Stage 3 contributes \(\dfrac{600}{1200}\times 60=30\) students; Key Stage 4 contributes \(\dfrac{400}{1200}\times 60=20\); the sixth form contributes \(\dfrac{200}{1200}\times 60=10\). Check the total: \(30+20+10=60\), as required. Each stratum's share of the sample matches its share of the school, so no year group is over- or under-represented — then you would take a simple random sample of the right size within each stage to fill the quota fairly.
CaseMaking inferences and critiquing a method
The point of a sample is inference: using the sample statistic as an estimate of the population parameter, while staying honest about the uncertainty. If a simple random sample of 200 Met Office daily readings from the Large Data Set has a mean maximum temperature of 14.2°C, that 14.2°C is our best estimate of the true mean for the whole population of days — but it is an estimate, and a different random sample would give a slightly different number. Larger, well-chosen samples give estimates that vary less; badly chosen samples give estimates that are confidently wrong, which is worse than being uncertain.
That is why the examinable skill is critique: given a described method, identify the source of bias and the population it fails to represent. The Literary Digest is the master class. Its sampling frame — telephone and car owners — excluded the poorer majority, so it was unrepresentative before a single ballot was returned. Worse, only 24% of the ballots came back, and people who feel strongly are likelier to reply, a second distortion called non-response bias. Two million responses could not rescue a frame that had already written the poor out of the population. When you critique a method, name the specific group it misses and the specific direction the result is skewed — 'this opportunity sample of gym-goers will overstate how much the town exercises' earns far more than a vague 'it might be biased'.
A researcher wants the average number of hours UK adults exercise per week and stands outside a gym at 8am on a Monday, asking the first 50 people who leave. Critique the method and suggest a better one.
This is an opportunity sample, and it is biased in a predictable direction. The population is 'all UK adults', but the sampling location — a gym doorway — guarantees that everyone surveyed already exercises, so the sample systematically overstates the population's activity; the sedentary majority has zero chance of selection, breaking the equal-chance principle entirely. The time (8am Monday) compounds it, catching committed early-morning gym-goers rather than casual users. A better design names a proper frame and randomises: take a stratified sample of the electoral register by region and age band, then a simple random sample within each stratum, so that non-exercisers are represented in their true proportion. State the bias, its direction, and the fix — that three-part answer is what the mark scheme rewards.
VocabularyKey terms the mark scheme pays for
TrapsMisconceptions that cost marks
ExamWhat examiners want
When asked to describe a sampling method, give the mechanism in steps, not just its name: for simple random sampling say 'number the population \(1\) to \(N\), generate random numbers, ignore repeats and any above \(N\)'; for systematic say how you compute \(k\) and that the start is random. Vague descriptions lose the method marks even when the label is right.
Critique questions are worth the most and follow a fixed shape: name the technique, identify the specific group the frame or method under- or over-represents, and state the direction the estimate is skewed — then propose a concrete improvement. 'It is biased' scores little; 'the opportunity sample at a gym overstates exercise because non-exercisers cannot be selected, so use a stratified sample of the electoral register instead' scores fully. Always give at least one advantage and one disadvantage when a method is compared, and tie both to the context in the question rather than reciting generic textbook points. Where the Large Data Set is involved, remember it is itself a sample of Met Office records, so the same representativeness questions apply to it.