HookTwo identical mice, one yellow and sick, one brown and well
In 2003, two researchers at Duke University in North Carolina, Randy Jirtle and Robert Waterland, took a strain of mouse that is normally fat and yellow and carries a high risk of diabetes and cancer, and turned its offspring into slim, brown, healthy animals — without altering a single letter of their DNA. All they changed was the mothers' diet during pregnancy, feeding them extra methyl donors: folic acid, vitamin B12 and choline. Those methyl groups settled onto the DNA around the agouti gene and switched it off. The pups were genetically identical to the yellow, sickly controls, yet they looked and lived completely differently. Same genome, different phenotype — decided not by the sequence of the genes but by which genes were allowed to speak.
That experiment is this whole section in miniature. Every cell in your body carries the same genome, yet a neurone and a white blood cell could hardly be more different, because gene expression is controlled — by transcription factors, by chemical tags that silence or wake genes, and by small RNA molecules that intercept the message before it is ever read. This section opens with what happens when the base sequence itself mutates, works through how identical genomes are switched to build specialised cells and how that control breaks down in cancer, and closes with the technology — sequencing, cutting, copying, probing and fingerprinting DNA — that now lets us read and rewrite the instructions. It is the most modern biology on the specification, and the part where exact, directional detail earns the marks.
ModelHow base-sequence mutations alter proteins
A gene mutation is a change in the base sequence of DNA. The commonest types are substitution (one base swapped for another), deletion (a base removed) and insertion or addition (a base added); bases can also be duplicated, inverted or translocated. What the exam cares about is how each type feeds through to the protein, and that turns on the fact that the genetic code is degenerate — most amino acids are coded by more than one triplet.
Because of that degeneracy, a substitution changes at most one codon and may not change the amino acid at all (a silent mutation). If it does change one amino acid it is a missense mutation; if it creates a stop codon it is a nonsense mutation that cuts the protein short. A deletion or insertion is far more destructive, because it causes a frameshift: every codon downstream is read in the wrong groups of three, so nearly the entire amino acid sequence after the change is wrong. A changed primary structure means hydrogen, ionic and disulfide bonds form in different places, so the tertiary structure — and therefore the function, such as an enzyme's active site — can be lost. Mutagenic agents such as ultraviolet light, ionising radiation and certain chemicals raise the mutation rate above its low spontaneous background.
Two mutations, two very different outcomes. Sickle-cell anaemia comes from a single substitution in the gene for the \(\beta\)-globin chain: the sixth codon of the mRNA changes from \(\text{GAG}\) to \(\text{GUG}\), so glutamic acid is replaced by valine. One amino acid — but valine is hydrophobic, the haemoglobin molecules stick together when oxygen is low, and the red cells sickle. Now contrast a frameshift. Read the sequence \(\text{AUG-CAU-CAU-CAU}\) as start-His-His-His. Delete the first base of the second codon and the ribosome now reads \(\text{AUG-AUC-AUC-AU...}\) — Met-Ile-Ile — every triplet after the deletion is different and the protein is unrecognisable. A single deleted base can wreck a protein that a single substitution would barely dent: that is why frameshifts are the more serious class.
ModelStem cells, totipotency and potency
Every specialised cell arises from a stem cell: an unspecialised cell that can keep dividing and can differentiate. What a stem cell is able to become defines its potency. Totipotent cells can become any cell type and the extra-embryonic tissue of the placenta, so they can form a whole organism — only the zygote and the cells of the very early embryo are truly totipotent (in plants, many mature cells stay totipotent, which is why a cutting can grow into a whole new plant). Pluripotent cells, found in embryos, can become any of the body's cell types but not a whole organism. Multipotent cells, such as the adult stem cells in bone marrow, form only a limited range — the various blood cells. Unipotent cells make just one type, as cardiomyocytes are produced from unipotent stem cells in the heart.
Differentiation happens because a cell expresses only part of its genome — the same DNA, but different genes switched on. The development that most excites medicine is the induced pluripotent stem cell (iPS cell): an ordinary adult body cell reprogrammed back to pluripotency by switching on specific transcription-factor genes. Because iPS cells can be made from a patient's own tissue, they sidestep both the immune-rejection problem and much of the ethical objection to using embryos, and they hold out treatments for conditions from spinal-cord injury and type-1 diabetes to age-related macular degeneration.
MechanismTranscription factors, epigenetics and RNA interference
Beyond the embryo, expression is tuned continuously by three mechanisms AQA wants in detail. First, transcription factors: proteins that move into the nucleus and bind to a specific base sequence, usually the promoter region in front of a gene, switching transcription on (by helping RNA polymerase bind) or off. Oestrogen shows how a hormone taps into this — being lipid-soluble it diffuses through the cell membrane and binds to a receptor that is itself a transcription factor, changing the receptor's shape so it can bind to DNA and stimulate transcription of specific genes.
Second, epigenetics: heritable changes in gene expression that do not change the base sequence, driven by chemical tags on the DNA and its histone proteins. Increased methylation of the DNA (methyl groups added to cytosine bases in the promoter) stops transcription factors binding and switches the gene off. The acetylation of histones works the other way: adding acetyl groups reduces the positive charge on the histones so they grip the negatively charged DNA less tightly, the chromatin loosens, and transcription increases; removing them condenses the chromatin and silences the gene. Because these tags respond to the environment and can be inherited, they explain the agouti mice of the introduction.
Third, RNA interference: small double-stranded RNA molecules (siRNA and miRNA) are cut into short lengths, one strand of which pairs with a complementary mRNA and targets it for breakdown or blocks its translation — so the gene's message is silenced after it has been transcribed.
CaseGene expression, mutation and tumours
Cancer is a disease of uncontrolled cell division, and it traces straight back to the control system in the previous block failing. Two gene classes matter. Proto-oncogenes normally stimulate division by coding for growth factors and their receptors; a mutation can turn one into an oncogene that is permanently switched on — a receptor that fires without a growth factor, say — so the cell divides relentlessly. Tumour suppressor genes do the opposite job, slowing division and triggering apoptosis in damaged cells; the best known is p53. A mutation that inactivates a tumour suppressor gene takes the brake off, and division runs on.
Crucially, expression can be lost with no change to the base sequence at all: hypermethylation of a tumour suppressor gene's promoter silences it epigenetically, while hypomethylation can activate an oncogene. An increased oestrogen concentration is linked to some breast cancers because oestrogen can promote transcription of genes that drive cell division in breast tissue. The resulting tumour is benign if it stays localised, encapsulated and slow-growing, or malignant if it invades surrounding tissue and sheds cells that spread in the blood or lymph — metastasis — to seed secondary tumours elsewhere. Distinguishing the two, and stating the oncogene and tumour-suppressor mechanisms as mirror images, is the marks-heavy skill here.
ModelUsing genome and proteome projects
Sequencing an organism's entire genome — its complete base sequence — lets biologists work towards its proteome, the full set of proteins it can make. In simpler organisms such as bacteria and viruses the step from genome to proteome is fairly direct, because their DNA has few non-coding stretches and no introns: sequence the genes and you can largely read off the proteins. That is why sequencing a pathogen is so valuable — it identifies the proteins and antigens on its surface, the raw material for vaccines and new drug targets.
In complex organisms the link is far harder to make. Eukaryotic DNA is full of non-coding regions — introns within genes and regulatory sequences between them — and a single gene can yield several different proteins depending on how its exons are spliced and later modified. So knowing the genome does not immediately reveal the proteome. Comparing genomes between species also maps evolutionary relationships, and, increasingly, sequencing an individual underpins personalised medicine, in which a person's own genome guides which drugs are likely to work and which to avoid.
MechanismRecombinant DNA technology
Recombinant DNA technology moves a gene from one organism into another, and it works only because the genetic code is universal and transcription and translation are essentially the same in all organisms — so a human gene placed in a bacterium is read and made into the very same protein, giving a transformed organism. The gene (a DNA fragment) can be obtained three ways: using reverse transcriptase to build a complementary DNA (cDNA) copy from an mRNA template (useful because mRNA is abundant in cells that make a lot of the protein, and the cDNA carries no introns); cutting it out with a restriction endonuclease, an enzyme that cuts DNA at a specific palindromic recognition sequence and can leave staggered sticky ends; or building it base by base with a gene machine.
The fragment is then amplified. In vivo, it is inserted into a vector — usually a plasmid cut with the same restriction enzyme so the sticky ends are complementary — joined with DNA ligase, and the recombinant plasmid is taken up by host bacteria (transformation); marker genes for antibiotic resistance or fluorescence then reveal which cells took it up. In vitro, the polymerase chain reaction (PCR) copies DNA without cells, cycling through 95°C to separate the strands, about 55°C for primers to anneal and 72°C for a heat-stable DNA polymerase to extend them. Each cycle doubles the DNA, and the applications — insulin from bacteria, GM crops, gene therapy — bring economic and ethical questions the exam expects you to weigh.
PCR doubles the number of DNA molecules every cycle, so from a single template you have \(2^n\) copies after \(n\) cycles. After 30 cycles that is
\[ 2^{30} = 1\,073\,741\,824 \approx 1.07 \times 10^{9} \]
— over a billion copies from one starting molecule, in under an hour. The growth is exponential, not linear: if a forensic sample began with 100 intact copies, after 20 cycles it would hold \(100 \times 2^{20} = 100 \times 1\,048\,576 \approx 1.05 \times 10^{8}\) copies. This is why a trace of DNA left at a crime scene, or a few cells in a diagnostic sample, is enough to work with — and why the number of cycles, far more than the starting amount, sets the final yield.
MechanismLocating alleles with DNA probes and labelling
To find out whether a particular allele is present — a disease allele, a mutated oncogene, an allele affecting how someone responds to a drug — you use a DNA probe: a short, single-stranded piece of DNA whose base sequence is complementary to the target, carrying a label so it can be detected. The label is either radioactive (for example a \(^{32}\text{P}\) tag revealed on photographic film) or, more commonly now, fluorescent (it glows under a specific wavelength of light).
The sample DNA is heated to make it single-stranded, the probe is added, and if the target sequence is present the probe binds to it by complementary base pairing — a process called hybridisation. Detecting the label then tells you the allele is there. Scaling this up gives the DNA microarray: thousands of different probes fixed in a grid on a slide, over which a fluorescently labelled sample is washed; wherever the sample hybridises, that spot fluoresces, so one chip can screen for thousands of alleles at once. This is the workhorse of genetic screening and a foundation of personalised medicine, flagging both inherited-disease risk and the likely response to particular drugs — with the evaluation caveat that a person may not want to know a risk they cannot yet treat.
DataGenetic fingerprinting and its applications
Genetic fingerprinting identifies individuals from the non-coding parts of their genome. Scattered through the introns and between genes are variable number tandem repeats (VNTRs) — short sequences repeated over and over, with the number of repeats varying enormously from person to person, and more so the less closely related they are. Because these regions do not code for proteins, they can vary freely without harm, which is exactly what makes them useful for telling people apart.
The method runs in stages. DNA is extracted and the repeat regions amplified by PCR. The fragments are then separated by gel electrophoresis: loaded into wells in a gel across which a voltage is applied, and because DNA is negatively charged (its phosphate groups), the fragments migrate towards the positive electrode, the smaller fragments travelling further through the gel. The separated fragments are made single-stranded, transferred to a membrane, and tagged with labelled probes so they show up as a pattern of bands — the fingerprint. Matching that pattern band for band underpins forensic identification, paternity testing, and the measurement of genetic variability and relationships within and between populations, as well as some medical diagnosis.
How sure can a court be that two matching fingerprints came from the same person? You multiply the population frequencies of the bands, assuming they are inherited independently. Suppose a profile is built from 15 bands, each of which about one person in four in the population happens to carry. The probability that an unrelated person matches all fifteen is
\[ \left(\frac{1}{4}\right)^{15} = \frac{1}{4^{15}} = \frac{1}{1\,073\,741\,824} \approx 9.3 \times 10^{-10} \]
— roughly one in a billion. That vanishingly small chance is what lets a match be presented as near-certain identification. The honest caveat, and a strong evaluation point, is that the multiplication assumes the bands really are independent and that the suspect is unrelated to the true source: close relatives share far more bands, so the one-in-a-billion figure does not apply to a brother.
VocabularyKey terms the mark scheme pays for
TrapsMisconceptions that cost marks
ExamWhat examiners want
This section is won on precision and direction. Examiners want exact mechanism language: a transcription factor binds to a specific base sequence / the promoter; increased methylation prevents transcription; increased acetylation loosens histone binding and increases transcription. Getting the direction backwards — saying methylation switches a gene on — throws away easy marks, so learn the pairs as opposites. State the oncogene and tumour-suppressor mechanisms as mirror images too: an oncogene is a proto-oncogene permanently activated; a tumour suppressor gene is one that has been inactivated or silenced by hypermethylation.
For recombinant DNA, sequence the steps in order and name the enzyme at each — reverse transcriptase, restriction endonuclease (the same one for gene and vector so the sticky ends match), DNA ligase — and be ready to calculate PCR yields as \(2^n\), carrying units and using standard form for the large numbers. The heaviest marks, especially on Paper 3, come from AO3 evaluation of the applications: give balanced, specific arguments on embryonic stem cells, GM organisms, genetic screening, gene therapy and fingerprinting, rather than a vague 'it raises ethical issues'. Interpreting data is AO2 — reading a gel (smaller fragments travel further) or a microarray (a fluorescing spot means the probe hybridised, so the allele is present). Because Paper 3 sets a 25-mark essay, practise linking gene expression to protein structure, cell specialisation and inheritance right across the specification.