Mutations and protein structure
Alteration of the sequence of bases in DNA can alter the structure of proteins and the regulation of gene expression 3.8.1
- Transcriptional control: transcription factors bind promoters/enhancers to alter a gene's transcription rate. Post-transcriptional control: alternative splicing (different exon combinations from one pre-mRNA) or changes to mRNA stability, acting after transcription.
- A coding-sequence mutation can change the protein itself (its amino acid sequence, folding and function) — the mutation types (substitution, insertion/deletion) already covered in Genetic information and variation.
- A control-region mutation (in a promoter or enhancer) instead changes how MUCH of an otherwise perfectly normal protein is produced, by altering transcription factor binding, without changing the protein's own sequence at all.
- Both mutation categories can have a real phenotypic effect, since phenotype depends on having the RIGHT AMOUNT of a functional protein present at the right time, not merely on having a protein that is theoretically capable of working — a perfectly normal protein made in the wrong quantity (or at the wrong time, or in the wrong cell type) can be just as phenotypically significant as an abnormal protein.
- This is the key distinction this subtopic adds beyond simple 'mutation changes protein structure' reasoning: gene expression itself, not just protein structure, is something a mutation can disrupt — and a control-region mutation doing so may leave the coding sequence, and so the protein structure prediction from it, completely unaffected.
Worked examples
Worked example 3.8.1 · 5 marks
Two mutations affecting the same gene are compared.
Mutation P lies within the gene's promoter and reduces the binding affinity of a transcription factor required for transcription.
Mutation Q lies within the gene's coding sequence and substitutes one amino acid for another at a position within the enzyme's active site.
Explain how the effect of each mutation on the resulting enzyme's activity would differ, and explain why both could still produce a similar reduction in a cell's overall rate of the reaction this enzyme catalyses.
Show worked solution
Mutation P does not alter the coding sequence, so enzyme molecules produced are normal — its effect is on quantity, via reduced transcription.
Mutation Q changes the active site's shape, reducing catalytic activity per molecule, regardless of quantity produced.
Overall rate depends on the product of enzyme quantity and activity per molecule; P reduces the first factor, Q reduces the second, and a large enough reduction in either can produce a comparable fall in the product.
Mark scheme · 5 marks
- States mutation P affects enzyme quantity, not structure 2 marks
- States mutation Q affects the active site's shape/activity per molecule, not quantity 1 mark
- States overall rate depends on the product of quantity and activity per molecule 1 mark
- Explains a sufficient reduction in either factor alone can produce a similar fall in overall rate 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.1 · 4 marks
A gene's promoter region becomes heavily methylated in a particular cell type, substantially reducing its transcription.
A separate cell of the same type, from a different individual, has a mutation deleting a large part of the same gene's coding sequence entirely.
Explain one key difference between these two changes in terms of whether the DNA base sequence itself has been altered, and explain why only one of the two changes could, in principle, be reversed without any further change to the DNA.
Show worked solution
Methylation does not alter the DNA base sequence — an epigenetic change.
The deletion is a genuine mutation, permanently altering the sequence.
Because methylation leaves the sequence intact, it is in principle reversible by demethylation; the deletion cannot be reversed this way, since the sequence itself has been permanently altered.
Mark scheme · 4 marks
- States methylation does not alter the DNA base sequence 1 mark
- States the deletion is a genuine, permanent mutation to the sequence 1 mark
- Explains methylation is reversible (demethylation) without further DNA change 1 mark
- Explains the deletion cannot be reversed without repairing/replacing the missing sequence 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Control of gene expression
Gene expression: from DNA to protein 3.8.2.1
- Most nucleated body cells in a multicellular organism share an essentially identical genome (the same complete set of genes), yet express different subsets of those genes — this is the central fact this whole subtopic explains: different cell types are different not because they contain different genes, but because they express different genes.
- Different sets of expressed genes produce different proteins, and different proteins are what actually give a specialised cell its particular structure and function — a muscle cell and a nerve cell differ because of which genes they express, not which genes they possess.
- A large proportion of the genome is non-coding DNA — much of it regulatory sequence controlling whether and how strongly nearby genes are transcribed, directly relevant to how gene expression is controlled at all.
Stem cells and cell potency 3.8.2.1
- Potency: the range of different cell types a single cell is capable of producing through division and differentiation.
- Totipotent: can produce all body cell types AND extra-embryonic tissue types — the zygote and cells of the earliest embryo.
- Pluripotent: can produce all body cell types, but not extra-embryonic tissue — cells of the early embryo's inner cell mass.
- Multipotent: can produce several related cell types — e.g. blood stem cells, which can form various blood cell types but not unrelated cell types.
- Unipotent: can produce only one specific cell type — e.g. cells of the cardiomyocyte (heart muscle) lineage.
- Potency generally decreases through this sequence (totipotent → pluripotent → multipotent → unipotent) as development proceeds — cells become progressively more specialised and lose the ability to form as many alternative cell types, as more and more genes become permanently switched off.
- Induced pluripotent stem (iPS) cells reverse this apparent one-way process artificially: an adult somatic cell has selected transcription factors introduced, reprogramming its gene expression back to a pluripotent state, capable (with directed differentiation) of forming a wide range of specialised cell types again.
- iPS cells have real potential uses — cell replacement therapy, disease modelling, and drug testing — but come with real evaluative concerns worth naming directly: control over differentiation is not perfect, there is a tumour risk (an incompletely differentiated or reprogrammed cell dividing uncontrolled), immune rejection is possible depending on the cell source, and (for genuinely embryonic-derived pluripotent cells specifically, not iPS cells) ethical concerns around embryo use.
- A specific practical limitation this raises: because iPS cells are derived from a patient's OWN adult cells (not embryos), they avoid the embryo-use ethical objection specifically, while still carrying the differentiation-control and tumour-risk concerns shared with other stem cell approaches.
Transcription factors and gene expression 3.8.2.2
- Transcription factor: a protein that binds to a specific DNA sequence (e.g. in a gene's promoter or a nearby response element) and thereby stimulates or inhibits that gene's transcription.
- Oestrogen (a steroid hormone) example: oestrogen, being lipid-soluble, diffuses directly through the cell-surface membrane and nuclear envelope → binds an oestrogen receptor, forming a receptor dimer that itself acts as a transcription factor → the oestrogen-bound receptor binds a specific response element near a target gene's promoter → stimulates RNA polymerase binding and transcription of that target gene.
- Some transcription factors, like the oestrogen receptor in this example, are activated by a signal entering the cell from outside (here, a steroid hormone diffusing directly across the membrane, rather than needing a cell-surface receptor and second-messenger relay); others are already present and active within the nucleus without needing any such external activating signal.
- The oestrogen receptor itself shuttles between cytoplasm and nucleus, and a substantial proportion of receptors are already nuclear even before oestrogen binds — hormone binding activates the receptor's transcription-factor function, rather than being solely what gets the receptor into the nucleus in the first place.
- This mechanism generalises well beyond oestrogen: any lipid-soluble signalling molecule capable of crossing the cell membrane directly can potentially act via an analogous receptor-as-transcription-factor route, distinct from the second-messenger cascade needed for a hormone (like glucagon, covered under Responses and coordination) that cannot cross the membrane itself.
Epigenetic control of gene expression 3.8.2.2
- Epigenetic control: a heritable change in gene expression that does not alter the underlying DNA base sequence at all — a genuinely separate category from mutation, which does change the sequence.
- DNA methylation: adding methyl groups (typically to cytosine bases) in a gene's promoter region — higher promoter methylation reduces access and commonly suppresses transcription; lower promoter methylation favours transcription factor binding and transcription.
- Histone acetylation: adding acetyl groups to histone protein tails — more acetylation loosens chromatin structure, giving more accessible DNA and favouring transcription; less acetylation gives more compact chromatin and reduces transcription.
- Because neither epigenetic mark changes the underlying DNA base sequence, both are, in principle, reversible — a genuinely different kind of change from a mutation, which (short of a further, independent mutation) is permanent.
- Gene-activity states set by these marks can pass through cell division without any change to the DNA sequence at all — this is precisely how a specialised cell type's particular pattern of gene expression is maintained stably across many rounds of mitosis, even though every daughter cell still carries the full, unchanged genome.
- Environmental factors (diet, stress, toxin exposure) can influence these marks, and in some cases such changes can even be inherited by offspring without any change to the underlying DNA sequence (transgenerational epigenetic inheritance) — an active, still-developing area of research.
- This is precisely why identical twins, despite starting with identical DNA, can develop measurably different epigenetic profiles (and so different gene expression patterns) over a lifetime, as they experience different environments.
- DNA-methylation and histone-modifying enzymes are themselves treatment targets in some cancers — since abnormal methylation patterns (e.g. reduced methylation activating a growth-promoting gene, or increased methylation silencing a tumour-suppressor gene) can directly contribute to cancer development, covered further in the next section.
RNA interference 3.8.2.2
- RNA interference (RNAi): a mechanism in which a small, single-stranded RNA guides a protein complex to a complementary target mRNA, leading to its cleavage and degradation — reducing translation of that specific gene's product, a form of post-transcriptional control.
- siRNA pathway: double-stranded RNA is cleaved by the enzyme Dicer into short siRNA fragments → a guide strand from the siRNA is loaded into a protein complex (RISC, RNA-induced silencing complex) → the guide strand base-pairs with a complementary region of the target mRNA → RISC cleaves the target mRNA, which is then degraded → less target mRNA is available, so less of that specific protein is translated.
- RNA interference acts AFTER transcription (on already-made mRNA), which is exactly why it is classed as post-transcriptional control, distinct from the transcriptional control exerted by transcription factors binding DNA directly.
- Because the guide strand's sequence determines which mRNA is targeted (through ordinary complementary base pairing), RNAi is highly specific — it can, in principle, be directed against essentially any chosen target gene's mRNA, which is part of why it has become a valuable laboratory tool for deliberately silencing a specific gene's expression to study its function.
- The siRNA cleavage pathway described here is one specific, well-characterised RNAi route; other related small-RNA pathways exist that can repress translation by other means (e.g. blocking translation without necessarily degrading the mRNA outright), though the core principle of a small guide RNA directing sequence-specific repression is shared.
Mutations, oncogenes and cancer 3.8.2.3
- Proto-oncogene: a normal gene promoting appropriately controlled cell growth and division; an activating mutation or excessive expression can convert it into an oncogene, driving excessive, inappropriate cell division — a stuck accelerator.
- Tumour-suppressor gene: normally restrains the cell cycle, or promotes DNA repair or apoptosis of damaged cells; an inactivating mutation or gene silencing removes this restraint — failed brakes.
- Benign tumour: stays localised (often encapsulated), does not invade surrounding tissue or spread elsewhere, though it may still physically harm surrounding tissue by its size or position. Malignant tumour: invades surrounding tissue directly and can spread via blood or lymph vessels to form secondary tumours elsewhere in the body (metastasis).
- Epigenetic routes to cancer: reduced promoter methylation can increase expression of a growth-promoting gene (mimicking an activating mutation's effect without changing the DNA sequence); promoter hypermethylation can silence a tumour-suppressor gene (mimicking an inactivating mutation's effect) — epigenetic changes and genetic mutations can produce functionally equivalent outcomes by entirely different mechanisms.
- Tumour development usually requires multiple genetic and epigenetic changes accumulating together in the same cell line, not a single mutation alone.
- Because both an accelerator-type change (oncogene activation) and a brakes-failure-type change (tumour-suppressor inactivation) typically need to occur, and multiple independent genes perform each role, a single mutation is rarely sufficient on its own to cause cancer — directly explaining why cancer risk rises with age, since more cell divisions over a longer lifetime give more opportunities for a sufficient combination of changes to accumulate in one cell line.
- Oestrogen and some breast cancers: in oestrogen-receptor-positive cells, increased oestrogen signalling can stimulate transcription of growth-promoting genes and cell division — directly connecting the transcription-factor mechanism covered earlier to a real clinical example, and explaining why treatments for these cancers can work by reducing oestrogen synthesis or blocking its receptor-mediated effects.
- When evaluating evidence linking a factor (such as oestrogen exposure, or any other proposed risk factor) to cancer, association alone does not establish causation — the same general evaluative caution that applies throughout this course to any observed correlation.
Worked examples
Worked example 3.8.2 · 4 marks
A biologist has two cell samples from the same early mammalian embryo: sample X, taken from the very earliest cell divisions, and sample Y, taken slightly later from the inner cell mass.
Sample X's cells can each give rise to every cell type in the body AND extra-embryonic tissue; sample Y's cells can each give rise to every body cell type but NOT extra-embryonic tissue.
Identify the potency of each sample's cells, and explain why sample Y's cells, despite this more limited potency, would still be more useful than sample X's for producing patient-specific replacement tissue without extra-embryonic complications.
Show worked solution
Sample X's cells are totipotent; sample Y's cells are pluripotent.
Sample Y's cells are more useful for replacement tissue precisely because they are already committed to body cell types only — directed differentiation reliably produces the specific tissue required without risking unwanted extra-embryonic tissue.
Mark scheme · 4 marks
- Identifies sample X's cells as totipotent 1 mark
- Identifies sample Y's cells as pluripotent 1 mark
- States pluripotent cells are committed to body cell types only 1 mark
- Explains this avoids the risk of unwanted extra-embryonic tissue 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.2 · 5 marks
A researcher wants to reduce the expression of a specific gene in cultured human cells, without permanently altering the cell's DNA sequence in any way.
Describe how RNA interference (RNAi) could achieve this, and explain why this approach specifically targets only the chosen gene's mRNA and not the mRNA of other, unrelated genes.
Show worked solution
A double-stranded RNA complementary to the target mRNA is introduced.
Dicer cleaves it into siRNA fragments; a guide strand loads into RISC and base-pairs with the target mRNA; RISC cleaves the bound mRNA, reducing translation of that gene's product, with no change to DNA.
Specificity arises because the guide strand's sequence, from complementary base pairing, only matches the intended target's mRNA.
Mark scheme · 5 marks
- States a double-stranded RNA complementary to the target mRNA is introduced 1 mark
- States Dicer cleaves it into siRNA, loaded into RISC 1 mark
- States the guide strand base-pairs with, and RISC cleaves, the target mRNA 1 mark
- States this acts at the mRNA level, with no change to DNA 1 mark
- Explains specificity arises from complementary base pairing to only the intended target 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.2 · 4 marks
A particular cell type requires an independent mutation to occur in each of three specific genes, A, B and C, before it can become cancerous.
Suppose, for illustration, that in a single cell division the probability of a mutation occurring in gene A is 1 in 10⁶, in gene B is 1 in 10⁷, and in gene C is 1 in 10⁵.
Calculate the probability that a single cell division produces all three mutations simultaneously in one cell, and explain why cancer nonetheless occurs, given that a person's body undergoes very many cell divisions over a lifetime.
Show worked solution
Probability of all three together:
extraordinarily small per division.
However, the body undergoes an enormous number of divisions over a lifetime, each an independent opportunity — multiplying a tiny per-division probability by a huge number of opportunities makes the event realistic in aggregate, and explains rising cancer risk with age.
Mark scheme · 4 marks
- Calculates the combined probability as 1 mark
- States this is extraordinarily small for any single division 1 mark
- States the body undergoes a very large number of independent divisions over a lifetime 1 mark
- Explains multiplying a tiny probability by many opportunities makes the event realistic in aggregate, explaining age-related risk 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Genome projects
Genome projects 3.8.3
- Bioinformatics: computational tools and databases used to store, organise, search and analyse large-scale biological data — searching gene sequences across genomes, predicting protein structure or function, and identifying disease-associated mutations by comparison against a reference database.
- Comparative genomics: whole-genome similarity provides evidence of evolutionary relationships (more similar sequence generally indicates more recent common ancestry); a single gene's sequence being conserved across widely diverged species is evidence of that gene's functional importance, since mutations there have generally been selected against over evolutionary time.
- Genome sequencing determines an organism's complete DNA base sequence — the Human Genome Project, an international collaborative effort largely completed in the early 2000s, sequenced the human genome and found that only a relatively modest proportion of it directly codes for protein.
- 'Does not code for protein' is not the same claim as 'has no function' — a substantial, already-identified portion of that non-coding DNA is regulatory sequence (promoters and enhancers) essential to whether and how strongly a nearby gene is expressed, without itself specifying any amino acid sequence.
- Personalised medicine (using an individual's own genome to inform their treatment) is a significant, still-developing practical application, made possible only because genome sequencing has become cheap and fast enough, and reference databases large enough, to make this kind of individual comparison routine.
Worked examples
Worked example 3.8.3 · 4 marks
Comparative genomics reveals that a particular short DNA sequence within a gene is almost identical across humans, mice, and zebrafish — three species whose lineages diverged from a common ancestor many millions of years ago.
A second region of the same gene, in an intron, shows far more sequence difference between the same three species.
Explain what the difference between these two regions' rates of change most likely indicates about their function.
Show worked solution
A sequence remaining almost unchanged across such divergent species indicates strong selection against mutation there, since harmful mutations were selected against independently in each lineage — evidence of an important, functionally constrained sequence.
The more variable intron region is consistent with weaker or no functional constraint.
Mark scheme · 4 marks
- States the conserved sequence shows strong selection against mutation 1 mark
- Explains this occurred independently in each lineage 1 mark
- Concludes this is evidence of functional importance 1 mark
- States the more variable region is consistent with weaker functional constraint 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.3 · 4 marks
The Human Genome Project found that a comparatively small proportion of the human genome directly codes for protein.
A student concludes from this that most of the human genome 'has no function and is essentially useless leftover DNA'.
Evaluate this conclusion.
Show worked solution
This is not well supported. 'Does not code for protein' is not the same as 'has no function' — much non-coding DNA consists of regulatory sequences (promoters, enhancers) essential to gene expression, or is transcribed into functional RNA (tRNA, rRNA). 'Not yet understood' and 'has no function' are different claims; the conclusion overstates current knowledge.
Mark scheme · 4 marks
- States 'not coding for protein' is not the same as 'having no function' 1 mark
- Identifies regulatory sequences (promoters/enhancers) as functional non-coding DNA 1 mark
- Identifies functional non-coding RNA (tRNA/rRNA) as a further example 1 mark
- Distinguishes 'not yet understood' from 'has no function', concluding the claim overstates current knowledge 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Gene technologies
Obtaining DNA fragments 3.8.4.1
- Copying mature mRNA: reverse transcriptase synthesises a single complementary DNA (cDNA) strand from a mature, spliced mRNA template → the RNA is removed and DNA polymerase synthesises the second strand → double-stranded cDNA — since it is made from already-spliced mRNA, this cDNA lacks introns entirely.
- Cutting genomic DNA: restriction endonucleases cut genomic DNA at specific recognition sequences, isolating a fragment containing the target gene directly from the genome (introns included, since this DNA was never spliced).
- Chemical synthesis: a known or computationally designed DNA sequence can be synthesised directly ('gene machine'), without needing any biological template at all.
- The near-universal genetic code is exactly what makes transferring a gene between species meaningful at all — a gene obtained by any of these three routes can, in principle, be expressed correctly in a different (suitable) host species, since that host reads the same codon-to-amino-acid assignments.
- cDNA (from mRNA) and directly-cut genomic DNA differ in one crucial respect: cDNA lacks introns (since it is copied from already-spliced mature mRNA), while genomic DNA retains them — this matters because many practical hosts used for gene expression (e.g. bacteria) cannot splice introns out of eukaryotic genes themselves, making cDNA generally the more reliable source for producing an intended protein product in such a host.
- Once obtained by any of these three routes, the target DNA fragment can be amplified further, either in vitro by PCR (covered next) or in vivo by inserting it into host cells that copy it as they themselves divide.
The polymerase chain reaction (PCR) 3.8.4.1
- PCR cycle: (1) Denaturation (~95°C) — hydrogen bonds between the two DNA strands break, separating them. (2) Primer annealing (~50–65°C) — short primers bind to complementary sequences flanking the target region on each separated strand, with their 3′ ends pointing into the target. (3) Extension (~72°C, using Taq polymerase) — a heat-tolerant DNA polymerase synthesises new complementary DNA 5′ to 3′ from each primer, producing two new double-stranded copies.
- Repeating the cycle (with newly made DNA itself also serving as template for the next cycle) gives, under ideal conditions, target copies after cycles — an exponential (doubling) increase.
- A heat-tolerant polymerase (Taq, originally isolated from a heat-tolerant bacterium) is specifically necessary because an ordinary polymerase would itself be denatured by the high-temperature step repeated every single cycle, needing to be replaced before each one — Taq polymerase survives this repeated heating, so the same polymerase molecules remain active and reusable throughout the whole reaction.
- The doubling assumption () holds only while reagents (primers, free nucleotides, polymerase) remain in excess relative to the amount of template present — real reactions plateau in later cycles once these components become limiting relative to the very large number of template molecules by then present, so PCR does not continue doubling indefinitely.
- Primers are chosen to be complementary to sequences flanking (either side of) the target region, not to the target region itself — their role is to define exactly which stretch of DNA gets amplified and to provide the free 3′ end that DNA polymerase requires to begin synthesis at all.
Recombinant DNA technology 3.8.4.1
- Sticky ends: short single-stranded overhangs left by many restriction endonucleases when they cut DNA at a specific (often staggered) recognition sequence — a sticky end base-pairs with a complementary sticky end from the same or a compatible enzyme, which is what makes joining two different DNA molecules together possible.
- Building a recombinant vector: a restriction enzyme cuts out the gene of interest, leaving sticky ends → the same (or a compatible) enzyme cuts a vector (commonly a plasmid), leaving complementary sticky ends → the compatible sticky ends anneal (base-pair) and DNA ligase seals the fragments together by forming phosphodiester bonds, producing a recombinant plasmid carrying a promoter, the target gene, and a terminator (which ends transcription — not a stop codon, which instead ends translation) → the recombinant vector is introduced into host cells (transformation) → a selectable marker gene identifies which cells actually took up the vector, and separate screening verifies the desired insert is present → successfully transformed cells are cultured, dividing and copying the vector DNA along with their own.
- Recombinant DNA technology combines DNA from different sources to produce a new combination not found in nature — genetic engineering, in its most literal sense.
- A marker gene is essential precisely because transformation efficiency is never 100% — without a way to select only the successfully transformed cells, the desired recombinant organisms would be lost among a majority of untransformed cells that never took up the vector at all.
- A vector's origin of replication is what allows it to be copied independently by the host cell's own machinery as that cell divides — without it, even a successfully inserted gene would not be reliably propagated to daughter cells.
Applications of gene technologies 3.8.4.1
- Somatic cell gene therapy: affects only the treated individual's own (non-reproductive) cells, and so is not passed on to any offspring. Germline gene therapy: would affect reproductive cells, and so could be passed to offspring — currently prohibited or highly restricted in humans in most jurisdictions.
- Medical protein production: engineered cells (e.g. bacteria carrying the human insulin gene) produce and are used to purify a useful protein at scale — bacterially produced human insulin has largely replaced earlier extraction from animal pancreases, matching the human sequence exactly (avoiding the immune response a slightly different animal sequence can provoke) and removing the supply limitation of animal pancreas availability.
- Crop improvement: a useful gene is introduced into a crop species (e.g. for pest resistance, herbicide resistance, or improved nutritional content), altering the gene present and allowing comparison of expression and resulting phenotype against the unmodified crop.
- Somatic gene therapy: a working copy of a gene is delivered into a patient's own somatic cells → the working gene is expressed → this may improve the affected cells' function, compensating for a faulty or non-functional endogenous copy — gene addition of this kind need not remove the original, faulty allele, it simply supplies a working copy alongside it.
- Germline gene therapy is restricted for both technical-safety and ethical reasons — the one clear line this whole topic draws between what is routinely practised (somatic applications) and what is not (germline modification of humans).
- Evaluating any of these applications requires weighing several genuinely distinct considerations together: benefits and safety, ecological effects (particularly relevant to crop modification and any release of modified organisms into the environment), cost and access (who can actually benefit from an expensive technology), and ownership (who holds rights to genetically modified organisms, sequences or techniques) — a well-reasoned evaluation names more than one of these, not just the most obvious safety concern alone.
- CRISPR-Cas9 (covered alongside recombinant DNA technology more broadly) extends this toolkit further: a guide RNA directs the Cas9 enzyme to cut DNA at essentially any chosen target sequence, offering far more flexible and precise targeting than a restriction endonuclease permanently fixed to one natural recognition sequence — it is the guide RNA, not Cas9 itself, that determines the target location, which is exactly why the guide RNA (not the enzyme) is redesigned for each new target.
DNA probes and hybridisation 3.8.4.2
- DNA probe: a short, single-stranded DNA sequence, labelled for detection, complementary to a specific target allele sequence — used to test whether that specific sequence is present in a DNA sample, via hybridisation (complementary base pairing).
- If the sample contains the complementary (matching) allele: the probe hybridises fully and is retained through the stringent wash, giving a positive signal, confirming the tested allele sequence is present.
- If the sample instead contains a mismatching (alternative) allele: the probe cannot hybridise as fully, and is removed by the stringent wash, giving no signal.
- Sample DNA is denatured (made single-stranded) → a labelled, single-stranded probe complementary to the target allele is added and allowed to hybridise (base-pair) with any matching sequence present → the sample is washed under stringent conditions → the probe is detected only if it remains bound.
- The stringent wash step is what actually gives the technique its specificity — wash conditions are deliberately set strict enough that a probe with even a small sequence mismatch to the target fails to remain bound, while a fully complementary probe does, converting a chemical hybridisation event into a reliable yes/no allele-detection result.
- Appropriate, validated hybridisation and wash conditions are essential for reliable allele discrimination — conditions that are too lenient risk a mismatched probe being retained (a false positive); conditions that are too stringent risk even a correctly matched probe failing to remain bound (a false negative).
- Applications include screening for inherited conditions, drug response, and health risk factors — results from this kind of screening can inform genetic counselling and personalised medicine, directly connecting this technique to the genome-project applications covered earlier.
Genetic fingerprinting 3.8.4.3
- VNTR (variable number tandem repeat): a locus where individuals can differ in how many times a short DNA sequence is repeated in a row — the basis of genetic fingerprinting, since fragment length at a VNTR locus varies between individuals depending on their specific repeat number.
- Fragment length at a VNTR locus: length = (constant flanking DNA) + (repeat unit length × number of repeats) — e.g. with 20 bp repeat units and 100 bp of total constant flanking DNA, 3 repeats give a 160 bp fragment while 5 repeats give a 200 bp fragment.
- PCR is carried out across several selected VNTR loci simultaneously, using primers that bind the constant flanking sequences either side of each repeat region (not the variable repeats themselves) → the resulting fragments (whose length depends directly on each individual's repeat number at that locus) are separated by gel electrophoresis, with negatively charged DNA fragments migrating toward the positive electrode, and smaller fragments travelling further than larger ones in a given time → resulting band patterns are compared between a questioned sample and one or more comparison samples across multiple loci.
- A single matching band at just one locus between two samples is comparatively weak evidence on its own, since a matching fragment length there could plausibly occur by chance — using MANY loci together, combined with statistical assessment of how likely a given multi-locus match is to occur by chance, is what actually gives genetic fingerprinting its practical power to distinguish individuals reliably.
- A match supports (is consistent with) a comparison, but does not by itself PROVE identity — the same evaluative caution (association/consistency is not proof) that recurs throughout this course whenever comparing evidence against a hypothesis.
- Identical twins usually share standard DNA profiles (since they arise from a single fertilised egg and so share essentially identical genomes), a genuine limitation of the technique worth naming directly when evaluating its reliability in specific contexts.
- Applications extend well beyond forensics alone: kinship testing, selective breeding programmes, and assessing genetic diversity within a population all use the same underlying VNTR-comparison principle.
Worked examples
Worked example 3.8.4 · 4 marks
A forensic sample contains 4 copies of a specific target DNA sequence.
The sample is amplified by PCR for 12 complete cycles, assumed 100% efficient throughout.
Calculate the number of copies of the target sequence present after amplification, and calculate how many additional complete cycles would be needed, from this point, to first exceed 10 million copies.
Show worked solution
After 12 cycles:
copies.
Need:
so further cycles (since ), giving cycles total and copies.
Mark scheme · 4 marks
- Calculates 16 384 copies after 12 cycles 1 mark
- Sets up the inequality 1 mark
- States further complete cycles are needed 1 mark
- Confirms the resulting total (≈16.8 million) exceeds 10 million 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.4 · 4 marks
A gene of interest is cut out of a donor organism's genome using restriction enzyme EcoRI, leaving sticky ends.
A bacterial plasmid vector is cut with the SAME enzyme, EcoRI.
Explain why using the same restriction enzyme for both the gene and the vector is essential for building a recombinant plasmid, and explain the role of DNA ligase in completing the process.
Show worked solution
The same enzyme leaves sticky ends of exactly matching, complementary sequence on both gene and vector, so they can anneal — a different enzyme would very likely leave incompatible sticky ends.
DNA ligase then catalyses phosphodiester bonds joining the two fragments' backbones together, sealing the gene into the plasmid.
Mark scheme · 4 marks
- States the same enzyme gives matching, complementary sticky ends 1 mark
- Explains a different enzyme would likely give incompatible sticky ends 1 mark
- States the sticky ends anneal by base pairing 1 mark
- States DNA ligase forms phosphodiester bonds, sealing the fragments together 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Worked example 3.8.4 · 5 marks
At a particular VNTR locus used in genetic fingerprinting, the repeat unit is 16 base pairs long, and the constant flanking DNA either side of the repeat region totals 84 base pairs.
Individual X has 7 repeats at this locus on one chromosome and 11 repeats on the homologous chromosome.
Calculate the length, in base pairs, of each of the two PCR-amplified fragments this locus would produce for individual X, and explain why this individual would show two separate bands (rather than one) at this locus on a genetic fingerprint gel.
Show worked solution
Fragment length = flanking DNA + (repeat length × repeats).
7 repeats:
bp.
11 repeats:
bp.
Two bands appear because individual X is heterozygous at this locus — the two differently-sized fragments travel different distances during gel electrophoresis.
Mark scheme · 5 marks
- Calculates 196 bp for the 7-repeat allele 1 mark
- Calculates 260 bp for the 11-repeat allele 1 mark
- States individual X is heterozygous at this locus 1 mark
- Explains the two different fragment lengths travel different distances on the gel, giving two bands 1 mark
- States a homozygous individual would instead show a single band 1 mark
Do not count matching words alone — ask whether your answer actually makes the same claim.
Per disputationem veritatem quaerimus