BiologyNucleic acids, genomes and protein synthesis › The genetic code and transcription

The genetic code and transcription

Four bases have to specify twenty amino acids, and they do it in threes. The rules that follow, 64 triplets for 20 amino acids, no overlaps, and almost the same table in every organism alive, explain a great deal about how mutations behave and why genetic engineering works at all.

Before this DNA and RNA structure · Complementary base pairing · The nuclear envelope and nuclear pores

COMMON MISCONCEPTION

The genetic code is degenerate, so it is a wasteful system evolution has not got round to tidying up.

What you should be able to do

Four properties, and what each one makes possible

The genetic code is the relationship between triplets of bases and amino acids. Four statements describe it, and each has a consequence you can be asked about directly.

PropertyWhat it meansWhat follows
TripletThree bases specify one amino acid64 possible triplets for 20 amino acids, so there is capacity to spare
DegenerateMost amino acids have more than one tripletMany base substitutions produce no change in the protein
Non-overlappingEach base belongs to one triplet onlyAn insertion or deletion not divisible by three shifts every triplet after it
UniversalAlmost all organisms read the same triplets the same wayA human coding sequence, freed of its introns and given bacterial control signals, directs a bacterium to make the human protein

Why three and not two

Show, using numbers, why a code based on pairs of bases could not specify twenty amino acids, and state how many triplets are spare once every amino acid has one.

Show the working

There are four bases. A code of one base gives 4¹ = 4 combinations, which is nowhere near enough. A code of two bases gives 4² = 16 combinations, still four short of the twenty amino acids that need specifying.

A code of three bases gives 4³ = 64 combinations, which is comfortably enough. Three of the 64 are stop signals, leaving 61 that specify amino acids.

Sixty-one triplets sharing twenty amino acids means 41 more than the minimum, and that surplus is what degeneracy consists of. Leucine, serine and arginine have six triplets each; methionine and tryptophan have only one.

Universality is worth stating carefully. The code is very nearly universal rather than absolutely so. Mitochondria and a few protoctists read a handful of triplets differently. Say 'universal' if a mark scheme asks for the property; say 'almost universal' if you are asked to be precise. Either way the consequence is the point: because a bacterium reads human triplets the same way a human cell does, human insulin can be manufactured in E. coli, and it has been since 1982.

Codon
A sequence of three bases in mRNA that specifies one amino acid or a stop signal.
Degenerate
Describing a code in which most amino acids are specified by more than one triplet.
Non-overlapping
Describing a code in which each base is part of one triplet only and triplets are read in sequence.

Genes, exons and introns

A gene is a sequence of DNA bases that codes for a polypeptide, or for a functional RNA molecule such as a tRNA. Its position on a chromosome is its locus. The genome is the complete set of DNA in a cell, and the proteome is the full range of proteins that cell can produce, a smaller and more changeable thing, since not every gene is expressed at once.

In eukaryotes a gene is not one continuous coding sequence. It is broken into exons, which are the parts that end up represented in the finished mRNA, separated by introns, which do not. Introns can be long: some human genes are more than 95% intron by length. Prokaryotic genes almost never contain introns, a few self-splicing exceptions aside, which is one of several reasons a human gene cannot simply be pasted into a bacterium and expected to work.

Gene
A sequence of DNA bases coding for a polypeptide or a functional RNA.
Exon
A section of a gene that is represented in the mature mRNA.
Intron
A non-coding section of a gene, transcribed into pre-mRNA and then removed before the mRNA leaves the nucleus.
Genome
The complete set of DNA in a cell, including the DNA that does not code for polypeptides.

A large proportion of eukaryotic DNA is non-coding: introns, repeated sequences between genes, and regions involved in switching genes on and off. Calling all of it useless was a fashion that did not survive contact with the evidence, and questions increasingly expect you to know that non-coding does not mean functionless.

Transcription, step by step

Transcription is the copying of one gene into mRNA, and it happens in the nucleus because that is where the DNA is. The enzyme is RNA polymerase, not DNA polymerase, which does a different job in a different process, and confusing the two is one of the quickest ways to lose a whole answer.

One strand is read and the other is ignored. Because the mRNA is complementary to the template, it ends up with the same sequence as the coding strand, except that every thymine is a uracil.

In order, then. RNA polymerase binds to a promoter region at the start of the gene and unwinds a short section of the double helix, breaking the hydrogen bonds between the paired bases. Only one of the two exposed strands is used: the template strand. Free RNA nucleotides in the nucleoplasm align against it by complementary base pairing, with uracil pairing to adenine wherever thymine would have gone. RNA polymerase joins the aligned nucleotides by condensation, forming phosphodiester bonds, and moves along the gene as it works. Behind it, the two DNA strands re-form their hydrogen bonds and the helix rewinds. When the enzyme reaches a terminator sequence at the end of the gene it detaches, and the finished pre-mRNA is released. The terminator is a transcription signal in the DNA, a different thing from the stop codons translation reads in the mRNA.

The strand that is not read is the coding strand, and it earns its name because its sequence is the same as the mRNA's, with T wherever the mRNA has U. That is useful in exam questions: if you are given the coding strand, you can write the mRNA by changing every T to a U. If you are given the template strand, you have to take the complement.

Splicing: cutting out what is not needed

The pre-mRNA that RNA polymerase produces is a copy of the whole gene, introns included. Before it can be used it goes through splicing: structures called spliceosomes cut the introns out and join the exons together, producing mature mRNA. Only then does the molecule leave the nucleus through a nuclear pore and travel to a ribosome.

In the default account every exon survives and stays in order while every intron is discarded; alternative splicing, described below, relaxes the first half of that. The mature mRNA is shorter than the gene it came from, which is why a gene and its mRNA are not the same length.

Splicing is not always done the same way on the same transcript. Under alternative splicing, different combinations of exons are retained, so one gene gives rise to several different mRNA molecules and therefore several different polypeptides. This is a large part of why humans manage a proteome of well over 100,000 proteins from around 20,000 protein-coding genes.

Pre-mRNA
The immediate product of transcription in a eukaryote, containing both exons and introns.
Splicing
The removal of introns from pre-mRNA and the joining of exons to produce mature mRNA.

Why degeneracy makes some mutations silent

Take the mRNA codon GAA, which specifies glutamic acid. Change its third base to G and you have GAG, which also specifies glutamic acid. The DNA has changed, the mRNA has changed, and the protein has not. A substitution of that kind is a silent mutation, and it is common precisely because the triplets sharing an amino acid tend to differ in the third base.

Compare that with a deletion. Remove one base and every triplet downstream is read from a new starting point, because the code is non-overlapping and read strictly in threes from a fixed start. A single deletion near the beginning of a gene can therefore change every amino acid after it and usually produces a stop codon early, giving a short and useless polypeptide. The contrast between a substitution that changes nothing and a deletion that changes everything is a standard exam comparison, and it comes straight out of two properties of the code.

TRY IT: Two mutations, two very different outcomes

A section of the coding strand of a gene reads TTA CGA GAA CCT. One sample has the fourth base changed from C to A. Another sample has the fourth base deleted altogether. Count along the sequence before you start. Explain why the second is likely to be far more damaging than the first, without needing a codon table.

Check your answer

The fourth base is the C at the start of the second triplet, so the substitution changes one triplet and nothing else. CGA becomes AGA, so at worst one amino acid in the finished polypeptide is different, and because the code is degenerate the new triplet may well specify the same amino acid anyway. Even if it does not, a single amino acid change in a region away from the active site or binding site often leaves the protein working normally.

The deletion is different in kind. The code is non-overlapping and read in threes from a fixed start, so removing one base pulls every subsequent base one place forward. Every triplet from that point on is read differently, which is a frame shift, and the amino acid sequence after the deletion bears no relation to the original.

A shifted reading frame also produces stop codons at random positions, so the polypeptide is usually cut short as well as scrambled. The protein almost never folds into a working shape. One base removed, an entire protein lost.

In the exam

Check yourself

A biotechnology company wants a human gene expressed in a bacterium. They isolate the gene directly from a human chromosome and insert it into E. coli, but no functional human protein is produced. Suggest why, and what they should have used instead.

Answer

A gene taken straight from a human chromosome contains introns. Bacteria do not carry out splicing, because their own genes almost never contain introns and they have no spliceosomes.

The bacterium therefore transcribes the whole inserted sequence, introns included, and translates it as though every base were coding. The introns are read in the reading frame that happens to follow, producing wrong amino acids and almost certainly premature stop codons, so no functional human protein appears.

The fix is to start from mature mRNA rather than from the chromosome. Isolate the mRNA from a human cell that expresses the gene heavily, then use reverse transcriptase to make a DNA copy of it. Because the introns have already been spliced out of that mRNA, the resulting DNA is a continuous coding sequence the bacterium can handle.

Note what this does not challenge: the code is still universal, so the bacterium reads each triplet exactly as a human cell would. The problem was never the code. It was the introns.

Questions

Written to the command words the boards use. Try them on paper before opening a scheme: the marks go to points made, not to length.

Question 15 marks

Describe how a molecule of pre-mRNA is produced from a gene in the nucleus of a eukaryotic cell.

Mark scheme
  1. B1 RNA polymerase binds to a promoter region at the start of the gene
  2. B1 the enzyme unwinds a short section of the double helix, breaking the hydrogen bonds between the paired bases
  3. B1 only one of the two exposed strands, the template strand, is read
  4. B1 free RNA nucleotides align against the template by complementary base pairing, with uracil pairing to adenine wherever thymine would have gone
  5. B1 RNA polymerase joins the aligned nucleotides by condensation into phosphodiester bonds, moving along the gene until it reaches a stop signal and releases the pre-mRNA

Question 24 marks

Explain why a mature mRNA molecule is shorter than the gene it was transcribed from, and explain how one gene can give rise to more than one polypeptide.

Mark scheme
  1. B1 the pre-mRNA is a copy of the whole gene, introns included
  2. B1 spliceosomes cut the introns out and join the exons together, so the introns are not represented in the mature mRNA
  3. B1 in alternative splicing, different combinations of exons are retained in the mature mRNA
  4. B1 so several different mRNA molecules, and therefore several different polypeptides, can come from the same gene

Question 34 marks

A single base is inserted into the coding sequence of a gene close to its start. The polypeptide produced is much shorter than normal and its amino acid sequence bears no relation to the original beyond the point of insertion. Suggest why, referring to two properties of the genetic code.

Mark scheme
  1. B1 the code is read in threes from a fixed starting point and is non-overlapping, so each base belongs to one triplet only
  2. B1 inserting a base pushes every subsequent base one place along, so every triplet after the insertion is read differently, which is a frame shift
  3. B1 the amino acid sequence after that point is therefore completely altered rather than altered in one place
  4. B1 a shifted reading frame produces stop codons at positions where there were none, so translation ends early and the polypeptide is short

Question 43 marks

The coding strand of a short section of a gene reads GAT CCA TTG. Give the base sequence of the template strand, and give the base sequence of the mRNA transcribed from this section.

Mark scheme
  1. B1 template strand CTA GGT AAC, each base complementary to the coding strand
  2. A1 mRNA GAU CCA UUG
  3. B1 the mRNA matches the coding strand, with uracil in place of every thymine

Question 52 marks

State what is meant by saying that the genetic code is degenerate, and state what is meant by saying that it is non-overlapping.

Mark scheme
  1. B1 degenerate: most amino acids are specified by more than one triplet
  2. B1 non-overlapping: each base belongs to one triplet only, and the triplets are read one after another

Worth remembering

  • Triplet, degenerate, non-overlapping, universal, and a consequence for each.
  • 4³ = 64 triplets, 61 coding and 3 stop, for 20 amino acids.
  • RNA polymerase transcribes; only the template strand is read.
  • The mRNA matches the coding strand with U in place of T.
  • Splicing removes introns and joins exons; alternative splicing lets one gene give several polypeptides.

CHECK YOUR PROGRESS

Rate how confident you are with each objective for this lesson. Ratings are kept in this browser, on this device, and are sent nowhere.

  • State the four properties of the genetic code and explain a consequence of each.
  • Explain, with numbers, why the code has to be read in threes rather than ones or twos.
  • Distinguish exons from introns, and a gene from the genome.
  • Describe transcription in sequence, naming the enzyme, the strand used and the bonds involved.
  • Explain what splicing does and why one gene can give rise to more than one polypeptide.
  • Explain why many base substitutions change nothing about the protein.

Open the full revision checklist to see every objective in the curriculum in one place.

WORKBOOK

The same questions as the player, on paper with room to work, and a separate book of mark schemes. Free to use; please do not redistribute or sell.