↓ Skip to main content
  1. Posts/

Codon Map

1361 words·7 mins·
Table of Contents ▼

Nucleotide Bases
#

The genetic code is the set of rules used by cells to translate nucleotide sequences into proteins. The code relies on the specific ordering of nitrogen-containing molecules called nucleotide bases. In DNA, these bases comprise Adenine ($\ce{A}$), Thymine ($\ce{T}$), Guanine ($\ce{G}$), and Cytosine ($\ce{C}$). Adenine and Guanine are both purines (a pyrimidine and an imidazole) while Cytosine and Thymine are both pyrimidines.

Structure of nucleotide bases

**A**denine
Adenine
**T**hymine (DNA)
Thymine (DNA)
**G**uanine
Guanine
**C**ytosine
Cytosine
**U**racil (RNA)
Uracil (RNA)

Codons & Amino Acids
#

Hydrogen bonding between purines and pyrimidines form the standard base pairs $\ce{A=T}$ (double bond) and $\ce{G#C}$ (triple bond). During the process of transcription, DNA serves as the template to synthesize messenger RNA (mRNA), where Uracil ($\ce{U}$) replaces Thymine to pair with Adenine. The linear sequence of these bases is read sequentially by cellular machinery in discrete, non-overlapping triplets known as codons.

  1. 1

    DNA to RNA (Transcription)

    RNA polymerase reads a specific gene on the DNA strand and builds a matching single-stranded copy called messenger RNA (mRNA). During this step, the RNA base uracil (U) replaces the DNA base thymine (T). The mRNA then leaves the nucleus and goes into the cytoplasm.
  2. 2

    RNA to Amino Acids (Translation)

    The ribosome reads the mRNA code in groups of three bases called codons. Another molecule called transfer RNA (tRNA) brings the matching amino acids to the ribosome one by one. Each codon matches a specific amino acid.
  3. 3

    Amino Acids to Proteins (Folding)

    The ribosome joins the amino acids together with peptide bonds into a long chain called a polypeptide. Once the chain is complete, it twists and folds into a complex, three-dimensional shape based on the traits of its amino acids. This finished shape becomes a working, functional protein.
Example sequence showing messenger RNA (mRNA) codons translated to amino acids.
Example sequence showing messenger RNA (mRNA) codons translated to amino acids.

Replication: DNA makes exact copies of itself so cells can divide and pass instructions to new cells.Transcription: The cell copies a section of DNA into a messenger RNA (mRNA) molecule inside the nucleus. Translation: Ribosomes read the mRNA code to link amino acids together and build a functional protein.

Because there are four distinct bases ($\ce{A}$, $\ce{T}$, $\ce{G}$, and $\ce{C}$), three-base combinations yield $4^3 = 64$ possible codons, where specific codons code for particular amino acids (or one of three stop codons). Since there are only $20$ amino acids1, we say there is redundancy, or degeneracy, in the genetic code. In other words, multiple codons can map to a single amino acid.

For example, Glutamine (Gln) is coded by $\ce{CAG}$ and $\ce{CAA}$ while Histidine (His) is coded by $\ce{CAC}$ and $\ce{CAU}$. In fact only methionine and tryptophan are specified by a single codon. The remaining amino acids are coded by at least two, and up to six, codons. Below is a lookup table that describes the relationship between codons and amino acids.

Standard codon table used to translate a three-letter mRNA sequence (codon) into a specific amino acid or a stop signal. NOTE. Methonine (Met) is the start codon.
Second base
First baseUCAGThird base
UPhe (F)Ser (S)Tyr (Y)Cys (C)U
C
Leu (L)StopStopA
StopTrp (W)G
CLeu (L)Pro (P)His (H)Arg (R)U
C
Gln (Q)A
G
AIle (I)Thr (T)Asn (N)Ser (S)U
C
Lys (K)Arg (R)A
Met (M)G
GVal (V)Ala (A)Asp (D)Gly (G)U
C
Glu (E)A
G
Note

Though the genetic code is degenerate (i.e., multiple codons may map to a single amino acid) the code is not ambiguous. In other words, a given codon always maps to exactly one amino acid

Degeneracy & Nucleotide Ambiguity Codes
#

To represent nucleotide degeneracy, or positions with more than one possible base in a codon, the IUPAC (International Union of Pure & Applied Chemistry) defines single-letter nucleic acid ambiguity codes. Below is the complete list of the IUPAC nucleic acid ambiguity codes.

    IUPAC Nucleic Acid Ambiguity Codes
    ------------------------------------
    R (Purines): A or G 
    Y (Pyrimidines): C or T/U 
    M (Amino): A or C 
    K (Keto): G or T/U 
    S (Strong): G or C 
    W (Weak): A or T/U 
    ------------------------------------
    H (not G): A, C, or T/U 
    B (not A): C, G, or T/U 
    V (not T/U): A, C, or G 
    D (not C): A, G, or T/U 
    ------------------------------------
    N (Any base): A, C, G, or T/U
Don’t get confused here. The single letter nucleic acid ambiguity codes are not the same as single letter abbreviations for amino acids.

Returning to the examples above, Glutamine (Gln) is coded by $\ce{CAG}$ and $\ce{CAA}$. Note that the third position of both codons is a purine ($\ce{G}$ or $\ce{A}$). Looking at the nucleic acid ambiguity table above we see that the code for purines is $\ce{R}$. So we can say the Glutamine is coded by $\ce{CAR}$. Similarly, Histidine (His), which is coded by $\ce{CAC}$ and $\ce{CAU}$, can be rewritten as $\ce{CAY}$, where $\ce{Y}$ is a pyrimidine. If instead we reference the codon lookup table we see that Serine (Ser) is coded for by 6 different codons: $\ce{UCA}$, $\ce{UCG}$, $\ce{UCC}$, $\ce{UCU}$, $\ce{AGU}$, and $\ce{AGC}$. Applying the ambiguity codes, we say that Serine (Ser) is coded by $\ce{UCN}$ and $\ce{AGY}$, where $\ce{N}$ stands for any base and $\ce{Y}$ corresponds to pyrimidine.

Mapping Codons to Amino Acids
#

In addition to codon tables, other methods have been developed to show the mapping between codons and their cooresponding amino acids. Memorizing codon to amino acid assignments is not a mandatory skill but it can significantly speed up data analysis, enable quick error detection in sequences, and generally deepen your understanding of sequence analysis, all without constantly relying on reference charts or tables.

Below are a few examples of different tools to show the relationship between codons and amino acids. The codon wheel for example is a circular chart where the inner ring is the first base of the three-letter codon. The middle ring is the second base, and the outer ring in the third base. The codon chart works similarly but also overlays additional information like amino acid properties (e.g., acidic, basic, etc.) and chemical structures.

Popular methods to map codons to amino acids

A Better Map?
#

flowchart LR
    A(Cytosine) --> B(Adenine)
    
    %% 1. Subgraph Layout with Markdown Titles
    subgraph 1st_base ["`**1st base**`"]
      A("`**C**
       Cytosine`")
    end
    subgraph 2nd_base ["`**2nd base**`"]
      B("`**A**
       Adenine`")
    end
    subgraph 3rd_base ["`**3rd base**`"]
      C(Guanine)
      D(Adenine)
      E(Cytosine)
      F(Uracil)
    end
    
    B --> C("`**G**
     Guanine`")
    B --> D("`**A**
     Adenine`")
    B --> E("`**C**
     Cytosine`")
    B --> F("`**U**
     Uracil`")
    
    subgraph Amino_Acids ["`**Amino Acids**`"]
      C -- "`**R**`" --> G((("`**Gln**
      Glutamine`")))
      D -- "`**R**`" --> G
      E -- "`**Y**`" --> H((("`**His**
      Histidine`"))) 
      F -- "`**Y**`" --> H
    end 

    %% ==========================================
    %% NODE STYLING (REUSABLE CLASSES)
    %% ==========================================
    classDef largeFont font-size:22px;
    classDef mediumFont font-size:16px;
    
    class A,B,G,H largeFont;
    class C,D,E,F mediumFont;

    %% ==========================================
    %% SUBGRAPH BOX STYLING
    %% ==========================================
    %% Syntax: style [subgraph_id] CSS_properties
    style 1st_base fill:currentColor,fill-opacity:0.07,stroke:currentColor,stroke-opacity:0.3,stroke-width:1px;
    style 2nd_base fill:currentColor,fill-opacity:0.07,stroke:currentColor,stroke-opacity:0.3,stroke-width:1px;
    style 3rd_base fill:none,stroke:none;
    style Amino_Acids fill:currentColor,fill-opacity:0.12,stroke:currentColor,stroke-opacity:0.5,stroke-width:2px;

    %% ==========================================
    %% LINK & ARROW TEXT STYLING
    %% ==========================================
    %% Edges are numbered sequentially (0, 1, 2...) in order of code appearance.
    %% Edges 0 through 4 handle the top layout branching rules.
    %% Edges 5 through 8 map the R and L paths inside Amino Acids.
    
    linkStyle 5 font-size:24px,font-weight:bold,fill:none;
    linkStyle 6 font-size:24px,font-weight:bold,fill:none;
    linkStyle 7 font-size:24px,font-weight:bold,fill:none;
    linkStyle 8 font-size:24px,font-weight:bold,fill:none;

I tried forever to memorize codon assignments using the tools described above with little luck–it just wouldn’t stick. If you are more of a visual learner like me, then tables are pretty useless. I tried the various wheel and chart representations but those didn’t help either.

Codon Trees

Conventional methods to map codons to amino acids

Codon Tree 5' G
Codon Tree 5’ G
Codon Tree 5' C
Codon Tree 5’ C
Codon Tree 5' A
Codon Tree 5’ A
Codon Tree 5' U
Codon Tree 5’ U

  1. In this artcile I discuss the standard genetic code but it is important to note that there are several variant translation systems used in certain organelles and microbes. Alternative genetic codes are variant translation systems where specific codons are reassigned to encode different amino acids or stop signals. ↩︎