Skip to content

DNA to Protein Translator: Genetic Code Codon Converter

Translate a DNA or RNA sequence into a protein sequence free online. Enter a sequence and reading frame to get the amino acid sequence and codon breakdown.

DNA to Protein Translator

Enter a DNA or RNA sequence. Use the bases A, C, G and T. U is read as T. Spaces, line breaks and numbers are ignored. Translation uses the Standard Genetic Code (NCBI table 1).

Reading frame

Begin translation at the first ATG in this reading frame, the way a ribosome starts reading an open reading frame, instead of at the very start of the frame.

Loading calculator...
📚

Documentation

DNA to Protein Translator

A DNA to protein translator converts a DNA or RNA sequence into the protein sequence it encodes. It reads the sequence in groups of three letters, called codons, and converts each codon into one amino acid using the standard genetic code.

What DNA translation is

A cell stores genetic instructions in DNA. To build a protein, the cell first copies a gene into messenger RNA (mRNA). This step is called transcription. Ribosomes then read the mRNA and link amino acids together in the order the mRNA specifies. This step is called translation. The finished chain of amino acids folds up into a working protein.

RNA differs from DNA in only one relevant way: RNA uses the base uracil (U) where DNA uses thymine (T). Because of this, a DNA sequence can be translated using the same codon table as an RNA sequence, as long as every U is first read as a T.

How to translate DNA to protein

The translator turns an input sequence into a protein sequence in four steps.

  1. Clean the sequence. Letters are converted to uppercase and every U is changed to T. Spaces, line breaks, digits and the punctuation used in pasted FASTA sequences (>, - and *) are removed. If nothing is left after cleaning, the translator reports the input as empty.
  2. Check the bases. Each remaining character must be A, C, G, or T. Any other character stops the process and produces an error message listing the invalid characters.
  3. Split into codons. Starting from the chosen reading frame, the sequence is cut into non-overlapping groups of three bases.
  4. Translate each codon. Each codon is matched to one amino acid using the genetic code table below. Translation stops at the first stop codon (TAA, TAG, or TGA). The stop codon itself is not added to the protein sequence.

If fewer than three bases are left after the last complete codon, those bases cannot be translated, and the translator reports them as leftover bases. Bases that sit after a stop codon are not counted as leftover, because translation has already ended.

Reading frames

A reading frame is the starting point used to cut a sequence into codons. Because each codon is three bases long, a sequence can be split starting at its first, second, or third base. These are called reading frame 1, 2, and 3. Each frame usually produces a different protein sequence. In a real gene, only one reading frame produces the correct protein. The other two frames often run into a stop codon quickly and produce a short, meaningless sequence.

Starting at the first ATG

Most genes begin with the codon ATG. It codes for methionine and marks the point where the protein starts. The translator has an option called "Scan to first start codon (ATG)". With the option on, the translator moves forward through the chosen frame until it reaches a codon that is exactly ATG, then begins translating there. Bases before that ATG are ignored. With the option off, translation begins at the first base of the frame.

The scan accepts only an ATG that lands on a codon boundary of the chosen frame. Take CATGGCCTAA. Frame 1 splits it into CAT GGC CTA, so the letters ATG never line up with a codon, and the translator reports that no start codon was found in this frame. Frame 2 starts at the second base and gives ATG GCC TAA, so translation begins straight away and produces MA.

The genetic code

The standard genetic code assigns each of the 64 possible codons to one of 20 amino acids or to a stop signal. This is the code used by most genes in most organisms, including humans.

Amino acidCodons
Alanine (A)GCT, GCC, GCA, GCG
Arginine (R)CGT, CGC, CGA, CGG, AGA, AGG
Asparagine (N)AAT, AAC
Aspartic acid (D)GAT, GAC
Cysteine (C)TGT, TGC
Glutamic acid (E)GAA, GAG
Glutamine (Q)CAA, CAG
Glycine (G)GGT, GGC, GGA, GGG
Histidine (H)CAT, CAC
Isoleucine (I)ATT, ATC, ATA
Leucine (L)TTA, TTG, CTT, CTC, CTA, CTG
Lysine (K)AAA, AAG
Methionine (M)ATG
Phenylalanine (F)TTT, TTC
Proline (P)CCT, CCC, CCA, CCG
Serine (S)TCT, TCC, TCA, TCG, AGT, AGC
Threonine (T)ACT, ACC, ACA, ACG
Tryptophan (W)TGG
Tyrosine (Y)TAT, TAC
Valine (V)GTT, GTC, GTA, GTG
StopTAA, TAG, TGA

ATG codes for methionine. It is also the codon most genes use to mark the start of the coding sequence.

Example: translating a DNA sequence

Take the sequence ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG in reading frame 1.

Split into codons: ATG GCC ATT GTA ATG GGC CGC TGA AAG GGT GCC CGA TAG

Translate each codon: ATG → M, GCC → A, ATT → I, GTA → V, ATG → M, GGC → G, CGC → R, TGA → stop.

The eighth codon, TGA, is a stop codon, so translation ends there. The protein sequence is MAIVMGR, seven amino acids long. The bases after the stop codon are not translated.

Example: a coding sequence with no stop codon

Take the sequence ATGAAGTACGCCTACATCGCCAAGCAGCGCCAG in reading frame 1. It is 33 bases long, which is exactly 11 codons.

Split into codons: ATG AAG TAC GCC TAC ATC GCC AAG CAG CGC CAG

Translate each codon: M, K, Y, A, Y, I, A, K, Q, R, Q.

No stop codon appears, so all 11 codons are translated. The protein sequence is MKYAYIAKQRQ, with no leftover bases.

Example: how reading frame changes the result

The short sequence ATGCATGCA shows why the reading frame matters.

  • Frame 1 (start at base 1): ATG CAT GCA → M, H, A. Protein: MHA.
  • Frame 2 (start at base 2): TGC ATG CA → C, M, with 2 leftover bases that cannot form a codon.
  • Frame 3 (start at base 3): GCA TGC A → A, C, with 1 leftover base.

Three different starting points give three different protein sequences from the same nine bases.

Uses of DNA to protein translation

Researchers translate DNA sequences to predict what protein a gene produces, to check that a cloned sequence is still in the correct reading frame after edits, and to see how a mutation changes a protein's amino acid sequence. Bioinformatics pipelines translate genome sequences automatically to help annotate genes and to search protein databases.

A short history

Francis Crick described the flow of genetic information from DNA to RNA to protein in 1958, calling it the central dogma of molecular biology. In 1961, Marshall Nirenberg and Heinrich Matthaei showed that the codon UUU codes for the amino acid phenylalanine, the first codon to be solved. By 1966, researchers including Nirenberg and Har Gobind Khorana had worked out all 64 codon assignments, giving the genetic code used by this translator today.

Frequently asked questions

What is the difference between DNA translation and transcription?

Transcription copies a DNA sequence into mRNA. Translation reads that mRNA and builds a protein from it. This tool skips the RNA step and translates a DNA sequence directly, applying the same codon rules.

Why does the translator offer three reading frames?

A codon is three bases long, so a sequence can be split into codons starting at its first, second, or third base. Only one of these three frames usually matches the real coding sequence of a gene, so checking all three can help find it.

What happens if my sequence length is not a multiple of three?

The extra bases at the end cannot form a complete codon. The translator reports the number of leftover bases and translates every complete codon that comes before them.

Can I paste an RNA sequence instead of DNA?

Yes. Every U in the input is automatically read as a T before translation, so an RNA sequence gives the same result as its DNA equivalent.

What does a stop codon do?

A stop codon (TAA, TAG, or TGA) marks the end of the coding sequence. The translator stops there and does not add the stop codon itself, or anything after it, to the protein sequence.

What does the start-codon option do?

It tells the translator to skip ahead to the first ATG in the chosen reading frame and start there, the way a ribosome starts reading at a start codon. With the option off, translation begins at the first base of the frame instead. If the frame contains no in-frame ATG, the translator says so rather than guessing.

Do all organisms use the same genetic code?

Almost all nuclear genes use the standard code shown above. A few exceptions exist, such as human mitochondrial DNA, where the codon TGA codes for tryptophan instead of acting as a stop signal.