DNA, RNA and how a gene becomes a protein
The structure of the molecule, the code it carries, and the two steps that turn instructions into working machinery.
01The structure
DNA is a double helix: two strands wound around each other, each built from repeating units called nucleotides. A nucleotide has three parts, a phosphate group, a deoxyribose sugar and one of four bases: adenine, thymine, cytosine and guanine.
The strands are held together by hydrogen bonds between bases, and the pairing is strict. Adenine pairs with thymine, cytosine with guanine. This complementary pairing is the whole trick: because each strand specifies the other, the molecule can be copied accurately, and a damaged strand can be rebuilt from its partner.
02Replication
Before a cell divides it copies its DNA. An enzyme unzips the helix, and DNA polymerase builds a new complementary strand along each old one. The result is two molecules, each containing one original strand and one new strand, which is why replication is described as semi-conservative.
Copying is extraordinarily accurate but not perfect. Polymerase proofreads as it goes, and repair systems catch most of what slips through. Errors that survive become mutations, which may be harmless, harmful or occasionally advantageous.
03Transcription
Genes are stretches of DNA that specify proteins, but DNA stays in the nucleus. The instructions are copied into messenger RNA, a single-stranded molecule that uses ribose instead of deoxyribose and uracil in place of thymine.
RNA polymerase binds a promoter region, unwinds the DNA and builds an RNA copy of one strand. In eukaryotes the initial transcript is edited before leaving the nucleus: non-coding stretches called introns are spliced out and the coding exons are joined.
04Translation
The messenger RNA travels to a ribosome, which reads it three bases at a time. Each triplet, called a codon, specifies one amino acid. Transfer RNA molecules carry the matching amino acids and recognise codons through a complementary anticodon.
The ribosome joins amino acids into a chain, which folds into a three-dimensional shape. That shape determines what the protein can do, which is why a single wrong amino acid can disable an enzyme. Reading starts at a start codon and stops at one of three stop codons.
05Why the code is redundant
There are sixty four possible codons and only twenty amino acids, so most amino acids are specified by more than one codon. The code is therefore described as degenerate, and the redundancy is protective: many single-base changes produce the same amino acid and have no effect at all.
The code is also nearly universal. The same codons mean the same amino acids in bacteria, plants and humans, which is strong evidence of shared ancestry and the reason genes can be transferred between species in biotechnology.
Test yourself
What does “Nucleotide” mean?
Which term matches this description: The rule that A pairs with T and C pairs with G.
What does “Codon” mean?
Which term matches this description: A non-coding section removed from the initial RNA transcript before translation.
What does “Semi-conservative” mean?
About this guide
An original guide written for Fathomly. © 2026 Fathomly, all rights reserved. Spotted an error? Send a correction.
Video: “Transcription and Translation: Protein Synthesis From DNA” by The Organic Chemistry Tutor, embedded from YouTube. The video belongs to its creator.