Understanding how DNA stores, copies, and expresses genetic information — from the double helix to the human genome
In the previous chapter, you learned about inheritance patterns and the genetic basis of such patterns. During Mendel's time, the nature of the "factors" regulating the pattern of inheritance was not clear. Over the next hundred years, the nature of the putative genetic material was investigated, culminating in the realisation that DNA — deoxyribonucleic acid — is the genetic material, at least for the majority of organisms.
DNA and RNA are the two types of nucleic acids found in living systems. DNA acts as the genetic material in most organisms. RNA, though it also acts as a genetic material in some viruses, mostly functions as a messenger. RNA has additional roles as well — it functions as an adapter, a structural molecule, and in some cases as a catalytic molecule.
In this chapter we discuss the structure of DNA, its replication, the process of making RNA from DNA (transcription), the genetic code that determines the sequences of amino acids in proteins, the process of protein synthesis (translation), and the elementary basis of their regulation. The determination of the complete nucleotide sequence of the human genome during the last decade has set in a new era of genomics.
DNA is a long polymer of deoxyribonucleotides. The length of DNA is usually defined as the number of nucleotides (or a pair of nucleotides referred to as base pairs) present in it. This is also the characteristic of an organism.
• Bacteriophage φ×174 — 5,386 nucleotides
• Bacteriophage lambda — 48,502 base pairs (bp)
• Escherichia coli — 4.6 × 10⁶ bp
• Haploid human DNA — 3.3 × 10⁹ bp
A nucleotide has three components: a nitrogenous base, a pentose sugar (ribose in RNA, deoxyribose in DNA), and a phosphate group. There are two types of nitrogenous bases — Purines (Adenine and Guanine) and Pyrimidines (Cytosine, Uracil, and Thymine). Cytosine is common to both DNA and RNA; Thymine is present only in DNA, while Uracil replaces Thymine in RNA.
A nitrogenous base is linked to the OH of 1′C pentose sugar through an N-glycosidic linkage to form a nucleoside (e.g., adenosine or deoxyadenosine, guanosine or deoxyguanosine, cytidine or deoxycytidine, uridine or deoxythymidine). When a phosphate group is linked to the OH of 5′C of a nucleoside through a phosphoester linkage, a nucleotide is formed. Two nucleotides are linked through a 3′–5′ phosphodiester linkage to form a dinucleotide. More nucleotides can be joined in this manner to form a polynucleotide chain.
A polynucleotide chain has a free phosphate moiety at 5′-end of the sugar (called the 5′-end) and a free OH of 3′C at the other end (called the 3′-end). The backbone of a polynucleotide chain is formed by sugar and phosphates, while the nitrogenous bases project from the backbone.
In RNA, every nucleotide residue has an additional –OH group at the 2′-position in ribose. Also, RNA contains uracil in place of thymine (5-methyl uracil).
DNA was first identified as an acidic substance in the nucleus by Friedrich Meischer in 1869, whom he named it "Nuclein." However, elucidation of its structure remained elusive due to technical limitations in isolating such a long polymer intact.
Base pairing confers a unique property to polynucleotide chains — they are complementary to each other. If the sequence of bases in one strand is known, the sequence in the other strand can be predicted. If each strand from a parental DNA acts as a template for synthesis of a new strand, the two daughter DNA molecules produced would be identical to the parental DNA molecule.
Francis Crick proposed the Central Dogma in molecular biology: genetic information flows from DNA → RNA → Protein. In some viruses, the flow is reversed (RNA → DNA), a process called reverse transcription.
Taking the distance between two consecutive base pairs as 0.34 nm, if the length of DNA double helix in a typical mammalian cell is calculated (6.6 × 10⁹ bp × 0.34 × 10⁻⁹ m/bp), it comes out to be approximately 2.2 metres — far greater than the dimension of a typical nucleus (approximately 10⁻⁶ m).
In prokaryotes (such as E. coli), the DNA is not scattered throughout the cell. DNA (being negatively charged) is held with some positively charged proteins in a region termed the nucleoid. The DNA in the nucleoid is organised in large loops held by proteins.
In eukaryotes, the organisation is much more complex. There is a set of positively charged, basic proteins called histones. Histones are rich in the basic amino acid residues lysine and arginine. They are organised to form a unit of eight molecules called the histone octamer.
The negatively charged DNA is wrapped around the positively charged histone octamer to form a structure called a nucleosome. A typical nucleosome contains 200 bp of DNA helix. Nucleosomes constitute the repeating unit of a structure in the nucleus called chromatin, which appears as "beads-on-string" when viewed under an electron microscope.
The beads-on-string structure in chromatin is packaged to form chromatin fibres that are further coiled and condensed at the metaphase stage of cell division to form chromosomes. Higher-level packaging requires additional proteins collectively referred to as Non-histone Chromosomal (NHC) proteins.
In a typical nucleus, some regions of chromatin are loosely packed and stain light — called euchromatin. Chromatin that is more densely packed and stains dark is called heterochromatin. Euchromatin is transcriptionally active, whereas heterochromatin is inactive.
Even though the discovery of nuclein by Meischer and the proposition of inheritance principles by Mendel were almost simultaneous, proving that DNA acts as genetic material took long. By 1926, the quest had reached the molecular level, but the question of what molecule was actually the genetic material had not been answered.
In 1928, Frederick Griffith, in a series of experiments with Streptococcus pneumoniae (bacterium responsible for pneumonia), witnessed a miraculous transformation in the bacteria.
When S. pneumoniae bacteria are grown on a culture plate, some produce smooth shiny colonies (S strain) while others produce rough colonies (R strain). The S strain bacteria have a mucous (polysaccharide) coat, while the R strain does not. Mice infected with the S strain (virulent) die from pneumonia, but mice infected with the R strain do not.
1. Heat-killed S strain injected into mice → mice did not die
2. Mixture of heat-killed S + live R bacteria injected → mice died
3. Living S bacteria recovered from dead mice
Griffith concluded that the R strain bacteria had been transformed by the heat-killed S strain. Some "transforming principle" had enabled the R strain to synthesise a smooth polysaccharide coat and become virulent.
Oswald Avery, Colin MacLeod, and Maclyn McCarty (1933–44) worked to determine the biochemical nature of the "transforming principle." They purified biochemicals (proteins, DNA, RNA, etc.) from heat-killed S cells to see which could transform live R cells into S cells.
They discovered that DNA alone from S bacteria caused R bacteria to become transformed. Protein-digesting enzymes (proteases) and RNA-digesting enzymes (RNases) did not affect transformation. Digestion with DNase did inhibit transformation. They concluded that DNA is the hereditary material, but not all biologists were convinced.
Unequivocal proof came from the experiments of Alfred Hershey and Martha Chase (1952). They worked with viruses that infect bacteria called bacteriophages.
1. Viruses grown on medium with radioactive phosphorus → contained radioactive DNA (not protein, since DNA contains P but protein does not)
2. Viruses grown on medium with radioactive sulfur → contained radioactive protein (not DNA, since protein contains S but DNA does not)
3. Radioactive phages attached to E. coli. Viral coats removed by agitating in a blender. Particles separated by centrifugation.
Bacteria infected with viruses carrying radioactive DNA were radioactive. Bacteria infected with viruses carrying radioactive protein were not radioactive. DNA is the genetic material passed from virus to bacteria.
While DNA is the predominant genetic material, RNA is the genetic material in some viruses (e.g., Tobacco Mosaic Virus, Qβ bacteriophage). A molecule that can act as genetic material must fulfill the following criteria:
Both DNA and RNA can direct their duplication. Stability is key: Griffith's "transforming principle" survived heat that killed bacteria. DNA's two complementary strands can separate by heating and come together when appropriate conditions are provided. The 2′-OH group present at every nucleotide in RNA makes it labile and easily degradable. RNA is also catalytic and hence reactive. Therefore, DNA is chemically less reactive and structurally more stable compared to RNA, making it the better genetic material.
Both DNA and RNA are able to mutate, but RNA being unstable mutates at a faster rate. Consequently, viruses with RNA genomes and shorter life spans mutate and evolve faster. RNA can directly code for protein synthesis, but DNA being more stable is preferred for storage of genetic information. For transmission of genetic information, RNA is better.
RNA was the first genetic material. There is now enough evidence to suggest that essential life processes (such as metabolism, translation, splicing, etc.) evolved around RNA. RNA used to act as a genetic material as well as a catalyst — there are important biochemical reactions in living systems catalysed by RNA catalysts and not by protein enzymes.
RNA being a catalyst was reactive and hence unstable. Therefore, DNA evolved from RNA with chemical modifications that make it more stable. DNA being double-stranded and having a complementary strand further resists changes by evolving a process of repair.
"It has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material." — Watson and Crick, 1953
The scheme suggested that the two strands would separate and act as a template for the synthesis of new complementary strands. After replication, each DNA molecule would have one parental and one newly synthesised strand. This scheme was termed semiconservative DNA replication.
It was first shown in Escherichia coli and subsequently in higher organisms. Matthew Meselson and Franklin Stahl (1958) worked on E. coli with the following steps:
The results were: after the first generation, only one band appeared (intermediate between ¹⁵N and ¹⁴N). After the second generation, two bands appeared — one intermediate and one light. This confirmed the semiconservative model.
This was also proved by Taylor and colleagues in 1958, who worked on Vicia faba (broad bean) using radioactive thymidine to detect distribution of chromosomes in root cells.
The process of replication requires a template (parental strand), the building blocks (deoxyribonucleoside triphosphates), enzyme DNA polymerase, and other components. Enzymes involved in replication include:
Unwinds the double helix by breaking hydrogen bonds between strands at the replication fork.
Relieves strain ahead of the replication fork by breaking, swiveling, and rejoining DNA strands.
Adds nucleotides in the 5′→3′ direction to the growing daughter strand using the parental strand as template.
Joins Okazaki fragments on the lagging strand by forming phosphodiester bonds.
Synthesises a short RNA primer to provide a 3′-OH group for DNA polymerase to begin synthesis.
Keeps DNA polymerase attached to the template, increasing processivity of synthesis.
Replication is continuous on the leading strand (synthesised in 5′→3′ direction toward the fork) and discontinuous on the lagging strand (synthesised away from the fork as short Okazaki fragments, later joined by DNA ligase).
DNA polymerase is extremely selective and incorporates only those nucleotides that form correct base pairs with the template. The enzyme also has a proofreading activity — it can detect incorrect nucleotides, excise them, and replace them. This results in an error rate of only about one mistake per billion nucleotides copied.
In transcription, only one strand of DNA is copied into RNA. The strand that is copied is called the template strand, and the other strand is called the coding strand. The DNA-dependent RNA polymerase catalyses the polymerisation of ribonucleotides in only one direction (5′→3′).
In bacteria, the mRNA produced does not require any processing (it is functional). In eukaryotes, the primary transcript (hnRNA) undergoes extensive processing. The coding strand has the same sequence as the mRNA (except T is replaced by U).
A transcription unit has three components:
1. A Promoter — the site where RNA polymerase binds and initiates transcription
2. The Structural Gene — the region transcribed into RNA
3. A Terminator — signals the end of transcription
In eukaryotes, the gene is split. The coding sequences called exons are interrupted by non-coding sequences called introns. Introns are removed and exons are joined to produce functional RNA by a process called splicing.
The sequence of bases in DNA (and subsequently mRNA) determines the amino acid sequence of a protein. Since there are only 4 bases and 20 amino acids to code for, the code should be a combination of bases taken together at a time. If two bases coded for one amino acid, then 4² = 16 combinations would be possible — not enough. Using three bases, 4³ = 64 combinations would be possible — more than sufficient.
This was experimentally proved by Marshall Nirenberg and Heinrich Matthaei (1961), who synthesised RNA molecules with defined combinations of bases. Nirenberg's cell-free system for protein synthesis and Severo Ochoa enzyme (polynucleotide phosphorylase) helped decipher the code.
| Amino Acid | Codons | Notes |
|---|---|---|
| Methionine (Met) | AUG | Also initiator codon |
| Phenylalanine (Phe) | UUU, UUC | — |
| Leucine (Leu) | UUA, UUG, CUU, CUC, CUA, CUG | 6 codons |
| Serine (Ser) | UCU, UCC, UCA, UCG, AGU, AGC | 6 codons |
| Tryptophan (Trp) | UGG | Single codon |
| Tyrosine (Tyr) | UAU, UAC | — |
| Cysteine (Cys) | UGU, UGC | — |
| Stop codons | UAA, UAG, UGA | No amino acid |
The relationships between genes and DNA are best understood by mutation studies. A classical example of a point mutation is a change of a single base pair in the gene for the beta-globin chain, resulting in the change of amino acid residue glutamate to valine — causing sickle cell anaemia.
Consider the statement: RAM HAS RED CAP (each word is a "triplet" like a codon).
If we insert the letter B between HAS and RED: RAM HAS BRE DCA P — the entire reading frame shifts.
Insertion or deletion of 1 or 2 bases causes frameshift mutation. However, insertion or deletion of three or its multiple bases inserts or deletes one or multiple codons, and the reading frame remains unaltered from that point onwards.
Francis Crick postulated the presence of an adapter molecule that would read the code and also link it to amino acids, because amino acids have no structural specialities to read the code uniquely. The tRNA (then called sRNA or soluble RNA) was known before the genetic code was postulated, but its role as an adapter molecule was assigned much later.
tRNA has an anticodon loop with bases complementary to the code, and it also has an amino acid acceptor end to which it binds to amino acids. tRNAs are specific for each amino acid. For initiation, there is a specific initiator tRNA. There are no tRNAs for stop codons.
In its secondary structure, tRNA looks like a clover-leaf. In actual structure, tRNA is a compact molecule that looks like an inverted L.
Translation refers to the process of polymerisation of amino acids to form a polypeptide. The order and sequence of amino acids are defined by the sequence of bases in mRNA. Amino acids are joined by a peptide bond, whose formation requires energy.
In the first phase, amino acids are activated in the presence of ATP and linked to their cognate tRNA — a process called charging of tRNA or aminoacylation of tRNA. If two such charged tRNAs are brought close enough, formation of a peptide bond between them is favoured energetically.
The cellular factory responsible for protein synthesis is the ribosome. It consists of structural RNAs and about 80 different proteins. In its inactive state, it exists as two subunits — a large subunit and a small subunit. The 23S rRNA in bacteria is the enzyme (ribozyme) for peptide bond formation.
A translational unit in mRNA is the sequence of RNA flanked by the start codon (AUG) and the stop codon, which codes for a polypeptide. mRNA also has untranslated regions (UTR) at both 5′-end (before start codon) and 3′-end (after stop codon), which are required for efficient translation.
For initiation, the ribosome binds to the mRNA at the start codon (AUG), recognised only by the initiator tRNA. During elongation, the ribosome moves from codon to codon along the mRNA. Amino acids are added one by one. At the end, a release factor binds to the stop codon, terminating translation and releasing the complete polypeptide from the ribosome.
Regulation of gene expression refers to a broad term that may occur at various levels. In eukaryotes, regulation could be exerted at:
In prokaryotes, control of the rate of transcriptional initiation is the predominant site for gene expression control. The accessibility of promoter regions is regulated by interaction of proteins with sequences called operators. Regulatory proteins can act both positively (activators) and negatively (repressors).
The elucidation of the lac operon was a result of close association between geneticist Francois Jacob and biochemist Jacques Monod. They were the first to elucidate a transcriptionally regulated system.
In the lac operon (lac refers to lactose), a polycistronic structural gene is regulated by a common promoter and regulatory genes. This arrangement is very common in bacteria and is referred to as an operon — examples include lac operon, trp operon, ara operon, his operon, val operon, etc.
Codes for the repressor of the lac operon. The "i" stands for inhibitor (not inducer). Synthesised constitutively (all the time).
Codes for β-galactosidase (β-gal), which catalyses hydrolysis of the disaccharide lactose into galactose and glucose.
Codes for permease, which increases permeability of the cell to β-galactosides, allowing lactose to enter.
Codes for transacetylase. All three enzymes are required for metabolism of lactose.
Lactose is the substrate for β-galactosidase and it regulates switching on and off of the operon — hence it is termed as the inducer. In the absence of a preferred carbon source like glucose, if lactose is provided in the growth medium, lactose is transported into the cells through the action of permease. (A very low level of expression of the lac operon has to be present in the cell all the time, otherwise lactose cannot enter the cells.)
1. The repressor is synthesised from the i gene and binds to the operator region, preventing RNA polymerase from transcribing the operon.
2. In the presence of an inducer (lactose or allolactose), the repressor is inactivated by interaction with the inducer.
3. This allows RNA polymerase access to the promoter and transcription proceeds.
Essentially, regulation of the lac operon can be visualised as regulation of enzyme synthesis by its substrate.
Regulation of the lac operon by the repressor is referred to as negative regulation. The lac operon is also under positive regulation, but that is beyond the scope at this level.
The Human Genome Project (HGP) was called a mega project. The human genome has approximately 3 × 10⁹ bp, and if the cost of sequencing were US $3 per bp, the total estimated cost would be approximately 9 billion US dollars. If the obtained sequences were stored in typed form in books (1000 letters per page, 1000 pages per book), 3,300 books would be required to store the information from a single human cell.
HGP was closely associated with the rapid development of a new area in biology called Bioinformatics.
The HGP was a 13-year project (1990–2003) coordinated by the U.S. Department of Energy and the National Institute of Health. The Wellcome Trust (U.K.) became a major partner, with contributions from Japan, France, Germany, China, and others.
Two major approaches were used:
Total DNA was isolated and converted into random fragments of relatively smaller sizes, cloned in suitable hosts using specialised vectors — BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes). Fragments were sequenced using automated DNA sequencers based on the method developed by Frederick Sanger. Sequences were arranged based on overlapping regions using specialised computer programs.
The sequence of chromosome 1 was completed in May 2006 — the last of the 24 human chromosomes (22 autosomes + X and Y) to be sequenced.
| Feature | Observation |
|---|---|
| Total base pairs | 3,164.7 million bp |
| Average gene size | 3,000 bases (largest: dystrophin at 2.4 million bases) |
| Total genes | ~30,000 (much lower than earlier estimates of 80,000–1,40,000) |
| Identical bases | 99.9% nucleotide bases are exactly the same in all humans |
| Unknown function | Over 50% of discovered genes have unknown functions |
| Protein-coding | Less than 2% of the genome codes for proteins |
| Repetitive sequences | Make up a very large portion of the genome (no direct coding functions) |
| Most genes | Chromosome 1 (2,968 genes); Fewest: Y (231 genes) |
| SNPs | ~1.4 million locations with single-base DNA differences (single nucleotide polymorphisms) |
Deriving meaningful knowledge from the DNA sequences will define research through coming decades. One of the greatest impacts of the HGP may be enabling a radically new approach to biological research. With whole-genome sequences and new high-throughput technologies, we can approach questions systematically and on a much broader scale — studying all genes in a genome, all transcripts in a particular tissue or organ, or how tens of thousands of genes and proteins work together in interconnected networks.
Besides providing clues to understanding human biology, learning about non-human organisms' DNA sequences can lead to understanding their natural capabilities applied toward solving challenges in health care, agriculture, energy production, and environmental remediation. Many non-human model organisms have also been sequenced, including bacteria, yeast, Caenorhabditis elegans, Drosophila, rice, and Arabidopsis.
As stated earlier, 99.9% of base sequence among humans is the same. Assuming the human genome as 3 × 10⁹ bp, differences would occur in approximately 0.1% of the sequences. It is these differences in DNA sequence that make every individual unique. DNA fingerprinting is a quick way to compare the DNA sequences of any two individuals.
DNA fingerprinting involves identifying differences in some specific regions in DNA called repetitive DNA. These repetitive DNA are separated from bulk genomic DNA as different peaks during density gradient centrifugation. The bulk DNA forms a major peak and the other small peaks are referred to as satellite DNA. Depending on base composition (A:T rich or G:C rich), length of segment, and number of repetitive units, satellite DNA is classified into many categories such as micro-satellites, mini-satellites, etc.
These sequences normally do not code for any proteins but form a large portion of the human genome. They show a high degree of polymorphism and form the basis of DNA fingerprinting. DNA from every tissue (blood, hair-follicle, skin, bone, saliva, sperm, etc.) from an individual shows the same degree of polymorphism, making it a very useful identification tool in forensic applications. As polymorphisms are inheritable from parents to children, DNA fingerprinting is the basis of paternity testing.
Polymorphism (variation at genetic level) arises due to mutations. Allelic sequence variation is described as DNA polymorphism if more than one variant (allele) at a locus occurs in the human population with a frequency greater than 0.01. If an inheritable mutation is observed in a population at high frequency, it is referred to as DNA polymorphism.
The probability of variation being observed in non-coding DNA sequences is higher, as mutations in these sequences may not have any immediate effect on an individual's reproductive ability. These mutations accumulate generation after generation, forming one of the bases of variability/polymorphism. There is a variety of different types of polymorphisms ranging from single nucleotide change to very large scale changes.
The technique was initially developed by Alec Jeffreys. He used a satellite DNA as probe that shows a very high degree of polymorphism, called Variable Number of Tandem Repeats (VNTR).
(i) Isolation of DNA
(ii) Digestion of DNA by restriction endonucleases
(iii) Separation of DNA fragments by electrophoresis
(iv) Transferring (blotting) of separated DNA fragments to synthetic membranes (nitrocellulose or nylon)
(v) Hybridisation using labelled VNTR probe
(vi) Detection of hybridised DNA fragments by autoradiography
The VNTR belongs to a class of satellite DNA referred to as mini-satellite. A small DNA sequence is arranged tandemly in many copy numbers. The copy number varies from chromosome to chromosome in an individual. The numbers of repeats show a very high degree of polymorphism. Consequently, the size of VNTR varies from 0.1 to 20 kb. After hybridisation with the VNTR probe, the autoradiogram gives many bands of differing sizes — a characteristic pattern for an individual's DNA. This pattern differs from individual to individual in a population, except in the case of monozygotic (identical) twins.
The sensitivity of the technique has been increased by use of polymerase chain reaction (PCR). Currently, many different probes are used to generate DNA fingerprints. DNA fingerprinting is also used in determining population and genetic diversities.
• Nucleic acids are long polymers of nucleotides. While DNA stores genetic information, RNA mostly helps in transfer and expression of information.
• DNA is chemically and structurally more stable than RNA, making it the better genetic material. However, RNA evolved first, and DNA was derived from RNA.
• The hallmark of the double-stranded helical DNA is hydrogen bonding between bases from opposite strands: A pairs with T through two H-bonds, G with C through three H-bonds.
• DNA replicates semiconservatively, guided by complementary H-bonding.
• During transcription, one strand of DNA acts as a template to direct synthesis of complementary RNA. In eukaryotes, introns are removed and exons are joined by splicing.
• The genetic code is read in triplets. tRNA acts as an adapter molecule, binding specific amino acids and pairing through H-bonding with mRNA codes via anticodons.
• Translation occurs on ribosomes, where one rRNA acts as a ribozyme for peptide bond formation.
• Regulation of transcription is the primary step for gene expression control. In bacteria, operons (like the lac operon) coordinate regulation of multiple genes.
• The Human Genome Project (1990–2003) sequenced all ~3 billion base pairs, revealing ~30,000 genes with less than 2% coding for proteins.
• DNA fingerprinting exploits VNTR polymorphisms to produce unique banding patterns for identification, forensics, and paternity testing.