Biology — Std 12
🧬

Gene: Its Nature, Expression and Regulation

Ch. 2Std 12

Easy Overview

So DNA is basically the instruction manual for your entire body. Every cell in your body — skin, liver, brain, bone — contains the same DNA, the same 3 billion base pairs, the same ~20,000 genes. Yet a skin cell is nothing like a liver cell. How does a molecule — literally just a chemical — encode the information needed to build and run a human being? And how do different cells read different parts of the manual? That's what this chapter is about: the nature of the gene, how it expresses itself, and how that expression is regulated. The story of DNA begins with its structure. In 1953, James Watson and Francis Crick, using X-ray crystallography data from Rosalind Franklin and Maurice Wilkins, proposed the double helix model. It's two strands of nucleotides running in opposite directions (antiparallel), held together by hydrogen bonds between complementary base pairs: adenine with thymine (two hydrogen bonds) and guanine with cytosine (three hydrogen bonds). The sugar-phosphate backbones form the rails, and the base pairs form the steps. This structure immediately suggested how DNA could replicate: just unzip the two strands and use each as a template. The central dogma of molecular biology, proposed by Francis Crick, states that genetic information flows from DNA to RNA to Protein. DNA is transcribed into mRNA, which is then translated into a polypeptide chain. This is the fundamental information pathway in all living cells. There are exceptions (retroviruses reverse the flow, using RNA to make DNA), but for most life, this dogma holds. DNA replication happens before cell division. The double helix unwinds with the help of enzymes like helicase, and DNA polymerase reads the template strand and adds complementary nucleotides. Because the two strands are antiparallel, replication is continuous on the leading strand but discontinuous on the lagging strand (Okazaki fragments). The result is two identical DNA molecules, each with one old strand and one new strand — semi-conservative replication. Meselson and Stahl proved this elegantly using nitrogen isotopes. Transcription is the process of making an RNA copy of a gene. RNA polymerase binds to the promoter region and moves along the DNA template, synthesizing a complementary RNA strand. In eukaryotes, the initial transcript (pre-mRNA) undergoes processing: 5' capping, 3' polyadenylation, and splicing (removing introns and joining exons). This mature mRNA then leaves the nucleus for the cytoplasm. Translation happens on ribosomes. The mRNA is read in triplets called codons, each coding for a specific amino acid. tRNA molecules act as adapters — they have an anticodon that pairs with the mRNA codon and carry the corresponding amino acid. The ribosome facilitates codon-anticodon pairing and catalyzes peptide bond formation. The genetic code is degenerate (multiple codons for the same amino acid), nearly universal, and read continuously without overlaps. Gene regulation determines which genes are expressed when. In bacteria, the lac operon is a classic model: when lactose is absent, a repressor protein binds to the operator and blocks transcription. When lactose is present, it binds the repressor and inactivates it, allowing transcription. In eukaryotes, regulation is more complex, involving transcription factors, enhancers, silencers, chromatin remodeling, and epigenetic modifications. Not all DNA codes for protein. Only about 1.5% of the human genome actually codes for proteins. The rest includes introns, regulatory sequences, repetitive DNA, satellite DNA, and mysterious 'junk DNA' that may have regulatory or structural functions. Introns are transcribed but spliced out before translation. Their discovery was surprising — why would cells waste energy making RNA they're going to chop up? We now know introns allow alternative splicing (one gene producing multiple proteins) and may have evolutionary significance. Mutations are changes in the DNA sequence. Point mutations affect a single base pair — silent (same amino acid), missense (different amino acid), or nonsense (premature stop codon). Frameshift mutations (insertions or deletions not in multiples of three) shift the reading frame and usually ruin the protein. Mutations can be spontaneous or induced by mutagens (radiation, chemicals). Cells have DNA repair mechanisms, but unrepaired mutations can cause cancer or genetic disease.

DNA structure — the Watson-Crick model

DNA is a double-stranded helix with two polynucleotide chains running antiparallel (one 5' to 3', the other 3' to 5'). Each nucleotide consists of a phosphate group, deoxyribose sugar, and a nitrogenous base (A, T, G, or C). The backbones are sugar-phosphate chains on the outside; bases project inward and pair via hydrogen bonds — A with T (2 bonds), G with C (3 bonds). This complementary base pairing (Chargaff's rule) is crucial for replication and transcription. The double helix has a major groove and minor groove where proteins bind.

Genomic organization — how DNA is packaged

If stretched out, your DNA would be about 2 meters long — yet it fits inside a nucleus just 6 microns wide. DNA wraps around histone proteins (H2A, H2B, H3, H4) to form nucleosomes — the 'beads on a string.' Nucleosomes coil into a 30-nm fiber, which loops and folds into chromosomes. Histone H1 holds the nucleosome structure together. This packaging also controls gene expression — tightly packed DNA (heterochromatin) is inactive; loosely packed (euchromatin) is active.

DNA replication — copying the blueprint

Replication is semi-conservative: each daughter DNA molecule has one parental strand and one newly synthesized strand. It starts at origins of replication. Helicase unwinds the DNA, single-strand binding proteins stabilize the strands, and topoisomerase relieves supercoiling. Primase lays down RNA primers. DNA polymerase III adds nucleotides 5' to 3'. On the leading strand, synthesis is continuous. On the lagging strand, it's discontinuous — Okazaki fragments joined by DNA ligase. Meselson and Stahl proved semi-conservative replication in 1958 using N-15 isotopes.

Transcription — DNA to RNA

Transcription produces an RNA copy of a gene. RNA polymerase binds to the promoter region (upstream of the gene) and unwinds the DNA. It reads the template strand 3' to 5' and synthesizes RNA 5' to 3', using complementary base pairing (A to U, T to A, G to C, C to G). Transcription stops at the terminator sequence. In bacteria, a single RNA polymerase does all transcription. In eukaryotes, RNA polymerase I (rRNA), II (mRNA), and III (tRNA) each handle different genes.

RNA processing — editing the transcript

In eukaryotes, the initial transcript (pre-mRNA) is processed before translation. A 5' cap (modified guanine) is added for stability and ribosome binding. A poly-A tail (100-250 adenine nucleotides) is added at the 3' end for stability and export. Introns (non-coding sequences) are removed and exons (coding sequences) are joined together by the spliceosome (a complex of snRNPs). Alternative splicing allows one gene to produce multiple protein isoforms.

The genetic code — the language of life

The genetic code is a set of 64 codons (three-nucleotide sequences) that specify amino acids. Three stop codons (UAA, UAG, UGA) signal termination; AUG is both the start codon and codes for methionine. The code has four key features: (1) Triplet — three bases per amino acid. (2) Degenerate — multiple codons for most amino acids (usually differing in the third base — the wobble position). (3) Universal — same code across almost all organisms. (4) Non-overlapping and comma-free — read sequentially without gaps.

Translation — from mRNA to protein

Translation occurs on ribosomes (rRNA + proteins). The small ribosomal subunit binds mRNA and finds the start codon (AUG). The initiator tRNA (carrying methionine) binds. Then the large subunit joins. The ribosome has three sites: A (aminoacyl — incoming tRNA binds), P (peptidyl — growing peptide chain), E (exit — empty tRNA leaves). Amino acids are linked by peptide bonds. Elongation continues until a stop codon enters the A site; then release factors trigger disassembly.

tRNA — the adapter molecule

tRNA is the interpreter between mRNA codons and amino acids. It has an anticodon (three bases complementary to the mRNA codon) at one end and the corresponding amino acid attached at the 3' end (acceptor stem). tRNA has a cloverleaf secondary structure with three loops — D loop, anticodon loop, and T-psi-C loop. The enzyme aminoacyl-tRNA synthetase attaches the correct amino acid to each tRNA with high specificity.

The lac operon — bacterial gene regulation

The lac operon is a cluster of genes in E. coli that digest lactose. It consists of three structural genes (lacZ, lacY, lacA), a promoter, an operator, and a regulatory gene (lacI). When lactose is absent, the lacI repressor binds to the operator, blocking transcription. When lactose is present, it's converted to allolactose, which binds the repressor and inactivates it. RNA polymerase can then transcribe the operon. This is inducible regulation. Glucose also regulates via catabolite repression (cAMP-CAP complex).

Eukaryotic gene regulation — layers of control

Eukaryotes regulate gene expression at multiple levels. (1) Chromatin level: DNA methylation (silences genes) and histone acetylation (activates genes). (2) Transcriptional level: transcription factors bind enhancers/silencers to activate or repress transcription. (3) Post-transcriptional level: alternative splicing, RNA editing, mRNA stability. (4) Translational level: regulatory proteins or microRNAs block translation. (5) Post-translational level: protein modification, folding, and degradation. This multi-layered control allows the same genome to produce hundreds of different cell types.

RNA interference (RNAi) — gene silencing

RNAi is a natural mechanism where small RNA molecules (siRNA, miRNA) bind to complementary mRNA and prevent translation or trigger degradation. miRNA (microRNA) is transcribed from the genome and regulates normal gene expression. siRNA (small interfering RNA) often comes from foreign RNA (viruses) and defends the genome. Both are processed by Dicer and loaded into the RISC complex, which finds and silences target mRNAs. RNAi is a powerful research tool and has therapeutic applications.

Mutations — changes in the genetic code

Mutations are permanent changes in DNA sequence. Point mutations affect a single base: silent (codon still codes for same amino acid due to degeneracy), missense (different amino acid — e.g., sickle cell anemia GAG to GTG, Glu to Val), nonsense (premature stop codon — truncated protein). Frameshift mutations (insertion/deletion of non-multiple-of-3 bases) shift the reading frame, altering all downstream codons — usually catastrophic. Mutagens include radiation (UV, X-rays), chemicals, and viruses.

DNA repair — the cell's proofreading system

Cells have multiple repair mechanisms. (1) Mismatch repair: after replication, enzymes scan for mismatched bases and correct them. (2) Excision repair (NER, BER): damaged bases are cut out and replaced. (3) Direct repair: enzymes directly reverse damage (photolyase fixes thymine dimers using light). (4) Double-strand break repair: non-homologous end joining (NHEJ) or homologous recombination. Defects in repair cause diseases — xeroderma pigmentosum (defect in NER) and HNPCC (colon cancer).

Introns, exons, and alternative splicing

Genes in eukaryotes are split — coding sequences (exons) interrupted by non-coding sequences (introns). Both are transcribed, but introns are spliced out. The spliceosome recognizes splice sites, cuts out introns, and joins exons. Alternative splicing allows one gene to produce multiple proteins by including different exon combinations. The human Dscam gene can theoretically produce 38,016 different isoforms! This explains how our ~20,000 genes can produce hundreds of thousands of different proteins.

Satellite DNA and repetitive sequences

Large portions of eukaryotic genomes consist of repetitive sequences. Satellite DNA: highly repetitive tandem repeats (10-100 bp) found in heterochromatin, used in DNA fingerprinting (VNTRs). Minisatellites: 10-60 bp repeats, also used in fingerprinting. Microsatellites (STRs): 2-6 bp repeats. Transposons: 'jumping genes' that move around the genome — they can cause mutations but also drive evolution. Alu elements are the most common transposons in humans, making up ~10% of our genome.

Key Points

  • DNA is a double helix with antiparallel strands; A=T (2 H-bonds), G=C (3 H-bonds) — Chargaff's rule
  • DNA packaging: nucleosomes (histones) to 30 nm fiber to loops to chromosomes
  • Semi-conservative replication: each new DNA has one old + one new strand (Meselson-Stahl experiment)
  • DNA replication enzymes: helicase, primase, DNA polymerase III, DNA ligase
  • Central dogma: DNA to RNA (transcription) to Protein (translation)
  • RNA polymerase binds promoter, reads template 3' to 5', synthesizes RNA 5' to 3'
  • Pre-mRNA processing: 5' cap, 3' poly-A tail, splicing (removing introns)
  • Genetic code: 64 codons, 61 sense + 3 stop; degenerate, universal, non-overlapping, triplet
  • Translation: ribosome reads mRNA codons, tRNA brings amino acids, polypeptide chain formed
  • Lac operon: repressor binds operator in absence of lactose; allolactose inactivates repressor
  • Eukaryotic gene regulation: chromatin remodeling, transcription factors, alternative splicing, miRNA
  • RNAi: Dicer processes dsRNA to siRNA/miRNA, RISC binds target mRNA, silencing occurs
  • Point mutations: silent (same AA), missense (different AA), nonsense (premature stop)
  • Frameshift mutations: insertion/deletion not multiple of 3, shifts reading frame, usually severe
  • DNA repair: mismatch repair, excision repair (NER/BER), direct repair, double-strand break repair
  • Alternative splicing: one gene to multiple proteins; explains proteome diversity
  • Satellite DNA: repetitive sequences used in DNA fingerprinting (VNTRs, STRs)

Practice Questions

  • Describe the Watson-Crick model of DNA structure. How does complementary base pairing support both replication and transcription?
  • Explain semi-conservative replication. How did Meselson and Stahl's experiment prove it?
  • Trace the path of a gene from DNA to a functional protein. Include transcription, RNA processing, and translation.
  • What is the genetic code? Explain its four key features. Why is it called degenerate?
  • Explain the lac operon model. How does it respond to the presence and absence of lactose?
  • Distinguish between missense, nonsense, silent, and frameshift mutations. Give an example of a disease caused by each type.
  • What is alternative splicing? How does it allow a single gene to produce multiple proteins?
  • Describe the process of eukaryotic gene regulation at the transcriptional and post-transcriptional levels.