Table of Contents
Introduction to Molecular Genetics 🧬
Molecular genetics focuses on how genetic information is stored, expressed, and altered at the molecular level. For USMLE purposes, this chapter introduces the core concepts that link DNA, RNA, and protein, and gives you the language and logic you will see repeatedly in questions. Detailed discussion of specific genetic diseases, inheritance patterns, or clinical syndromes is covered in other chapters, so here we stay with the molecular framework that underlies them.
Structure and Organization of DNA 🧫
Human DNA is composed of nucleotides, each containing a sugar, a phosphate group, and a nitrogenous base. The DNA double helix has two antiparallel strands, one running 5′ to 3′ and the other 3′ to 5′. Base pairing is specific, adenine pairs with thymine using two hydrogen bonds, and guanine pairs with cytosine using three hydrogen bonds. Higher G–C content increases the melting temperature of DNA because of these extra hydrogen bonds and stronger base stacking.
In eukaryotic cells, DNA is linear and organized into chromosomes. DNA wraps around histone proteins to form nucleosomes, which further coil to form chromatin. Tightly packed heterochromatin is transcriptionally inactive, while loosely packed euchromatin is transcriptionally active.
Important rule: Higher G–C content increases DNA stability and melting temperature. Heterochromatin is generally inactive, euchromatin is generally active.
DNA Replication 🔁
DNA replication is the process of copying the genetic material before cell division. It is semiconservative, which means that each daughter DNA molecule contains one parental strand and one newly synthesized strand. Replication begins at origins of replication, which are specific DNA sequences where the double helix unwinds. Eukaryotic chromosomes have multiple origins to allow faster replication of long DNA molecules.
Helicase unwinds the DNA, creating a replication fork. Single-stranded binding proteins prevent the separated strands from reannealing. DNA polymerase adds nucleotides only in the 5′ to 3′ direction, using the parental strand as a template. Because of the antiparallel nature of DNA, the leading strand is synthesized continuously in the same direction as the replication fork, while the lagging strand is synthesized discontinuously as short Okazaki fragments, which are later joined by DNA ligase.
Primase synthesizes short RNA primers that are necessary because DNA polymerases cannot start synthesis de novo, they can only extend an existing strand. Topoisomerases relieve torsional stress ahead of the replication fork by creating controlled breaks and rejoining the DNA.
In eukaryotes, telomerase extends the ends of linear chromosomes using an RNA template. This is especially important in rapidly dividing cells. Many cancer cells reactivate or overexpress telomerase to maintain indefinite proliferative capacity.
Key rules:
- DNA polymerase synthesizes DNA only in the 5′ to 3′ direction and reads the template 3′ to 5′.
- Leading strand is continuous, lagging strand is discontinuous and forms Okazaki fragments.
- Topoisomerases prevent excessive supercoiling, and telomerase maintains telomere length in eukaryotes.
DNA Repair Mechanisms 🛠️
DNA is constantly damaged by endogenous and exogenous factors, but cells possess repair systems that maintain genetic stability. Molecular defects in these systems are important in many diseases and cancers, but here we focus on the types of repair mechanisms themselves.
Base excision repair corrects small, non-helix-distorting base lesions, such as deaminated bases or oxidized bases. A specific glycosylase removes the damaged base, creating an abasic site. An endonuclease cuts the DNA backbone, the sugar phosphate is removed, and DNA polymerase fills the gap, followed by ligation.
Nucleotide excision repair deals with bulky, helix-distorting lesions, such as thymine dimers caused by ultraviolet light. A multiprotein complex recognizes the distortion, removes a short single-stranded DNA segment containing the lesion, and DNA polymerase fills in the correct sequence followed by ligation.
Mismatch repair corrects replication errors that escape proofreading, such as base-base mismatches or small insertion-deletion loops. The newly synthesized strand is recognized, a segment containing the mismatch is removed, then DNA polymerase and ligase restore the correct sequence.
Double-strand break repair pathways include nonhomologous end joining, which directly ligates broken DNA ends, and homologous recombination, which uses a homologous DNA sequence as a template for accurate repair. Homologous recombination is more accurate but requires a sister chromatid and is therefore mainly active in late S and G2 phases.
High-yield rule:
Base excision repair fixes small base damage, nucleotide excision repair fixes bulky lesions, mismatch repair corrects replication errors, and double-strand break repair relies on nonhomologous end joining or homologous recombination.
From DNA to RNA: Transcription 📝
Transcription is the synthesis of RNA from a DNA template. It begins at promoter regions, which are specific DNA sequences near the start site of a gene. In eukaryotes, the core promoter often contains a TATA box. Transcription factors bind promoter elements and help position RNA polymerase correctly.
RNA polymerase synthesizes RNA in the 5′ to 3′ direction, reading the DNA template strand 3′ to 5′. Only one strand of DNA serves as the template for a particular gene, and the RNA sequence is complementary to this template and almost identical to the coding strand, except that RNA uses uracil instead of thymine.
In eukaryotes, there are three major nuclear RNA polymerases. RNA polymerase I makes most ribosomal RNA, RNA polymerase II synthesizes messenger RNA and some small nuclear RNAs, and RNA polymerase III produces transfer RNA and some small RNAs. This division of labor is specific to eukaryotes.
Transcription ends at terminator sequences. The newly formed primary RNA transcript in eukaryotes is initially pre-mRNA, which requires extensive processing before it becomes mature mRNA.
Core facts:
- RNA polymerase synthesizes RNA 5′ to 3′ from a 3′ to 5′ DNA template.
- In eukaryotes, RNA polymerase II is responsible for mRNA synthesis.
- Promoters control where transcription starts, terminators control where it stops.
RNA Processing in Eukaryotes ✂️
Pre-mRNA in eukaryotic cells must undergo several processing steps before it is translated. These steps occur in the nucleus and are crucial for mRNA stability, nuclear export, and correct translation.
At the 5′ end, a 7-methylguanosine cap is added. This 5′ cap protects mRNA from degradation and is required for ribosome binding during translation initiation. At the 3′ end, polyadenylate polymerase adds a poly(A) tail, a sequence of many adenine residues. The poly(A) tail also protects mRNA from degradation and plays a role in nuclear export and translation efficiency.
One of the most important processing events is splicing. Eukaryotic genes are usually interrupted by noncoding sequences called introns, which lie between coding segments called exons. The spliceosome, a complex of proteins and small nuclear RNAs, recognizes specific splice sites at intron-exon boundaries and removes the introns while joining the exons.
Alternative splicing allows a single gene to produce multiple mRNA isoforms by combining exons in different patterns. This greatly increases protein diversity without increasing the number of genes.
Essential rules:
- Eukaryotic pre-mRNA receives a 5′ cap and a 3′ poly(A) tail and undergoes splicing.
- Introns are removed, exons are joined.
- Alternative splicing from one gene can produce multiple different proteins.
Types of RNA and Their Roles 📦
Molecular genetics distinguishes several functional classes of RNA. Messenger RNA carries the genetic code from DNA to the ribosome. Transfer RNA delivers specific amino acids to the ribosome according to the codon sequence in the mRNA. Ribosomal RNA forms the structural and catalytic core of ribosomes, which perform protein synthesis.
In addition, small nuclear RNAs are components of the spliceosome and are involved in pre-mRNA splicing. MicroRNAs and small interfering RNAs are examples of regulatory RNAs that modulate gene expression post-transcriptionally, often by binding to target mRNAs and promoting their degradation or inhibiting their translation.
Although all of these RNAs are transcribed from DNA, only mRNA directly encodes proteins. The others are functional RNAs that perform structural or regulatory tasks.
From RNA to Protein: Translation 🍽️
Translation is the process of protein synthesis from mRNA at the ribosome. The genetic code is read in sets of three nucleotides called codons. Each codon specifies either an amino acid or a stop signal. With four nucleotides and three positions, the total number of possible codons is $4^3 = 64$. Most amino acids are encoded by more than one codon, which means the genetic code is degenerate.
Transfer RNA molecules have an anticodon loop that base pairs with the codon in mRNA and an acceptor stem that binds a specific amino acid. Aminoacyl tRNA synthetases are enzymes that attach the correct amino acid to its corresponding tRNA using ATP. This is a critical specificity step in translation. Errors here lead to incorporation of the wrong amino acid.
Ribosomes read the mRNA from the 5′ to 3′ direction. Translation usually begins at an AUG start codon, which codes for methionine in eukaryotes. Stop codons signal termination and do not encode an amino acid. During elongation, the ribosome catalyzes peptide bond formation between amino acids, and the polypeptide chain grows until a stop codon is reached.
High-yield translation rules:
- The genetic code is read in non-overlapping triplets called codons.
- Standard start codon is AUG. Stop codons signal termination and encode no amino acid.
- Aminoacyl tRNA synthetases ensure correct amino acid attachment to tRNAs, which is crucial for translation accuracy.
Regulation of Gene Expression 🧪
Molecular genetics also studies how cells control gene expression at different steps between DNA and protein. In eukaryotes, chromatin structure is a major regulatory level. Acetylation of histone tails generally opens chromatin and promotes transcription, while deacetylation and certain methylation patterns promote a more closed, transcriptionally inactive state.
Regulation also occurs at the level of transcription through activators, repressors, and enhancers. Enhancers are DNA sequences that can increase transcription of a gene, often acting at a distance and independent of orientation, through binding of specific transcription factors. Silencers exert the opposite effect.
Post-transcriptional regulation includes control of splicing patterns, mRNA stability, and export from the nucleus. MicroRNAs can reduce translation of specific mRNAs by binding to complementary sequences. At the translational level, cells can modify initiation, elongation, or the availability of ribosomes. Post-translational modifications such as phosphorylation, glycosylation, and ubiquitination alter protein function, localization, or stability, although these details will be discussed more in biochemistry.
Mutations and Their Molecular Consequences 🧪⚠️
Mutations are heritable changes in DNA sequence. At the molecular level, point mutations involve alteration of a single nucleotide. A silent mutation changes a codon but does not change the amino acid due to redundancy of the genetic code. A missense mutation changes one amino acid to another. A nonsense mutation converts a codon for an amino acid into a stop codon, which usually leads to a truncated protein.
Insertions or deletions of nucleotides can produce frameshift mutations when the number of nucleotides added or removed is not a multiple of three. This shifts the reading frame downstream of the mutation, often leading to a completely different amino acid sequence and a premature stop codon.
Large-scale mutations include insertions, deletions, inversions, or translocations of bigger DNA segments, and whole chromosomal changes. Those structural alterations belong more to cytogenetics and inheritance chapters, but they all originate from molecular-level events such as double-strand breaks and misrepair.
Key mutation principles:
- Silent: DNA change, same amino acid.
- Missense: DNA change, different amino acid.
- Nonsense: DNA change, premature stop codon.
- Frameshift: insertion or deletion not in multiples of three alters the reading frame.
Recombinant DNA and Basic Molecular Techniques 🧪🧫
Molecular genetics provides the conceptual basis for many laboratory techniques you will encounter in USMLE-style questions. These techniques rely on standard properties of nucleic acids such as base pairing, enzymatic cutting, or polymerase activity.
Restriction enzymes cut DNA at specific short sequences and generate fragments that can be ligated to other DNA pieces, which allows the construction of recombinant DNA molecules. Plasmids are circular DNA molecules often used as vectors to carry and replicate foreign DNA in bacteria. A typical plasmid vector contains an origin of replication, a selectable marker such as an antibiotic resistance gene, and a multiple cloning site where foreign DNA can be inserted.
Polymerase chain reaction, or PCR, is used to amplify specific DNA sequences in vitro. It uses sequence-specific primers, heat-stable DNA polymerase, and cycles of denaturation, annealing, and extension. Quantitative PCR can measure the amount of DNA or cDNA and therefore assess gene expression levels indirectly.
Gel electrophoresis separates DNA, RNA, or proteins based on size and charge. Smaller DNA fragments move faster through the gel matrix. Blotting techniques transfer nucleic acids or proteins from a gel to a membrane and then use labeled probes or antibodies for detection. DNA sequencing allows determination of the exact nucleotide order within a DNA fragment.
While the technical details of each method will be explored in other contexts, what is unique to molecular genetics is the understanding that all of these methods exploit the central properties of DNA and RNA, such as complementary base pairing, specific recognition sequences, and enzymatic synthesis in the 5′ to 3′ direction.