Kahibaro
Discord Login Register

2.3.2 Molecular Biology

Overview of Molecular Biology 🧬

Molecular biology in USMLE Step 1 focuses on how information flows inside the cell, from DNA to RNA to protein. You will see questions that link these molecular processes to diseases, drugs, and laboratory methods. The goal in this chapter is to understand how DNA is organized and replicated, how RNA is made and processed, and how proteins are synthesized and regulated, together with the key lab techniques that test-makers love.

Molecular biology is highly conceptual but also very detail oriented. Many questions are built around subtle defects in enzymes, polymerases, or repair proteins. You must be comfortable with vocabulary and with the directionality and dependencies of each step in the central dogma of molecular biology.

Central dogma of molecular biology:
DNA $\to$ RNA $\to$ Protein
Genetic information is usually stored in DNA, transcribed into RNA, and translated into protein.

DNA Structure and Organization 🧫

At the USMLE level, you must distinguish structural features of DNA that explain both normal function and susceptibility to damage. Human DNA is double stranded, antiparallel, and uses complementary base pairing. Each strand runs in opposite directions, one 5' to 3' and the other 3' to 5'. Phosphodiester bonds link nucleotides between the 3' hydroxyl of one sugar and the 5' phosphate of the next.

Hydrogen bonds hold the two strands together. Adenine pairs with thymine using two hydrogen bonds, while guanine pairs with cytosine using three hydrogen bonds. Regions rich in GC base pairs are therefore harder to separate and are important at some regulatory and structural sites. You should remember that DNA stability increases with higher GC content and stronger hydrogen bonding.

Eukaryotic DNA is linear and is packaged into chromosomes. The basic unit of packaging is the nucleosome, a segment of DNA wrapped around a core of histone proteins. Histones are rich in positively charged amino acids, such as lysine and arginine, which interact with the negatively charged phosphate backbone of DNA. Nucleosomes further coil and fold to form chromatin. Chromatin that is more open and transcriptionally active is called euchromatin, while tightly packed and transcriptionally silent chromatin is called heterochromatin. This structural organization is crucial for gene regulation.

Chemical modifications of histones, especially acetylation, and of DNA itself, especially methylation of cytosine residues, modulate how accessible the DNA is. Increased DNA methylation usually leads to gene silencing, while histone acetylation generally relaxes chromatin and promotes transcription. Many cancers involve abnormal patterns of DNA methylation.

Another structural feature that appears often in questions is the telomere, a repetitive DNA sequence at the ends of linear chromosomes. Telomeres protect chromosome ends from degradation and from being recognized as DNA breaks. Each cell division tends to shorten telomeres, and the enzyme telomerase can restore them in certain cells such as germ cells and some stem cells. Malignant cells frequently reactivate telomerase.

DNA Replication 🧪

DNA replication is semiconservative, which means each daughter DNA molecule contains one parental strand and one newly synthesized strand. Replication begins at specific origin sites and proceeds in both directions. In eukaryotes, there are multiple origins of replication on each chromosome that allow rapid copying of the large genome. At each origin, local unwinding creates a replication bubble with two replication forks.

At the replication fork, helicase unwinds the double helix, and single strand binding proteins stabilize the separated strands. Ahead of the fork, topoisomerases relieve torsional stress produced by unwinding. In bacteria, DNA gyrase is a relevant topoisomerase targeted by fluoroquinolone antibiotics. In eukaryotes, related topoisomerases are important for preventing DNA tangling and breakage.

DNA polymerase can only add nucleotides to an existing 3' hydroxyl group and synthesizes new DNA in the 5' to 3' direction. A primase first lays down a short RNA primer that provides the free 3' hydroxyl. On the leading strand, DNA synthesis is continuous because it proceeds in the same direction as the unwinding fork. On the lagging strand, synthesis is discontinuous and produces short DNA fragments called Okazaki fragments. Each fragment requires its own RNA primer.

In prokaryotes, different DNA polymerases have specialized roles. In eukaryotes, a family of polymerases handles leading and lagging strands, primer removal, and some repair. All polymerases require a template and a primer and have intrinsic accuracy supported by proofreading activity.

Key rule for directionality:
DNA synthesis always occurs by adding nucleotides to the 3' end.
Template is read 3' to 5', new strand is synthesized 5' to 3'.

Proofreading depends on 3' to 5' exonuclease activity. When a wrong base is inserted, it is removed and replaced. Additional repair mechanisms, such as mismatch repair and nucleotide excision repair, correct errors that escape proofreading or remove damaged bases created by ultraviolet light or chemicals. Loss of these systems is associated with specific hereditary cancer syndromes and increases mutation rates.

Because eukaryotic chromosomes are linear, their ends pose a challenge for replication. Telomerase solves this by extending the 3' end of the template strand using an internal RNA template, then allowing normal DNA polymerase to fill in. Without effective telomerase, repeated divisions cause progressive telomere shortening, which limits cell replication potential.

Transcription and RNA Processing 🧾

Transcription is the synthesis of RNA from a DNA template. It is carried out by RNA polymerases that read DNA in the 3' to 5' direction and synthesize RNA in the 5' to 3' direction. Transcription starts at promoters, specific DNA sequences where transcription factors and RNA polymerase bind. The most commonly tested eukaryotic promoter elements include the TATA box and similar motifs. Promoter mutations can reduce or eliminate gene expression.

In eukaryotes, different RNA polymerases produce different classes of RNA. One polymerase mainly synthesizes ribosomal RNA, another produces messenger RNA, and a third produces transfer RNA and some small RNAs. Many antibiotics selectively inhibit the bacterial RNA polymerase and so interfere with transcription in prokaryotes without blocking eukaryotic enzymes.

Transcription proceeds through initiation, elongation, and termination. Regulatory sequences, such as enhancers and silencers, can be located far from the gene itself and can strongly increase or decrease transcription when bound by specific transcription factors. These sequences are used in tissue specific expression and in response to hormones or signals.

The primary RNA transcript in eukaryotes called pre mRNA undergoes extensive processing before becoming mature mRNA. One early modification is addition of a 5' cap a modified guanine nucleotide that protects RNA from degradation and helps initiate translation. At the 3' end, a polyadenylate tail of many adenine residues is added after cleavage at a specific site. This poly A tail stabilizes the RNA and influences export from the nucleus to the cytoplasm.

Introns are noncoding sequences that interrupt coding exons. They are removed by the spliceosome, a complex of small nuclear RNAs and proteins that recognizes specific splice sites. Mistakes in splicing can create nonfunctional proteins or even new, abnormal proteins. Many diseases arise from splicing defects, and some antisense therapies work by modifying splicing patterns.

Alternative splicing permits a single gene to generate multiple different mRNAs and therefore different proteins by selecting different combinations of exons. This greatly increases protein diversity without increasing gene number and is a major way to vary function in different tissues. Many exam questions use alternative splicing to explain why a mutation has tissue specific effects.

Mature mRNA contains a 5' untranslated region, a coding region beginning with a start codon, a stop codon, and a 3' untranslated region. It is exported from the nucleus through nuclear pores and then used for translation.

Genetic Code and tRNA 🎯

The genetic code links nucleotide triplets in mRNA called codons to amino acids in proteins. Each codon consists of three bases and specifies one amino acid or a stop signal. There are 64 possible codons, but only 20 standard amino acids, so the code is redundant. Many amino acids are encoded by more than one codon, usually differing in the third base. This helps buffer against some mutations, because a change in the third position can still encode the same amino acid.

Start and stop codons are important to know. The main start codon in eukaryotes codes for methionine. Stop codons do not encode any amino acid but instead signal termination of translation. Mutations that create premature stop codons produce truncated proteins, and mutations that destroy normal stop codons cause abnormally long proteins.

Transfer RNA is the adaptor molecule that brings amino acids to the ribosome during translation. Each tRNA has an anticodon loop that base pairs with the codon on mRNA and a 3' end where the corresponding amino acid is attached. The 3' terminal sequence is universally conserved and provides the attachment site. The correct amino acid is loaded onto each tRNA by a specific aminoacyl tRNA synthetase. These enzymes are extremely important for translation accuracy. If the wrong amino acid is attached, the tRNA will still pair with the codon, and the ribosome will incorporate the wrong amino acid into the growing peptide.

tRNA molecules are folded into a characteristic cloverleaf secondary structure with several arms. In addition to the anticodon loop and the acceptor stem at the 3' end, there are extra loops that help interact with the ribosome and with the synthetase. Some antibiotics and toxins interfere with aminoacyl tRNA synthetases or with tRNA function and therefore block protein synthesis.

Translation and Protein Synthesis 🧬

Translation is the process of building a polypeptide chain from an mRNA template. It occurs on ribosomes, large ribonucleoprotein complexes that exist as a small and a large subunit. In prokaryotes and eukaryotes, ribosomes have different sizes and rRNA compositions, which allows selective targeting by some drugs.

Translation has three stages, initiation, elongation, and termination. Initiation begins when the small ribosomal subunit binds to the mRNA near the 5' end and scans for the start codon. In eukaryotes, this process requires initiation factors and the 5' cap. The initiator tRNA carrying methionine pairs with the start codon in the P site of the ribosome. Then the large subunit joins, and the full ribosome is assembled.

The ribosome has three binding sites for tRNA, named A, P, and E. During elongation, a charged tRNA carrying the next amino acid binds to the A site according to base pairing between its anticodon and the codon in the mRNA. A peptidyl transferase activity, part of the large subunit, catalyzes formation of a peptide bond between the amino acid in the A site and the growing peptide in the P site. The peptide is transferred to the tRNA in the A site. The ribosome then translocates along the mRNA so that the tRNA with the growing chain shifts to the P site, and the now empty tRNA moves to the E site to exit. This cycle repeats, adding amino acids one by one.

Several antibiotics act at well defined points in this translation cycle, often by binding to either the 30S or 50S bacterial ribosomal subunit, blocking tRNA binding, peptide bond formation, or translocation. The USMLE expects you to connect the site of action of these drugs to their effect on protein synthesis.

When a stop codon enters the A site, no tRNA binds. Instead, a release factor protein recognizes the stop codon and promotes hydrolysis of the bond between the polypeptide and the tRNA in the P site. The completed polypeptide is released, and the ribosomal subunits dissociate from the mRNA.

Multiple ribosomes can translate a single mRNA simultaneously, forming a structure known as a polyribosome. This increases efficiency and allows rapid production of larger amounts of protein.

After translation, many proteins undergo post translational modifications. These include folding with the help of chaperone proteins, cleavage of signal peptides, formation of disulfide bonds, glycosylation, phosphorylation, and other chemical changes. Some of these modifications are critical for correct function and for proper localization of proteins within the cell.

Gene Regulation in Eukaryotes 🔁

Eukaryotic gene expression is tightly regulated at several levels. Regulation at the level of transcription is especially important. Promoters, enhancers, and silencers, together with transcription factors and coactivators, determine when a gene is transcribed and at what rate. A single transcription factor can regulate many genes that share a common binding element in their promoters or enhancers.

Chromatin structure is a central component of gene regulation. DNA methylation, histone modification, and chromatin remodeling complexes all influence whether transcription machinery can access a gene. Methylation of promoter CpG islands is a common feature of gene silencing. Histone acetylation usually opens chromatin and is associated with active gene transcription. Drugs that inhibit histone deacetylases or DNA methyltransferases can reactivate silenced genes and are used in some malignancies.

Post transcriptional regulation includes alternative splicing, RNA editing, control of mRNA stability, and microRNA mediated repression. MicroRNAs are small noncoding RNAs that pair imperfectly with target mRNAs, usually in the 3' untranslated region, and reduce their translation or lead to their degradation. Altered microRNA expression patterns are found in many cancers and other diseases.

Gene expression is also controlled at the level of translation initiation. Some mRNAs are stored in a silent state and only translated when conditions change, for example in response to nutrient deprivation or stress. The presence of specific sequence elements in the untranslated regions can recruit proteins that either promote or repress translation.

Finally, protein abundance can be regulated by controlled degradation. Many proteins are tagged with ubiquitin and then destroyed by the proteasome. This system rapidly removes misfolded or damaged proteins and also regulates levels of key signaling proteins and cell cycle regulators.

Recombinant DNA and Cloning Techniques 🧫

USMLE questions frequently involve recombinant DNA technology, which combines DNA from different sources in order to clone, express, or modify genes. A central element is the use of restriction endonucleases, enzymes that cut DNA at specific short sequences, often creating sticky ends that can base pair with complementary sequences. By cutting both a plasmid vector and a DNA fragment with the same restriction enzyme, their sticky ends can anneal and be joined by DNA ligase to create recombinant DNA.

Plasmids are small circular DNA molecules that replicate independently in bacteria and can carry inserted foreign DNA. In a typical cloning experiment, a gene of interest is inserted into a plasmid that also contains an antibiotic resistance gene. Bacteria are transformed with this plasmid, and only those that take up the plasmid survive on antibiotic containing media. The bacteria then replicate, amplifying the inserted gene. This principle is used to produce large quantities of DNA or protein.

Complementary DNA libraries are generated from mature mRNA using reverse transcriptase to synthesize DNA copies. These libraries contain only expressed sequences without introns. Because bacteria cannot remove introns, cDNA constructs are necessary for bacterial expression of eukaryotic proteins. Genomic DNA libraries, in contrast, contain fragments of all genomic sequences, with introns and regulatory regions included.

Recombinant vectors can be designed for different purposes. Some are optimized for protein expression, others for high copy replication, and still others for gene delivery to mammalian cells, including viral vectors. Basic understanding of cloning, ligation, and selection is necessary for interpreting experimental vignettes.

Polymerase Chain Reaction (PCR) and DNA Analysis 🧪

PCR is a powerful technique that amplifies specific DNA sequences exponentially. It requires template DNA, two primers that flank the region of interest, heat stable DNA polymerase, deoxynucleotide triphosphates, and a suitable buffer with magnesium ions. The reaction undergoes repeated cycles of denaturation, annealing, and extension in a thermal cycler.

In the denaturation step, high temperature separates the two DNA strands. During annealing, the temperature is lowered so that the primers can base pair with their complementary target sequences. In the extension step, the polymerase extends each primer by adding nucleotides to the 3' end, synthesizing new DNA. Each cycle doubles the number of DNA copies, so the amount grows as $2^n$, where $n$ is the number of cycles, until reagents become limiting.

PCR can detect very small amounts of pathogen DNA, identify specific mutations, or amplify DNA for sequencing or cloning. Variants of PCR, such as reverse transcriptase PCR, begin from RNA templates that are first converted to cDNA. Real time quantitative PCR measures the amount of DNA produced in each cycle using fluorescent signals and allows quantification of initial template amounts.

Other important DNA analysis methods include Southern blotting, which detects specific DNA fragments separated by electrophoresis and then transferred to a membrane. Northern blotting similarly detects RNA, and Western blotting detects proteins using labeled antibodies. These classic techniques often appear in questions to connect shifts in band size or intensity with deletions, insertions, or changes in expression level.

Gel electrophoresis separates nucleic acids or proteins based on size, using an electric field that moves charged molecules through a porous gel matrix. Smaller fragments migrate faster and farther. Restriction fragment length polymorphism analysis and some forensic methods rely on characteristic band patterns on gels.

DNA Sequencing and Genomics 🧬

DNA sequencing identifies the exact order of nucleotides in a DNA fragment. Traditional Sanger sequencing uses chain terminating dideoxynucleotides that lack a 3' hydroxyl group. When these are incorporated during synthesis, elongation stops. By running many parallel reactions and separating fragments by size, the sequence can be read. Modern automated methods detect fluorescent labels on terminating nucleotides and assemble the sequence by computer.

Next generation sequencing technologies can process millions of fragments simultaneously. They are useful for whole genome sequencing, exome sequencing, and targeted gene panels. In clinical practice, these methods detect mutations in tumor DNA, identify inherited disease causing variants, and analyze microbial communities. USMLE questions often give a brief description of a sequencing output and ask you to interpret a point mutation, insertion, deletion, or frameshift.

Microarrays analyze gene expression patterns or DNA variation across thousands of loci at once. In expression microarrays, cDNA from patient and control samples hybridizes to probes on a chip, and differences in signal intensity reflect changes in mRNA levels. In genomic microarrays, such as single nucleotide polymorphism arrays or comparative genomic hybridization arrays, large scale deletions or duplications are detected based on gain or loss of signal in chromosomal regions.

Genome wide association studies compare many common single nucleotide polymorphisms in large groups of individuals with and without a disease. They identify loci that are statistically associated with disease risk, even when the functional variant is not yet known. These studies rely on high density genotyping arrays and advanced biostatistics, which you will meet elsewhere in the course.

Mutations and DNA Repair 🧯

Mutations are permanent changes in the DNA sequence. They can affect a single base pair or large segments of chromosomes. Point mutations alter a single nucleotide and can be silent, missense, or nonsense, depending on their effect on the encoded amino acid. Silent mutations do not change the amino acid, usually by altering the third base of a codon. Missense mutations change one amino acid and may alter protein function. Nonsense mutations convert a codon for an amino acid into a stop codon, leading to premature termination.

Insertions or deletions of nucleotides that are not in multiples of three cause frameshift mutations. These shift the reading frame, changing every downstream codon and usually producing a truncated and nonfunctional protein. Insertions or deletions that are multiples of three add or remove amino acids without changing the reading frame, which may still be harmful if they disturb critical domains.

At the chromosomal level, mutations include deletions, duplications, inversions, and translocations. These structural changes can disrupt genes, create fusion genes, or alter gene dosage. Some cancers are driven by specific translocations that create oncogenic fusion proteins. The molecular details of those chromosomal disorders are discussed elsewhere, but you should recognize that they arise from double strand breaks and abnormal repair.

Cells are constantly exposed to DNA damaging agents such as ultraviolet light, ionizing radiation, reactive oxygen species, and chemicals. To maintain genomic stability, cells use multiple DNA repair pathways. Base excision repair removes single damaged bases, such as those formed by oxidation or deamination, and replaces them using a template strand. Nucleotide excision repair removes bulky lesions, including pyrimidine dimers caused by ultraviolet light. Mismatch repair corrects errors that escape proofreading after replication.

Defects in specific repair pathways cause characteristic syndromes with high cancer risk or sensitivity to sunlight and radiation. For USMLE, you must connect the type of damage, the impaired repair mechanism, and the clinical consequence. For example, deficiency in nucleotide excision repair predisposes to severe photosensitivity, while defective mismatch repair leads to increased rates of certain internal malignancies.

Double strand breaks are particularly dangerous, since both strands are cut. Cells repair them using either homologous recombination, which uses the intact sister chromatid as a template and is relatively accurate, or nonhomologous end joining, which joins broken ends directly and is error prone. Some immunologic processes such as V(D)J recombination use controlled double strand breaks and repair to generate receptor diversity.

Laboratory Diagnostics Using Molecular Biology 🧬

Modern diagnostics rely heavily on molecular biology methods. Polymerase chain reaction and its variants are used to detect viral load in infections, to identify specific bacterial pathogens, and to check for minimal residual disease in leukemia by finding characteristic fusion transcripts. PCR based tests can rapidly diagnose some genetic disorders by detecting known mutations or deletions.

Fluorescence in situ hybridization uses labeled DNA probes that hybridize to complementary sequences on chromosomes in cells. Under a fluorescence microscope, specific chromosomal regions light up. FISH can detect microdeletions, duplications, translocations, and copy number changes that are too small to see with standard karyotyping. Different colors help distinguish multiple probes in a single assay.

Allele specific oligonucleotide probes and DNA chips can identify single nucleotide variants associated with disease or drug response. Pharmacogenetic testing uses such methods to predict how a patient will metabolize certain drugs or their risk of severe adverse reactions.

Enzyme linked immunosorbent assay is a protein based method that detects antigens or antibodies using antibody antigen interactions and enzyme linked secondary antibodies. Although it is not a DNA technique, it often appears alongside molecular methods in questions about laboratory diagnosis, especially for viral infections and autoimmune diseases.

Molecular methods are also central to prenatal diagnosis and carrier screening. Fetal DNA can be obtained by chorionic villus sampling or amniocentesis and tested by PCR, FISH, or sequencing. Noninvasive prenatal testing analyzes fetal DNA fragments in maternal blood using high throughput sequencing. Positive results are usually confirmed with more direct tests.

Emerging Molecular Therapies 💊

Molecular biology is not only diagnostic but also therapeutic. Gene therapy attempts to introduce functional copies of genes into patient cells, often using viral vectors. These vectors can be integrating, which insert the gene into the host genome, or nonintegrating, which maintain the gene as an extra chromosomal element. Integration risks insertional mutagenesis, so much effort has gone into improving safety.

Genome editing technologies allow direct modification of endogenous DNA sequences. The best known platform uses RNA guided nucleases to create double strand breaks at specific genomic sites, which are then repaired by the cell. By providing a repair template, precise changes can be introduced. Potential applications include correction of monogenic diseases, but off target effects and delivery challenges remain concerns.

Antisense oligonucleotides are short synthetic nucleic acids designed to bind specific mRNA sequences and alter splicing or block translation. Several approved drugs use this strategy to partially restore normal protein production in diseases with splicing defects. Small interfering RNA approaches use similar principles to degrade target mRNAs.

Monoclonal antibodies and small molecule inhibitors often target proteins in signaling pathways that originate from mutations uncovered by molecular analyses. By linking a molecular abnormality to a targeted therapy, precision medicine tailors treatment to the genetic profile of a tumor or individual.

Although the underlying technology can be complex, USMLE questions typically describe these therapies in simplified terms and ask you to recognize the basic principle, the level at which they act, or a key risk such as immune reaction to viral vectors or unintended mutation.

Views: 15

Comments

Please login to add a comment.

Don't have an account? Register now!