Lecture Notes in Genome Bioinformatics

Lecture Notes in Genome Bioinformatics

by Prof. Subhashini Srinivasan

Chapter 4.6: Lecture Notes in Genome Bioinformatics

Epigenomics examines heritable or persistent changes in gene regulation that do not require changes in the underlying DNA sequence. Major epigenetic features include DNA methylation, histone modifications, chromatin accessibility, nucleosome positioning, and three-dimensional chromatin interactions. Because each feature requires a different experimental assay, the computational analysis also differs.

Chapter 4.5: Lecture Notes in Genome Bioinformatics

The early 1990s marked the era of expressed sequence tag (EST) sequencing, an ingenious strategy aimed at identifying human genes without waiting for the complete human genome sequence. Instead of sequencing the entire genome, researchers sequenced short portions of cDNA derived from expressed transcripts. These ESTs provided sequence tags for expressed genes and rapidly expanded the catalogue of human genes. The success of the EST approach laid the foundation for microarray technology, which enabled genome-wide measurement of gene expression across different tissues, developmental stages, disease states, and experimental conditions. Microarrays became one of the dominant technologies for transcriptome profiling for more than a decade. RNA sequencing (RNA-seq) has largely replaced microarrays for transcriptome analysis because sequencing provides a more direct and comprehensive measurement of RNA abundance. Microarrays are fundamentally hybridization-based detection technologies: a transcript can be detected only if a corresponding probe is already represented on the array. Consequently, transcripts that are novel, poorly annotated, highly divergent, or expressed as previously unknown isoforms may be missed. Microarrays also have limitations in their quantitative range. Fluorescence intensity is used as a proxy for transcript abundance, but the relationship between transcript concentration and measured fluorescence is not perfectly linear over the entire dynamic range. At high transcript concentrations, probe spots can become saturated, placing an upper limit on the measurable signal. At the other end of the spectrum, weak signals can be difficult to distinguish from background fluorescence. Because genes can differ by several orders of magnitude in expression level, capturing both very highly and very weakly expressed transcripts accurately in a single hybridization experiment is challenging. RNA-seq addresses many of these limitations by counting sequenced reads derived from RNA molecules rather than measuring hybridization intensity. It does not require a predefined probe for every transcript and can therefore detect novel transcripts, alternative splice isoforms, allele-specific expression, and previously unannotated genes, provided sufficient sequencing depth and appropriate analysis methods are used. Its digital nature also provides a substantially broader dynamic range than microarray fluorescence measurements. Thus, the progression from EST sequencing → microarrays → RNA-seq represents a broader evolution in transcriptomics: from identifying individual expressed sequences, to measuring predefined transcripts simultaneously, and finally to directly sampling and quantifying the transcriptome through high-throughput sequencing.

Chapter 4.4: Lecture Notes in Genome Bioinformatics

Resequencing refers to sequencing the genomes of multiple individuals from a species for the purpose of identifying genetic variation by comparing their sequences with an already assembled reference genome of that species. Unlike de novo genome assembly, resequencing does not require reconstructing the entire genome from scratch. Instead, sequencing reads are aligned to the reference genome, and differences such as SNPs, small insertions and deletions (indels), structural variants, and copy-number variations can be identified. The availability of a high-quality reference genome dramatically reduces the computational and sequencing effort required to study genetic diversity across individuals and populations. Resequencing has therefore become one of the most powerful applications of NGS, particularly for species in which a reference genome is already available.

Chapter 4.3: Lecture Notes in Genome Bioinformatics

Because biological function is ultimately what makes a genomic sequence useful, a major objective of genome projects is to identify transcriptional units, determine their exon–intron structures, and assign functions to the proteins and non-coding RNAs encoded by the genome. Genome annotation has therefore evolved from simple ORF finding into a multi-layered process that integrates ab initio prediction, homology, transcript evidence, protein evidence, comparative genomics, and functional databases. In early genome projects, gene identification relied heavily on de novo gene prediction, often using Generalized Hidden Markov Models (GHMMs). These programs infer genes from intrinsic properties of genomic DNA, including coding potential, codon usage, splice-site signals, and the expected organization of exons and introns. Although predictions from these methods are imperfect, they remain valuable, particularly when little or no experimental information is available for a newly sequenced organism. Modern annotation pipelines increasingly combine several independent sources of evidence. RNA-seq and full-length transcript sequencing can provide direct evidence for transcription and exon–intron boundaries, while long-read transcript technologies such as PacBio Iso-Seq and Oxford Nanopore cDNA/direct-RNA sequencing can resolve complete transcript structures and alternative isoforms. Protein homology provides another powerful layer of evidence, allowing predicted genes to be compared with experimentally characterized proteins from related organisms. Comparative genomics can further identify conserved coding regions that are difficult to recognize from sequence composition alone. Annotation has also expanded beyond protein-coding genes. Modern genome projects routinely identify and annotate non-coding RNAs, regulatory elements, repetitive sequences, transposable elements, pseudogenes, and structural features of chromosomes. For complex eukaryotic genomes, annotation may therefore involve separate tracks for genes, transcripts, promoters, enhancers, ncRNAs, repeats, and other functional elements. Several automated annotation platforms are widely used. RAST (Rapid Annotation using Subsystem Technology) has been particularly influential for bacterial and archaeal genomes, where its subsystem-based approach provides rapid identification and functional assignment of genes. Other widely used approaches include Prokka and Bakta for prokaryotic genomes and BRAKER, MAKER, and Funannotate for eukaryotic genomes. Large genome resources such as Ensembl and NCBI RefSeq combine computational annotation with extensive comparative and experimental evidence. Despite these advances, manual curation remains essential for high-confidence annotation. Ensembl distinguishes automatic annotation as the genome-wide determination of transcripts from manual curation, in which individual gene models are reviewed and corrected on a case-by-case basis. UniProtKB/Swiss-Prot remains a classic example of high-quality manually curated protein annotation, in contrast to automatically annotated entries in TrEMBL. Curators can resolve errors in exon boundaries, distinguish closely related paralogs, identify alternative transcripts, and assign functions based on experimental evidence that automated pipelines may miss. The central principle has therefore shifted from “predicting genes” to “integrating evidence.” A high-quality annotation is no longer defined simply by the number of predicted genes, but by how convincingly independent lines of evidence support each gene model and its proposed biological function.

Chapter 4.2: Lecture Notes in Genome Bioinformatics

For assembled genomes to become biologically useful resources, it is essential to identify and annotate the regions that encode proteins and other functional elements. It is important to note that only about 1.5% of the three-billion-base human genome consists of protein-coding sequences, while the majority comprises regulatory regions, introns, repetitive elements, non-coding RNAs, and other genomic features. In contrast, genomes of prokaryotes are much more compact, with a substantially higher proportion of coding DNA. Protein-coding genes in bacteria and archaea are often densely packed, with relatively short intergenic regions, and in some cases, genes may even overlap. Therefore, genome annotation, the process of identifying genes, predicting their structures, and assigning potential functions, is a critical step after assembly. While the quality of the assembly determines the accuracy with which genomic regions can be reconstructed, annotation transforms the raw sequence into a functional genome resource that can be used for comparative genomics, evolutionary studies, and understanding the genetic basis of biological traits.

Chapter 4.1: Lecture Notes in Genome Bioinformatics

Imagine a machine that could walk along the 2-meter-long DNA molecules packed inside the nucleus of every cell and report the identity of every base it encounters. There would be little need for genome assembly and much of this chapter would become unnecessary. Sequencing machines can accurately read only a limited stretch of DNA before the signal becomes too noisy. If a machine can reliably read 100–200 bases before it begins to “blabber,” the result is a short sequencing read, such as the approximately 150-base reads commonly generated by Illumina platforms. Now imagine deploying a billion Lilliputians, each landing at a random location on the nuclear DNA from many cells and walking as far as it can while accurately reporting the bases it encounters. We would obtain a billion short reads, each representing a small fragment of the genome. If these reads were distributed randomly across the genome, many would overlap with one another. These overlaps provide the clues needed to reconstruct progressively longer stretches of DNA, called contigs. This is the fundamental challenge of genome assembly: reconstructing a long DNA sequence from millions or billions of short, overlapping observations. The quality of an assembly is not simply an all-or-none measure. A perfect telomere-to-telomere (T2T) assembly is the ultimate goal of assembling genomes of any organism, but obtaining such an assembly can require substantially more data and sophisticated technologies. The human genome draft published in 2001 contained thousands of gaps, yet it was enormously valuable and transformed our ability to study genes, genomic variation, and the functional organization of the genome. Thus, a genome assembly does not need to be perfect to be useful. An assembly that produces sufficiently long contigs and scaffolds with meaningful genomic context can already provide a powerful foundation for gene discovery, comparative genomics, variant analysis, and aiding many biological applications.

Chapter 2: Lecture Notes in Genome Bioinformatics

A genome is the complete genetic blueprint of an organism, and genomics is the study of the structure, function, organization, and evolution of entire genomes. In cellular organisms, the genome consists of deoxyribonucleic acid (DNA), whereas some viruses use ribonucleic acid (RNA) as their genetic material. Although the terms genome and are often used interchangeably, they represent distinct concepts. The genome encompasses all the hereditary information required to build, maintain, and reproduce an organism, whereas the genome sequence is simply the linear arrangement of nucleotide bases that encodes this information. Like words and sentences in a language, DNA and RNA sequences consist of ordered strings of discrete nucleotide units. These nucleotide sequences form the fundamental genomic elements that collectively determine the structure, regulation, and function of living organisms. The classical central dogma of molecular biology—DNA → RNA → Protein—was formulated largely from studies in bacteria and provides a partial description of the genetic program of complex eukaryotes. Subsequent large-scale investigations, including the exhaustive analysis of approximately 1% of the human genome by the ENCODE pilot project, revealed an unexpectedly rich landscape of functional genomic elements1. Although only about 1.5% of the human genome is translated to encodes proteins, genome-wide studies using technologies such as EST sequencing, tiling microarrays and RNA sequencing have demonstrated that a substantial fraction of the genome is transcribed. Many of these transcribed sequences function as regulatory RNAs that influence gene expression, chromatin organization, development, and disease. For many years, non-protein-coding regions were dismissed as "junk DNA" or referred to as genomic "dark matter" because their biological functions were poorly understood. It is now evident that many of these regions contain functional elements, including promoters, enhancers, non-coding RNAs, cis-regulatory elements, and structural features that play essential roles in regulating gene expression and cellular function. The ability to identify sequence variation within these genomic elements has transformed biology and medicine. Once variants associated with specific biological traits or diseases are identified, they can be exploited for diagnostics, therapeutics, crop improvement, and vector control. Furthermore, genome-editing technologies such as CRISPR-Cas systems now allow many of these elements to be modified directly, creating unprecedented opportunities for functional studies and precision genetic engineering. This chapter introduces the major classes of genomic and epigenomic elements currently investigated using high-throughput technologies, including genes and transcripts, promoters, enhancers, non-coding RNAs, small interfering RNAs (siRNAs), cis-regulatory elements, single nucleotide polymorphisms (SNPs), structural variants, and epigenetic modifications to find causative genotype under a phenotype of interest.

Chapter 3.4: Lecture Notes in Genome Bioinformatics

Scaffolding the very long contigs generated from PacBio and Oxford Nanopore assemblies requires long-range information that extends far beyond the capabilities of conventional mate-pair libraries. While mate-pair sequencing can provide links across genomic distances of several kilobases, it is not practical for connecting contigs separated by hundreds of kilobases or millions of bases. To overcome this limitation, technologies originally developed to study chromosome organization have been adapted for genome assembly. Together, Hi-C and optical mapping have transformed genome assembly by providing chromosome-scale scaffolding information, bridging the gap between long-read assembly and complete reference-quality genomes. These technologies have been particularly important in generating telomere-to-telomere and near-complete genome assemblies for complex organisms.

Chapter 3.3: Lecture Notes in Genome Bioinformatics

Single-molecule real-time (SMRT) sequencing, pioneered by Pacific Biosciences and complemented by Nanopore sequencing technologies, represents a major advance over short-read sequencing approaches. The key breakthrough was the elimination of the DNA amplification step, allowing individual DNA molecules to be sequenced directly and avoiding biases introduced during PCR amplification. This direct sequencing approach also enabled the generation of much longer reads, typically in the range of 10–20 kb and beyond, providing valuable long-range information for genome assembly, structural variant detection, and resolving repetitive regions. However, early single-molecule sequencing technologies were associated with substantially higher error rates, approaching 10–15% compared with the approximately 99.9% base-level accuracy of Illumina short reads. These errors were largely random and could be eliminated by sequencing the same molecule multiple times and alimenting them by improved consensus algorithms. Recent advances, particularly high-fidelity (HiFi) sequencing from PacBio, have dramatically improved accuracy while retaining long-read capabilities, bridging the gap between the accuracy of short reads and the genomic resolution offered by long reads.

Chapter 3.2: Lecture Notes in Genome Bioinformatics

Frederick Sanger first developed methods for sequencing proteins in the 1950s, establishing the foundation for determining the primary structure of biological macromolecules47. However, following the landmark discovery of the DNA double-helical structure by Watson and Crick in 1953, the focus of molecular biology rapidly shifted toward developing technologies for sequencing DNA. This transition was driven by several fundamental insights. First, the central dogma of molecular biology established that genetic information flows from DNA to RNA and ultimately to proteins, with messenger RNA acting as the intermediate carrier of information48. Therefore, understanding DNA sequences provided a direct view of the genetic blueprint underlying biological function. Second, DNA emerged as a more attractive molecule for technological manipulation and sequencing. DNA is chemically more stable than proteins, can be amplified, copied, and manipulated using enzymatic approaches, and contains information that can be interpreted using a universal genetic code49. Importantly, the relationship between DNA and protein sequences is directional and asymmetric because of the degeneracy of the genetic code. For example, a protein sequence cannot uniquely determine its corresponding DNA sequence because multiple codons can encode the same amino acid. For example, the short peptide sequence ALKRST can be encoded by approximately 6,912 different DNA sequences because several amino acids in this peptide have multiple synonymous codons. In contrast, once a DNA sequence is known, the encoded protein sequence can usually be predicted unambiguously, except for the presence of multiple possible reading frames in an unknown DNA segment. This asymmetry made DNA sequencing a far more powerful approach for understanding biology. A single DNA sequence provides access not only to the encoded protein but also to regulatory regions, non-coding elements, and evolutionary information embedded within the genome.
2 of 3