Lecture Notes in Genome Bioinformatics

Lecture Notes in Genome Bioinformatics

por Prof. Subhashini Srinivasan

Chapter 8: Lecture Notes in Genome Bioinformatics

Below, we share the stories of three projects at IBAB, each of which began with a societal gap area and was built, end to end, into an experiment designed to answer it through genomics; a philosophy of advancing training and research in tandem, at a time when the field itself was evolving fast. NGS liberated developing countries from their dependence on the West for defining and pursuing their own research priorities. For the first time, the ability to generate and analyze genomic data was becoming accessible enough that countries could build biological resources around questions of national importance. For me, one such question emerged very early: Could genomics help India address protein malnutrition, control malaria, find hidden cause for rare genetic disorders within families? But How? Where to begin?

Chapter 7: Lecture Notes in Genome Bioinformatics

Bioinformatics is arguably one of the fields most naturally prepared for modern AI because the field has spent decades converting biology into large, structured, machine-readable datasets. In many ways, AI arrived in biology after bioinformatics had already built the infrastructure that AI needed.

Chapter 6: Lecture Notes in Genome Bioinformatics

UNIX, Command-line and pipelines

Chapter 5: Lecture Notes in Genome Bioinformatics

For most of its history, biology was an observational and experimental science. Scientists studied organisms, cells, tissues, and molecules by looking at them, manipulating them, and recording what happened. The information generated by these experiments was often descriptive and, for the most part, remained closely tied to the laboratory in which it was produced. The transformation of biology into a data-driven science began when biological information could be measured at scale, represented digitally, stored systematically, and analyzed computationally. Microarrays transformed gene expression into numerical matrices containing measurements for thousands of genes simultaneously. The arrival of high-throughput technologies accelerated this transformation dramatically. The next-generation sequencing revolution made it possible to generate millions and eventually billions of short DNA fragments in a single experiment enabling measurement of genome-scale gene expression, variations, epigenetic changes and proteomics across diverse biological contexts. Structural biology produced increasingly large repositories of experimentally determined molecular structures. Every technological advance added another layer to the digital representation of biology. Yet generating digital data was only the beginning. The data had to be organized, interpreted, and connected to biological meaning. This was the problem that gave rise to modern bioinformatics. Early bioinformatics developed algorithms that translated biological questions into computational procedures. Sequence alignment determined how two or more sequences could be compared. Dynamic programming provided systematic solutions to alignment problems. Genome assembly reconstructed long DNA sequences from millions of fragments. Gene-prediction programs such as GENSCAN used probabilistic models to recognize genes within genomic DNA. Multiple sequence alignment transformed collections of related sequences into representations of evolutionary conservation. Public databases allowed researchers to search previously generated information rather than repeat experiments that had already been performed elsewhere. Traditional bioinformatics generally required humans to tell the computer what to look for. A programmer defined the states, rules, scoring systems, features, or statistical models, and the computer searched for the best solution. Whether assembling a genome, aligning sequences, predicting genes, or identifying conserved residues, the underlying biological assumptions were largely specified in advance. Modern artificial intelligence changes this relationship. Instead of explicitly defining every feature that may be biologically important, AI can learn patterns from enormous collections of biological data. A model can encounter millions of protein sequences and learn relationships among amino acids without being explicitly programmed with the rules of protein evolution. It can process vast quantities of genomic sequence and learn sequence patterns associated with genes, regulatory elements, or other biological features. It can learn representations that connect sequence to structure, structure to function, and genetic variation to phenotype. The distinction is profound. Classical bioinformatics primarily encoded human knowledge into algorithms; modern AI increasingly allows algorithms to extract knowledge from the data themselves. This chapter follows that transition—from biological information being digitized, to biological data being computationally analyzed, and finally to biological knowledge being learned by machines. It provides the bridge between the algorithmic era of bioinformatics and the emerging era of AI-driven biology.

Chapter 4.10: Lecture Notes in Genome Bioinformatics

The proteome is the complete set of proteins produced by a biological system at a particular time and under a particular condition. Unlike the genome, which is relatively stable, the proteome is highly dynamic. It changes with cell type, developmental stage, environmental conditions, disease, nutrition, and other physiological states. This dynamic nature makes proteins especially valuable for understanding what a cell is doing. Proteins are the primary functional molecules of the cell. They act as enzymes, receptors, transporters, structural components, signaling molecules, and regulators of gene expression. Although the genome provides the blueprint for producing these molecules, the presence of a gene does not necessarily indicate that its corresponding protein is produced, in what quantity, or in what functional state. Proteomics therefore provides a layer of biological information that lies closer to phenotype than the genome or transcriptome. The systematic study of proteins began long before the word proteomics was coined. Individual proteins were purified, characterized, and sequenced using biochemical methods throughout the twentieth century. The development of mass spectrometry, together with advances in protein separation, peptide chemistry, chromatography, and computational analysis, transformed this field. Instead of studying one protein at a time, it became possible to identify and quantify thousands of proteins in a biological sample simultaneously. A typical modern proteomics experiment begins with extraction of proteins from a biological sample such as a cell, tissue, blood, plant, or microbial community. The proteins are usually digested into peptides, commonly using the enzyme trypsin. The resulting peptides are separated by liquid chromatography and introduced into a mass spectrometer. The instrument measures the mass-to-charge ratio (m/z) of peptide ions and, through tandem mass spectrometry, generates fragmentation patterns that can be used to identify the peptides and, consequently, the proteins from which they originated. The computational component is central to modern proteomics. Observed peptide spectra can be compared with theoretical spectra generated from protein databases derived from genome or transcriptome sequences. Matching peptides provide evidence for the presence of proteins. The number or intensity of peptide signals can then be used to estimate relative or absolute protein abundance, depending on the experimental method. An important advantage of proteomics is that it can reveal biological changes that cannot be inferred from DNA sequence alone. Two organisms may have nearly identical genomes but produce very different amounts of proteins under different environmental conditions. Even within the same cell, proteins can undergo post-translational modifications (PTMs) such as phosphorylation, acetylation, glycosylation, and ubiquitination. These modifications can alter protein activity, localization, stability, or interactions without changing the underlying DNA sequence. Proteomics can therefore be viewed as another form of high-dimensional biological measurement. Just as RNA-seq converts gene expression into a gene-by-sample matrix, proteomics can generate a protein-by-sample matrix in which each sample is represented as a vector of protein abundances. These vectors can subsequently be compared using correlation, distance measures, PCA, clustering, machine learning, and other computational approaches.

Chapter 4.9: Lecture Notes in Genome Bioinformatics

One of the ambitious goals of whole-metagenome sequencing (WMGS) is to obtain full genome assemblies of hundreds or thousands of novel unculturable microbes. It is challenging because the sequencing data contain DNA from hundreds or thousands of organisms with vastly different abundances. Highly abundant organisms generate large numbers of reads and can dominate the dataset, while low-abundance organisms may have insufficient coverage for reliable assembly. In addition, closely related organisms can contribute highly similar sequences, making it difficult to assign reads and assembled contigs to the correct organism. These challenges have led to the development of a variety of metagenome assembly strategies, broadly including alignment-based (reference-guided) and composition-based (de novo) approaches. The choice of strategy depends on the availability of reference genomes, the complexity of the microbial community, and the abundance and divergence of the organisms present.

Chapter 4.8-part3: Lecture Notes in Genome Bioinformatics

One of the ambitious goals of whole-metagenome sequencing (WMGS) is to obtain full genome assemblies of hundreds or thousands of novel unculturable microbes. It is challenging because the sequencing data contain DNA from hundreds or thousands of organisms with vastly different abundances. Highly abundant organisms generate large numbers of reads and can dominate the dataset, while low-abundance organisms may have insufficient coverage for reliable assembly. In addition, closely related organisms can contribute highly similar sequences, making it difficult to assign reads and assembled contigs to the correct organism. These challenges have led to the development of a variety of metagenome assembly strategies, broadly including alignment-based (reference-guided) and composition-based (de novo) approaches. The choice of strategy depends on the availability of reference genomes, the complexity of the microbial community, and the abundance and divergence of the organisms present.

Chapter 4.8-part2: Lecture Notes in Genome Bioinformatics

A particularly important transition is therefore occurring in microbial taxonomy. 16S rRNA moved microbial identification from culture-dependent phenotyping to sequence-based classification; whole-genome sequencing is now moving taxonomy from a single-marker system toward genome-wide phylogeny.

Chapter 4.8: Lecture Notes in Genome Bioinformatics

Prokaryotes are ubiquitous and form a fundamental component of the biosphere, contributing substantially to global biomass and playing essential roles in carbon, nitrogen, and other biogeochemical cycles. The microbial world is extraordinarily diverse with only a small fraction of the estimated millions to billions of microbial species have been formally described, and an even smaller fraction has been cultured and experimentally characterized. Microorganisms live in intimate association with plants and animals and can profoundly influence their metabolism, development, immunity, and health. The collective genomes of microorganisms associated with a host are therefore sometimes referred to as its “second genome.” The composition and abundance of microbial communities vary dramatically with their environment. A teaspoon of soil, for example, can contain an enormous diversity of microorganisms, often including thousands of bacterial and archaeal taxa, with highly uneven abundances. The human gut contains a much smaller but still remarkably diverse community, comprising hundreds to more than a thousand bacterial species depending on the individual and the criteria used to define a species. Plant roots are particularly rich microbial habitats because they interact directly with soil and release nutrients that support specialized microbial communities. Marine environments are similarly diverse with even a milliliter of seawater containing hundreds of thousands to millions of microbial cells representing thousands of different taxa. Before the advent of next-generation sequencing (NGS), microbiology depended heavily on isolating microorganisms and growing them in culture. Individual organisms could then be studied for their biochemical properties, or their genomes could be sequenced. This approach, however, was fundamentally limited by cultivability. Many microorganisms cannot readily be grown under standard laboratory conditions. The first complete bacterial genome to be sequenced was that of Haemophilus influenzae, published in 1995. For many years thereafter, genome sequencing remained largely an organism-by-organism exercise and was constrained by the cost and labor required for Sanger sequencing. The enormous microbial diversity present in natural environments far exceeds the diversity that can be cultured in the laboratory. This created a major blind spot in traditional microbiology leading to organisms that could not be isolated could not easily be studied. Genome sequencing of cultured bacteria nevertheless revealed extensive diversity in gene content and provided a foundation for assigning functions to conserved genes. Many genes retain significant homology across bacterial species, particularly at the protein level. However, because of the degeneracy of the genetic code, nucleotide sequences can diverge considerably while encoding similar or even identical proteins. Consequently, DNA-level homology can be insufficient for designing universal PCR primers for a gene family. One solution is to target genomic regions that are sufficiently conserved across diverse microbial groups. Ribosomal RNA (rRNA) genes are particularly valuable because ribosomes are essential for protein synthesis and are present in all cellular organisms. Several regions of rRNA genes are highly conserved, while other regions evolve sufficiently rapidly allowing to distinguish between related organisms. This combination of conserved and variable regions makes rRNA genes excellent molecular markers for microbial identification and community profiling.

Chapter 4.7: Lecture Notes in Genome Bioinformatics

The comparison of genomes within a species and/or across different species has played a pivotal role in modern genomic research. Broadly, comparative genomics can be viewed at two levels. The first is a genome-wide comparison of basic features, such as genome size, chromosome number, GC content, gene density, number of coding sequences (CDS), and repetitive DNA. The second is a high-resolution comparison of the DNA sequences themselves, allowing individual genes, exons, regulatory elements, structural variants, and other genomic features to be examined in detail. High-resolution comparative genomics is particularly powerful because evolutionary conservation provides an important clue to biological function. DNA sequences that remain conserved across related organisms are more likely to be functionally important, whereas rapidly diverging regions may be under weaker functional constraint. Comparison of related genomes can therefore help identify protein-coding genes, exons, regulatory elements, conserved non-coding regions, and other functional elements, even when these elements are difficult to recognize from a single genome alone. Conversely, comparison can also reveal lineage-specific sequences, gene losses, duplications, and other evolutionary changes. A simple and intuitive method of genome comparison is a dot plot, in which two chromosomes or genomic sequences are compared by plotting regions of sequence similarity against their genomic coordinates. A continuous diagonal indicates that the sequences occur in the same linear order, whereas breaks, inversions, or displaced diagonal segments can reveal rearrangements, insertions, deletions, and inversions. For comparison of multiple genomes, tools such as MAUVE can identify locally collinear blocks and reveal large-scale rearrangements, inversions, and other structural differences. Such analyses are particularly useful for examining synteny, the conservation of the relative order of genes or other genomic elements between chromosomes or species. Comparative genomics is largely computational, but interpretation remains essential. A computationally detected similarity does not automatically imply identical biological function, and differences in genome assembly quality, annotation, repetitive DNA, and evolutionary distance can strongly influence the results. Thus, visualization tools such as dot plots and genome browsers, together with sequence alignment and phylogenetic analysis, are often used to interpret the observed similarities and differences. One of the striking observations from comparative genomics is that organisms often retain many of the same genes while their order and chromosomal locations can change substantially during evolution. Orthologous genes that occur together on human chromosome 1, for example, may be distributed across several chromosomes in another species because of chromosome rearrangements, translocations, inversions, and fusion or fission events. Comparative genomics therefore provides a powerful way to reconstruct the evolutionary history of chromosomes while simultaneously identifying conserved genomic elements that are likely to be functionally important.
1 de 3