Chapter 3.2: Lecture Notes in Gen...

Chapter 3.2: Lecture Notes in Genome Bioinformatics

Lecture Notes in Genome Bioinformatics por Prof. Subhashini Srinivasan
15 sept 2026
22:46

Notas del episodio

Frederick Sanger first developed methods for sequencing proteins in the 1950s, establishing the foundation for determining the primary structure of biological macromolecules47. However, following the landmark discovery of the DNA double-helical structure by Watson and Crick in 1953, the focus of molecular biology rapidly shifted toward developing technologies for sequencing DNA.

This transition was driven by several fundamental insights. First, the central dogma of molecular biology established that genetic information flows from DNA to RNA and ultimately to proteins, with messenger RNA acting as the intermediate carrier of information48. Therefore, understanding DNA sequences provided a direct view of the genetic blueprint underlying biological function.

Second, DNA emerged as a more attractive molecule for technological manipulation and sequencing. DNA is chemically more stable than proteins, can be amplified, copied, and manipulated using enzymatic approaches, and contains information that can be interpreted using a universal genetic code49. Importantly, the relationship between DNA and protein sequences is directional and asymmetric because of the degeneracy of the genetic code.

For example, a protein sequence cannot uniquely determine its corresponding DNA sequence because multiple codons can encode the same amino acid. For example, the short peptide sequence ALKRST can be encoded by approximately 6,912 different DNA sequences because several amino acids in this peptide have multiple synonymous codons. In contrast, once a DNA sequence is known, the encoded protein sequence can usually be predicted unambiguously, except for the presence of multiple possible reading frames in an unknown DNA segment.

This asymmetry made DNA sequencing a far more powerful approach for understanding biology. A single DNA sequence provides access not only to the encoded protein but also to regulatory regions, non-coding elements, and evolutionary information embedded within the genome.

Palabras clave

Illumina, Short Redas, Mate-pair, Long reads, PacBio

Sobre qué lugar trata este episodio