Chapter 4.2: Lecture Notes in Gen...

Chapter 4.2: Lecture Notes in Genome Bioinformatics

Lecture Notes in Genome Bioinformatics por Prof. Subhashini Srinivasan
15 sept 2026
22:28

Notas del episodio

For assembled genomes to become biologically useful resources, it is essential to identify and annotate the regions that encode proteins and other functional elements. It is important to note that only about 1.5% of the three-billion-base human genome consists of protein-coding sequences, while the majority comprises regulatory regions, introns, repetitive elements, non-coding RNAs, and other genomic features. In contrast, genomes of prokaryotes are much more compact, with a substantially higher proportion of coding DNA. Protein-coding genes in bacteria and archaea are often densely packed, with relatively short intergenic regions, and in some cases, genes may even overlap.

Therefore, genome annotation, the process of identifying genes, predicting their structures, and assigning potential functions, is a critical step after assembly. While the quality of the assembly determines the accuracy with which genomic regions can be reconstructed, annotation transforms the raw sequence into a functional genome resource that can be used for comparative genomics, evolutionary studies, and understanding the genetic basis of biological traits.

Palabras clave

Hidden Markov Model, Augustus, GeneScan, GenMark

Sobre qué lugar trata este episodio