Chapter 4.3: Lecture Notes in Gen...

Chapter 4.3: Lecture Notes in Genome Bioinformatics

Lecture Notes in Genome Bioinformatics por Prof. Subhashini Srinivasan
15 sept 2026
20:20

Notas del episodio

Because biological function is ultimately what makes a genomic sequence useful, a major objective of genome projects is to identify transcriptional units, determine their exon–intron structures, and assign functions to the proteins and non-coding RNAs encoded by the genome. Genome annotation has therefore evolved from simple ORF finding into a multi-layered process that integrates ab initio prediction, homology, transcript evidence, protein evidence, comparative genomics, and functional databases.

In early genome projects, gene identification relied heavily on de novo gene prediction, often using Generalized Hidden Markov Models (GHMMs). These programs infer genes from intrinsic properties of genomic DNA, including coding potential, codon usage, splice-site signals, and the expected organization of exons and introns. Although predictions from these methods are imperfect, they remain valuable, particularly when little or no experimental information is available for a newly sequenced organism.

Modern annotation pipelines increasingly combine several independent sources of evidence. RNA-seq and full-length transcript sequencing can provide direct evidence for transcription and exon–intron boundaries, while long-read transcript technologies such as PacBio Iso-Seq and Oxford Nanopore cDNA/direct-RNA sequencing can resolve complete transcript structures and alternative isoforms. Protein homology provides another powerful layer of evidence, allowing predicted genes to be compared with experimentally characterized proteins from related organisms. Comparative genomics can further identify conserved coding regions that are difficult to recognize from sequence composition alone.

Annotation has also expanded beyond protein-coding genes. Modern genome projects routinely identify and annotate non-coding RNAs, regulatory elements, repetitive sequences, transposable elements, pseudogenes, and structural features of chromosomes. For complex eukaryotic genomes, annotation may therefore involve separate tracks for genes, transcripts, promoters, enhancers, ncRNAs, repeats, and other functional elements.

Several automated annotation platforms are widely used. RAST (Rapid Annotation using Subsystem Technology) has been particularly influential for bacterial and archaeal genomes, where its subsystem-based approach provides rapid identification and functional assignment of genes. Other widely used approaches include Prokka and Bakta for prokaryotic genomes and BRAKER, MAKER, and Funannotate for eukaryotic genomes. Large genome resources such as Ensembl and NCBI RefSeq combine computational annotation with extensive comparative and experimental evidence.

Despite these advances, manual curation remains essential for high-confidence annotation. Ensembl distinguishes automatic annotation as the genome-wide determination of transcripts from manual curation, in which individual gene models are reviewed and corrected on a case-by-case basis. UniProtKB/Swiss-Prot remains a classic example of high-quality manually curated protein annotation, in contrast to automatically annotated entries in TrEMBL. Curators can resolve errors in exon boundaries, distinguish closely related paralogs, identify alternative transcripts, and assign functions based on experimental evidence that automated pipelines may miss.

The central principle has therefore shifted from “predicting genes” to “integrating evidence.” A high-quality annotation is no longer defined simply by the number of predicted genes, but by how convincingly independent lines of evidence support each gene model and its proposed biological function.

Sobre qué lugar trata este episodio