
Note sull'episodio
For most of its history, biology was an observational and experimental science. Scientists studied organisms, cells, tissues, and molecules by looking at them, manipulating them, and recording what happened. The information generated by these experiments was often descriptive and, for the most part, remained closely tied to the laboratory in which it was produced. The transformation of biology into a data-driven science began when biological information could be measured at scale, represented digitally, stored systematically, and analyzed computationally.
Microarrays transformed gene expression into numerical matrices containing measurements for thousands of genes simultaneously. The arrival of high-throughput technologies accelerated this transformation dramatically. The next-generation sequencing revolution made it possible to generate millions and eventually billions of short DNA fragments in a single experiment enabling measurement of genome-scale gene expression, variations, epigenetic changes and proteomics across diverse biological contexts. Structural biology produced increasingly large repositories of experimentally determined molecular structures. Every technological advance added another layer to the digital representation of biology.
Yet generating digital data was only the beginning. The data had to be organized, interpreted, and connected to biological meaning. This was the problem that gave rise to modern bioinformatics.
Early bioinformatics developed algorithms that translated biological questions into computational procedures. Sequence alignment determined how two or more sequences could be compared. Dynamic programming provided systematic solutions to alignment problems. Genome assembly reconstructed long DNA sequences from millions of fragments. Gene-prediction programs such as GENSCAN used probabilistic models to recognize genes within genomic DNA. Multiple sequence alignment transformed collections of related sequences into representations of evolutionary conservation. Public databases allowed researchers to search previously generated information rather than repeat experiments that had already been performed elsewhere.
Traditional bioinformatics generally required humans to tell the computer what to look for. A programmer defined the states, rules, scoring systems, features, or statistical models, and the computer searched for the best solution. Whether assembling a genome, aligning sequences, predicting genes, or identifying conserved residues, the underlying biological assumptions were largely specified in advance.
Modern artificial intelligence changes this relationship.
Instead of explicitly defining every feature that may be biologically important, AI can learn patterns from enormous collections of biological data. A model can encounter millions of protein sequences and learn relationships among amino acids without being explicitly programmed with the rules of protein evolution. It can process vast quantities of genomic sequence and learn sequence patterns associated with genes, regulatory elements, or other biological features. It can learn representations that connect sequence to structure, structure to function, and genetic variation to phenotype.
The distinction is profound. Classical bioinformatics primarily encoded human knowledge into algorithms; modern AI increasingly allows algorithms to extract knowledge from the data themselves.
This chapter follows that transition—from biological information being digitized, to biological data being computationally analyzed, and finally to biological knowledge being learned by machines. It provides the bridge between the algorithmic era of bioinformatics and the emerging era of AI-driven biology.