Biological systems can be characterised by their ability to store and transmit information. This is made possible, amongst other things, by the arrangement of nucleotides and amino acids, as well as by their composition within genomes and proteomes. More specifically, these properties influence how genetic information is encoded and passed on to the next generation, how genes are translated into proteins, and how molecular functions arise at various biological levels.
A central aspect of this research area is the question of how GC content shapes the statistical properties of genetic and proteomic sequences. In pseudo-random and natural sequences, the GC content influences the frequency of synonymous codons (i.e. triplets that code for the same amino acid), the occurrence of stop codons, the frequency and prediction of open reading frames, amino acid distributions, and the information potential stored in genetic codes.
Overall, this project examines how sequence structure shapes biological information at various molecular levels: from DNA, through the coding of triplets in the genetic code, to proteins. By applying concepts from statistics, bioinformatics and comparative approaches, it demonstrates how patterns in sequence composition can reveal evolutionary influences, structural relationships and biologically relevant signals. In a broader sense, this research area contributes to a sequence-centred understanding of all aspects of life in which information is shaped by both chemical conditions and evolutionary history.