Fine-structured multi-scaling long-range correlations in completely sequenced genomes-features, origin, and classification

Tobias Knoch, M Goker, R Lohner, A (Anis) Abu-Seiris, Frank Grosveld

Research output: Contribution to journalArticleAcademicpeer-review

14 Citations (Scopus)
10 Downloads (Pure)

Abstract

The sequential organization of genomes, i.e. the relations between distant base pairs and regions within sequences, and its connection to the three-dimensional organization of genomes is still a largely unresolved problem. Long-range power-law correlations were found using correlation analysis on almost the entire observable scale of 132 completely sequenced chromosomes of 0.5 x 10(6) to 3.0 x 10(7) bp from Archaea, Bacteria, Arabidopsis thaliana, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster, and Homo sapiens. The local correlation coefficients show a species-specific multi-scaling behaviour: close to random correlations on the scale of a few base pairs, a first maximum from 40 to 3,400 bp (for Arabidopsis thaliana and Drosophila melanogaster divided in two submaxima), and often a region of one or more second maxima from 10(5) to 3 x 10(5) bp. Within this multi-scaling behaviour, an additional fine-structure is present and attributable to codon usage in all except the human sequences, where it is related to nucleosomal binding. Computer-generated random sequences assuming a block organization of genomes, the codon usage, and nucleosomal binding explain these results. Mutation by sequence reshuffling destroyed all correlations. Thus, the stability of correlations seems to be evolutionarily tightly controlled and connected to the spatial genome organization, especially on large scales. In summary, genomes show a complex sequential organization related closely to their three-dimensional organization.
Original languageUndefined/Unknown
Pages (from-to)757-779
Number of pages23
JournalEuropean Biophysics Journal
Volume38
Issue number6
DOIs
Publication statusPublished - 2009

Research programs

  • EMC MGC-02-13-02

Cite this