- AutorIn
- Christopher Jasper Schulmerich Klapproth
- Titel
- On the conservation and identification of non-coding RNAs
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:15-qucosa2-1051390
- Datum der Einreichung
- 07.11.2025
- Datum der Verteidigung
- 02.06.2026
- Abstract (DE)
- RNA is a fundamental molecule in any living cell, serving as the transmitter of information from the genetic code to the ribosome. Beyond this canonical view of RNA as the blueprint for protein synthesis, it has become apparent that non-coding RNA fulfills many crucial biological roles, particularly in terms of gene regulation. Over the last decades it has been suggested that non-coding transcripts might in fact surpass coding RNA in terms of raw quantity and diversity. A very large class of these non-coding RNAs are the so-called long non-coding RNAs (lncRNAs), typically defined as over 200 nucleotides in length, that have been found to be involved with a large number of regulatory networks. Disturbances in lncRNA biological pathways have been shown to be associated with health complications and disease, most notably cancer. Considerable research efforts have been made to identify non-coding RNAs against a background of coding genes, which was shown to be challenging due to the wide range of sizes, shapes and functions they are encountered in. In this work, the properties of lncRNAs are studied primarily from a conservation angle. The secondary structure of non-coding RNA is often considered to be one of the primary drivers of their biological activity. Based on this, a reasonable assumption to make would be that the evolutionary conservation of these structures is a good approximation of their biological importance. A software framework for using retrainable machine learning models for the genome-wide annotation of RNA genes with secondary structure conservation is introduced. The tools are then utilized to study the conservation levels of lncRNAs in plant species, in particular Arabidopsis thaliana. Here it is found that, while displaying higher levels of conservation signals than the genomic background, the secondary structure is not ubiquitous, suggesting regulatory lncRNAs to be very heterogeneous in their mode of function. A particularly interesting example of lncRNA that has gained attention in contemporary research is the telomerase RNA (TR). This particular transcript serves as the template for telomere elongation carried out by the telomerase reverse transcriptase, which is generally considered to be evolutions’s primary answer to the end replication problem of DNA. While much research has been carried out on the telomerase machinery, the TR in particular has been found to be both elusive and of widely different form in many species. One reason for this is that TRs, while fundamentally fulfilling a highly conserved role, are biologically very heterogeneous. In most animals, the TR is a snoRNA of 300 to 500 nucleotides with a H/ACA-box, transcribed by the RNA Polymerase III. However, many other examples exist, such as TRs that are C/D-box snoRNAs in Ciliates. In insects it was recently found that the TR switched to a RNA Polymerase II transcription, a TR biogenesis pathway that is usually observed in plants. However, generally, almost all species seem to follow a rule of exactly one TR gene with one telomere motif per species. In this work, it was discovered that the mining bee genus Andrena expresses multiple copies of the TR gene in almost all of it’s sequenced species. Intriguingly, these extra copies show another diversion from the canonical TR by displaying inner-species variations in the template site with corresponding variations of telomeric tandem repeats. This marks a notable diversion from the previously assumed dogma and is the first instance of this phenomenon described in animals. Interestingly, this seems to be a relatively new invention in insect evolutionary history, as these findings could not be reproduced in related species so far. It does, however, mach similar isolated observations previously made in plants, raising the question of convergent evolution across very long time scales and phylum boundaries. Expanding on the analysis of TRs, a gene that had previously stayed elusive is the TR gene in Caenorhabditis elegans. A particular challenge here is the impossibility of annotating this gene via traditional homology search due to large evolutionary distances and it’s intrinsic heterogeneous nature. Using a reproducible and robust bioinformatics pipeline based on an exhaustive filtering approach, it is shown that the gene can be located using RNAseq, genome and reference annotation data, leading to the first description of the gene in this species that we are aware of. It is furthermore shown that the TR gene is highly conserved in at least 16 other Caenorhabditis species. This serves as a powerful example to show the capabilities of this method in species where a homology search is fundamentally unviable for the stated reasons. In summary, this work analyzes long non-coding RNAs from an evolutionary conservation angle and contributes significant new results to our understanding of the boundaries of telomere biology in animals. In particular it shows telomerase RNAs are substantially more heterogeneous than previously thought, with variant telomeric tandem repeats in the same animal species being biologically viable.
- Freie Schlagwörter (DE)
- Bioinformatics, Genomics, RNA, Machine Learning, Telomerase
- Klassifikation (DDC)
- 500
- Den akademischen Grad verleihende / prüfende Institution
- Universität Leipzig, Leipzig
- Version / Begutachtungsstatus
- publizierte Version / Verlagsversion
- URN Qucosa
- urn:nbn:de:bsz:15-qucosa2-1051390
- Veröffentlichungsdatum Qucosa
- 08.06.2026
- Dokumenttyp
- Dissertation
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis
CC BY 4.0