Video walkthrough

ULTRA-effective labeling of tandem repeats in genomic sequence

Curious 4:16 CC AI

paperi.ai
0:00 / 0:00

Daniel R. Olson, Travis J. Wheeler

Some of the most important parts of a genome are made from repeated DNA—but once those repeats become damaged and irregular, computer programs can mistake them for something else entirely. This paper introduces a way to keep finding them.

Abstract

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions. Availability and implementation: ULTRA is released under an open source license, and is available for download at https://github.com/TravisWheelerLab/ULTRA.

Transcript

Some of the most important parts of a genome are made from repeated DNA—but once those repeats become damaged and irregular, computer programs can mistake them for something else entirely. This paper introduces a way to keep finding them. Long-read sequencing has made accurate repetitive DNA available where short-read sequencing once left only unmappable fragments.

Repetitive DNA can now be examined instead of simply treated as missing. That creates a practical need: repetitive sequence must be marked so researchers can understand newly uncovered DNA and reduce errors in other genome software.

ULTRA is a tool for identifying and labeling these locally repetitive regions. ULTRA is fast enough for an efficient analysis pipeline, covers repetitive regions containing many mutations reliably, and provides statistics and labels that can be interpreted.

The central problem is deciding which DNA is repetitive enough to hide from other analyses. A low cutoff hides too much and misses real repetitive sequence. A high cutoff creates the opposite problem: it hides too little, so worn-down repeats can produce false matches in the annotation.

One possible way around that choice is not to hide repetitive DNA at all, but to include a better understanding of repetition directly in the analysis. ULTRA was designed to improve both the ability to find repetitive DNA and the ability to avoid wrongly labeling ordinary DNA, especially when mutation has made repetition hard to recognize.

It was also designed to provide meaningful scores, consistent and understandable descriptions of repetitive patterns, settings that can be changed for particular genomes, and an easy user experience. The user-ready release demonstrates its labeling and scoring, then compares its performance with two other tools, reporting high genome coverage and a low rate of false labels.

Mutations can also insert or delete letters, shifting where the repeating pattern seems to fall. ULTRA models both the missing or added letter and the temporary shift that follows it. It uses separate internal states for inserted letters, deleted letters, and the offset after an insertion or deletion, keeping track of the repeating pattern after the text slips out of alignment.

This matters because it shows how a repeat can be split precisely where its internal pattern changes, while similar neighboring patterns remain joined. That lets ULTRA preserve both the full repetitive region and its two meaningful subrepeats.

Across both standard and optimized settings, ULTRA covers more repetitive DNA than the two compared tools. For most genomes, it also has the smallest rate of false labels. There is an important exception: one genome with an unusually strong bias toward two DNA letters did not fit ULTRA’s default expectation, and its false-label rate was higher.

Changing the settings did not always help. In many genomes it increased false labels, especially when looking for very long repeating patterns, because the tuning disabled the states that handle inserted and deleted letters. Accurately separating two interwoven repeat patterns remains possible as mutations accumulate, but eventually becomes unreliable.

The turning point depends strongly on the pattern’s length: shorter patterns lose accuracy sooner, while longer ones remain useful deeper into heavily altered sequences. ULTRA was also the only compared tool able to split its work across multiple processing threads.

Its memory use can be reduced in a predictable way by using fewer threads. In the reported whole-genome processing test, another tool failed to process every chromosome without crashing, even after attempts to adjust its settings and give it more memory.

ULTRA finds more repetitive DNA than the compared tools, including regions changed by many mutations, while usually making fewer false labels. That can make newly readable genomes easier to interpret and safer for other software to analyze.

A derivative work by Paperi · AI-generated script, voice and captions · pages and figures unaltered

Made with Paperi.

Drop in a research PDF — get a narrated video walkthrough like this one, with highlights that follow the narration. Free to start.

Try it with your paper →

More in Biochemistry, Genetics and Molecular Biology

Combined SERS-Raman screening of HER2-overexpressing or silenced breast cancer cell lines 4:31

Combined SERS-Raman screening of HER2-overexpressing or silenced breast cancer cell lines

Breast cancers that look alike can behave very differently. This study uses light to find a key cancer marker on individual cells—and then reveals a surprising change that remains even after that marker is silenced.

Alteration of the N6-methyladenosine methylation landscape in a mouse model of polycystic ovary syndrome 4:26

Alteration of the N6-methyladenosine methylation landscape in a mouse model of polycystic ovary syndrome

Polycystic ovary syndrome can affect periods, fertility, and long-term health, yet there is no single treatment for its underlying cause. This study points to a tiny chemical mark on genetic messages as one possible piece of the puzzle.

Exploitation of Quercetin’s Antioxidative Properties in Potential Alternative Therapeutic Options for Neurodegenerative Diseases 3:14

Exploitation of Quercetin’s Antioxidative Properties in Potential Alternative Therapeutic Options for Neurodegenerative Diseases

A compound found in fruits and vegetables may help protect nerve cells from damage linked to aging brain diseases. But the evidence is promising, uneven, and not yet a treatment people can rely on.

All 20 papers in Biochemistry, Genetics and Molecular Biology →