Derivation of oligonucleotide barcodes that are absent from natural sequences
Abstract
Abstract DNA barcodes are short synthetic sequences used to uniquely identify target molecules, samples, or objects, and are essential for a wide range of biotechnology applications including high-throughput sequencing, lineage tracing, genetic screens, massively parallel reporter assays, and DNA-based data storage. However, randomly generated barcodes are prone to off-target hybridization and cross-reactivity with endogenous biological sequences, which can compromise experimental specificity and interpretation. Barcodes derived from sequences that are entirely absent from natural genomes, termed DNA primes, offer an ideal solution by eliminating the possibility of unintended interactions with biological material. Here, we present barcodesDB, a comprehensive database of synthetic DNA barcodes systematically identified by scanning 403,199 complete organismal genome assemblies across the tree of life, together with 215 Gbp of raw metagenomic sequencing reads spanning marine, soil, polar and host-associated environments. These barcodes are absent, on either strand, from every assembly and sequencing read examined in the release described, providing maximal specificity and minimizing cross-reactivity for downstream applications. We provide an open-access web application that enables researchers to search for barcodes satisfying user-defined constraints, including GC content and substring requirements, and to query whether candidate sequences occur in nature. This resource supports robust and scalable barcode design for diverse experimental and applied contexts. Our database is publicly available at: https://barcodesdb.com/.
// Source
Authors: Michail Patsakis, Kimonas Provatas, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
Institutions: The University of Texas at Austin, Austin College