A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays
Abstract
Abstract: Regulatory DNA contains sequence elements that guide transcription factor (TF) binding and chromatin accessibility, that in turn control the expression of nearby genes. Most current descriptions of such regulatory elements are based on classical, statistical enrichment-based motif discovery methods applied to TF binding signals. Here, we present ENCODE GRAMMAR (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations and Rules): one of the largest collections of regulatory DNA deep learning models trained on TF binding and chromatin accessibility to date, and ENCODE MotifCompendium: the first comprehensive lexicon of predictive regulatory motifs derived from the models. For ENCODE GRAMMAR, we trained BPNet and ChromBPNet models across 2,339 TF ChIP-seq datasets in 788 TFs, 1,143 DNase-seq datasets in 287 samples, and 369 ATAC-seq datasets in 204 samples. From the models, we extracted 286,836 sequence motifs that quantitatively predict regulatory signal in the cellular context of each dataset. To consolidate the discovered motifs into a single, non-redundant union across all cell contexts, we developed MotifCompendium—a GPU-accelerated motif management package, that can perform an accelerated calculation of motif-optimized pairwise similarity between 10,000 motifs in just 6 seconds on a 12GB GPU, flag for noisy, undesirable motifs, cluster the motifs, and provide other utilities to facilitate motif analyses. Using MotifCompendium, we consolidated the discovered motifs into a single, unified lexicon of 3,384 motifs that are predicted to drive TF binding and chromatin accessibility across cell contexts (ENCODE MotifCompendium). From ENCODE MotifCompendium, we observed that motifs from chromatin accessibility can be highly orthogonal to those from TF ChIP-seq, especially in different cell contexts (primary cells vs. cell lines). The lexicon also shows substantial overlap with existing TF binding motif databases, recapitulating, on average, ~93% of TF motifs from previous curated databases, while 10% of the lexicon represent entirely novel binding modes not captured by existing databases. We identify previously uncharacterized zinc finger-like motifs, composite motifs with two or more sub-motifs in preferential spacing, and cell-type-specific variants of motifs. This work provides a foundational resource of predictive models, a scalable computational framework for extracting sequence features, and the first unified lexicon of model-derived regulatory motifs.
// Source
Authors: Chang Min Yun, Salil Deshpande, Vivekanandan Ramalingam, Vivian Hecht, Aman Patel, Anusri Pampari, Selin Jessa, Ryan Zhao, Austin T. Wang, Anshul Kundaje
Institutions: Stanford University