Analysis of theDrosophilaand human DPR elements reveals a distinct human variant whose specificity can be enhanced by machine learning

Abstract
The RNA polymerase II core promoter is the site of convergence of the signals that lead to the initiation of transcription. Here, we performed a comparative analysis of the downstream core promoter region (DPR) in Drosophila and humans by using machine learning. These studies revealed a distinct human-specific version of the DPR and led to the use of machine learning models for the identification of synthetic extreme DPR motifs with specificity for human transcription factors relative to Drosophila factors and vice versa. More generally, machine learning models could similarly be used to design synthetic DNA elements with customized functional properties.
Funding Information
  • National Institutes of Health (R35 GM118060)