TY - JOUR T1 - Paired-end analysis of transcription start sites in Arabidopsis reveals plant-specific promoter signatures. JF - Plant Cell Y1 - 2014 A1 - Morton, Taj A1 - Petricka, Jalean A1 - Corcoran, David L A1 - Li, Song A1 - Winter, Cara M A1 - Carda, Alexa A1 - Benfey, Philip N A1 - Ohler, Uwe A1 - Megraw, Molly KW - Arabidopsis KW - Arabidopsis Proteins KW - Binding Sites KW - Cluster Analysis KW - DNA, Plant KW - Gene Expression Regulation, Plant KW - Genome, Plant KW - Models, Genetic KW - Nucleotide Motifs KW - Plant Roots KW - Promoter Regions, Genetic KW - RNA, Messenger KW - RNA, Plant KW - Sequence Analysis, DNA KW - Species Specificity KW - TATA Box KW - Transcription Factors KW - Transcription Initiation Site AB -

Understanding plant gene promoter architecture has long been a challenge due to the lack of relevant large-scale data sets and analysis methods. Here, we present a publicly available, large-scale transcription start site (TSS) data set in plants using a high-resolution method for analysis of 5' ends of mRNA transcripts. Our data set is produced using the paired-end analysis of transcription start sites (PEAT) protocol, providing millions of TSS locations from wild-type Columbia-0 Arabidopsis thaliana whole root samples. Using this data set, we grouped TSS reads into "TSS tag clusters" and categorized clusters into three spatial initiation patterns: narrow peak, broad with peak, and weak peak. We then designed a machine learning model that predicts the presence of TSS tag clusters with outstanding sensitivity and specificity for all three initiation patterns. We used this model to analyze the transcription factor binding site content of promoters exhibiting these initiation patterns. In contrast to the canonical notions of TATA-containing and more broad "TATA-less" promoters, the model shows that, in plants, the vast majority of transcription start sites are TATA free and are defined by a large compendium of known DNA sequence binding elements. We present results on the usage of these elements and provide our Plant PEAT Peaks (3PEAT) model that predicts the presence of TSSs directly from sequence.

[Link to Additional Data and Supplementary Materials]

VL - 26 IS - 7 ER - TY - JOUR T1 - Sustained-input switches for transcription factors and microRNAs are central building blocks of eukaryotic gene circuits. JF - Genome Biol Y1 - 2013 A1 - Megraw, Molly A1 - Mukherjee, Sayan A1 - Ohler, Uwe KW - Algorithms KW - Animals KW - Arabidopsis KW - Computational Biology KW - Drosophila melanogaster KW - Gene Expression Regulation KW - Gene Regulatory Networks KW - Humans KW - MicroRNAs KW - Molecular Sequence Annotation KW - Nucleic Acid Conformation KW - Software KW - Transcription Factors AB -

WaRSwap is a randomization algorithm that for the first time provides a practical network motif discovery method for large multi-layer networks, for example those that include transcription factors, microRNAs, and non-regulatory protein coding genes. The algorithm is applicable to systems with tens of thousands of genes, while accounting for critical aspects of biological networks, including self-loops, large hubs, and target rearrangements. We validate WaRSwap on a newly inferred regulatory network from Arabidopsis thaliana, and compare outcomes on published Drosophila and human networks. Specifically, sustained input switches are among the few over-represented circuits across this diverse set of eukaryotes.

VL - 14 IS - 8 ER - TY - JOUR T1 - A stele-enriched gene regulatory network in the Arabidopsis root. JF - Mol Syst Biol Y1 - 2011 A1 - Brady, Siobhan M A1 - Zhang, Lifang A1 - Megraw, Molly A1 - Martinez, Natalia J A1 - Jiang, Eric A1 - Yi, Charles S A1 - Liu, Weilin A1 - Zeng, Anna A1 - Taylor-Teeples, Mallorie A1 - Kim, Dahae A1 - Ahnert, Sebastian A1 - Ohler, Uwe A1 - Ware, Doreen A1 - Walhout, Albertha J M A1 - Benfey, Philip N KW - Arabidopsis KW - Arabidopsis Proteins KW - Gene Expression Profiling KW - Gene Regulatory Networks KW - MicroRNAs KW - Plant Roots KW - Reproducibility of Results KW - Systems Biology KW - Transcription Factors KW - Two-Hybrid System Techniques AB -

Tightly controlled gene expression is a hallmark of multicellular development and is accomplished by transcription factors (TFs) and microRNAs (miRNAs). Although many studies have focused on identifying downstream targets of these molecules, less is known about the factors that regulate their differential expression. We used data from high spatial resolution gene expression experiments and yeast one-hybrid (Y1H) and two-hybrid (Y2H) assays to delineate a subset of interactions occurring within a gene regulatory network (GRN) that determines tissue-specific TF and miRNA expression in plants. We find that upstream TFs are expressed in more diverse cell types than their targets and that promoters that are bound by a relatively large number of TFs correspond to key developmental regulators. The regulatory consequence of many TFs for their target was experimentally determined using genetic analysis. Remarkably, molecular phenotypes were identified for 65% of the TFs, but morphological phenotypes were associated with only 16%. This indicates that the GRN is robust, and that gene expression changes may be canalized or buffered.

VL - 7 ER - TY - JOUR T1 - MicroRNA promoter analysis. JF - Methods Mol Biol Y1 - 2010 A1 - Megraw, Molly A1 - Hatzigeorgiou, Artemis G KW - MicroRNAs KW - Promoter Regions, Genetic KW - Transcription Factors AB -

In this chapter, we present a brief overview of current knowledge about the promoters of plant microRNAs (miRNAs), and provide a step-by-step guide for predicting plant miRNA promoter elements using known transcription factor binding motifs. The approach to promoter element prediction is based on a carefully constructed collection of Positional Weight Matrices (PWMs) for known transcription factors (TFs) in Arabidopsis. A key concept of the method is to use scoring thresholds for potential binding sites that are appropriate to each individual transcription factor. While the procedure can be applied to search for Transcription Factor Binding Sites (TFBSs) in any pol-II promoter region, it is particularly practical for the case of plant miRNA promoters where upstream sequence regions and binding sites are not readily available in existing databases. The majority of the material described in this chapter is available for download at http://microrna.gr.

[Link to Tools and Supplementary Materials]

VL - 592 ER - TY - JOUR T1 - miRGen 2.0: a database of microRNA genomic information and regulation. JF - Nucleic Acids Res Y1 - 2010 A1 - Alexiou, Panagiotis A1 - Vergoulis, Thanasis A1 - Gleditzsch, Martin A1 - Prekas, George A1 - Dalamagas, Theodore A1 - Megraw, Molly A1 - Grosse, Ivo A1 - Sellis, Timos A1 - Hatzigeorgiou, Artemis G KW - 3' Untranslated Regions KW - Algorithms KW - Animals KW - Cell Line, Tumor KW - Computational Biology KW - Databases, Genetic KW - Databases, Nucleic Acid KW - Humans KW - Information Storage and Retrieval KW - Internet KW - Mice KW - MicroRNAs KW - Polymorphism, Single Nucleotide KW - Software KW - Transcription Factors AB -

MicroRNAs are small, non-protein coding RNA molecules known to regulate the expression of genes by binding to the 3'UTR region of mRNAs. MicroRNAs are produced from longer transcripts which can code for more than one mature miRNAs. miRGen 2.0 is a database that aims to provide comprehensive information about the position of human and mouse microRNA coding transcripts and their regulation by transcription factors, including a unique compilation of both predicted and experimentally supported data. Expression profiles of microRNAs in several tissues and cell lines, single nucleotide polymorphism locations, microRNA target prediction on protein coding genes and mapping of miRNA targets of co-regulated miRNAs on biological pathways are also integrated into the database and user interface. The miRGen database will be continuously maintained and freely available at http://www.microrna.gr/mirgen/.

VL - 38 IS - Database issue ER - TY - JOUR T1 - A transcription factor affinity-based code for mammalian transcription initiation. JF - Genome Res Y1 - 2009 A1 - Megraw, Molly A1 - Pereira, Fernando A1 - Jensen, Shane T A1 - Ohler, Uwe A1 - Hatzigeorgiou, Artemis G KW - Base Composition KW - Databases, Genetic KW - DNA KW - Gene Expression Regulation KW - Genome, Human KW - Humans KW - Promoter Regions, Genetic KW - RNA Polymerase II KW - TATA Box KW - Transcription Factors KW - Transcription Initiation Site KW - Transcription, Genetic AB -

The recent arrival of large-scale cap analysis of gene expression (CAGE) data sets in mammals provides a wealth of quantitative information on coding and noncoding RNA polymerase II transcription start sites (TSS). Genome-wide CAGE studies reveal that a large fraction of TSS exhibit peaks where the vast majority of associated tags map to a particular location ( approximately 45%), whereas other active regions contain a broader distribution of initiation events. The presence of a strong single peak suggests that transcription at these locations may be mediated by position-specific sequence features. We therefore propose a new model for single-peaked TSS based solely on known transcription factors (TFs) and their respective regions of positional enrichment. This probabilistic model leads to near-perfect classification results in cross-validation (auROC = 0.98), and performance in genomic scans demonstrates that TSS prediction with both high accuracy and spatial resolution is achievable for a specific but large subgroup of mammalian promoters. The interpretable model structure suggests a DNA code in which canonical sequence features such as TATA-box, Initiator, and GC content do play a significant role, but many additional TFs show distinct spatial biases with respect to TSS location and are important contributors to the accurate prediction of single-peak transcription initiation sites. The model structure also reveals that CAGE tag clusters distal from annotated gene starts have distinct characteristics compared to those close to gene 5'-ends. Using this high-resolution single-peak model, we predict TSS for approximately 70% of mammalian microRNAs based on currently available data.

[Links to Tools and Supplementary Materials]

VL - 19 IS - 4 ER - TY - JOUR T1 - MicroRNA promoter element discovery in Arabidopsis. JF - RNA Y1 - 2006 A1 - Megraw, Molly A1 - Baev, Vesselin A1 - Rusinov, Ventsislav A1 - Jensen, Shane T A1 - Kalantidis, Kriton A1 - Hatzigeorgiou, Artemis G KW - Arabidopsis KW - Base Sequence KW - Binding Sites KW - Databases, Genetic KW - Feedback, Physiological KW - Genes, Plant KW - MicroRNAs KW - Promoter Regions, Genetic KW - TATA Box KW - Transcription Factors KW - Transcription Initiation Site AB -

In this study we present a method of identifying Arabidopsis miRNA promoter elements using known transcription factor binding motifs. We provide a comparative analysis of the representation of these elements in miRNA promoters, protein-coding gene promoters, and random genomic sequences. We report five transcription factor (TF) binding motifs that show evidence of overrepresentation in miRNA promoter regions relative to the promoter regions of protein-coding genes. This investigation is based on the analysis of 800-nucleotide regions upstream of 63 experimentally verified Transcription Start Sites (TSS) for miRNA primary transcripts in Arabidopsis. While the TATA-box binding motif was also previously reported by Xie and colleagues, the transcription factors AtMYC2, ARF, SORLREP3, and LFY are identified for the first time as overrepresented binding motifs in miRNA promoters.

VL - 12 IS - 9 ER -