<?xml version="1.0" encoding="UTF-8"?><xml><records><record><source-app name="Biblio" version="7.x">Drupal-Biblio</source-app><ref-type>17</ref-type><contributors><authors><author><style face="normal" font="default" size="100%">Morton, Taj</style></author><author><style face="normal" font="default" size="100%">Wong, Weng-Keen</style></author><author><style face="normal" font="default" size="100%">Megraw, Molly</style></author></authors></contributors><titles><title><style face="normal" font="default" size="100%">TIPR: transcription initiation pattern recognition on a genome scale.</style></title><secondary-title><style face="normal" font="default" size="100%">Bioinformatics</style></secondary-title><alt-title><style face="normal" font="default" size="100%">Bioinformatics</style></alt-title></titles><keywords><keyword><style  face="normal" font="default" size="100%">Algorithms</style></keyword><keyword><style  face="normal" font="default" size="100%">Genomics</style></keyword><keyword><style  face="normal" font="default" size="100%">Machine Learning</style></keyword><keyword><style  face="normal" font="default" size="100%">Molecular Sequence Annotation</style></keyword><keyword><style  face="normal" font="default" size="100%">Sequence Analysis, DNA</style></keyword><keyword><style  face="normal" font="default" size="100%">Software</style></keyword><keyword><style  face="normal" font="default" size="100%">Transcription Initiation Site</style></keyword><keyword><style  face="normal" font="default" size="100%">Transcription Initiation, Genetic</style></keyword></keywords><dates><year><style  face="normal" font="default" size="100%">2015</style></year><pub-dates><date><style  face="normal" font="default" size="100%">2015 Dec 1</style></date></pub-dates></dates><volume><style face="normal" font="default" size="100%">31</style></volume><pages><style face="normal" font="default" size="100%">3725-32</style></pages><language><style face="normal" font="default" size="100%">eng</style></language><abstract><style face="normal" font="default" size="100%">&lt;p&gt;&lt;b&gt;MOTIVATION: &lt;/b&gt;The computational identification of gene transcription start sites (TSSs) can provide insights into the regulation and function of genes without performing expensive experiments, particularly in organisms with incomplete annotations. High-resolution general-purpose TSS prediction remains a challenging problem, with little recent progress on the identification and differentiation of TSSs which are arranged in different spatial patterns along the chromosome.&lt;/p&gt;&lt;p&gt;&lt;b&gt;RESULTS: &lt;/b&gt;In this work, we present the Transcription Initiation Pattern Recognizer (TIPR), a sequence-based machine learning model that identifies TSSs with high accuracy and resolution for multiple spatial distribution patterns along the genome, including broadly distributed TSS patterns that have previously been difficult to characterize. TIPR predicts not only the locations of TSSs but also the expected spatial initiation pattern each TSS will form along the chromosome-a novel capability for TSS prediction algorithms. As spatial initiation patterns are associated with spatiotemporal expression patterns and gene function, this capability has the potential to improve gene annotations and our understanding of the regulation of transcription initiation. The high nucleotide resolution of this model locates TSSs within 10 nucleotides or less on average.&lt;/p&gt;&lt;p&gt;&lt;b&gt;CONTACT: &lt;/b&gt;megrawm@science.oregonstate.edu.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;http://megraw-dev.cgrb.oregonstate.edu/TIPR&quot;&gt;[Software and Supplementary Materials Link]&lt;/a&gt;&lt;/p&gt;</style></abstract><issue><style face="normal" font="default" size="100%">23</style></issue></record><record><source-app name="Biblio" version="7.x">Drupal-Biblio</source-app><ref-type>17</ref-type><contributors><authors><author><style face="normal" font="default" size="100%">Morton, Taj</style></author><author><style face="normal" font="default" size="100%">Petricka, Jalean</style></author><author><style face="normal" font="default" size="100%">Corcoran, David L</style></author><author><style face="normal" font="default" size="100%">Li, Song</style></author><author><style face="normal" font="default" size="100%">Winter, Cara M</style></author><author><style face="normal" font="default" size="100%">Carda, Alexa</style></author><author><style face="normal" font="default" size="100%">Benfey, Philip N</style></author><author><style face="normal" font="default" size="100%">Ohler, Uwe</style></author><author><style face="normal" font="default" size="100%">Megraw, Molly</style></author></authors></contributors><titles><title><style face="normal" font="default" size="100%">Paired-end analysis of transcription start sites in Arabidopsis reveals plant-specific promoter signatures.</style></title><secondary-title><style face="normal" font="default" size="100%">Plant Cell</style></secondary-title><alt-title><style face="normal" font="default" size="100%">Plant Cell</style></alt-title></titles><keywords><keyword><style  face="normal" font="default" size="100%">Arabidopsis</style></keyword><keyword><style  face="normal" font="default" size="100%">Arabidopsis Proteins</style></keyword><keyword><style  face="normal" font="default" size="100%">Binding Sites</style></keyword><keyword><style  face="normal" font="default" size="100%">Cluster Analysis</style></keyword><keyword><style  face="normal" font="default" size="100%">DNA, Plant</style></keyword><keyword><style  face="normal" font="default" size="100%">Gene Expression Regulation, Plant</style></keyword><keyword><style  face="normal" font="default" size="100%">Genome, Plant</style></keyword><keyword><style  face="normal" font="default" size="100%">Models, Genetic</style></keyword><keyword><style  face="normal" font="default" size="100%">Nucleotide Motifs</style></keyword><keyword><style  face="normal" font="default" size="100%">Plant Roots</style></keyword><keyword><style  face="normal" font="default" size="100%">Promoter Regions, Genetic</style></keyword><keyword><style  face="normal" font="default" size="100%">RNA, Messenger</style></keyword><keyword><style  face="normal" font="default" size="100%">RNA, Plant</style></keyword><keyword><style  face="normal" font="default" size="100%">Sequence Analysis, DNA</style></keyword><keyword><style  face="normal" font="default" size="100%">Species Specificity</style></keyword><keyword><style  face="normal" font="default" size="100%">TATA Box</style></keyword><keyword><style  face="normal" font="default" size="100%">Transcription Factors</style></keyword><keyword><style  face="normal" font="default" size="100%">Transcription Initiation Site</style></keyword></keywords><dates><year><style  face="normal" font="default" size="100%">2014</style></year><pub-dates><date><style  face="normal" font="default" size="100%">2014 Jul</style></date></pub-dates></dates><volume><style face="normal" font="default" size="100%">26</style></volume><pages><style face="normal" font="default" size="100%">2746-60</style></pages><language><style face="normal" font="default" size="100%">eng</style></language><abstract><style face="normal" font="default" size="100%">&lt;p&gt;Understanding plant gene promoter architecture has long been a challenge due to the lack of relevant large-scale data sets and analysis methods. Here, we present a publicly available, large-scale transcription start site (TSS) data set in plants using a high-resolution method for analysis of 5&amp;#39; ends of mRNA transcripts. Our data set is produced using the paired-end analysis of transcription start sites (PEAT) protocol, providing millions of TSS locations from wild-type Columbia-0 Arabidopsis thaliana whole root samples. Using this data set, we grouped TSS reads into &amp;quot;TSS tag clusters&amp;quot; and categorized clusters into three spatial initiation patterns: narrow peak, broad with peak, and weak peak. We then designed a machine learning model that predicts the presence of TSS tag clusters with outstanding sensitivity and specificity for all three initiation patterns. We used this model to analyze the transcription factor binding site content of promoters exhibiting these initiation patterns. In contrast to the canonical notions of TATA-containing and more broad &amp;quot;TATA-less&amp;quot; promoters, the model shows that, in plants, the vast majority of transcription start sites are TATA free and are defined by a large compendium of known DNA sequence binding elements. We present results on the usage of these elements and provide our Plant PEAT Peaks (3PEAT) model that predicts the presence of TSSs directly from sequence.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;http://megraw-dev.cgrb.oregonstate.edu/3PEAT&quot;&gt;[Link to Additional Data and Supplementary Materials]&lt;/a&gt;&lt;/p&gt;</style></abstract><issue><style face="normal" font="default" size="100%">7</style></issue></record></records></xml>