{"database": "metadata", "table": "run_metadata", "rows": [[36362, "SRR489487", "SRX143564", "SRS310285", "SRP012376", "PRJNA160143", "Extensive alternative polyadenylation during zebrafish development", "GSE37453", "Transcriptome Analysis", "The post transcriptional fate of messenger RNAs mRNAs is largely dictated by their three prime' untranslated regions three prime'UTRs  which are defined by cleavage and polyadenylation CPA of pre mRNAs. We used polyA position profiling by sequencing 3P Seq to map polyA sites at eight developmental stages and tissues in the zebrafish. Analysis of over 60 million 3P Seq reads substantially increased and improved existing three prime'UTR annotations  resulting in confidently identified three prime'UTRs for more than 78.79% of the annotated protein coding genes in zebrafish. Most zebrafish genes undergo alternative CPA with more than a thousand genes using different dominant three prime'UTRs at different stages. three prime'UTRs tend to be shortest in the ovaries and longest in the brain. Isoforms with some of the shortest three prime'UTRs are highly expressed in the ovary yet absent in the maternally contributed RNAs of the embryo  perhaps because their three prime'UTRs are too short to accommodate a uridine rich motif required for stability of the maternal mRNA. At two hpf  thousands of unique polyA sites appear at locations lacking a typical polyadenylation signal  which suggests a wave of widespread cytoplasmic polyadenylation of mRNA degradation intermediates. Our insights into the identities  formation  and evolution of zebrafish three prime'UTRs provide a resource for studying gene regulation during vertebrate development. Overall design: 3P Seq was used to map the three prime' ends of protein coding genes in the zebrafish genome", null, "pubmed:22722342", null, "3P Seq Ovary", "GSM919970", null, "source name:female adults|genotype/variation:wild type|tissue:ovary|development stage:adult", "3P Seq Ovary", "For 3P Seq: Reads were reverse complemented and aligned to the D. rerio genome Zv9/danRer7 using Bowtie. Reads that aligned to up to four genomic locus and had one or more mismatches at their three prime end within a terminal adenylate run were carried forward as 3P tags. Reads mapping to the same locus with the same number of terminal adenylates were consolidated in the processed data file. The BED file is as in Jan et al. GSE24924 For 3P PE Seq: In each read pair  read #1 captured the three prime end of the polyA tail in the antisense orientation  and read #2 captured a portion of the three prime region of the transcript  and occasionally  the beginning of the polyA tail  in the sense orientation. Read #2 began with an adapter of 26 bases. Reads with more than 10 mismatches to the adapter were discarded and the first 26 bases were removed before further processing.  A read pair was considered informative only if read #1 began with Ts and read #2 contained 2\u201339 terminal As. post leading Ts and terminal As were removed from reads #1 and #2  respectively  they were mapped to the genome using Bowtie  allowing for up to two mismatches and requiring a unique mapping position in the zebrafish genome. The length of the tail encoded in read #2 was defined as the maximum number of trailing As allowing for up to one mismatch. Only cases in which this number was larger than the number of As encoded in the genome at the predicted cleavage position by at least 2 bases were carried forward. Cleavage and polyadenylation position was defined as the last non A base in read #2. The length of the polyA tail at that position was estimated using the corresponding read #1  and defined as the maximal i for which >90% of the bases in the first i bases of read #1 were Ts. This criterion was used to allow for some sequencing errors expected when sequencing long homopolymers. Genome build: danRer7 Supplementary files format and content: The .bed files contain the positions where 3P Seq reads were mapped  and the name of each track is the number of reads mapping to that position. Supplementary files format and content: The bigWig .bw files contain a summary of the number of reads whose three prime ends mapped to each position in the genome. Supplementary files format and content: For the paired end sequencing for mapping polyA tail lengths  the text files contain the following columns : chromosome  position  strand  number of polyA length reads  followed by the individual measurements lower bounds on polyA tail length", "female adults", "Ovaries and testes were obtained as described  in Gupta and Mullins 2010.", "3P Seq; see http://web.wi.mit.edu/bartel/pub/protocols.html", "Zebrafish embryos or adults grown under standard condition", "genotype/variation:wild type|tissue:ovary|developmental stage:adult", "GSM919970", "GSM919970: 3P Seq Ovary; Danio rerio; RNA Seq", "GSM919970 1", "GSM919970: 3P Seq Ovary", "1", null, "GEO Accession:GSM919970", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "SINGLE", "ILLUMINA", "Illumina Genome Analyzer II", "<SPOT_DESCRIPTOR><SPOT_DECODE_SPEC><SPOT_LENGTH>36</SPOT_LENGTH><READ_SPEC><READ_INDEX>0</READ_INDEX><READ_CLASS>Application Read</READ_CLASS><READ_TYPE>Forward</READ_TYPE><BASE_COORD>1</BASE_COORD></READ_SPEC></SPOT_DECODE_SPEC></SPOT_DESCRIPTOR>", "SRP012376", null, null, "3P_Seq_Ovary.fastq", "fastq", 644470380.0, 17901955.0, "GSM919970 r1", "0:36", "A:263871252;C:97757369;G:94972134;T:187683515;N:186110", 36, null, null, null, 263871252, 97757369, 94972134, 187683515, 186110, "SRX143564", "SRS310285", "SRA051955", "GEO", "Whitehead Institute for Biomedical Research", 1, 0.51356, null, 0.04868, null, 0.821, null, 0.48788, null, 36, null, "B", null, "usable mapping rate", "illumina", "early_illumina", "3prime", "cdna_unspecified", "unknown", "bulk", "unknown", "unknown", null, "United States", "2012-04-20", "Adult", "Adult", "Gonad", "Reproductive System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["36362"], "units": {}, "query_ms": 13.730912000028184}