run_metadata: 36365
This data as json
| rowid | run.accession | experiment.accession | sample.accession | study.accession | bioproject | study.title | study.alias | study.type | study.abstract | study.attributes | study.PMIDs | sample.description | sample.title | sample.alias | sample.centername | sample.attributes | GEOsample.title | GEOsample.dataprocessing | GEOsample.source | GEOsample.treatmentprotocol | GEOsample.extractprotocol | GEOsample.growthprotocol | GEOsample.characteristics | GEOsample.accession | experiment.title | experiment.alias | experiment.library_name | experiment.design_description | experiment.library_construction_protocol | experiment.attributes | experiment.library_strategy | experiment.library_source | experiment.library_selection | experiment.library_layout | experiment.platform | experiment.instrument_model | experiment.spot_descriptor | experiment.study_ref | run.title | run.attributes | run.filename | run.semantic_name | run.total_bases | run.total_spots | run.alias | run.read_lengths | run.base_counts | run.r1_length | run.r2_length | run.r3_length | run.r4_length | run.Acount | run.Ccount | run.Gcount | run.Tcount | run.Ncount | run.experiment | run.pool_member | submission.accession | submission.srasource | submission.bioprojectsource | seqdetective.n_mates | seqdetective.mapping_rate.mate1 | seqdetective.mapping_rate.mate2 | seqdetective.nofeature_rate.mate1 | seqdetective.nofeature_rate.mate2 | seqdetective.sparsity.mate1 | seqdetective.sparsity.mate2 | seqdetective.pos_strand_rate.mate1 | seqdetective.pos_strand_rate.mate2 | seqdetective.readlen.mate1 | seqdetective.readlen.mate2 | seqdetective.judgement.mate1 | seqdetective.judgement.mate2 | seqdetective.judgement.reason | platform_family | instrument_generation | read_bias | selection_class | prep_kit | sc_or_bulk | tech_class | technology | tech_variant | submission.bioprojectsource.country | earliest_date | devstage_curation | devstage_curation_coarse | tissue_curation | tissue_curation_coarse |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 36365 | SRR489484 | SRX143561 | SRS310282 | SRP012376 | PRJNA160143 | Extensive alternative polyadenylation during zebrafish development | GSE37453 | Transcriptome Analysis | The post transcriptional fate of messenger RNAs mRNAs is largely dictated by their three prime' untranslated regions three prime'UTRs which are defined by cleavage and polyadenylation CPA of pre mRNAs. We used polyA position profiling by sequencing 3P Seq to map polyA sites at eight developmental stages and tissues in the zebrafish. Analysis of over 60 million 3P Seq reads substantially increased and improved existing three prime'UTR annotations resulting in confidently identified three prime'UTRs for more than 78.79% of the annotated protein coding genes in zebrafish. Most zebrafish genes undergo alternative CPA with more than a thousand genes using different dominant three prime'UTRs at different stages. three prime'UTRs tend to be shortest in the ovaries and longest in the brain. Isoforms with some of the shortest three prime'UTRs are highly expressed in the ovary yet absent in the maternally contributed RNAs of the embryo perhaps because their three prime'UTRs are too short to accommodate a uridine rich motif required for stability of the maternal mRNA. At two hpf thousands of unique polyA sites appear at locations lacking a typical polyadenylation signal which suggests a wave of widespread cytoplasmic polyadenylation of mRNA degradation intermediates. Our insights into the identities formation and evolution of zebrafish three prime'UTRs provide a resource for studying gene regulation during vertebrate development. Overall design: 3P Seq was used to map the three prime' ends of protein coding genes in the zebrafish genome | pubmed:22722342 | 3P Seq PreMZT 3 | GSM919967 | source name:whole embryo at 1.5 hpf 2 hpf type|tissue:whole embryo|development stage:1.5 hpf 2 hpf | 3P Seq PreMZT 3 | For 3P Seq: Reads were reverse complemented and aligned to the D. rerio genome Zv9/danRer7 using Bowtie. Reads that aligned to up to four genomic locus and had one or more mismatches at their three prime end within a terminal adenylate run were carried forward as 3P tags. Reads mapping to the same locus with the same number of terminal adenylates were consolidated in the processed data file. The BED file is as in Jan et al. GSE24924 For 3P PE Seq: In each read pair read #1 captured the three prime end of the polyA tail in the antisense orientation and read #2 captured a portion of the three prime region of the transcript and occasionally the beginning of the polyA tail in the sense orientation. Read #2 began with an adapter of 26 bases. Reads with more than 10 mismatches to the adapter were discarded and the first 26 bases were removed before further processing. A read pair was considered informative only if read #1 began with Ts and read #2 contained 2–39 terminal As. post leading Ts and terminal As were removed from reads #1 and #2 respectively they were mapped to the genome using Bowtie allowing for up to two mismatches and requiring a unique mapping position in the zebrafish genome. The length of the tail encoded in read #2 was defined as the maximum number of trailing As allowing for up to one mismatch. Only cases in which this number was larger than the number of As encoded in the genome at the predicted cleavage position by at least 2 bases were carried forward. Cleavage and polyadenylation position was defined as the last non A base in read #2. The length of the polyA tail at that position was estimated using the corresponding read #1 and defined as the maximal i for which >90% of the bases in the first i bases of read #1 were Ts. This criterion was used to allow for some sequencing errors expected when sequencing long homopolymers. Genome build: danRer7 Supplementary files format and content: The .bed files contain the positions where 3P Seq reads were mapped and the name of each track is the number of reads mapping to that position. Supplementary files format and content: The bigWig .bw files contain a summary of the number of reads whose three prime ends mapped to each position in the genome. Supplementary files format and content: For the paired end sequencing for mapping polyA tail lengths the text files contain the following columns : chromosome position strand number of polyA length reads followed by the individual measurements lower bounds on polyA tail length | whole embryo at 1.5 hpf 2 hpf | Anesthetized decorioneted embryos were washed three times in PBS 137 mM NaCl 2.7 mM KCl 1.5 mM KH2PO4 8 mM Na2HP04 pH 7.4 and suspended in PBS containing 1% freshly added formaldehyde. Embryos were transferred to a dounce homogenizer dounced several times and incubated at room temperature for 15 min. Formaldehyde was quenched by adding 1/20 volume 2.5 M glycine. Cells were pelleted at 400 x g for 5 min. The supernatant was removed and pellets were rinsed twice with PBS flash frozen in liquid nitrogen and stored at 80C. | 3P Seq; see http://web.wi.mit.edu/bartel/pub/protocols.html | Zebrafish embryos or adults grown under standard condition | genotype/variation:wild type|tissue:whole embryo|developmental stage:1.5 hpf 2 hpf | GSM919967 | GSM919967: 3P Seq PreMZT 3; Danio rerio; RNA Seq | GSM919967 1 | GSM919967: 3P Seq PreMZT 3 | 1 | GEO Accession:GSM919967 | RNA-Seq | TRANSCRIPTOMIC | cDNA | SINGLE | ILLUMINA | Illumina HiSeq 2000 | <SPOT_DESCRIPTOR><SPOT_DECODE_SPEC><SPOT_LENGTH>40</SPOT_LENGTH><READ_SPEC><READ_INDEX>0</READ_INDEX><READ_CLASS>Application Read</READ_CLASS><READ_TYPE>Forward</READ_TYPE><BASE_COORD>1</BASE_COORD></READ_SPEC></SPOT_DECODE_SPEC></SPOT_DESCRIPTOR> | SRP012376 | 3P_Seq_PreMZT_3.fastq | fastq | 2568831200.0 | 64220780.0 | GSM919967 r1 | 0:40 | A:928593325;C:457629269;G:462667409;T:719889166;N:52031 | 40 | 928593325 | 457629269 | 462667409 | 719889166 | 52031 | SRX143561 | SRS310282 | SRA051955 | GEO | Whitehead Institute for Biomedical Research | 1 | 0.71138 | 0.10487 | 0.79904 | 0.51337 | 40 | B | usable mapping rate | illumina | hiseq_era | 3prime | cdna_unspecified | unknown | bulk | unknown | unknown | United States | 2012-04-20 | Cleavage | Embryo | Whole Organism | All anatomical structures |