{"database": "metadata", "table": "run_metadata", "rows": [[42961, "SRR5865381", "SRX3033431", "SRS2382227", "SRP113522", "PRJNA395690", "High resolution annotation of Zebrafish transcriptome using long read sequencing", "GSE101843", "Transcriptome Analysis", "With the emergence of zebrafish as an important model organism  a concerted effort has been made to study its transcriptome. This effort is limited by gaps in zebrafish annotation  which is especially pronounced concerning transcripts dynamically expressed during zygotic genome activation ZGA. To date  short read sequencing has been the principal technology for zebrafish transcriptome annotation. In part because these sequence reads are too short for assembly methods to resolve the full complexity of the transcriptome  the current annotation is rudimentary. By providing direct observation of full length transcripts  recently refined long read sequencing platforms can dramatically improve annotation coverage and accuracy. Here  we leveraged the SMRT platform to study the early ZGA stage zebrafish transcriptome. Our analysis revealed additional novelty and complexity in the zebrafish transcriptome  identifying 2748 high confidence novel transcripts that originated from previously unannotated loci and 1835 new isoforms in previously annotated genes. Overall design: Pooled RNA of a amanitin / untreated embryos were collected and profiled with long read sequencing. Temporally corresponding pre/post ZGA pooled embryonic RNA samples were profiled with short read RNA seq. Long read raw data were assembled into transcripts using IsoSeq \u00a0PMID: 27407110  mapped to the reference GRCz10 genome using GMAP [PMID:15728110] and annotated against the reference transcriptome using Cuffcompare [PMC3334321]. Novel transcripts were compared to short read data and computationally validated in constructing a final long read augmented transcriptome.", null, "pubmed:30061115", null, "long reads preZGA", "GSM2717135", null, "tissue:256 cell|treatment:a amanitin|developmental stage:6 hpf", "long reads preZGA", "Samples were run on the HiSeq2500 System  and base calling and Q scoring was performed with Illumina\u2019s Real Time Analysis RTA version 1.18.64. bcl2fastq 1.8.3 was used to demulitplex and convert files to the fastq format isoforms were quantified using kallisto Genome build: GRCz10", "256 cell", "The procedure for achieving suppression of ZGA is treatment of one to four cell embryos by injection of 0.2nmol of the RNA Polymerase inhibitor \u03b1 amanitin. Pools of 15 embryos each for both untreated wild type embryos and embryos injected with 0.2nmol \u03b1 amanitin to abolish zygotic transcription  as described previously Zamir et al. 1997", "RNA was isolated from zebrafish embryos via standard Trizol protocol as described Kent et al. 2016. Reverse transcription was accomplished using the Superscript III kit from Invitrogen with 2ug total RNA. RT PCR reactions were prepared with 50ng cDNA per reaction and RedTaq reverse transcriptase mix Sigma. Annealing temperature was 58\u00b0 C and extension time was 3 minutes. RT qPCR reactions were prepared with 20ng cDNA per sample and set up according to the Promega GoTaq 2 Step SybrGreen\u2122 kit using a fast 2 step protocol on the Agilent Mx3000P qPCR system and initial melting time of 3 minutes. For short read sequencing  RNA was isolated from sample sets of approximately 20 embryos at 2.5 hpfand 5.25 hpf stages. RNA quality was assessed by Agilent Bioanalyzer. Samples with a RIN score of 9 or greater were used to generate cDNA libraries following treatment with the Ribo Zero rRNA removal kit Illumina", "Zebrafish were maintained according to standard protocols and embryos were obtained during natural spawning of either AB  TAB14  or TAB5 WT adults. The IACUC committee of the Icahn School of Medicine at Mount Sinai approved all protocols.  Fertilized eggs were collected and staged based on morphological criteria to identify embryos at pre 256 cell stage and post ZGA", "treatment:a amanitin|developmental stage:6 hpf", "GSM2717135", "GSM2717135: long reads preZGA; Danio rerio; RNA Seq", "GSM2717135", null, "1", "RNA was isolated from zebrafish embryos via standard Trizol protocol as described Kent et al. 2016. Reverse transcription was accomplished using the Superscript III kit from Invitrogen with 2ug total RNA. RT PCR reactions were prepared with 50ng cDNA per reaction and RedTaq reverse transcriptase mix Sigma. Annealing temperature was 58\u00b0 C and extension time was 3 minutes. RT qPCR reactions were prepared with 20ng cDNA per sample and set up according to the Promega GoTaq 2 Step SybrGreen\u2122 kit using a fast 2 step protocol on the Agilent Mx3000P qPCR system and initial melting time of 3 minutes. For short read sequencing  RNA was isolated from sample sets of approximately 20 embryos at 2.5 hpfand 5.25 hpf stages. RNA quality was assessed by Agilent Bioanalyzer. Samples with a RIN score of 9 or greater were used to generate cDNA libraries following treatment with the Ribo Zero rRNA removal kit Illumina", "GEO Accession:GSM2717135", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "SINGLE", "PACBIO_SMRT", "PacBio RS", null, "SRP113522", null, "loader:fastq load.py", "Zebrafish_a_aman.fastq", "pacbio_native", 90521885.0, 73347.0, "GSM2717135 r1", null, "A:25481448;C:19609613;G:20975956;T:24454868;N:0", null, null, null, null, 25481448, 19609613, 20975956, 24454868, 0, "SRX3033431", "SRS2382227", "SRA591628", "GEO", "Sealfon, Department of Neurology, Icahn School of Medicine at Mount Sinai", 2, 0.61069, 0.61069, 0.0, 0.0, 0.99857, 0.99857, 0.43661, 0.43661, 1668, 1668, "B", "B", "biological fallback assumption", "pacbio", "pacbio_early", "full_length", "rrna_depletion", "ribozero", "bulk", "unknown", "unknown", null, "United States", "2017-07-25", "Gastrula", "Embryo", "Embryo Imprecise", "All anatomical structures"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["42961"], "units": {}, "query_ms": 7.849515000998508}