{"database": "metadata", "table": "run_metadata", "rows": [[44109, "SRR6250341", "SRX3357265", "SRS2656347", "SRP123526", "PRJNA416939", "Single cell RNAseq SMART seq2 of wild type TLAB and MZoep tz57 zebrafish embryos at 50% epiboly stage", "GSE106466", "Transcriptome Analysis", "SMART seq2 was performed on single cells isolated from visually staged zebrafish embryos. Overall design: Samples were all sequenced in one batch. Some were generated with a five prime' UMI tagged method  and others are full length SMART seq2.", "parent bioproject:PRJNA417291", "pubmed:29700225", null, "wt p3 S11", "GSM2838427", null, "tissue:Embryonic cell|developmental stage:50% epiboly|library preparation batch:p3|genotype:Wild type TLAB", "wt p3 S11", "Bases 1 8 the UMI and template switch were trimmed from the remaining reads  trimmed stretches of polyA or polyT longer than 6bp and removed any reads that were shorter than 15 bp post trimming. We identified and depleted mtRNA and rRNA reads from the remaining reads by using Bowtie2 to align all reads against a transcriptome of mtRNA and rRNA  using the parameters  q   phred33  N 1  I 1  X 2000  k 200   score min L 3.36  0.617   no discordant   no mixed   no unal. Remaining reads were then aligned to a modified version of the Ensembl Zv10 release 81 zebrafish transcriptome using Bowtie2 with parameters  q   phred33  N 1  I 1  X 2000  k 200   score min L 3.36  0.617   no discordant   no mixed. The transcriptome start site of each gene was padded by 100bp to increase alignment rates  since reads were generated from the 5\u2019 end of the transcript  and not all transcripts had correct annotations for their 5\u2019 end. Additionally  pseudogenes were removed from the transcriptome as they sometimes interfered with real genes  given our 5\u2019 coordinate selection described below. We chose the best alignment per read as the read with the most 5\u2019 alignment coordinate to a transcript among those that aligned to the correct strand and had an alignment score of no more than 4 below the best alignment score overall. We then iteratively collapsed pairs of UMIs  if the sum of the minimum Phred quality scores across non matching bases was at most 30 to account for mis matched bases that were likely due to sequencing errors. A collapsed pair was represented for future comparisons by the UMI with a larger prior number of aligned reads unless the other UMI had at least 75% as many reads aligned and had equal or better UMI quality scores for those alignments evaluated as the softmin over aligned reads of the softmax over quality scores for the UMI bases in an aligned read. post collapsing  the number of transcripts observed per gene was calculated. This value was then adjusted based on the probability of a collision the drawing at random of the same UMI multiple times  which was calculated from the number of UMIs not detected for a given gene. The full length fragmented reads without xxx\u2019 enrichment or UMI were aligned using RSEM against a reference transcriptome  which is depleted of mtRNA  rRNA  and pseudogenes as described above. The command \u2018rsem calculate expression\u2019 was used with parameters  p 4   paired end   seed length 24 for performing the alignment. The resulting gene levels in FPKM was used to build the expression matrix. Genome build: GRCz10 dr81 Supplementary files format and content: Tab delimited text files include gene levels UMI counts or FKPM for each sample.", "Embryonic cell", null, "50% epiboly embryos were manually deyolked and mechanically dissociated. Single cells were manually picked under a dissecting microscope and flash frozen on dry ice in 4uL lysis buffer. Libraries were prepared using a SMART seq2 protocol Picelli et al.  2014 with custom TSO oligos and indexing adapters Satija et al.  2015. Reaction volumn for tagmentation steps were reduced to 15% of the volumn indicated in the protocol. 88 MZoep libraries were built with customized indices for 5\u2019 enriched transcripts with unique molecular identifiers UMIs. The remaining libraries were built without xxx  and contained fragments along full length transcripts.", "Fish of each genotype were incrossed and embryos were collected. Embryos were dechorionated and allowed to grow at 28\u00b0C in standard fish water containing methylene blue till 50% epiboly stage.", "developmental stage:50% epiboly|library preparation batch:p3|genotype:Wild type TLAB", "GSM2838427", "GSM2838427: wt p3 S11; Danio rerio; RNA Seq", "GSM2838427", null, "1", "50% epiboly embryos were manually deyolked and mechanically dissociated. Single cells were manually picked under a dissecting microscope and flash frozen on dry ice in 4uL lysis buffer. Libraries were prepared using a SMART seq2 protocol Picelli et al.  2014 with custom TSO oligos and indexing adapters Satija et al.  2015. Reaction volumn for tagmentation steps were reduced to 15% of the volumn indicated in the protocol. 88 MZoep libraries were built with customized indices for five prime enriched transcripts with unique molecular identifiers UMIs. The remaining libraries were built without xxx  and contained fragments along full length transcripts.", null, "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "Illumina HiSeq 2500", null, "SRP123526", null, null, "wt_p3_S11.R1.fastq.gz wt_p3_S11.R2.fastq.gz", "fastq fastq", 2408.0, 43.0, "GSM2838427 r1", "0:24 1:32", "A:568;C:630;G:651;T:559;N:0", 24, 32, null, null, 568, 630, 651, 559, 0, "SRX3357265", "SRS2656347", "SRA627905", "GEO", "Schier Lab, Molecular and Cellular Biology, Harvard University", 2, 0.7742, 0.38462, 0.35483, 0.07692, 0.99977, 0.99991, 0.6923, 1.0, 24, 32, "B", "B", "mate2-mate1 similar by mapping diff", "illumina", "hiseq_era", "full_length", "poly_a", "unknown", "sc", "single_cell_plate", "smartseq", null, "United States", "2017-11-02", "Gastrula", "Embryo", "Embryo Imprecise", "All anatomical structures"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["44109"], "units": {}, "query_ms": 10.384487002738751}