{"database": "metadata", "table": "run_metadata", "rows": [[44000, "SRR6211492", "SRX3320767", "SRS2626340", "SRP121343", "PRJNA415636", "Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars", "GSE106121", "Other", "A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However  a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here  we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes  we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses  LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types  or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes.", null, "pubmed:29644996", null, "Heart 2 mRNA", "GSM2830062", null, "source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult", "Heart 2 mRNA", "Alignment and transcript counting of libraries were done using Cell Ranger 2.0.2. Cell numbers to be extracted were set at a minimum of 6000 but were increased if there were substantially more cells with more than 500 unique transcripts. Exact numbers can be found in Supplementary Table 1 in publication. Every sequencing read consists of a cellular barcode  a UMI  and a transcript sequence originating from an mRNA molecule. These transcript sequences were aligned using bwa aln3 with setting ' q 50' to a reference transcriptome constructed from Ensembl release 74 www.ensembl.org with extended three prime UTR regions. We filtered out all unmapped reads and all reads that were not uniquely mapped. post alignment  we determined which cellular barcodes corresponded to cells. We defined a cell to be a cellular barcode with at least five hundred uniquely mapped molecules. For each cellular barcode  we counted the number of molecules mapped to each gene  using the UMI correction method described by Gr\u00fcn et al. This method corrects for the possibility of the same UMI being used for two different transcripts in the same cell with the formula t =  K ln1 \u2013 k o/K  with t the final number of transcripts  k o the observed UMIs  and K the total number of UMIs possible. As protection against barcode sequencing errors  we counted the occurrence of each nucleotide for each barcode and filtered out barcodes in which one nucleotide occurred ten or more times. Furthermore  we filtered out barcodes that were one nucleotide substitution removed from a barcode with at least eight times as many transcripts. Genome build: GRCz10   release 90 Supplementary files format and content: * matrix.mtx: Single cell transcript count table; * barcodes.tsv: List of cell barcodes.", "Heart and blood", null, "Single cell dissociation. 10X Genomics Chromium", null, "strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult", "GSM2830062", "GSM2830062: Heart 2 mRNA; Danio rerio; RNA Seq", "GSM2830062", null, "1", "Single cell dissociation. 10X Genomics Chromium", "GEO Accession:GSM2830062", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP121343", null, null, "H6_wt_R2.fastq.gz H6_wt_R1.fastq.gz", "fastq fastq", 19466673438.0, 141062851.0, "GSM2830062 r1", "0:28 1:110", "A:5788453893;C:4266605110;G:4479895656;T:4930437501;N:1281278", 28, 110, null, null, 5788453893, 4266605110, 4479895656, 4930437501, 1281278, "SRX3320767", "SRS2626340", "SRA623333", "GEO", "Max Delbr\u00fcck Center", 2, 0.00114, 0.94531, 0.00026, 0.06835, 0.99819, 0.88308, 0.68098, 0.70993, 28, 110, "T", "B", "sc-like readlen", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_droplet", "10x", null, "Germany", "2017-10-24", "Adult", "Adult", "Multi-tissue", "Multi-system"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["44000"], "units": {}, "query_ms": 10.568444995442405}