{"database": "metadata", "table": "run_metadata", "rows": [[43988, "SRR6811827", "SRX3768867", "SRS3023384", "SRP121343", "PRJNA415636", "Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars", "GSE106121", "Other", "A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However  a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here  we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes  we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses  LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types  or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes.", null, "pubmed:29644996", null, "Pancreas 3 endo scar", "GSM3032170", null, "source name:Primary pancreatic islet|strain/background:Zebrabow M|tissue:Primary pancreatic islet|developmental stage:Adult", "Pancreas 3 endo scar", "Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode  a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped  had an incorrect barcode  or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors  we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step  we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step  we aimed to remove easily recognizable sequencing errors. To this end  we consecutively considered scar sequences that have the same cellular barcode and UMI  UMIs that have the same cellular barcode and scar sequence  and cellular barcodes that have the same UMI and scar sequence. In each step  we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI  or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus  but information about scar expression levels was not required in our downstream analysis. In the third filtering step  we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight times as many reads. Scar sequences in the same cell that were one Hamming distance apart but had a read ratio less than eight were tested on three criteria if both of them occurred at least twice in the scar library: Do both scars have more than one transcript?  Do both scars occur in cells independently from each other?  Do the UMIs of both scars have Hamming distance of two or more? If two of these criteria were true  the scars were kept and the sequences were placed on a list of validated scars that  if they occurred in the same cell in another library  did not have to be tested anymore. If one or zero criteria were true  the scar that had only one transcript  or the scar that did not occur independently  were filtered out. Genome build: N/A Supplementary files format and content: List of scar transcripts with cell barcode  scar name  cell type  scar probability and number of organisms that have this cell.", "Primary pancreatic islet", null, "Single cell dissociation. 10X Genomics Chromium", null, "strain/background:Zebrabow M|tissue:Primary pancreatic islet|developmental stage:Adult", "GSM3032170", "GSM3032170: Pancreas 3 endo scar; Danio rerio; OTHER", "GSM3032170", null, "1", "Single cell dissociation. 10X Genomics Chromium", "GEO Accession:GSM3032170", "OTHER", "TRANSCRIPTOMIC", "other", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP121343", null, null, "P7endo_scar_R1.fastq.gz P7endo_scar_R2.fastq.gz", "fastq fastq", 2537781396.0, 20465979.0, "GSM3032170 r1", "0:26 1:98", "A:741745261;C:801599713;G:561531527;T:431666561;N:1238334", 26, 98, null, null, 741745261, 801599713, 561531527, 431666561, 1238334, "SRX3768867", "SRS3023384", "SRA623333", "GEO", "Max Delbr\u00fcck Center", 2, 0.00014, 0.00289, 0.00012, 0.00028, 0.99995, 0.99602, 1.0, 0.68478, 26, 98, "T", "T", "mates < 9% mapping rate", "illumina", "nextseq", "unknown", "other", "unknown", "sc", "single_cell_droplet", "10x", null, "Germany", "2018-03-06", "Adult", "Adult", "Pancreas", "Endocrine System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["43988"], "units": {}, "query_ms": 11.132815998280421}