run_metadata
8 rows where experiment.library_strategy = "OTHER", tissue_curation = "Multi-tissue" and tissue_curation_coarse = "Multi-system"
This data as json, CSV (advanced)
| Link | rowid ▼ | run.accession | experiment.accession | sample.accession | study.accession | bioproject | study.title | study.alias | study.type | study.abstract | study.attributes | study.PMIDs | sample.description | sample.title | sample.alias | sample.centername | sample.attributes | GEOsample.title | GEOsample.dataprocessing | GEOsample.source | GEOsample.treatmentprotocol | GEOsample.extractprotocol | GEOsample.growthprotocol | GEOsample.characteristics | GEOsample.accession | experiment.title | experiment.alias | experiment.library_name | experiment.design_description | experiment.library_construction_protocol | experiment.attributes | experiment.library_strategy | experiment.library_source | experiment.library_selection | experiment.library_layout | experiment.platform | experiment.instrument_model | experiment.spot_descriptor | experiment.study_ref | run.title | run.attributes | run.filename | run.semantic_name | run.total_bases | run.total_spots | run.alias | run.read_lengths | run.base_counts | run.r1_length | run.r2_length | run.r3_length | run.r4_length | run.Acount | run.Ccount | run.Gcount | run.Tcount | run.Ncount | run.experiment | run.pool_member | submission.accession | submission.srasource | submission.bioprojectsource | seqdetective.n_mates | seqdetective.mapping_rate.mate1 | seqdetective.mapping_rate.mate2 | seqdetective.nofeature_rate.mate1 | seqdetective.nofeature_rate.mate2 | seqdetective.sparsity.mate1 | seqdetective.sparsity.mate2 | seqdetective.pos_strand_rate.mate1 | seqdetective.pos_strand_rate.mate2 | seqdetective.readlen.mate1 | seqdetective.readlen.mate2 | seqdetective.judgement.mate1 | seqdetective.judgement.mate2 | seqdetective.judgement.reason | platform_family | instrument_generation | read_bias | selection_class | prep_kit | sc_or_bulk | tech_class | technology | tech_variant | submission.bioprojectsource.country | earliest_date | devstage_curation | devstage_curation_coarse | tissue_curation | tissue_curation_coarse |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 43987 | 43987 | SRR6811828 | SRX3768868 | SRS3023414 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 3 exo scar | GSM3032171 | source name:Pancreas except primary islet liver|strain/background:Zebrabow M|tissue:Pancreas except primary islet liver|developmental stage:Adult | Pancreas 3 exo scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas except primary islet liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas except primary islet liver|developmental stage:Adult | GSM3032171 | GSM3032171: Pancreas 3 exo scar; Danio rerio; OTHER | GSM3032171 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM3032171 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P7exo_scar_R1.fastq.gz P7exo_scar_R2.fastq.gz | fastq fastq | 4291410352.0 | 34608148.0 | GSM3032171 r1 | 0:26 1:98 | A:1266869162;C:1345773385;G:950601988;T:726078376;N:2087441 | 26 | 98 | 1266869162 | 1345773385 | 950601988 | 726078376 | 2087441 | SRX3768868 | SRS3023414 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00014 | 0.00241 | 0.00012 | 0.00024 | 0.99995 | 0.99691 | 0.0 | 0.65306 | 26 | 98 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2018-03-06 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 43989 | 43989 | SRR6811826 | SRX3768866 | SRS3023382 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 3 scar | GSM3032169 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 3 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM3032169 | GSM3032169: Heart 3 scar; Danio rerio; OTHER | GSM3032169 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM3032169 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H7_scar_R1.fastq.gz H7_scar_R2.fastq.gz | fastq fastq | 3999397296.0 | 32253204.0 | GSM3032169 r1 | 0:26 1:98 | A:1170970784;C:1255843391;G:881154893;T:689473188;N:1955040 | 26 | 98 | 1170970784 | 1255843391 | 881154893 | 689473188 | 1955040 | SRX3768866 | SRS3023382 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00018 | 0.00228 | 0.00015 | 0.00019 | 0.99991 | 0.99659 | 0.8 | 0.67235 | 26 | 98 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2018-03-06 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 44008 | 44008 | SRR6211484 | SRX3320759 | SRS2626332 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 2 scar | GSM2830055 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 2 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM2830055 | GSM2830055: Heart 2 scar; Danio rerio; OTHER | GSM2830055 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830055 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H6_scar_R1.fastq.gz H6_scar_R2.fastq.gz | fastq fastq | 2079107172.0 | 15065994.0 | GSM2830055 r1 | 0:28 1:110 | A:580640871;C:641740953;G:501505035;T:355092832;N:127481 | 28 | 110 | 580640871 | 641740953 | 501505035 | 355092832 | 127481 | SRX3320759 | SRS2626332 | SRA623333 | GEO | Max Delbrück Center | 2 | 5e-05 | 0.00253 | 4e-05 | 0.00023 | 1.0 | 0.99778 | 0.58419 | 28 | 110 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44009 | 44009 | SRR6211483 | SRX3320758 | SRS2626348 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 2 scar | GSM2830054 | source name:Pancreas and liver|strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | Pancreas 2 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas and liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | GSM2830054 | GSM2830054: Pancreas 2 scar; Danio rerio; OTHER | GSM2830054 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830054 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P6_scar_R1.fastq.gz P6_scar_R2.fastq.gz | fastq fastq | 1026715860.0 | 7439970.0 | GSM2830054 r1 | 0:28 1:110 | A:287816843;C:320879355;G:243598630;T:174359452;N:61580 | 28 | 110 | 287816843 | 320879355 | 243598630 | 174359452 | 61580 | SRX3320758 | SRS2626348 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00029 | 0.00123 | 0.00028 | 6e-05 | 1.0 | 0.99914 | 0.25714 | 28 | 110 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44010 | 44010 | SRR6211482 | SRX3320757 | SRS2626331 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 1 scar | GSM2830053 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 1 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM2830053 | GSM2830053: Heart 1 scar; Danio rerio; OTHER | GSM2830053 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830053 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H5_scar_R1.fastq.gz H5_scar_R2.fastq.gz | fastq fastq | 1287546624.0 | 10218624.0 | GSM2830053 r1 | 0:26 1:100 | A:365272675;C:413037677;G:296495693;T:212570153;N:170426 | 26 | 100 | 365272675 | 413037677 | 296495693 | 212570153 | 170426 | SRX3320757 | SRS2626331 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00023 | 0.00044 | 0.00022 | 3e-05 | 1.0 | 0.99955 | 0.32653 | 26 | 100 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44012 | 44012 | SRR6211480 | SRX3320755 | SRS2626329 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 1 scar | GSM2830051 | source name:Pancreas and liver|strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | Pancreas 1 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas and liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | GSM2830051 | GSM2830051: Pancreas 1 scar; Danio rerio; OTHER | GSM2830051 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830051 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P5_scar_R1.fastq.gz P5_scar_R2.fastq.gz | fastq fastq | 1192602600.0 | 9465100.0 | GSM2830051 r1 | 0:26 1:100 | A:346841760;C:383266222;G:268916933;T:193424756;N:152929 | 26 | 100 | 346841760 | 383266222 | 268916933 | 193424756 | 152929 | SRX3320755 | SRS2626329 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00021 | 0.00091 | 0.00014 | 2e-05 | 0.99997 | 0.99892 | 0.0 | 0.23931 | 26 | 100 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 59501 | 59501 | SRR11924323 | SRX8469995 | SRS6770646 | SRP265951 | PRJNA637293 | The shift from early to late types of ribosomes in zebrafish development involves changes at a subset of rRNA 2' O Me sites | GSE151797 | Other | A sequencing based profiling method RiboMeth seq for ribose methylations was used to study methylation patterns during Zebrafish Danio rerio development Overall design: All samples were analyzed in biological triplicates except for adult tail trunk that was in duplicate. | pubmed:32912962 | Adult tail 2 | GSM4591062 | source name:adult tail trunk|tissue:adult tail trunk|rna fraction:size fractionated 20 40 nt whole cell RNA | Adult tail 2 | Library strategy: RiboMeth seq Barcode separation using python script Adaptor trimming using Cutadapt v. 2.0 Mapping to rRNA reference sequence using Bowtie2 v. 2.3.4.1 Counting read ends and calculating RiboMeth seq scores using python scripts The output FASTA files from small RNA seq were merged and used as the basis of the SNORD search and rRNA interaction prediction. Initially SNORDs were identified by running the merged FASTA file through snoScan Schattner et al. 2005 against zebrafish early and late rRNA reference sequences Locati et al. 2017. Genome build: early and late zebrafish rRNA locati et al. The reference sequences are available in the FASTA file on the series record. Supplementary files format and content: MS Excel file contains five prime and three prime read count and calculated RiboMeth seq score at all positions in the rRNA sequence. | adult tail trunk | Tissues were homogenized and whole cell RNA was extracted using Qiazol Qiagen according to the manufacturer. RiboMeth seq: 5 10 ug of RNA was partially degraded by alkaline at denaturing temperatures. The size fraction 20 40 nt was purified on gels and linkers added using a system relying on a modified Arabidopsis tRNA ligase joining 2' three prime cyclic phosphate and five prime phosphate ends. The library fragments were then sequenced on the Ion Proton platform. See Birkedal U Christensen Dalsgaard M Krogh N Sabarinathan R Gorodkin J Nielsen H. Profiling of ribose methylations in RNA by high throughput sequencing. Angewandte Chemie. 2015;542:451 5 for detailed description | tissue:adult tail trunk|rna fraction:size fractionated 20 40 nt whole cell RNA | GSM4591062 | GSM4591062: Adult tail 2; Danio rerio; OTHER | GSM4591062 | 1 | Tissues were homogenized and whole cell RNA was extracted using Qiazol Qiagen according to the manufacturer. RiboMeth seq: 5 10 ug of RNA was partially degraded by alkaline at denaturing temperatures. The size fraction 20 40 nt was purified on gels and linkers added using a system relying on a modified Arabidopsis tRNA ligase joining 2' three prime cyclic phosphate and five prime phosphate ends. The library fragments were then sequenced on the Ion Proton platform. See Birkedal U Christensen Dalsgaard M Krogh N Sabarinathan R Gorodkin J Nielsen H. Profiling of ribose methylations in RNA by high throughput sequencing. Angewandte Chemie. 2015;542:451 5 for detailed description | GEO Accession:GSM4591062 | OTHER | TRANSCRIPTOMIC | other | SINGLE | ION_TORRENT | Ion Torrent Proton | SRP265951 | intentional duplicate | Adult_tail_2.bam GSE151797_Reference_sequence.fa | bam bam | 156647953.0 | 5517527.0 | GSM4591062 r1 | 0:28.39 | A:32903673;C:49609293;G:39096775;T:35038212;N:0 | 28 | 32903673 | 49609293 | 39096775 | 35038212 | 0 | SRX8469995 | SRS6770646 | SRA1083099 | GEO | RNA Group - Prof. Henrik Nielsen, Department of Cellular and Molecular Medicine, University of Copenhagen | 1 | 0.80662 | 0.20653 | 0.88722 | 0.58074 | 37 | B | usable mapping rate | ion_torrent | ion_torrent | 5prime | small_rna | unknown | bulk | unknown | unknown | Denmark | 2020-06-04 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||||||||
| 59502 | 59502 | SRR11924321 | SRX8469994 | SRS6770645 | SRP265951 | PRJNA637293 | The shift from early to late types of ribosomes in zebrafish development involves changes at a subset of rRNA 2' O Me sites | GSE151797 | Other | A sequencing based profiling method RiboMeth seq for ribose methylations was used to study methylation patterns during Zebrafish Danio rerio development Overall design: All samples were analyzed in biological triplicates except for adult tail trunk that was in duplicate. | pubmed:32912962 | Adult tail 1 | GSM4591061 | source name:adult tail trunk|tissue:adult tail trunk|rna fraction:size fractionated 20 40 nt whole cell RNA | Adult tail 1 | Library strategy: RiboMeth seq Barcode separation using python script Adaptor trimming using Cutadapt v. 2.0 Mapping to rRNA reference sequence using Bowtie2 v. 2.3.4.1 Counting read ends and calculating RiboMeth seq scores using python scripts The output FASTA files from small RNA seq were merged and used as the basis of the SNORD search and rRNA interaction prediction. Initially SNORDs were identified by running the merged FASTA file through snoScan Schattner et al. 2005 against zebrafish early and late rRNA reference sequences Locati et al. 2017. Genome build: early and late zebrafish rRNA locati et al. The reference sequences are available in the FASTA file on the series record. Supplementary files format and content: MS Excel file contains five prime and three prime read count and calculated RiboMeth seq score at all positions in the rRNA sequence. | adult tail trunk | Tissues were homogenized and whole cell RNA was extracted using Qiazol Qiagen according to the manufacturer. RiboMeth seq: 5 10 ug of RNA was partially degraded by alkaline at denaturing temperatures. The size fraction 20 40 nt was purified on gels and linkers added using a system relying on a modified Arabidopsis tRNA ligase joining 2' three prime cyclic phosphate and five prime phosphate ends. The library fragments were then sequenced on the Ion Proton platform. See Birkedal U Christensen Dalsgaard M Krogh N Sabarinathan R Gorodkin J Nielsen H. Profiling of ribose methylations in RNA by high throughput sequencing. Angewandte Chemie. 2015;542:451 5 for detailed description | tissue:adult tail trunk|rna fraction:size fractionated 20 40 nt whole cell RNA | GSM4591061 | GSM4591061: Adult tail 1; Danio rerio; OTHER | GSM4591061 | 1 | Tissues were homogenized and whole cell RNA was extracted using Qiazol Qiagen according to the manufacturer. RiboMeth seq: 5 10 ug of RNA was partially degraded by alkaline at denaturing temperatures. The size fraction 20 40 nt was purified on gels and linkers added using a system relying on a modified Arabidopsis tRNA ligase joining 2' three prime cyclic phosphate and five prime phosphate ends. The library fragments were then sequenced on the Ion Proton platform. See Birkedal U Christensen Dalsgaard M Krogh N Sabarinathan R Gorodkin J Nielsen H. Profiling of ribose methylations in RNA by high throughput sequencing. Angewandte Chemie. 2015;542:451 5 for detailed description | GEO Accession:GSM4591061 | OTHER | TRANSCRIPTOMIC | other | SINGLE | ION_TORRENT | Ion Torrent Proton | SRP265951 | intentional duplicate | Adult_tail_1.bam GSE151797_Reference_sequence.fa | bam bam | 50029636.0 | 1968329.0 | GSM4591061 r1 | 0:25.42 | A:9571393;C:14825016;G:13276309;T:12356918;N:0 | 25 | 9571393 | 14825016 | 13276309 | 12356918 | 0 | SRX8469994 | SRS6770645 | SRA1083099 | GEO | RNA Group - Prof. Henrik Nielsen, Department of Cellular and Molecular Medicine, University of Copenhagen | 1 | 0.46312 | 0.12127 | 0.91806 | 0.67424 | 44 | B | usable mapping rate | ion_torrent | ion_torrent | 5prime | small_rna | unknown | bulk | unknown | unknown | Denmark | 2020-06-04 | Adult | Adult | Multi-tissue | Multi-system |
Advanced export
JSON shape: default, array, newline-delimited
CREATE TABLE run_metadata("run.accession" VARCHAR, "experiment.accession" VARCHAR, "sample.accession" VARCHAR, "study.accession" VARCHAR, bioproject VARCHAR, "study.title" VARCHAR, "study.alias" VARCHAR, "study.type" VARCHAR, "study.abstract" VARCHAR, "study.attributes" VARCHAR, "study.PMIDs" VARCHAR, "sample.description" VARCHAR, "sample.title" VARCHAR, "sample.alias" VARCHAR, "sample.centername" VARCHAR, "sample.attributes" VARCHAR, "GEOsample.title" VARCHAR, "GEOsample.dataprocessing" VARCHAR, "GEOsample.source" VARCHAR, "GEOsample.treatmentprotocol" VARCHAR, "GEOsample.extractprotocol" VARCHAR, "GEOsample.growthprotocol" VARCHAR, "GEOsample.characteristics" VARCHAR, "GEOsample.accession" VARCHAR, "experiment.title" VARCHAR, "experiment.alias" VARCHAR, "experiment.library_name" VARCHAR, "experiment.design_description" VARCHAR, "experiment.library_construction_protocol" VARCHAR, "experiment.attributes" VARCHAR, "experiment.library_strategy" VARCHAR, "experiment.library_source" VARCHAR, "experiment.library_selection" VARCHAR, "experiment.library_layout" VARCHAR, "experiment.platform" VARCHAR, "experiment.instrument_model" VARCHAR, "experiment.spot_descriptor" VARCHAR, "experiment.study_ref" VARCHAR, "run.title" VARCHAR, "run.attributes" VARCHAR, "run.filename" VARCHAR, "run.semantic_name" VARCHAR, "run.total_bases" DOUBLE, "run.total_spots" DOUBLE, "run.alias" VARCHAR, "run.read_lengths" VARCHAR, "run.base_counts" VARCHAR, "run.r1_length" BIGINT, "run.r2_length" BIGINT, "run.r3_length" BIGINT, "run.r4_length" BIGINT, "run.Acount" BIGINT, "run.Ccount" BIGINT, "run.Gcount" BIGINT, "run.Tcount" BIGINT, "run.Ncount" BIGINT, "run.experiment" VARCHAR, "run.pool_member" VARCHAR, "submission.accession" VARCHAR, "submission.srasource" VARCHAR, "submission.bioprojectsource" VARCHAR, "seqdetective.n_mates" BIGINT, "seqdetective.mapping_rate.mate1" DOUBLE, "seqdetective.mapping_rate.mate2" DOUBLE, "seqdetective.nofeature_rate.mate1" DOUBLE, "seqdetective.nofeature_rate.mate2" DOUBLE, "seqdetective.sparsity.mate1" DOUBLE, "seqdetective.sparsity.mate2" DOUBLE, "seqdetective.pos_strand_rate.mate1" DOUBLE, "seqdetective.pos_strand_rate.mate2" DOUBLE, "seqdetective.readlen.mate1" BIGINT, "seqdetective.readlen.mate2" BIGINT, "seqdetective.judgement.mate1" VARCHAR, "seqdetective.judgement.mate2" VARCHAR, "seqdetective.judgement.reason" VARCHAR, platform_family VARCHAR, instrument_generation VARCHAR, read_bias VARCHAR, selection_class VARCHAR, prep_kit VARCHAR, sc_or_bulk VARCHAR, tech_class VARCHAR, technology VARCHAR, tech_variant VARCHAR, "submission.bioprojectsource.country" VARCHAR, earliest_date DATE, devstage_curation VARCHAR, devstage_curation_coarse VARCHAR, tissue_curation VARCHAR, tissue_curation_coarse VARCHAR);;
CREATE INDEX idx_run_bioproject ON run_metadata(bioproject);;
CREATE INDEX idx_run_run_accession ON run_metadata("run.accession");;