run_metadata
9 rows where experiment.library_strategy = "OTHER", experiment.platform = "ILLUMINA" and tissue_curation_coarse = "Multi-system"
This data as json, CSV (advanced)
| Link | rowid ▼ | run.accession | experiment.accession | sample.accession | study.accession | bioproject | study.title | study.alias | study.type | study.abstract | study.attributes | study.PMIDs | sample.description | sample.title | sample.alias | sample.centername | sample.attributes | GEOsample.title | GEOsample.dataprocessing | GEOsample.source | GEOsample.treatmentprotocol | GEOsample.extractprotocol | GEOsample.growthprotocol | GEOsample.characteristics | GEOsample.accession | experiment.title | experiment.alias | experiment.library_name | experiment.design_description | experiment.library_construction_protocol | experiment.attributes | experiment.library_strategy | experiment.library_source | experiment.library_selection | experiment.library_layout | experiment.platform | experiment.instrument_model | experiment.spot_descriptor | experiment.study_ref | run.title | run.attributes | run.filename | run.semantic_name | run.total_bases | run.total_spots | run.alias | run.read_lengths | run.base_counts | run.r1_length | run.r2_length | run.r3_length | run.r4_length | run.Acount | run.Ccount | run.Gcount | run.Tcount | run.Ncount | run.experiment | run.pool_member | submission.accession | submission.srasource | submission.bioprojectsource | seqdetective.n_mates | seqdetective.mapping_rate.mate1 | seqdetective.mapping_rate.mate2 | seqdetective.nofeature_rate.mate1 | seqdetective.nofeature_rate.mate2 | seqdetective.sparsity.mate1 | seqdetective.sparsity.mate2 | seqdetective.pos_strand_rate.mate1 | seqdetective.pos_strand_rate.mate2 | seqdetective.readlen.mate1 | seqdetective.readlen.mate2 | seqdetective.judgement.mate1 | seqdetective.judgement.mate2 | seqdetective.judgement.reason | platform_family | instrument_generation | read_bias | selection_class | prep_kit | sc_or_bulk | tech_class | technology | tech_variant | submission.bioprojectsource.country | earliest_date | devstage_curation | devstage_curation_coarse | tissue_curation | tissue_curation_coarse |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 32781 | 32781 | SRR29411985 | SRX24925451 | SRS21630866 | SRP513930 | PRJNA1124008 | Cell state transitions are decoupled from cell division during early embryo development [II] | GSE269848 | Other | Paper abstract: As tissues develop cells divide and differentiate concurrently. Conflicting evidence shows that cell division is either dispensable or required for formation of cell types. To determine the role of cell division in differentiation we arrested the cell cycle in zebrafish embryos using two independent approaches and profiled them at single cell resolution. We show that cell division is dispensable for differentiation of all embryonic tissues during initial cell type differentiation from early gastrulation to the end of segmentation. However in the absence of cell division differentiation slows down in some cell types and cells exhibit global stress responses. While differentiation is robust to blocking cell division the proportions of cells across cell states are not but show evidence of partial compensation. This work clarifies our understanding of the role of cell division in development and showcases the utility of combining embryo wide perturbations with single cell RNA sequencing to uncover the role of common biological processes across multiple tissues. Overall design: Design of experiment for this particular dataset: Forty tails from zebrafish embryos at 24 38 hpf and 48 hpf were collected. Cells were dissociated from these tails and multi seq tags were added to the cells for each time point. Cells from all time points were pooled and single cell sequencing was performed using Indrops. Four gene expression GEX libraries were prepared for transcriptomic information and four multiseq tag libraries TAG were prepared for reading out the multiseq tags. We provide the raw data .fastq files processed data both raw counts.tsv.gz and filtered by total UMI counts filtered.h5ad files and the entire filtered data post removing doublets as all data.h5ad. We also provide counts for multi seq tags example TAG 1.counts.csv. Information about pooling libraries and library indices is provided in Pooling Samples.xlsx. Only 24 hpf data was used for Kukreja et. al. 2024 and hence the cell state informa… | pubmed:37546736 | 24 38 hpf and 48 hpf tails | GSM8328864 | source name:embryo tail|tissue:embryo tail|geo loc name:missing|collection date:missing | 24 38 hpf and 48 hpf tails | Reads were mapped onto the Zebrafish genome as described in Kukreja et al. 2024 manuscript. Multi seq barcodes for each sample were identified using a custom pipeline available here: https://github.com/AllonKleinLab/klunctions/. Barcode abundance was used to manually remove multiplet populations as well as assign cells to their appropriate sample. Data from only 24 hpf tails were used for the manuscript. We provide 8 FASTQ files 2 lanes x 4 reads per lane. The read files from an inDrops run have the following content: * R1 001.fastq.gz : contains the three prime UTR cDNA read * R2 001.fastq.gz : contains the first half of the cell barcode * R3 001.fastq.gz : contains the library index for demultiplexing libraries * R4 001.fastq.gz : contains the second half of the barcode and the UMI. All subsequent processing steps from FASTQ files to count matrixes were performed using the custom pipeline for inDrops data analysis github.com/indrops. Single cell transcriptomes were barcoded using inDrops Klein et al Cell 2015. Standard transcriptome RNA seq libraries were processed as reported in Zilionis et al. Nature Protocol 2016 using inDrops v3 protocol. The transcriptome libraries were sequenced on reads Illumina NextSeq 500. Libraries used standard Illumina sequencing primers and 61 cycles for Read1 14 cycles for Read2 8 cycles each for IndexRead1 and IndexRead2. Raw fastq files was processed using inDrops.py pipeline github.com/indrops/indrops. Sequenced reads were mapped to a zebrafish reference transcriptome built from the zebrafish GRCz10 genome assembly Assembly Accession: GCF 000002035.5 using bowtie version 1.1.143. To obtain the final counts matrix used for data analysis total count filters were applied as described in Kukreja et. al. 2024 method section "Single cell RNA seq Data preprocessing" Assembly: GRCz10 genome assembly Assembly Accession: GCF 000002035.5 Supplementary files format and content: Raw counts for transcriptome tsv.gz: cell barcode gene names and raw counts for all cells Supplementary f… | embryo tail | Forty zebrafish tail samples collected at 24 38 hpf and 48 hpf were dissected between the yolk ball and extension and dissociated according to a modified version of the protocol described by Bresciani et al. 2018. Briefly the tails were dissociated with a mixture of DNaseI 20µg/mL Collagenase/Dispase 8 mg/mL and 0.25% Trypsin EDTA at 30.5°C for 15 minutes. The proteases were quenched with DMEM + 10% FBS. Samples were washed and resuspend in PBS. Cells from each timepoint were hashed with a unique lipid modified oligo using Multi seq. The barcoded samples were subsequently pooled washed with 1% BSA + PBS and resuspended in 0.1% BSA + 18% Optiprep in PBS at a final concentration of 300 000 cells/mL. Single cell transcriptomes were captured using inDrops. NGS libraries were prepared by the Single cell Core at Harvard Medical School and sequenced using an Illumina NovaSeq kit. | tissue:embryo tail | GSM8328864 | GSM8328864: 24 38 hpf and 48 hpf tails; Danio rerio; OTHER | GSM8328864 r1 | GSM8328864 | 1 | Forty zebrafish tail samples collected at 24 38 hpf and 48 hpf were dissected between the yolk ball and extension and dissociated according to a modified version of the protocol described by Bresciani et al. 2018. Briefly the tails were dissociated with a mixture of DNaseI 20µg/mL Collagenase/Dispase 8 mg/mL and 0.25% Trypsin EDTA at 30.5°C for 15 minutes. The proteases were quenched with DMEM + 10% FBS. Samples were washed and resuspend in PBS. Cells from each timepoint were hashed with a unique lipid modified oligo using Multi seq. The barcoded samples were subsequently pooled washed with 1% BSA + PBS and resuspended in 0.1% BSA + 18% Optiprep in PBS at a final concentration of 300 000 cells/mL. Single cell transcriptomes were captured using inDrops. NGS libraries were prepared by the Single cell Core at Harvard Medical School and sequenced using an Illumina NovaSeq kit. | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | Illumina NovaSeq 6000 | SRP513930 | Undetermined_S0_L002_R1_001.fastq.gz Undetermined_S0_L002_R2_001.fastq.gz Undetermined_S0_L002_R3_001.fastq.gz Undetermined_S0_L002_R4_001.fastq.gz | fastq fastq fastq fastq | 58603546956.0 | 505202991.0 | GSM8328864 r1 | 0:86 1:8 2:8 3:14 | A:13191204330;C:9162028793;G:9431065581;T:11662090182;N:1068340 | 86 | 8 | 8 | 14 | 13191204330 | 9162028793 | 9431065581 | 11662090182 | 1068340 | SRX24925451 | SRS21630866 | SRA1899240 | Harvard University | Harvard University | B | usable mapping rate | illumina | novaseq_era | unknown | other | trueseq | sc | single_cell_droplet | indrops | United States | 2024-06-14 | Multi-stage | Embryo | Tail | Multi-system | ||||||||||||||||||||||
| 32782 | 32782 | SRR29411986 | SRX24925451 | SRS21630866 | SRP513930 | PRJNA1124008 | Cell state transitions are decoupled from cell division during early embryo development [II] | GSE269848 | Other | Paper abstract: As tissues develop cells divide and differentiate concurrently. Conflicting evidence shows that cell division is either dispensable or required for formation of cell types. To determine the role of cell division in differentiation we arrested the cell cycle in zebrafish embryos using two independent approaches and profiled them at single cell resolution. We show that cell division is dispensable for differentiation of all embryonic tissues during initial cell type differentiation from early gastrulation to the end of segmentation. However in the absence of cell division differentiation slows down in some cell types and cells exhibit global stress responses. While differentiation is robust to blocking cell division the proportions of cells across cell states are not but show evidence of partial compensation. This work clarifies our understanding of the role of cell division in development and showcases the utility of combining embryo wide perturbations with single cell RNA sequencing to uncover the role of common biological processes across multiple tissues. Overall design: Design of experiment for this particular dataset: Forty tails from zebrafish embryos at 24 38 hpf and 48 hpf were collected. Cells were dissociated from these tails and multi seq tags were added to the cells for each time point. Cells from all time points were pooled and single cell sequencing was performed using Indrops. Four gene expression GEX libraries were prepared for transcriptomic information and four multiseq tag libraries TAG were prepared for reading out the multiseq tags. We provide the raw data .fastq files processed data both raw counts.tsv.gz and filtered by total UMI counts filtered.h5ad files and the entire filtered data post removing doublets as all data.h5ad. We also provide counts for multi seq tags example TAG 1.counts.csv. Information about pooling libraries and library indices is provided in Pooling Samples.xlsx. Only 24 hpf data was used for Kukreja et. al. 2024 and hence the cell state informa… | pubmed:37546736 | 24 38 hpf and 48 hpf tails | GSM8328864 | source name:embryo tail|tissue:embryo tail|geo loc name:missing|collection date:missing | 24 38 hpf and 48 hpf tails | Reads were mapped onto the Zebrafish genome as described in Kukreja et al. 2024 manuscript. Multi seq barcodes for each sample were identified using a custom pipeline available here: https://github.com/AllonKleinLab/klunctions/. Barcode abundance was used to manually remove multiplet populations as well as assign cells to their appropriate sample. Data from only 24 hpf tails were used for the manuscript. We provide 8 FASTQ files 2 lanes x 4 reads per lane. The read files from an inDrops run have the following content: * R1 001.fastq.gz : contains the three prime UTR cDNA read * R2 001.fastq.gz : contains the first half of the cell barcode * R3 001.fastq.gz : contains the library index for demultiplexing libraries * R4 001.fastq.gz : contains the second half of the barcode and the UMI. All subsequent processing steps from FASTQ files to count matrixes were performed using the custom pipeline for inDrops data analysis github.com/indrops. Single cell transcriptomes were barcoded using inDrops Klein et al Cell 2015. Standard transcriptome RNA seq libraries were processed as reported in Zilionis et al. Nature Protocol 2016 using inDrops v3 protocol. The transcriptome libraries were sequenced on reads Illumina NextSeq 500. Libraries used standard Illumina sequencing primers and 61 cycles for Read1 14 cycles for Read2 8 cycles each for IndexRead1 and IndexRead2. Raw fastq files was processed using inDrops.py pipeline github.com/indrops/indrops. Sequenced reads were mapped to a zebrafish reference transcriptome built from the zebrafish GRCz10 genome assembly Assembly Accession: GCF 000002035.5 using bowtie version 1.1.143. To obtain the final counts matrix used for data analysis total count filters were applied as described in Kukreja et. al. 2024 method section "Single cell RNA seq Data preprocessing" Assembly: GRCz10 genome assembly Assembly Accession: GCF 000002035.5 Supplementary files format and content: Raw counts for transcriptome tsv.gz: cell barcode gene names and raw counts for all cells Supplementary f… | embryo tail | Forty zebrafish tail samples collected at 24 38 hpf and 48 hpf were dissected between the yolk ball and extension and dissociated according to a modified version of the protocol described by Bresciani et al. 2018. Briefly the tails were dissociated with a mixture of DNaseI 20µg/mL Collagenase/Dispase 8 mg/mL and 0.25% Trypsin EDTA at 30.5°C for 15 minutes. The proteases were quenched with DMEM + 10% FBS. Samples were washed and resuspend in PBS. Cells from each timepoint were hashed with a unique lipid modified oligo using Multi seq. The barcoded samples were subsequently pooled washed with 1% BSA + PBS and resuspended in 0.1% BSA + 18% Optiprep in PBS at a final concentration of 300 000 cells/mL. Single cell transcriptomes were captured using inDrops. NGS libraries were prepared by the Single cell Core at Harvard Medical School and sequenced using an Illumina NovaSeq kit. | tissue:embryo tail | GSM8328864 | GSM8328864: 24 38 hpf and 48 hpf tails; Danio rerio; OTHER | GSM8328864 r1 | GSM8328864 | 1 | Forty zebrafish tail samples collected at 24 38 hpf and 48 hpf were dissected between the yolk ball and extension and dissociated according to a modified version of the protocol described by Bresciani et al. 2018. Briefly the tails were dissociated with a mixture of DNaseI 20µg/mL Collagenase/Dispase 8 mg/mL and 0.25% Trypsin EDTA at 30.5°C for 15 minutes. The proteases were quenched with DMEM + 10% FBS. Samples were washed and resuspend in PBS. Cells from each timepoint were hashed with a unique lipid modified oligo using Multi seq. The barcoded samples were subsequently pooled washed with 1% BSA + PBS and resuspended in 0.1% BSA + 18% Optiprep in PBS at a final concentration of 300 000 cells/mL. Single cell transcriptomes were captured using inDrops. NGS libraries were prepared by the Single cell Core at Harvard Medical School and sequenced using an Illumina NovaSeq kit. | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | Illumina NovaSeq 6000 | SRP513930 | Undetermined_S0_L001_R1_001.fastq.gz Undetermined_S0_L001_R2_001.fastq.gz Undetermined_S0_L001_R3_001.fastq.gz Undetermined_S0_L001_R4_001.fastq.gz | fastq fastq fastq fastq | 44072422880.0 | 379934680.0 | GSM8328864 r2 | 0:86 1:8 2:8 3:14 | A:9919784228;C:6894965731;G:7094749745;T:8764100624;N:782152 | 86 | 8 | 8 | 14 | 9919784228 | 6894965731 | 7094749745 | 8764100624 | 782152 | SRX24925451 | SRS21630866 | SRA1899240 | Harvard University | Harvard University | B | usable mapping rate | illumina | novaseq_era | unknown | other | trueseq | sc | single_cell_droplet | indrops | United States | 2024-06-14 | Multi-stage | Embryo | Tail | Multi-system | ||||||||||||||||||||||
| 43987 | 43987 | SRR6811828 | SRX3768868 | SRS3023414 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 3 exo scar | GSM3032171 | source name:Pancreas except primary islet liver|strain/background:Zebrabow M|tissue:Pancreas except primary islet liver|developmental stage:Adult | Pancreas 3 exo scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas except primary islet liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas except primary islet liver|developmental stage:Adult | GSM3032171 | GSM3032171: Pancreas 3 exo scar; Danio rerio; OTHER | GSM3032171 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM3032171 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P7exo_scar_R1.fastq.gz P7exo_scar_R2.fastq.gz | fastq fastq | 4291410352.0 | 34608148.0 | GSM3032171 r1 | 0:26 1:98 | A:1266869162;C:1345773385;G:950601988;T:726078376;N:2087441 | 26 | 98 | 1266869162 | 1345773385 | 950601988 | 726078376 | 2087441 | SRX3768868 | SRS3023414 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00014 | 0.00241 | 0.00012 | 0.00024 | 0.99995 | 0.99691 | 0.0 | 0.65306 | 26 | 98 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2018-03-06 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 43989 | 43989 | SRR6811826 | SRX3768866 | SRS3023382 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 3 scar | GSM3032169 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 3 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM3032169 | GSM3032169: Heart 3 scar; Danio rerio; OTHER | GSM3032169 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM3032169 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H7_scar_R1.fastq.gz H7_scar_R2.fastq.gz | fastq fastq | 3999397296.0 | 32253204.0 | GSM3032169 r1 | 0:26 1:98 | A:1170970784;C:1255843391;G:881154893;T:689473188;N:1955040 | 26 | 98 | 1170970784 | 1255843391 | 881154893 | 689473188 | 1955040 | SRX3768866 | SRS3023382 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00018 | 0.00228 | 0.00015 | 0.00019 | 0.99991 | 0.99659 | 0.8 | 0.67235 | 26 | 98 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2018-03-06 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 44008 | 44008 | SRR6211484 | SRX3320759 | SRS2626332 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 2 scar | GSM2830055 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 2 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM2830055 | GSM2830055: Heart 2 scar; Danio rerio; OTHER | GSM2830055 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830055 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H6_scar_R1.fastq.gz H6_scar_R2.fastq.gz | fastq fastq | 2079107172.0 | 15065994.0 | GSM2830055 r1 | 0:28 1:110 | A:580640871;C:641740953;G:501505035;T:355092832;N:127481 | 28 | 110 | 580640871 | 641740953 | 501505035 | 355092832 | 127481 | SRX3320759 | SRS2626332 | SRA623333 | GEO | Max Delbrück Center | 2 | 5e-05 | 0.00253 | 4e-05 | 0.00023 | 1.0 | 0.99778 | 0.58419 | 28 | 110 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44009 | 44009 | SRR6211483 | SRX3320758 | SRS2626348 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 2 scar | GSM2830054 | source name:Pancreas and liver|strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | Pancreas 2 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas and liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | GSM2830054 | GSM2830054: Pancreas 2 scar; Danio rerio; OTHER | GSM2830054 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830054 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P6_scar_R1.fastq.gz P6_scar_R2.fastq.gz | fastq fastq | 1026715860.0 | 7439970.0 | GSM2830054 r1 | 0:28 1:110 | A:287816843;C:320879355;G:243598630;T:174359452;N:61580 | 28 | 110 | 287816843 | 320879355 | 243598630 | 174359452 | 61580 | SRX3320758 | SRS2626348 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00029 | 0.00123 | 0.00028 | 6e-05 | 1.0 | 0.99914 | 0.25714 | 28 | 110 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44010 | 44010 | SRR6211482 | SRX3320757 | SRS2626331 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Heart 1 scar | GSM2830053 | source name:Heart and blood|strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | Heart 1 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Heart and blood | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Heart and blood|developmental stage:Adult | GSM2830053 | GSM2830053: Heart 1 scar; Danio rerio; OTHER | GSM2830053 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830053 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | H5_scar_R1.fastq.gz H5_scar_R2.fastq.gz | fastq fastq | 1287546624.0 | 10218624.0 | GSM2830053 r1 | 0:26 1:100 | A:365272675;C:413037677;G:296495693;T:212570153;N:170426 | 26 | 100 | 365272675 | 413037677 | 296495693 | 212570153 | 170426 | SRX3320757 | SRS2626331 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00023 | 0.00044 | 0.00022 | 3e-05 | 1.0 | 0.99955 | 0.32653 | 26 | 100 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | |||||||||||||
| 44012 | 44012 | SRR6211480 | SRX3320755 | SRS2626329 | SRP121343 | PRJNA415636 | Simultaneous lineage tracing and cell type identification using CRISPR/Cas9 induced genetic scars | GSE106121 | Other | A key goal of developmental biology is to understand how a single cell transforms into a full grown organism consisting of many different cell types. Single cell RNA sequencing scRNA seq has become a widely used method due to its ability to identify all cell types in a tissue or organ in a systematic manner. However a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses LINNAEUS LINeage tracing by Nuclease Activated Editing of Ubiquitous Sequences can be used as a systematic approach for identifying the lineage origin of novel cell types or of known cell types under different conditions. Overall design: Combining scRNA seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes. | pubmed:29644996 | Pancreas 1 scar | GSM2830051 | source name:Pancreas and liver|strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | Pancreas 1 scar | Library strategy: Targeted amplification Scar reads have the same structure as transcript reads: they consist of a barcode a UMI and a scar. The scar sequences were aligned using bwa mem6 to a reference of RFP. We defined a cell as a barcode with at least 500 reads. We removed reads that were unmapped had an incorrect barcode or did not start with the exact PCR primer we used. We truncated all scar sequences to 75 nucleotides and filtered out shorter sequences. To correct for sequencing errors we implemented several rounds of scar filtering Supplementary Fig. 2 in publication. We started by counting the number of times each molecule was sequenced. Sequencing errors will typically have fewer reads than the actual scars they originate from. As a first filtering step we therefore removed all molecules only seen once to reduce the complexity in the dataset for consecutive filtering steps. In the second filtering step we aimed to remove easily recognizable sequencing errors. To this end we consecutively considered scar sequences that have the same cellular barcode and UMI UMIs that have the same cellular barcode and scar sequence and cellular barcodes that have the same UMI and scar sequence. In each step we kept only the molecule with the highest number of reads. The rationale behind this is that it is very improbable to have two valid scar sequences in the same cell with the same UMI or to have a scar sequence with the same UMI appear in two different cells. The observation of two different UMIs for the same scar in the same cell is much more likely and corresponds to detection of multiple transcripts from the same locus but information about scar expression levels was not required in our downstream analysis. In the third filtering step we specifically targeted sequencing errors within each cell. We compared the scar sequences found within a cell to each other. We filtered out sequences that had a Hamming distance of 2 or less to another scar sequence in the same cell that occurred in at least eight tim… | Pancreas and liver | Single cell dissociation. 10X Genomics Chromium | strain/background:Zebrabow M|tissue:Pancreas and liver|developmental stage:Adult | GSM2830051 | GSM2830051: Pancreas 1 scar; Danio rerio; OTHER | GSM2830051 | 1 | Single cell dissociation. 10X Genomics Chromium | GEO Accession:GSM2830051 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | NextSeq 500 | SRP121343 | P5_scar_R1.fastq.gz P5_scar_R2.fastq.gz | fastq fastq | 1192602600.0 | 9465100.0 | GSM2830051 r1 | 0:26 1:100 | A:346841760;C:383266222;G:268916933;T:193424756;N:152929 | 26 | 100 | 346841760 | 383266222 | 268916933 | 193424756 | 152929 | SRX3320755 | SRS2626329 | SRA623333 | GEO | Max Delbrück Center | 2 | 0.00021 | 0.00091 | 0.00014 | 2e-05 | 0.99997 | 0.99892 | 0.0 | 0.23931 | 26 | 100 | T | T | mates < 9% mapping rate | illumina | nextseq | unknown | other | unknown | sc | single_cell_droplet | 10x | Germany | 2017-10-24 | Adult | Adult | Multi-tissue | Multi-system | ||||||||||||
| 50874 | 50874 | SRR8356798 | SRX5167561 | SRS4175446 | SRP173971 | PRJNA510841 | Transcription factor induction of vascular blood stem cell niches in vivo [TOMO Seq] | GSE124150 | Other | We report RNA tomography tomo seq gene expression data for the tail region of a 72 hpf zebrafish embryo. The portion of the tail containing the caudal hematopoietic tissue CHT was isolated by manual dissection snap froze and then cryosections 40 in total 8 µm thick were collected along the dorsal ventral axis. Each cryosection was placed into a tube and the RNA was exctracted and barcoded during a reverse trancription step prior to library synthesis and sequencing. Overall design: The region of the tail containing the CHT was isolated using a scalpel and then snap frozen on dry ice. The RNA from individual cryosections was extracted using Trizol reagent and then barcoded during a reverse transcription step prior to library synthesis according to the previously published method Junker et al. 2014. | parent bioproject:PRJNA510836 | Tomo seq zebrafish tail at 72 hpf | GSM3521717 | source name:mpx:GFP transgenic zebrafish embryo wild type Casper background|tissue:tail|developmental stage:72 hpf | Tomo seq zebrafish tail at 72 hpf | Library strategy: RNA tomography tomo seq Paired end reads were aligned to the transcriptome using bwa version 0.6.2 with default parameters. The zebrafish transcriptome was based on genome release zv9 and contained improved gene annotations as described in the Junker et al. 2014. The right mate of each read pair was mapped to the ensemble of all transcripts and to the set of 92 ERCC spike ins in sense direction. Reads mapping equally to multiple loci were discarded. Mapped reads were assigned to sections based on barcodes according to the CEL seq protocol Hashimshony et al. Cell Reports 2012 Mapping and demultiplexing were done as described in Grün et al. Validation of noise models for single cell transcriptomics Nature Methods 2014. The script needed to process the data is available in https://www.dropbox.com/sh/7s59vvocwtn2ct4/AAAR6pWte8xOzObCYONAFPvIa?dl=0 where script tomo.sh is an example for how to run the mapping scripts. Genome build: zv9 | mpx:GFP transgenic zebrafish embryo wild type Casper background | No treatment | The total RNA from individual cryosections 40 in total each 8 µm thick was extracted using Trizol reagent. The RNA from individual cryosections was barcoded during a reverse transcription step prior to library synthesis according to the method previously described Junker et al. 2014. | mpx:GFP transgenic zebrafish were incrossed and embryos were grown under standard conditions at 28C in E3 buffer until 72 hpf. Embryos were screened for transgene expression and then euthanized by tricaine overdose. The region of the tail containing the CHT was isolated using a scalpel and then snap frozen on dry ice. | tissue:tail|developmental stage:72 hpf | GSM3521717 | GSM3521717: Tomo seq zebrafish tail at 72 hpf; Danio rerio; OTHER | GSM3521717 | 1 | The total RNA from individual cryosections 40 in total each 8 µm thick was extracted using Trizol reagent. The RNA from individual cryosections was barcoded during a reverse transcription step prior to library synthesis according to the method previously described Junker et al. 2014. | GEO Accession:GSM3521717 | OTHER | TRANSCRIPTOMIC | other | PAIRED | ILLUMINA | Illumina HiSeq 2500 | SRP173971 | 3762222874.0 | 24959737.0 | GSM3521717 r1 | 0:75.36 1:75.37 | A:1160852277;C:485683733;G:556965887;T:1558649449;N:71528 | 75 | 75 | 1160852277 | 485683733 | 556965887 | 1558649449 | 71528 | SRX5167561 | SRS4175446 | SRA825234 | GEO | Oncology/Hematology, Boston Children's Hospital | 2 | 0.20644 | 0.76788 | 0.14917 | 0.11551 | 0.98169 | 0.8382 | 0.50072 | 0.52728 | 76 | 76 | T | B | mate1 technical by mapping diff | illumina | hiseq_era | unknown | other | unknown | sc | single_cell_plate | celseq | United States | 2018-12-19 | Larval | Larval | Tail | Multi-system |
Advanced export
JSON shape: default, array, newline-delimited
CREATE TABLE run_metadata("run.accession" VARCHAR, "experiment.accession" VARCHAR, "sample.accession" VARCHAR, "study.accession" VARCHAR, bioproject VARCHAR, "study.title" VARCHAR, "study.alias" VARCHAR, "study.type" VARCHAR, "study.abstract" VARCHAR, "study.attributes" VARCHAR, "study.PMIDs" VARCHAR, "sample.description" VARCHAR, "sample.title" VARCHAR, "sample.alias" VARCHAR, "sample.centername" VARCHAR, "sample.attributes" VARCHAR, "GEOsample.title" VARCHAR, "GEOsample.dataprocessing" VARCHAR, "GEOsample.source" VARCHAR, "GEOsample.treatmentprotocol" VARCHAR, "GEOsample.extractprotocol" VARCHAR, "GEOsample.growthprotocol" VARCHAR, "GEOsample.characteristics" VARCHAR, "GEOsample.accession" VARCHAR, "experiment.title" VARCHAR, "experiment.alias" VARCHAR, "experiment.library_name" VARCHAR, "experiment.design_description" VARCHAR, "experiment.library_construction_protocol" VARCHAR, "experiment.attributes" VARCHAR, "experiment.library_strategy" VARCHAR, "experiment.library_source" VARCHAR, "experiment.library_selection" VARCHAR, "experiment.library_layout" VARCHAR, "experiment.platform" VARCHAR, "experiment.instrument_model" VARCHAR, "experiment.spot_descriptor" VARCHAR, "experiment.study_ref" VARCHAR, "run.title" VARCHAR, "run.attributes" VARCHAR, "run.filename" VARCHAR, "run.semantic_name" VARCHAR, "run.total_bases" DOUBLE, "run.total_spots" DOUBLE, "run.alias" VARCHAR, "run.read_lengths" VARCHAR, "run.base_counts" VARCHAR, "run.r1_length" BIGINT, "run.r2_length" BIGINT, "run.r3_length" BIGINT, "run.r4_length" BIGINT, "run.Acount" BIGINT, "run.Ccount" BIGINT, "run.Gcount" BIGINT, "run.Tcount" BIGINT, "run.Ncount" BIGINT, "run.experiment" VARCHAR, "run.pool_member" VARCHAR, "submission.accession" VARCHAR, "submission.srasource" VARCHAR, "submission.bioprojectsource" VARCHAR, "seqdetective.n_mates" BIGINT, "seqdetective.mapping_rate.mate1" DOUBLE, "seqdetective.mapping_rate.mate2" DOUBLE, "seqdetective.nofeature_rate.mate1" DOUBLE, "seqdetective.nofeature_rate.mate2" DOUBLE, "seqdetective.sparsity.mate1" DOUBLE, "seqdetective.sparsity.mate2" DOUBLE, "seqdetective.pos_strand_rate.mate1" DOUBLE, "seqdetective.pos_strand_rate.mate2" DOUBLE, "seqdetective.readlen.mate1" BIGINT, "seqdetective.readlen.mate2" BIGINT, "seqdetective.judgement.mate1" VARCHAR, "seqdetective.judgement.mate2" VARCHAR, "seqdetective.judgement.reason" VARCHAR, platform_family VARCHAR, instrument_generation VARCHAR, read_bias VARCHAR, selection_class VARCHAR, prep_kit VARCHAR, sc_or_bulk VARCHAR, tech_class VARCHAR, technology VARCHAR, tech_variant VARCHAR, "submission.bioprojectsource.country" VARCHAR, earliest_date DATE, devstage_curation VARCHAR, devstage_curation_coarse VARCHAR, tissue_curation VARCHAR, tissue_curation_coarse VARCHAR);;
CREATE INDEX idx_run_bioproject ON run_metadata(bioproject);;
CREATE INDEX idx_run_run_accession ON run_metadata("run.accession");;