rowid,run.accession,experiment.accession,sample.accession,study.accession,bioproject,study.title,study.alias,study.type,study.abstract,study.attributes,study.PMIDs,sample.description,sample.title,sample.alias,sample.centername,sample.attributes,GEOsample.title,GEOsample.dataprocessing,GEOsample.source,GEOsample.treatmentprotocol,GEOsample.extractprotocol,GEOsample.growthprotocol,GEOsample.characteristics,GEOsample.accession,experiment.title,experiment.alias,experiment.library_name,experiment.design_description,experiment.library_construction_protocol,experiment.attributes,experiment.library_strategy,experiment.library_source,experiment.library_selection,experiment.library_layout,experiment.platform,experiment.instrument_model,experiment.spot_descriptor,experiment.study_ref,run.title,run.attributes,run.filename,run.semantic_name,run.total_bases,run.total_spots,run.alias,run.read_lengths,run.base_counts,run.r1_length,run.r2_length,run.r3_length,run.r4_length,run.Acount,run.Ccount,run.Gcount,run.Tcount,run.Ncount,run.experiment,run.pool_member,submission.accession,submission.srasource,submission.bioprojectsource,seqdetective.n_mates,seqdetective.mapping_rate.mate1,seqdetective.mapping_rate.mate2,seqdetective.nofeature_rate.mate1,seqdetective.nofeature_rate.mate2,seqdetective.sparsity.mate1,seqdetective.sparsity.mate2,seqdetective.pos_strand_rate.mate1,seqdetective.pos_strand_rate.mate2,seqdetective.readlen.mate1,seqdetective.readlen.mate2,seqdetective.judgement.mate1,seqdetective.judgement.mate2,seqdetective.judgement.reason,platform_family,instrument_generation,read_bias,selection_class,prep_kit,sc_or_bulk,tech_class,technology,tech_variant,submission.bioprojectsource.country,earliest_date,devstage_curation,devstage_curation_coarse,tissue_curation,tissue_curation_coarse 43939,SRR6176649,SRX3287366,SRS2596840,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW067 f2,GSM2813936,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW067 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813936,GSM2813936: DEW067 f2; Danio rerio; RNA Seq,GSM2813936,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813936,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-067_S10.R1.fastq.gz FC1_DEW-067_S10.R2.fastq.gz,fastq fastq,2551096506.0,30568258.0,GSM2813936 r1,,A:557861215;C:486414729;G:614056373;T:892106619;N:657570,,,,,557861215,486414729,614056373,892106619,657570,SRX3287366,SRS2596840,SRA619743,GEO,Harvard University,2,0.74967,0.00628,0.22106,0.006,0.79527,0.99945,0.55023,0.46875,34,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43940,SRR6176650,SRX3287366,SRS2596840,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW067 f2,GSM2813936,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW067 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813936,GSM2813936: DEW067 f2; Danio rerio; RNA Seq,GSM2813936,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813936,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-067_S10.R1.fastq.gz FC2_DEW-067_S10.R2.fastq.gz,fastq fastq,1155281120.0,13849592.0,GSM2813936 r2,,A:233187378;C:202743748;G:364218355;T:354503618;N:628021,,,,,233187378,202743748,364218355,354503618,628021,SRX3287366,SRS2596840,SRA619743,GEO,Harvard University,2,0.75419,0.02424,0.22148,0.02238,0.79415,0.99943,0.55145,0.9,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43941,SRR6176651,SRX3287366,SRS2596840,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW067 f2,GSM2813936,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW067 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813936,GSM2813936: DEW067 f2; Danio rerio; RNA Seq,GSM2813936,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813936,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-067_S10.R1.fastq.gz FC3_DEW-067_S10.R2.fastq.gz,fastq fastq,2630905456.0,31531644.0,GSM2813936 r3,,A:571354506;C:501012015;G:621212959;T:936052979;N:1272997,,,,,571354506,501012015,621212959,936052979,1272997,SRX3287366,SRS2596840,SRA619743,GEO,Harvard University,2,0.75366,0.0038,0.22017,0.00351,0.79265,0.99931,0.5547,0.63414,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43942,SRR6176652,SRX3287366,SRS2596840,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW067 f2,GSM2813936,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW067 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813936,GSM2813936: DEW067 f2; Danio rerio; RNA Seq,GSM2813936,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813936,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-067_S10.R1.fastq.gz FC4_DEW-067_S10.R2.fastq.gz,fastq fastq,2426720147.0,29074926.0,GSM2813936 r4,,A:521871975;C:458708124;G:576717935;T:869109171;N:312942,,,,,521871975,458708124,576717935,869109171,312942,SRX3287366,SRS2596840,SRA619743,GEO,Harvard University,2,0.75679,0.00616,0.22081,0.00577,0.79074,0.99896,0.55926,0.67692,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43943,SRR6176653,SRX3287366,SRS2596840,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW067 f2,GSM2813936,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW067 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813936,GSM2813936: DEW067 f2; Danio rerio; RNA Seq,GSM2813936,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813936,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC5_DEW-067_S4.R1.fastq.gz FC5_DEW-067_S4.R2.fastq.gz,fastq fastq,7689824768.0,92165385.0,GSM2813936 r5,,A:1694218914;C:1478360575;G:1791783776;T:2724399126;N:1062377,,,,,1694218914,1478360575,1791783776,2724399126,1062377,SRX3287366,SRS2596840,SRA619743,GEO,Harvard University,2,0.75345,0.00154,0.22254,0.00137,0.79263,0.99967,0.55441,0.57142,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43944,SRR6176644,SRX3287365,SRS2596839,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW066 f2,GSM2813935,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW066 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813935,GSM2813935: DEW066 f2; Danio rerio; RNA Seq,GSM2813935,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813935,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-066_S9.R1.fastq.gz FC1_DEW-066_S9.R2.fastq.gz,fastq fastq,2140044898.0,25651167.0,GSM2813935 r1,,A:467615827;C:408349284;G:516714028;T:746812741;N:553018,,,,,467615827,408349284,516714028,746812741,553018,SRX3287365,SRS2596839,SRA619743,GEO,Harvard University,2,0.74758,0.00524,0.2137,0.00493,0.7987,0.99941,0.55876,0.52631,34,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43945,SRR6176645,SRX3287365,SRS2596839,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW066 f2,GSM2813935,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW066 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813935,GSM2813935: DEW066 f2; Danio rerio; RNA Seq,GSM2813935,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813935,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-066_S9.R1.fastq.gz FC2_DEW-066_S9.R2.fastq.gz,fastq fastq,965058666.0,11574298.0,GSM2813935 r2,,A:195003483;C:170100194;G:303694091;T:295746532;N:514366,,,,,195003483,170100194,303694091,295746532,514366,SRX3287365,SRS2596839,SRA619743,GEO,Harvard University,2,0.75025,0.01899,0.21414,0.01762,0.7977,0.99943,0.5594,0.24832,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43946,SRR6176646,SRX3287365,SRS2596839,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW066 f2,GSM2813935,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW066 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813935,GSM2813935: DEW066 f2; Danio rerio; RNA Seq,GSM2813935,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813935,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-066_S9.R1.fastq.gz FC3_DEW-066_S9.R2.fastq.gz,fastq fastq,2086254612.0,25012693.0,GSM2813935 r3,,A:452988540;C:397479617;G:494445479;T:740347977;N:992999,,,,,452988540,397479617,494445479,740347977,992999,SRX3287365,SRS2596839,SRA619743,GEO,Harvard University,2,0.75183,0.00319,0.21451,0.0029,0.79429,0.99935,0.56364,0.58536,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43947,SRR6176647,SRX3287365,SRS2596839,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW066 f2,GSM2813935,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW066 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813935,GSM2813935: DEW066 f2; Danio rerio; RNA Seq,GSM2813935,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813935,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-066_S9.R1.fastq.gz FC4_DEW-066_S9.R2.fastq.gz,fastq fastq,1923868374.0,23057073.0,GSM2813935 r4,,A:413966332;C:363922629;G:458666692;T:687070548;N:242173,,,,,413966332,363922629,458666692,687070548,242173,SRX3287365,SRS2596839,SRA619743,GEO,Harvard University,2,0.75485,0.00489,0.21313,0.00454,0.79281,0.99908,0.55638,0.54237,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43948,SRR6176648,SRX3287365,SRS2596839,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW066 f2,GSM2813935,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW066 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813935,GSM2813935: DEW066 f2; Danio rerio; RNA Seq,GSM2813935,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813935,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC5_DEW-066_S3.R1.fastq.gz FC5_DEW-066_S3.R2.fastq.gz,fastq fastq,6104866273.0,73196289.0,GSM2813935 r5,,A:1345028270;C:1173765932;G:1427797225;T:2157425704;N:849142,,,,,1345028270,1173765932,1427797225,2157425704,849142,SRX3287365,SRS2596839,SRA619743,GEO,Harvard University,2,0.75337,0.00127,0.21518,0.00111,0.79437,0.99961,0.54968,0.65,33,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43949,SRR6176639,SRX3287364,SRS2596838,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW065 f2,GSM2813934,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW065 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813934,GSM2813934: DEW065 f2; Danio rerio; RNA Seq,GSM2813934,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813934,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-065_S8.R1.fastq.gz FC1_DEW-065_S8.R2.fastq.gz,fastq fastq,2299056187.0,27557923.0,GSM2813934 r1,,A:501686139;C:441327292;G:550895560;T:804557888;N:589308,,,,,501686139,441327292,550895560,804557888,589308,SRX3287364,SRS2596838,SRA619743,GEO,Harvard University,2,0.73689,0.00588,0.20924,0.00551,0.80273,0.99926,0.5692,0.55555,33,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43950,SRR6176640,SRX3287364,SRS2596838,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW065 f2,GSM2813934,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW065 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813934,GSM2813934: DEW065 f2; Danio rerio; RNA Seq,GSM2813934,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813934,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-065_S8.R1.fastq.gz FC2_DEW-065_S8.R2.fastq.gz,fastq fastq,1037253750.0,12441178.0,GSM2813934 r2,,A:209154151;C:183891879;G:324728467;T:318916935;N:562318,,,,,209154151,183891879,324728467,318916935,562318,SRX3287364,SRS2596838,SRA619743,GEO,Harvard University,2,0.74123,0.01962,0.21258,0.01799,0.80318,0.99926,0.5667,0.8983,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43951,SRR6176641,SRX3287364,SRS2596838,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW065 f2,GSM2813934,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW065 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813934,GSM2813934: DEW065 f2; Danio rerio; RNA Seq,GSM2813934,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813934,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-065_S8.R1.fastq.gz FC3_DEW-065_S8.R2.fastq.gz,fastq fastq,2440058543.0,29255489.0,GSM2813934 r3,,A:529090031;C:467802013;G:573700898;T:868289954;N:1175647,,,,,529090031,467802013,573700898,868289954,1175647,SRX3287364,SRS2596838,SRA619743,GEO,Harvard University,2,0.74341,0.00364,0.21329,0.00338,0.79983,0.99949,0.56272,0.45945,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43952,SRR6176642,SRX3287364,SRS2596838,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW065 f2,GSM2813934,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW065 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813934,GSM2813934: DEW065 f2; Danio rerio; RNA Seq,GSM2813934,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813934,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-065_S8.R1.fastq.gz FC4_DEW-065_S8.R2.fastq.gz,fastq fastq,2248915479.0,26953029.0,GSM2813934 r4,,A:482877461;C:428438228;G:531631238;T:805679337;N:289215,,,,,482877461,428438228,531631238,805679337,289215,SRX3287364,SRS2596838,SRA619743,GEO,Harvard University,2,0.74561,0.00511,0.21378,0.00469,0.79876,0.99898,0.56089,0.54285,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43953,SRR6176643,SRX3287364,SRS2596838,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW065 f2,GSM2813934,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW065 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813934,GSM2813934: DEW065 f2; Danio rerio; RNA Seq,GSM2813934,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813934,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC5_DEW-065_S2.R1.fastq.gz FC5_DEW-065_S2.R2.fastq.gz,fastq fastq,7093789012.0,85056439.0,GSM2813934 r5,,A:1558003473;C:1373608340;G:1645617228;T:2515581430;N:978541,,,,,1558003473,1373608340,1645617228,2515581430,978541,SRX3287364,SRS2596838,SRA619743,GEO,Harvard University,2,0.74182,0.00129,0.21487,0.00114,0.80154,0.99965,0.55567,0.5,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43954,SRR6176634,SRX3287363,SRS2596837,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW064 f2,GSM2813933,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW064 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813933,GSM2813933: DEW064 f2; Danio rerio; RNA Seq,GSM2813933,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813933,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-064_S7.R1.fastq.gz FC1_DEW-064_S7.R2.fastq.gz,fastq fastq,2382079622.0,28603442.0,GSM2813933 r1,,A:531702920;C:455344015;G:577305407;T:817119830;N:607450,,,,,531702920,455344015,577305407,817119830,607450,SRX3287363,SRS2596837,SRA619743,GEO,Harvard University,2,0.71162,0.00613,0.2075,0.00587,0.80081,0.99945,0.55002,0.625,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43955,SRR6176635,SRX3287363,SRS2596837,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW064 f2,GSM2813933,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW064 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813933,GSM2813933: DEW064 f2; Danio rerio; RNA Seq,GSM2813933,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813933,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-064_S7.R1.fastq.gz FC2_DEW-064_S7.R2.fastq.gz,fastq fastq,1103522256.0,13268132.0,GSM2813933 r2,,A:228512922;C:195157576;G:347887797;T:331375946;N:588015,,,,,228512922,195157576,347887797,331375946,588015,SRX3287363,SRS2596837,SRA619743,GEO,Harvard University,2,0.70781,0.02127,0.20824,0.01963,0.79776,0.99922,0.56118,0.89265,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43956,SRR6176636,SRX3287363,SRS2596837,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW064 f2,GSM2813933,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW064 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813933,GSM2813933: DEW064 f2; Danio rerio; RNA Seq,GSM2813933,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813933,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-064_S7.R1.fastq.gz FC3_DEW-064_S7.R2.fastq.gz,fastq fastq,2482547494.0,29821327.0,GSM2813933 r3,,A:550735946;C:473925826;G:590960648;T:865736614;N:1188460,,,,,550735946,473925826,590960648,865736614,1188460,SRX3287363,SRS2596837,SRA619743,GEO,Harvard University,2,0.71704,0.00423,0.20978,0.00397,0.79618,0.99945,0.5645,0.60526,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43957,SRR6176637,SRX3287363,SRS2596837,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW064 f2,GSM2813933,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW064 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813933,GSM2813933: DEW064 f2; Danio rerio; RNA Seq,GSM2813933,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813933,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-064_S7.R1.fastq.gz FC4_DEW-064_S7.R2.fastq.gz,fastq fastq,2339730381.0,28089056.0,GSM2813933 r4,,A:515317502;C:444338876;G:558021847;T:821750487;N:301669,,,,,515317502,444338876,558021847,821750487,301669,SRX3287363,SRS2596837,SRA619743,GEO,Harvard University,2,0.71723,0.00493,0.20856,0.00455,0.79417,0.999,0.53427,0.5,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43958,SRR6176638,SRX3287363,SRS2596837,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW064 f2,GSM2813933,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW064 f2,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813933,GSM2813933: DEW064 f2; Danio rerio; RNA Seq,GSM2813933,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813933,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC5_DEW-064_S1.R1.fastq.gz FC5_DEW-064_S1.R2.fastq.gz,fastq fastq,7092166646.0,85201470.0,GSM2813933 r5,,,,,,,,,,,,SRX3287363,SRS2596837,SRA619743,GEO,Harvard University,2,0.70151,0.00121,0.20646,0.00113,0.7917,0.99971,0.56301,0.53333,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43959,SRR6176630,SRX3287362,SRS2596836,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW063 f1,GSM2813932,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW063 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813932,GSM2813932: DEW063 f1; Danio rerio; RNA Seq,GSM2813932,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813932,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-063_S6.R1.fastq.gz FC1_DEW-063_S6.R2.fastq.gz,fastq fastq,2703662317.0,32411788.0,GSM2813932 r1,,A:585487740;C:518930190;G:654356649;T:944203697;N:684041,,,,,585487740,518930190,654356649,944203697,684041,SRX3287362,SRS2596836,SRA619743,GEO,Harvard University,2,0.73487,0.005,0.20711,0.00478,0.8057,0.99953,0.56192,0.5,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43960,SRR6176631,SRX3287362,SRS2596836,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW063 f1,GSM2813932,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW063 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813932,GSM2813932: DEW063 f1; Danio rerio; RNA Seq,GSM2813932,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813932,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-063_S6.R1.fastq.gz FC2_DEW-063_S6.R2.fastq.gz,fastq fastq,1228068654.0,14732853.0,GSM2813932 r2,,,,,,,,,,,,SRX3287362,SRS2596836,SRA619743,GEO,Harvard University,2,0.72354,0.01516,0.20566,0.01388,0.79793,0.9992,0.57243,0.8492,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43961,SRR6176632,SRX3287362,SRS2596836,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW063 f1,GSM2813932,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW063 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813932,GSM2813932: DEW063 f1; Danio rerio; RNA Seq,GSM2813932,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813932,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-063_S6.R1.fastq.gz FC3_DEW-063_S6.R2.fastq.gz,fastq fastq,2865383991.0,34360219.0,GSM2813932 r3,,A:616333652;C:548965462;G:681196824;T:1017502380;N:1385673,,,,,616333652,548965462,681196824,1017502380,1385673,SRX3287362,SRS2596836,SRA619743,GEO,Harvard University,2,0.74018,0.00341,0.20674,0.00318,0.80188,0.99953,0.56249,0.51612,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43962,SRR6176633,SRX3287362,SRS2596836,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW063 f1,GSM2813932,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW063 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813932,GSM2813932: DEW063 f1; Danio rerio; RNA Seq,GSM2813932,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813932,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-063_S6.R1.fastq.gz FC4_DEW-063_S6.R2.fastq.gz,fastq fastq,2661201306.0,31897292.0,GSM2813932 r4,,A:567795696;C:506539944;G:635339481;T:951184460;N:341725,,,,,567795696,506539944,635339481,951184460,341725,SRX3287362,SRS2596836,SRA619743,GEO,Harvard University,2,0.74623,0.00451,0.20997,0.00422,0.79949,0.99928,0.56605,0.4375,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43963,SRR6176626,SRX3287361,SRS2596835,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW062 f1,GSM2813931,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW062 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813931,GSM2813931: DEW062 f1; Danio rerio; RNA Seq,GSM2813931,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813931,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-062_S5.R1.fastq.gz FC1_DEW-062_S5.R2.fastq.gz,fastq fastq,2694116042.0,32276313.0,GSM2813931 r1,,A:589473486;C:513444615;G:648306227;T:942198524;N:693190,,,,,589473486,513444615,648306227,942198524,693190,SRX3287361,SRS2596835,SRA619743,GEO,Harvard University,2,0.75006,0.00665,0.22251,0.00642,0.79776,0.99953,0.56954,0.53571,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43964,SRR6176627,SRX3287361,SRS2596835,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW062 f1,GSM2813931,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW062 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813931,GSM2813931: DEW062 f1; Danio rerio; RNA Seq,GSM2813931,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813931,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-062_S5.R1.fastq.gz FC2_DEW-062_S5.R2.fastq.gz,fastq fastq,1221053333.0,14634722.0,GSM2813931 r2,,A:246914186;C:214804950;G:383568470;T:375108039;N:657688,,,,,246914186,214804950,383568470,375108039,657688,SRX3287361,SRS2596835,SRA619743,GEO,Harvard University,2,0.7539,0.0214,0.22301,0.02015,0.79681,0.99945,0.56759,0.82089,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43965,SRR6176628,SRX3287361,SRS2596835,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW062 f1,GSM2813931,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW062 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813931,GSM2813931: DEW062 f1; Danio rerio; RNA Seq,GSM2813931,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813931,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-062_S5.R1.fastq.gz FC3_DEW-062_S5.R2.fastq.gz,fastq fastq,2732077263.0,32737784.0,GSM2813931 r3,,A:593509115;C:520050361;G:645630165;T:971575716;N:1311906,,,,,593509115,520050361,645630165,971575716,1311906,SRX3287361,SRS2596835,SRA619743,GEO,Harvard University,2,0.75396,0.00435,0.22336,0.00415,0.79257,0.99957,0.566,0.62068,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43966,SRR6176629,SRX3287361,SRS2596835,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW062 f1,GSM2813931,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW062 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813931,GSM2813931: DEW062 f1; Danio rerio; RNA Seq,GSM2813931,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813931,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-062_S5.R1.fastq.gz FC4_DEW-062_S5.R2.fastq.gz,fastq fastq,2513049040.0,30103618.0,GSM2813931 r4,,A:540807825;C:475246041;G:596916568;T:899752561;N:326045,,,,,540807825,475246041,596916568,899752561,326045,SRX3287361,SRS2596835,SRA619743,GEO,Harvard University,2,0.7577,0.00581,0.22002,0.00546,0.79131,0.99912,0.5594,0.69491,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43967,SRR6176622,SRX3287360,SRS2596834,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW061 f1,GSM2813930,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW061 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813930,GSM2813930: DEW061 f1; Danio rerio; RNA Seq,GSM2813930,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813930,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-061_S4.R1.fastq.gz FC1_DEW-061_S4.R2.fastq.gz,fastq fastq,3154139976.0,37801014.0,GSM2813930 r1,,A:693516487;C:600830347;G:760784710;T:1098206077;N:802355,,,,,693516487,600830347,760784710,1098206077,802355,SRX3287360,SRS2596834,SRA619743,GEO,Harvard University,2,0.74925,0.00596,0.22537,0.00573,0.79271,0.99961,0.55809,0.5,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43968,SRR6176623,SRX3287360,SRS2596834,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW061 f1,GSM2813930,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW061 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813930,GSM2813930: DEW061 f1; Danio rerio; RNA Seq,GSM2813930,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813930,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-061_S4.R1.fastq.gz FC2_DEW-061_S4.R2.fastq.gz,fastq fastq,1440125590.0,17268172.0,GSM2813930 r2,,A:292214170;C:252368952;G:455430577;T:439332662;N:779229,,,,,292214170,252368952,455430577,439332662,779229,SRX3287360,SRS2596834,SRA619743,GEO,Harvard University,2,0.75669,0.02314,0.22842,0.02167,0.79393,0.99943,0.55096,0.91666,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43969,SRR6176624,SRX3287360,SRS2596834,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW061 f1,GSM2813930,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW061 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813930,GSM2813930: DEW061 f1; Danio rerio; RNA Seq,GSM2813930,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813930,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-061_S4.R1.fastq.gz FC3_DEW-061_S4.R2.fastq.gz,fastq fastq,3242023471.0,38863564.0,GSM2813930 r3,,A:708050133;C:616691319;G:767462847;T:1148237613;N:1581559,,,,,708050133,616691319,767462847,1148237613,1581559,SRX3287360,SRS2596834,SRA619743,GEO,Harvard University,2,0.75359,0.00376,0.22736,0.00349,0.79123,0.99939,0.55827,0.54054,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43970,SRR6176625,SRX3287360,SRS2596834,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW061 f1,GSM2813930,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW061 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813930,GSM2813930: DEW061 f1; Danio rerio; RNA Seq,GSM2813930,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813930,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-061_S4.R1.fastq.gz FC4_DEW-061_S4.R2.fastq.gz,fastq fastq,2998511059.0,35932183.0,GSM2813930 r4,,A:648015203;C:566278950;G:715347339;T:1068484671;N:384896,,,,,648015203,566278950,715347339,1068484671,384896,SRX3287360,SRS2596834,SRA619743,GEO,Harvard University,2,0.75921,0.00595,0.22572,0.00561,0.78782,0.9991,0.5551,0.54545,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43971,SRR6176618,SRX3287359,SRS2596833,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW060 f1,GSM2813929,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW060 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813929,GSM2813929: DEW060 f1; Danio rerio; RNA Seq,GSM2813929,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813929,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-060_S3.R1.fastq.gz FC1_DEW-060_S3.R2.fastq.gz,fastq fastq,2408131518.0,28860483.0,GSM2813929 r1,,A:529387121;C:457366350;G:579506465;T:841263578;N:608004,,,,,529387121,457366350,579506465,841263578,608004,SRX3287359,SRS2596833,SRA619743,GEO,Harvard University,2,0.74623,0.00668,0.22553,0.00648,0.79776,0.99959,0.56194,0.33333,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43972,SRR6176619,SRX3287359,SRS2596833,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW060 f1,GSM2813929,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW060 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813929,GSM2813929: DEW060 f1; Danio rerio; RNA Seq,GSM2813929,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813929,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-060_S3.R1.fastq.gz FC2_DEW-060_S3.R2.fastq.gz,fastq fastq,1150468859.0,13794506.0,GSM2813929 r2,,A:233434027;C:201107426;G:362797690;T:352510992;N:618724,,,,,233434027,201107426,362797690,352510992,618724,SRX3287359,SRS2596833,SRA619743,GEO,Harvard University,2,0.75233,0.02671,0.22646,0.0248,0.79454,0.99955,0.55277,0.91351,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43973,SRR6176620,SRX3287359,SRS2596833,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW060 f1,GSM2813929,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW060 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813929,GSM2813929: DEW060 f1; Danio rerio; RNA Seq,GSM2813929,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813929,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-060_S3.R1.fastq.gz FC3_DEW-060_S3.R2.fastq.gz,fastq fastq,2490214284.0,29850978.0,GSM2813929 r3,,A:542990500;C:472486099;G:588451265;T:885110481;N:1175939,,,,,542990500,472486099,588451265,885110481,1175939,SRX3287359,SRS2596833,SRA619743,GEO,Harvard University,2,0.75217,0.00417,0.22495,0.00392,0.79371,0.99945,0.55373,0.57575,33,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43974,SRR6176621,SRX3287359,SRS2596833,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW060 f1,GSM2813929,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW060 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813929,GSM2813929: DEW060 f1; Danio rerio; RNA Seq,GSM2813929,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813929,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-060_S3.R1.fastq.gz FC4_DEW-060_S3.R2.fastq.gz,fastq fastq,2335848886.0,27991106.0,GSM2813929 r4,,A:504358369;C:440180845;G:555523061;T:835493929;N:292682,,,,,504358369,440180845,555523061,835493929,292682,SRX3287359,SRS2596833,SRA619743,GEO,Harvard University,2,0.75612,0.00625,0.22331,0.00595,0.7907,0.99933,0.55242,0.55555,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43975,SRR6176614,SRX3287358,SRS2596832,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW059 f1,GSM2813928,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW059 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813928,GSM2813928: DEW059 f1; Danio rerio; RNA Seq,GSM2813928,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813928,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-059_S2.R1.fastq.gz FC1_DEW-059_S2.R2.fastq.gz,fastq fastq,2468427036.0,29617911.0,GSM2813928 r1,,A:542065912;C:474767208;G:599126438;T:851830072;N:637406,,,,,542065912,474767208,599126438,851830072,637406,SRX3287358,SRS2596832,SRA619743,GEO,Harvard University,2,0.72834,0.00622,0.21554,0.00591,0.79888,0.99941,0.54791,0.62162,33,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43976,SRR6176615,SRX3287358,SRS2596832,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW059 f1,GSM2813928,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW059 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813928,GSM2813928: DEW059 f1; Danio rerio; RNA Seq,GSM2813928,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813928,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-059_S2.R1.fastq.gz FC2_DEW-059_S2.R2.fastq.gz,fastq fastq,1116808176.0,13413942.0,GSM2813928 r2,,A:227203028;C:198647634;G:352135639;T:338217127;N:604748,,,,,227203028,198647634,352135639,338217127,604748,SRX3287358,SRS2596832,SRA619743,GEO,Harvard University,2,0.72547,0.02085,0.21465,0.01937,0.79878,0.99949,0.55459,0.91719,28,49,T,B,sc-like readlen,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43977,SRR6176616,SRX3287358,SRS2596832,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW059 f1,GSM2813928,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW059 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813928,GSM2813928: DEW059 f1; Danio rerio; RNA Seq,GSM2813928,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813928,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-059_S2.R1.fastq.gz FC3_DEW-059_S2.R2.fastq.gz,fastq fastq,2433240009.0,29205364.0,GSM2813928 r3,,A:530808152;C:467393483;G:580237225;T:853631720;N:1169429,,,,,530808152,467393483,580237225,853631720,1169429,SRX3287358,SRS2596832,SRA619743,GEO,Harvard University,2,0.73343,0.00393,0.21608,0.00374,0.79446,0.99961,0.55546,0.66666,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43978,SRR6176617,SRX3287358,SRS2596832,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW059 f1,GSM2813928,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW059 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813928,GSM2813928: DEW059 f1; Danio rerio; RNA Seq,GSM2813928,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813928,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-059_S2.R1.fastq.gz FC4_DEW-059_S2.R2.fastq.gz,fastq fastq,2254112176.0,27041760.0,GSM2813928 r4,,A:488078877;C:430617586;G:538610970;T:796514276;N:290467,,,,,488078877,430617586,538610970,796514276,290467,SRX3287358,SRS2596832,SRA619743,GEO,Harvard University,2,0.73283,0.00521,0.21537,0.0049,0.79375,0.9991,0.5546,0.56603,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43979,SRR6176610,SRX3287357,SRS2596831,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW058 f1,GSM2813927,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW058 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813927,GSM2813927: DEW058 f1; Danio rerio; RNA Seq,GSM2813927,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813927,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC1_DEW-058_S1.R1.fastq.gz FC1_DEW-058_S1.R2.fastq.gz,fastq fastq,2403771793.0,28792952.0,GSM2813927 r1,,A:517810269;C:461480878;G:579161807;T:844718642;N:600197,,,,,517810269,461480878,579161807,844718642,600197,SRX3287357,SRS2596831,SRA619743,GEO,Harvard University,2,0.75451,0.00612,0.2178,0.00596,0.79953,0.99963,0.5419,0.5,34,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43980,SRR6176611,SRX3287357,SRS2596831,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW058 f1,GSM2813927,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW058 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813927,GSM2813927: DEW058 f1; Danio rerio; RNA Seq,GSM2813927,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813927,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC2_DEW-058_S1.R1.fastq.gz FC2_DEW-058_S1.R2.fastq.gz,fastq fastq,1112151214.0,13326194.0,GSM2813927 r2,,,,,,,,,,,,SRX3287357,SRS2596831,SRA619743,GEO,Harvard University,2,0.75157,0.01674,0.22743,0.01541,0.79145,0.99924,0.55163,0.86311,34,49,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43981,SRR6176612,SRX3287357,SRS2596831,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW058 f1,GSM2813927,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW058 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813927,GSM2813927: DEW058 f1; Danio rerio; RNA Seq,GSM2813927,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813927,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC3_DEW-058_S1.R1.fastq.gz FC3_DEW-058_S1.R2.fastq.gz,fastq fastq,2422140902.0,29018774.0,GSM2813927 r3,,A:518176419;C:464271902;G:572892842;T:865632668;N:1167071,,,,,518176419,464271902,572892842,865632668,1167071,SRX3287357,SRS2596831,SRA619743,GEO,Harvard University,2,0.76201,0.00368,0.21966,0.00343,0.79334,0.99951,0.55787,0.57142,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 43982,SRR6176613,SRX3287357,SRS2596831,SRP120009,PRJNA414416,Simultaneous single cell profiling of lineages and cell types in the vertebrate brain,GSE105010,Other,The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries,,pubmed:29608178,,DEW058 f1,GSM2813927,,source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf,DEW058 f1,Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored.,zebrafish brain,,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,,tissue:brain|developmental stage:23 25dpf,GSM2813927,GSM2813927: DEW058 f1; Danio rerio; RNA Seq,GSM2813927,,1,Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed and fragmented. The three prime fragments were reverse transcribed and prepared for sequencing Libraries were prepared as described in Zilionis et al. 2017 Nature Protocols PMID = 27929523,GEO Accession:GSM2813927,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP120009,,loader:fastq load.py|options: appendBCtoName,FC4_DEW-058_S1.R1.fastq.gz FC4_DEW-058_S1.R2.fastq.gz,fastq fastq,2262034665.0,27092828.0,GSM2813927 r4,,A:479344618;C:430199692;G:539208487;T:812990390;N:291478,,,,,479344618,430199692,539208487,812990390,291478,SRX3287357,SRS2596831,SRA619743,GEO,Harvard University,2,0.76665,0.00556,0.22471,0.00524,0.79123,0.99935,0.55873,0.64814,34,50,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2017-10-16,Larval,Larval,Brain,Nervous System 48402,SRR7240617,SRX4146448,SRS3360068,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,cell line B fish,GSM3167490,,tissue:melanoma cell line cells in a fish|time:NA|type:CEL Seq,cell line B fish,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,melanoma cell line cells in a fish,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:NA|type:CEL Seq,GSM3167490,GSM3167490: cell line B fish; Danio rerio; RNA Seq,GSM3167490,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167490,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,Efish_R2_001.fastq.gz Efish_R1_001.fastq.gz,fastq fastq,7254277529.0,93829840.0,GSM3167490 r1,0:21.77 1:55.55,A:2013320976;C:1304302014;G:1275408938;T:2658668719;N:2576882,21,55,,,2013320976,1304302014,1275408938,2658668719,2576882,SRX4146448,SRS3360068,SRA713129,GEO,"Yanai, NYU",2,0.35315,0.81249,0.33399,0.25257,0.99222,0.89526,0.4613,0.58245,26,58,T,B,sc-like readlen,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48403,SRR7240616,SRX4146447,SRS3360067,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,cell line B cultured,GSM3167489,,tissue:melanoma cell line cells|time:NA|type:CEL Seq,cell line B cultured,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,melanoma cell line cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:NA|type:CEL Seq,GSM3167489,GSM3167489: cell line B cultured; Danio rerio; RNA Seq,GSM3167489,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167489,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,Edish_R1_001.fastq.gz Edish_R2_001.fastq.gz,fastq fastq,3315328468.0,40362508.0,GSM3167489 r1,0:24.95 1:57.19,A:870288472;C:583199815;G:546968405;T:1313319011;N:1552765,24,57,,,870288472,583199815,546968405,1313319011,1552765,SRX4146447,SRS3360067,SRA713129,GEO,"Yanai, NYU",2,0.4358,0.77687,0.41224,0.1804,0.98673,0.92372,0.39048,0.55564,26,58,T,B,sc-like readlen,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48404,SRR7240615,SRX4146446,SRS3360066,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,cell line A fish,GSM3167488,,tissue:melanoma cell line cells in a fish|time:NA|type:CEL Seq,cell line A fish,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,melanoma cell line cells in a fish,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:NA|type:CEL Seq,GSM3167488,GSM3167488: cell line A fish; Danio rerio; RNA Seq,GSM3167488,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167488,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,Afish_R1_001.fastq.gz Afish_R2_001.fastq.gz,fastq fastq,1760981184.0,27515331.0,GSM3167488 r1,0:13 1:51,A:527501645;C:354257934;G:378832308;T:500290272;N:99025,13,51,,,527501645,354257934,378832308,500290272,99025,SRX4146446,SRS3360066,SRA713129,GEO,"Yanai, NYU",2,0.0,0.86857,0.0,0.2721,1.0,0.84869,,0.53458,13,51,T,B,sc-like readlen,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48405,SRR7240614,SRX4146445,SRS3360065,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,cell line A cultured,GSM3167487,,tissue:melanoma cell line cells|time:NA|type:CEL Seq,cell line A cultured,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,melanoma cell line cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:NA|type:CEL Seq,GSM3167487,GSM3167487: cell line A cultured; Danio rerio; RNA Seq,GSM3167487,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167487,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,Adish_R1_001.fastq.gz Adish_R2_001.fastq.gz,fastq fastq,2758577088.0,43102767.0,GSM3167487 r1,0:13 1:51,A:901840607;C:537128308;G:580843998;T:738597709;N:166466,13,51,,,901840607,537128308,580843998,738597709,166466,SRX4146445,SRS3360065,SRA713129,GEO,"Yanai, NYU",2,0.0,0.73258,0.0,0.1063,1.0,0.85865,,0.58942,13,51,T,B,sc-like readlen,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48406,SRR7240613,SRX4146444,SRS3360064,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF3 ST,GSM3167486,,tissue:tumor tissue section|time:NA|type:ST,ZF3 ST,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor tissue section,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:NA|type:ST,GSM3167486,GSM3167486: ZF3 ST; Danio rerio; RNA Seq,GSM3167486,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167486,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00715A_S4_R1_001.fastq.gz BS00715A_S4_R2_001.fastq.gz,fastq fastq,8346596335.0,108397355.0,GSM3167486 r1,0:31 1:46,A:2031003864;C:1619238171;G:2052969755;T:2622927017;N:20457528,31,46,,,2031003864,1619238171,2052969755,2622927017,20457528,SRX4146444,SRS3360064,SRA713129,GEO,"Yanai, NYU",2,0.0219,0.7716,0.01972,0.18796,0.99429,0.83086,0.55276,0.5937,31,46,T,B,mate1 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48407,SRR7240612,SRX4146443,SRS3360063,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF2 time point 2,GSM3167485,,tissue:tumor cells|time:2017 06 08|type:inDrop,ZF2 time point 2,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 08|type:inDrop,GSM3167485,GSM3167485: ZF2 time point 2; Danio rerio; RNA Seq,GSM3167485,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167485,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00418A_S2_R1_001.fastq.gz BS00418A_S2_R2_001.fastq.gz,fastq fastq,18879921248.0,219533968.0,GSM3167485 r1,0:35 1:51,A:4915572934;C:4065059620;G:5109553396;T:4789479927;N:255371,35,51,,,4915572934,4065059620,5109553396,4789479927,255371,SRX4146443,SRS3360063,SRA713129,GEO,"Yanai, NYU",2,0.22339,0.01677,0.04998,0.01242,0.904,0.99795,0.64786,0.76205,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48408,SRR7240611,SRX4146442,SRS3360062,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF2 time point 1,GSM3167484,,tissue:tumor cells|time:2017 05 25|type:inDrop,ZF2 time point 1,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 05 25|type:inDrop,GSM3167484,GSM3167484: ZF2 time point 1; Danio rerio; RNA Seq,GSM3167484,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167484,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00417A_S1_R1_001.fastq.gz BS00417A_S1_R2_001.fastq.gz,fastq fastq,15195762260.0,176694910.0,GSM3167484 r1,0:35 1:51,A:3730288221;C:3100069649;G:3917397047;T:4447806244;N:201099,35,51,,,3730288221,3100069649,3917397047,4447806244,201099,SRX4146442,SRS3360062,SRA713129,GEO,"Yanai, NYU",2,0.45584,0.02194,0.10866,0.01434,0.8508,0.99594,0.56562,0.74387,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48409,SRR7240610,SRX4146441,SRS3360061,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 2 time point 3,GSM3167483,,tissue:tumor cells|time:2017 06 22|type:inDrop,ZF1 tumor 2 time point 3,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 22|type:inDrop,GSM3167483,GSM3167483: ZF1 tumor 2 time point 3; Danio rerio; RNA Seq,GSM3167483,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167483,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00516A_S3_R1_001.fastq.gz BS00516A_S3_R2_001.fastq.gz,fastq fastq,9387695500.0,109159250.0,GSM3167483 r1,0:35 1:51,A:1996385781;C:1819073805;G:2713262409;T:2858501572;N:471933,35,51,,,1996385781,1819073805,2713262409,2858501572,471933,SRX4146441,SRS3360061,SRA713129,GEO,"Yanai, NYU",2,0.78336,0.12929,0.17042,0.07562,0.83116,0.98196,0.62361,0.7167,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48410,SRR7240609,SRX4146440,SRS3360060,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 2 time point 2,GSM3167482,,tissue:tumor cells|time:2017 06 15|type:inDrop,ZF1 tumor 2 time point 2,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 15|type:inDrop,GSM3167482,GSM3167482: ZF1 tumor 2 time point 2; Danio rerio; RNA Seq,GSM3167482,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167482,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00515A_S2_R1_001.fastq.gz BS00515A_S2_R2_001.fastq.gz,fastq fastq,7645768166.0,88904281.0,GSM3167482 r1,0:35 1:51,A:1659830380;C:1483879485;G:1955626018;T:2546036303;N:395980,35,51,,,1659830380,1483879485,1955626018,2546036303,395980,SRX4146440,SRS3360060,SRA713129,GEO,"Yanai, NYU",2,0.78059,0.04637,0.15414,0.03167,0.81156,0.99125,0.5786,0.82113,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48411,SRR7240608,SRX4146439,SRS3360059,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 2 time point 1,GSM3167481,,tissue:tumor cells|time:2017 06 02|type:inDrop,ZF1 tumor 2 time point 1,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 02|type:inDrop,GSM3167481,GSM3167481: ZF1 tumor 2 time point 1; Danio rerio; RNA Seq,GSM3167481,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167481,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,BS00514A_S1_R1_001.fastq.gz BS00514A_S1_R2_001.fastq.gz,fastq fastq,9759139906.0,113478371.0,GSM3167481 r1,0:35 1:51,A:2243979852;C:1912724883;G:2573948349;T:3027979990;N:506832,35,51,,,2243979852,1912724883,2573948349,3027979990,506832,SRX4146439,SRS3360059,SRA713129,GEO,"Yanai, NYU",,,,,,,,,,,,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48412,SRR7240607,SRX4146438,SRS3360058,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 1 time point 4,GSM3167480,,tissue:tumor cells|time:2017 06 22|type:inDrop,ZF1 tumor 1 time point 4,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 22|type:inDrop,GSM3167480,GSM3167480: ZF1 tumor 1 time point 4; Danio rerio; RNA Seq,GSM3167480,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167480,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,t4_R1.fastq.gz t4_R2.fastq.gz,fastq fastq,13285828250.0,154486375.0,GSM3167480 r1,0:35 1:51,A:3106404784;C:2642499289;G:3548717541;T:3987624595;N:582041,35,51,,,3106404784,2642499289,3548717541,3987624595,582041,SRX4146438,SRS3360058,SRA713129,GEO,"Yanai, NYU",2,0.588,0.04697,0.13733,0.02918,0.8308,0.99119,0.59025,0.75017,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48413,SRR7240606,SRX4146437,SRS3360057,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 1 time point 3,GSM3167479,,tissue:tumor cells|time:2017 06 15|type:inDrop,ZF1 tumor 1 time point 3,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 15|type:inDrop,GSM3167479,GSM3167479: ZF1 tumor 1 time point 3; Danio rerio; RNA Seq,GSM3167479,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167479,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,t3_R1.fastq.gz t3_R2.fastq.gz,fastq fastq,12949174972.0,150571802.0,GSM3167479 r1,0:35 1:51,A:3045479995;C:2532873724;G:3340677661;T:4029571801;N:571791,35,51,,,3045479995,2532873724,3340677661,4029571801,571791,SRX4146437,SRS3360057,SRA713129,GEO,"Yanai, NYU",2,0.56814,0.02383,0.11626,0.01895,0.82599,0.99569,0.55364,0.73492,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48414,SRR7240605,SRX4146436,SRS3360056,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 1 time point 2,GSM3167478,,tissue:tumor cells|time:2017 06 09|type:inDrop,ZF1 tumor 1 time point 2,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 09|type:inDrop,GSM3167478,GSM3167478: ZF1 tumor 1 time point 2; Danio rerio; RNA Seq,GSM3167478,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167478,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,t2_R1.fastq.gz t2_R2.fastq.gz,fastq fastq,24232410834.0,281772219.0,GSM3167478 r1,0:35 1:51,A:5863941889;C:4979860831;G:6744452007;T:6643024623;N:1131484,35,51,,,5863941889,4979860831,6744452007,6643024623,1131484,SRX4146436,SRS3360056,SRA713129,GEO,"Yanai, NYU",2,0.36566,0.03272,0.07206,0.03013,0.85719,0.99673,0.55683,0.65034,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 48415,SRR7240604,SRX4146435,SRS3360055,SRP149420,PRJNA473915,Single cell analysis of tumor progression reveals the function structure and evolution of cancer archetypes,GSE115140,Transcriptome Analysis,The classic cancer evolution model posits that driver mutations sweep the population sequentially as the complete set of hallmarks are assembled by the neoplastic clone. However recent work has challenged this model revealing that most tumors contain highly complex dynamics with genetic diversity reflecting distinct clonal architectures. The functional and phenotypic heterogeneity also has been shown to have a crucial influence on the fate of the tumor5–10. However it is not well understood how distinct tumor clonal populations coexist and function. Here we study tumor architecture at the level of individual cells by sampling a zebrafish melanoma tumor over time and space. We found that cancer transcriptional programs can be classified to three archetypes each exploiting the existing neural crest mature melanocytes and stress modules and distinct intra tumor locations. Strikingly these archetypes are conserved in human melanoma. Further we found that the cancer cells are comprised of two distinct clones where one expresses a unique archetype. Over time we found that the cells of this clone adapt by exhibiting a more similar profile to the corresponding archetype. Overall design: Single cell RNA sequencing of zebrafish tumor cells from 2 zebrafish at multiple time points.,,,,ZF1 tumor 1 time point 1,GSM3167477,,tissue:tumor cells|time:2017 06 02|type:inDrop,ZF1 tumor 1 time point 1,Illumina RTA v2 software was used for basecalling and quality determination. Raw sequencing data obtained from the inDrop method was processed using a custom built pipeline available at https://github.com/flo compbio/singlecell. Briefly the location of the known “W1” adapter sequence of the inDrop RT primer was located in the barcode read read 2. Reads for which the W1 sequence could not be detected were discarded. The start position of the W1 sequence was then used to infer the length of the first part of the inDrop cell barcode in each read which can range from 8 11 bp as well as the start position of the second part of the inDrop cell barcode which is 8 bp long. Cell barcode sequences were mapped to the known list of 384 barcode sequences for each read. The resulting barcode combination was used to identify the cell from which the fragment originated. Finally UMI sequence was extracted and reads with low confidence base calls for the six bases comprising the UMI sequence minimum PHRED score less than 20 were discarded. The reads containing the mRNA sequence read 1 were mapped using STAR with parameter “—outSAMmultNmax 1” and default settings otherwise27. Expression was quantified by counting the number of reads mapped to each gene and correcting for UMI as described previously Grün et al. 2014 The genome and gff file used included the zebrafish genome and the BRAF human vector. Single cell transcriptomes with UMIs>750 mitochondrial transcripts < 20% and ribosomal transcripts < 30% were retained for analysis. Raw sequencing data obtained from the Spatial Transcriptomics ST method were processed using a publicly available pipeline https://github.com/jfnavarro/st pipeline. Briefly quality trimming is performed to remove low quality bases and reads with long nucleotide stretches > 15. Read 2 transcript sequence is mapped with STAR 2.5.1 and Read 1 spatial barcode is demultiplexed with Taggd. Reads that contain both a valid spatial barcode and are correctly map are kept. UMIs are then counted with htseq count; to get a final read count annotated reads are grouped by spatial barcode. Raw sequencing data obtained from the CEL Seq2 method were processed using a publicly available pipeline https://github.com/yanailab/celseq2. Genome build: Danio rerio.GRCz10 Supplementary files format and content: For inDrop data: tab delimited files with genes as rows and cells as columns. For ST data: tab delimited files with spots x and y coordinates for the ST array as columns and genes as rows. For CEL Seq data: csv files with genes as rows and cells as columns.,tumor cells,,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,,time:2017 06 02|type:inDrop,GSM3167477,GSM3167477: ZF1 tumor 1 time point 1; Danio rerio; RNA Seq,GSM3167477,,1,For inDrop and CEL Seq samples were obtained using a dissecting forceps and placing the desired tissue sample in 1.5 mL Eppendorf tube followed by addition of 500 uL 0.25% Trypsin EDTA for digestion. The digestion was carried out at 37ºC in thermomixer for 15 30 min to soften tissue every 5 minutes mashing the tissue using a disposable pestle to break up softened tissue. Upon completion of incubation at 37ºC 500 uL of DMEM10 were added to deactivate the trypsin. Cells were washed three times by spinning down the sample at 500 rcf for 5 minutes and resuspended in PBS. Then sample was filtered twice using 5 mL polystyrene round bottom tube with 35um cell strainer. Viability and single cell consistency were checked prior to encapsulation of the cells using the inDrop system for each biopsy taken. For the ST sample sections from zebrafish melanoma tumors were obtained by sectioning the entire tumor with its surrounding. Tissue was gently washed with cold 1X PBS and 4 5 mm3 cubes were removed with a scalpel for OCT embedding. Tissue was transferred from 1X PBS to a dry sterile 10 cm dish and gently dried prior to equilibration in cold OCT for 2 minutes. The tissue was then transferred to a tissue mold with OCT and snap frozen in liquid nitrogen chilled isopentane. Tissue blocks were stored at 80°C until further use. Prior to cryosectioning the cryostat was cleaned with 100% ethanol and equilibrated to an internal temperature of 18°C for 30 minutes. Once equilibrated OCT embedded tissue blocks were mounted onto the chuck and equilibrated to the cryostat temperature for 15 20 minutes prior to trimming. ST slide was also placed inside cryostat to keep the slide cold and minimize RNase activity. Sections were cut at 10 µm sections and mounted onto the ST arrays and stored at 80°C until use maximum of two weeks. Prior to fixation and staining the ST array was removed from the 80C and into a RNase free biosafety hood for 5 minutes to bring to room temperature followed by warming on a 37°C heat block for 1 minute. Tissue was fixed for 10 minutes with 3.6% formaldehyde in 1X PBS and subsequently rinsed in 1x PBS. Next the tissue was dehydrated with isopropanol for 1 minute followed by staining with hematoxylin and eosin. Slides were mounted in 65 µl 80% glycerol and brightfield images were taken on a Leica 397 SCN400 F whole slide scanner at 40X resolution. For inDrop library construction the cells and reverse transcription RT reaction was carried out as previously described in Klein et al. 2015. RNA amplification and library preparation were carried out according to this protocol incorporating the changes introduced in Zilionis et al. 2017. Briefly RNA was reverse transcribed RT with SuperScript III Invitrogen in droplets. Droplet emulsions were broken and post RT material underwent second strand syntehsis and in vitro transcription using the T7 High Yield Enzyme mix New England Biolabs. RNA was fragmented for 3 minutes with 1X fragmentation reagent prior to RT with random hexamers eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For spatial transcriptomics ST library construction we followed the previously published ST protocol from Stahl et al. Science 2016 with minor changes. Briefly post brightfield imaging of stained tissue tissue was permeabilized with collagenase and 0.1% pepsin prior to an overnight RT step with SuperScript III on the ST slide. Tissue was digested away with proteinase K and 1% 2 mercaptoethanol prior to cleavage of probes from the slide surface with USER enzyme. Second strand synthesis was performed with DNA Pol I and RNase H New england biolabs followed by an in vitro transcription amplification step with the MEGAScript T7 kit Thermofisher. Amplified RNA then underwent a second RT with random hexamers and SuperScript II Invitrogen eliminating the need for an adaptor ligation step. Cycles required for final library amplification was assessed by quantitative PCR qPCR with KAPA HiFi Hot Start PCR Mix KAPA Biosystems and EvaGreen dye Biotium. Final libraries were amplified with KAPA HiFi Hot Start PCR mix for 9 to 13 cycles. inDrop library size assessed on a DNA BioAnalyzer chip following the manufacturer's instructions Agilent. For CEL Seq2 sample cells were sorted using fluorescence activated cell sorting FACS into 384 well plate contacting 1.2ul primer mix and library was constructed according to the CEL Seq2 protocol Hashimshony et al. 2016.,GEO Accession:GSM3167477,RNA-Seq,TRANSCRIPTOMIC,cDNA,PAIRED,ILLUMINA,NextSeq 500,,SRP149420,,,t1_R1.fastq.gz t1_R2.fastq.gz,fastq fastq,9252930920.0,107592220.0,GSM3167477 r1,0:35 1:51,A:2335222288;C:1982182390;G:2612836318;T:2322125944;N:563980,35,51,,,2335222288,1982182390,2612836318,2322125944,563980,SRX4146435,SRS3360055,SRA713129,GEO,"Yanai, NYU",2,0.23022,0.0283,0.05144,0.02308,0.90254,0.99608,0.58852,0.70642,35,51,B,T,mate2 technical by mapping diff,illumina,nextseq,unknown,random_priming,unknown,sc,single_cell_droplet,indrops,,United States,2018-05-31,Undetermined,Undetermined,Cancer or Tumor,Cancer or Tumor 75092,SRR24239717,SRX20035842,SRS17374857,SRP433739,PRJNA958104,The secreted neuronal signal Spock1 regulates the blood brain barrier,GSE230236,Transcriptome Analysis,The blood brain barrier BBB is a unique set of properties of the brain vasculature which severely restricts its permeability to proteins and small molecules. Classic chick quail chimera studies showed that these properties are not intrinsic to the brain vasculature but rather are induced by surrounding neural tissue. Here we identify Spock1 as a candidate neuronal signal for regulating BBB permeability in zebrafish and mice. Mosaic genetic analysis shows that neuronally expressed Spock1 is cell non autonomously required for a functional BBB. Leakage in spock1 mutants is associated with altered extracellular matrix ECM increased endothelial transcytosis and altered pericyte endothelial interactions. Furthermore a single dose of recombinant SPOCK1 into spock1 mutants quenches gelatinase activity restores vascular expression of BBB genes including mcamb and partially restores barrier function. These analyses support a model in which neuronally secreted Spock1 induces BBB properties by altering the ECM thereby regulating pericyte endothelial interactions and downstream vascular gene expression. Overall design: Bulk RNAseq Libraries 66 71 of leaky mutant and wild type siblings to identify the genetic lesion responsible for the leaky phenotype. These analyses revealed linkage to chr14 and more specifically to the spock1 gene. scRNAseq Library scDRBrain of dissected spock1 mutant and wild type brains was then used to identify all cell type specific changes in gene expression in the mutant background.,,pubmed:37437574,,scRNA seq for WT and hm41 larval heads 3 and 5dpf,GSM7208220,,tissue:mixed|cell type:mixed|genotype:mixed|time:3 dpf 5 dpf|geo loc name:missing|collection date:missing,scRNA seq for WT and hm41 larval heads 3 and 5dpf,Sequencing reads were mapped to the Zebrafish GRCz11 R101 genome assembly using a custom python pipeline as previously described see Zilionis et al. Nature Protocols 2017 and https://github.com/indrops/. Multi seq hashtags were identified using custom code available at: https://github.com/AllonKleinLab/klunctions/tree/master/Ignas/Hashing. We first removed background cell barcodes by only considering transcriptomes with greater than 350 UMIs for downstream analysis. In order to remove dead cells transcriptomes were further filtered by mitochondrial read percentage >20%. Cell demultiplexing was performed by manual inspection. Specifically thresholds were drawn to delineate single cells from background and multiplet populations. The resulting counts matrix was normalized to the mean UMIs per cell in the dataset. Assembly: GRCz11 Supplementary files format and content: The gene counts matrix is an output from rsem differential gene expression analysis and is a raw counts estimate not normalized for each gene. The first column is gene name and then post that each column represents an individual sample WT for the first 3 and then MUT for the last 3. Supplementary files format and content: The h5ad file contains the counts matrix for the demultiplexed single cell data. This file also holds relevant genotype and timepoint annotations as well as Multi seq barcode counts for each cell.,mixed,,Dissected brain tissues were dissociated using a modified protocol adapted from Bresciani et. al. 2018 PMID: 30364607. Briefly chemical dissociations were performed at 30.5°C using a mixture of 0.25% Trypsin EDTA Collagenase/Dispase 8 mg/mL and DNaseI 20 µg/mL for 15 20 minutes with gentle pipetting every 2.5 minutes. The dissociations were quenched using DMEM + 10% fetal bovine serum and filtered through a 40 µM cell strainer. The dissociation mixtures were spun down twice at 700g for 5 min and washed with PBS. The mixtures were then resuspended in PBS and barcoded using Multi seq as described in McGinnis et. al. 2019 PMID: 31209384 with slight modifications. For each sample 80 pmoles of Lipid modified oligos LMOs were used to hash every 500k cells. The hashing reaction was quenched using PBS + 1% BSA. The barcoded samples were pooled into a single tube and washed twice with PBS + 1% BSA 700g for 5 min.. The pooled cell mixture was resuspended in PBS + 0.1%BSA + 18% Optiprep at a final concentration of 300k cells/mL prior to single cell capture with inDrops. Single cell transcriptomes were captured by the Single cell Core SCC at the Harvard Medical School as previously described Zilionis et al. Nature Protocols 2017 using the inDrops V3 chemistry. The Single cell Core at the Harvard Medical School prepared the gene expression inDrops v3 chemistry and Multi seq libraries.,,cell type:mixed|genotype:mixed|time:3 dpf 5 dpf,GSM7208220,GSM7208220: scRNA seq for WT and hm41 larval heads 3 and 5dpf; Danio rerio; RNA Seq,GSM7208220 r1,GSM7208220,1,Dissected brain tissues were dissociated using a modified protocol adapted from Bresciani et. al. 2018 PMID: 30364607. Briefly chemical dissociations were performed at 30.5°C using a mixture of 0.25% Trypsin EDTA Collagenase/Dispase 8 mg/mL and DNaseI 20 µg/mL for 15 20 minutes with gentle pipetting every 2.5 minutes. The dissociations were quenched using DMEM + 10% fetal bovine serum and filtered through a 40 µM cell strainer. The dissociation mixtures were spun down twice at 700g for 5 min and washed with PBS. The mixtures were then resuspended in PBS and barcoded using Multi seq as described in McGinnis et. al. 2019 PMID: 31209384 with slight modifications. For each sample 80 pmoles of Lipid modified oligos LMOs were used to hash every 500k cells. The hashing reaction was quenched using PBS + 1% BSA. The barcoded samples were pooled into a single tube and washed twice with PBS + 1% BSA 700g for 5 min.. The pooled cell mixture was resuspended in PBS + 0.1%BSA + 18% Optiprep at a final concentration of 300k cells/mL prior to single cell capture with inDrops. Single cell transcriptomes were captured by the Single cell Core SCC at the Harvard Medical School as previously described Zilionis et al. Nature Protocols 2017 using the inDrops V3 chemistry. The Single cell Core at the Harvard Medical School prepared the gene expression inDrops v3 chemistry and Multi seq libraries.,,RNA-Seq,TRANSCRIPTOMIC SINGLE CELL,cDNA,PAIRED,ILLUMINA,Illumina NovaSeq 6000,,SRP433739,,loader:fastq load.py|options: readTypes=BTBT read1PairFiles=Undetermined S0 L001 R1 001.fastq.gz read2PairFiles=Undetermined S0 L001 R2 001.fastq.gz read3PairFiles=Undetermined S0 L001 R3 001.fastq.gz read4PairFiles=Undetermined S0 L001 R4 001.fastq.gz,,,548653863228.0,4729774683.0,GSM7208220 r1,,,,,,,,,,,,SRX20035842,SRS17374857,SRA1728506,"Megason Lab, Systems Biology, Harvard Medical School","Megason Lab, Systems Biology, Harvard Medical School",2,0.83589,0.0,0.23127,0.0,0.7723,1.0,0.46689,,86,8,B,T,sc-like readlen,illumina,novaseq_era,unknown,cdna_unspecified,unknown,sc,single_cell_droplet,indrops,,United States,2023-04-21,Larval,Larval,Multi-tissue,Multi-system