{"database": "metadata", "table": "run_metadata", "is_view": false, "human_description_en": "where technology = \"celseq\" and tissue_curation = \"Fin\"", "rows": [[41278, "SRR6039223", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate01_R1.fastq.gz FinC_plate01_R2.fastq.gz", "fastq fastq", 4156649914.0, 27498834.0, "GSM2781033 r1", "0:75.66 1:75.50", "A:1184667722;C:812111779;G:946047137;T:1213491214;N:332062", 75, 75, null, null, 1184667722, 812111779, 946047137, 1213491214, 332062, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.13191, 0.2761, 0.11086, 0.21281, 0.97997, 0.9586, 0.48699, 0.512, 76, 76, "B", "B", "mate1-mate2 similar by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41279, "SRR6039224", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinE_plate07_R1.fastq.gz FinE_plate07_R2.fastq.gz", "fastq fastq", 7045442122.0, 46667895.0, "GSM2781033 r10", "0:75.65 1:75.32", "A:1992737539;C:1481998828;G:1584855122;T:1984679859;N:1170774", 75, 75, null, null, 1992737539, 1481998828, 1584855122, 1984679859, 1170774, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.08863, 0.2093, 0.07731, 0.18045, 0.97934, 0.95655, 0.51176, 0.52167, 76, 76, "B", "B", "mate1-mate2 similar by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41280, "SRR6039225", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinE_plate08_R1.fastq.gz FinE_plate08_R2.fastq.gz", "fastq fastq", 18530449971.0, 122665653.0, "GSM2781033 r11", "0:75.68 1:75.38", "A:4667348302;C:4333487274;G:4681098112;T:4848415162;N:101121", 75, 75, null, null, 4667348302, 4333487274, 4681098112, 4848415162, 101121, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.05595, 0.12249, 0.04637, 0.10243, 0.98415, 0.96451, 0.51093, 0.52879, 75, 75, "B", "B", "mate1-mate2 similar by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41281, "SRR6039226", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate02_R1.fastq.gz FinC_plate02_R2.fastq.gz", "fastq fastq", 5961095878.0, 39431358.0, "GSM2781033 r2", "0:75.66 1:75.51", "A:1678045732;C:1185915367;G:1391829821;T:1704836561;N:468397", 75, 75, null, null, 1678045732, 1185915367, 1391829821, 1704836561, 468397, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.13015, 0.27027, 0.11023, 0.2177, 0.97985, 0.95964, 0.5491, 0.51821, 76, 76, "B", "B", "mate1-mate2 similar by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41282, "SRR6039227", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate03_R1.fastq.gz FinC_plate03_R2.fastq.gz", "fastq fastq", 6170232577.0, 40846092.0, "GSM2781033 r3", "0:75.56 1:75.50", "A:1866777846;C:1148020988;G:1151073641;T:2003993265;N:366837", 75, 75, null, null, 1866777846, 1148020988, 1151073641, 2003993265, 366837, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.21215, 0.40881, 0.18316, 0.33198, 0.95919, 0.91102, 0.49711, 0.53044, 75, 76, "T", "B", "mate1 technical by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41283, "SRR6039228", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate04_R1.fastq.gz FinC_plate04_R2.fastq.gz", "fastq fastq", 11725606474.0, 77611342.0, "GSM2781033 r4", "0:75.62 1:75.46", "A:3255674293;C:2567072026;G:2847190480;T:3054954945;N:714730", 75, 75, null, null, 3255674293, 2567072026, 2847190480, 3054954945, 714730, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.06975, 0.13988, 0.05976, 0.119, 0.98533, 0.97165, 0.50512, 0.46229, 75, 75, "B", "B", "mate1-mate2 similar by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41284, "SRR6039229", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate05_R1.fastq.gz FinC_plate05_R2.fastq.gz", "fastq fastq", 10721896827.0, 71080622.0, "GSM2781033 r5", "0:75.44 1:75.41", "A:3568478635;C:1747554798;G:1675757438;T:3729472140;N:633816", 75, 75, null, null, 3568478635, 1747554798, 1675757438, 3729472140, 633816, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.32303, 0.52716, 0.2796, 0.43703, 0.94276, 0.90057, 0.51543, 0.52404, 75, 75, "T", "B", "mate1 technical by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41285, "SRR6039230", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinC_plate06_R1.fastq.gz FinC_plate06_R2.fastq.gz", "fastq fastq", 4861560177.0, 32218130.0, "GSM2781033 r6", "0:75.46 1:75.44", "A:1606516776;C:766379277;G:687731226;T:1800636189;N:296709", 75, 75, null, null, 1606516776, 766379277, 687731226, 1800636189, 296709, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.40204, 0.66185, 0.35053, 0.55748, 0.94637, 0.89148, 0.48866, 0.50482, 75, 75, "T", "B", "mate1 technical by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41286, "SRR6039231", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinE_plate02_R1.fastq.gz FinE_plate02_R2.fastq.gz", "fastq fastq", 7812360278.0, 51697480.0, "GSM2781033 r7", "0:75.54 1:75.58", "A:2098454376;C:1641311761;G:1943503636;T:2128270555;N:819950", 75, 75, null, null, 2098454376, 1641311761, 1943503636, 2128270555, 819950, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.0105, 0.02592, 0.00852, 0.02011, 0.99626, 0.98754, 0.33762, 0.53206, 76, 75, "T", "T", "mates < 9% mapping rate", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41287, "SRR6039232", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinE_plate03_R1.fastq.gz FinE_plate03_R2.fastq.gz", "fastq fastq", 6704494628.0, 44436630.0, "GSM2781033 r8", "0:75.49 1:75.39", "A:2087008571;C:1124579474;G:1252171872;T:2240541544;N:193167", 75, 75, null, null, 2087008571, 1124579474, 1252171872, 2240541544, 193167, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.24968, 0.47949, 0.2198, 0.40276, 0.95154, 0.89238, 0.48976, 0.54355, 75, 75, "T", "B", "mate1 technical by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"], [41288, "SRR6039233", "SRX3187382", "SRS2515225", "SRP082370", "PRJNA339266", "Single cell sequencing reveals dissociation induced gene expression in tissue subpopulations", "GSE85755", "Other", "In many gene expression studies  cells are extracted by tissue dissociation and Fluorescence Activated Cell Sorting FACS  but the effect of these protocols on cellular transcriptomes is not well characterized and often ignored. Here  we applied single cell mRNA sequencing scRNA seq to muscle stem cells  and unexpectedly found a subpopulation that is strongly affected by the widely used dissociation protocol that we employed. One implication of this finding is that several published transcriptomics studies may need to be reinterpreted. Importantly  we detected similar subpopulations in other single cell datasets  suggesting that cells from other tissues might be affected by this artefact as well. Overall design: Mouse satellite cells and zebrafish fin cells were extracted from Tibialis Anterior muscles of Pax7nGFP mice and wildtype zebrafish fins  respectively. For cell extraction  traditional Supplementary Methods dissociation protocols that combine mechanical and enzymatic dissociation were employed  and live cells were subsequently sorted into plates using FACS. Next  single cell mRNA sequencing CEL Seq or SORT Seq robotized version of CEL Seq2 was applied  and data was analyzed with RaceID2 to identify clusters. CEL Seq samples: Manual CEL Seq; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 1h collagenase treated default dissociation protocol; 96 cells per plate with 96 different barcodes see \"Cel seq barcodes 96.csv\"; some primes numbers are bulk samples see \"BulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\"; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count tables; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged CEL Seq AllMiceAndLibrariesMerged.csv\" file is count table with reads from all mice and libraries merged Annotation of columns: Zx.y  where Z = mouse  x = library and y = cell barcode; bulk samples are not included any more in this file; See Supplementary Methods for details. SORT Seq 1h and 2h dissociated samples: Robotized CEL Seq2; Satellite cells unstained; Male Pax7nGFP mice 5 mpf 7 mpf; 8 muscles from 4 mice; One plate of 1h default dissociation protocol and one plate of 2h collagenase treated cells; 384 cells per plate with each of the 96 barcodes see \"Cel seq barcodes 96.csv\" used 4 times per plate therefore  each plate has 4 libraries; No bulk samples included; Spike ins included see \"ERCC92.fa\"; No mitochondrial reads in count table; In some wells  we sorted no cell internal negative control; barcodes #95 and #96 were used for empty wells; Sequencing lanes not concatenated in fastq files uploaded here; \"Merged SORT Seq DissociationTimecourse.csv\" file is count table were reads from all dissociation timepoints are merged Annotation columns: DZhx y  where Z = 1 or 2 hours collagenase treated  x = library and y = cell barcode; See Supplementary Methods for details. SORT Seq MitoTracker stained samples pilot and repeat: Robotized CEL Seq2 samples; Satellite cells stained with MitoTracker; Female Pax7nGFP mice 1 4.7 mpf mouse for pilot experiment; 3 mpf 6 mpf mice for repeat experiment; 1h collagenase treated default dissociation protocol; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: pilot experiment has 263 cells so plate was partly empty  repeat experiment done with 4 full plates; No bulk samples included; Spike ins included see \"ERCC92.fa\"; Mitochondrial reads rows named \"*  chrM\" included in count tables these were removed prior to RaceID2; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; No merged file was generated for pilot experiment as only one library  \"Merged MitoTracker Repeat.csv\" file is count table were reads from all plates of repeat experiment were merged Annotation of columns: Plx Welly  where x = plate number 1 4 and x = cell barcode; See Supplementary Methods for details. SORT Seq zebrafish fin samples: Robotized CEL Seq2; Fin cells unstained; all live cells; Wildtype zebrafish; Dissociated using default fin dissociation protocol Supplementary Methods; 384 cells per plate with each of the 384 barcodes see \"Cel seq barcodes 384.csv\" used 1 times per plate therefore  each plate has 1 library; Note: only merged count table file \"fin C E count table.csv\"  Annotation columns: Xx.py.prim.finZ  where x = cell barcode  y = plate number and Z is fish C or E and no individual library count table files were uploaded to GEO for zebrafish fin data; No bulk samples included; Spike ins not included in merged count tables file; Mitochondrial reads not included in merged count tables file; In some wells  we sorted no cell barcodes #357 #360 and #381 #384 were used for empty wells; Sequencing lanes concatenated in fastq files uploaded here; See Supplementary Methods for details.", null, "pubmed:28960196", null, "SORT Seq zebrafish fin merged", "GSM2781033", null, "tissue:All cells from caudal fin|strain:Wildtype", "SORT Seq zebrafish fin merged", "Reads 2 were mapped to the reference transcriptome created from the genomes downloaded from the UCSC genome browser; ERCC Spike in sequences and mitochondrial sequences were added in sense direction using bwa version 0.6.2 r126 with default parameters. All isoforms of the same gene were merged to a single gene locus and reads mapping to multiple loci in the transcriptome were discarded. Reads 1 contains the cel specific barcode information first 8 bases followed by a UMI sequence 4bp for unstained satellite cell data; 6bp for MitoTracker Stained satellite cells and zebrafish data and a polyT stretch. Reads 1 were thus used to extract the cell barcode sequences see \u201cCel seq barcodes 96.csv\u201d file for sequences used for unstained satellite cell data and see \"Cel seq barcodes 384\" for MitoTracker stained satellite cells and zebrafish data and UMIs; see Supplementary Methods for details. post mapping  a UMI correction was applied to the read and barcode counts files to generate unique transcript count tables as described before see Supplementary Methods section for details. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells. Genome build: mus musculus: mm10; danio rerio: Zv9 Ensembl release 74; both were extended with spike ins see \"ERCC92.fa\" Supplementary files format and content: *.coutt.csv  *TranscriptCounts.tsv and *count table.csv: tab separated data files listing how many reads of which transcripts were detected in all sequenced cells post UMI correction. Column names refer to cells sequenced in this library numbers refer the the CEL Seq primer barcode used for that cell. The first column lists official gene symbols followed by the chromosome name  separated by a double underscore. Note that processed files for satellite cell data here still contain ERCC Spike in molecules  bulk samples bulk samples are only included in CEL Seq experiments; see \u201cBulkSamples BarcodesAndNrOfCellsUsed perCEL Seq1 library\u201d file for description of how many cells were used as bulk and which primer barcode sequence was used for bulk sample in each library and mitochondrial reads only for MitoTracker stained satellite cells.", "All cells from caudal fin", null, "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", null, "strain:Wildtype", "GSM2781033", "GSM2781033: SORT Seq zebrafish fin merged; Danio rerio; RNA Seq", "GSM2781033", null, "1", "Cells were sorted into Vapor Lock Qiagen containing a droplet with primers and dNTPs. Cells were lysed at 65 degrees Celsius  post which SORT Seq Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017 was applied. As in SORT Seq protocol Muraro et al.  2016  with minor modifications as desribed in Supplementary Methods of van den Brink et al.  2017.", "GEO Accession:GSM2781033", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP082370", null, null, "FinE_plate04_R2.fastq.gz FinE_plate04_R1.fastq.gz", "fastq fastq", 8402262674.0, 55692887.0, "GSM2781033 r9", "0:75.48 1:75.39", "A:2618499207;C:1444061167;G:1613776766;T:2725680678;N:244856", 75, 75, null, null, 2618499207, 1444061167, 1613776766, 2725680678, 244856, "SRX3187382", "SRS2515225", "SRA453335", "GEO", "Alexander van Oudenaarden, Hubrecht Institute", 2, 0.26069, 0.43322, 0.23296, 0.36903, 0.95663, 0.90542, 0.50653, 0.52921, 76, 75, "T", "B", "mate1 technical by mapping diff", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "celseq", null, "Netherlands", "2017-09-12", "Undetermined", "Adult", "Fin", "Surface Structure"]], "truncated": false, "filtered_table_rows_count": 11, "expanded_columns": [], "expandable_columns": [], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": [], "units": {}, "query": {"sql": "select rowid, [run.accession], [experiment.accession], [sample.accession], [study.accession], bioproject, [study.title], [study.alias], [study.type], [study.abstract], [study.attributes], [study.PMIDs], [sample.description], [sample.title], [sample.alias], [sample.centername], [sample.attributes], [GEOsample.title], [GEOsample.dataprocessing], [GEOsample.source], [GEOsample.treatmentprotocol], [GEOsample.extractprotocol], [GEOsample.growthprotocol], [GEOsample.characteristics], [GEOsample.accession], [experiment.title], [experiment.alias], [experiment.library_name], [experiment.design_description], [experiment.library_construction_protocol], [experiment.attributes], [experiment.library_strategy], [experiment.library_source], [experiment.library_selection], [experiment.library_layout], [experiment.platform], [experiment.instrument_model], [experiment.spot_descriptor], [experiment.study_ref], [run.title], [run.attributes], [run.filename], [run.semantic_name], [run.total_bases], [run.total_spots], [run.alias], [run.read_lengths], [run.base_counts], [run.r1_length], [run.r2_length], [run.r3_length], [run.r4_length], [run.Acount], [run.Ccount], [run.Gcount], [run.Tcount], [run.Ncount], [run.experiment], [run.pool_member], [submission.accession], [submission.srasource], [submission.bioprojectsource], [seqdetective.n_mates], [seqdetective.mapping_rate.mate1], [seqdetective.mapping_rate.mate2], [seqdetective.nofeature_rate.mate1], [seqdetective.nofeature_rate.mate2], [seqdetective.sparsity.mate1], [seqdetective.sparsity.mate2], [seqdetective.pos_strand_rate.mate1], [seqdetective.pos_strand_rate.mate2], [seqdetective.readlen.mate1], [seqdetective.readlen.mate2], [seqdetective.judgement.mate1], [seqdetective.judgement.mate2], [seqdetective.judgement.reason], platform_family, instrument_generation, read_bias, selection_class, prep_kit, sc_or_bulk, tech_class, technology, tech_variant, [submission.bioprojectsource.country], earliest_date, devstage_curation, devstage_curation_coarse, tissue_curation, tissue_curation_coarse from run_metadata where \"technology\" = :p0 and \"tissue_curation\" = :p1 order by rowid limit 101", "params": {"p0": "celseq", "p1": "Fin"}}, "facet_results": {"experiment.library_strategy": {"name": "experiment.library_strategy", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "RNA-Seq", "label": "RNA-Seq", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&experiment.library_strategy=RNA-Seq", "selected": false}], "truncated": false}, "experiment.library_source": {"name": "experiment.library_source", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "TRANSCRIPTOMIC", "label": "TRANSCRIPTOMIC", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&experiment.library_source=TRANSCRIPTOMIC", "selected": false}], "truncated": false}, "experiment.library_selection": {"name": "experiment.library_selection", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "cDNA", "label": "cDNA", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&experiment.library_selection=cDNA", "selected": false}], "truncated": false}, "experiment.library_layout": {"name": "experiment.library_layout", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "PAIRED", "label": "PAIRED", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&experiment.library_layout=PAIRED", "selected": false}], "truncated": false}, "experiment.platform": {"name": "experiment.platform", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "ILLUMINA", "label": "ILLUMINA", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&experiment.platform=ILLUMINA", "selected": false}], "truncated": false}, "devstage_curation_coarse": {"name": "devstage_curation_coarse", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "Adult", "label": "Adult", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&devstage_curation_coarse=Adult", "selected": false}], "truncated": false}, "devstage_curation": {"name": "devstage_curation", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "Undetermined", "label": "Undetermined", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&devstage_curation=Undetermined", "selected": false}], "truncated": false}, "tissue_curation_coarse": {"name": "tissue_curation_coarse", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "Surface Structure", "label": "Surface Structure", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin&tissue_curation_coarse=Surface+Structure", "selected": false}], "truncated": false}, "tissue_curation": {"name": "tissue_curation", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "Fin", "label": "Fin", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?technology=celseq", "selected": true}], "truncated": false}, "technology": {"name": "technology", "type": "column", "hideable": false, "toggle_url": "/metadata/run_metadata.json?technology=celseq&tissue_curation=Fin", "results": [{"value": "celseq", "label": "celseq", "count": 11, "toggle_url": "http://metadata.rnaquarium.org/metadata/run_metadata.json?tissue_curation=Fin", "selected": true}], "truncated": false}}, "suggested_facets": [], "next": null, "next_url": null, "private": false, "allow_execute_sql": true, "query_ms": 91.29026000118756}