{"database": "metadata", "table": "run_metadata", "rows": [[60943, "SRR12661676", "SRX9142643", "SRS7383947", "SRP282920", "PRJNA664124", "Emergence of neuronal diversity during vertebrate brain development", "GSE158142", "Transcriptome Analysis", "Neurogenesis comprises many steps from progenitor proliferation to neuronal differentiation and maturation. These processes are highly regulated  but the landscape of transcriptional changes underlying brain development are poorly characterized. Here  we describe a developmental single cell RNA seq catalog of 220 000 zebrafish brain cells encompassing 12 stages from 12 hpf to 15 dpf We characterize known and novel gene markers for 800 clusters and provide an overview of the diversification of neurons and progenitors across these timepoints. We also introduce an optimized version of the GESTALT lineage recorder that enables higher expression and recovery of Cas9 edited barcodes to query lineage segregation. Cell type characterization indicates that most embryonic neural progenitor states are transitory and transcriptionally distinct from neural progenitors of post embryonic stages. Reconstruction of cell specification trajectories reveals that late stage retinal neural progenitors transcriptionally overlap cell states observed in the embryo. The zebrafish brain development atlas provides a resource to define and manipulate specific subsets of neurons and to uncover the molecular mechanisms underlying vertebrate neurogenesis. Overall design: 10X Genomics v2 scRNA seq libraries and scGESTALT", null, "pubmed:33068532", null, "Gest zBr15dpf12", "GSM4793260", null, "source name:zebrafish brain|tissue:brain|developmental stage:15dpf", "Gest zBr15dpf12", "Single cell RNA Sequencing data FASTQ files were processed using Cell Ranger v2.0.2 to generate count matrices Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. File uploaded as Danio rerio.GRCz10.86.modified.gtf.gz. Some libraries were sequenced twice as technical replicates marked as *b in raw files and the processing was done using the combination of all data processed data files have comb* prefix scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 Supplementary files format and content: *barcodes.tsv files contain 10X Genomics cell barcodes  *genes.tsv files contain gene names  *matrix.mtx contain transcriptome count data  *web summary.html contain summary statistics for transcript mapping  .rds files are the processed R objects  .txt files contain marker genes identified for each cluster in the dataset  URD*.rds files are the processed R objects for performing URD cell trajectory analysis. tree*.rds files are the final URD objects with cell trajectory trees for the retina and hypothalamus. Supplementary files format and content: scGESTALT data:  Gest zBr15dpf8 and Gest zBr15dpf9 correspond to two samples from the same zebrafish ZF1. Gest zBr15dpf10 and Gest zBr15dpf11 correspond to two samples from the same zebrafish ZF2. Gest zBr15dpf12 and Gest zBr15dpf13 correspond to two samples from the same zebrafish ZF3. Gest zBr15dpf14 and Gest zBr15dpf15 correspond to two samples from the same zebrafish ZF4. Supplementary files format and content: scGESTALT data: *GestMaster.txt contains the 10X Genmoics cell identifiers CellBarcode and CellUMI that were used to matchlineage barcodes to transcriptomes. They also contain lineage barcode sequences HMID column   the t SNE cluster membership number ClusterIdent column  and unique identifier Readname. zfAllMerge FILTEREDUSEME.xlsx contains the merged statistics of all samples  filtered to contain high confidence barcodes  the barcode sequence aligned to a reference unedited sequence mergedRead column  mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. Further details of columns can be found at https://github.com/shendurelab/Cas9FateMapping", "zebrafish brain", null, "Gestalt barcode was PCR amplified. Sample indices and flow cell adaptors were then added by PCR. Single cell suspensions were processed through 10X Genomics v2 kits according to manufacturer's protocol  to generate single cell cDNA libraries and prepared for sequencing", null, "tissue:brain|developmental stage:15dpf", "GSM4793260", "GSM4793260: Gest zBr15dpf12; Danio rerio; RNA Seq", "GSM4793260", null, "1", "Gestalt barcode was PCR amplified. Sample indices and flow cell adaptors were then added by PCR. Single cell suspensions were processed through 10X Genomics v2 kits according to manufacturer's protocol  to generate single cell cDNA libraries and prepared for sequencing", "GEO Accession:GSM4793260", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "ILLUMINA", "NextSeq 500", null, "SRP282920", null, "loader:fastq load.py|options:  platform=Illumina   readTypes=TTB   read1PairFiles=zBr15dpf12 I1 001.fastq.gz   read2PairFiles=zBr15dpf12 R1 001.fastq.gz   read3PairFiles=zBr15dpf12 R2 001.fastq.gz", "zBr15dpf12_I1_001.fastq.gz zBr15dpf12_R1_001.fastq.gz zBr15dpf12_R2_001.fastq.gz", "fastq fastq fastq", 1077176310.0, 3663865.0, "GSM4793260 r1", "0:8 1:26 2:260", "A:267515738;C:322832426;G:276056686;T:210748138;N:23322", 8, 26, 260, null, 267515738, 322832426, 276056686, 210748138, 23322, "SRX9142643", "SRS7383947", "SRA1127180", "GEO", "Harvard University", 1, 2e-05, null, 0.0, null, 0.99995, null, 0.5, null, 260, null, "T", null, "under 1.2% mapping rate", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_droplet", "10x", null, "United States", "2020-09-17", "Larval", "Larval", "Brain", "Nervous System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["60943"], "units": {}, "query_ms": 10.324526003387291}