{"database": "metadata", "table": "run_metadata", "rows": [[36242, "SRR33873799", "SRX29085358", "SRS25297146", "SRP590547", "PRJNA1272846", "Organism wide contributions to systemic skeletal muscle repair in larval zebrafish.", "GSE299146", "Transcriptome Analysis", "We developed a systemic muscle injury model in zebrafish. Single cell transcriptomic analysis of muscle and non muscle tissues revealed that systemic and local muscle injuries elicit distinct cellular molecular responses  both quantitatively and qualitatively. These studies suggest that large  and small scale muscle injuries activate different regenerative programs  resulting in either systemic or local repair. Overall design: To characterize the regenerative responses and signaling pathways activated post systemic or local muscle injury  we generated Tgactc1b:NTR mCherry transgenic zebrafish lines.At 4 dpf  transgenic larvae were treated with MTZ overnight to systemically injure all myofibers  or were subjected to local mechanical injury needlestick of 2 3 somites.Uninjured and systemically injured larvae were collected at 1  2  and 4 xxx post injury. Mechanically injured larvae were collected at 2 xxx post injury. Sequencing results were compared between uninjured  systemically injured  and mechanically injured siblings.", null, null, null, "larval trunk  2dpi Needle Stick", "GSM9034433", null, "source name:Trunk|tissue:Trunk|genotype:Tgactc1b:NTR mCherry|treatment:Needle Stick|batch:4/20/2022|geo loc name:missing|collection date:missing", "larval trunk  2dpi Needle Stick", "post sequencing  the Illumina output was processed using the CellRanger v8.0.1 pipeline to generate gene barcode count matrices. A custom reference genome was made with the \u201ccellranger mkref\u201d command  using the fasta file of zebrafish reference genome GRCz11 constructed from the Ensemble genome build https://useast.ensembl.org/Danio rerio/Info/Index and the sorted Gene Transfer Format file v4.3.2 from the improved zebrafish transcriptome annotation Lawson et al.  2020. Base call files for each sample from Illumina were demultiplexed into FASTQ reads. Then  the \u201ccellranger count\u201d pipeline was used to align sequencing reads in FASTQ files to the custom reference genome. Both exon and intron sequences were included during the alignment. The filtered gene barcode count matrices generated by \u201ccellranger count\u201d was used for downstream analysis. Datasets were integrated and analyzed using Seurat v4.4.0 package with R v4.4.2Stuart etal.  2019; Team  2024. Each sample count matrix was filtered for genes that were expressed in at least 3 cells and cells expressing at least 200 genes  followed by cell quality assessment usingcommonly used QC matrixes Ilicic et al.  2016. Cells having a unique number of genes between 200 and 6500  mitochondrial gene percentage <5 and total number of reads UMIs between 300and 35000 were used for downstream processing. Each dataset was independently normalizedand scaled using the \u201cSCTransform\u201d function  which is an improved method for normalization  that performs a variance stabilizing transformation using negative binomial regression Hafemeisterand Satija  2019. Standard integration workflow of Seurat was used to identify shared sources of variation across experiments as well as mutual nearest neighbors Butler et al.  2018; Haghverdiet al.  2018. Integration features were selected based on the top 6000 highly variable features using \u201cSelectIntegrationFeatures\u201d function nfeatures\u2009=\u20096000  which was used as input for the\u201canchor.features\u201d argument of the \u201cFindIntegrationAnchors\u201d function. PCA analysis was performed on the 6000 variable features and the top 50 principal components selected based on the elbow plot heuristic  which measures the contribution of variation in each component. These 50 principal components were used in \u201cFindNeighbors\u201d and \u201cFindClusters\u201d functions to perform graph based clustering on a shared nearest neighbor graph Levine et al.  2015; Xu and Su 2015. Louvain algorithm was used for modularity optimization in cell clustering using \u201cFindClusters\u201d function. The resolution parameter res\u2009=\u20090.3 that determines the granularity of clusteringwas selected by visually inspecting clusters with resolutions ranging between 0.1 and 2.0 as well as clustree graph Zappia and Oshlack  2018. Uniform Manifold Approximation and Reduction UMAP was used for non linear dimensional reduction of the first 50 principal components and visualize the data using \u201cRunUMAP\u201d function Becht et al.  2018. Data was graphed using different plot functions  such as \u201cDimPlot\u201d  \u201cVlnPlot\u201d  \u201cFeaturePlot\u201d  \u201cDotplot\u201d and \u201cDoHeatmap\u201d  to view the cell cluster identity and marker gene expression. Cell proportions were extracted using the \u201ctable\u201d and \u201cprop.table\u201d functions. Differential gene expression for individual clusters was identified using Wilcoxon rank sum test in the \u201cFindAllMarkers\u201d function. Marker genes detected in at least 25% of the clustered cells and had a positive average log2FC were reported. Muscle cells identified from the complete dataset were subclustered using the \u201csubset\u201d function for subcluster analysis. The muscle subset was again normalized and scaled using the \u201cSCTransform\u201d function with glmGamPoi method Ahlmann Eltze and Huber  2021. Fifty principal components were used and the resolution parameter was set to 0.4. Downstream analysis was done as described above for the integrated analysis. We used a zebrafish single atlas Sur et al.  2023 to generate a database of markers for each tissue and cell type Table S1. The differentially expressed markers of each cluster were crossreferenced with our compiled database using our \u201cDE Marker Scoring\u201d algorithm Saraswathy et al.  2024. For every matching marker gene  one point was given to the respective cluster under the column name with matching cell identity. Iteration over every marker gene was performed to generate a scoring matrix with varying points for each cluster against the different cell identities compiled in the database Table S2  sheet: scoring. The \u201cphyper\u201d function in R was then used to calculate one tailed binomial probabilities using hypergeometric distribution for the total score obtained from each cluster against each cell identity in the database Table S2  sheet: Binomial probability. \u2212log10 of probability values were obtained for plotting the heatmap Table S2  sheet:\u2212log10P. The resulting values were scaled from 0 to 100 and plotted as a heatmap using GraphPad prism. Each cluster was given an identity based on the maximum \u2212log10 p score obtained in the heatmap. The top DE markers of clusters with ambiguous scores were manually annotated using the \u201cident genes\u201d from the zebrafish single atlas Sur et al.  2023 and literature search McKellar et al.  2021. Top DE markers generated for each cluster is given in Table S3  sheet:topDEmarkers.all.clusters \u201cRenameIdents\u201d function was used to assign identity to each cluster. To confirm the assigned cluster identities  enrichment of classical markers of respective cell types were tested using Dot plot. The R package CellChat v2.1.2 was used to evaluate regenerative cell cell interactions post local and systemic muscle injury Jin et al.  2021. CellChat models the probability of cell cell communication by integrating our gene expression data with a database of known interaction between signaling ligands  receptors  and their cofactors CellChatDB. The RNA data was used to create CellChat object using \u201ccreateCellChat\u201d function  followed by the recommended preprocessing functions with default parameters for the analysis of individual datasets. Truncated mean method with 10% trimmed observation was used to compute average gene expression per cell group for the complete dataset. The default trimean method was used to compute average gene expression for the CellChat analysis. CellChatDB.zebrafish was used to infer cell cell communication. All categories of ligand receptor interactions in the database were used in the analysis. Communications involving less than 10 cells were excluded. The \u201cnetAnalysis computeCentrality\u201d function was used to calculate network centrality scores at each time point. Functions such as \u201cnetVisual circle\u201d  \u201cnetAnalysis contribution\u201d  \u201cnetVisual aggregate\u201d  \u201cnetVisual bubble\u201d  \u201cnetAnalysis signalingRole heatmap\u201d  and \u201cnetAnalysis signalingRole scatter\u201d were used to generate different plots used in this paper. Assembly: zebrafish reference genome GRCz11 and the sorted Gene Transfer Format file v4.3.2 from the improved zebrafish transcriptome annotation Lawson et al.  2020. Supplementary files format and content: Compressed Filtered Feature and Barcode files as .tsv file and count matrices in .mtx file format.", "Trunk", "Fish were treated with 12 hours of metronidazole or were stuck locally with a needle", "15 larvae per cohort were dissected to remove the head  caudal fin  and internal organs and then processed as described Farnsworth DB 459 2020. For tissue lysis  500\u03bcL of dissociation buffer 2ug/ul Collagenase P  2mM CaCl2  0.2 ul DNaseI  0.25% Trypsin  in 1X PBS was added to the samples and incubated at 28 \u00b0C until samples were visually dissociated 15min. Dissociation was stopped with equal volume of stop buffer 10% FBS  0.001mM EDTA  in 1X PBS. Samples were centrifuged at 750rcf for 7min at 4\u00b0C  the supernatant was discarded  and the pellet was resuspended in 100\u03bcL of resuspension buffer 1% FBS  2mM CaCl2  1X Penicillin/Streptomycin  in DMEM. The suspension was strained through 20\u03bcm cell strainer Pluriselect USA  43 10020  40  centrifuged once 500rcf  1 min  4\u00b0C to pass cells through the strainer  and centrifuged a second time 750rcf  7 min  4\u00b0C to pellet the cells. The cells were resuspended in 100\u03bcL of resuspension buffer  centrifugated 750rcf  7min  4\u00b0C  and the pellet was resuspended in 50\u03bcL 0.04% BSA in PBS. 5\u03bcl of cells were then combined with 5\u03bcl of Hoechst Invitrogen  H3570  incubated for 10min  and transferred to a hemocytometer Bulldogbio  NC1731934 to determine cell concentration. An additional 5\u03bcl of cells were incubated with Trypan blue Sigma  T8154 for 5min  transferred to a hemocytometer  and counted to assess cell viability. Samples with > 70% viability were then submitted for single cell sequencing. For snRNA seq  30 \u00b5l of isolated nuclei at a concentration of 1000 nuclei/\u00b5l was submitted to Genome Technology Access Center at McDonnel Genome Institute of Washington University. Two biological replicates of each genotype were used. cDNA was prepared post the GEM generation and barcoding  followed by the GEM RT reaction and bead cleanup steps.  cDNA was amplified for 11 13 cycles then purified using SPRIselect beads. Purified cDNA samples were then run on a Bioanalyzer to determine the cDNA concentration. GEX libraries were prepared as recommended by the 10x Genomics Chromium Single Cell 3\u2019 Reagent Kits User Guide v3.1 Chemistry Dual Index with appropriate modifications to the PCR cycles based on the calculated cDNA concentration. For sample preparation on the 10x Genomics platform  the Chromium Next GEM Single Cell 3\u2019 Kit v3.1  16 rxns PN 1000268  Chromium Next GEM Chip G Single Cell Kit  48 rxns PN 1000120  and Dual Index Kit TT Set A  96 rxns PN 1000215 were used. The concentration of each library was accurately determined through qPCR utilizing the KAPA library Quantification Kit according to the manufacturer\u2019s protocol KAPA Biosystems/Roche to produce cluster counts appropriate for the Illumina NovaSeq6000 instrument. Normalized libraries were sequenced on a NovaSeq6000 S4 Flow Cell using the XP workflow and a 50x10x16x150 sequencing recipe according to manufacturer protocol. A median sequencing depth of 50 000 reads/cell was targeted for each Gene Expression Library.", "Larval zebrafish of 4 dpf underwent drug treatments or needle stick injuries for muscle ablation", "tissue:Trunk|genotype:Tgactc1b:NTR mCherry|treatment:Needle Stick|batch:4/20/2022", "GSM9034433", "GSM9034433: larval trunk  2dpi Needle Stick; Danio rerio; RNA Seq", "GSM9034433 r1", "GSM9034433", "1", "15 larvae per cohort were dissected to remove the head  caudal fin  and internal organs and then processed as described Farnsworth DB 459 2020. For tissue lysis  500\u03bcL of dissociation buffer 2ug/ul Collagenase P  2mM CaCl2  0.2 ul DNaseI  0.25% Trypsin  in 1X PBS was added to the samples and incubated at 28 \u00b0C until samples were visually dissociated 15min. Dissociation was stopped with equal volume of stop buffer 10% FBS  0.001mM EDTA  in 1X PBS. Samples were centrifuged at 750rcf for 7min at 4\u00b0C  the supernatant was discarded  and the pellet was resuspended in 100\u03bcL of resuspension buffer 1% FBS  2mM CaCl2  1X Penicillin/Streptomycin  in DMEM. The suspension was strained through 20\u03bcm cell strainer Pluriselect USA  43 10020  40  centrifuged once 500rcf  1 min  4\u00b0C to pass cells through the strainer  and centrifuged a second time 750rcf  7 min  4\u00b0C to pellet the cells. The cells were resuspended in 100\u03bcL of resuspension buffer  centrifugated 750rcf  7min  4\u00b0C  and the pellet was resuspended in 50\u03bcL 0.04% BSA in PBS. 5\u03bcl of cells were then combined with 5\u03bcl of Hoechst Invitrogen  H3570  incubated for 10min  and transferred to a hemocytometer Bulldogbio  NC1731934 to determine cell concentration. An additional 5\u03bcl of cells were incubated with Trypan blue Sigma  T8154 for 5min  transferred to a hemocytometer  and counted to assess cell viability. Samples with > 70% viability were then submitted for single cell sequencing. For snRNA seq  30 \u00b5l of isolated nuclei at a concentration of 1000 nuclei/\u00b5l was submitted to Genome Technology Access Center at McDonnel Genome Institute of Washington University. Two biological replicates of each genotype were used. cDNA was prepared post the GEM generation and barcoding  followed by the GEM RT reaction and bead cleanup steps.  cDNA was amplified for 11 13 cycles then purified using SPRIselect beads. Purified cDNA samples were then run on a Bioanalyzer to determine the cDNA concentration. GEX libraries were prepared as recommended by the 10x Genomics Chromium Single Cell three prime Reagent Kits User Guide v3.1 Chemistry Dual Index with appropriate modifications to the PCR cycles based on the calculated cDNA concentration. For sample preparation on the 10x Genomics platform  the Chromium Next GEM Single Cell three prime Kit v3.1  16 rxns PN 1000268  Chromium Next GEM Chip G Single Cell Kit  48 rxns PN 1000120  and Dual Index Kit TT Set A  96 rxns PN 1000215 were used. The concentration of each library was accurately determined through qPCR utilizing the KAPA library Quantification Kit according to the manufacturer's protocol KAPA Biosystems/Roche to produce cluster counts appropriate for the Illumina NovaSeq6000 instrument. Normalized libraries were sequenced on a NovaSeq6000 S4 Flow Cell using the XP workflow and a 50x10x16x150 sequencing recipe according to manufacturer protocol. A median sequencing depth of 50 000 reads/cell was targeted for each Gene Expression Library.", null, "RNA-Seq", "TRANSCRIPTOMIC SINGLE CELL", "cDNA", "PAIRED", "ILLUMINA", "Illumina NovaSeq 6000", null, "SRP590547", null, null, "AJAH-ACTC_NTR_Mcherry-Needlestick-2dpi-NS-2dpi-lib1_S1_L003_R1_001.fastq.gz AJAH-ACTC_NTR_Mcherry-Needlestick-2dpi-NS-2dpi-lib1_S1_L003_R2_001.fastq.gz", "fastq fastq", 73442728978.0, 412599601.0, "GSM9034433 r1", "0:28 1:150", "A:21212027449;C:16424751525;G:17879456991;T:17925559458;N:933555", 28, 150, null, null, 21212027449, 16424751525, 17879456991, 17925559458, 933555, "SRX29085358", "SRS25297146", "SRA2144659", "Johnson Lab, Developmental Biology, Washington University in St.Louis", "Johnson Lab, Developmental Biology, Washington University in St.Louis", null, null, null, null, null, null, null, null, null, null, null, "T", "B", "sc-like readlen", "illumina", "novaseq_era", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_droplet", "10x", null, "United States", "2025-06-06", "Larval", "Larval", "Trunk", "Surface Structure"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["36242"], "units": {}, "query_ms": 8.305294999445323}