{"database": "metadata", "table": "run_metadata", "rows": [[46305, "SRR7662119", "SRX4522743", "SRS3641059", "SRP131808", "PRJNA432257", "Single cell RNA sequencing of zebrafish beta cells from various stages", "GSE109881", "Transcriptome Analysis", "Age associated deterioration of cellular physiology leads to pathological conditions  and detection of premature aging could provide a window for preventive therapies against age related diseases. For this  methods that accurately evaluate cellular age are required. However  such techniques are currently limited and based on post hoc evaluation using a limited set of histological markers. Development of a technique capable of predicting cellular age  and its modifiers  requires a framework that can robustly handle the noise associated with single cell sampling protocols. Here  we implement GERAS GEnetic Reference for Age of Single cell  a machine learning based framework capable of assigning individual cells to chronological stages based on their transcriptomes.  GERAS displayed greater than 90% accuracy in predicting the chronological stage of zebrafish beta cells and human pancreatic cells. The framework demonstrates robustness against biological and technical noise  as evaluated by its performance on independent samplings of single cells. Additionally  GERAS enabled the evaluation of differences in calorie intake and body mass index on the aging of zebrafish and human cells  respectively.  We further harnessed the predictive power of GERAS to identify genome wide molecular factors that correlate with aging. We show that one of these factors  junb  which declines in expression with aging  is necessary to maintain the proliferative state of juvenile beta cells. Our results showcase the applicability of a machine learning framework to predict the chronological stage of heterogeneous cell populations. The study demonstrates the utility of stage classifiers in assessing pro aging factors  and uncovering candidate genes associated with premature aging. Overall design: We used fluorescence activated cell sorting FACS coupled with next generation RNA Sequencing to profile beta cells from various stages. Cells were sorted into a 96 well plates and single cell library prepared using SMART Seq v4 Ultra Low Input RNA Kit . Sequencing was performed on llumina Nextseq500 aiming at an average sequencing depth of 0.5 million reads per cell. Reads were splice aligned to the zebrafish genome  GRCz10  using HISAT2. htseq count was used to assign reads to exons thus eventually getting counts per gene.", null, "pubmed:30464314", null, "4mpf IF 49", "GSM3325365", null, "tissue:beta cells|age:4mpf|strain:Tgins:BB1.0L|feeding:Intermittent feeding", "4mpf IF 49", "Trimming using trim galore using default parameters Mapping using HISAT2 with default parameters Counts per gene generated using htseq count with default parameters Genome build: Zebrafish GRCz10", "beta cells", null, "FACS SMART Seq v4", null, "age:4mpf|strain:Tgins:BB1.0L|feeding:Intermittent feeding", "GSM3325365", "GSM3325365: 4mpf IF 49; Danio rerio; RNA Seq", "GSM3325365", null, "1", "FACS SMART Seq v4", "GEO Accession:GSM3325365", "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "SINGLE", "ILLUMINA", "NextSeq 500", null, "SRP131808", null, null, "4mpf_IF_49.fastq.gz", "fastq", 15299560.0, 201310.0, "GSM3325365 r1", "0:76 1:0", "A:4227832;C:3362859;G:3368691;T:4339906;N:272", 76, 0, null, null, 4227832, 3362859, 3368691, 4339906, 272, "SRX4522743", "SRS3641059", "SRA654064", "GEO", "Ninov Lab, Center for Regenerative Therapies Dresden", 1, 0.84899, null, 0.19228, null, 0.95572, null, 0.61479, null, 76, null, "B", null, "usable mapping rate", "illumina", "nextseq", "unknown", "cdna_unspecified", "unknown", "sc", "single_cell_plate", "smartseq", null, "Germany", "2018-08-08", "Adult", "Adult", "Pancreas", "Endocrine System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["46305"], "units": {}, "query_ms": 10.922784997092094}