{"database": "metadata", "table": "run_metadata", "rows": [[70589, "SRR20001192", "SRX16042026", "SRS13721694", "SRP385134", "PRJNA856272", "Synonymous codon massive reporter library in zebrafish embryos", "GSE207584", "Transcriptome Analysis", "Messenger RNA mRNA stability substantially impacts steady state gene expression levels in a cell. mRNA stability is strongly affected by codon composition in a translation dependent manner across species  through a mechanism termed codon optimality. We have developed iCodon www.iCodon.org  an algorithm for customizing mRNA expression through the introduction of synonymous codon substitutions into the coding sequence. iCodon is optimized for four vertebrate transcriptomes: mouse  human  frog  and fish. Users can predict the mRNA stability of any coding sequence based on its codon composition and subsequently generate more stable optimized or unstable deoptimized variants encoding for the same protein. Further  we show that codon optimality predictions correlate with both mRNA stability using a massive reporter library and expression levels using fluorescent reporters and analysis of endogenous gene expression in zebrafish embryos and/or human cells. Therefore  iCodon will benefit basic biological research  as well as a wide range of applications for biotechnology and biomedicine. Overall design: We designed a library of 1 600 sequences containing 100 different 100 amino acid long coding sequences capturing the average composition of the zebrafish proteome  with 10 variants each designed using iCodon to have 5 bins of increasing predicted stability with 2 sequences in each bin. The groups were referred as iCodon 1  iCodon 2  iCodon 3  iCodon 4  and iCodon 5  from less to more stable mRNA predictions. Additionally  5 synonymous sequences were designed using 5 independent iterations of IDT's codon optimization tool and 1 synonymous sequence using Genewiz's codon optimization tool as this method returned the exact same sequences in each independent run. These sequences were cloned downstream of 11 fixed codons to control for similar translation initiation and the reporters contain constant five prime and three prime UTR sequences. 1395 of the 1600 designed sequences were ordered  in vitro synthetized mRNA was injected into 1 cell stage zebrafish embryos  and the mRNA decay post 2  5  and 8 hours post injection was calculated by targeted RNA sequencing.", null, "pubmed:35840631", null, "zf library 2h 2", "GSM6300332", null, "source name:Zebrafish embryo|tissue:Zebrafish embryo|time point:2h|genotype:WT|treatment:Injected zebrafish embryos", "zf library 2h 2", "Paired end reads were merged with FLASH using default parameters  reads that could not be merged were discarded. Next  using Cutadapt  the constant regions 5\u2019 = ATGGTGAGCAAGGGCGAGGAGCTGTCCCTCGAG and 3\u2019 = TCTAGATAG were removed to keep only the library variable region. Reads without xxx regions were discarded. Then  reads were separated by length read length = 297  expected length based on sequence design and mapped them to the library with salmon to quantify expression at each time point Mapped reads represent the perfect sequences that we designed. Then  the reads that were not mapped read length = 297  nucleotide substitutions were merged with the reads that have a different sequence length read length < 297 or > 297  insertions and/or deletions. All these reads were called the \u201cimperfect\u201d set. The constant regions that were removed in the first steps of the analysis were added back and therefore every unique read contains a coding sequence which starts in the first ATG start codon. The codon composition was counted as consecutive triplet until it finds a stop codon. The occurrences of each unique read were counted at every time point to calculate transcript abundance transcripts per million. Assembly: reference.fasta Supplementary files format and content: csv files include TPM values for each sample Library strategy: Targeted Deep Sequencing", "Zebrafish embryo", "10 pg of library mRNA were injected into 1 cell stage zebrafish embryos in 3 independent replicates. Each replicate utilized an independent needle  injection mix  and breeding. 25 embryos were collected at 2  5  and 8 hours post injection", "Injected zebrafish embryos were collected in Buffer RLT QIAGEN RNeasy kit  followed by vortexing for 1 min and stored at  70\u00b0C until extraction. RNA was thawed at 37\u00b0C  extracted following the manufacturer\u2019s protocol  and eluted in 50 \u03bcL. mRNA was isolated with Dynabeads oligo dT mRNA purification kit following the manufacturer\u2019s protocol. mRNA was eluted in 10 \u03bcl of 10 mM Tris HCl  pH 7.5. 5 \u03bcL of purified mRNA were used for cDNA synthesis using the SuperScript IV Reverse Transcriptase kit  utilizing a specific oligo that binds to the 3\u2019 Illumina adapter GTAGTGACTGGAGTTCAGAC. The cDNA was cleaned with the Qiagen MinElute PCR purification kit and eluted in 20 \u03bcL of water. We tested different PCR cycles to amplify the library using specific oligos that target the surrounding Illumina adapters. We selected 16 cycles for the samples at 2 and 5 hpi  and 18 cycles for 8 hpi. The barcoded libraries were cleaned using AMPure XP beads and resuspended in 20 \u03bcL of buffer EB Qiagen. The libraries were sequenced in an Illumina MiSeq instrument 250bp pair end with 10% of PhiX as spike in.\u00a0", null, "tissue:Zebrafish embryo|time point:2h|genotype:WT|treatment:Injected zebrafish embryos", "GSM6300332", "GSM6300332: zf library 2h 2; Danio rerio; OTHER", "GSM6300332 r1", "GSM6300332", "1", "Injected zebrafish embryos were collected in Buffer RLT QIAGEN RNeasy kit  followed by vortexing for 1 min and stored at  70\u00b0C until extraction. RNA was thawed at 37\u00b0C  extracted following the manufacturer's protocol  and eluted in 50 \u03bcL. mRNA was isolated with Dynabeads oligo dT mRNA purification kit following the manufacturer's protocol. mRNA was eluted in 10 \u03bcl of 10 mM Tris HCl  pH 7.5. 5 \u03bcL of purified mRNA were used for cDNA synthesis using the SuperScript IV Reverse Transcriptase kit  utilizing a specific oligo that binds to the three prime Illumina adapter GTAGTGACTGGAGTTCAGAC. The cDNA was cleaned with the Qiagen MinElute PCR purification kit and eluted in 20 \u03bcL of water. We tested different PCR cycles to amplify the library using specific oligos that target the surrounding Illumina adapters. We selected 16 cycles for the samples at 2 and 5 hpi  and 18 cycles for 8 hpi. The barcoded libraries were cleaned using AMPure XP beads and resuspended in 20 \u03bcL of buffer EB Qiagen. The libraries were sequenced in an Illumina MiSeq instrument 250bp pair end with 10% of PhiX as spike in.", null, "OTHER", "TRANSCRIPTOMIC", "other", "PAIRED", "ILLUMINA", "Illumina MiSeq", null, "SRP385134", null, null, "zf_library_2h_2_1.fastq.gz zf_library_2h_2_2.fastq.gz", "fastq fastq", 947154022.0, 1886761.0, "GSM6300332 r1", "0:251 1:251", "A:238912401;C:233180719;G:242163235;T:232651663;N:246004", 251, 251, null, null, 238912401, 233180719, 242163235, 232651663, 246004, "SRX16042026", "SRS13721694", "SRA1449828", "Stowers Institute for Medical Research", "Stowers Institute for Medical Research", 2, 1e-05, 0.0, 0.0, 0.0, 0.99997, 1.0, 1.0, null, 251, 251, "T", "T", "mates < 9% mapping rate", "illumina", "miseq", "unknown", "poly_a", "unknown", "bulk", "unknown", "unknown", null, "United States", "2022-07-06", "Undetermined", "Embryo", "Embryo Imprecise", "All anatomical structures"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["70589"], "units": {}, "query_ms": 9.132562001468614}