{"database": "metadata", "table": "run_metadata", "rows": [[39612, "SRR1873568", "SRX915248", "SRS870226", "SRP056014", "PRJNA277780", "Identification of small non coding RNAs in zebrafish", "GSE66718", "Transcriptome Analysis", "MicroRNAs miRNAs are a new class of small RNAs of approximately 22 nucleotides in length that control eukaryotic gene expression by fine tuning mRNA translation. They regulate a wide variety of biological processes  namely developmental timing  cell differentiation  cell proliferation  immune response and infection. For this reason  their identification is essential to understand eukaryotic biology. Their small size  low abundance and high instability complicated early identification; however  cloning/Sanger sequencing and new generation genome sequencing approaches overcame most technical hurdles and are being used for rapid miRNA identification in many eukaryotes. We have applied 454 DNA pyrosequencing technology to miRNA discovery in zebrafish Danio rerio. For this  a series of cDNA libraries were prepared from small non coding RNAs isolated at different embryonic time points and from fully developed organs. Each cDNA library was tagged with specific sequences and was sequenced using the Roche FLX genome sequencer. This approach retrieved 90% of the 192 miRNAs previously identified by cloning/Sanger sequencing and bioinformatics and 25 novel miRNAs were predicted. Overall design: Small RNA libraries were prepared from different zebrafish developmental stages  namely  24 hpf  72 hpf  96 hpf  5 dpf dpf  45 dpf  young adult and from adult brain  eyes  gills  heart  skin and fins.", null, "pubmed:26694924", null, "gills", "GSM1630512", null, "source name:adult gills 1 year old|strain/background:AB|genotype/variation:wild type|tissue:gills|developmental stage:adult|age:1 year", "gills", "Base calling and quality trimming of sequence reads was carried out using the Genome Sequencer FLX software. Raw images were processed to remove background noise and the data was normalized. TAGs and adapter sequences of zebrafish developmental and adult tissues samples were then identified and trimmed and those reads with correct TAGs and adapters > 15 nt were retrieved for downstream analysis using the miRDeep software https://www.mdc berlin.de/8551903/en/. miRDeep aligned the sequences against the zebrafish genome using megaBlast with seed length set at 12  the traditional blast output  and minimum local identity set at 100. The blast output was then parsed for miRDeep uploading and aligned sequences with a maximum of 2 mismatches in the three prime end were retrieved. Reads that matched more than 10 different genome loci were discarded and only those with one or more alignments were kept  and using the remaining alignments as guidelines  the potential precursors were excised from the genome. The secondary structure of putative precursors was predicted using RNAfold and signatures were created by retaining reads that aligned perfectly with those putative precursors to generate the signature format. Finally  miRDeep predicted miRNAs by discarding non plausible Dicer products and scoring plausible ones. To assess seed conservation  plausible Dicer processing sequences were blasted against a local version of mature miRNAs from miRBase 12.0 that lacked zebrafish miRNA sequences. Borderline miRNA candidates were also resolved by determining their relative stability using Randfold. To distinguish between novel and known miRNAs  selected pre miRNAs were blasted against Danio rerio stem loop sequences miRBase and those that did not produce any or produced imperfect alignments were scored as novel miRNAs. Pairs of signatures and structures were used to estimate the number of false positives by randomly permutation  using miRDeep. To overcome the inherent lack of sensitivity of miRDeep  novel transcripts encoding miRNAs predicted by bioinformatics were retrieved from Ensembl 5.2 using BioMart and from literature predictions. These sequences were then used to perform a megaBlast search against our data with seed length set at 12. The transcripts with perfect matches and alignment length larger than 18 nt were kept for further processing. These transcripts were then compared with the mature miRNAs present in miRBase 12.0 and those that produced imperfect alignments or did not produce alignments were considered new miRNAs. Read numbers were normalized as described by Chen and colleagues Genes & Development 2005  19:1288 1293  and a miRNA expression profile  using identical number of reads for each sample  was generated. The number of reads between samples was normalized as indicated below:  Expression Reads = [1000 x NRmiRNAXY]/ TNRmiRNAsY  where NRmiRNAXY is the number of reads of miRNAX X = any miRNA in sample Y  and TNRmiRNAsY is the total number of miRNAs in sample Y. 1000 is an arbitrary number of reads. The data was transformed into log2 scale to build the heat map using the MeV 4.0 software package http://www.tm4.org/mev.html. Genome build: Zv8 Supplementary files format and content: Tab delimited text files include the miRNA IDs  sequences and respective raw counts post data processing for each sample.", "adult gills 1 year old", null, "100 \u03bcg of total RNA from each sample was isolated using TRIzol\u00ae and small RNAs were enriched by differential precipitation using polyethylene glycol. Total RNAs were fractioned using 12% denaturing PAGE and small RNAs of 15 30 nt were gel isolated using Gel Filtration cartridges from Edge Biosystems. For cDNA synthesis  the small RNA molecules previously isolated were first ligated to a three prime adapter AMP five primep five primep/CTGTAGGCACCATCAATdi deoxyC  three prime in absence of ATP and gel excised in the range of 35 and 50 nt. A second ligation was performed with the five prime adapter \"Nelson's linker\" five primeATCGTrArGrGrCrArCrCrUrGrArArA three prime  for 1 hour at 37\u00b0C  followed by phenol extraction. First strand cDNA synthesis was then performed using a specific three prime primer and Superscript\u2122 III reverse transcriptase Invitrogen. RNase H treated cDNA was PCR amplified with adapter specific primers. Each sample contained a specific TAG constituted by 3 nucleotides  as detailed next. 24hpf   ATC; 72hpf   ACT; 96hpf   CAG; 5dpf   ATG; 45dpf   CCG; entire adult   GTA; brain   CGG; Heart   CTG; Eyes   GCT; fins   GTT; skin   TAC; gills   TCC. PCR products were then run on 10% denaturing PAGE containing 7 M urea and the corresponding band 100 nt was eluted from the gel with Probe Elution Buffer from Ambion  at 37\u00b0C overnight. These products were used for the emulsion PCR. Parallel DNA pyrosequencing was performed using the Genome Sequencer FLX Roche  following established protocols for DNA library sequencing.", "Wild type AB zebrafish strain was maintained at 28\u00baC on a 14 h light/10 h dark cycle.", "strain/background:AB|genotype/variation:wild type|tissue:gills|developmental stage:adult|age:1 year", "GSM1630512", "GSM1630512: gills; Danio rerio; miRNA Seq", "GSM1630512", null, "1", "100 \u03bcg of total RNA from each sample was isolated using TRIzol\u00ae and small RNAs were enriched by differential precipitation using polyethylene glycol. Total RNAs were fractioned using 12% denaturing PAGE and small RNAs of 15 30 nt were gel isolated using Gel Filtration cartridges from Edge Biosystems. For cDNA synthesis  the small RNA molecules previously isolated were first ligated to a three prime adapter AMP five primep five primep/CTGTAGGCACCATCAATdi deoxyC  three prime in absence of ATP and gel excised in the range of 35 and 50 nt. A second ligation was performed with the five prime adapter \"Nelson's linker\" five primeATCGTrArGrGrCrArCrCrUrGrArArA three prime  for 1 hour at 37\u00b0C  followed by phenol extraction. First strand cDNA synthesis was then performed using a specific three prime primer and Superscript\u2122 III reverse transcriptase Invitrogen. RNase H treated cDNA was PCR amplified with adapter specific primers. Each sample contained a specific TAG constituted by 3 nucleotides  as detailed next. 24hpf   ATC; 72hpf   ACT; 96hpf   CAG; 5dpf   ATG; 45dpf   CCG; entire adult   GTA; brain   CGG; Heart   CTG; Eyes   GCT; fins   GTT; skin   TAC; gills   TCC. PCR products were then run on 10% denaturing PAGE containing 7 M urea and the corresponding band 100 nt was eluted from the gel with Probe Elution Buffer from Ambion  at 37\u00b0C overnight. These products were used for the emulsion PCR. Parallel DNA pyrosequencing was performed using the Genome Sequencer FLX Roche  following established protocols for DNA library sequencing.", "GEO Accession:GSM1630512", "miRNA-Seq", "TRANSCRIPTOMIC", "size fractionation", "SINGLE", "LS454", "454 GS FLX", null, "SRP056014", null, null, null, null, 189248.0, 3563.0, "GSM1630512 r1", "0:4 1:49.11", "A:48212;C:55198;G:44765;T:40711;N:362", 4, 49, null, null, 48212, 55198, 44765, 40711, 362, "SRX915248", "SRS870226", "SRA246117", "GEO", "University of Aveiro", 1, 0.0, null, 0.0, null, 1.0, null, null, null, 44, null, "T", null, "under 1.2% mapping rate", "legacy", "early", "3prime", "size_fractionation", "unknown", "bulk", "other_seq", "454", null, "Portugal", "2015-03-09", "Adult", "Adult", "Gill", "Respiratory System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["39612"], "units": {}, "query_ms": 8.872156002325937}