{"database": "metadata", "table": "run_metadata", "rows": [[34877, "SRR32303701", "SRX27639471", "SRS24046086", "SRP563065", "PRJNA1221811", "Targeted selection of experimental replicate groups is the likely driver of differential expression artifacts", "GSE289151", "Transcriptome Analysis", "The design of a scientific experiment forms the foundation for rigorous hypothesis testing and robust interpretations. In RNA seq analyses where common model organisms are used to study the effects of gene perturbations  breeding schemes are typically designed to produce siblings with either mutant or wildtype alleles of a gene. Siblings are expected to possess reduced genetic variability  allowing the biological responses of a gene perturbation to be observed accurately  and above that of minimal background noise. However  as previously reported Baer et al. 2024  inter  and intra genetic variability of the parental subjects allows for the possibility of such between the progeny. What matters is not so much the genetic diversity  but more specifically the differences in polymorphic loci represented between replicate sample groups  which are constructed to perform comparative statistical tests. This is largely due to chance when selecting a subset of progeny at random to ultimately represent the RNA sequencing samples. For many RNA sequencing experiments this subset is often small due to minimise costs  and therefore the opportunity of skewing is likely. Additionally  when the subset of samples is not selected at random  particularly when defined based on genotype at a genetic locus  further skewing may artificially and unexpectedly be introduced. The end result is expression differences due to unequal representations of genetic features that influence gene expression  i.e. expression quantitative trai loci eQTLs  which manifest as noise obscuring the desired signal and therefore confound interpretations. In this study we intentionally implemented an experimental design that opposes our advice from previous findings. We analyse the impacts of strain specific differences in zebrafish  by outcrossing two different strains of zebrafish. Strain specific differences were quantified by performing Differential Allelic Representation DAR analysis  consequently providing additional evidence that DAR contributes to false positives in differential expression analysis. We first assess the impact of DAR in isolation  where functional differences due to mutation are mitigated by equal representation within each group subject to comparison. We then extend our investigation by rearranging the same set of samples to form additional distinct experimental groups  allowing functional comparisons similar to our previous approach  but in the presence of elevated DAR. Overall design: mRNA sequencing of 3 mpf sibling zebrafish whole brains  consisting of 27 samples that were genotyped at two loci on chromosomes 14 and 17  with n=3 for each genotype combination. Genotypes at chromosome 14 are homozygous for Tubingen TU strain  homozygous for PK strain  and heterozygous TU/PK. Genotypes at chromosome 17 are heterozygous for the psen1 T428del allele modelling early onset familial Alzheimer's disease  heterozygous for the psen1 W233fs allele modelling familial Acne Inversa  and wild type.", null, null, null, "Whole brain  chr17 T428del/+  chr14 TU/PK  rep3", "GSM8785378", null, "source name:Whole brain|tissue:Whole brain|genotype:chr17 T428del/+; chr14 TU/PK|batch:1|geo loc name:missing|collection date:missing", "Whole brain  chr17 T428del/+  chr14 TU/PK  rep3", "Raw reads were trimmed with fastp v0.23.4 with parameters:   detect adapter for pe   qualified quality phred 20   length required 35   trim poly g Trimmed reads were aligned to GRCz11 Ensembl release 111 with STAR v2.7.11b with parameters:   sjdbOverhang 149   outSAMtype BAM SortedByCoordinate   twopassMode Basic Aligned reads were summarised to the gene level using Subread featureCounts v2.0.6 with parameters:  p   countReadPairs   fracOverlap 1   minOverlap 35 Assembly: GRCz11 Supplementary files format and content: Processed data contains gene level counts as output by featureCounts in TSV format.", "Whole brain", null, "RNA was extracted from whole brain tissue using Qiagen RNeasy columns. Stranded polyA libraries prepared according to the Nugen Universal Plus mRNA seq protocol and included 13 cycles of amplification.", null, "tissue:Whole brain|genotype:chr17 T428del/+; chr14 TU/PK|batch:1", "GSM8785378", "GSM8785378: Whole brain  chr17 T428del/+  chr14 TU/PK  rep3; Danio rerio; RNA Seq", "GSM8785378 r1", "GSM8785378", "1", "RNA was extracted from whole brain tissue using Qiagen RNeasy columns. Stranded polyA libraries prepared according to the Nugen Universal Plus mRNA seq protocol and included 13 cycles of amplification.", null, "RNA-Seq", "TRANSCRIPTOMIC", "cDNA", "PAIRED", "DNBSEQ", "DNBSEQ-G400", null, "SRP563065", null, null, "SAGCFN_23_00423_S27_R1_001.fastq.gz SAGCFN_23_00423_S27_R2_001.fastq.gz", "fastq fastq", 21983734400.0, 109918672.0, "GSM8785378 r1", "0:100 1:100", "A:6067597899;C:4889387435;G:4903291823;T:6123135261;N:321982", 100, 100, null, null, 6067597899, 4889387435, 4903291823, 6123135261, 321982, "SRX27639471", "SRS24046086", "SRA2076182", "The University of Adelaide", "The University of Adelaide", null, null, null, null, null, null, null, null, null, null, null, "B", "B", "biological fallback assumption", "bgi", "bgi", "unknown", "poly_a", "unknown", "bulk", "unknown", "unknown", null, "Australia", "2025-02-10", "Undetermined", "Adult", "Brain", "Nervous System"]], "columns": ["rowid", "run.accession", "experiment.accession", "sample.accession", "study.accession", "bioproject", "study.title", "study.alias", "study.type", "study.abstract", "study.attributes", "study.PMIDs", "sample.description", "sample.title", "sample.alias", "sample.centername", "sample.attributes", "GEOsample.title", "GEOsample.dataprocessing", "GEOsample.source", "GEOsample.treatmentprotocol", "GEOsample.extractprotocol", "GEOsample.growthprotocol", "GEOsample.characteristics", "GEOsample.accession", "experiment.title", "experiment.alias", "experiment.library_name", "experiment.design_description", "experiment.library_construction_protocol", "experiment.attributes", "experiment.library_strategy", "experiment.library_source", "experiment.library_selection", "experiment.library_layout", "experiment.platform", "experiment.instrument_model", "experiment.spot_descriptor", "experiment.study_ref", "run.title", "run.attributes", "run.filename", "run.semantic_name", "run.total_bases", "run.total_spots", "run.alias", "run.read_lengths", "run.base_counts", "run.r1_length", "run.r2_length", "run.r3_length", "run.r4_length", "run.Acount", "run.Ccount", "run.Gcount", "run.Tcount", "run.Ncount", "run.experiment", "run.pool_member", "submission.accession", "submission.srasource", "submission.bioprojectsource", "seqdetective.n_mates", "seqdetective.mapping_rate.mate1", "seqdetective.mapping_rate.mate2", "seqdetective.nofeature_rate.mate1", "seqdetective.nofeature_rate.mate2", "seqdetective.sparsity.mate1", "seqdetective.sparsity.mate2", "seqdetective.pos_strand_rate.mate1", "seqdetective.pos_strand_rate.mate2", "seqdetective.readlen.mate1", "seqdetective.readlen.mate2", "seqdetective.judgement.mate1", "seqdetective.judgement.mate2", "seqdetective.judgement.reason", "platform_family", "instrument_generation", "read_bias", "selection_class", "prep_kit", "sc_or_bulk", "tech_class", "technology", "tech_variant", "submission.bioprojectsource.country", "earliest_date", "devstage_curation", "devstage_curation_coarse", "tissue_curation", "tissue_curation_coarse"], "primary_keys": ["rowid"], "primary_key_values": ["34877"], "units": {}, "query_ms": 8.584996001445688}