run_metadata
13 rows where experiment.library_layout = "SINGLE", experiment.library_strategy = "OTHER" and tissue_curation_coarse = "Nervous System"
This data as json, CSV (advanced)
| Link | rowid ▼ | run.accession | experiment.accession | sample.accession | study.accession | bioproject | study.title | study.alias | study.type | study.abstract | study.attributes | study.PMIDs | sample.description | sample.title | sample.alias | sample.centername | sample.attributes | GEOsample.title | GEOsample.dataprocessing | GEOsample.source | GEOsample.treatmentprotocol | GEOsample.extractprotocol | GEOsample.growthprotocol | GEOsample.characteristics | GEOsample.accession | experiment.title | experiment.alias | experiment.library_name | experiment.design_description | experiment.library_construction_protocol | experiment.attributes | experiment.library_strategy | experiment.library_source | experiment.library_selection | experiment.library_layout | experiment.platform | experiment.instrument_model | experiment.spot_descriptor | experiment.study_ref | run.title | run.attributes | run.filename | run.semantic_name | run.total_bases | run.total_spots | run.alias | run.read_lengths | run.base_counts | run.r1_length | run.r2_length | run.r3_length | run.r4_length | run.Acount | run.Ccount | run.Gcount | run.Tcount | run.Ncount | run.experiment | run.pool_member | submission.accession | submission.srasource | submission.bioprojectsource | seqdetective.n_mates | seqdetective.mapping_rate.mate1 | seqdetective.mapping_rate.mate2 | seqdetective.nofeature_rate.mate1 | seqdetective.nofeature_rate.mate2 | seqdetective.sparsity.mate1 | seqdetective.sparsity.mate2 | seqdetective.pos_strand_rate.mate1 | seqdetective.pos_strand_rate.mate2 | seqdetective.readlen.mate1 | seqdetective.readlen.mate2 | seqdetective.judgement.mate1 | seqdetective.judgement.mate2 | seqdetective.judgement.reason | platform_family | instrument_generation | read_bias | selection_class | prep_kit | sc_or_bulk | tech_class | technology | tech_variant | submission.bioprojectsource.country | earliest_date | devstage_curation | devstage_curation_coarse | tissue_curation | tissue_curation_coarse |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 43842 | 43842 | SRR6176750 | SRX3287416 | SRS2596889 | SRP120009 | PRJNA414416 | Simultaneous single cell profiling of lineages and cell types in the vertebrate brain | GSE105010 | Other | The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries | pubmed:29608178 | ZF3 scGSTLT | GSM2813986 | source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf | ZF3 scGSTLT | Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored. | zebrafish brain | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | tissue:brain|developmental stage:23 25dpf | GSM2813986 | GSM2813986: ZF3 scGSTLT; Danio rerio; OTHER | GSM2813986 | 1 | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | GEO Accession:GSM2813986 | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina MiSeq | SRP120009 | loader:fastq load.py|options: appendBCtoName | F6_UMI.merged.fq.gz | fastq | 786192951.0 | 2901081.0 | GSM2813986 r1 | 0:271 | A:215364033;C:176865797;G:224444690;T:169518396;N:35 | 271 | 215364033 | 176865797 | 224444690 | 169518396 | 35 | SRX3287416 | SRS2596889 | SRA619743 | GEO | Harvard University | 1 | 0.00904 | 0.0 | 0.99997 | 0.0 | 271 | T | under 1.2% mapping rate | illumina | miseq | unknown | other | unknown | sc | single_cell_droplet | indrops | United States | 2017-10-16 | Larval | Larval | Brain | Nervous System | ||||||||||||||||||
| 43843 | 43843 | SRR6176749 | SRX3287415 | SRS2596888 | SRP120009 | PRJNA414416 | Simultaneous single cell profiling of lineages and cell types in the vertebrate brain | GSE105010 | Other | The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries | pubmed:29608178 | ZF2 scGSTLT | GSM2813985 | source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf | ZF2 scGSTLT | Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored. | zebrafish brain | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | tissue:brain|developmental stage:23 25dpf | GSM2813985 | GSM2813985: ZF2 scGSTLT; Danio rerio; OTHER | GSM2813985 | 1 | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | GEO Accession:GSM2813985 | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina MiSeq | SRP120009 | loader:fastq load.py|options: appendBCtoName | F5_UMI.merged.fq.gz | fastq | 106145626.0 | 391949.0 | GSM2813985 r1 | 0:270.81 | A:29282656;C:22944139;G:29276360;T:24642464;N:7 | 270 | 29282656 | 22944139 | 29276360 | 24642464 | 7 | SRX3287415 | SRS2596888 | SRA619743 | GEO | Harvard University | 1 | 2e-05 | 0.0 | 0.99995 | 0.0 | 270 | T | under 1.2% mapping rate | illumina | miseq | unknown | other | unknown | sc | single_cell_droplet | indrops | United States | 2017-10-16 | Larval | Larval | Brain | Nervous System | ||||||||||||||||||
| 43844 | 43844 | SRR6176748 | SRX3287414 | SRS2596887 | SRP120009 | PRJNA414416 | Simultaneous single cell profiling of lineages and cell types in the vertebrate brain | GSE105010 | Other | The lineage relationships among the hundreds of cell types generated during development are difficult to reconstruct. A recent method GESTALT used CRISPR–Cas9 barcode editing for large scale lineage tracing but was restricted to early development and did not identify cell types. Here we present scGESTALT which combines the lineage recording capabilities of GESTALT with cell type identification by single cell RNA sequencing. The method relies on an inducible system that enables barcodes to be edited at multiple time points capturing lineage information from later stages of development. Sequencing of 60 000 transcriptomes from the juvenile zebrafish brain identified >100 cell types and marker genes. Using these data we generate lineage trees with hundreds of branches that help uncover restrictions at the level of cell types brain regions and gene expression cascades during differentiation. scGESTALT can be applied to other multicellular organisms to simultaneously characterize molecular identities and lineage histories of thousands of cells during development and disease. Overall design: inDrops libraries of single cell transcriptomes and scGESTALT barcodes and genomic DNA GESTALT libraries | pubmed:29608178 | ZF1 scGSTLT | GSM2813984 | source name:zebrafish brain|tissue:brain|developmental stage:23 25dpf | ZF1 scGSTLT | Single cell RNA Sequencing data FASTQ files were processed using the inDrops.py bioinformatics pipeline available at https://github.com/indrops/indrops. Transcriptome libraries were mapped to a zebrafish reference built from a custom GTF file and the zebrafish GRCz10 release 86 genome assembly. Bowtie version1.1.1 was used with parameter –e 200; UMI quantification was used with parameter –u 2 counts were ignored from UMIs split between more than 2 genes. genomic DNA GESTALT and scGESTALT libraries were processed using a custom pipeline available at https://github.com/shendurelab/Cas9FateMapping Genome build: GRCz10 CSV files for transcriptome data were generated using the inDrops pipeline. Each column in the CSV files contains a cell identifier and each row contains expression values for genes. Txt files for genomic DNA GESTALT libraries *allReadCounts contain lineage barcode sequences HMID column for each cell and their proportion in the sequenced libraries. scGESTALT data *GestMaster.txt contains the inDrops cell identifiers CellBarcode and BarcodeKey that were used to match barcodes to transcriptomes. They also contain lineage barcode sequences HMID column for each cell with a corresponding inDrops single cell gene expression profile as well as the the t SNE cluster membership number ClusterIdent column Txt files ending in *stats.txt contain information about the each individually captured UMI or cell per sample. The barcode sequence aligned to a reference unedited sequence mergedRead column mutations at each target site target[X] columns and the edited sequences at each target site sequence[X] columns were used for downstream analysis. inDropsExpMatrix noQ txt file is the gene expression matrix for the full dataset. Columns are individual cells from different batches of whole brains f1 f2 f3 f4 f5 f6 or brain regions fore mid hind . fall.inDrops.Robj is the processed Seurat R object which can be loaded into R and explored. | zebrafish brain | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | tissue:brain|developmental stage:23 25dpf | GSM2813984 | GSM2813984: ZF1 scGSTLT; Danio rerio; OTHER | GSM2813984 | 1 | Single cell suspensions were processed through inDrops to generate single cell cDNA libraries. cDNAs were in vitro transcribed reverse transcribed and prepared for sequencing Gestalt barcode was PCR amplified by two step PCR. Sample indices and flow cell adaptors were then added by PCR. | GEO Accession:GSM2813984 | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina MiSeq | SRP120009 | loader:fastq load.py|options: appendBCtoName | F3_UMI.merged.fq.gz | fastq | 817055786.0 | 3014966.0 | GSM2813984 r1 | 0:271 | A:235846657;C:170311176;G:222336783;T:188561170;N:0 | 271 | 235846657 | 170311176 | 222336783 | 188561170 | 0 | SRX3287414 | SRS2596887 | SRA619743 | GEO | Harvard University | 1 | 4e-05 | 0.0 | 0.99995 | 0.4 | 271 | T | under 1.2% mapping rate | illumina | miseq | unknown | other | unknown | sc | single_cell_droplet | indrops | United States | 2017-10-16 | Larval | Larval | Brain | Nervous System | ||||||||||||||||||
| 70382 | 70382 | SRR19762944 | SRX15807635 | SRS13499606 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | ZF Ko H5 EW20 | GSM6256878 | source name:cpsf6 / 6dpf head|tissue:head|genotype:cpsf6 / |treatment:N1 | ZF Ko H5 EW20 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | cpsf6 / 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:cpsf6 / |treatment:N1 | GSM6256878 | GSM6256878: ZF Ko H5 EW20; Danio rerio; OTHER | GSM6256878 r1 | GSM6256878 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | ZF_Ko_H5_EW20_R1.fastq.gz | fastq | 3464837420.0 | 28635020.0 | GSM6256878 r1 | 0:121 | A:1494034798;C:622079032;G:645639126;T:703047177;N:37287 | 121 | 1494034798 | 622079032 | 645639126 | 703047177 | 37287 | SRX15807635 | SRS13499606 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.47903 | 0.15095 | 0.82883 | 0.6863 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70383 | 70383 | SRR19762945 | SRX15807634 | SRS13499605 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | ZF Ko H4 EW19 | GSM6256877 | source name:cpsf6 / 6dpf head|tissue:head|genotype:cpsf6 / |treatment:N1 | ZF Ko H4 EW19 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | cpsf6 / 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:cpsf6 / |treatment:N1 | GSM6256877 | GSM6256877: ZF Ko H4 EW19; Danio rerio; OTHER | GSM6256877 r1 | GSM6256877 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | ZF_Ko_H4_EW19_R1.fastq.gz | fastq | 3251893150.0 | 26875150.0 | GSM6256877 r1 | 0:121 | A:1344795862;C:598411565;G:620697978;T:687947810;N:39935 | 121 | 1344795862 | 598411565 | 620697978 | 687947810 | 39935 | SRX15807634 | SRS13499605 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.50383 | 0.16013 | 0.82217 | 0.6919 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70384 | 70384 | SRR19762946 | SRX15807633 | SRS13499604 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | ZF Ko H3 EW18 | GSM6256876 | source name:cpsf6 / 6dpf head|tissue:head|genotype:cpsf6 / |treatment:N1 | ZF Ko H3 EW18 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | cpsf6 / 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:cpsf6 / |treatment:N1 | GSM6256876 | GSM6256876: ZF Ko H3 EW18; Danio rerio; OTHER | GSM6256876 r1 | GSM6256876 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | ZF_Ko_H3_EW18_R1.fastq.gz | fastq | 2614428850.0 | 21606850.0 | GSM6256876 r1 | 0:121 | A:1106782184;C:471490720;G:482307751;T:553818128;N:30067 | 121 | 1106782184 | 471490720 | 482307751 | 553818128 | 30067 | SRX15807633 | SRS13499604 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.49216 | 0.17707 | 0.824 | 0.67519 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70385 | 70385 | SRR19762947 | SRX15807632 | SRS13499603 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | ZF Ko H2 EW17 | GSM6256875 | source name:cpsf6 / 6dpf head|tissue:head|genotype:cpsf6 / |treatment:N1 | ZF Ko H2 EW17 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | cpsf6 / 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:cpsf6 / |treatment:N1 | GSM6256875 | GSM6256875: ZF Ko H2 EW17; Danio rerio; OTHER | GSM6256875 r1 | GSM6256875 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | ZF_Ko_H2_EW17_R1.fastq.gz | fastq | 2829766984.0 | 23386504.0 | GSM6256875 r1 | 0:121 | A:1154124199;C:523756445;G:540103572;T:611750443;N:32325 | 121 | 1154124199 | 523756445 | 540103572 | 611750443 | 32325 | SRX15807632 | SRS13499603 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.53128 | 0.17359 | 0.81479 | 0.67717 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70386 | 70386 | SRR19762948 | SRX15807631 | SRS13499602 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | ZF Ko H1 EW16 | GSM6256874 | source name:cpsf6 / 6dpf head|tissue:head|genotype:cpsf6 / |treatment:N1 | ZF Ko H1 EW16 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | cpsf6 / 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:cpsf6 / |treatment:N1 | GSM6256874 | GSM6256874: ZF Ko H1 EW16; Danio rerio; OTHER | GSM6256874 r1 | GSM6256874 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | ZF_Ko_H1_EW16_R1.fastq.gz | fastq | 3127010139.0 | 25843059.0 | GSM6256874 r1 | 0:121 | A:1303638776;C:565310993;G:598748714;T:659275918;N:35738 | 121 | 1303638776 | 565310993 | 598748714 | 659275918 | 35738 | SRX15807631 | SRS13499602 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.49184 | 0.16601 | 0.82201 | 0.67657 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70387 | 70387 | SRR19762949 | SRX15807630 | SRS13499601 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | Zf Wt H5 EW15 | GSM6256873 | source name:wild type 6dpf head|tissue:head|genotype:wild type|treatment:N1 | Zf Wt H5 EW15 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | wild type 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:wild type|treatment:N1 | GSM6256873 | GSM6256873: Zf Wt H5 EW15; Danio rerio; OTHER | GSM6256873 r1 | GSM6256873 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | Zf_Wt_H5_EW15_R1.fastq.gz | fastq | 3020456208.0 | 24962448.0 | GSM6256873 r1 | 0:121 | A:1299290531;C:541614248;G:555748976;T:623766444;N:36009 | 121 | 1299290531 | 541614248 | 555748976 | 623766444 | 36009 | SRX15807630 | SRS13499601 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.50369 | 0.18129 | 0.83769 | 0.67627 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70388 | 70388 | SRR19762950 | SRX15807629 | SRS13499600 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | Zf Wt H4 EW14 | GSM6256872 | source name:wild type 6dpf head|tissue:head|genotype:wild type|treatment:N1 | Zf Wt H4 EW14 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | wild type 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:wild type|treatment:N1 | GSM6256872 | GSM6256872: Zf Wt H4 EW14; Danio rerio; OTHER | GSM6256872 r1 | GSM6256872 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | Zf_Wt_H4_EW14_R1.fastq.gz | fastq | 2964716348.0 | 24501788.0 | GSM6256872 r1 | 0:121 | A:1296624390;C:527132924;G:548516138;T:592409475;N:33421 | 121 | 1296624390 | 527132924 | 548516138 | 592409475 | 33421 | SRX15807629 | SRS13499600 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.46081 | 0.15564 | 0.85048 | 0.6808 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70389 | 70389 | SRR19762951 | SRX15807628 | SRS13499598 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | Zf Wt H3 EW13 | GSM6256871 | source name:wild type 6dpf head|tissue:head|genotype:wild type|treatment:N1 | Zf Wt H3 EW13 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | wild type 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:wild type|treatment:N1 | GSM6256871 | GSM6256871: Zf Wt H3 EW13; Danio rerio; OTHER | GSM6256871 r1 | GSM6256871 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | Zf_Wt_H3_EW13_R1.fastq.gz | fastq | 3442300565.0 | 28448765.0 | GSM6256871 r1 | 0:121 | A:1485526892;C:608224693;G:650560614;T:697949308;N:39058 | 121 | 1485526892 | 608224693 | 650560614 | 697949308 | 39058 | SRX15807628 | SRS13499598 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.47353 | 0.1692 | 0.84898 | 0.67905 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70390 | 70390 | SRR19762952 | SRX15807627 | SRS13499599 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | Zf Wt H2 EW12 | GSM6256870 | source name:wild type 6dpf head|tissue:head|genotype:wild type|treatment:N1 | Zf Wt H2 EW12 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | wild type 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:wild type|treatment:N1 | GSM6256870 | GSM6256870: Zf Wt H2 EW12; Danio rerio; OTHER | GSM6256870 r1 | GSM6256870 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | Zf_Wt_H2_EW12_R1.fastq.gz | fastq | 2948425634.0 | 24367154.0 | GSM6256870 r1 | 0:121 | A:1211451419;C:546445066;G:560022243;T:630468316;N:38590 | 121 | 1211451419 | 546445066 | 560022243 | 630468316 | 38590 | SRX15807627 | SRS13499599 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.53462 | 0.18347 | 0.82309 | 0.66989 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System | |||||||||||||||||
| 70391 | 70391 | SRR19762953 | SRX15807626 | SRS13499597 | SRP382883 | PRJNA851381 | Loss of CPSF6 causes developmental disease via bimodal changes in polyadenylation site usage and protein expression [Zebrafish] | GSE206558 | Other | Most pre messenger RNA pre mRNA undergo extensive processing to create distinct transcripts from the same gene. One of these processes alternative polyadenylation involves over twenty proteins to bind and cleave the pre mRNA at polyA sites that can lie within the three prime UTR introns or exons; this can modulate protein function but the effect of choosing a site internal to the gene vs. within the three prime UTR remains unclear. Here we show that reduced expression of CPSF6 one of the proteins involved in site selection derails development in both humans and zebrafish by causing a bidirectional shift in polyA site usage. CPSF6 insufficiency favors the use of intronic polyA sites in neuronal genes reducing mRNA and protein abundance but promotes three prime UTR site usage in cardiovascular and skeletal genes upregulating mRNA and protein.These data thus provides a long sought link between APA and gene expression and shows that polyA site selection influences development. Overall design: Comparative analysis of alternative polyadenylation using polyA click seq PAC seq on whole larva and head of cpsf6 / Danio rerio compared to stage matched wt controls. | parent bioproject:PRJNA851372 | pubmed:36800428 | Zf Wt H1 EW11 | GSM6256869 | source name:wild type 6dpf head|tissue:head|genotype:wild type|treatment:N1 | Zf Wt H1 EW11 | Raw reads were trimed using fastp Trimed reads were aligned to the reference genome using bowtie2 PCR duplicates were removed using umi tools Samtools were used to sort convert and index alignment files Deeptools were used to generate bigwig files Assembly: GRCz11 Supplementary files format and content: bigwig files Library strategy: PAC seq | wild type 6dpf head | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | tissue:head|genotype:wild type|treatment:N1 | GSM6256869 | GSM6256869: Zf Wt H1 EW11; Danio rerio; OTHER | GSM6256869 r1 | GSM6256869 | 1 | RNA was harvested using Rneasy mini kit Qiagen. 2 ug of total RNA was used for the construction of sequencing libraries. We reverse transcribed 1 ug of total RNA with the partial P7 adapter Illumina 4N 21T and dNTPs with the addition of spiked in azido nucleotides AzVTPs at 5:1. We click ligated the p5 adapter IDT to the 5′ end of the cDNA with CuAAC. The p5 adaptor contained a UMI. We then amplified the cDNA for 17 cycles with five prime and 3′ indexing primer and purified it on a 2% agarose gel by extracting amplicon from 200 300 base pairs. We pooled the libraries and sequenced single end 100 base pair reads on a NovaSeq Illumina. | OTHER | TRANSCRIPTOMIC | other | SINGLE | ILLUMINA | Illumina HiSeq 4000 | SRP382883 | loader:fastq load.py | Zf_Wt_H1_EW11_R1.fastq.gz | fastq | 2689461797.0 | 22226957.0 | GSM6256869 r1 | 0:121 | A:1082451613;C:513056920;G:525960217;T:567960717;N:32330 | 121 | 1082451613 | 513056920 | 525960217 | 567960717 | 32330 | SRX15807626 | SRS13499597 | SRA1440631 | Baylor College of Medicine | Baylor College of Medicine | 1 | 0.55368 | 0.16655 | 0.82828 | 0.69871 | 121 | B | usable mapping rate | illumina | hiseq_era | unknown | poly_a | unknown | bulk | unknown | unknown | United States | 2022-06-21 | Larval | Larval | Head | Nervous System |
Advanced export
JSON shape: default, array, newline-delimited
CREATE TABLE run_metadata("run.accession" VARCHAR, "experiment.accession" VARCHAR, "sample.accession" VARCHAR, "study.accession" VARCHAR, bioproject VARCHAR, "study.title" VARCHAR, "study.alias" VARCHAR, "study.type" VARCHAR, "study.abstract" VARCHAR, "study.attributes" VARCHAR, "study.PMIDs" VARCHAR, "sample.description" VARCHAR, "sample.title" VARCHAR, "sample.alias" VARCHAR, "sample.centername" VARCHAR, "sample.attributes" VARCHAR, "GEOsample.title" VARCHAR, "GEOsample.dataprocessing" VARCHAR, "GEOsample.source" VARCHAR, "GEOsample.treatmentprotocol" VARCHAR, "GEOsample.extractprotocol" VARCHAR, "GEOsample.growthprotocol" VARCHAR, "GEOsample.characteristics" VARCHAR, "GEOsample.accession" VARCHAR, "experiment.title" VARCHAR, "experiment.alias" VARCHAR, "experiment.library_name" VARCHAR, "experiment.design_description" VARCHAR, "experiment.library_construction_protocol" VARCHAR, "experiment.attributes" VARCHAR, "experiment.library_strategy" VARCHAR, "experiment.library_source" VARCHAR, "experiment.library_selection" VARCHAR, "experiment.library_layout" VARCHAR, "experiment.platform" VARCHAR, "experiment.instrument_model" VARCHAR, "experiment.spot_descriptor" VARCHAR, "experiment.study_ref" VARCHAR, "run.title" VARCHAR, "run.attributes" VARCHAR, "run.filename" VARCHAR, "run.semantic_name" VARCHAR, "run.total_bases" DOUBLE, "run.total_spots" DOUBLE, "run.alias" VARCHAR, "run.read_lengths" VARCHAR, "run.base_counts" VARCHAR, "run.r1_length" BIGINT, "run.r2_length" BIGINT, "run.r3_length" BIGINT, "run.r4_length" BIGINT, "run.Acount" BIGINT, "run.Ccount" BIGINT, "run.Gcount" BIGINT, "run.Tcount" BIGINT, "run.Ncount" BIGINT, "run.experiment" VARCHAR, "run.pool_member" VARCHAR, "submission.accession" VARCHAR, "submission.srasource" VARCHAR, "submission.bioprojectsource" VARCHAR, "seqdetective.n_mates" BIGINT, "seqdetective.mapping_rate.mate1" DOUBLE, "seqdetective.mapping_rate.mate2" DOUBLE, "seqdetective.nofeature_rate.mate1" DOUBLE, "seqdetective.nofeature_rate.mate2" DOUBLE, "seqdetective.sparsity.mate1" DOUBLE, "seqdetective.sparsity.mate2" DOUBLE, "seqdetective.pos_strand_rate.mate1" DOUBLE, "seqdetective.pos_strand_rate.mate2" DOUBLE, "seqdetective.readlen.mate1" BIGINT, "seqdetective.readlen.mate2" BIGINT, "seqdetective.judgement.mate1" VARCHAR, "seqdetective.judgement.mate2" VARCHAR, "seqdetective.judgement.reason" VARCHAR, platform_family VARCHAR, instrument_generation VARCHAR, read_bias VARCHAR, selection_class VARCHAR, prep_kit VARCHAR, sc_or_bulk VARCHAR, tech_class VARCHAR, technology VARCHAR, tech_variant VARCHAR, "submission.bioprojectsource.country" VARCHAR, earliest_date DATE, devstage_curation VARCHAR, devstage_curation_coarse VARCHAR, tissue_curation VARCHAR, tissue_curation_coarse VARCHAR);;
CREATE INDEX idx_run_bioproject ON run_metadata(bioproject);;
CREATE INDEX idx_run_run_accession ON run_metadata("run.accession");;