diff --git a/.spelling b/.spelling index 2ee7acc0..d75eb443 100644 --- a/.spelling +++ b/.spelling @@ -597,6 +597,7 @@ deconvoluting DMG-H3 gemcitabine Hematopoiesis +iDAT Illumina in-vitro in-vivo diff --git a/content/3.genomics-platform/1.getting-started/3.making-a-data-request.md b/content/3.genomics-platform/1.getting-started/3.making-a-data-request.md index 40b17b09..a674c365 100644 --- a/content/3.genomics-platform/1.getting-started/3.making-a-data-request.md +++ b/content/3.genomics-platform/1.getting-started/3.making-a-data-request.md @@ -124,3 +124,15 @@ If you receive an email from us that your DAA is incomplete, you may edit your D ## Managing your Data Request Go to our [Managing Data Overview](/genomics-platform/managing-data/overview) documentation page to learn how to check the status of your data request, complete an EDAA draft, upload a revised DAA, and ultimately access your data from your [My Dashboard](https://platform.stjude.cloud/requests/manage) page. + +## Unrestricted Data + +Certain Data within Genomics Platform is unrestricted, meaning that access is available to all requestors and does not require a data access agreement. + +To access this data, please complete the following steps: + +1. Create an account on or log in to Genomics Platform. +2. Narrow your selection by filtering to Feature Count Files only and/or iDAT files only and then selecting Request Data at the bottom right of the screen. +3. Choose to vend the data to a new or existing project +4. Submit the request. +5. The data will be vended to your selected project in a folder labeled with the date the data was requested. diff --git a/content/3.genomics-platform/2.about-our-data/1.data-sets-and-data-access-units.md b/content/3.genomics-platform/2.about-our-data/1.data-sets-and-data-access-units.md index 21713fa5..d6cd40ac 100644 --- a/content/3.genomics-platform/2.about-our-data/1.data-sets-and-data-access-units.md +++ b/content/3.genomics-platform/2.about-our-data/1.data-sets-and-data-access-units.md @@ -9,6 +9,7 @@ title: Data Sets and Data Access Units - [Data Set](#data-set) - [Data Access Committee (DAC)](#data-access-committee-dac) - [Embargo Date](#embargo-date) + - [Unrestricted Data](#unrestricted-data) - [List of DAUs](#list-of-daus) - [List of Data Sets](#list-of-data-sets) @@ -62,6 +63,16 @@ Publishing using any of the files _before_ the embargo date has passed is strict Some Data, including Data funded by the NIH, are not subject to embargo. Applicable Embargo Dates can be found in [Genomics Platform Metadata](https://platform.stjude.cloud/api/v1/manifest.tsv){target="_blank"} in the `SJ_Embargo_Date` column. +### Unrestricted Data + +Certain data within the Genomics Platform is unrestricted. Unrestricted data is not subject to DAC-reviewed approval before a user can obtain it. +The unrestricted dataset on St. Jude Cloud currently includes: + +- Feature count files +- [COMET](https://comet.stjude.org/) iDAT files + +Steps to access unrestricted data can be found [here](http://docs.stjude.cloud/genomics-platform/getting-started/making-a-data-request#unrestricted-data). + --- ## List of DAUs @@ -158,7 +169,7 @@ The following data set(s) are included within SJLIFE: ## List of Data Sets -We currently have 21 [Data Sets](#data-set) listed below. +We currently have 22 [Data Sets](#data-set) listed below. Additional information can also be seen including which [Data Access Units (DAU)](#data-access-unit-dau) the Data Set belongs to, tissue type, sequencing type, number of samples, additional links, and a brief description. | Data Set | DAU(s) | Tissue Type | Sequencing | Samples | @@ -167,6 +178,7 @@ Additional information can also be seen including which [Data Access Units (DAU) | [CCSS](#childhood-cancer-survivor-study) | CCSS | Germline Only | WGS | 2,912 | | [CICERO Benchmark](#cicero-benchmark) | PCGP, Clinical Genomics | Paired Tumor-Normal | RNA-Seq | 124 | | [Clinical Pilot](#clinical-pilot) | PCGP, Clinical Genomics | Paired Tumor-Normal | WGS, WES, RNA-Seq | 155 | +| [COMET](#comet) | Unrestricted | iDAT file | — | 4269 | | [CReATe](#clinical-research-in-als-and-related-disorders-for-therapeutic-development-consortium) | CReATe | PBMC Germline DNA | WGS | 705 | | [CSTN](#childhood-solid-tumor-network) | PCGP, Clinical Genomics | Paired Tumor-Normal | WGS, WES, RNA-Seq | 143 | | [G4K](#genome-4-kids) | PCGP, Clinical Genomics | Paired Tumor-Normal | WGS, WES, RNA-Seq | 565 | @@ -183,7 +195,7 @@ Additional information can also be seen including which [Data Access Units (DAU) | [RTCG](#real-time-clinical-genomics) | PCGP, Clinical Genomics | Paired Tumor-Normal | WGS, WES, RNA-Seq | 7,767 | | [SGP](#sickle-cell-genome-project) | SGP | Germline Only | WGS | 807 | | [SJLIFE](#st-jude-life) | SJLIFE | Germline Only | WGS, WES | 4,838 | -| [SJLIFE_ClonalHematopoiesis](#st-jude-life-clonal-hematopoiesis) | SJLIFE | — | Single Cell-WGS, Targeted | 3,192 | +| [SJLIFE_ClonalHematopoiesis](#st-jude-life-clonal-hematopoiesis) | PCGP | — | Single Cell-WGS, Targeted | 3,192 | | [tMN](#pediatric-therapy-related-myeloid-neoplasms-tmn) | PCGP | Paired Tumor-Normal | WGS, WES, RNA-Seq | 206 | ### Atypical Teratoid / Rhabdoid Tumor-derived Tumoroid Models @@ -254,6 +266,14 @@ In addition to patients enrolled in the PGB1 Cohort (primary participants), the This dataset includes WGS data from N=705 in PGB1, including N=472 ALS/ALS-FTD, N=20 PMA, N=47 PLS, N=162 HSP, and N=4 with other related disorders. The findings of the project were published in [Translational Neurodegeneration](https://translationalneurodegeneration.biomedcentral.com/articles/10.1186/s40035-025-00516-2). +### COMET + +**DAU**: - | **Tissue Type**: - | **Assay Type**: Illumina Infinium 850K array | **Samples**: 4,629| **[Additional Information About COMET](https://www.stjude.org/research/departments/computational-biology/comet.html)** + +The solid tumor COmprehensive METhylation (COMET) database is a searchable repository of pediatric solid tumor DNA methylation and copy number variant (CNV) profiles, generated using the Illumina Infinium 850K array, paired with matched whole slide histology images (WSI). +It is the largest and most comprehensive extracranial pediatric solid tumor epigenetic reference dataset in the world, offering DNA methylation profiles across 20 different types of pediatric solid tumors along with a comparative collection of patient-derived orthotopic xenografts, cell lines, adult sarcomas, and normal tissues. +See [Unrestricted Data](#unrestricted-data) for more details on requesting access to this data set. + ### DMG-H3K27a Clonal Evolution **DAU**: PCGP | **Tissue Type**: — | **Sequencing Type**: WGS, WES | **Samples**: 70 diff --git a/content/3.genomics-platform/2.about-our-data/2.file-formats-and-sequencing-information.md b/content/3.genomics-platform/2.about-our-data/2.file-formats-and-sequencing-information.md index 6f16b861..d83cff3d 100644 --- a/content/3.genomics-platform/2.about-our-data/2.file-formats-and-sequencing-information.md +++ b/content/3.genomics-platform/2.about-our-data/2.file-formats-and-sequencing-information.md @@ -13,6 +13,7 @@ St. Jude Cloud hosts both raw genomic data files and processed results files: | Somatic VCF | Curated list of somatic variants produced by the St. Jude somatic variant analysis pipeline. | [Click here](#somatic-vcf-files) | | CNV | List of somatic copy number alterations produced by St. Jude CONSERTING pipeline. | [Click here](#cnv-files) | | Feature Counts | Curated list of read counts mapped to each gene produced by [HTSeq](https://htseq.readthedocs.io/en/master/) | [Click here](#feature-counts-files) | +| iDAT | Raw, paired intensity files output by an Illumina microarray scanner for a single sample — one per fluorescence channel — before normalization or genotype/methylation calling. | [Click here](#idat-files) | ### BAM files @@ -195,6 +196,11 @@ The files are tab-delimited text and contain the feature key and read count for [rnaseq-rfc]: https://stjudecloud.github.io/rfcs/0001-rnaseq-workflow-v2.0.0.html#specification [gencode]: https://www.gencodegenes.org/human/release_31.html +### iDAT files + +These are the raw, paired intensity files output by an Illumina microarray for a single sample. +Each sample includes two IDAT files — one per fluorescence channel (Green and Red) containing the raw, unprocessed probe intensity signal from the array before any normalization or genotype/methylation calling. + ## Sequencing Information ### Whole Genome and Whole Exome diff --git a/content/4.pecan/1.overview/1.getting-started.md b/content/4.pecan/1.overview/1.getting-started.md index def358a5..879a1d6d 100644 --- a/content/4.pecan/1.overview/1.getting-started.md +++ b/content/4.pecan/1.overview/1.getting-started.md @@ -98,6 +98,16 @@ Data Facets represent a distinct type of post-processed genomic data for collect +
Methylation landscape of over 4,400 pediatric cancer samples in PeCan.
+