Example datasets¶
The combined LAML oncoplot uses load_dataset("tcga_laml_combined_oncoplot");
the p53 sequence comparison loads aligned FASTA text with
load_dataset("p53_sequence_comparison", as_format="text"). Their
sources and processing are described in the
combined oncoplot and
p53 comparison gallery pages.
The package ships the tables that the gallery examples use, so you can try the API on real data without downloading anything.
Data retain their own terms, independently of the source packages’ software licenses. See the third-party notices for file-level status, attribution, and terms.
from genome_spy.datasets import available_datasets, load_dataset
available_datasets()
counts = load_dataset("airway_scaledcounts")
load_dataset() returns a pandas DataFrame for tabular files and parsed JSON
for JSON files. Pass as_format="text" to get the raw file contents instead.
Dataset |
Contents |
|---|---|
|
Gallery airway statistics and sample counts for downstream gene review |
|
Sample table for the airway RNA-seq experiment |
|
Rounded, length-scaled gene counts for the same eight samples |
|
HapMap coordinates with simulated p-values and effect sizes |
|
Somatic mutation calls for one TCGA breast-tumor sample |
|
Historical UniProt feature counts and Pfam protein domains |
|
Plotly Dash Bio alteration fixture |
|
Project-authored synthetic reference window |
|
UCSC hg38 reference sequence with interval metadata |
|
Prepared recurrent PIK3CA mutations and protein domains |
|
Somatic mutation calls for TCGA acute myeloid leukemia |
|
Clinical annotations for those leukemia samples |
|
Prepared mutation, copy-number, clinical, pathway, VAF, and MutSig tables |
|
34 p53 protein sequences aligned with MAFFT L-INS-i, compressed FASTA |
|
Alteration matrix for TCGA lung adenocarcinoma samples |
|
GISTIC2 copy-number scores for TCGA ovarian tumors |
|
GISTIC2 peak regions for the same cohort |
|
Assembly-wide hg19 and hg38 RefSeq gene bodies |
|
Prepared RNF7 direct-RNA coverage, reference bases, read-alignment events, exon intervals, and m6Anet site probabilities |
Data use and provenance
airway_metadataandairway_scaledcountsdescribe the airway smooth muscle RNA-seq experiment of Himes et al., PLoS One 2014 (GEO GSE52778), independently reprocessed by Stephen Turner for the Bioconnector workshops with kallisto and tximportlengthScaledTPM. The bundled CSVs exactly match those workshop files; the workshop declares CC BY-NC-SA 4.0.hapmap_gwasis theHapMapexample table from the manhattanly R package. HapMap supplied the build-36 map coordinates and rs identifiers; UCSC hg18 supplied gene annotations. The association statistics are simulated or derived from simulated p-values. The table contains no individual genotypes.brca_maf,tcga_laml_maf, andtcga_laml_annotationsare the example files bundled with maftools and contain TCGA mutation calls or clinical annotations.pik3ca_tcga_brca_lollipopcontains the chart-ready named datasets from GenomeSpy’s official TCGA-BRCA PIK3CA lollipop example. Mutation counts come from GDC masked somatic MAFs and protein domains from UniProt P42336; see the gallery example for full provenance.pyoncoprint_tcgais the example alteration table from pyoncoprint as a cBioPortal TCGA LUAD export. See the cBioPortal data policy and the study-specific notice linked in our third-party notices. Its microbiome track comes from a study retracted in 2024 and is not a validated finding.tcga_ov_gistic_scoresandtcga_ov_gistic_lesionsare the completescores.gisticandall_lesions.conf_99.txttables used by the official GenomeSpy example. They are open-access TCGA OV-TP GISTIC2 output produced by the Broad Institute TCGA Genome Data Analysis Center, Firehose run 2016-01-28 (source archive).refseq_gene_bodiesis independently prepared from the official hg19 and hg38 UCSCrefGenetables. Overlapping transcripts are collapsed by symbol, chromosome, and strand; transcript counts prioritize colliding labels. See UCSC’s conditions for use.rnf7_direct_rnacombines a reduced transcript-window extract from the CC BY 4.0 xPore demo archive, RNF7 rows from m6Anet Supplementary Table 6, and three GRCh38 exon sequences from Ensembl. Python preparation pseudonymized molecule labels, calculated coverage, and derived mismatches and CIGAR events for the direct-RNA gallery example.
Results shown with the TCGA-derived tables are in whole or part based upon data generated by the TCGA Research Network.