The Data Science Garage

Using Data Science to Fight Cancer

The Lab

Our lab brings together computer engineering, statistical analysis, and biology to understand complex biological systems and cancer. Based at the OHSU Knight Cancer Institute, we develop computational approaches for systems biology, integrative analysis, and cancer research.

Our research spans reliable data systems, statistical and machine-learning methods, and collaborative challenges that help researchers address difficult biomedical questions. We emphasize open, reproducible tools and analyses that make complex biological data more useful for discovery.

Our work is supported by the National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI).

Projects

Cools projects

CALYPR
A hybrid cloud and on-prem genomics platform for data integration, workflow execution, and analysis.
GDAN Tumor Molecular Pathology
Pretrained models classify non-TCGA cancer samples into molecular subtypes defined by TCGA.
AnVIL
A cloud platform for accessing, managing, and analyzing genomic data through NHGRI’s AnVIL data commons.
SMC Het
A DREAM challenge to compare methods for reconstructing tumor subclones and their evolution.
GRIP
A graph database and query platform that works across multiple storage back ends.
BMEG
A graph model for connecting biomedical data and evidence to support cross-domain discovery.
DREAM-SMC
Somatic Mutation Calling
DREAM-SMC-RNA
RNA Fusion Calling and Isoform Quantification
Funnel
A distributed task execution system for cloud and high-performance computing workflows.
MC3
A uniform reanalysis of TCGA exome data to support cross-cancer mutation studies.

People

Researchers

Avatar

Brian Walsh

Senior Research Software Engineer

Graph Databases, Machine Learning, Pan Cancer, Cross Project Data Harmonization

Avatar

Isabel Quesada

Computational Biologist

Avatar

Matthew Peterkort

Research Software Engineer

Graph Databases, JSON Schema, Data Science

Principal Investigators

Avatar

Kyle Ellrott

Associate Professor

Computational Biology, Data Science, Machine Learning, Precision Medicine, Cancer Early Detection

Alumni

Avatar

Malisa Smith

Student Intern

Avatar

Jeena Lee

Student Intern

Avatar

Ryan Spangler

Research Software Engineer

Avatar

Adam Struck

Research Software Engineer

Graph Databases, Machine Learning, Data Science

Avatar

Brian Karlberg

PhD student

Computational Biology, Precision Oncology

Avatar

Jordan Tagle

Computational Biologist

Machine Learning, Precision Oncology, Cancer Early Detection

Avatar

Liam Beckman

Research Software Engineer

Software Development, Computational Biology

Avatar

Nasim Sanati

Computational Biologist

Biomedical Knowledge Graphs, Data Seience

Avatar

Quinn Wai Wong

Research Software Engineer

Software Development, Systems Design, Computational Biology

Recent Publications

A compendium of next-generation patient-derived models for diverse cancers

The development of new therapeutics and the validation of pathogenetic cancer mechanisms require representative laboratory models1,2. However, existing collections represent only a fraction of the diversity observed in human cancer2-4. Recent technologies have enabled efficient in vitro model derivation (for example, tumour organoids)5. However, whether these maintain essential properties of patient tumours during long-term expansion has not been systematically investigated. Here we present results of a large-scale international programme-the Human Cancer Models Initiative-which involved the generation of a resource of 665 next-generation models from 2,780 donors with 25 cancer types and integrated tumour-model whole genome, exome, methylome and transcriptome analyses. The resource provides 522 models with comprehensive clinical data, 153 models of rare cancers and 71 models from participants with non-European ancestry. Analyses of 421 matched tumour-model pairs reveal high genetic (97.8%) and epigenetic (95%) concordance and define correlates of model discordance. Single-nucleus RNA sequencing of tumour-model pairs reveals subsets of models in which culture conditions significantly influence cell states. Finally, we characterize model preservation of extrachromosomal DNA and post-treatment mutational signatures to provide opportunities to study therapeutic resistance. This model repository is being made available to the community-including multimodal molecular profiling, clinical information and integrative software tools-thus providing a valuable resource for preclinical investigation of cancer pathogenesis and treatment response.

Pooling multimodal cancer data across unaligned embedding spaces maintains tumor of origin signal

AI-based embeddings offer the possibilities of encoding complex biological data into low-dimensional spaces, called embedding spaces, that maintain the relationships between entities. Vector pooling is the process of aggregating an array of embedded vectors, either usually by summing or averaging, to summarize the total movement with the embedding space. Embedded vector pooling allows sampling of an arbitrary number of points to be summarized into a fixed sized vector, and is frequently used to sample networks of embedded values or to summarize protein language model vectors. There is an open question about the compatibility of embedding spaces that are created without any coordination. It has been assumed that signals in these unaligned embedding spaces would be destroyed if vectors were pooled into summed values. To challenge this idea, we created a number of benchmarks that utilized unaligned embedded values and pooled them into heterogeneous vectors to test information retrieval. To power this benchmark, we trained embedding models across different cancer data modalities and tested how well pooled heterogeneous vectors were able to retain biologically relevant information. Our research shows that signal from unaligned embedded values is conserved and able to still be used for learning tasks, such as data modality and tumor of origin recognition. All code and computational experiments related to this publication can be found at https://github.com/EllrottLab/heterogeneous-embedding-vectors.

A decentralized future for the open-science databases

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR’s federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

Robust software development practices improve citations of RNA-seq tools

RNA sequencing (RNA-seq) has emerged as an exemplary technology in biology and clinical applications, offering a crucial complement to other transcriptomic profiling protocols due to its high sensitivity, precision, and accuracy in characterizing transcriptomes. However, the rapid proliferation of RNA-seq tools necessitates the adoption of robust software development practices. Such development underscores the critical need to examine how RNA-seq tools are developed, maintained, and distributed; and whether the data they generate is reproducible as all of these factors are essential for ensuring software reliability, transparency, and trust in scientific findings. We conducted a comprehensive assessment of 434 RNA-seq tools developed between 2008 and 2024, categorizing them based on the type of analysis they perform. Our evaluation encompassed their software development and distribution methodologies, as well as the attributes contributing to their widespread adoption and dependability within the biomedical community, which were quantified by factors such as package manager availability, containerization, multithreading support, documentation quality, and inclusion of example datasets. Our findings establish the first documented positive association between rigorous software development practices and their adoption of published RNA-seq tools as measured by citations (Mann-Whitney U test, p-value = 4.9e-26). By identifying key characteristics of widely adopted software, our findings guide developing robust and user-friendly RNA-seq tools, thereby reinforcing the call for rigorous community-wide standards.

Pan-cancer immune and stromal deconvolution predicts clinical outcomes and mutation profiles

Traditional gene expression deconvolution methods assess a limited number of cell types, therefore do not capture the full complexity of the tumor microenvironment (TME). Here, we integrate nine deconvolution tools to assess 79 TME cell types in 10,592 tumors across 33 different cancer types, creating the most comprehensive analysis of the TME. In total, we found 41 patterns of immune infiltration and stroma profiles, identifying heterogeneous yet unique TME portraits for each cancer and several new findings. Our findings indicate that leukocytes play a major role in distinguishing various tumor types, and that a shared immune-rich TME cluster predicts better survival in bladder cancer for luminal and basal squamous subtypes, as well as in melanoma for RAS-hotspot subtypes. Our detailed deconvolution and mutational correlation analyses uncover 35 therapeutic target and candidate response biomarkers hypotheses (including CASP8 and RAS pathway genes).