Our research focuses on processing, annotation, and interpretation of cutting-edge omics datasets to understand human diseases such as cancers. Our approaches include computational/statistical modeling, machine learning, large data integration, and close wet-lab collaboration.

Integrative Omics

Data accumulation in multi-omics enables us to better interpret gene/protein functions through integrative approaches. Our group develops computational methods to address data integration with: gene expression (RNA-seq, scRNA-seq, spatial transcriptome), chromatin status/accessibility (ChIP-seq, ATAC-seq, scATAC-seq), 3D chromatin looping (HiC & HiChIP), genome editing (CRISPR screens) and protein/metabolite expression (spatial proteomics/MALDI).

Diagram of a super-enhancer's ATAC/ChIP-seq signal, its Hi-C loop to a target gene, and the resulting rise in RNA-seq expression, all aligned to the same genomic position. Diagram of noisy, batch-specific sequencing tracks passing through a statistical correction model to produce signal that is comparable across samples.

Cancer Biomarkers

Cancers are one of the main biological settings where we apply our computational methodologies . Our recent work focuses on understanding oncogenesis mechanisms based on close collaboration with investigators from diverse background. We study oncogenesis and therapeutics associated to oncogenic viruses (e.g. Epstein–Barr virus and human papillomavirus etc.) and clonal hematopoiesis.

Diagram of a host chromatin loop before viral infection versus after, when a distinctly-shaped EBV particle and HPV capsid rewire the loop to switch on an oncogene. Diagram of clonal hematopoiesis: one mutant blood stem cell clonally expanding, alongside a scatter plot of CH driver genes by cancer type and age at detection.