Senior Principal Scientist in Biostatistics · PhD
Senior Principal Scientist in Biostatistics at Johnson & Johnson Innovative Medicine with over six years of industry experience and a PhD from Ghent University. I design and deploy end-to-end data science solutions — spanning multi-omics analysis, predictive modelling, machine learning, AI infrastructure, R Shiny web tools, and R software packages — to accelerate cancer immunotherapy drug discovery. Author of 12+ peer-reviewed publications with 390+ citations in Bioinformatics, Genome Biology, and BMC Genomics.
Alemu Takele Assefa, PhD
Senior Principal Scientist in Biostatistics
Johnson & Johnson Innovative Medicine · Beerse, Belgium
ORCID
0000-0002-7773-0621My current primary focus is developing statistical frameworks that leverage spatial transcriptomics (ST) to navigate the tumor microenvironment (TME) and identify robust, immune-accessible drug targets for cancer immunotherapy. ST preserves tissue architecture and spatial context — crucial for understanding cellular heterogeneity and immune cell–tumor interactions that are invisible to bulk or single-cell sequencing alone.
Profile
I am a Senior Principal Scientist in Biostatistics at Johnson & Johnson Innovative Medicine, where I have progressed through four successive roles since December 2019. My academic foundation — a PhD in Statistical Data Analysis from Ghent University (2016–2020), an MSc in Biostatistics with Distinction from Hasselt University, and a BSc in Statistics & Computer Science with a perfect GPA from Aksum University — has equipped me with deep, cross-disciplinary expertise that bridges statistical theory, computational biology, and applied data science.
My technical repertoire spans the full modern data science stack for biomedical research. In multi-omics analysis, I integrate signals across bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, flow cytometry, CyTOF, and high-content cell imaging to uncover biologically meaningful insights. I apply predictive modelling and machine learning tools for biomarker discovery, dose–response characterisation, and compound screening, and have built AI infrastructure leveraging large language models for automated, data-driven target identification in oncology. I author and maintain R software packages and engineer R Shiny web applications and interactive dashboards that automate complex analytical pipelines and deliver results to scientific stakeholders at scale.
This breadth of expertise is validated by over ten peer-reviewed publications cited more than 321 times in journals including Bioinformatics, Genome Biology, BMC Genomics, and SLAS Discovery, alongside seven invited presentations at international conferences across Europe and North America.
Core Domains
Career
Johnson & Johnson Innovative Medicine · Beerse, Belgium
Institute of Public Health and Epidemiology, University of Gondar
Faculty of Computational Science, Aksum University
Ongoing Work
Alongside day-to-day drug discovery support, I am actively pursuing the following research directions at the interface of statistics, computational biology, and artificial intelligence.
Multi-Omics Drug Discovery
Continued development of multi-omics based analytical frameworks to support target identification, biomarker discovery, and therapeutic decision-making in cancer immunotherapy.
AI & ML-Driven Analysis Automation
Building AI and machine learning infrastructure to automate end-to-end data analysis and biological interpretation pipelines, reducing manual effort and accelerating scientific decision cycles.
Single-Cell Experimental Design Guidelines
Developing prospective and retrospectively applicable guidelines for optimising the design of single-cell experiments — covering sample size, batch structure, and power considerations to maximise reproducibility and biological signal.
Biomarker Discovery Study Design Tools
Creating computational tools for designing and evaluating biomarker discovery studies grounded in multi-omics data, enabling rigorous statistical planning and post-hoc assessment of discovery reliability.
RIPTAC Target & Patient Stratification
Improving tumour target discovery and patient population identification for RIPTAC (Receptor-Induced Proximity Targeting & Cancer) therapy development, integrating multi-omics signals to refine candidate selection and patient stratification in oncology.
Open Source
Across 15 public repositories, my open-source work spans five interconnected areas of computational biology and biostatistics. Several key packages are hosted under collaborative GitHub organisations — CenterForStatistics-UGent (Ghent University) and koenvandenberge — reflecting academic partnerships formed during and after the PhD. Differential gene expression methodology is represented by tools for bulk RNA-seq, single-cell RNA-seq, and long non-coding RNA data — from simulation frameworks (SPsimSeq) and Shiny-based exploration apps (AppDGE) to non-parametric model implementations (PIMseq). Experimental design & data utility is covered by tools for RNA subsampling (RNAsub) and optimal sample pooling strategies. Single-cell & multi-omics analysis is the focus of voomCLR, an R package for differential cell composition analysis across scRNA-seq, flow cytometry, and CyTOF. High-content imaging & compound screening is addressed by the HCIcellPaintingAssayQCtool, which automates quality control of Cell Painting assays using 2D biosignature prediction intervals. Finally, reproducible research & supplementary archives house code, extended results, and additional figures for published manuscripts from the PhD thesis and subsequent journal articles.
R Packages & Core Tools
R package for differential cell composition analysis in single-cell studies. Applies centred log-ratio (CLR) transformation with voom-based precision weighting and bias correction to scRNA-seq, flow cytometry, and CyTOF data. Hosted under collaborator Koen Van den Berge's account. Companion to the 2025 Bioinformatics publication.
Lightweight R package for subsampling RNA-seq datasets to a target library size or total count per sample. Streamlined from the subSeq Bioconductor package for flexible integration in custom analysis and simulation pipelines.
R package implementing Probabilistic Index Models (PIM) for differential expression testing in bulk and single-cell RNA-seq data. A distribution-free, rank-based approach that handles complex experimental designs without strong parametric assumptions. Hosted under the Center for Statistics UGent organisation.
Simulation & Experimental Design
Semi-parametric simulator for bulk and single-cell RNA-seq data. Uses an exponential-family density estimator and Gaussian copulas to learn distributional properties — including gene-gene correlations and zero-inflation — from real data, then generates realistic synthetic datasets for benchmarking DGE tools. Also available on Bioconductor. Published in Bioinformatics (2020). Hosted under the Center for Statistics UGent organisation.
Reproducible code and supplementary results for the BMC Genomics (2020) paper on RNA sample pooling strategies — evaluating how pooling affects statistical power and cost-efficiency across different experimental designs.
Interactive Tools & Shiny Apps
An interactive R Shiny application guiding researchers through differential gene expression analysis workflows. Provides an accessible interface for RNA-seq data exploration, tool selection, and comparison of DGE methods across multiple statistical frameworks.
High-Content Imaging & Drug Screening
Automated quality control tool for high-content imaging (Cell Painting) data. Constructs 2D prediction intervals on reference biosignatures to detect aberrant assay plates in compound screening campaigns — supporting the published SLAS Discovery (2023) methodology.
Reproducible Research & PhD Archives
Supplementary results, extended analyses, and reproducible scripts accompanying the Ghent University PhD thesis on statistical methods for bulk and single-cell digital transcriptomics. Includes additional figures and benchmarking code supporting all methodology chapters.
Research Output
Recognition
Recipient of three Innovation Leadership Awards at Johnson & Johnson Innovative Medicine, recognising the development of novel statistical and computational tools that advance drug discovery and single-cell data analysis.
Academic Training
Technical Competencies
Programming
Data Modalities
Statistical Methods
Tools & Platforms
Languages
Scientific Engagement
Beyond the Lab
Science is what I do, but curiosity, warmth, and a love for the world around me are who I am. Here is a glimpse of life beyond the data.
Get in Touch
Open to scientific collaborations, consulting, and conversations around biostatistics, computational biology, and drug discovery data science. Find me on any of the platforms below.
Guestbook
Notes & Comments
Leave a note, a question, or a thought on my work. Comments are powered by GitHub Discussions, so a free GitHub sign-in keeps the conversation spam-free and public.
Loading comments… If nothing appears, the guestbook is still being set up — feel free to email me in the meantime.