Senior Principal Scientist in Biostatistics · PhD

Alemu Takele, Assefa

Biostatistics & Computational Biology

Senior Principal Scientist in Biostatistics at Johnson & Johnson Innovative Medicine with over six years of industry experience and a PhD from Ghent University. I design and deploy end-to-end data science solutions — spanning multi-omics analysis, predictive modelling, machine learning, AI infrastructure, R Shiny web tools, and R software packages — to accelerate cancer immunotherapy drug discovery. Author of 12+ peer-reviewed publications with 390+ citations in Bioinformatics, Genome Biology, and BMC Genomics.

View Publications Get in Touch
ATA

Alemu Takele Assefa, PhD

Senior Principal Scientist in Biostatistics
Johnson & Johnson Innovative Medicine · Beerse, Belgium

390+
Citations
12+
Publications
6+
Yrs Industry

ORCID

0000-0002-7773-0621


Current Research Focus
Presented at EuroBioc2026 · June 3–5, 2026 · Turku, Finland

Spatial Transcriptomics for Tumor Microenvironment Characterisation & Drug Target Prioritisation

My current primary focus is developing statistical frameworks that leverage spatial transcriptomics (ST) to navigate the tumor microenvironment (TME) and identify robust, immune-accessible drug targets for cancer immunotherapy. ST preserves tissue architecture and spatial context — crucial for understanding cellular heterogeneity and immune cell–tumor interactions that are invisible to bulk or single-cell sequencing alone.

Spatial target heterogeneity. Quantifying within- and between-donor variability of tumor target positivity across tissue islands using CV-based metrics, to prioritise spatially consistent, druggable candidates.
Immune cell infiltration (ICI) scoring. Proposing quantitative proximity-based scores — overall infiltration fraction, fraction of tumour cells with ≥ n immune neighbours, and mean immune neighbour count — to map tumour accessibility to immune effectors.
Multi-target integration. Combining spatial heterogeneity and ICI scores across tumour targets and immune cell types (CD8+ T cells, NK cells, macrophages) to rank targets by their therapeutic actionability.
View Code on GitHub
~168K
Cells retained after QC filtering
100
Spatial islands per tissue sample (k-means)
CVW & CVB
Within- & between-donor variability metrics
4
Immune cell subtypes profiled spatially
TUMOR ISLANDS Target+ Target− IMMUNE INFILTRATION R = radius threshold ● Tumor cell ● Immune cell -- Radius R
ST Data10x Genomics FFPE
Island QCk-means · 100 regions
CV ScoringCVW & CVB
ICI ScoreProximity metrics
Target RankPrioritisation

Research & Expertise

I am a Senior Principal Scientist in Biostatistics at Johnson & Johnson Innovative Medicine, where I have progressed through four successive roles since December 2019. My academic foundation — a PhD in Statistical Data Analysis from Ghent University (2016–2020), an MSc in Biostatistics with Distinction from Hasselt University, and a BSc in Statistics & Computer Science with a perfect GPA from Aksum University — has equipped me with deep, cross-disciplinary expertise that bridges statistical theory, computational biology, and applied data science.

My technical repertoire spans the full modern data science stack for biomedical research. In multi-omics analysis, I integrate signals across bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, flow cytometry, CyTOF, and high-content cell imaging to uncover biologically meaningful insights. I apply predictive modelling and machine learning tools for biomarker discovery, dose–response characterisation, and compound screening, and have built AI infrastructure leveraging large language models for automated, data-driven target identification in oncology. I author and maintain R software packages and engineer R Shiny web applications and interactive dashboards that automate complex analytical pipelines and deliver results to scientific stakeholders at scale.

This breadth of expertise is validated by over ten peer-reviewed publications cited more than 321 times in journals including Bioinformatics, Genome Biology, BMC Genomics, and SLAS Discovery, alongside seven invited presentations at international conferences across Europe and North America.

Bulk & single-cell transcriptomics (RNA-seq, scRNA-seq)
Spatial transcriptomics
Proteomics
Cell phenotype imaging — Cell Painting
Flow cytometry & CyTOF
Statistical method development
Cancer immunotherapy & drug discovery
Dose–response & biomarker analysis
AI / LLM-assisted target identification
R Shiny development & data visualization
Clinical trial & pre-clinical translation

Work Experience

Expertise Map

Hover or tap a node to trace connections · drag to rearrange
Core profile Research domain Software / tool Method & application
Dec 2019 — Present
Beerse, Belgium
Johnson & Johnson
EXPRESSION LEVELS AI JOHNSON & JOHNSON INNOVATIVE MEDICINE

Senior Principal Scientist in Biostatistics

Johnson & Johnson Innovative Medicine · Beerse, Belgium

Career progression ladder Four successive J&J roles shown on a rising diagonal, from Senior Scientist in 2019 to Senior Principal Scientist in 2026. Senior Principal Scientist Jul 2026 — Present CURRENT Principal Scientist II Dec 2024 — Jul 2026 Principal Scientist Dec 2022 — Dec 2024 Senior Scientist Dec 2019 — Dec 2022
  • Integrated LLM-based AI to support target identification for immunotherapy antibody design in oncology.
  • Developed a database integrating pre-clinical and clinical drug discovery data to assist translation of in-vitro and in-vivo findings for safety and efficacy prediction in clinical trials.
  • Automated image analysis for Colony Formation Unit (CFU) counting from clonogenic assays for drug screening.
  • Developed the R package voomCLR for differential cell composition analysis of single-cell experiments (scRNA-seq, flow cytometry, CyTOF).
  • Developed automated positive-control-based quality assessment of high-content image cell phenotype data for compound screening.
  • Automated a metric for optimization of bi-specific antibody design for cancer immunotherapy using scRNA-seq.
  • Improved methodology for immune cells' polyfunctionality detection using the polyfunctionality strength index (PSI) from flow cytometry.
  • Developed pipelines for differential gene expression analysis of scRNA-seq data from complex experimental designs.
Oct 2015 — Jan 2016
Gondar, Ethiopia
SURVIVAL ANALYSIS POPULATION HEALTH BIOSTATISTICS UNIVERSITY OF GONDAR · INSTITUTE OF PUBLIC HEALTH

Lecturer in Biostatistics

Institute of Public Health and Epidemiology, University of Gondar

Sep 2010 — Sep 2013
Aksum, Ethiopia
ŷ = β₀ + β₁x₁ + ε H₀: μ₁ = μ₂ vs H₁: μ₁ ≠ μ₂ P(A|B) = P(B|A)·P(A) / P(B) ∑ xᵢ / n = x̄ σ² = E[(X - μ)²] AKSUM UNIVERSITY PROBABILITY μ AKSUM UNIVERSITY · FACULTY OF COMPUTATIONAL SCIENCE

Assistant Lecturer in Statistics & Mathematics

Faculty of Computational Science, Aksum University


Current Research Activities

Alongside day-to-day drug discovery support, I am actively pursuing the following research directions at the interface of statistics, computational biology, and artificial intelligence.

🔬

Multi-Omics Drug Discovery

Continued development of multi-omics based analytical frameworks to support target identification, biomarker discovery, and therapeutic decision-making in cancer immunotherapy.

🤖

AI & ML-Driven Analysis Automation

Building AI and machine learning infrastructure to automate end-to-end data analysis and biological interpretation pipelines, reducing manual effort and accelerating scientific decision cycles.

🧫

Single-Cell Experimental Design Guidelines

Developing prospective and retrospectively applicable guidelines for optimising the design of single-cell experiments — covering sample size, batch structure, and power considerations to maximise reproducibility and biological signal.

📊

Biomarker Discovery Study Design Tools

Creating computational tools for designing and evaluating biomarker discovery studies grounded in multi-omics data, enabling rigorous statistical planning and post-hoc assessment of discovery reliability.

🎯

RIPTAC Target & Patient Stratification

Improving tumour target discovery and patient population identification for RIPTAC (Receptor-Induced Proximity Targeting & Cancer) therapy development, integrating multi-omics signals to refine candidate selection and patient stratification in oncology.


GitHub Projects

Across 15 public repositories, my open-source work spans five interconnected areas of computational biology and biostatistics. Several key packages are hosted under collaborative GitHub organisations — CenterForStatistics-UGent (Ghent University) and koenvandenberge — reflecting academic partnerships formed during and after the PhD. Differential gene expression methodology is represented by tools for bulk RNA-seq, single-cell RNA-seq, and long non-coding RNA data — from simulation frameworks (SPsimSeq) and Shiny-based exploration apps (AppDGE) to non-parametric model implementations (PIMseq). Experimental design & data utility is covered by tools for RNA subsampling (RNAsub) and optimal sample pooling strategies. Single-cell & multi-omics analysis is the focus of voomCLR, an R package for differential cell composition analysis across scRNA-seq, flow cytometry, and CyTOF. High-content imaging & compound screening is addressed by the HCIcellPaintingAssayQCtool, which automates quality control of Cell Painting assays using 2D biosignature prediction intervals. Finally, reproducible research & supplementary archives house code, extended results, and additional figures for published manuscripts from the PhD thesis and subsequent journal articles.

DGE Methodology SPsimSeq · PIMseq · AppDGE Experimental Design RNAsub · RNA-pooling Single-cell & Multi-omics voomCLR HCI & Drug Screening HCIcellPaintingAssayQCtool R Reproducible Research PhD thesis · Supplementary 15 PUBLIC REPOS R Bioconductor

R Packages & Core Tools

R package for differential cell composition analysis in single-cell studies. Applies centred log-ratio (CLR) transformation with voom-based precision weighting and bias correction to scRNA-seq, flow cytometry, and CyTOF data. Hosted under collaborator Koen Van den Berge's account. Companion to the 2025 Bioinformatics publication.

RNAsub R

Lightweight R package for subsampling RNA-seq datasets to a target library size or total count per sample. Streamlined from the subSeq Bioconductor package for flexible integration in custom analysis and simulation pipelines.

PIMseq R

R package implementing Probabilistic Index Models (PIM) for differential expression testing in bulk and single-cell RNA-seq data. A distribution-free, rank-based approach that handles complex experimental designs without strong parametric assumptions. Hosted under the Center for Statistics UGent organisation.

Simulation & Experimental Design

Semi-parametric simulator for bulk and single-cell RNA-seq data. Uses an exponential-family density estimator and Gaussian copulas to learn distributional properties — including gene-gene correlations and zero-inflation — from real data, then generates realistic synthetic datasets for benchmarking DGE tools. Also available on Bioconductor. Published in Bioinformatics (2020). Hosted under the Center for Statistics UGent organisation.

Reproducible code and supplementary results for the BMC Genomics (2020) paper on RNA sample pooling strategies — evaluating how pooling affects statistical power and cost-efficiency across different experimental designs.

Interactive Tools & Shiny Apps

AppDGE R

An interactive R Shiny application guiding researchers through differential gene expression analysis workflows. Provides an accessible interface for RNA-seq data exploration, tool selection, and comparison of DGE methods across multiple statistical frameworks.

High-Content Imaging & Drug Screening

Automated quality control tool for high-content imaging (Cell Painting) data. Constructs 2D prediction intervals on reference biosignatures to detect aberrant assay plates in compound screening campaigns — supporting the published SLAS Discovery (2023) methodology.

Reproducible Research & PhD Archives

Supplementary results, extended analyses, and reproducible scripts accompanying the Ghent University PhD thesis on statistical methods for bulk and single-cell digital transcriptomics. Includes additional figures and benchmarking code supporting all methodology chapters.

View all 15 repositories on GitHub

Selected Publications

View my complete publication record including citation metrics, co-author networks, and full-text links on my academic profiles.

2025
Bioinformatics · Oxford University Press
Assessing differential cell composition in single-cell studies using voomCLR
Alemu Takele Assefa, Bie Verbist, Koen Van den Berge
2025
BMC Genomics · Springer
Differential detection workflows for multi-sample single-cell RNA-seq data
Jeroen Gilis, Laura Perin, Milan Malfait, Helena L Crowell, Koen Van den Berge, Alemu Takele Assefa, Bie Verbist, Davide Risso, Lieven Clement
preprint
bioRxiv · Cold Spring Harbor Laboratory, 2024
Strategies for addressing pseudoreplication in multi-patient scRNA-seq data
Milan Malfait, Jeroen Gilis, Koen Van den Berge, Alemu Takele Assefa, Bie Verbist, Lieven Clement
2024
Blood · American Society of Hematology
JNJ-87801493 (CD20×CD28), a potential first-in-class CD20-targeted CD28 costimulatory bispecific antibody, enhances the activity of B-cell-targeting T-cell engagers in preclinical models
Lorena Fontan, Adam Zwolak, Irene Guimerans-Lorenzo, Mariette Bekkers, Nicholas Hein, Nele Vloemans, Emanuele Trella, Tina Smets, Ivo Cornelissen, Alemu Assefa, Bie Verbist, et al.
SLAS Discovery · Elsevier, 2023
Automated quality control tool for high-content imaging data by building 2D prediction intervals on reference biosignatures
Alemu Takele Assefa, Bie Verbist, Emmanuel Gustin, Danielle Peeters
Bioinformatics · Oxford University Press, 2020
SPsimSeq: semi-parametric simulation of bulk and single-cell RNA-sequencing data
Alemu Takele Assefa, Jo Vandesompele, Olivier Thas
BMC Genomics · Springer, 2020
On the utility of RNA sample pooling to optimize cost and statistical power in RNA sequencing experiments
Alemu Takele Assefa, Jo Vandesompele, Olivier Thas
Genome Biology · BioMed Central, 2018
Differential gene expression analysis tools exhibit substandard performance for long non-coding RNA-sequencing data
Alemu Takele Assefa, Katrijn De Paepe, Celine Everaert, Pieter Mestdagh, Olivier Thas, Jo Vandesompele
Integrated Blood Pressure Control · Taylor & Francis, 2018
Blood pressure control status and associated factors among adult hypertensive patients on outpatient follow-up at University of Gondar Referral Hospital
Yaregal Animut, Alemu Takele Assefa, Dereseh Gezie Lemma

Cited By

The most influential papers that build on my publications, ranked by their own citation count.

Honors & Awards

Recipient of three Innovation Leadership Awards at Johnson & Johnson Innovative Medicine, recognising the development of novel statistical and computational tools that advance drug discovery and single-cell data analysis.

2022
Innovation Leadership Award
Johnson & Johnson Innovative Medicine
For the development of a novel quality-control tool for high-content imaging data from cell painting assays.
Reference: Automated quality control tool for high-content imaging data by building 2D prediction intervals on reference biosignatures.
2023
Innovation Leadership Award
Johnson & Johnson Innovative Medicine
For the development of a pipeline for co-expression analysis of single-cell RNA-seq data to identify paired tumour targets using OR and AND logic combinations.
2025
Innovation Leadership Award
Johnson & Johnson Innovative Medicine
For developing the voomCLR tool for effective analysis of cell composition by correcting for compositional bias.
Reference: Assessing differential cell composition in single-cell studies using voomCLR.

Education & Degrees

PhD in Statistical Data Analysis
Faculty of Science, Ghent University
February 2016 — December 2020 · Ghent, Belgium
Statistical methods for the analysis of bulk and single-cell digital transcriptomics data. Developed novel statistical methods and optimized experimental designs for differential gene expression analysis. Teaching assistance for Statistical Genomics; supervised master theses.
Research: scRNA-seq · DGE analysis · Simulation methods
MSc in Biostatistics
Data Science Institute, Hasselt University
October 2013 — September 2015 · Hasselt, Belgium
Advanced statistical inference, regression, survival and longitudinal modelling, multivariate techniques, infectious disease modelling. Thesis: Modelling CD4+ cell counts and hemoglobin concentration for HIV-1 patients on ART in Mildmay Uganda, using joint mixed-effect models.
Graduated with Distinction (75%)
BSc in Statistics & Computer Science
Faculty of Computational Science, Aksum University
November 2007 — August 2010 · Aksum, Ethiopia
Statistics, probability theory, experimental design, time series, multivariate analysis. Computer science including OOP in C++, data structures, database systems, and system design.
Graduated with Very Great Distinction (CGPA 4.0/4.0)

Skills & Tools

Programming

RPython SASC++Git

Data Modalities

Bulk RNA-seq scRNA-seq Spatial transcriptomics Proteomics Flow cytometry CyTOF Cell Painting

Statistical Methods

DGE analysis Mixed-effect models Compositional DA Dose–response modelling Biomarker discovery Dimensionality reduction

Tools & Platforms

R ShinyBioconductor LaTeXDashboard design LLM / AI tools

Languages

English (professional) Amharic (native) Dutch (basic)

Conferences & Presentations

2026
EuroBioc2026 — European Bioconductor Conference
Poster: Spatial Transcriptomics Enables Navigation of the Tumor Microenvironment and Evaluation of Tumor Accessibility for Robust Drug Target Prioritization · June 3–5, 2026
Turku, Finland
2024
NCS2024 — Non-Clinical Statistics Conference
September 23–25, 2024
Wiesbaden, Germany
2023
EuroBioc2023 — European Bioconductor Conference
September 20–22, 2023
Ghent, Belgium
2022
NCS2022 — Non-Clinical Statistics
Oral: Scoring and visualizing polyfunctionality in flow cytometry data
Louvain-la-Neuve, Belgium
2019
ISCB — 40th Annual Conference
Oral: Probabilistic index models for testing DGE in scRNA-seq data
Leuven, Belgium
2019
ENAR Spring Meeting
Oral: Probabilistic index models for testing DGE in scRNA-seq data
Philadelphia, USA
2017
IBSXXIX — International Biometric Conference
Oral: DGE tools evaluation for long non-coding RNA sequencing data
Barcelona, Spain
2017
CNC6th — 6th Channel Network Conference, IBS
Poster: Non-parametric simulation for DGE methods in lncRNA-seq cancer studies
Hasselt, Belgium

Life & Interests

Science is what I do, but curiosity, warmth, and a love for the world around me are who I am. Here is a glimpse of life beyond the data.

✈️
Travel
From the highlands of Ethiopia to the heart of Europe, I find inspiration in discovering new cultures, landscapes, and perspectives. Every journey broadens the mind as much as any study.
📚
Reading
A dedicated reader across genres — from scientific literature and methodology texts to history, biography, and contemporary fiction. Books have always been a companion and a teacher.
Coffee Culture
Ethiopia is the birthplace of coffee, and I carry that heritage with pride. The traditional Ethiopian coffee ceremony — bunna — is a ritual of community, conversation, and connection that I cherish deeply.
👨‍👩‍👧‍👦
Family Time
At the heart of everything is family. Whether sharing meals, exploring Belgium together, or staying close with loved ones back in Ethiopia, family time restores and grounds me.

Let's Connect

Open to scientific collaborations, consulting, and conversations around biostatistics, computational biology, and drug discovery data science. Find me on any of the platforms below.


Notes & Comments

Leave a note, a question, or a thought on my work. Comments are powered by GitHub Discussions, so a free GitHub sign-in keeps the conversation spam-free and public.

Loading comments… If nothing appears, the guestbook is still being set up — feel free to email me in the meantime.