Throughout the past decade, the field of computational histopathology has witnessed remarkable advancements, both in quality and efficiency, largely influenced or accelerated by contemporary advancements in Deep Learning (DL) tools and methods. This progress has brought the clinical adoption of DL-backed technology nearer, promising significant benefits to patients and service providers. Such technology offers valuable insights into the diagnosis process, streamlines repetitive and complex tasks, alleviates the risk of inconsistency and bias in clinical decisions, aids in discovering novel treatment methods and plans, and helps devise ways to predict the efficacy of existing ones. Nonetheless, an infamous challenge facing DL models is their strong reliance on expansive datasets. Large, powerful models can contain billions of parameters; supervising their training effectively requires vast labeled datasets to achieve reliable generalization and noise resilience. Amassing such extensive labeled data collections places additional demands on field experts, a burden exacerbated in histopathology, incurring higher costs and hindering the field’s progress. Motivated by the great success in deep learning and the commensurate need to address data scarcity in computational histopathology, this dissertation aims to scout, develop, and devise advanced solutions to this problem, while obviating the cumbersome expansion of datasets through expert annotation. Instead, we explore methods from within the DL field that rely on inexpensive, accessible, applicable, and generalizable alternatives to exhaustive supervision. To tackle this problem, we make several contributions, each targeting a different histopathology task (tissue segmentation, nuclei detection, slide classification, patch classification, and nuclei segmentation), all unified under the overarching theme of mitigating data scarcity. Specifically, we make the following contributions: 1) we undertake a comprehensive review of the state of the art, organizing existing approaches into discrete categories and highlighting potential research gaps. This review was designed to be both comprehensive and informative, serving as a guide for this dissertation’s progress and for other researchers interested in the domain. 2) We develop a new method to synthesize histopathology images from existing ones, effectively doubling data size without requiring physical acquisition or expert annotation. This approach enables us to train a nucleus detection model that outperforms several state-of-the-art detection models, affirmed by visual and quantitative comparisons. 3) We develop two modules attachable to Multiple Instance Learning methods, the primary techniques for weakly-supervised Whole Slide Image classification. These modules integrate topological information from the tumor environment, contrasting with previous works that rely solely on slide morphology or chromaticity. We confirm, visually and quantitatively, that these modules can improve existing methods by up to 22% without additional labeled data. 4) We address the pressing issue of class imbalance, diagnosing it as a symptom of (selective) data scarcity, and develop an algorithmic remedy that complements other data-driven methods proposed herein. We showcase this method’s efficacy on a patch classification task, where it outperforms relevant existing methods. 5) We investigate and implement a novel semi-supervised learning approach that integrates unlabeled samples from external datasets after passing them through a selection mechanism. We demonstrate the benefit of this selection mechanism in filtering appropriate external data sources and the overall pipeline’s efficacy in enhancing tissue segmentation in histopathology. In addition to the above contributions, we propose a novel self-supervised method, detailed in an appendix as future work, specifically designed for the downstream task of nuclei segmentation. This method involves a creative surrogate task intended to confer the learner with features highly relevant to nuclei segmentation, offering a targeted alternative to generic self-supervised approaches. With these contributions, we hope to address the issue of data scarcity in Deep Learning applied to histopathology from multiple angles, thereby making a significant contribution towards promoting the clinical adoption of these valuable tools.
| Date of Award | 2025 |
|---|
| Original language | American English |
|---|
| Supervisor | Naoufel Werghi (Supervisor) |
|---|
- Computational Histopathology
- Deep Learning
- Data Scarcity
- Data-Efficient Learning
- Semi-Supervised Learning
- Weakly-supervised Learning
Deep Learning in Computational Histopatholgoy Under the Constraint of Data Scarcity
Obeid, A. (Author). 2025
Student thesis: Doctoral Thesis