Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, their giga-pixel size creates two intertwined bottlenecks: computational inefficiency, as prohibitive GPU memory makes standard end-to-end (E2E) training infeasible, and label inefficiency, as expert annotation is tedious and expensive. This dissertation confronts both challenges through novel architectures, training paradigms, and self-supervised learning methods for efficient WSI analysis.
To improve computational efficiency, this dissertation first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework training only a few prompts to guide large pre-trained models, reducing trainable parameters, memory, and training time. Third, this work proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy in 1D models. Fourth, it presents Locally Bi-directional Mamba (LBMamba), whose hardware-aware local backward scan integrates bi-directional scanning into a single forward pass, improving the throughput-performance trade-off of Mamba models.
To improve label efficiency, this dissertation proposes a precise location based matching strategy for self-supervised dense contrastive learning, which allows a local patch in one augmented view to match multiple overlapping patches in another, producing more accurate correspondences and superior features for dense prediction tasks like segmentation and detection. Additionally, to better scale multi-channel cell imaging modalities, this dissertation introduces ChannelSFormer, a channel-agnostic vision transformer that disentangles spatial and channel-wise reasoning through divided attention and channel class token, enabling effective representation learning across variable channel configurations in both self-supervised and supervised settings.
In summary, this dissertation presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.
Speaker: Jingwei Zhang
Location: NCS 220
Zoom:
https://stonybrook.zoom.us/j/93175806292?pwd=xbtxnQyYGoThz5B1DyJxJxPF9lxiJE.1Meeting ID: 931 7580 6292
Passcode: 314091