Chapter Four · failure evidence

What Image Segmentation & Computer Vision got wrong, from 95 dissertations

The records document practical breakdowns and engineering trade-offs encountered across computer vision and image segmentation systems. Across these trials, deep neural models and heuristic tools frequently fail due to domain shifts, heavy computational demands, boundary confusion, and underperformance relative to simpler baselines. These records come from PhD theses at 32 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Complex neural architectures often underperform simpler models or established baselines

21 theses · 15 institutions

Attempts to deploy diffusion models, promptable foundation models, adversarial training, and single-stage instance detectors often yielded higher error rates and worse segmentation accuracy than standard discriminative baselines. In addition, simpler classifiers, basic U-Nets, and traditional clustering algorithms regularly matched or exceeded the performance of these larger designs.

Tried and failed

diffusion models applied to deterministic semantic segmentation. Outcome: worse than baseline. Reason: did not outperform standard discriminative models on deterministic prediction benchmarks

Fast and Future: Towards Efficient Forecasting in Video Semantic Segmentation · EPFL

Tried and failed

promptable foundation model post-processing for segmentation applied to color-sensitive semantic segmentation. Outcome: worse than baseline. Reason: expanded false positives by segmenting non-target objects sharing color features

A Machine Learning Approach to Recognize Environmental Features Associated with Social Factors · Virginia Tech

Lost to a baseline

A simple 1% labeled-frame threshold baseline beat the complex combined labeling-and-segmentation dual-threshold method in prognostic stratification (85.96% vs 83.33%).

Development of Lung Ultrasound Quantitative Approaches and Automatic Semi-Quantitative Strategies: In Silico, In Vitro, and Clinical Studies · IRIS - UNITN - prod

Lost to a baseline

Original SAM achieved higher recall (0.53 vs 0.41) than the fine-tuned cropping approach on bare cropland, though driven by over-segmentation/false positives.

A Unified Framework for Advancing Soil Erosion and Flood Assessment Through Deep Learning and Process-Based Modeling · Publikationssystem UB Tuebingen

Lost to a baseline

Direct 2D segmentation without priors (Deep-Single) had higher surface error outliers on noisy B-scans compared to the graph-based Aura baseline.

RETINAL OCT IMAGE ANALYSIS USING DEEP LEARNING · JScholarship

Lost to a baseline

ResNet with ASPP module on radargram segmentation achieved overall accuracy slightly lower than literature SVM baselines

Advanced methods for simulation-based performance assessment and analysis of radar sounder data · IRIS - UNITN - prod

Lost to a baseline

Epistemic uncertainty data selection (Hausdorff 4.15 mm) lost to baseline fully supervised (Hausdorff 3.97 mm) in MRI segmentation

Improving deep-learning segmentation performance in 3D neuroimaging with minimal manual annotations · Oxford

Tried and failed

Double deep image prior unsupervised segmentation applied to pathological tissue segmentation. Outcome: worse than baseline. Reason: yielded very low overlap and aggregated Jaccard index unsuitable for complex tissue structures

Recognition, retrieval, and harmonisation for multicentre clinical data analysis · Imperial

Lost to a baseline

Adv-trained ResDSN Coarse 3D medical image segmentation performance on clean data (79.09% accuracy/Dice) lost to baseline ResDSN Coarse (87.84%).

TOWARDS DEEP LEARNING ROBUSTNESS FOR COMPUTER VISION IN THE REAL WORLD · JScholarship

Lost to a baseline

Custom-trained ilastik pixel classifier underperformed generic pre-trained deep learning models (Cellpose and StarDist) on nuclear segmentation of crowded/touching cells

Exploring tumour heterogeneity and responses to therapy using single-cell resolved microscopy · Imperial

Lost to a baseline

DeepLabv3 achieved lower accuracy metrics (F1 score 0.751 vs 0.809) on visible image segmentation compared to FPN despite having over 2x the parameter count.

Condition Assessment of Civil Infrastructure and Materials Using Deep Learning · Virginia Tech

Lost to a baseline

SegResNet demonstrated slightly higher sensitivity in overall liver tumor segmentation than SmoothSegNet despite having lower accuracy and Dice score due to over-segmentation.

Knowledge-Informed Weakly-Supervised Deep Learning Models for Cancer Applications · Georgia Tech

Lost to a baseline

SDXL-Turbo achieved lower semantic segmentation performance (36.99% overall Dice-AUC) compared to full SDXL (38.27% overall Dice-AUC).

METHODOLOGICAL ADVANCES IN REAL-WORLD COMPUTER VISION: ROBUSTNESS, EFFICIENCY AND HUMAN PREFERENCE ALIGNMENT · Cornell

Considered and rejected

Considered and rejected: Rejected using a dual 3D segmentation head ensemble for domain generalization in LiDOG as it performed worse than the auxiliary 2D BEV projection decoder (30.07 vs 44.18 mIoU).

Handling Domain Shift in 3D Point Cloud Perception · IRIS - UNITN - prod

Tried and failed

real-time one-stage instance segmentation applied to sidewalk semantic segmentation. Reason: produced inaccurate segmentations and exhibited high sensitivity to input image resolution without specific backbone tuning

Robust and precise sidewalk detection with ensemble learning: Enhancing road safety and facilitating curb space management · Iowa State

Tried and failed

single-stage real-time instance segmentation applied to onboard robotic vision. Outcome: worse than baseline. Reason: produced lower-quality segmentation masks compared to an optimised two-stage instance segmentation model

Autonomous exploration and object reconstruction with an MAV · Imperial

Tried and failed

fine-tuned real-time instance segmentation models applied to anatomical structures in surgical video. Outcome: did not generalise. Reason: model suffered extreme false positive rates and low specificity on unseen test video frames

Artificial Intelligence to Minimize Intraoperative Complications: Using Segmentation for Surgical Anatomy Recognition · Harvard

Lost to a baseline

U-Net beat U-KAN on Sentinel-1 crop field segmentation Recall (85.56% vs 77.50%)

Spatio-Temporal Machine Learning for Ecology and Crisis Management · IRIS - POLITO - prod

Tried and failed

hardcoding domain-specific anatomical constraints into neural network applied to image segmentation models. Outcome: worse than baseline. Reason: rigid boundary rules reduced segmentation accuracy and introduced software instability

A study of the relationship between monocotyledonous plant anatomy and water · University of Nottingham Repository

Lost to a baseline

CoBEVT achieved 60.4% mIoU on vehicle BEV segmentation and 63.0% on drivable area, beating CoBEVFusion (59.5% and 61.7% mIoU)

Enhancing Perception for Autonomous Vehicles · Queens University Institutional Repository

Lost to a baseline

K-means achieved higher Precision in traffic frame segmentation (0.89 vs 0.87 for HHGATSD)

Novel optimization methods and model for improving sustainability and efficiency in last-mile logistics · DeustoTeka

Cross-dataset distribution shifts and synthetic data fail to generalize to real clinical and field imagery

13 theses · 9 institutions

Segmentation networks trained on synthetic objects, specific scanner resolutions, or narrow geographic datasets experienced severe performance drops and merge errors when transferred to real target domains. Generative augmentations and heuristic synthetic adjustments failed to bridge these protocol gaps and occasionally created artifacts that further degraded test accuracy.

Tried and failed

U-Net segmentation cross-dataset transfer applied to histology image segmentation. Outcome: did not generalise. Reason: Domain shift from differing staining and imaging protocols reduced accuracy to simple thresholding baseline

Computational modelling of diffusion magnetic resonance imaging based on cardiac histology · Imperial

Tried and failed

direct cross-resolution neural network deployment applied to medical image segmentation. Outcome: did not generalise. Reason: resolution and signal-to-noise mismatch between high-resolution multi-average training data and lower-resolution single acquisitions

Deep learning-based analysis of multiple sclerosis lesions with high and ultra-high field MRI · EPFL

Tried and failed

generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data

Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial

Tried and failed

off-the-shelf object segmentation without domain-specific fine-tuning applied to surgical video anatomy detection. Outcome: did not generalise. Reason: models pre-trained on common natural images fail to recognize specialized surgical anatomy

XMARCUS: A Pathway Towards Remote Robotic Surgery Coaching · Virginia Tech

Tried and failed

direct transfer of synthetic-trained 3D CNNs applied to real 3D volumetric image segmentation. Outcome: did not generalise. Reason: domain gap between synthetic training data and noisy real data caused discontinuous predictions and low recall

Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech

Tried and failed

direct cross-dataset model transfer without domain adaptation applied to electron microscopy 3D segmentation. Outcome: did not generalise. Reason: acquisition variations across datasets caused severe undersegmentation and high merge errors

Grounded: Inference via Local Signals & Learned Representations by Organic & Artificial Systems · Harvard

Tried and failed

single-modality deep learning auto-segmentation for error screening applied to radiotherapy target volume contour validation. Outcome: did not generalise. Reason: CT-only segmentation produced an unacceptably high false positive rate for anomaly detection

INTEGRATION OF ARTIFICIAL INTELLIGENCE IN QUALITY ASSURANCE FOR HEAD AND NECK CANCER RADIOTHERAPY CLINICAL TRIALS · Penn

Tried and failed

mixup data augmentation applied to medical image segmentation across population shifts. Outcome: worse than baseline. Reason: linear interpolation created unrealistic images that degraded generalization across pathological groups

Improving the domain generalization and robustness of neural networks for medical imaging · Imperial

Tried and failed

adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance

Improving the domain generalization and robustness of neural networks for medical imaging · Imperial

Lost to a baseline

Heuristic TEA degraded performance for DeepMedic on cross-site prostate segmentation (Site B) compared to no TEA (67.4% vs 71.7% DSC with learned class-specific TRA).

Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial

Lost to a baseline

Models trained with synthetic long cells showed decreased segmentation performance on in-distribution short cells compared to models trained exclusively on short cells.

Bacterial Deepfakes: Generating Synthetic Microscopy Data to Improve Adaptability of Deep Learning-Based Segmentation Models · ResearchWorks

Considered and rejected

Considered and rejected: Rejected fully simulated datasets as sole ground truth for defect segmentation because synthetic defects fail to capture real shape complexity, boundary definition, and noise interactions.

End-to-End Artificial Intelligence-Based Pipeline for Quantitative Reporting of Lung V/Q Scintigraphy—VQ-SPRINT: Segmentation, Pseudo-planar Generation, and Registration INtegration Tool · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Rejected supervised machine-learning organ segmentation trained on external datasets (CT-ORG, AAPM Thoracic) due to domain shifts and severe disease exclusion in prior benchmarks, choosing unsupervised morphological processing instead.

Towards Fully Automated Interpretation of Volumetric Medical Images with Deep Learning · DukeSpace

Tried and failed

semantic segmentation model trained on urban driving applied to aerial building facade and ground segmentation. Outcome: did not generalise. Reason: domain shift between urban street-level perspective and aerial or campus imagery degraded feature recognition

Aerial Cadastral and Flood Assessment for Disaster Risk Management in Appalachia · Virginia Tech

Tried and failed

synthetic object injection for anomaly segmentation applied to road scene obstacle detection. Outcome: did not generalise. Reason: biases model toward large nearby objects, missing small distant obstacles and causing false positives

Detecting Anomalies and Obstacles in Road Scenes · EPFL

High computational complexity and inference latency prevent real-time deployment

14 theses · 12 institutions

Volumetric 3D networks, dense vision transformers, multi-modal fusion heads, and connected component post-processing created excessive runtime delays and hardware strain. Because of these processing bottlenecks, researchers rejected pixel-wise architectures or selected lighter models to maintain acceptable frame rates in robotics and clinical imaging.

Tried and failed

pointwise Gaussian process regression applied to point cloud segmentation. Outcome: too slow. Reason: computational cost was too high for real-time processing and struggled with occlusion

Online vehicle trajectory extraction based on LiDAR data · Texas Tech

Tried and failed

star-convex object detection for tracking and segmentation applied to large-scale high-throughput image sequences. Outcome: too slow. Reason: per-frame inference caused excessive overall processing times on long high-throughput video sequences

Towards systems biophotonics in microscopy, medicine, and robotics · Georgia Tech

Tried and failed

connected-component analysis post-processing applied to neural network image segmentation masks. Outcome: too slow. Reason: Increased inference time tenfold without providing any meaningful improvement in segmentation accuracy

Automated cardiac segmentation pipeline and motion analysis in rodent models of pulmonary hypertension · Imperial

Lost to a baseline

Single-image U-Net and DeepLabV3 segmentation models significantly outperformed all fusion architectures (CMNeXt, MMSFormer, StitchFusion) in inference latency (5.0–5.2 ms vs 18.2–83.9 ms).

Autonomous System for Identifying and Capturing Floating Waste · Georgia Tech

Lost to a baseline

The proposed hybrid segmentation algorithm required double the processing time of LSA and OHRH baseline methods due to dynamic per-loop metric updates

An improved segmentation and classification method for building extraction from RGB images using GEOBIA framework · Queens University Institutional Repository

Lost to a baseline

VGG-16 semantic segmentation achieved only 2 fps with frequent misidentifications and lost to manual contact determination

Characterisation of a Novel Bioadhesive Used by the Ctenophore Pleurobrachia pileus · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected CNN-based face segmentation models for the real-time wearable implementation due to high computational complexity.

Augmented reality for the sight impaired · Oxford

Considered and rejected

Considered and rejected: Direct 3D deep learning segmentation rejected due to excessive computational expense and hardware demands for real-time ultrasound imaging, opting for 2D slice-by-slice U-Net.

Forward-Viewing Ultrasound Guidance of a Robotically-Steered Guidewire for Peripheral Interventions · Georgia Tech

Considered and rejected

Considered and rejected: Rejected instance object segmentation (e.g., Mask R-CNN/UNet) in favor of bounding-box object detection (YOLOv4) due to inference latency and higher data annotation burden.

Vision-Based Force Planning and Voice-Based Human-Machine Interface of an Assistive Robotic Exoskeleton Glove for Brachial Plexus Injuries · Virginia Tech

Considered and rejected

Considered and rejected: Rejected pixel-wise semantic/instance segmentation and GrabCut for object extraction due to high computational overhead in robotics.

Scene understanding via scene graph Szenenverständnis mittels Szenengraph · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected 3D CNNs and R-CNN architectures for dynamic frame segmentation because driving video frames lack strict temporal dependence and U-Net was more computationally efficient.

Methods for Classifying Driver Engagement in Autonomous Vehicles Using Physiological Sensors · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Decided against transformer architectures for semantic segmentation due to extreme computational expense and large data requirements compared to CNNs.

Integration of machine learning for enhanced digital rock physics workflows · UT Austin

Considered and rejected

Considered and rejected: Decided against semantic segmentation architectures (pixel-wise segmentation) due to high computational cost and because area composition estimation does not require precise item boundary segmentation.

Development of a method to classify and analyse the composition of mixed waste materials in real-time. · Cranfield

Considered and rejected

Considered and rejected: Rejected standard 3D CNNs from scratch for organ segmentation due to training instability, lack of 3D pre-trained weights, and high computational cost

Towards Robust Deep Learning for Medical Image Analysis · JScholarship

Semantic segmentation struggles with overlapping boundaries and topological continuity

13 theses · 11 institutions

Pixel-wise semantic classification frequently merged adjacent or intersecting instances into single undifferentiated masks and produced jagged or fragmented lines. Standard loss functions and uncropped single-stage networks failed to maintain instance boundaries, fine curvilinear structures, and topological continuity.

Tried and failed

centerline and distance-based segmentation loss functions applied to thin curvilinear structure segmentation. Outcome: worse than baseline. Reason: degraded segmentation performance compared to standard Dice loss

Pavement Crack Segmentation with Dense Local Geometry Features and Boundary Enhancement Loss · Georgia Tech

Lost to a baseline

Single-class semantic segmentation achieved 85% Mean IoU, whereas two-class training achieved lower Mean IoU (80%) due to irregular, jagged predicted boundaries on metric bars.

Integration and classification of spatial data for 3D modelling and monitoring of built heritage · IRIS - POLITO - prod

Lost to a baseline

Mesmer whole-cell segmentations produced more rounded boundaries, under-capturing irregular membrane/cytoplasmic extensions compared to traditional watershed segmentation.

Computational Methods to Process and Analyze High-dimensional Imaging Mass Cytometry Datasets in Pathological Tissue Samples · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected semantic segmentation (pixel-wise shape classification) because precise item borders are unneeded for mass/area composition estimation and fails on cluttered waste

Development of a method to classify and analyse the composition of mixed waste materials in real-time · Cranfield

Considered and rejected

Considered and rejected: Rejected relying purely on monocular semantic segmentations without instance labels because overlapping obstacles merge into single impassable blocks.

Perception Enabled Planning for Autonomous Systems · Cornell

Considered and rejected

Considered and rejected: Semantic segmentation via standard U-Net rejected due to inability to differentiate multiple or intersecting instruments without extensive post-processing

Detection and 3D Localization of Surgical Instruments for Image-Guided Surgery · JScholarship

Considered and rejected

Considered and rejected: Rejected text annotations from the WiSe dataset because its semantic segmentation masks lacked instance granularity (words/lines) needed for tracking.

Lecture Video Summarization by Detection and Representation of Content · DSpace at SUNY Buffalo

Considered and rejected

Considered and rejected: Directly using semantic segmentation 'otherprop' class alone for object extraction without plane detection (rejected due to ScanNet under-segmentation causing poor object boundary precision)

Object change detection for autonomous indoor robots in open-world settings · DSpace-CRIS at TU Wien

Considered and rejected

Considered and rejected: Single-stage semantic segmentation network on uncropped images was rejected due to unacceptable false positive rates on internal organelles and out-of-focus boundaries.

Quantifying RTK Signal Transduction Processes Across The Plasma Membrane · JScholarship

Tried and failed

semantic segmentation of spatially adjacent overlapping features applied to medical image segmentation of multiple findings. Outcome: did not generalise. Reason: models confuse boundary distinctions when multiple distinct localized patterns co-occur in close spatial proximity

AI Systems for Understanding and Grounding Radiology Reports · Harvard

Tried and failed

pixel-wise loss functions for continuous line segmentation applied to crack detection in structural images. Reason: standard pixel losses fail to preserve topological continuity, producing fragmented segmentations

Automated post-earthquake damage assessment of stone masonry buildings integrating machine learning, computer vision, and physics-based modeling · EPFL

Considered and rejected

Considered and rejected: Rejected stateful panoptic fusion for streaming sectors because it causes under-segmentation for objects spanning into upcoming unseen sectors.

On the Use of Vision and Range Data for Scene Understanding · JScholarship

Considered and rejected

Considered and rejected: Segmentation-only turning lane validation was rejected because it generated noisy, disconnected segmentation masks for valid lanes

Enriching Digital Maps with Aerial Imagery and GPS Data · MIT

Intensity and heuristic thresholding techniques fail in noisy and variable environments

11 theses · 9 institutions

Global, adaptive, and batch thresholding strategies proved ineffective when target objects were obscured by confounding hyper-intensities, uneven illumination, or low contrast. These rigid heuristic cutoffs caused extensive false negatives on subtle structures and misclassified background noise as foreground targets.

Tried and failed

intensity thresholding segmentation applied to anatomical structure segmentation in medical imaging. Outcome: worse than baseline. Reason: confounding hyper-intensities prevent accurate separation of target structures from background tissue

Deep Learning for Localizing and Segmenting Anatomies in Medical Imaging · Cornell

Tried and failed

Gaussian mixture model automatic thresholding applied to time-lapse microscopy image segmentation. Outcome: worse than baseline. Reason: Failed to accurately segment features across time-lapse frames compared to a static threshold

Precipitation Dynamics at the Solution-Solution Interface in Confined Geometries, and the Effects of Organics on Precipitate Evolution · Georgia Tech

Tried and failed

color thresholding and skeletonization applied to overlapping root system segmentation. Reason: heuristic segmentation failed to separate complex overlapping structures from background

The development and deployment of machine learning based high throughput phenotyping of soybean nodules and root system architecture traits for agronomic, physiological and genetic · Iowa State

Tried and failed

simple intensity thresholding for segmentation applied to subcellular molecular cluster detection. Reason: missed low-intensity functional clusters and failed to resolve closely spaced assemblies

MECHANISMS OF TRANSCRIPTION FACTOR HUB FORMATION AND FUNCTION DURING EMBRYONIC DEVELOPMENT · Penn

Tried and failed

Automated intensity thresholding segmentation applied to subcellular focal adhesion quantification. Outcome: worse than baseline. Reason: Batch thresholding masks failed to detect features accurately, producing high false-negative rates versus manual quantification

Mechanotaxis and mechanoresistance in cancer · Imperial

Tried and failed

fixed-threshold normalized difference index segmentation applied to surface water mapping across regions. Outcome: did not generalise. Reason: a single threshold caused overestimation at some sites while completely omitting target features at others

Observing the Seasonal Evolution of Supraglacial Ponds in High Mountain Asia: A Supervised Classification Approach · Cambridge

Tried and failed

heuristic contrast adjustment and thresholding preprocessing applied to grayscale electron microscopy segmentation. Outcome: worse than baseline. Reason: degraded image quality and disrupted deep learning feature extraction

Deep Learning Approach for Cell Nuclear Pore Detection and Quantification over High Resolution 3D Data · Virginia Tech

Considered and rejected

Considered and rejected: Rejected global and adaptive threshold-based segmentation baselines (e.g., Otsu's method) because histogram-based thresholding inherently fails to distinguish root from non-root foreground noise.

ITErRoot: High Throughput Segmentation of 2-Dimensional Root System Architecture · HARVEST

Considered and rejected

Considered and rejected: Thresholding pixel values alone for laser line extraction (lacked repeatable precision required for structured light sensing, necessitating regional segmentation / binary image conversion).

Volumetric flow monitoring of biomass through an industrial grinder · Iowa State

Considered and rejected

Considered and rejected: Rejected using global thresholding segmentation due to poor multiclass classification compared to watershed and Bayesian/machine-learning approaches

Wettability characterisation of sandstone and carbonate rocks using X-ray micro-CT imaging · Imperial

Considered and rejected

Considered and rejected: Rejected purely automated thresholding (Otsu/RenyiEntropy) for final segmentation because overestimation and rigidity included background and cell bodies instead of true synaptic puncta.

Mapping of excitatory/inhibitory synaptic ratio in the lateral geniculate and mediodorsal thalamic nuclei · OpenBU

Lack of spatial, multimodal, or temporal context degrades segmentation accuracy

9 theses · 7 institutions

Omitting elevation models, historical viewpoints, high-contrast modalities, or bidirectional volumetric context deprived models of critical structural priors. These localized or unimodal approaches suffered from perspective ambiguities, high false positive rates, and truncated peripheral context.

Tried and failed

omitting digital elevation data in visual segmentation applied to aerial imagery semantic segmentation. Outcome: worse than baseline. Reason: lacks crucial topographic priors, causing false positive segmentations above natural elevation thresholds

Monitoring and understanding treeline dynamics in the Swiss Alps from 80 years of aerial imagery · EPFL

Tried and failed

omitting specialized high-contrast input modalities in segmentation applied to cortical lesion detection. Outcome: worse than baseline. Reason: standard input sequences lacked sufficient contrast to resolve subtle cortical boundaries without specialized imaging sequences

Deep learning-based analysis of multiple sclerosis lesions with high and ultra-high field MRI · EPFL

Tried and failed

fixed margin bounding box cropping and resizing applied to semantic image segmentation. Outcome: worse than baseline. Reason: fixed margin crops either truncate peripheral features or introduce irrelevant distracting background context

Deep face tracking and parsing in the wild · Imperial

Considered and rejected

Considered and rejected: Single-scan batch inference for online segmentation was rejected due to lower accuracy from lacking historical viewpoint context.

3D Segmentation and Damage Analysis from Robotic Scans of Disaster Sites · Georgia Tech

Tried and failed

fine-tuning 2D models for volumetric data applied to 3D medical image segmentation. Outcome: worse than baseline. Reason: lack of bidirectional volumetric context across slices

AI Systems for Understanding and Grounding Radiology Reports · Harvard

Considered and rejected

Considered and rejected: Rejected standard CNN local feature extraction alone for overlap segmentation because it loses global context, resulting in subpar segmentation

Cervical Cell Separation using Deep Learning Techniques · Carleton University Institutional Repository

Tried and failed

pure vision transformer with smaller patch size applied to 3D spatiotemporal microstructural degradation prediction. Outcome: worse than baseline. Reason: Lacks inductive spatial biases of CNNs for complex 3D volumetric sequences

TransVNet: Predicting bone degradation using ViT and virtual dataset of cellular microstructures · Iowa State

Tried and failed

deep feature classification on unmasked raw images applied to pairwise visual disambiguation. Outcome: worse than baseline. Reason: pretrained visual features failed to distinguish ambiguous relationships without explicit foreground segmentation masks

Pushing the Boundaries of 3D Spatial Understanding · Cornell

Tried and failed

unimodal visual feature segmentation applied to spatial accident risk prediction. Outcome: worse than baseline. Reason: spatial dispersion and high false positive rates without auxiliary structural and dynamic context

Enhancing autonomous vehicle decision-making through scenario-based traffic rule integration · Imperial

Automated segmentation pipelines fall short of manual or semi-automated expert workflows

8 theses · 8 institutions

Fully automated deep learning systems exhibited inconsistent contouring failures, lesion spiculation biases, and error propagation compared to expert manual masking. These models struggled to encode subjective clinical expertise and institutional protocol differences, forcing researchers to retain manual intervention.

Lost to a baseline

Active contour segmentation was more biased than commercial semi-automatic segmentation on low-spiculation lesions at 2.5 mm slice thickness

Truth-based Radiomics for Prediction of Lung Cancer Prognosis · DukeSpace

Tried and failed

deep learning auto-segmentation for automated quality assurance applied to clinical target volume delineation. Outcome: did not generalise. Reason: Protocol variability, error propagation from upstream inputs, and inability to encode subjective clinical judgment.

INTEGRATION OF ARTIFICIAL INTELLIGENCE IN QUALITY ASSURANCE FOR HEAD AND NECK CANCER RADIOTHERAPY CLINICAL TRIALS · Penn

Lost to a baseline

Def-RgDL SMG segmentation (DSC 53.8%) was inferior to DL alone without guidance (DSC 63.5%) in a case where registration propagated a false vessel contour.

Efficient and Intelligent Radiotherapy Planning and Adaptation · DSpace at UTSWMED

Lost to a baseline

Manual clover dry matter fraction prediction from Hansen et al. [101] achieved 7.8% standard deviation, surpassing the thesis's ERFNet grass/legumes segmentation in complex outdoor lighting

Mobile vision system for estimation of soil and plant properties Mobiles Bildverarbeitungssystem für die Schätzung von Boden- und Pflanzeneigenschaften · DSpace-CRIS at TU Wien

Lost to a baseline

Manual full-organ segmentation outperformed the simpler, faster ROI-based baseline method in test-retest repeatability across nearly all sequences.

Multiparametrische Magnetresonanztomographie der Nieren in einer prospektiven Probanden- und Patientenstudie: Retest-Reliabilität funktioneller Gewebeparameter und Evaluation einer Deep Learning-basierten Organsegmentierung · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Excluded automated computer-vision machine learning classification in favor of manual human coding due to current ML limitations in recognizing complex semantic scene context.

A Collective Sense of Place and the Image of the City @ Urban Public Spaces: Analysis on People's Perception of User-Generated Image Content and Hashtags on Instagram · Virginia Tech

Lost to a baseline

Automated nnU-NET SAT segmentation suffered complete or partial failure in 13% of datasets, requiring manual correction compared to reliable manual masking.

Magnetic Resonance Imaging and Spectroscopy Methods for Studying Obesity: Applications for Bariatric Surgery · University of Nottingham Repository

Lost to a baseline

CardioINSIGHT automated CT segmentation inconsistently failed compared to manual multi-slice 2D contour interpolation.

Feasibility of improving risk stratification in the inherited cardiac conditions · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Benchmark TEDM against foundation models like SAM for semi-supervised medical image segmentation under domain shift and out-of-distribution conditions. Blocker: None

Label-efficient medical image segmentation: the benefits and limitations of semi-supervised generative models · Imperial

Left open

Evaluate and improve unsupervised instance segmentation failure cases caused by background clutter and edge ambiguity. Blocker: None

Enhancing Unsupervised Instance Segmentation with Exemplars · Carleton University Institutional Repository

Left open

Develop advanced post-processing techniques specifically tailored for gradient-based weakly-supervised semantic segmentation methods. Blocker: The thesis provides no specific design, algorithm, or mathematical formulation for the intended post-processing techniques.

Learning without Expert Labels for Multimodal Data · Virginia Tech

Left open

Implement and evaluate alternative classification algorithms beyond K-NN for historical document character classification in the press variant identification pipeline. Blocker: Lack of the specific historical print dataset and upstream segmentation pipeline output used in the thesis.

Automated identification of press variants in old documents · De Montfort Open Research Archive (DORA)

Left open

Develop physiology-aware machine learning models to reduce segmentation errors in multiplexed tissue images. Blocker: Lacks specific definition, architecture, or formulation for what constitutes 'physiology-aware' models

Single-cell Methods and Spatial Analysis for Highly Multiplexed Tissue Images · Harvard

Left open

Develop generalist deep learning segmentation architectures beyond U-Net to segment multiple cellular compartments across imaging modalities. Blocker: Lacks specific architectural design requirements, target benchmark datasets, and evaluation metrics

Grounded: Inference via Local Signals & Learned Representations by Organic & Artificial Systems · Harvard

Left open

Extend weakly supervised satellite semantic segmentation to self-supervised frameworks with automatic error detection for refining low-resolution labels. Blocker: None

Weak-Supervised Deep Learning Methods for the Analysis of Multi-Source Satellite Remote Sensing Images · IRIS - UNITN - prod

Left open

Combine pixel-wise, region-wise, and fuzzy boundary-wise loss functions and evaluate segmentation performance across standard metrics. Blocker: None

Incorporating fuzzy-based methods to deep learning models for semantic segmentation · University of Nottingham Repository

Left open

Develop a transfer-learning module to generalize brain metastasis segmentation across varied institutional MRI acquisition protocols. Blocker: Access to diverse multi-institutional brain metastasis MRI datasets with expert segmentations

Advancing Radiotherapy Treatment Through Artificial Intelligence-Driven Approaches · DSpace at UTSWMED

Left open

Adapt the PrinCut-Auto segmentation framework to work across multiple imaging modalities such as MRI, two-photon, and confocal microscopy. Blocker: Lacks specific algorithmic techniques or concrete benchmarks for achieving generalizability across modalities

Animal Internal Motion Analysis with Unsupervised Machine Learning Methods · Virginia Tech

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.