Chapter Four · failure evidence
What Semantic & Object Recognition got wrong, from 33 dissertations
Across these thesis records, semantic segmentation and object recognition methods frequently encounter obstacles including severe computational overhead, extensive manual labeling demands, and architectural performance degradations. Researchers also face significant performance drops caused by distribution shifts, inappropriate input sampling or loss functions, and an inability to separate adjacent or overlapping object boundaries. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Dense pixel-wise segmentation models cause excessive computational overhead and high latency
Deploying pixel-level semantic segmentation and vision transformer architectures creates substantial computational bottlenecks on embedded devices and high-resolution aerial datasets. As a result, systems struggle to maintain real-time frame rates and researchers often abandon dense segmentation in favor of lighter bounding-box centroids or manual methods.
Considered and rejected
Considered and rejected: Semantic segmentation was dropped in favor of object detection bounding-box centroids for downstream goal generation due to the computational overhead and latency of multi-channel segmentation fusion models on embedded hardware.
Autonomous System for Identifying and Capturing Floating Waste · Georgia Tech
Lost to a baseline
VGG-16 semantic segmentation achieved only 2 fps with frequent misidentifications and lost to manual contact determination
Characterisation of a Novel Bioadhesive Used by the Ctenophore Pleurobrachia pileus · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected pixel-wise semantic/instance segmentation and GrabCut for object extraction due to high computational overhead in robotics.
Scene understanding via scene graph Szenenverständnis mittels Szenengraph · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Pixel-level semantic segmentation of aerial imagery, rejected due to extreme computational overhead on ultra-high-resolution images and variable scale/resolution distortions across oblique aerial perspectives.
Considered and rejected
Considered and rejected: Decided against semantic segmentation architectures (pixel-wise segmentation) due to high computational cost and because area composition estimation does not require precise item boundary segmentation.
Development of a method to classify and analyse the composition of mixed waste materials in real-time. · Cranfield
Considered and rejected
Considered and rejected: Decided against transformer architectures for semantic segmentation due to extreme computational expense and large data requirements compared to CNNs.
Integration of machine learning for enhanced digital rock physics workflows · UT Austin
Suboptimal input cropping, context selection, and background sampling harm segmentation accuracy
Fixed margin bounding box cropping and resizing either cuts off peripheral features or incorporates misleading background context. Additionally, omitting topographic elevation priors, relying on global shape cues, or using random background sampling degrades boundary classification and produces false positive errors.
Tried and failed
omitting digital elevation data in visual segmentation applied to aerial imagery semantic segmentation. Outcome: worse than baseline. Reason: lacks crucial topographic priors, causing false positive segmentations above natural elevation thresholds
Monitoring and understanding treeline dynamics in the Swiss Alps from 80 years of aerial imagery · EPFL
Tried and failed
fixed margin bounding box cropping and resizing applied to semantic image segmentation. Outcome: worse than baseline. Reason: fixed margin crops either truncate peripheral features or introduce irrelevant distracting background context
Deep face tracking and parsing in the wild · Imperial
Tried and failed
Fully convolutional networks using global shape cues applied to satellite semantic segmentation. Outcome: no signal. Reason: Global shape cues failed to provide sufficient discriminative features for accurate boundary classification
Tried and failed
random background sampling for semantic segmentation applied to satellite imagery land cover classification. Outcome: worse than baseline. Reason: underperformed compared to boundary-targeted and feature-based sampling strategies
Deeper backbones and complex fusion architectures degrade model performance
Increasing convolutional depth in encoder-decoder networks harms performance across all evaluation metrics compared to shallower baselines, while one-stage instance segmentation exhibits extreme sensitivity to input resolution. Complex graph fusion networks also incur severe graph generation overhead and class fragmentation, leading practitioners to prefer alternative models like Unet++.
Tried and failed
deeper convolutional backbone in encoder-decoder network applied to semantic image segmentation. Outcome: worse than baseline. Reason: deeper architecture degraded performance across all evaluation metrics compared to shallower baseline
3D Panoptic Segmentation with Unsupervised Clustering for Visual Perception in Autonomous Driving · Cranfield
Considered and rejected
Considered and rejected: Rejected DeepLabV3 and standard Unet architectures in favor of Unet++ for colon semantic segmentation due to superior IOU accuracy.
Fluorescence Imaging and Signal Quantification for Enhanced Surgical Precision and Real-Time Decision Making · Research Repository UCD
Tried and failed
real-time one-stage instance segmentation applied to sidewalk semantic segmentation. Reason: produced inaccurate segmentations and exhibited high sensitivity to input image resolution without specific backbone tuning
Considered and rejected
Considered and rejected: Decided against using the fusion CNN+GNN architecture for semantic vessel segmentation due to grid-graph generation overhead and class fragmentation
Application of machine learning methods to the analysis of x-ray angiography images · Oxford
Limited training data and manual ground truth annotation bottlenecks restrict model utility
Training semantic segmentation networks on small datasets causes accuracy to drop and fluctuate unpredictably, while manual mask creation for large image sets is prohibitively labor-intensive. Furthermore, adding synthetic training data for majority classes produces minimal performance returns when adequate real features are already present.
Tried and failed
U-Net semantic segmentation with small training sets applied to aerial agricultural object detection. Outcome: data insufficient. Reason: Training set size below 500 images caused accuracy to drop and fluctuate unpredictably
Application of unmanned aerial systems and deep learning in high-throughput plant phenotyping · Texas Tech
Considered and rejected
Considered and rejected: Rejected relying purely on automated deep-learning semantic segmentation for 15,988 canola images due to the prohibitive human cost of generating manual ground truth masks
Using procedural models to improve image-based plant phenotyping · HARVEST
Tried and failed
synthetic data augmentation for majority classes applied to semantic image segmentation. Outcome: no signal. Reason: sufficient real features were already present, yielding minimal to diminishing performance returns
Image synthesis with class-aware semantic diffusion models for surgical scene segmentation · Imperial
Considered and rejected
Considered and rejected: Rejected semantic segmentation for patch extraction due to extreme labeling labor requirements compared to SIFT/PHOW clustering.
Unravelling the Spatial Distribution of Individual-Level Abandoned Houses at Large Scale Using Open-Access Remotely Sensed Data · DSpace at SUNY Buffalo
Semantic segmentation fails to separate overlapping boundaries and distinct clustered instances
Models struggle with boundary delineation when distinct localized patterns co-occur in close spatial proximity or appear within cluttered environments. Relying on standard semantic segmentation without instance separation merges adjacent physical objects into a single continuous class, obstructing robot path planning.
Tried and failed
semantic segmentation of spatially adjacent overlapping features applied to medical image segmentation of multiple findings. Outcome: did not generalise. Reason: models confuse boundary distinctions when multiple distinct localized patterns co-occur in close spatial proximity
AI Systems for Understanding and Grounding Radiology Reports · Harvard
Tried and failed
semantic segmentation without instance segmentation for obstacle mapping applied to robot path planning in outdoor environments. Reason: distinct obstacles were merged into a single continuous class, blocking navigable paths
Perception Enabled Planning for Autonomous Systems · Cornell
Considered and rejected
Considered and rejected: Rejected semantic segmentation (pixel-wise shape classification) because precise item borders are unneeded for mass/area composition estimation and fails on cluttered waste
Development of a method to classify and analyse the composition of mixed waste materials in real-time · Cranfield
Models fail to generalize across domain shifts, novel objects, and complex scenes
Object recognition fails when dealing with novel or unquantifiable objects, and models trained on handheld imagery collapse when transferred to aerial drone viewpoints. Specialized training objectives like recall loss also fail on complex scenes with intra-class similarities while saturating simple scenes.
Considered and rejected
Considered and rejected: Rejected using object recognition labels (YOLOv5s) directly for grasp selection, because household items are unquantifiable and novel objects fail.
Tried and failed
recall loss objective applied to semantic image segmentation. Outcome: did not generalise. Reason: failed on highly complex scenes with intra-class similarities and saturated simple scenes
Robustness under distribution shifts in computer vision · Georgia Tech
Tried and failed
semantic segmentation models trained on handheld imagery applied to aerial bridge inspection imagery. Outcome: did not generalise. Reason: distribution shift between human-collected handheld images and unmanned aerial vehicle perspective and quality
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop automated registration, AI feature detection, and semantic HBIM modeling algorithms for architectural documentation point clouds. Blocker: The thesis provides only broad directions without technical specifications, datasets, or concrete algorithmic architectures
Optimizing The Digital Documentation Process by Integrating Laser Scanning, Hand Measurements, And Photographic Data · Texas Tech
Left open
Evaluate Dynamic Focal Loss on object detection and semantic segmentation benchmarks using standard computer vision datasets. Blocker: None
Advancing Industrial Monitoring with Deep Learning: From Bottleneck Analysis to Safety Compliance · Scholarship at UWindsor Institutional Repository
Left open
Apply the thesis's quality quantification framework to transformer-based semantic segmentation models such as ViT and Swin Transformer. Blocker: None
Incorporating fuzzy-based methods to deep learning models for semantic segmentation · University of Nottingham Repository
Left open
Weight unit-semantic pairs by an IoU threshold greater than 0.25 within GAN dissection unit count comparisons. Blocker: None
Visualizing Semiotics with Generative Adversarial Networks · Harvard
Left open
Implement semantic segmentation-based network architectures to enable pixel-wise anomaly localization for unmanned aerial systems imagery. Blocker: None
Adversarial Learning based framework for Anomaly Detection in the context of Unmanned Aerial Systems · Virginia Tech
Left open
Extend weakly supervised satellite semantic segmentation to self-supervised frameworks with automatic error detection for refining low-resolution labels. Blocker: None
Weak-Supervised Deep Learning Methods for the Analysis of Multi-Source Satellite Remote Sensing Images · IRIS - UNITN - prod
Left open
Implement semantic segmentation to extract complex iceberg morphologies from coastal time-lapse photography instead of using bounding rectangles. Blocker: Access to the specific coastal time-lapse photography datasets and ground-truth segmentation masks.
Iceberg Stability Investigations Using Machine Learning for Alaska and Greenland · Virginia Tech
Left open
Disentangle semantic learning from geographic memorization and spatial overfitting caused by dominant prototype words in multimodal land cover mapping models. Blocker: No specific approach, metrics, or evaluation protocol are specified for disentangling spatial memorization.
Multimodal Land Cover Mapping from Remote Sensing Imagery, Species Observations, and Language · EPFL
Left open
Evaluate vision transformer architectures for off-road terrain semantic segmentation using the proposed FCIoUv2 loss function. Blocker: None
Class-Based Approach to Mitigating Class Imbalance Issues in Vision-Based Off-Road Terrain Type Recognition · Carleton University Institutional Repository
Left open
Weight navigation graph candidate paths using long-range 2D semantic segmentation masks to penalize obstacle-dense regions. Blocker: None
Perception and Planning for Autonomous Navigation in Unstructured Environments · Cornell
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.