Chapter Four · failure evidence
What End-to-End Deep Learning got wrong, from 51 dissertations
Across diverse problem domains, end-to-end deep learning frequently suffers from severe sample inefficiency, optimization convergence issues, and poor generalization under distribution shifts. Consequently, practitioners often reject or replace end-to-end systems with modular pipelines, physics-informed architectures, and classical machine learning baselines that achieve superior performance. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
End-to-end reinforcement learning and policy models suffer from extreme sample inefficiency and long-horizon control failures
Direct sensor-to-control and policy models struggle with high sample complexity, intractable simulation requirements, and compounding imitation errors over extended horizons. They frequently converge to poor local optima, overfit simulated dynamics, and fail to transfer reliably to physical systems.
Tried and failed
direct end-to-end feedforward neural network policies applied to contact-rich robotic manipulation control. Outcome: worse than baseline. Reason: fell into local minima and overfit without structured mechanics-based constraints
Considered and rejected
Considered and rejected: Rejected model-free reinforcement learning algorithms (e.g., A3C) for end-to-end driving due to poor sample efficiency compared to conditional imitation learning
Towards Recognition as a Regularizer in Autonomous Driving · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected monolithic end-to-end vision-language-action policies for off-road autonomy due to lack of interpretability, vulnerability to compounding imitation learning errors over long horizons, and inability to train from data lacking low-level control labels
Off-Road Navigation Under Sensing Uncertainty · ResearchWorks
Considered and rejected
Considered and rejected: Rejected end-to-end DRL learning of guidance and low-level control directly from states to actuator forces/torques because policy overfits simulated dynamics and fails during sim-to-reality transfer.
Deep Reinforcement Learning as Guidance for Aerospace Robotics · Carleton University Institutional Repository
Tried and failed
end-to-end deep reinforcement learning applied to multi-agent sensor trajectory planning. Outcome: data insufficient. Reason: mapping high-dimensional system states directly to control actions required intractable amounts of training simulation data
Mobile sensors management algorithms for environmental monitoring · UT Austin
Tried and failed
end-to-end deep reinforcement learning applied to long-horizon sequential control tasks. Outcome: did not converge. Reason: Controllers got trapped in poor local optima across various model capacities.
Considered and rejected
Considered and rejected: Rejected pure end-to-end policy learning (RL/BC), because credit assignment and sample complexity fail over long robotics planning horizons.
Reasoning over Hierarchical Abstractions for Long-Horizon Planning in Robotics · MIT
Considered and rejected
Considered and rejected: Rejected direct end-to-end RL on raw tactile images due to sample inefficiency requiring 8 hours of physical robot wear and poor out-of-distribution generalization
Considered and rejected
Considered and rejected: Rejected end-to-end scaling of foundation models directly to robot control due to high inference latency, real-time control constraints, and data bottlenecks.
Physical AI via Hierarchical Decision Processes · JScholarship
Considered and rejected
Considered and rejected: Rejected end-to-end continuous action RL for redundant manipulator priority control in favor of discrete priority permutation selection to fix sample inefficiency
Adaptive Control in Robotics via Deep Reinforcement Learning: From Autonomous Task Management to Zero-Shot Human Assistance · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected end-to-end sensor-to-control models (e.g. imitation/reinforcement learning) because lack of structural modularity prevents reuse and requires excessive data.
Neural scene representations for dense-semantic SLAM · Imperial
Considered and rejected
Considered and rejected: End-to-end learning directly from raw camera frames was rejected to avoid sim-to-real transfer failure, maintain multi-UAV generalizability, and improve sample efficiency.
Energy-aware 3D path planning via Reinforcement Learning for aerial object detection and mapping · Iowa State
Considered and rejected
Considered and rejected: Rejected end-to-end neural policy learning for multimodal demonstration understanding in favor of modular neuro-symbolic program synthesis
Scaling Human Supervision for Robotic Manipulation · ResearchWorks
Modular pipelines and simpler statistical baselines consistently outperform end-to-end models
End-to-end architectures frequently underperform classical algorithms such as gradient-boosted decision trees, random forests, linear regression, and modular pipelines with differentiable components. In multiple evaluation settings, end-to-end models yield higher prediction error, worse convergence rates, or lower classification accuracy than simpler baselines.
Tried and failed
end-to-end deep neural networks applied to markerless relative pose estimation. Outcome: worse than baseline. Reason: produced unacceptably high relative orientation and yaw errors
Distributed Predictive Formation Control of Autonomous Rotary-Wing Micro Aerial Vehicles · EPFL
Tried and failed
end-to-end adaptive decoder training applied to neural signal decoding for gait. Outcome: worse than baseline. Reason: end-to-end optimization reduced decoding performance compared to modular training with differentiable components
TOWARDS EFFICIENT AND SCALABLE MACHINE LEARNING FOR FUTURE NEURAL INTERFACES · Cornell
Tried and failed
end-to-end decision-focused learning applied to contextual linear optimization. Outcome: worse than baseline. Reason: achieves slower regret convergence rates than estimate-then-optimize under low dual degeneracy
Lost to a baseline
Data-driven neural networks trained end-to-end on neural recordings were outperformed by task-driven neural network models.
Reverse engineering primate sensorimotor control with machine learning · EPFL
Lost to a baseline
When all task-relevant features are fully known a priori, feature-based Bayesian inference baselines (Coactive, FERL, DemPref) match or outperform end-to-end unstructured reward learning
Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots · Virginia Tech
Lost to a baseline
End-to-End deep superpixel models (Pipelines F-I, 77.3%-83.2% mIoU) performed worse than standalone RGB baseline (88.8% mIoU) on 512x512 gLitter
Deep Learning Using Tiny Domain-Specific Datasets with Sparse Labels · University of Nottingham Repository
Lost to a baseline
End-to-end Neural Network with GNN embeddings underperformed Random Forest and XGBoost baselines across all metrics (Accuracy 0.5836, F1 0.4873, AUC-ROC 0.5405, AUC-PR 0.4672)
Predicting psychological treatment dropout using graph neural networks · OpenBU
Considered and rejected
Considered and rejected: Rejected using the GNN as an end-to-end prediction model, using it instead as an embedding preprocessor because GBDTs outperform deep learning on tabular data
Predicting psychological treatment dropout using graph neural networks · OpenBU
Considered and rejected
Considered and rejected: Rejected end-to-end multimodal deep networks in favor of modular 2-stage feature extraction + gradient-boosted trees for usability and data efficiency.
Lost to a baseline
End-to-end softmax dense layers achieved lower accuracy (94.8%) compared to back-end PLDA scoring (96.0%) on 2-second x-vector voice quality identification.
Intra-speaker Voice Quality Recognition for Voice Therapy · Georgia Tech
Considered and rejected
Considered and rejected: Rejected training an end-to-end WER-optimizing neural network or complex fusion model for fusing MiDaS depth and MediaPipe z-axis to avoid overfitting and alignment complexity, choosing linear regression instead.
Sign Language Recognition Using Wearable Motion Sensing and Video Co-Training · DSpace at SUNY Buffalo
Bypassing intermediate representations and domain physics causes models to plateau and overfit
Mapping raw sensor inputs directly to final targets without intermediate physical states or domain constraints forces networks to require massive datasets and violates geometric inductive biases. Relying purely on end-to-end loss functions fails to guarantee constraint satisfaction and impedes interpretability.
Tried and failed
task-guided end-to-end image-to-image translation applied to cross-domain depth completion. Outcome: worse than baseline. Reason: downstream task feedback provided insufficient constraint on image translation quality
Domain adaptation for semantic and 3D tasks · Imperial
Considered and rejected
Considered and rejected: Rejected pure end-to-end deep learning from raw symbols/samples because it requires massive datasets for every modem operating mode, fails to generalize, and is difficult to debug compared to observable physical transductions.
Noise Metrology in Optical Communication Systems · Cambridge
Considered and rejected
Considered and rejected: Rejected pure end-to-end data-driven neural regression models in favor of task-driven transfer models due to poor out-of-distribution neural explainability.
Reverse engineering primate sensorimotor control with machine learning · EPFL
Tried and failed
end-to-end training without intermediate representation supervision applied to multi-stage 3D object detection. Outcome: did not generalise. Reason: optimizing downstream loss alone creates arbitrary representations that violate geometric inductive bias
Pseudo-LiDAR: Camera-based 3D object detection for autonomous driving · Cornell
Considered and rejected
Considered and rejected: Rejected end-to-end deep learning from raw multiplexed interferograms without crude phase estimation due to needing significantly more training data and epochs to learn wave propagation
Single-shot quantitative interferometric microscopy for imaging high-speed dynamics · MIT
Considered and rejected
Considered and rejected: Rejected 'end-to-end' machine learning models lacking rheological domain knowledge because imposing physical laws via loss functions increases training time without guaranteeing invariance or constraint satisfaction on test data
Mathematics, Methods, and Models for Data-Driven Rheology · MIT
Considered and rejected
Considered and rejected: Rejected end-to-end deep learning from raw fluoroscopic images to head parameters in favor of ML (ResNet-50 + XGBoost) due to limited dataset size (~hundreds of images).
A Method to Determine Patient Eye-Lens Dose During Fluoroscopically-Guided Neuro-Interventional Procedures · DSpace at SUNY Buffalo
Tried and failed
direct end-to-end regression without intermediate state representations applied to composite failure prediction. Outcome: did not generalise. Reason: direct input-output mapping lacks sufficient intermediate physical information compared to indirect prediction, causing performance to plateau
Machine learning for predictive virtual testing of composite airframes · Imperial
Considered and rejected
Considered and rejected: End-to-end raw time-series machine learning models were rejected in favor of physically-interpretable low-dimensional feature pairs to prevent overfitting
Hemodynamics of Native and Bioprosthetic Aortic Valves: Insights from a Reduced Degree-of-Freedom Model · JScholarship
Considered and rejected
Considered and rejected: Rejected using end-to-end raw audio spectrograms without domain-specific feature engineering (as done in VGGVox) due to lack of interpretability.
INTERPRETABILITY FOR ARTIFICIAL INTELLIGENCE IN SPEAKER RECOGNITION TASKS · Calhoun
End-to-end networks fail to generalize under environment shifts and dynamic conditions
End-to-end models trained on fixed conditions suffer severe performance degradation, aliasing, and hallucinations when deployed in novel environments or under dynamic channel shifts. Without explicit intermediate feature masking or data consistency mechanisms, models overfit to their training distributions.
Tried and failed
end-to-end learning from raw images applied to perception failure prediction. Outcome: did not generalise. Reason: failed to generalize failure prediction across novel deployment environments without explicit intermediate feature masking
Introspective perception for mobile robots · UT Austin
Tried and failed
end-to-end differentiable optical-digital design applied to extended depth-of-field imaging. Outcome: did not generalise. Reason: training exclusively on planar targets prevented reconstruction of sharp edges across multiple scene depths
Programmable Optics for Computational Photography · Publikationssystem UB Tuebingen
Tried and failed
end-to-end communication training under fixed channel conditions applied to transmission across dynamic physical channels. Outcome: did not generalise. Reason: models overfit to specific channel conditions and degraded under mismatched dynamic channel environments
Deep learning enabled semantic communications with speech recognition and synthesis · Imperial
Considered and rejected
Considered and rejected: Rejected pure image-domain end-to-end deep learning networks due to lack of data consistency layers, leading to hallucinations and poor generalizability.
Optimizing reconstruction and segmentation of free-breathing whole-heart CMR to enable clinical implementation · Georgia Tech
Considered and rejected
Considered and rejected: Rejected end-to-end supervised deep learning inversion models for MRI due to severe degradation and aliasing artifacts under test-time sampling pattern or anatomy shifts.
Compressed sensing using generative models : theory and applications · UT Austin
Considered and rejected
Considered and rejected: Rejected end-to-end trained black-box CNNs for pupil phase retrieval because they require massive labeled datasets, fail outside training distributions, and lack physical interpretability.
Adaptive optics for corrections of phase and polarisation state aberrations in microscopes · Oxford
Considered and rejected
Considered and rejected: Rejected direct end-to-end learning of task failure probability p(f|z) from raw images due to extreme sample scarcity and severe overfitting in novel environments.
Introspective perception for mobile robots · UT Austin
Lost to a baseline
End-to-end models with GRU units performed worse than simple RNNs under state sequence auxiliary supervision when transition distribution shift exceeded 0.8.
COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS · Penn
Mathematical and structural obstacles impede end-to-end gradient backpropagation and convergence
Direct end-to-end optimization can completely stall when layers encounter non-differentiable operations like argmax selection or pure noise from random measurement matrices. Furthermore, gradient projection phenomena and lack of curriculum scheduling cause training to collapse or trap optimization in suboptimal states.
Tried and failed
direct end-to-end training without curriculum learning applied to continuous sequence-to-sequence recognition. Outcome: worse than baseline. Reason: learning full-length continuous sequences directly from scratch caused optimization difficulties and performance degradation
Deep audio-visual speech recognition · Imperial
Tried and failed
end-to-end learning with optimization problem layers applied to decision-making systems. Outcome: did not converge. Reason: gradient projection phenomenon impedes effective gradient backpropagation during training
Machine Learning in Decision-Making Systems: Fairness, Robustness, and Data Bias · EPFL
Considered and rejected
Considered and rejected: Rejected end-to-end trained models with random measurement matrices for each image because the network receives pure noise and fails to learn.
Compressed sensing using generative models : theory and applications · UT Austin
Tried and failed
differentiable greedy decoding via argmax selection applied to end-to-end speech and language pipelines. Reason: argmax probability selection and token mapping operations are inherently non-differentiable
Deep learning enabled semantic communications with speech recognition and synthesis · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Optimize hyperparameters end-to-end across the entire bidirectional CycleGAN domain adaptation network instead of tuning individual sub-networks separately. Blocker: Requires the proprietary exoskeleton sensor datasets and specific simulation-to-real training pipeline developed in the thesis.
Enabling Scalable, Versatile, and Robust Control for Robotic Exoskeletons · Georgia Tech
Left open
Incorporate end-to-end deep learning tracking architectures without separate detection heads for traffic signal operations. Blocker: None
Machine learning application powering automation of efficient traffic operation · Iowa State
Left open
Develop an end-to-end deep neural network that predicts 3D depth maps directly from fringe phase maps without post-processing reconstruction. Blocker: None
High speed 3D photomechanics testing via additional temporal sampling · Iowa State
Left open
Learn transport mappings directly via neural optimal transport to unify matching and posterior inference into an end-to-end process. Blocker: None
Informed machine learning models for advancing cardiac disease prognosis · EPFL
Left open
Train an end-to-end deep learning model to predict air pollution metrics directly from raw street view images. Blocker: None
Leveraging Street View and Remote Sensing Imagery to Enhance Air Quality Modeling through Computer Vision and Machine Learning · Virginia Tech
Left open
Train end-to-end deep learning models directly on crack pattern images using expanded experimental or numerical simulation datasets. Blocker: Requires further experimental testing apparatus or complex high-fidelity numerical simulation data not yet generated
Damage Assessment of Stone Masonry Piers Using Imaged Surface Cracks · EPFL
Left open
Combine explicit depth estimation with end-to-end learning for bird's-eye view 3D object detection and map prediction. Blocker: None
Learning Birds-Eye View Representations for Autonomous Driving · Cambridge
Left open
Extend the Implicit AutoEncoder to jointly learn trainable implicit representations end-to-end within the autoencoder architecture. Blocker: None
Representation learning for point cloud understanding · UT Austin
Left open
Develop fully automated end-to-end whole slide image analysis pipelines that generalize across clinical datasets without model drift over time. Blocker: The goal is a broad research direction without specific targets, datasets, or defined methodological approaches.
Tumor Profiling from Pathology Images using Deep Learning · Harvard
Left open
Evaluate end-to-end deep learning architectures across multimodal streams for learning-centered emotion classification. Blocker: No specific multimodal dataset or architecture is defined for this broad task
Metodología para la identificación de emociones en un ambiente educativo con aprendizaje computacional · Repositorio Institucional BUAP
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.