Chapter Four · failure evidence
What Convolutional Neural Networks got wrong, from 67 dissertations
Across the collected theses, convolutional neural networks frequently encountered difficulties when applied to non-spatial representations, temporal sequences, and small or shifted datasets. In numerous domains, these networks underperformed simpler linear models, suffered from architectural limitations, or were rejected due to excessive training complexity and operational overhead. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Flawed architectural configurations and input representations degrade convolutional network performance
Networks suffered substantial accuracy drops when trained on inappropriate input encodings, such as text without embedding layers, scrambled word order, or overly coarse pixel grids. Performance also deteriorated when architectures removed essential components like global pooling or attention, failed to concatenate multi-scale features, or used single-layer formulations.
Tried and failed
CNN with embedding layer applied to Arabic text classification. Outcome: worse than baseline. Reason: Standard convolutional architecture failed to capture semantic and syntactic patterns in Arabic thesis abstracts
Towards Logical Reasoning and Learning in Open and Dynamic Environments · Virginia Tech
Tried and failed
direct regression using convolutional neural network applied to spatial transfer function estimation from measurements. Outcome: worse than baseline. Reason: reduced low-dimensional morphological feature set lacked sufficient information for direct continuous mapping
Estimation of Hearing Aid Head Related Transfer Functions using Anthropometric Features · Harvard
Considered and rejected
Considered and rejected: Rejected shallow networks with only a single convolutional layer because they performed significantly worse on the F0 estimation task and produced poorer matches to human psychophysics than multi-convolutional-layer architectures.
Task-optimized models of human hearing link perception and neural coding · MIT
Considered and rejected
Considered and rejected: Rejected convolutional layer biases due to reduced network generalization capacity.
Advancing astronomical image deconvolution with deep learning · EPFL
Tried and failed
CNN without embedding layer applied to text classification. Outcome: worse than baseline. Reason: Lacked semantic representation of words prior to convolutional feature extraction
Improving the Accessibility of Arabic Electronic Theses and Dissertations (ETDs) with Metadata and Classification · Virginia Tech
Tried and failed
concatenating conditioning embeddings to raw inputs applied to image classification with convolutional networks. Outcome: worse than baseline. Reason: input-level concatenation degrades prospective risk compared to conditioning at deeper network layers
Tried and failed
one-hot CNN with randomized word order applied to text classification. Outcome: worse than baseline. Reason: destroying sequential order prevents convolutional filters from learning meaningful local n-gram patterns compared to simple bag-of-words
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Tried and failed
simplifying activation functions and removing attention layers applied to compact convolutional neural networks. Outcome: worse than baseline. Reason: naive removal of squeeze-and-excitation layers and substituting hard-swish with ReLU severely degraded classification accuracy
Efficient model adaptation and compression for edge intelligence · UT Austin
Tried and failed
removing global pooling in neural architecture search applied to evolving convolutional neural network topologies. Outcome: worse than baseline. Reason: evolved networks collapsed to using only pooling layers and failed to train effectively
Image Classification with Evolved Convolutional Neural Networks · Harvard
Tried and failed
convolutional neural network for multiscale spatial modeling applied to spatial mixture population property modeling. Outcome: did not generalise. Reason: biased toward learning shorter-scale spatial continuity while averaging out long-range continuity structures
Deep learning for spatial nonstationarity : evaluation, mitigation, and generation · UT Austin
Considered and rejected
Considered and rejected: Rejected standard deep CNN architectures without feature concatenation because deep convolutional layers gradually saturate and lose low-concentration metabolite information while shallow layers lack high-concentration metabolite details.
Considered and rejected
Considered and rejected: 2D Convolutional Neural Network for 4x6 NaI array classification (deprecated because 24 pixels were too coarse and low-multiplicity muons mimicked signal topology)
Studies of the Electron Neutrino Charged-current Interaction on $^{127}$I · DukeSpace
Convolutional models fail to outperform simpler linear models and classical baselines
Convolutional neural networks were repeatedly matched or beaten by simpler methods, including linear and quadratic discriminants, multinomial logistic regression, and standard feed-forward networks. Practitioners noted that the additional architectural complexity produced no noticeable accuracy gains while incurring higher computational demands.
Tried and failed
convolutional neural networks with small training sets applied to time-series classification. Outcome: worse than baseline. Reason: underperformed compared to linear and quadratic discriminants under sparse data training conditions
Power Distribution System Topology Identification and Prediction Employing Privacy Preserving Measurements · Texas Tech
Tried and failed
convolutional neural networks and k-nearest neighbors applied to epigenomic data imputation. Outcome: worse than baseline. Reason: consistently underperformed deep tensor factorization and gradient boosted decision tree models
Gene-regulatory circuitry of disease risk and progression · MIT
Tried and failed
standard convolutional neural networks for event classification applied to particle track image identification. Outcome: worse than baseline. Reason: High false-positive rate compared to parameter-based geometric convex hull analysis
The MIGDAL experiment · Imperial
Considered and rejected
Considered and rejected: Rejected convolutional neural networks for charge density analysis because feed-forward networks achieved comparable results while training faster.
MACHINE LEARNING AND MONTE CARLO STUDIES OF STRONGLY CORRELATED SYSTEMS · Cornell
Tried and failed
convolutional neural networks applied to sequence-based classification scoring functions. Outcome: worse than baseline. Reason: Nonlinear models offered no substantial performance benefit over linear regularized logistic regression.
Lost to a baseline
Gradient boosting and Convolutional Neural Networks (max accuracy 51-53%) failed to outperform simpler multinomial logistic regression, which had comparable accuracy and superior computation speed.
Lost to a baseline
On the small convolutional architecture for MNIST (6 and 16 5x5 filters), convolutional ENN achieved 97.65% accuracy compared to 98.70% for the GDN baseline.
The Capabilities of Neural Systems Depend on a Hierarchically Structured World · DSpace at UTSWMED
Tried and failed
convolutional neural network backward models applied to dry-contact ear-EEG speech stimulus reconstruction. Outcome: worse than baseline. Reason: None
Decoding auditory EEG responses to speech via deep neural networks · Imperial
Lost to a baseline
Simple majority-class baseline (predicting Haydn) achieves 71.8% accuracy, while the proposed convolutional models failed to reliably exceed 80% on Haydn vs Mozart string quartet attribution.
Leveraging Generative Models for Music and Signal Processing · ResearchWorks
Lost to a baseline
4-layer Feedforward Deep Neural Network (DNN) outperformed Convolutional Neural Networks (CNN) in estimating metasurface inverse design parameters from diffraction profiles.
Integrated bio-photonic devices : sensors, imagers, and beyond · MIT
Convolutional networks struggle on tabular and non-spatial data lacking grid structure
Convolutional layers failed or were rejected when applied to tabular features, genomic matrices, coordinate arrays, and multi-omics datasets that lack spatial locality or translation equivariance. In these contexts, convolutional models could not learn meaningful structural relationships and were abandoned in favor of multi-layer perceptrons or hand-crafted tabular features.
Tried and failed
1D convolutional neural networks on raw tabular matrices applied to genomic sequence variant classification. Outcome: worse than baseline. Reason: High dimensionality and lack of spatial inductive bias without dimensionality reduction
Applications of Machine Learning in Source Attribution and Gene Function Prediction · Virginia Tech
Tried and failed
convolutional neural network architectures applied to multi-omics data embedding and translation. Outcome: worse than baseline. Reason: lack of natural spatial locality or grid structure in omics feature ordering
Learning from multi-omics data of cancer · Imperial
Tried and failed
1D convolutional networks on tabular parameters applied to spatially dependent passive component modeling. Outcome: did not generalise. Reason: tabular scalar features failed to capture spatial layout dependencies compared to image-based representations
Machine learning-driven design automation for high-frequency circuits and packaging · UT Austin
Considered and rejected
Considered and rejected: Rejected image-based matrix encoding with Convolutional Neural Networks (CNNs) in favor of tabular hardware-aware features and small MLPs due to high conversion and training overheads.
Design and Management Strategies for Hardware Accelerators · DukeSpace
Considered and rejected
Considered and rejected: Rejected Convolutional Neural Networks (used in Dev et al. 2022) in favor of MLPs because CNNs are suited for spatial/image/time-series data rather than tabular clinical data.
Predictive Modeling on Cardiovascular Health · UT Austin
Tried and failed
PCA coordinate embeddings for convolutional networks applied to all-atom molecular structure generation. Outcome: worse than baseline. Reason: Coordinates lacked translation-equivariant image structure needed for convolutional architectures
Modeling Biomolecular Interactions with Generative Models · MIT
Considered and rejected
Considered and rejected: Rejected Fully Connected Neural Networks (FCNNs) and standard Convolutional Neural Networks (CNNs) for large-scale distributed routing because CNNs fail to generalize to arbitrary graphs and lack permutation invariance, while FCNNs/CNNs cannot be implemented in a distributed manner
GRAPH NEURAL NETWORKS FOR COMMUNICATION IN MULTI-AGENT SYSTEMS · Penn
Considered and rejected
Considered and rejected: Excluded convolutional layers from the neural network architecture (which were used in the reference amino acid ordering models) due to the lack of spatial/locational meaning in the binary rower vector.
Convolutional architectures fail to capture temporal dynamics and sequential dependencies
Feedforward convolutional networks and temporal convolutional layers struggled to capture long-term temporal dependencies and dynamic spatiotemporal interactions in time series. As a consequence, these models were routinely outperformed by recurrent architectures such as LSTMs and GRUs or failed to meet performance thresholds.
Tried and failed
non-recurrent 2D convolutional network applied to spatiotemporal neuroimaging classification. Outcome: worse than baseline. Reason: inability to exploit spatiotemporal dynamics compared to recurrent architectures
Machine learning in resting-state and naturalistic fMRI analysis · Cornell
Tried and failed
1D convolutional neural network for sequential forecasting applied to meteorological time series forecasting. Outcome: worse than baseline. Reason: 1D-CNN failed to capture long-term temporal dependencies compared to recurrent architectures
Forecasting Behind the Meter Solar Energy Production using Machine Learning Algorithms · Texas Tech
Tried and failed
standard recurrent and convolutional sequence architectures applied to multivariate traffic time-series forecasting. Reason: architectures failed to meet target error thresholds on complex multivariate patterns
Communication-efficient personalization in federated learning for edge devices · Iowa State
Tried and failed
dilated causal convolutional neural network applied to single-channel time series anomaly detection. Outcome: worse than baseline. Reason: underperformed human inter-rater agreement at fixed decision thresholds
QUANTITATIVE METHODS AND NEUROSTIMULATION MAPPING TO GUIDE PRECISION EPILEPSY THERAPIES · Penn
Tried and failed
non-recurrent feedforward and convolutional neural networks applied to spatiotemporal extreme climate event prediction. Outcome: did not generalise. Reason: lack of temporal recurrence failed to capture temporal dynamics required for predicting extreme events
Identifying drought in total water storage data using deep learning methods · UT Austin
Considered and rejected
Considered and rejected: Rejected Convolutional Neural Networks (CNNs) because they are invariant to the ordering of features, making them unsuitable for multi-limb time-series sequences.
Agile Tripedal Locomotion for Damaged Quadruped Robots · Penn
Tried and failed
temporal convolutional networks instead of recurrent layers applied to voice activity detection. Outcome: worse than baseline. Reason: yielded poor performance despite faster training time
Electroencephalography (EEG) based speech technologies using neural networks · UT Austin
Lost to a baseline
Phoneme classification from the last feedforward convolutional layer before the RNN achieved 52.4% accuracy, losing to CLAPP-s with GRU (61.7%)
Biologically plausible unsupervised learning in shallow and deep neural networks · EPFL
Convolutional models suffer from severe overfitting and poor out-of-distribution generalization
Under sparse training data, environmental mismatches, or distributional shifts across simulators, convolutional networks exhibited severe overfitting and failed to generalize to unseen test samples. These models generated near-constant or over-smoothed predictions and lagged behind hand-crafted feature baselines and regularized regression models.
Tried and failed
convolutional neural networks applied to voltammetric sensor signal regression. Outcome: did not generalise. Reason: yielded near-constant predictions on out-of-probe data, performing significantly worse than elastic net models
Novel Electrochemical Methods for Human Neurochemistry · Virginia Tech
Tried and failed
1D convolutional neural network applied to electrophysiological signal classification. Outcome: worse than baseline. Reason: insufficient training data caused poor generalization compared to handcrafted feature models
TOWARDS EFFICIENT AND SCALABLE MACHINE LEARNING FOR FUTURE NEURAL INTERFACES · Cornell
Tried and failed
convolutional neural network regression on covariance matrices applied to acoustic source range estimation. Outcome: worse than baseline. Reason: model over-smoothed predictions under environmental parameter mismatch without improving on matched-field processing accuracy
Ambient acoustics as indicator of environmental change in the Beaufort Sea: experiments & methods for analysis · Woods Hole
Tried and failed
2D deep convolutional neural network classification applied to 3D volumetric medical image classification. Outcome: did not generalise. Reason: Severe overfitting causing poor external validation performance compared to hand-crafted feature baselines
Tried and failed
transfer learning with pre-trained convolutional neural networks applied to droplet deposition density map classification. Outcome: overfit. Reason: models underperformed severely on test data and suffered from poor generalization and overfitting
Automated Exploration of High-Mix, Low Volume Direct Write Design Spaces Through Artificial Intelligence · Georgia Tech
Tried and failed
convolutional neural network on simulated physical fields applied to parameter inference across simulators. Outcome: did not generalise. Reason: subtle distributional shift and modeling differences between simulators biased out-of-distribution predictions
Tried and failed
raw waveform convolutional neural networks applied to speech state and trait detection. Outcome: did not generalise. Reason: poor generalization performance compared to baseline feature representations on unseen test data
Novel Methods For Detection And Analysis Of Atypical Aspects In Speech · EPFL
Tried and failed
convolutional neural network on raw binary images applied to porous media permeability estimation. Outcome: overfit. Reason: lack of geometric features or hierarchical multiscale representation
Machine learning applications for porous media · UT Austin
Graph convolutional networks fail to model complex relational data or surpass baselines
Graph convolutional networks underperformed linear baselines, sequence models, and equivariant transformers when applied to raw pixel grids, biological graphs, and crystal structures. In molecular potential energy modeling, continuous-filter graph convolutions proved unstable and failed to generate smooth energy surfaces across distance scans.
Tried and failed
graph convolutional networks with linear classifier applied to cosmological raw pixel image classification. Outcome: worse than baseline. Reason: Graph CNN representation on raw pixels lacked discriminative power for linear classification
Leveraging topology, geometry, and symmetries for efficient Machine Learning · EPFL
Tried and failed
graph convolutional network without pretraining applied to molecular property prediction. Outcome: worse than baseline. Reason: failed to exceed sequence-based models and barely matched simple baselines
Photoionization Detection of Volatile Organic Compounds · Harvard
Tried and failed
contrastive learning and early-fusion graph convolutional networks applied to heterogeneous dense and sparse biological graphs. Outcome: worse than baseline. Reason: None
Towards Network-Guided Large-Scale Foundation Models on Single-Cell Transcriptomics · Virginia Tech
Tried and failed
continuous-filter graph convolutional network applied to intermolecular potential energy surface modeling. Outcome: unstable. Reason: failed to generate smooth potential energy surfaces across distance scans
Reactive Molecular Dynamics in Ionic Media · Georgia Tech
Tried and failed
graph convolutional networks applied to discrete choice prediction. Outcome: worse than baseline. Reason: None
Computational Perspectives on Individual and Collective Decision-Making · Cornell
Tried and failed
graph convolutional network encoder-decoder architecture applied to human motion prediction. Outcome: worse than baseline. Reason: did not improve or slightly degraded performance compared to simple linear layers
Efficient Depth-based Deep Learning Methods for Multi-Party Pose Estimation · EPFL
Lost to a baseline
Standard Graph Convolutional Network (GCN) baseline was outperformed by both SE(3)-Transformer and EGNN on crystal structure data.
Convolutional neural networks are rejected due to computational overhead and operational rigidity
Authors rejected convolutional architectures because of excessive setup complexity, lengthy heuristic convergence times, and substantial training overheads on small datasets. Additionally, these models lacked interpretability and modularity, requiring complete re-training for differing encoding conditions.
Considered and rejected
Considered and rejected: Rejected using scalogram-based convolutional neural networks (e.g., WT_CNN) for IR thermal histories due to lack of interpretability, high computational overhead, and need for large training datasets.
High throughput process-materials framework for repairing Ni-based superalloys · Georgia Tech
Considered and rejected
Considered and rejected: Rejected deeper multi-layer CNN architectures for raw signal processing in favor of a single convolutional layer to avoid heavy computation and difficult training on small datasets.
Signal Processing Based Pathological Voice Detection Techniques · Scholarship at UWindsor Institutional Repository
Considered and rejected
Considered and rejected: Rejected the character convolutional neural network architecture from ELMo for general contextualized language models due to lack of an adequate output text generation method
Integrating Distributional, Compositional, and Relational Approaches to Neural Word Representations · Georgia Tech
Considered and rejected
Considered and rejected: Rejected Deep Learning and Convolutional Neural Networks (CNNs) due to setup complexity, long heuristic convergence, and lack of transparency.
Reduced Order Non-INtrusive (RONIN) Modeling for Strategic Defense Planning · Georgia Tech
Considered and rejected
Considered and rejected: Rejected standard supervised deep convolutional networks for image-to-excitation reconstruction due to lack of modularity and requiring complete re-training for every visual encoding condition.
Sensory representations optimized for the natural environment · Penn
Left open by the authors
Problems the authors named and did not get to.
Left open
Fuse Random Forest models with convolutional neural networks to classify SEM images for solid dosage formulation analysis. Blocker: Access to the thesis's specific formulation SEM image dataset
Overcoming challenges of solid dosage formulation development by using emerging technologies · UT Austin
Left open
Implement a fully convolutional neural network for object-based remote sensing classification to reduce pixel noise from Random Forest outputs. Blocker: None
Tectonic and Climatic Controls on Continental River Systems · MIT
Left open
Train convolutional neural networks to predict convergence and intermediary trajectory efficiency between initial and final robot swarm configurations. Blocker: None
Analytical and Machine Learning Methods for the Complete Safe Coordination of Astrobot Swarms · EPFL
Left open
Compare the performance of recurrent versus convolutional GNN architectures on code type inference benchmarks. Blocker: None
Leveraging machine learning for enhancing code performance and programming productivity · Georgia Tech
Left open
Benchmark the runtime of tropical convolutional layers against standard layers across fixed hardware setups. Blocker: None
APPLICATION OF TROPICAL CONVOLUTIONAL LAYERS IN NEURAL NETWORKS · Calhoun
Left open
Prove the minimax optimal rate of convergence for deep convolutional neural networks in least squares regression. Blocker: None
Statistical Learning Theory of Deep Neural Networks: A Generalization Viewpoint · Georgia Tech
Left open
Develop a feature selection method to select features automatically extracted by convolutional neural networks for probabilistic load forecasting. Blocker: None
Left open
Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Left open
Develop and evaluate deep learning models like convolutional neural networks on imbalanced crash severity datasets to capture nonlinear relationships. Blocker: None
Evaluating Factors Contributing to Crash Severity Among Older Drivers: Statistical Modeling and Machine Learning Approaches · Virginia Tech
Left open
Develop a methodological framework incorporating convolutional and recurrent neural networks for land use and cover change prediction using multi-temporal satellite imagery. Blocker: None
Using Data Science for Understanding and Predicting Long-term Land Use and Cover Changes: A Study Case in Mexico · IRIS - SNS - prod
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.