Chapter Four · failure evidence

What Convolutional Neural Networks got wrong, from 67 dissertations

Across the collected theses, convolutional neural networks frequently encountered difficulties when applied to non-spatial representations, temporal sequences, and small or shifted datasets. In numerous domains, these networks underperformed simpler linear models, suffered from architectural limitations, or were rejected due to excessive training complexity and operational overhead. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Flawed architectural configurations and input representations degrade convolutional network performance

12 theses · 8 institutions

Networks suffered substantial accuracy drops when trained on inappropriate input encodings, such as text without embedding layers, scrambled word order, or overly coarse pixel grids. Performance also deteriorated when architectures removed essential components like global pooling or attention, failed to concatenate multi-scale features, or used single-layer formulations.

Tried and failed

CNN with embedding layer applied to Arabic text classification. Outcome: worse than baseline. Reason: Standard convolutional architecture failed to capture semantic and syntactic patterns in Arabic thesis abstracts

Towards Logical Reasoning and Learning in Open and Dynamic Environments · Virginia Tech

Tried and failed

direct regression using convolutional neural network applied to spatial transfer function estimation from measurements. Outcome: worse than baseline. Reason: reduced low-dimensional morphological feature set lacked sufficient information for direct continuous mapping

Estimation of Hearing Aid Head Related Transfer Functions using Anthropometric Features · Harvard

Considered and rejected

Considered and rejected: Rejected shallow networks with only a single convolutional layer because they performed significantly worse on the F0 estimation task and produced poorer matches to human psychophysics than multi-convolutional-layer architectures.

Task-optimized models of human hearing link perception and neural coding · MIT

Considered and rejected

Considered and rejected: Rejected convolutional layer biases due to reduced network generalization capacity.

Advancing astronomical image deconvolution with deep learning · EPFL

Tried and failed

CNN without embedding layer applied to text classification. Outcome: worse than baseline. Reason: Lacked semantic representation of words prior to convolutional feature extraction

Improving the Accessibility of Arabic Electronic Theses and Dissertations (ETDs) with Metadata and Classification · Virginia Tech

Tried and failed

concatenating conditioning embeddings to raw inputs applied to image classification with convolutional networks. Outcome: worse than baseline. Reason: input-level concatenation degrades prospective risk compared to conditioning at deeper network layers

The Principles of Learning on Multiple Tasks · Penn

Tried and failed

one-hot CNN with randomized word order applied to text classification. Outcome: worse than baseline. Reason: destroying sequential order prevents convolutional filters from learning meaningful local n-gram patterns compared to simple bag-of-words

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Tried and failed

simplifying activation functions and removing attention layers applied to compact convolutional neural networks. Outcome: worse than baseline. Reason: naive removal of squeeze-and-excitation layers and substituting hard-swish with ReLU severely degraded classification accuracy

Efficient model adaptation and compression for edge intelligence · UT Austin

Tried and failed

removing global pooling in neural architecture search applied to evolving convolutional neural network topologies. Outcome: worse than baseline. Reason: evolved networks collapsed to using only pooling layers and failed to train effectively

Image Classification with Evolved Convolutional Neural Networks · Harvard

Tried and failed

convolutional neural network for multiscale spatial modeling applied to spatial mixture population property modeling. Outcome: did not generalise. Reason: biased toward learning shorter-scale spatial continuity while averaging out long-range continuity structures

Deep learning for spatial nonstationarity : evaluation, mitigation, and generation · UT Austin

Considered and rejected

Considered and rejected: Rejected standard deep CNN architectures without feature concatenation because deep convolutional layers gradually saturate and lose low-concentration metabolite information while shallow layers lack high-concentration metabolite details.

Machine learning for magnetic resonance spectroscopy: modeling in the preclinical development process · OpenBU

Considered and rejected

Considered and rejected: 2D Convolutional Neural Network for 4x6 NaI array classification (deprecated because 24 pixels were too coarse and low-multiplicity muons mimicked signal topology)

Studies of the Electron Neutrino Charged-current Interaction on $^{127}$I · DukeSpace

Convolutional models fail to outperform simpler linear models and classical baselines

10 theses · 7 institutions

Convolutional neural networks were repeatedly matched or beaten by simpler methods, including linear and quadratic discriminants, multinomial logistic regression, and standard feed-forward networks. Practitioners noted that the additional architectural complexity produced no noticeable accuracy gains while incurring higher computational demands.

Tried and failed

convolutional neural networks with small training sets applied to time-series classification. Outcome: worse than baseline. Reason: underperformed compared to linear and quadratic discriminants under sparse data training conditions

Power Distribution System Topology Identification and Prediction Employing Privacy Preserving Measurements · Texas Tech

Tried and failed

convolutional neural networks and k-nearest neighbors applied to epigenomic data imputation. Outcome: worse than baseline. Reason: consistently underperformed deep tensor factorization and gradient boosted decision tree models

Gene-regulatory circuitry of disease risk and progression · MIT

Tried and failed

standard convolutional neural networks for event classification applied to particle track image identification. Outcome: worse than baseline. Reason: High false-positive rate compared to parameter-based geometric convex hull analysis

The MIGDAL experiment · Imperial

Considered and rejected

Considered and rejected: Rejected convolutional neural networks for charge density analysis because feed-forward networks achieved comparable results while training faster.

MACHINE LEARNING AND MONTE CARLO STUDIES OF STRONGLY CORRELATED SYSTEMS · Cornell

Tried and failed

convolutional neural networks applied to sequence-based classification scoring functions. Outcome: worse than baseline. Reason: Nonlinear models offered no substantial performance benefit over linear regularized logistic regression.

Investigating how sequence variation in T cell receptors and human leukocyte antigens shapes T cell development · Harvard

Lost to a baseline

Gradient boosting and Convolutional Neural Networks (max accuracy 51-53%) failed to outperform simpler multinomial logistic regression, which had comparable accuracy and superior computation speed.

Essays on the role of networks and networking in entrepreneurship and innovation: social structure as a playfield for networking · Imperial

Lost to a baseline

On the small convolutional architecture for MNIST (6 and 16 5x5 filters), convolutional ENN achieved 97.65% accuracy compared to 98.70% for the GDN baseline.

The Capabilities of Neural Systems Depend on a Hierarchically Structured World · DSpace at UTSWMED

Tried and failed

convolutional neural network backward models applied to dry-contact ear-EEG speech stimulus reconstruction. Outcome: worse than baseline. Reason: None

Decoding auditory EEG responses to speech via deep neural networks · Imperial

Lost to a baseline

Simple majority-class baseline (predicting Haydn) achieves 71.8% accuracy, while the proposed convolutional models failed to reliably exceed 80% on Haydn vs Mozart string quartet attribution.

Leveraging Generative Models for Music and Signal Processing · ResearchWorks

Lost to a baseline

4-layer Feedforward Deep Neural Network (DNN) outperformed Convolutional Neural Networks (CNN) in estimating metasurface inverse design parameters from diffraction profiles.

Integrated bio-photonic devices : sensors, imagers, and beyond · MIT

Convolutional networks struggle on tabular and non-spatial data lacking grid structure

8 theses · 7 institutions

Convolutional layers failed or were rejected when applied to tabular features, genomic matrices, coordinate arrays, and multi-omics datasets that lack spatial locality or translation equivariance. In these contexts, convolutional models could not learn meaningful structural relationships and were abandoned in favor of multi-layer perceptrons or hand-crafted tabular features.

Tried and failed

1D convolutional neural networks on raw tabular matrices applied to genomic sequence variant classification. Outcome: worse than baseline. Reason: High dimensionality and lack of spatial inductive bias without dimensionality reduction

Applications of Machine Learning in Source Attribution and Gene Function Prediction · Virginia Tech

Tried and failed

convolutional neural network architectures applied to multi-omics data embedding and translation. Outcome: worse than baseline. Reason: lack of natural spatial locality or grid structure in omics feature ordering

Learning from multi-omics data of cancer · Imperial

Tried and failed

1D convolutional networks on tabular parameters applied to spatially dependent passive component modeling. Outcome: did not generalise. Reason: tabular scalar features failed to capture spatial layout dependencies compared to image-based representations

Machine learning-driven design automation for high-frequency circuits and packaging · UT Austin

Considered and rejected

Considered and rejected: Rejected image-based matrix encoding with Convolutional Neural Networks (CNNs) in favor of tabular hardware-aware features and small MLPs due to high conversion and training overheads.

Design and Management Strategies for Hardware Accelerators · DukeSpace

Considered and rejected

Considered and rejected: Rejected Convolutional Neural Networks (used in Dev et al. 2022) in favor of MLPs because CNNs are suited for spatial/image/time-series data rather than tabular clinical data.

Predictive Modeling on Cardiovascular Health · UT Austin

Tried and failed

PCA coordinate embeddings for convolutional networks applied to all-atom molecular structure generation. Outcome: worse than baseline. Reason: Coordinates lacked translation-equivariant image structure needed for convolutional architectures

Modeling Biomolecular Interactions with Generative Models · MIT

Considered and rejected

Considered and rejected: Rejected Fully Connected Neural Networks (FCNNs) and standard Convolutional Neural Networks (CNNs) for large-scale distributed routing because CNNs fail to generalize to arbitrary graphs and lack permutation invariance, while FCNNs/CNNs cannot be implemented in a distributed manner

GRAPH NEURAL NETWORKS FOR COMMUNICATION IN MULTI-AGENT SYSTEMS · Penn

Considered and rejected

Considered and rejected: Excluded convolutional layers from the neural network architecture (which were used in the reference amino acid ordering models) due to the lack of spatial/locational meaning in the binary rower vector.

Rowing Against the Wind: An Analysis of the Impact of Variable Wind Conditions on Current and Prospective Rowing Selection Methods · Harvard

Convolutional architectures fail to capture temporal dynamics and sequential dependencies

8 theses · 6 institutions

Feedforward convolutional networks and temporal convolutional layers struggled to capture long-term temporal dependencies and dynamic spatiotemporal interactions in time series. As a consequence, these models were routinely outperformed by recurrent architectures such as LSTMs and GRUs or failed to meet performance thresholds.

Tried and failed

non-recurrent 2D convolutional network applied to spatiotemporal neuroimaging classification. Outcome: worse than baseline. Reason: inability to exploit spatiotemporal dynamics compared to recurrent architectures

Machine learning in resting-state and naturalistic fMRI analysis · Cornell

Tried and failed

1D convolutional neural network for sequential forecasting applied to meteorological time series forecasting. Outcome: worse than baseline. Reason: 1D-CNN failed to capture long-term temporal dependencies compared to recurrent architectures

Forecasting Behind the Meter Solar Energy Production using Machine Learning Algorithms · Texas Tech

Tried and failed

standard recurrent and convolutional sequence architectures applied to multivariate traffic time-series forecasting. Reason: architectures failed to meet target error thresholds on complex multivariate patterns

Communication-efficient personalization in federated learning for edge devices · Iowa State

Tried and failed

dilated causal convolutional neural network applied to single-channel time series anomaly detection. Outcome: worse than baseline. Reason: underperformed human inter-rater agreement at fixed decision thresholds

QUANTITATIVE METHODS AND NEUROSTIMULATION MAPPING TO GUIDE PRECISION EPILEPSY THERAPIES · Penn

Tried and failed

non-recurrent feedforward and convolutional neural networks applied to spatiotemporal extreme climate event prediction. Outcome: did not generalise. Reason: lack of temporal recurrence failed to capture temporal dynamics required for predicting extreme events

Identifying drought in total water storage data using deep learning methods · UT Austin

Considered and rejected

Considered and rejected: Rejected Convolutional Neural Networks (CNNs) because they are invariant to the ordering of features, making them unsuitable for multi-limb time-series sequences.

Agile Tripedal Locomotion for Damaged Quadruped Robots · Penn

Tried and failed

temporal convolutional networks instead of recurrent layers applied to voice activity detection. Outcome: worse than baseline. Reason: yielded poor performance despite faster training time

Electroencephalography (EEG) based speech technologies using neural networks · UT Austin

Lost to a baseline

Phoneme classification from the last feedforward convolutional layer before the RNN achieved 52.4% accuracy, losing to CLAPP-s with GRU (61.7%)

Biologically plausible unsupervised learning in shallow and deep neural networks · EPFL

Convolutional models suffer from severe overfitting and poor out-of-distribution generalization

8 theses · 8 institutions

Under sparse training data, environmental mismatches, or distributional shifts across simulators, convolutional networks exhibited severe overfitting and failed to generalize to unseen test samples. These models generated near-constant or over-smoothed predictions and lagged behind hand-crafted feature baselines and regularized regression models.

Tried and failed

convolutional neural networks applied to voltammetric sensor signal regression. Outcome: did not generalise. Reason: yielded near-constant predictions on out-of-probe data, performing significantly worse than elastic net models

Novel Electrochemical Methods for Human Neurochemistry · Virginia Tech

Tried and failed

1D convolutional neural network applied to electrophysiological signal classification. Outcome: worse than baseline. Reason: insufficient training data caused poor generalization compared to handcrafted feature models

TOWARDS EFFICIENT AND SCALABLE MACHINE LEARNING FOR FUTURE NEURAL INTERFACES · Cornell

Tried and failed

convolutional neural network regression on covariance matrices applied to acoustic source range estimation. Outcome: worse than baseline. Reason: model over-smoothed predictions under environmental parameter mismatch without improving on matched-field processing accuracy

Ambient acoustics as indicator of environmental change in the Beaufort Sea: experiments & methods for analysis · Woods Hole

Tried and failed

2D deep convolutional neural network classification applied to 3D volumetric medical image classification. Outcome: did not generalise. Reason: Severe overfitting causing poor external validation performance compared to hand-crafted feature baselines

Development of artificial intelligence models to classify pulmonary nodules and improve lung cancer early diagnosis · Imperial

Tried and failed

transfer learning with pre-trained convolutional neural networks applied to droplet deposition density map classification. Outcome: overfit. Reason: models underperformed severely on test data and suffered from poor generalization and overfitting

Automated Exploration of High-Mix, Low Volume Direct Write Design Spaces Through Artificial Intelligence · Georgia Tech

Tried and failed

convolutional neural network on simulated physical fields applied to parameter inference across simulators. Outcome: did not generalise. Reason: subtle distributional shift and modeling differences between simulators biased out-of-distribution predictions

Towards Precision Measurements Of The Optical Depth To Reionization Using 21 Cm Data And Machine Learning · Penn

Tried and failed

raw waveform convolutional neural networks applied to speech state and trait detection. Outcome: did not generalise. Reason: poor generalization performance compared to baseline feature representations on unseen test data

Novel Methods For Detection And Analysis Of Atypical Aspects In Speech · EPFL

Tried and failed

convolutional neural network on raw binary images applied to porous media permeability estimation. Outcome: overfit. Reason: lack of geometric features or hierarchical multiscale representation

Machine learning applications for porous media · UT Austin

Graph convolutional networks fail to model complex relational data or surpass baselines

7 theses · 6 institutions

Graph convolutional networks underperformed linear baselines, sequence models, and equivariant transformers when applied to raw pixel grids, biological graphs, and crystal structures. In molecular potential energy modeling, continuous-filter graph convolutions proved unstable and failed to generate smooth energy surfaces across distance scans.

Tried and failed

graph convolutional networks with linear classifier applied to cosmological raw pixel image classification. Outcome: worse than baseline. Reason: Graph CNN representation on raw pixels lacked discriminative power for linear classification

Leveraging topology, geometry, and symmetries for efficient Machine Learning · EPFL

Tried and failed

graph convolutional network without pretraining applied to molecular property prediction. Outcome: worse than baseline. Reason: failed to exceed sequence-based models and barely matched simple baselines

Photoionization Detection of Volatile Organic Compounds · Harvard

Tried and failed

contrastive learning and early-fusion graph convolutional networks applied to heterogeneous dense and sparse biological graphs. Outcome: worse than baseline. Reason: None

Towards Network-Guided Large-Scale Foundation Models on Single-Cell Transcriptomics · Virginia Tech

Tried and failed

continuous-filter graph convolutional network applied to intermolecular potential energy surface modeling. Outcome: unstable. Reason: failed to generate smooth potential energy surfaces across distance scans

Reactive Molecular Dynamics in Ionic Media · Georgia Tech

Tried and failed

graph convolutional networks applied to discrete choice prediction. Outcome: worse than baseline. Reason: None

Computational Perspectives on Individual and Collective Decision-Making · Cornell

Tried and failed

graph convolutional network encoder-decoder architecture applied to human motion prediction. Outcome: worse than baseline. Reason: did not improve or slightly degraded performance compared to simple linear layers

Efficient Depth-based Deep Learning Methods for Multi-Party Pose Estimation · EPFL

Lost to a baseline

Standard Graph Convolutional Network (GCN) baseline was outperformed by both SE(3)-Transformer and EGNN on crystal structure data.

Deep learning algorithms for predicting association between antibody sequence, structure, and antibody properties · Oxford

Convolutional neural networks are rejected due to computational overhead and operational rigidity

5 theses · 3 institutions

Authors rejected convolutional architectures because of excessive setup complexity, lengthy heuristic convergence times, and substantial training overheads on small datasets. Additionally, these models lacked interpretability and modularity, requiring complete re-training for differing encoding conditions.

Considered and rejected

Considered and rejected: Rejected using scalogram-based convolutional neural networks (e.g., WT_CNN) for IR thermal histories due to lack of interpretability, high computational overhead, and need for large training datasets.

High throughput process-materials framework for repairing Ni-based superalloys · Georgia Tech

Considered and rejected

Considered and rejected: Rejected deeper multi-layer CNN architectures for raw signal processing in favor of a single convolutional layer to avoid heavy computation and difficult training on small datasets.

Signal Processing Based Pathological Voice Detection Techniques · Scholarship at UWindsor Institutional Repository

Considered and rejected

Considered and rejected: Rejected the character convolutional neural network architecture from ELMo for general contextualized language models due to lack of an adequate output text generation method

Integrating Distributional, Compositional, and Relational Approaches to Neural Word Representations · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Deep Learning and Convolutional Neural Networks (CNNs) due to setup complexity, long heuristic convergence, and lack of transparency.

Reduced Order Non-INtrusive (RONIN) Modeling for Strategic Defense Planning · Georgia Tech

Considered and rejected

Considered and rejected: Rejected standard supervised deep convolutional networks for image-to-excitation reconstruction due to lack of modularity and requiring complete re-training for every visual encoding condition.

Sensory representations optimized for the natural environment · Penn

Left open by the authors

Problems the authors named and did not get to.

Left open

Fuse Random Forest models with convolutional neural networks to classify SEM images for solid dosage formulation analysis. Blocker: Access to the thesis's specific formulation SEM image dataset

Overcoming challenges of solid dosage formulation development by using emerging technologies · UT Austin

Left open

Implement a fully convolutional neural network for object-based remote sensing classification to reduce pixel noise from Random Forest outputs. Blocker: None

Tectonic and Climatic Controls on Continental River Systems · MIT

Left open

Train convolutional neural networks to predict convergence and intermediary trajectory efficiency between initial and final robot swarm configurations. Blocker: None

Analytical and Machine Learning Methods for the Complete Safe Coordination of Astrobot Swarms · EPFL

Left open

Compare the performance of recurrent versus convolutional GNN architectures on code type inference benchmarks. Blocker: None

Leveraging machine learning for enhancing code performance and programming productivity · Georgia Tech

Left open

Benchmark the runtime of tropical convolutional layers against standard layers across fixed hardware setups. Blocker: None

APPLICATION OF TROPICAL CONVOLUTIONAL LAYERS IN NEURAL NETWORKS · Calhoun

Left open

Prove the minimax optimal rate of convergence for deep convolutional neural networks in least squares regression. Blocker: None

Statistical Learning Theory of Deep Neural Networks: A Generalization Viewpoint · Georgia Tech

Left open

Develop a feature selection method to select features automatically extracted by convolutional neural networks for probabilistic load forecasting. Blocker: None

A Novel Embedded Feature Selection Framework for Probabilistic Load Forecasting With Sparse Data via Bayesian Inference · HARVEST

Left open

Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Left open

Develop and evaluate deep learning models like convolutional neural networks on imbalanced crash severity datasets to capture nonlinear relationships. Blocker: None

Evaluating Factors Contributing to Crash Severity Among Older Drivers: Statistical Modeling and Machine Learning Approaches · Virginia Tech

Left open

Develop a methodological framework incorporating convolutional and recurrent neural networks for land use and cover change prediction using multi-temporal satellite imagery. Blocker: None

Using Data Science for Understanding and Predicting Long-term Land Use and Cover Changes: A Study Case in Mexico · IRIS - SNS - prod

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.