Chapter Four · failure evidence

What Deep Learning & Pretraining got wrong, from 23 dissertations

The records evaluate deep learning and pretrained language models across biomedical, tabular, language, and forecasting applications. Although deep neural methods offer powerful capabilities, they frequently suffer from overfitting on small datasets, fail against simpler classical baselines, or face domain adaptation and deployment limitations. These records come from PhD theses at 15 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Deep learning models are rejected due to data scarcity and high overfitting risk in small sample regimes

5 theses · 5 institutions

Deep learning models were rejected across behavioral, biomedical, tabular, chemometric, and electroencephalogram studies because small sample sizes created a severe risk of overfitting. Authors noted that deep neural networks required excessive amounts of data and computational overhead while offering poor interpretability compared to classical methods.

Considered and rejected

Considered and rejected: Rejected using complex deep learning techniques without multi-task framing on small-sample behavioral datasets due to severe overfitting and poor generalizability.

Towards Human-Centered Behavioral Sensing for Student Support · ResearchWorks

Considered and rejected

Considered and rejected: Rejected Deep Learning (DL) models due to high risk of overfitting on small biomedical sample sizes, computational overhead, and lack of interpretability.

Machine learning prediction of atezolizumab treatment response in metastatic urothelial carcinoma using gene expression and clinical data · Imperial

Considered and rejected

Considered and rejected: Rejected OpenNN and deep learning frameworks (Keras/TensorFlow) due to narrow focus on neural networks and excessive data requirements for small tabular data

Fatigue Monitoring System · SUPSI - ARIS

Considered and rejected

Considered and rejected: Rejected deep learning approaches due to limited training data and high susceptibility to overfitting in small EEG sample regimes.

Instrumentation for daily-life brain-computer interfaces · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Deep learning / ANN models were deprioritized over PLS/SVM due to high chemometric collinearity and small sample size ('small n, large p' regime) risking overfitting without massive datasets.

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Deep architectures underperform or are rejected in favor of simpler classical baselines

4 theses · 3 institutions

Deep learning approaches were beaten by or rejected in favor of simpler models such as Random Forest, logistic regression, and optical flow interpolation. These classical baselines provided lower whole-image error, better false positive reduction, competitive tabular performance, and faster polynomial-time explainability.

Tried and failed

deep learning on raw signals and scalograms applied to arterial Doppler waveform classification. Outcome: worse than baseline. Reason: deep models underperformed handcrafted feature-based logistic regression classifier

Testing for peripheral arterial disease in diabetes · Imperial

Considered and rejected

Considered and rejected: Deep learning architectures rejected in favor of Random Forest due to competitive tabular performance and fast polynomial-time SHAP computation via TreeExplainer.

ZeoSyn: A Comprehensive Zeolite Synthesis Dataset Enabling Machine-Learning Rationalization of Hydrothermal Parameters · MIT

Considered and rejected

Considered and rejected: Rejected using deep learning architecture over Random Forest classifier because Random Forest was simpler and reduced false positives by 50% vs 35% for DNN.

Molecular simulation on two-dimensional nanoporous material for energy-efficient separation · EPFL

Lost to a baseline

Linear and optical flow interpolation yielded lower whole-image MSE than deep learning methods (R2UNet/proposed) because the latter deprioritized background non-ROI pixels.

Controllable synthetic algorithms and evaluations for annotated clinical image synthesis · Imperial

Domain adaptation and specialized pretrained representations degrade downstream performance

3 theses · 3 institutions

Applying domain fine-tuned language models to time series forecasting led to performance degradation across all horizons and metrics relative to standard deep learning models. Similarly, domain-adapted language representations underperformed general BERT and GPT-2 on summarization and degraded sentiment classification recall when trained on raw chatbot text.

Tried and failed

domain fine-tuned zero-shot language model applied to time series forecasting. Outcome: worse than baseline. Reason: exhibited performance degradation across all horizons and metrics compared to standard deep learning models

Deep Learning Methods for Built Environment Operational Management · Virginia Tech

Tried and failed

Domain-specific pretrained language model embeddings applied to Extractive text summarization. Outcome: worse than baseline. Reason: Domain-adapted BERT underperformed general domain BERT and GPT-2 on extractive summarization metrics

Computational acquisition of knowledge in small-data environments: a case study in the field of energetics · Imperial

Tried and failed

pretrained language model embeddings on raw text applied to chatbot text sentiment classification. Outcome: worse than baseline. Reason: raw chatbot-generated text representations degraded classification recall compared to customer text alone

Three essays on strategic digital initiatives in modern organizations · Georgia Tech

Fine-tuning and scaling pretrained transformers causes overfitting or underperforms smaller models

2 theses · 2 institutions

Fine-tuning pretrained transformers for continuous regression suffered from fluctuating validation loss in later epochs and failed to generalize to unseen data. Furthermore, larger pretrained transformer embeddings underperformed smaller DistilBERT representations across all evaluated downstream text classification tasks.

Tried and failed

fine-tuning pretrained transformer for continuous regression applied to short text evaluation scores. Outcome: overfit. Reason: validation loss fluctuated in later epochs, failing to generalize to unseen data

Enhancing Sensory Panel Decision-Making: A GenAI Approach Using RoBERTa to Quantify Free-Form Comments · Cornell

Tried and failed

larger pretrained transformer embeddings applied to fine-grained text classification. Outcome: worse than baseline. Reason: DistilBERT outperformed full BERT representations across all tested downstream classifiers.

Spatiotemporal Event Forecasting and Analysis with Ubiquitous Urban Sensors · Virginia Tech

Pretrained models and large deep architectures are rejected due to temporal leakage and latency limits

2 theses · 2 institutions

Off-the-shelf contemporary large language models were rejected for historical financial forecasting because severe temporal pretraining leakage invalidated their use. In addition, full-scale GPU and cloud deep learning baselines were rejected because they flattened features and failed to optimize for low edge latency.

Considered and rejected

Considered and rejected: Rejected off-the-shelf contemporary LLMs for historical financial forecasting due to severe temporal pretraining leakage.

Essays in Finance, Technology, and Behavior · Harvard

Considered and rejected

Considered and rejected: Rejected comparing against full-scale GPU/cloud deep learning baselines (Transformers/CNN-LSTM) because they treat input features as flat vectors and do not optimize for low edge latency

LIDIT: Low-Latency Intrusion Detection in IoMT Devices using TinyML · Queens University Institutional Repository

Left open by the authors

Problems the authors named and did not get to.

Left open

Explore and evaluate improved deep learning architectures for classifying Electronic Theses and Dissertations. Blocker: None

Classifying ETDs · Virginia Tech

Left open

Tune hyperparameters and architectures for deep learning emergency department arrival forecasting models across 1h to 48h horizons. Blocker: Requires the emergency department arrival dataset used in the thesis, which is likely private hospital data.

Predicting the Unpredictable: Comparing Statistical Forecasting and Deep Learning Models for Forecasting Emergency Department Arrivals · Harvard

Left open

Evaluate pretrained transformer models with subword tokenization and transfer learning on scraped job postings to predict bankruptcy and corporate growth. Blocker: Scraped job listings dataset may not be publicly archived with the thesis

Assessing Corporate Growth and Bankruptcy Risk Using Public Data Proxies · Harvard

Left open

Apply transfer learning and transformer architectures to improve generalization of deep learning models across diverse seismic datasets. Blocker: The objective is a broad research direction lacking specific datasets, target tasks, or evaluation metrics.

Improving accuracy and efficiency of seismic data analysis using deep learning · UT Austin

Left open

Develop deep learning methods to reduce the computational burden of real-time SDRE feedback stabilization for nonlinear PDEs. Blocker: The proposal lacks any concrete model architecture, training formulation, or target performance metrics.

State-dependent Riccati equation feedback stabilization for nonlinear PDEs · Imperial

Left open

Evaluate deep learning architectures against shallow models for GHG and odour emissions from biological nutrient removal wastewater data. Blocker: Lacks access to the specific cold-region wastewater treatment plant field and laboratory monitoring dataset.

ESTIMATION OF GREENHOUSE GAS AND ODOUR EMISSIONS FROM COLD REGION MUNICIPAL BIOLOGICAL NUTRIENT REMOVAL WASTEWATER TREATMENT PROCESSES · HARVEST

Left open

Develop tabular deep learning architectures or evaluate feature selection across wider combinations of cardinality and coefficient of variation. Blocker: None

Data Value Analytics for Feature Selection and Dataset Retrieval · Research Repository UCD

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.