Chapter Four · failure evidence

What Recurrent Neural Networks got wrong, from 47 dissertations

Recurrent neural networks frequently encounter practical difficulties ranging from gradient pathologies and error accumulation over long sequences to inferior performance compared to tree-based baselines on tabular sequential tasks. Practitioners also reject or abandon these architectures due to computational overhead, lack of hidden state interpretability, and poor generalization under distribution shifts or inadequate regularization. These records come from PhD theses at 18 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Recurrent neural networks underperform simpler baselines and tree-based models on tabular and sequential features

9 theses · 7 institutions

Deep sequence models repeatedly perform worse than gradient boosted trees, random forests, support vector machines, and simple feedforward networks on sequential regression and classification tasks. These performance deficits are especially pronounced on tabular sensor readings, short clinical trajectories, and time-series data with limited training sample sizes.

Tried and failed

LSTM on sentiment and technical features applied to financial price direction prediction. Outcome: worse than baseline. Reason: Recurrent neural networks underperformed traditional support vector machines on multimodal tabular and sentiment features

Machine Learning for Financial Market Forecasting · Harvard

Tried and failed

ARIMA and recurrent neural networks applied to transient emission time series prediction. Outcome: worse than baseline. Reason: severe non-linearity and limited training sample size caused overfitting and high variance compared to gradient boosting

A deterministic model for wear of piston ring and liner and a machine learning-based model for engine oil emissions · MIT

Tried and failed

recurrent and time-delay neural networks applied to aerodynamic time-series surrogate modeling. Outcome: worse than baseline. Reason: training sensitivity caused recurrent architectures to underperform simpler feedforward networks on time-series data

A Controller Development Methodology Incorporating Unsteady, Coupled Aerodynamics and Flight Control Modeling for Atmospheric Entry Vehicles · Georgia Tech

Tried and failed

recurrent neural networks for time-series regression applied to tabular time-series behavior forecasting. Outcome: worse than baseline. Reason: Deep sequence models underperformed gradient boosted trees on tabular sequential regression features.

Understanding and Predicting Sit-Stand Desk Usage Patterns and Willingness among Knowledge Workers: A Data-Driven Approach · Virginia Tech

Tried and failed

Stacked Bi-LSTM applied to time-series anomaly detection. Outcome: worse than baseline. Reason: Tree-based ensembles outperformed the recurrent neural network on the tabular sensor features.

Anomaly Detection and Prevention in Smart Grid Transmission System Using Machine Learning · Texas Tech

Tried and failed

recurrent neural networks for clinical trajectory modelling applied to short tabular longitudinal laboratory trajectories. Outcome: worse than baseline. Reason: failed to outperform simple logistic regression and exhibited substantially wider confidence intervals

Clinical risk prediction of progression to severe dengue illness during the febrile phase in primary healthcare settings · Imperial

Tried and failed

recurrent neural networks applied to time-series trigger classification. Outcome: worse than baseline. Reason: simpler feedforward neural network performed better or was preferred for classification

Three Terrestrial Planets Transiting Mid-to-Late M Dwarfs and Constraints on the Occurrence of Such Worlds · Harvard

Tried and failed

recurrent neural networks for tabular wear forecasting applied to component degradation time series regression. Outcome: worse than baseline. Reason: deep sequence models did not outperform tree-based regressors despite higher computational cost

Digital Twin-Driven Condition Monitoring Approach for Aircraft Carbon Brakes · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Recurrent Neural Network (RNN) approach for future streamflow modeling due to superior objective function metrics achieved by Random Forest

Climate change impact on the spatial distribution of droughts in Kirindi oya and Maduru oya dry zone river basins in Sri Lanka · Institutional Repository University of Moratuwa

Vanishing and exploding gradients impede training across long sequences and stiff dynamics

9 theses · 9 institutions

Standard recurrent networks suffer from severe gradient vanishing and exploding during backpropagation through time over extended sequences or deep computational graphs. These gradient pathologies prevent models from retaining long-term dependencies and cause dynamic tracking instabilities or training failure on stiff differential equations.

Considered and rejected

Considered and rejected: Rejected standard Recurrent Neural Networks (RNNs) due to vanishing gradient issues with long-term sequential dependencies in power load cycles.

IMPROVING CYBER RESILIENCE OF SHIPBOARD POWER SYSTEMS USING MACHINE LEARNING · Calhoun

Tried and failed

Standard recurrent neural networks applied to long-term time series prediction. Reason: Gradient vanishing prevents retaining early sequence features over long horizons

Wind Power Prediction and Uncertainty Modeling for Power System Operation · HARVEST

Tried and failed

physics-informed neural networks and recurrent neural networks applied to stiff differential equations. Outcome: did not converge. Reason: severe gradient pathologies during training on stiff multiscale dynamics

Approximation of Large Stiff Acausal Models · MIT

Tried and failed

Vanilla recurrent neural network applied to Dynamic actuator system identification. Outcome: unstable. Reason: Vanishing and exploding gradients caused output fluctuations during periodic dynamic tracking

LSTM sequence-to-sequence based System Identification and feedforward-feedback control of piezoelectric actuators · Iowa State

Tried and failed

Training standard recurrent neural networks with backpropagation applied to long audio sequence recognition. Outcome: unstable. Reason: Exploding gradients during backpropagation through time over long sequences.

Biologically Inspired Spiking Neural Networks for Speech Recognition · EPFL

Considered and rejected

Considered and rejected: Rejected basic Simple Recurrent Neural Networks (RNNs) due to vanishing gradients that prevented learning long-range word dependencies across long sequences.

Data driven strategies for product design · UT Austin

Considered and rejected

Considered and rejected: Rejected Recurrent Neural Networks (RNNs) due to excessive resource consumption hindering real-time performance and vanishing gradient problems from frequent zero-velocity fixations

Leveraging the multi-modal nature of communication in immersive XR environments · Imperial

Considered and rejected

Considered and rejected: Rejected feedforward neural networks (FFNN) and standard recurrent neural networks (Elman and Jordan RNNs) for hydrologic sequence modeling due to inability to retain sequential memory and vulnerability to the vanishing gradient problem.

Machine Learning Methods for Modeling Streamflow in Intermittent Rivers and Ephemeral Streams · Texas Tech

Considered and rejected

Considered and rejected: Rejected standard recurrent neural networks (RNNs) for Processing Cells in favor of LSTM cells to avoid the vanishing gradient problem over deep logical circuits.

Pushing the frontier of quantum many-body simulation using classical computers · Cornell

Sequential error accumulation and history drift degrade multi-step predictions

6 theses · 4 institutions

Autoregressive rollouts and stateful recurrence accumulate errors across successive forecasting steps, leading to severe drift and decaying correlation over time. Sequence models also fail to retain early historical context over long horizons or suffer divergence when evaluating optimization steps without resetting hidden states.

Tried and failed

recurrent neural networks for time-series forecasting applied to long-term reservoir production forecasting. Outcome: worse than baseline. Reason: Severe error accumulation and inability to capture long-term decline trends compared to standard empirical baselines

Asset Development and Sweet Spot Identification in Unconventional Reservoirs Using Machine Learning Approaches · Texas Tech

Tried and failed

autoregressive LSTM for multi-step time series forecasting applied to daily aggregated emergency department arrivals. Outcome: worse than baseline. Reason: fine-grained recurrent models accumulate errors when aggregating over longer daily horizons compared to short hourly predictions

Predicting the Unpredictable: Comparing Statistical Forecasting and Deep Learning Models for Forecasting Emergency Department Arrivals · Harvard

Tried and failed

Standard recurrent neural network applied to sequential degradation data modeling. Outcome: did not generalise. Reason: Captures only recent memory and fails to address long-term dependencies in the sequence

Advanced data-driven methods for prognostics and life extension of assets using condition monitoring and sensor data. · Cranfield

Tried and failed

recurrent neural network with convolutional layers applied to spatiotemporal fluid flow PDE dynamics. Outcome: did not generalise. Reason: fails to capture source patterns/location after initial steps and suffers decaying correlation over time

Deep Learning for Dynamical Systems: Modeling, Prediction, and Control · Georgia Tech

Tried and failed

recurrent neural networks for sequential decision making applied to error recovery action suggestion. Outcome: worse than baseline. Reason: human error accumulation over sequential steps degraded accuracy and robustness relative to non-recurrent models

Facilitating Reliable Autonomy with Human-Robot Interaction · Georgia Tech

Tried and failed

stateful recurrent neural networks in iterative solvers applied to coupled nonlinear equilibrium problems. Outcome: did not converge. Reason: evaluating sequence models across optimization steps without resetting hidden states causes fictitious history accumulation and divergence

Physics enabled Data-driven structural analysis for mechanical components and assemblies · Georgia Tech

High computational overhead and architectural constraints make recurrent models inferior to alternatives

5 theses · 5 institutions

Recurrent architectures face high inference state buffering overheads, lengthy training times, and numerical instability during higher-order hypergradient computation. Consequently, practitioners favor convolutional networks or transformers due to better spatial feature extraction, easier handling of missing data, and superior sequence modeling performance.

Considered and rejected

Considered and rejected: Rejected using Recurrent Neural Networks (RNNs) with backpropagation through time in GF due to numerical instability during higher-order hypergradient computation

On machine learning methods for time series with financial applications · Oxford

Considered and rejected

Considered and rejected: Rejected using recurrent neural networks (RNNs) for marker prediction due to difficulty handling missing values without discontinuities and higher computational cost.

From finger animation to full-body embodiment of avatars with different morphologies and proportions · EPFL

Considered and rejected

Considered and rejected: Rejected Recurrent Neural Networks (RNNs) because they use a single weight matrix across units, limiting spatial feature detection and requiring inordinate training time compared to CNNs.

Predicting Phenotypes From Novel Genomic Markers Using Deep Learning · HARVEST

Considered and rejected

Considered and rejected: BiLSTM-CRF architectures were rejected for named entity recognition in favor of BERT-based models because transformer-based models consistently outperform sequential recurrent neural networks.

A pipeline for data and knowledge extraction from material science literature to accelerate scientific discovery · Georgia Tech

Tried and failed

recurrent neural networks applied to hardware branch prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and high inference-engine state buffering overheads compared to CNN architectures

Using convolutional neural networks to improve branch prediction · UT Austin

Recurrent models fail to generalize across domain shifts and specialized data distributions

5 theses · 3 institutions

Continuous recurrent networks generate invalid negative values when applied to zero-inflated intermittent time series, and anomaly classifiers fail to distinguish physically similar sensor fault signatures. Additionally, models struggle with cross-task transfer, acoustic noise degradation, and synthetic text generated by modern architectures.

Tried and failed

continuous recurrent neural network regression applied to zero-inflated intermittent time series. Outcome: did not generalise. Reason: assuming a single continuous distribution failed to predict true zeros and generated invalid negative values

Machine Learning Methods for Modeling Streamflow in Intermittent Rivers and Ephemeral Streams · Texas Tech

Tried and failed

transfer learning using pretrained recurrent layers applied to cross-task text classification. Outcome: worse than baseline. Reason: None

Towards Explainable Event Detection and Extraction · Virginia Tech

Tried and failed

recurrent neural network text classifier applied to transformer-generated synthetic text detection. Outcome: did not generalise. Reason: architectures trained on older generation models fail to detect modern transformer-based outputs

Defending Against Misuse of Synthetic Media: Characterizing Real-world Challenges and Building Robust Defenses · Virginia Tech

Tried and failed

neuro-inspired feature extraction with recurrent networks applied to speech recognition under acoustic degradation. Outcome: did not generalise. Reason: models degraded on noise-vocoded speech but tolerated periodic speech, contradicting human perceptual robustness

Features of hearing: applications of machine learning to uncover the building blocks of hearing · Imperial

Tried and failed

recurrent neural network for multi-class anomaly classification applied to cyber-physical system sensor incident classification. Outcome: did not generalise. Reason: Failed to distinguish between physically similar fault signatures of missing vs corrupted sensor readings

H2OGAN: A Deep Learning Approach for Detecting and Generating Cyber-Physical Anomalies · Virginia Tech

Regularization strategies and layered structural configurations degrade recurrent performance

4 theses · 3 institutions

Applying dropout regularization to recurrent networks fails to mitigate complexity in stacked multi-layer setups and harms performance when datasets are already large. Similarly, layering unconstrained recurrent units before constrained units or adding redundant sensor features causes performance degradation without improving generalization.

Tried and failed

Dropout regularization in deep multi-layer LSTM networks applied to time series imputation. Outcome: worse than baseline. Reason: Failed to mitigate added complexity and error inflation caused by stacking multiple recurrent layers

Data Driven Early Stage Design Support for Offshore Wind Farms · Research Repository UCD

Tried and failed

applying dropout regularization to recurrent neural networks applied to time series prediction on large datasets. Outcome: worse than baseline. Reason: regularization degraded performance without improving generalization due to the sufficiently large training dataset size

Traffic Signal Phase and Timing Prediction: A Machine Learning and Controller Logic Hybrid Approach · Virginia Tech

Tried and failed

layering unconstrained recurrent units before constrained units applied to acoustic feature modeling in speech recognition. Outcome: worse than baseline. Reason: None

Biologically Inspired Spiking Neural Networks for Speech Recognition · EPFL

Tried and failed

adding redundant sensor features to recurrent networks applied to kinematic trajectory prediction. Outcome: worse than baseline. Reason: additional sensor features introduced redundancy that caused performance degradation or offered no improvement

Gait Phase Estimation and Foot Trajectory Prediction During Dynamic Walking Using Gated Recurrent Units · Virginia Tech

Opaque hidden states reduce model interpretability compared to traditional methods

3 theses · 3 institutions

Recurrent neural networks are rejected in time series applications because their hidden state representations lack interpretability for multivariate processes. Practitioners instead select traditional models such as SARIMA or Gaussian mixture hidden Markov models to retain transparent dynamics and superior efficiency.

Considered and rejected

Considered and rejected: Rejected sequence models (Gated Recurrent Units / GRU) and SDEs for temperature, HVAC consumption, and electricity spot price processes in favor of traditional time series models (SARIMA, Holt-Winters, ARX-GJR-GARCH) for superior interpretability and efficiency.

Three Essays on Energy Markets · JScholarship

Considered and rejected

Considered and rejected: Rejected recurrent neural networks (LSTMs) for temporal modeling due to opaque hidden states lacking interpretability for multivariate data, choosing Gaussian mixture HMMs instead

Introducing Productive Engagement for Social Robots Supporting Learning · EPFL

Considered and rejected

Considered and rejected: Rejected using recurrent networks to learn change detection metrics because it reduces model interpretability.

Detecting and Leveraging Changes in Temporal Data · Georgia Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Benchmark and compare the dual temporal attention recurrent model against time-series Transformers for many-to-many sequence prediction. Blocker: None

Recurrence and Temporal Attention Synergy for Optimal Time-Series Modeling and Interpretability · TXST Digital Repository

Left open

Implement recurrent architectures like LSTMs using historical sequential sensor data for RUL estimation and anomaly reconstruction. Blocker: None

Toward The Democratization of Industrial Machine Learning · Georgia Tech

Left open

Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None

Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard

Left open

Develop a robust training method for recurrent neural networks processing static images with both recurrence and internal noise. Blocker: None

EXPLAINING FEATURES OF SIMPLE HUMAN DECISIONS USING BAYESIAN NEURAL NETWORKS · Georgia Tech

Left open

Add a recurrent decoder or sliding attention mechanism over sequential window embeddings to incorporate temporal context in stereo-EEG seizure detection. Blocker: Requires clinical stereo-EEG dataset with sub-second channel-level seizure annotations used in the thesis

Channel-Level Seizure Onset and Offset Detection in Stereo-EEG with Sub-Second Temporal Precision Using Machine Learning · Harvard

Left open

Develop recurrent neural networks or RNN-based GANs to map time-based geomechanical oilfield data. Blocker: No specific architecture, dataset, or performance targets are specified

Enhanced Oil Field Data-Wrangling using Machine Learning · Texas Tech

Left open

Implement and evaluate a recurrent neural network architecture conditioned on state history for the dynamics error predictor in Error-Aware Policies. Blocker: None

Learning Control Policies for Fall Prevention and Safety in Bipedal Locomotion · Georgia Tech

Left open

Train an LSTM-based recurrent neural network to automate coarse visual categorization of astronomical seed light curves. Blocker: None

ExoSpotter: Few Shot Relevance Feedback For Learning High Recall Exoplanet Search · MIT

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.