Chapter Four · failure evidence

What Probabilistic Latent Modeling got wrong, from 67 dissertations

The records document recurring failures and trade-offs encountered when applying probabilistic latent variable models across text, spatial, tabular, and dynamical data. Researchers frequently find that standard latent modeling approaches struggle with data sparsity, distribution misspecification, optimization instability, and competition from simpler deterministic baselines. These records come from PhD theses at 24 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Latent Dirichlet Allocation fails on short or sparse text documents

10 theses · 8 institutions

Short texts such as tweets, survey responses, comments, and search queries lack sufficient word co-occurrence and document length for Latent Dirichlet Allocation to infer coherent topic representations. As a result, pairwise similarities diverge from human judgment, topic coherence degrades, and the method is consistently rejected in favor of mixture models or dense neural embeddings.

Tried and failed

Latent Dirichlet Allocation topic modelling applied to short-text survey responses. Outcome: worse than baseline. Reason: standard LDA struggles with sparsity and lack of word co-occurrence in short texts

Examining structural, systemic, and enabling approaches to sustainable food system transformations · Imperial

Tried and failed

Latent Dirichlet Allocation topic modeling applied to short social media texts. Outcome: worse than baseline. Reason: sparse word co-occurrences in short documents led to poor topic coherence compared to NMF and neural embeddings

Assessing mental wellbeing in urban areas using social media data: understanding when and where urbanites stress and de-stress · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) text topic features due to poor perplexity/interpretability on 89 profile tweet aggregations.

Personality-Driven Social Media Curation: How Personality Traits Affect Following Decisions on Twitter · Scholars' Bank

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for short sanction text messages, choosing biterm topic modeling instead.

Three Papers on Peer Sanctioning, its Evaluation, and its Justification · DukeSpace

Tried and failed

Latent Dirichlet Allocation for semantic similarity applied to short text categorization and clustering. Outcome: no signal. Reason: LDA pairwise similarity matrices differed substantially from human judgment correlation matrices

Understanding Cyber Attacks Using Text Mining · Texas Tech

Tried and failed

Latent Dirichlet Allocation topic modeling applied to short user comment text corpus. Outcome: no signal. Reason: failed to reveal distinct latent themes or nuanced categorization across user comments

Public Use of Open Access Research: Evidence from the National Academies and Harvard DASH Repository · Georgia Tech

Tried and failed

Latent Dirichlet Allocation topic modeling applied to short conversational text queries. Outcome: no signal. Reason: text length was too brief and sparse to extract coherent topics

Evaluating the explainability of AI-driven clinical decision support tools · Imperial

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) in favor of the Gibbs Sampling Dirichlet Multinomial Mixture (GSDMM) model for topic modeling due to GSDMM's superior handling of short, dense tweet text

Decision support for sustainable energy systems, energy economics, urban mobility, and emerging methods in information systems research · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for short survey priority responses due to poor performance on documents with fewer than 50 words.

Examining structural, systemic, and enabling approaches to sustainable food system transformations · Imperial

Considered and rejected

Considered and rejected: Rejected standard Latent Dirichlet Allocation (LDA) for clustering short extracted action spans due to severe sparsity and inability of synonyms to co-occur in 1-10 word spans.

Three Essays on Natural Language Processing and Information Extraction with Applications to Political Violence and International Security · MIT

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for query topic modeling because short search queries (avg 5.2 words) yield sparse vectors and ignore word sequence order.

EFFECTS OF INFORMATION ON DIGITAL PLATFORMS AND ONLINE ECOSYSTEMS · Penn

Linear factor models and principal component analysis fail to capture genuine latent factor structures

9 theses · 7 institutions

Linear projections and principal component methods frequently capture sample noise instead of true weak factors, resulting in poor out-of-sample generalization and failure to converge to true components. Furthermore, deterministic principal components fail to account for measurement error, time-series structure, or parameter uncertainty when modeling underlying latent factors.

Tried and failed

non-negative and empirical Bayes matrix factorization applied to variant-phenotype cluster decomposition. Outcome: worse than baseline. Reason: failed to improve biological relevance of clusters compared to Latent Dirichlet Allocation

Characterizing Regulatory Elements and Non-Coding Variants in the Human Genome · Harvard

Tried and failed

projecting latent factor decompositions to recover unprojected components applied to stochastic discount factor decomposition. Outcome: did not converge. Reason: projected weak-factor and idiosyncratic components failed to converge to their true unprojected counterparts

Asset pricing with unsystematic risk · Imperial

Tried and failed

retaining high-order principal components for weak signals applied to linear factor asset pricing models. Outcome: did not generalise. Reason: high-order components captured primarily sample-specific noise rather than true weak latent factors, causing severe out-of-sample degradation

Asset pricing with unsystematic risk · Imperial

Lost to a baseline

Linear PCA dimensionality reduction on agent-based stochastic data found a latent space size of n - 1 (25x larger than non-linear DR), resulting in poor reducibility compared to non-linear methods.

Reduced Order Non-INtrusive (RONIN) Modeling for Strategic Defense Planning · Georgia Tech

Considered and rejected

Considered and rejected: Principal component analysis was rejected in favor of common-factor analysis (unweighted least squares) because error variance was unknown and generic latent constructs were sought

Measuring job performance associated to the direct effect of the global context (the emergence of the job-context model). · oURspace

Considered and rejected

Considered and rejected: Rejected non-parametric Principal Component Analysis (PCA) for latent factor extraction due to lack of time-series structure and inability to propagate parameter uncertainty

Parameter Uncertainty, Cashflow Betas, and Earnings Announcement Premia · Scholars' Bank

Considered and rejected

Considered and rejected: Rejected principal component analysis (PCA) because it cannot model underlying latent factors via covariance

An Exploration of Fear of Sleep and Experiential Avoidance in the Context of PTSD and Insomnia Symptoms · Scholars' Bank

Considered and rejected

Considered and rejected: Rejected factor analysis (FA) in favor of PCA for factor score generation to describe empirical dataset variance without imposing latent variable normality assumptions.

Whole-water systems modelling for sustainable catchment management · Imperial

Considered and rejected

Considered and rejected: Deterministic PCA with arbitrary truncation threshold (rejected because it incorporates undesired noise and lacks probabilistic modeling of model reduction error)

Data-driven Multi-Scale Modeling with Uncertainty Quantification of Damage in Unidirectional and Woven Composites Using Parametrically-Upscaled Continuum Damage Mechanics (PUCDM) Model · JScholarship

Considered and rejected

Considered and rejected: Rejected Principal Component Analysis (PCA) in favor of Exploratory Factor Analysis (EFA) because PCA does not model latent constructs with conceptual meaning and assumes zero measurement error.

Evaluating the Efficacy of Talent Identification and Development in the National Hockey League Entry Draft · YorkSpace

Variational autoencoders underperform simpler deterministic and linear baselines

10 theses · 8 institutions

Probabilistic latent representations from variational autoencoders frequently underperform simpler alternatives like deterministic autoencoders, principal component analysis, and nearest neighbor baselines in reconstruction and prediction tasks. Adding variational inference overhead often increases prediction uncertainty, degrades cluster separation, and fails to provide benefits over standard point-estimate baselines.

Tried and failed

sequence-to-sequence variational autoencoder applied to sketch to latent space mapping. Outcome: worse than baseline. Reason: failed to show significant improvement over simpler vector KNN regression despite higher training complexity

Computational Gestural Making: A framework for exploring the creative potential of gestures, materials, and computational tools · MIT

Tried and failed

surrogate modeling on variational autoencoder latent space applied to property prediction from spatial representations. Outcome: worse than baseline. Reason: nonlinear latent embedding degraded forward predictions and increased uncertainty compared to linear dimensionality reduction

Neural Inverse Microstructure Design with Bayesian Scale-Bridging · Georgia Tech

Tried and failed

variational probabilistic embedding applied to human comparison choice modeling. Outcome: worse than baseline. Reason: the underlying choice likelihood model is identical to point-estimate maximum likelihood estimation

Stochastic Models for Comparison-based Search · EPFL

Lost to a baseline

In pure embedding quality (Recall@1), probabilistic models slightly lagged behind standard supervised Cross-Entropy baseline on ViT-Medium

Uncertainties of Latent Representations in Computer Vision · Publikationssystem UB Tuebingen

Lost to a baseline

VAE + Gaussian process regression on latent variables lost to a simple null baseline (predicting test QY using the most similar sequence in the training set), which achieved matching correlation.

INVESTIGATING THE STRUCTURAL REQUISITES OF PHOTOPHYSICAL PROPERTIES IN FLUORESCENT BIOMOLECULES · JScholarship

Lost to a baseline

VAE reconstruction MSE underperformed PCA and NMF when using ≥ 5 latent parameters

Characterising & classifying the local population of ultracool dwarfs with Gaia DR2 and EDR3 · Imperial

Lost to a baseline

AE and VAE Latent Noise Segmentation achieved lower ARI (0.926 and 0.918) than simple noisy reconstruction clustering (0.982 and 0.977) on the Gradient Occlusion dataset.

Differences in Visual Perception in Humans and Deep Neural Networks · EPFL

Lost to a baseline

Large VAE with latent size 128 (44.3 AUPRC) was beaten by PCA difference vector baseline (67.3 AUPRC) on the Floods dataset with k=3 history.

Intelligent decision making on-board satellites · Oxford

Considered and rejected

Considered and rejected: Rejected stochastic latent sampling in VAEs, choosing deterministic autoencoders with explicit decoder regularization and ex-post density estimation.

Reining in the Deep Generative Models · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Rejected standard VAE in favor of deterministic autoencoder because generative latent distribution modeling was unnecessary for embedding compression.

Neural Compression for Scalable Question-Answer Retrieval · DalSpace

Variational latent variable models suffer from training instability, sampling issues, and posterior collapse

8 theses · 7 institutions

Enforcing sparsity, autoregressive covariance, or stochastic sampling in variational latent spaces often leads to unstable training and fails to resolve posterior collapse. In addition, random latent sampling in sparse data regimes causes low generation success rates, while inference via Monte Carlo sampling or latent inversion introduces high computational complexity and convergence failures.

Tried and failed

L1 or L2 regularization on posterior parameters applied to variational autoencoder latent representations. Outcome: worse than baseline. Reason: Failed to increase representation sparsity compared to vanilla variational autoencoder baseline.

Injecting Inductive Biases into Distributed Representations of Text · Cambridge

Tried and failed

variational autoencoder for sparse latent representation applied to text sentence embeddings. Outcome: unstable. Reason: struggled to achieve steady and consistent sparsity across hyperparameter configurations

Injecting Inductive Biases into Distributed Representations of Text · Cambridge

Tried and failed

autoregressive latent covariance structure applied to variational autoencoders. Reason: did not prevent posterior collapse when it occurred in standard VAE baseline

Deep Time: Deep Learning Extensions to Time Series Factor Analysis with Applications to Uncertainty Quantification in Economic and Financial Modeling · Virginia Tech

Tried and failed

i.i.d. latent sampling in variational autoencoders applied to set and graph generation. Reason: It optimizes an excessively loose lower bound on the true evidence lower bound.

Equivariant Neural Architectures for Representing and Generating Graphs · EPFL

Lost to a baseline

Separately trained predictor and VAE required significantly more optimization steps and exhibited higher variance during latent space search compared to simultaneously trained VAE with predictor

Integrating AI across the Chemistry Discovery Cycle: Advancing Sustainable Chemistry through Digital Methods · EPFL

Considered and rejected

Considered and rejected: Rejected Monte Carlo sampling from the marginal word distribution during CVAE inference in favor of feeding expected latent alignments to avoid high computational complexity

LOOKING INTO ACTORS, OBJECTS AND THEIR INTERACTIONS FOR VIDEO UNDERSTANDING · JScholarship

Tried and failed

latent space sampling in variational autoencoder applied to random 3D mesh generation. Outcome: data insufficient. Reason: training data sparsity led to low valid generation success rate when sampling the latent space randomly

Towards human-centered generative design : cross-modal synthesis for three-dimensional design concept generation · UT Austin

Tried and failed

latent space inversion via optimization and encoder applied to deep progressive generative adversarial networks. Outcome: did not converge. Reason: optimization and encoder-based methods fail to accurately invert deep multi-scale generative models

Dissection of Deep Neural Networks · MIT

Considered and rejected

Considered and rejected: Rejected Variational Autoencoders (VAEs) for sound-masking noise generation because GANs avoid deterministic bias and perform better with discrete latent variables.

Robust Defenses Against Adversarial Machine Learning in Internet of Things Security · Carleton University Institutional Repository

Standard Gaussian latent priors cause distributional mismatch and over-smoothing

8 theses · 6 institutions

Assuming continuous Gaussian priors creates dense and overlapping latent spaces that hinder distinct cluster separation and degrade performance on non-Gaussian or tabular systems. In dynamical and fluid modeling, the enforcing of smooth Gaussian latent spaces leads to underpredicted kinetic energy and loss of high-frequency spectral content.

Lost to a baseline

Vanilla autoencoders achieved slightly higher silhouette scores for cluster separation compared to VAE due to VAE's continuous, overlapping latent space.

Machine Understanding of Architectural Space: From Analytical to Generative Applications · EPFL

Considered and rejected

Considered and rejected: Rejected enforcing a Gaussian prior via adversarial discriminators on the style latent bottleneck due to training instability and degraded RMSE

Using wearable sensors and machine learning to enrich lower limb assistive devices · ResearchWorks

Tried and failed

conditional variational autoencoder with factorised latent space applied to 3D multi-object scene completion. Reason: failed to learn a structured prior, yielding incomplete reconstructions when sampling from Gaussian noise

Scene understanding for 3D multi-object scenes: labelling, reasoning and decomposing · Imperial

Tried and failed

variational autoencoders with Gaussian priors applied to systems with strongly non-Gaussian statistics. Outcome: did not generalise. Reason: standard Gaussian latent priors poorly represent strongly non-Gaussian statistics

Physics-Driven Machine Learning for Applications in Geophysical Fluid Dynamics · MIT

Tried and failed

variational autoencoder for dimensionality reduction before clustering applied to high-dimensional mass spectrometry imaging data. Reason: Gaussian prior forced an overly dense latent space, preventing separation into distinct clusters

Mass spectral imaging of clinical samples using deep learning · Imperial

Tried and failed

variational autoencoder transformer reduced order model applied to turbulent fluid flow prediction. Reason: Smooth latent representations underpredicted kinetic energy and lost high-frequency spectral content

A Dynamic Model of Unpiloted Aerial Systems for Complex Ship Airwake Environments · Georgia Tech

Lost to a baseline

Variational Autoencoders (VAEs) lose to standard Autoencoders (AEs) when training SNR and testing SNR perfectly match due to gaps in the latent space

Model and data driven approaches to wireless image transmission · Imperial

Considered and rejected

Considered and rejected: Rejected Variational Autoencoders (VAEs) in FASTER-CE to avoid assuming Gaussian distributions over tabular latent spaces

Multi-objective approaches towards trustworthy machine learning · UT Austin

Latent Dirichlet Allocation yields uninterpretable, unstable, or poorly aligned topic structures

7 theses · 7 institutions

Unsupervised topic models often generate topics filled with generic, loosely related words that fail to map onto predefined domain categories or theoretical constructs. Models also exhibit instability over longitudinal corpora, where topic tracking is sensitive to initial text sequence ordering and fails to produce consistent results across time.

Tried and failed

Latent Dirichlet Allocation topic modeling applied to organizational text corpora. Outcome: no signal. Reason: generated uninterpretable topics with generic words failing to map to distinct theoretical constructs

Embracing community-centered approaches: Institutional logics and organizational responses in the U.S. museum field · Cornell

Tried and failed

unsupervised Latent Dirichlet Allocation applied to text classification against predefined taxonomy. Reason: unsupervised topic clusters did not correspond one-to-one with predefined target taxonomy categories

Analytics-Enabled Quality and Safety Management Methods for High-Stakes Manufacturing Applications · MIT

Considered and rejected

Considered and rejected: Rejected using Latent Dirichlet Allocation (LDA) for document topic modeling due to non-accurate and inconsistent inferring results.

Using Machine Learning And Natural Language Processing To Improve Scientific Processes · Penn

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for topic modeling due to reliance on bag-of-words, requirement of prior topic number determination, and ignoring word ordering/semantics in favor of Top2Vec.

Three Essays on the Use of Health Information Technology to Improve Patient Safety: Cases of Technology Design and Participation in Professional-Only Healthcare Forums · DSpace at SUNY Buffalo

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for issue report topic extraction and review aspect extraction because it produces loosely-related words when vocabulary is large

Machine Learning Driven Guidance for Software Maintenance: Enhancing Code Management and Features · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) in favor of Non-negative Matrix Factorization (NMF) due to NMF's higher topic coherence and model stability.

Essays on central bank communication · UT Austin

Tried and failed

undirected latent Dirichlet allocation applied to longitudinal corporate annual report text. Outcome: unstable. Reason: critically dependent on initial text sequence and could not reliably track recurring topics across decades

The application of quantitative methods to the study of firm-level adaptation and investment preferences in the energy sector · Imperial

Gaussian process latent models and Bayesian optimization surrogates fail in dynamic or complex settings

5 theses · 5 institutions

Approximations to deep Gaussian process posteriors produce poor uncertainty quantification and inaccurate variance estimates for acquisition functions in Bayesian optimization. Standard Gaussian process regression also fails to track time-varying environmental shifts and incurs severe hyperparameter complexity in high-dimensional spaces.

Tried and failed

Gaussian approximation of deep Gaussian process posteriors applied to Bayesian optimization uncertainty estimation. Reason: yielded poor uncertainty quantification and inaccurate variance estimates for acquisition functions

Physics-informed Machine Learning for Digital Twins of Metal Additive Manufacturing · Virginia Tech

Tried and failed

Standard Bayesian optimization without contextual modeling applied to optimization under dynamic periodic environmental drift. Outcome: did not converge. Reason: Standard Gaussian process regression could not track or account for dynamic time-varying environmental context shifts

Machine Learning Applications for Improving Accelerator Operations · Cornell

Lost to a baseline

Gaussian Process surrogate underperformed GradientBoosting during Bayesian optimization, dipping below p=0.05 around iteration 50.

Data-Driven Design of Recycling-Friendly Aluminium Alloys · MIT

Lost to a baseline

One latent GP framework performed worse (median RMS = 32.9 m s-1) than the baseline FF' model (RMS = 28.5 m s-1) on synchronous synthetic test data.

The epoch of giant planet migration : searching for young planets within the stellar noise · UT Austin

Considered and rejected

Considered and rejected: Rejected Gaussian process latent variable models (Bayesian approach) for mocap synthesis due to hyperparameter complexity, kernel selection issues, and loss of nuances in high-dimensional space

Predicting lower limb kinematics and kinetics from internal measurement units using deep learning · Imperial

Left open by the authors

Problems the authors named and did not get to.

Left open

Perform biological interpretation and functional annotation of latent topics and theme k-mers identified from shotgun metagenomic topic models. Blocker: None

Microbiome and Metagenomics: Statistical Methods, Computation and Applications · Penn

Left open

Investigate the causal implications and identification strategies of latent topics extracted from text data. Blocker: Lacks specific causal hypotheses, targets, identification strategies, or experimental frameworks

LEARNING AND INFERENCE IN DIGITAL MARKETS: METHODS AND APPLICATIONS · Cornell

Left open

Benchmark co-segregation-based Bayesian linear regression feature selection against standard methods like PCA and random forests for microbial phenotypic GWAS. Blocker: None

Probing Genomic, Metabolic, And Phenotypic Evolution In Microbes Using Comparative And Experimental Evolution Data · Georgia Tech

Left open

Develop a probabilistic model of microbial biomass growth to scale column-level dynamics to ecosystem-level microbial community behavior. Blocker: Lack of mathematical formulation, target datasets, and specific probabilistic framework defined in the thesis

Investigation of microbial community dynamics in soil during variable hydrological forcing · EPFL

Left open

Extract literature corpora from academic databases beyond Scopus to analyze topic trends via Latent Dirichlet Allocation. Blocker: None

Error component models for traffic crash severity analysis · Texas Tech

Left open

Incorporate monotonicity, fairness, or probabilistic/Bayesian priors into constrained low-rank approximation frameworks. Blocker: Lacks specific mathematical formulations, optimization objectives, and evaluation datasets or metrics.

Algorithms for Data Fusion, Representation Learning, and Scalable Clustering based on Constrained Low-Rank Approximation · Georgia Tech

Left open

Initialize variational autoencoders or graph-based multi-scale clustering algorithms with latent factors learned from Integrative Hierarchical Poisson Factorisation. Blocker: None

Robust machine learning methods for high-dimensional datasets with applications in genomics and finance · Imperial

Left open

Extract high-level semantic features from visual stimuli to model and explain latent aesthetic taste criteria among collaborative filtering user cohorts. Blocker: None

TOWARDS UNDERSTANDING AND PREDICTING INDIVIDUALIZED AESTHETIC RESPONSES TO ECOLOGICALLY VALID VISUAL STIMULI · Cornell

Left open

Develop unsupervised methods to select subsets of latent dimensions in variational autoencoders for anomaly detection. Blocker: None

Anomaly Detection via Latent Variables Learned by Variational Autoencoders · Queens University Institutional Repository

Left open

Model agent intent using a continuous latent space with variational autoencoders for approximate inference in multiagent reinforcement learning. Blocker: None

Robust and Scalable Multiagent Reinforcement Learning in Adversarial Scenarios · MIT

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.