Chapter Four · failure evidence
What Hierarchical Clustering got wrong, from 66 dissertations
The records document numerous operational and methodological difficulties encountered when applying hierarchical clustering across diverse scientific domains. Researchers frequently abandoned or rejected hierarchical algorithms due to prohibitive computational scaling, vulnerability to noise, rigid irreversible merges, and inferior cluster recovery compared to simpler or alternative baselines. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
High computational runtime and memory overhead make hierarchical clustering impractical on large datasets
Agglomerative and divisive hierarchical algorithms suffered from steep polynomial or exponential time complexity, full pairwise distance calculation bottlenecks, and memory exhaustion that caused repeated crashes on large graphs and images. Authors rejected these methods for real-time edge prediction, large search graphs, and large-scale bipartite projections in favor of faster partitioning baselines.
Considered and rejected
Considered and rejected: Rejected agglomerative hierarchical clustering with dendrograms because it crashed repeatedly due to dataset size
A MACHINE LEARNING APPROACH TO NETWORK SECURITY CLASSIFICATION USING NETFLOW DATA · Calhoun
Considered and rejected
Considered and rejected: Rejected divisive hierarchical clustering due to its algorithmic time complexity of O(2^N) compared to O(N^2 log N) - O(N^3) for agglomerative clustering
A machine learning approach to develop turbulence closures using clustering and neural networks · Imperial
Tried and failed
hierarchical agglomeration modularity maximization applied to large-scale bipartite projection graphs. Outcome: too slow. Reason: computational time and memory scaled poorly on very large dense projected constraint graphs
Data-Driven Decision Analytics in Complex Systems: Information Infrastructure, Problem Decomposition, and Algorithmic Design · Georgia Tech
Lost to a baseline
Hierarchical clustering with embedded cluster evaluation matches T-SNE/DBSCAN speed on small systems (5-bus) but scales as O(n^3) while T-SNE/DBSCAN remains linear/polynomial
Operational moving target defences for improved power system cyber-physical security · Imperial
Considered and rejected
Considered and rejected: Rejected divisive hierarchical clustering strategies due to computational expense (2^{N-1} - 1 possible partitions).
Robust Bayesian Anomaly Detection Methods for Large Scale Sensor Systems · Virginia Tech
Considered and rejected
Considered and rejected: Rejected hierarchical agglomerative clustering because computing the full set of query results is intractable over large search graphs.
Considered and rejected
Considered and rejected: Decided against classical clustering (k-medoids, hierarchical clustering) for online property grouping because of O(n^2) distance matrix computation and inability to predict optimal group counts.
Model checking large design spaces: Theory, tools, and experiments · Iowa State
Considered and rejected
Considered and rejected: Standard Euclidean hierarchical clustering for mass spectrometry imaging pixel data due to high computational overhead (billions of pairwise distance calculations).
Comparative analysis of sporadic and colitis associated cancers · Imperial
Considered and rejected
Considered and rejected: Rejected AI/generative models (e.g., Ward hierarchical clustering, generative networks) for real-time edge/cloud prediction due to high O(n^2)-O(n^3) computational overhead.
Contribuciones al ahorro de energía en redes de sensores inalámbricos aplicados a la piscicultura · accedaCRIS
Considered and rejected
Considered and rejected: Rejected Ward's minimum variance hierarchical clustering method due to excessive computing time compared to K-means.
Considered and rejected
Considered and rejected: Rejected divisive hierarchical clustering and density-based clustering (DBSCAN) due to computational expense and sensitivity to MinPts/Eps parameters.
Noise, outliers, and dominant confounding features corrupt hierarchical cluster formation
Hierarchical clustering frequently allowed additive noise, weak sensor outliers, and dominant spatial or patient-level variances to overpower meaningful signals and misassign atypical data points. Combining heterogeneous feature sets or running clustering directly on unsmoothed time series produced impure, semantically incoherent, or structurally distorted groupings.
Considered and rejected
Considered and rejected: Rejected single-linkage and complete-linkage hierarchical clustering due to extreme sensitivity to noise/outliers when devices share few overlapping points.
Characterizing and Detecting Physical Layer Issues in Cable Broadband Networks · DukeSpace
Considered and rejected
Considered and rejected: Rejected using latitude and longitude features in hierarchical clustering for rainfall up-scaling because spatial coordinates dominated and degraded precipitation pattern grouping.
Underpinning research on the dynamical aspects of one-dimensional nonlinear lattices · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected K-means and hierarchical clustering for final spatio-temporal hotspot analysis because they cannot remove noise points and require pre-specifying cluster counts.
IDENTIFYING PROBABLE MARITIME PIRACY EVENTS USING MARITIME INCIDENT DATA · Calhoun
Tried and failed
density-based and hierarchical clustering algorithms applied to low-energy sensor impact classification. Outcome: unstable. Reason: overly sensitive to hyperparameter selection and outlier noise in weak signals
Tried and failed
hierarchical clustering on raw noise residuals applied to synthetic image detection. Outcome: no signal. Reason: additive random noise dominated residual space causing impure mixed clusters
Defending Against Misuse of Synthetic Media: Characterizing Real-world Challenges and Building Robust Defenses · Virginia Tech
Tried and failed
Hierarchical clustering on unsmoothed time series applied to Short lifecycle demand curves. Outcome: no signal. Reason: Unsmoothed noise prevented the dendrogram from separating curves according to underlying structural patterns.
Optimizing Sales Forecasting, Inventory, Pricing and Sourcing Decisions · EPFL
Tried and failed
hierarchical clustering with combined feature sets applied to heterogeneous behavioral time-series data. Outcome: no signal. Reason: Simultaneous clustering of mixed intensity and temporal variables obscured meaningful segment structure.
Tried and failed
Hierarchical stochastic block modeling for community detection applied to social network co-following graphs. Outcome: no signal. Reason: Produced noisy and semantically incoherent cluster groupings combining unrelated topics
Tried and failed
hierarchical clustering and t-SNE dimensionality reduction applied to high-dimensional perturbation log-fold change data. Outcome: no signal. Reason: Unsupervised clustering failed to separate distinct biological groups based on variable feature fold changes
Studying the tissue-specificity of cancer driver genes through KRAS and genetic dependency screens. · Harvard
Tried and failed
agglomerative hierarchical clustering with Ward's linkage applied to quantitative MRI tumor habitat segmentation. Outcome: worse than baseline. Reason: misassigned outlier high-value voxels to low-value clusters rather than isolating distinct habitats
Toward advances in data acquisition and analysis for quantitative multi-contrast magnetic resonance imaging · UT Austin
Considered and rejected
Considered and rejected: Standard unsupervised clustering (K-means, Hierarchical) on raw patient images or GLM-residualized profiles because it captures dominant non-pathology variance (e.g., brain size, age) and assumes linear covariate relationships.
Analyzing Heterogeneity In Neuroimaging With Probabilistic Multivariate Clustering Approaches · Penn
Irreversible hierarchical splits and merges generate severe cluster size imbalances and lock in early errors
Agglomerative and divisive clustering procedures cannot iteratively reassign points, locking in suboptimal early merges and producing highly skewed partitions dominated by single massive clusters or unreliably small fragments. Authors also reported that rigid hierarchical branching hindered direct parameter comparisons and caused minority trajectory groups to be dropped as noise.
Considered and rejected
Considered and rejected: Rejected hierarchical clustering for the result consistency heat maps because it grouped parameter combinations in a way that hindered direct comparison.
A Modular Data Analytic Pipeline for Feature Selection in High Dimensional Microbial Data Sets · HARVEST
Tried and failed
k-means and hierarchical clustering applied to pairwise co-occurrence and embedding distances. Reason: Produced highly unbalanced partitions dominated by one or two massive clusters
Tried and failed
k-means clustering initialized with hierarchical clustering applied to spatial tracking coordinate data. Outcome: worse than baseline. Reason: Lloyd iterations failed to improve upon the initial hierarchical assignments across all k values.
New methods in home-range overlap and clustering · Iowa State
Tried and failed
hierarchical agglomerative clustering applied to clinical feature subtyping. Reason: produced highly uneven cluster size distributions
A MULTI-MODAL APPROACH TO THE COMPLEXITIES OF ENDOMETRIOSIS: ENVIRONMENTAL, CLINICAL, AND MULTI-OMIC · Penn
Tried and failed
relaxed cluster radius termination thresholds applied to hierarchical bisecting k-means clustering. Outcome: worse than baseline. Reason: larger radius thresholds yielded no significant improvements in clustering quality or efficiency
Clustering and fast search of high dimensional big data · Texas Tech
Considered and rejected
Considered and rejected: Rejected discrete mode clustering via DBSCAN, k-medoids, spectral clustering, and hierarchical clustering in favor of k-means for simplicity and explicit cluster count control.
Modeling and Control of Networked Systems: Applications to Air Transportation · MIT
Considered and rejected
Considered and rejected: Avoided standard k-means 1D splits in hierarchical bisection when cluster sizes/densities differ significantly, preferring KDE-based thresholding.
Algorithms for Data Fusion, Representation Learning, and Scalable Clustering based on Constrained Low-Rank Approximation · Georgia Tech
Considered and rejected
Considered and rejected: Rejected hierarchical clustering because divisive/agglomerative splits/merges cannot be reassigned iteratively, opting instead for discrete partitioning via k-medoids (PAM).
Exposure and Attitudes: Informing Upstream Barriers to Safe Systems via Machine Learning, Bayesian Estimation, and Attitudinal Research · Georgia Tech
Considered and rejected
Considered and rejected: Rejected hierarchical clustering as the backbone engine for TDC because it cannot undo suboptimal early merges and has high computational complexity
Data-Driven Mobile Traffic Modelling in Cellular Networks · IRIS - POLITO - prod
Tried and failed
density-based hierarchical trajectory clustering applied to spatial trajectory datasets. Outcome: data insufficient. Reason: Low-volume minority trajectory patterns fell below the minimum cluster size threshold and were excluded as noise.
Hierarchical Behavior Models for Characterizing Trajectories within Terminal Airspace · MIT
Considered and rejected
Considered and rejected: Rejected stations/clusters containing <15 data points in hierarchical cluster analysis because they produce unrepresentatively small within-cluster variances.
Crustal seismic structure of the eastern Mediterranean: evidence from broadband seismology · Imperial
Specific linkage criteria fail to capture cluster structure or violate mathematical assumptions
Single, complete, and average linkages frequently produced overly sparse divisions or failed to account for underlying cluster geometry, while centroid linkage violated merge monotonicity. In addition, Ward linkage proved mathematically incompatible with Bray-Curtis dissimilarity and failed on varying cluster densities, while exponential linkage offered no benefit over standard average linkage.
Tried and failed
exponential linkage hierarchical agglomerative clustering applied to named entity normalization. Outcome: worse than baseline. Reason: failed to outperform simple average linkage
A pipeline for data and knowledge extraction from material science literature to accelerate scientific discovery · Georgia Tech
Tried and failed
exponential linkage hierarchical agglomerative clustering applied to entity normalization. Outcome: worse than baseline. Reason: failed to outperform standard average linkage despite interpolating between linkage criteria
A pipeline for data and knowledge extraction from material science literature to accelerate scientific discovery · Georgia Tech
Considered and rejected
Considered and rejected: Rejected single linkage and complete linkage hierarchical clustering because they fail to account for cluster structure.
Considered and rejected
Considered and rejected: Centroid linkage criterion for hierarchical clustering was rejected due to violation of monotonicity of merges.
A systematic design for applying personalised herbal medicine recommendations for type 2 diabetes · Imperial
Considered and rejected
Considered and rejected: Hierarchical clustering on the RCC genetic divergence matrix using Complete or Single linkage algorithms instead of Ward.D.
Statistical Methods For Multi-Omics Inference From Single Cell Transcriptome · Penn
Considered and rejected
Considered and rejected: Rejected using Ward's linkage method directly for hierarchical agglomerative cluster analysis because it is mathematically incompatible with Bray-Curtis dissimilarity.
Floristics, Conservation, and Restoration of Virginia's Piedmont Grasslands · Virginia Tech
Tried and failed
Hierarchical clustering with Ward's linkage applied to Vibration spectral feature embeddings. Reason: Failed to handle clusters with varying shapes and densities, dropping recall on minority classes
Vibration Event Detection and Classification in an Instrumented Building · Virginia Tech
Considered and rejected
Considered and rejected: Complete and average data linkage methods were evaluated for hierarchical clustering but were rejected because they resulted in more sparsely divided data.
Considered and rejected
Considered and rejected: Rejected complete, single, and average hierarchical clustering methods in favor of Ward's method based on agglomerative coefficient comparisons.
A Pedagogical Approach to Create and Assess Domain-Specific Data Science Learning Materials in the Biomedical Sciences · Virginia Tech
Considered and rejected
Considered and rejected: Rejected complete linkage, single linkage, and UPGMA hierarchical clustering algorithms in favor of Ward's minimum variance method due to Ward's higher agglomerative coefficient.
Large sample statistical estimation of streamflow sensitivity to climate and land cover change · Oxford
Hierarchical clustering achieves inferior separation and accuracy compared to simpler or graph-based baselines
Across spatio-temporal trajectories, retinal cell profiles, and benchmark mixture distributions, hierarchical clustering produced less distinct cluster separation and worse objective scores than k-means or level-set tree clustering. In other settings, hierarchical clustering attained high precision only at low recall compared to graph-based Leiden clustering or generated less explainable partitions than log-linear regression residuals.
Tried and failed
hierarchical clustering applied to spatio-temporal time series trajectories. Outcome: worse than baseline. Reason: produced less distinct cluster separations compared to k-means
Diagnosing Prevailing Trends and Disparate Impacts of COVID-19 at the County-level · Harvard
Tried and failed
approximate pairwise distance estimation for hierarchical clustering applied to genome-scale phylogenetic reconstruction. Outcome: worse than baseline. Reason: heuristic distance approximations failed to accurately preserve hierarchical topology and cophenetic correlation across ranks
Tackling the current limitations of bacterial taxonomy with genome-based classification and identification on a crowdsourcing Web service · Virginia Tech
Lost to a baseline
Hierarchical clustering of PCA-projected phenotype profile correlations achieved higher precision at lower recall (<0.2) compared to Leiden clustering on the PHATE diffusion affinity graph.
Image-based pooled genetic screens for complex cellular phenotypes · MIT
Lost to a baseline
Hierarchical clustering of phenotype profile correlations demonstrated higher precision at lower recall (<0.2) than Leiden clustering on the PHATE diffusion affinity graph.
Image-based pooled genetic screens for complex cellular phenotypes · MIT
Lost to a baseline
Ward's hierarchical clustering on non-linear CUT GAN manifold produced less explainable and poorly separated phenotypic clusters compared to the simpler log-linear regression residual baseline
Brain Metabolic Responses To Alzheimer Pathologies With Molecular Imaging And Machine Learning · Penn
Lost to a baseline
K-means, K-medoid, DBSCAN, and hierarchical clustering (single and Ward linkage) were beaten by level-set tree clustering (~95% cumulative accuracy on baseline gene OCR cluster recovery).
Lost to a baseline
Visual stimulus-based hierarchical clustering performed worse than electrical STA-based clustering at separating distinct ON visual types in rd10 degenerated retina (blending all ON cells into OFF clusters).
Classifying Retinal Ganglion Cells for Bionic Vision · Publikationssystem UB Tuebingen
Lost to a baseline
Hierarchical clustering with Ward's method achieved a suboptimal criterion sum of BDs (7.09 for k=2, 4.67 for k=3) compared to k-means algorithms (6.46 and 4.48/4.56).
New methods in home-range overlap and clustering · Iowa State
Lost to a baseline
On the 4-Gaussian mixture dataset in R^2 (n=6,000), Average Linkage hierarchical clustering (0.664 NMI) lost to k-Means (0.797 NMI), Spectral Clustering (0.805 NMI), Divisive Spectral (0.820 NMI), and Tangles (0.829 NMI).
Graphs in Unsupervised Machine Learning · Publikationssystem UB Tuebingen
High heterogeneity and profile instability prevent hierarchical clustering from recovering true biological or empirical structures
High transcriptional heterogeneity and inter-sample variation caused hierarchical clustering to lose biological lineage organization and produce clusters that lacked functional pathway enrichment. High cluster counts or noisy empirical station data similarly caused the hierarchy to collapse into uninterpretable micro-clusters or completely homogeneous groupings lacking statistical differences.
Tried and failed
hierarchical clustering of pseudotime-dependent gene expression applied to single-cell transcriptome trajectories. Outcome: no signal. Reason: high transcriptional heterogeneity prevented meaningful ordering by genotype or time point
Tried and failed
hierarchical clustering of cross-condition effect correlations applied to functional pathway discovery from QTLs. Outcome: no signal. Reason: most clusters lacked functional pathway enrichment across diverse environments
Tried and failed
hierarchical pseudobulk differential abundance analysis applied to single-cell RNA-seq treatment response data. Outcome: did not generalise. Reason: high inter-sample tissue variation produced counterintuitive marker enrichment across response groups
Lost to a baseline
Standard hierarchical agglomerative clustering (Ward's and average linkage) failed to identify higher-order lineage organization (5 of 6 lineage subgroups) in healthy airway scRNAseq cell types compared to K2Taxonomer.
Tried and failed
McQuitty hierarchical clustering applied to multivariate spatial raster data. Outcome: worse than baseline. Reason: Collapsed into a single cluster or generated clusters without statistically significant differences.
The essential role of soil maps to support precision agriculture · Iowa State
Tried and failed
hierarchical clustering with high cluster numbers applied to voxel response profile patterns. Outcome: unstable. Reason: subsequent cluster divisions yielded unstable or tiny clusters that were uninterpretable
Considered and rejected
Considered and rejected: Rejected using pure unsupervised clustering (hierarchical clustering + elbow/silhouette) because it optimizes variance rather than ensuring genetic homogeneity.
Considered and rejected
Considered and rejected: Decided against traditional hierarchical clustering for classifying all genes because it ignored metrics from clustering when genes lacked OCR clusters.
Tried and failed
resampled parameter stacking with hierarchical clustering applied to crustal seismic receiver functions. Outcome: unstable. Reason: produced unreliable results across diverse real-world stations despite prior literature reports
Crustal seismic structure of the eastern Mediterranean: evidence from broadband seismology · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop hierarchical metrics to quantify clusterness and trajectoriness on subsets and individual clusters in single-cell datasets. Blocker: None
Developing Graph-based Computational Algorithms for Single-cell Data Science · Georgia Tech
Left open
Analyze why Ward hierarchical clustering outperforms Neighbor-Joining on specific metric-learned cell embedding spaces. Blocker: None
Reconstructing Cell Lineage Trees from Phenotypic Features with Metric Learning · Penn
Left open
Compare the satellite clustering performance of k-means against hierarchical clustering algorithms like CHAMELEON or HDBSCAN on the longitudinal tracking dataset. Blocker: None
Left open
Evaluate hierarchical and density-based clustering algorithms against existing approaches for grouping user interests in news recommender systems. Blocker: None
Exploring Media Bias in News Recommender Systems: From Detection to Mitigation · Research Repository UCD
Left open
Compare urban land-use clustering solutions generated by DBSCAN, HDBSCAN, and hierarchical clustering algorithms on spatial data. Blocker: None
Left open
Develop quantitative stopping criteria to assess representativeness and quality of training samples selected via hierarchical clustering. Blocker: None
Left open
Develop improved methods for selecting representative data points within clusters to better estimate posterior distributions in hierarchical dataset selection. Blocker: The task is described purely as a vague direction without specific proposed alternative sampling methods or concrete targets.
Hierarchical Bayesian Dataset Selection · Virginia Tech
Left open
Integrate alternative clustering objectives beyond K-means and hierarchical clustering into the multi-task learning cancer subtyping framework. Blocker: The specific clustering algorithms and their mathematical integration into the multi-task objective are unspecified.
Supervised-Unsupervised Cancer Subtyping Based on Multi-Task Learning · DSpace at SUNY Buffalo
Left open
Extend scalable hierarchical clustering theoretical guarantees and algorithms to the Massively Parallel Computation model with strongly-sublinear local memory O(n^gamma) where gamma < 1. Blocker: None
Towards Scalability and Robustness for Ranking, Clustering, and Multi-Armed Bandits · Penn
Left open
Develop dynamic clustering extensions for non-metric or high-dimensional graph representations beyond approximate hierarchical agglomerative clustering. Blocker: No concrete algorithmic formulation, data structures, or performance target specifications are provided.
Parallel Algorithms, Optimizations, and Benchmarks for Metric and Graph Clustering · MIT
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.