Chapter Four · failure evidence

What Transfer Learning & Domain Adaptation got wrong, from 56 dissertations

The records document various failures of transfer learning and domain adaptation across diverse imaging, sequential, biological, and physical tasks. In many settings, transferred representations and adaptation algorithms introduced negative transfer, failed to bridge domain shifts, or were beaten by simpler target-specific baselines. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Pretrained natural image representations fail to transfer to specialized non-natural image modalities

9 theses · 5 institutions

Models pretrained on natural images or external datasets struggled to transfer effectively to specialized visual domains like thermal images, microscopy, RF spectrograms, and malware graphs. In many instances, the pretraining representations introduced inductive biases that caused overfitting or underperformed relative to training from scratch.

Tried and failed

transfer learning for image segmentation applied to noisy infrared thermal images. Outcome: worse than baseline. Reason: struggled to segment low signal-to-noise regions compared to hybrid fuzzy clustering

Autonomous Experimentation to Accelerate Boiling Heat Transfer Research · MIT

Tried and failed

transfer learning with natural image pretrained weights applied to satellite land cover image segmentation. Reason: None

Practical methods for the advancement of precision conservation via land cover classification and conformal prediction · Iowa State

Tried and failed

ImageNet pre-trained transfer learning applied to calcium imaging neural decoding. Outcome: worse than baseline. Reason: natural image visual representations transfer poorly to fluorescence microscopy signals compared to training from scratch

Optimizing sensorimotor behaviors through information integration and mental simulation · MIT

Tried and failed

transfer learning with natural image pretraining applied to malware graph visual representation classification. Outcome: worse than baseline. Reason: None

Developing Robust Models, Algorithms, Databases and Tools With Applications to Cybersecurity and Healthcare · Georgia Tech

Tried and failed

transfer learning with natural image pre-trained models applied to simulated microscopy image classification. Outcome: worse than baseline. Reason: Pre-training bias hindered adaptation compared to training from random initialization

Mitigating effects of data heterogeneity in decentralized learning and addressing challenges of adapting foundation models · Iowa State

Tried and failed

transfer learning with natural image pre-trained weights applied to RF spectral eigengram object detection. Outcome: worse than baseline. Reason: natural image visual features did not transfer well compared to random initialization for non-visual spectrogram representations

Intelligently Leveraging Multi-Channel Image Processing Neural Networks for Multi-View Co-Channel Signal Detection · Virginia Tech

Tried and failed

ImageNet pre-trained transfer learning applied to grayscale image classification. Outcome: did not generalise. Reason: Natural image features do not transfer effectively to grayscale, manufacturing, or medical domains

Engineering-Driven Learning Approaches for Bio-Manufacturing and Personalized Medicine · Georgia Tech

Tried and failed

transfer learning with pre-trained convolutional neural networks applied to droplet deposition density map classification. Outcome: overfit. Reason: models underperformed severely on test data and suffered from poor generalization and overfitting

Automated Exploration of High-Mix, Low Volume Direct Write Design Spaces Through Artificial Intelligence · Georgia Tech

Considered and rejected

Considered and rejected: Rejected using transfer learning from external datasets for the 1000-species classification model to avoid introducing domain training biases.

Deep Learning and Continual Learning Techniques for Plant Image Analysis Tiefes Lernen und kontinuierliche Lernmethoden für die Analyse von Pflanzenbildern · open_UMR Marburg DSpace 10.0

Transfer learning and domain adaptation underperform training from scratch or local target baselines

8 theses · 8 institutions

Across multiple applications, transfer learning models failed to improve upon simple baselines trained solely on the target dataset or simpler non-transfer representations. Complex adaptation schemes and pretraining were routinely matched or beaten by standard empirical risk minimization, random weight initialization, and simple dataset merging.

Tried and failed

transfer learning via large-scale pretraining applied to industrial semantic segmentation. Outcome: worse than baseline. Reason: pretraining on related domain datasets underperformed training from scratch on the target domain

Beyond standard benchmarking : towards robust and trustworthy robotics for industrial and nuclear applications · UT Austin

Lost to a baseline

Transfer learning on Goodfellow CNN (10s) on ICU test set achieved F1=0.85, matching zero-shot ECG-FM v1 (F1=0.85) and beaten by non-transfer Inception v3 recurrence plots (F1=0.88).

Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients: A Comparison of Artificial Intelligence Approaches · Queens University Institutional Repository

Lost to a baseline

Transfer learning models for zinc(II) salphen emission energy (Fingerprint + Ridge R2 = 0.831 ± 0.095, MAE = 0.0769 ± 0.0216 eV) showed no significant improvement over the baseline model trained purely on the 49 zinc(II) salphen dataset (Fingerprint + Ridge R2 = 0.825 ± 0.216, MAE = 0.0706 ± 0.0429 eV)

Discovery of photoactive anti-cancer platinum(II) complexes: synthesis, screening and machine learning approaches · Imperial

Considered and rejected

Considered and rejected: Rejected pure transfer learning due to decreased classification performance compared to retraining dataset-specific models.

Automated Design and Optimization of Metallic Alloys · unevada

Tried and failed

domain adaptation and generalization algorithms applied to tabular neuroimaging classification across sites. Outcome: worse than baseline. Reason: Specialized domain adaptation methods failed to significantly outperform standard empirical risk minimization.

Fair and Generalizable Machine Learning for Neuroimaging · Penn

Lost to a baseline

Multi-source domain adaptation methods were outperformed by simple merging on the Heintz-Buschart real-data target domain.

Statistical Methods for Outcome Measurement Error Correction, And Multi-study Prediction And Causal Inference under Study Heterogeneity · Harvard

Tried and failed

transfer learning from lower-fidelity pre-trained neural networks applied to neural network force field training. Outcome: worse than baseline. Reason: pre-trained weights provided no faster convergence or accuracy gain over random initialization

Machine-learning models for analysis of biomass reactions and prediction of reaction energies · Georgia Tech

Lost to a baseline

Smaller fully-connected neural network models (e.g., 5-2 and 5-5) performed worse when trained with transfer learning (train RMSE 2.80 mm and 2.49 mm) than when trained only with self-supervised learning (2.20 mm and 1.72 mm) or supervised learning (0.97 mm and 1.48 mm).

Enabling Shape-Based Approaches for Autonomous Percutaneous Interventions with Sensorized Needles · JScholarship

Lost to a baseline

Transfer learning models for platinum(II) NNC emission energy (Mordred + RF R2 = 0.760 ± 0.123, test R2 = 0.18, test MAE = 0.14 eV) showed no significant improvement over the baseline model trained purely on the platinum(II) NNC dataset (Mordred + RF R2 = 0.723 ± 0.241, MAE = 0.0927 ± 0.0279 eV)

Discovery of photoactive anti-cancer platinum(II) complexes: synthesis, screening and machine learning approaches · Imperial

Distribution alignment and unsupervised domain adaptation fail to bridge domain gaps or cause negative transfer

9 theses · 8 institutions

Adversarial and distribution alignment techniques frequently failed to bridge domain differences in tasks like depth completion and histopathology, sometimes even degrading accuracy compared to unadapted baselines. When domain shift was minimal or target data distributions were already balanced, domain adaptation mechanisms provided no meaningful signal and produced negative transfer.

Tried and failed

unsupervised domain adaptation with distribution alignment applied to large-scale image classification. Outcome: worse than baseline. Reason: negative transfer occurred due to poor target label estimation when source data was plentiful

TOWARDS ROBUST VISUAL PERCEPTION SYSTEMS IN REAL-WORLD ENVIRONMENTS · Cornell

Tried and failed

unsupervised domain adaptation with adversarial training applied to intra-position gesture classification. Reason: source and target distributions within identical positions lacked sufficient domain shift for adaptation benefits

Adaptive gesture recognition for human-robot interface using mechanomyography (MMG) · Imperial

Tried and failed

adversarial feature and output alignment applied to LiDAR depth completion domain adaptation. Outcome: no signal. Reason: provided minimal performance impact and failed to bridge the domain gap

Domain adaptation for semantic and 3D tasks · Imperial

Tried and failed

increasing backbone neural network depth applied to unsupervised domain adaptation. Outcome: did not generalise. Reason: deeper networks did not resolve domain discrepancy and slightly reduced feature transferability across domains

Domain Adaptation with a Classifier Trained by Robust Pseudo-Labels · Virginia Tech

Tried and failed

unsupervised domain adaptation applied to tasks with small domain discrepancy. Outcome: worse than baseline. Reason: source and target domains already had minimal discrepancy, rendering adaptation alignment ineffective

Overcoming Distribution Shifts and Imbalance Challenges in Representation Learning for Deep Regression Models · EPFL

Tried and failed

unsupervised domain adaptation for multiple instance learning applied to out-of-distribution histopathology slide classification. Outcome: did not generalise. Reason: feature and pixel adaptation failed to provide statistically significant improvements on distinct external target cohort

Domain Adaptation for Breast Cancer Computational Pathology: Evaluating Feature-Level and Pixel-Level Approaches Under Leave-One-Domain-Out Framework · Harvard

Tried and failed

learning prior class distribution in optimal transport applied to unsupervised domain adaptation. Outcome: no signal. Reason: the target domain had a balanced class distribution matching a uniform prior

A prototype-oriented framework for deep transfer learning applications · UT Austin

Lost to a baseline

Unsupervised domain adaptation alone using CycleGAN without BWE (11.50% EER / 0.532 minDCF on SRE16-YUE-eval40) degraded performance relative to the unadapted/no-BWE baseline (7.46% EER / 0.382 minDCF).

Robust Speaker Recognition using Perceptual and Adversarial Speech Enhancement · JScholarship

Considered and rejected

Considered and rejected: Adversarial domain adaptation was rejected in favor of pseudo-shot learning because adversarial training aligns only domain-level distributions without ensuring correct class-level feature alignment

Coping with the distribution change in soil classification with LIBS · oURspace

Negative transfer occurs across severely mismatched tasks, environments, and biological systems

9 theses · 7 institutions

Severe discrepancies between source and target settings, such as mismatched signal-to-noise ratios, distinct geographic environments, and disparate genomic architectures, led to complete transfer failure. Direct cross-domain transfers from preclinical models or synthetic datasets degraded predictive accuracy because the underlying representations lacked shared inductive structures.

Tried and failed

transfer learning with synthetic pretraining and finetuning applied to subsurface seismic fault segmentation. Outcome: did not generalise. Reason: negative transfer and inductive bias mismatch between distinct synthetic seismic geological models

Multiscale Integration of Cross-Modal Subsurface Data for Reservoir Characterization under Label-Constrained Environments · Georgia Tech

Lost to a baseline

At very low task similarity rho in transfer learning, standard learning with no transfer (delta = 0) beats both hard and soft transfer.

Estimation and Learning via Convex Optimization: Asymptotics, Phase Transitions, and New Algorithms · Harvard

Tried and failed

zero-shot transfer learning with pretrained audio embeddings applied to idiosyncratic atypical vocalization classification. Outcome: no signal. Reason: generic pretrained representations failed to capture idiosyncratic acoustic patterns of atypical vocalizations

Foundations of Cognitive, Affective, and Communicative Systems for Neurodiverse Individuals · MIT

Tried and failed

cross-region transfer learning for acoustic classification applied to audio event detection. Outcome: did not generalise. Reason: models trained on external or online datasets failed to transfer to a new geographic acoustic environment

Listening in on the forest: use of bioacoustics to preserve soundscapes and rare species · Imperial

Lost to a baseline

Transferring from easier (high SNR) to harder (low SNR) synthetic domains was beaten by the limited-data baseline when SNR ranges did not overlap

Foundations of Radio Frequency Transfer Learning · Virginia Tech

Lost to a baseline

Generic zero-shot AudioSet transfer learning achieved only 51.1% accuracy on self-talk, performing barely above chance

Interfaces and Models for Improved Understanding of Real-World Communicative and Affective Nonverbal Vocalizations by Minimally Speaking Individuals · MIT

Tried and failed

direct cross-dataset model transfer without domain adaptation applied to cross-study drug sensitivity prediction. Outcome: did not generalise. Reason: distribution mismatch between source and target datasets degraded predictive performance without domain mapping

Application of advanced machine learning based approaches in cancer precision medicine · Texas Tech

Tried and failed

direct transfer learning from preclinical cell lines applied to clinical patient drug response prediction. Outcome: did not generalise. Reason: distribution shift and biological discrepancies between in vitro cell models and in vivo human tumors

Predicting cancer patient response to chemotherapy using machine learning from small data · Imperial

Tried and failed

transfer learning across related domains applied to polygenic risk score prediction. Outcome: did not generalise. Reason: insufficient shared genetic architecture between auxiliary and target traits

TRANSFER LEARNING IN CLASSIFICATION AND REGRESSION WITH SUMMARY STATISTICS · Penn

Transfer learning is outperformed by simpler heuristics and standard control models in sequential and dynamical systems

8 theses · 6 institutions

In time series, physical tracking, and degradation forecasting, transfer learning architectures were beaten by basic persistent baselines, model-based controllers, and null models. These models failed to extrapolate dynamic trends such as late-life battery aging and performed worse than basic single-cluster or semi-supervised baselines.

Lost to a baseline

Zero-information transfer learning (testing source classifiers directly on target memory data with simple score averaging) performed significantly worse than typical unidimensional memory classification (t(42) = 1.90, p = 0.032)

More than sum of its parts : investigating episodic memory as a multidimensional cognitive process · UT Austin

Tried and failed

transfer learning for long-term degradation prediction applied to battery state of health estimation. Outcome: did not generalise. Reason: model failed to extrapolate capacity loss during late-life aging phases beyond trained cycle horizons

Modeling and Simulation of Power System with High Penetration of Inverter-based Resources · Georgia Tech

Lost to a baseline

Model-free tracking MSE on target SISO system after transfer learning (2.135e-5) was slightly worse than model-based control with an accurate predictor.

Control of Agentic Systems Using the Newton-Raphson Controller · Georgia Tech

Lost to a baseline

Transfer learning on 2 clusters (MAE 0.127) performed worse than the K=1 base model (MAE 0.123) for LSTM on the 4-week cell-A sample.

A Framework for Generalizing Uncertainty in Mobile Network Traffic Prediction · Virginia Tech

Lost to a baseline

Multimodal transfer learning flood prediction achieved 0.783 accuracy (1-year), lost to naive persistent baseline (0.895 accuracy).

Multimodal Machine Learning for Climate Adaptation · MIT

Lost to a baseline

Cold-started and warm-started transfer learning CNN models on M. buryatense copper response performed no better than null baseline models trained on shuffled sequences (F1 ~0.34 vs shuffled F1 ~0.33; Pearson r ~0.26 vs shuffled r ~0.19).

Towards sustainable biomolecule production: computational approaches to accelerate genetic tool development for engineering metabolism in microorganisms · ResearchWorks

Lost to a baseline

Real-time unsupervised domain adaptation performed statistically worse than the semi-supervised baseline on level ground walking R2 (0.64 vs 0.77).

Enabling Scalable, Versatile, and Robust Control for Robotic Exoskeletons · Georgia Tech

Lost to a baseline

Transfer learning of eqt-pnw was outperformed by semblance ensembling in cross-domain picking accuracy on noisy data

Observing Seismic Variations by Earth and Lab Fluids and Fractures · Harvard

Constrained transfer strategies and frozen feature representations degrade downstream performance

6 theses · 6 institutions

Restricting transfer learning to frozen early layers or fixed feature representations resulted in lower prediction quality than full model fine-tuning. Furthermore, transferring linear projections to nonlinear kernels and extracting representations without target label awareness discarded predictive information and failed to mitigate representation bias.

Considered and rejected

Considered and rejected: Rejected standard transfer learning that updates only pre-trained model weights without retaining source dataset access, due to suboptimality under the data processing inequality.

Fair and Generalizable Machine Learning for Neuroimaging · Penn

Tried and failed

transfer learning with frozen feature layers applied to time series classification. Outcome: worse than baseline. Reason: None

Statistical Machine Learning on Time Series with Applications to Manufacturing and Healthcare · Texas Tech

Considered and rejected

Considered and rejected: Transfer learning by freezing/constraining early neural network layers was rejected because it degraded prediction performance relative to full-network fine-tuning.

Building Blocks of Neural Network Intermolecular Interaction Potentials · Georgia Tech

Tried and failed

transferring linear representation projections to nonlinear kernels applied to molecular property prediction. Outcome: did not generalise. Reason: linear optimization does not account for higher-order feature interactions introduced by polynomial kernel powers

A general and efficient framework for atomistic machine learning · EPFL

Tried and failed

weight decay regularization applied to fixed-feature transfer learning. Outcome: did not generalise. Reason: does not mitigate transfer of representation bias

Towards ML Models That We Can Deploy Confidently · MIT

Considered and rejected

Considered and rejected: Rejected transfer component analysis (TCA) for domain adaptation because it extracts representations without utilizing source domain labels, which can discard features with moderate domain mismatch but high predictive power

Towards Robust Machine Learning for Health Applications · Publikationssystem UB Tuebingen

Left open by the authors

Problems the authors named and did not get to.

Left open

Apply formal transfer learning algorithms to measure and mitigate negative transfer between simulated genomic data and empirical target datasets. Blocker: None

The Inference of Selective Sweep Parameters from their Genomic Footprint · Cornell

Left open

Scale the multi-step Vision Transformer transfer learning framework to broader medical image classification and segmentation tasks across modalities. Blocker: None

Detecting Covid-19 Effectively With Transformers And CNN-based Deep Learning Mechanisms · Texas Tech

Left open

Implement transfer learning using pre-trained computer vision models on low-light plant image datasets for leaf segmentation. Blocker: Requires the private low-light and EMCCD luminescence plant imaging dataset from the thesis lab

Automation of Luminescence Quantitation for High-Throughput Plant Phenotyping Using Image Processing and Deep Learning · TXST Digital Repository

Left open

Develop transfer learning and domain adaptation methods to relax transportability assumptions between clinical trial surrogate marker datasets. Blocker: No specific statistical framework or target algorithm is defined beyond general concepts

Heterogeneous surrogate markers in clinical trials and real-world settings · UT Austin

Left open

Develop domain adaptation algorithms combined with spectral data calibration for LIBS soil classification across distribution shifts. Blocker: Access to the thesis's specific calibrated LIBS soil spectral datasets.

Coping with the distribution change in soil classification with LIBS · oURspace

Left open

Develop a theoretical framework analyzing how contrastive learning affects feature distribution alignment in domain adaptation. Blocker: Lacks specific mathematical framework, formulation, or concrete hypotheses to test

Exploring Deep Representation Learning on Vision and Language Intelligence · DukeSpace

Left open

Analyze theoretical and empirical conditions under which simple multi-source data merging outperforms domain adaptation in regression settings. Blocker: None

Statistical Methods for Outcome Measurement Error Correction, And Multi-study Prediction And Causal Inference under Study Heterogeneity · Harvard

Left open

Investigate domain adaptation methods for short-sequence LSTM ensembles applied to significantly different source and target dynamic systems. Blocker: The objective is a broad research direction without specific target datasets, adaptation approaches, or evaluation metrics.

Parameter optimization for enhanced modeling of dynamic systems · Iowa State

Left open

Evaluate pretrained transformer models with subword tokenization and transfer learning on scraped job postings to predict bankruptcy and corporate growth. Blocker: Scraped job listings dataset may not be publicly archived with the thesis

Assessing Corporate Growth and Bankruptcy Risk Using Public Data Proxies · Harvard

Left open

Develop unified time-series foundation models across diverse datasets using transfer learning and domain adaptation techniques. Blocker: None

Graph-based Time-series Forecasting in Deep Learning · Virginia Tech

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.