Chapter Four · failure evidence

What Contrastive & Representation Learning got wrong, from 65 dissertations

The records document multiple failure modes across contrastive and representation learning in computer vision, graph reasoning, audio processing, and biomedical domains. Authors repeatedly observe that complex representation frameworks suffer from sampling bias, collapse during multi-objective optimization, and often fail to outperform simpler baseline representations. These records come from PhD theses at 21 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Pretrained and self-supervised representations underperform simpler or supervised baselines on specialized tasks

11 theses · 6 institutions

Self-supervised contrastive models and learned embeddings repeatedly failed to match traditional supervised models, tree-based baselines, or raw matrix inputs across clinical, genomic, and acoustic benchmarks. Generic pretrained representations struggled to capture fine-grained domain characteristics like pathological vocal deviations, biological graph structures, and chemical reaction properties.

Lost to a baseline

Self-supervised pre-training methods (Contrastive and MLM) achieved lower average ranks than supervised pre-training across all downstream data regimes on MetaMIMIC

Applications of Optimization and Machine Learning to Healthcare · ResearchWorks

Lost to a baseline

Self-supervised contrastive localization model had 272% larger error at 40° elevation vs 0°, underperforming the supervised baseline model

Modeling and Evaluating Human Sound Localization in the Natural Environment · MIT

Lost to a baseline

End-to-end Neural Network with GNN embeddings underperformed Random Forest and XGBoost baselines across all metrics (Accuracy 0.5836, F1 0.4873, AUC-ROC 0.5405, AUC-PR 0.4672)

Predicting psychological treatment dropout using graph neural networks · OpenBU

Lost to a baseline

SVM and Logistic Regression lost to simple adjacency matrix representations when using node2vec embeddings on BioGRID-physical.

Benchmarking Methods For Predicting Phenotype Gene Associations · Virginia Tech

Lost to a baseline

Transformer language model reaction embeddings failed to outperform Random Forest, Extreme Gradient Boosting, or Feedforward Neural Networks using standard fingerprint or DFT features.

Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT

Tried and failed

dense supervised contrastive learning representations applied to affinity prediction. Outcome: worse than baseline. Reason: linear probing or fine-tuning failed to beat standard cross-entropy baseline

Label-Efficient Deep Learning Methods for Event Detection, Object Segmentation and Cell Tracking in Biological Microscopy Images · EPFL

Tried and failed

Word Mover's Distance on learned sequence embeddings applied to cross-language code clone detection. Outcome: worse than baseline. Reason: significantly degraded detection effectiveness compared to Siamese bidirectional LSTM networks

Helping Developers Migrate their Code across Programming Languages · Virginia Tech

Tried and failed

temporal contrastive loss in time series representation applied to neural spike train representations. Outcome: worse than baseline. Reason: None

Building a foundation model for neuroscience · Georgia Tech

Lost to a baseline

CGLVM achieved a slightly higher Mutant Cell Line Silhouette (0.980±0.007) than contrastiveVI (0.966±0.001) on the salient latent representation of MIX-seq data.

Structured Deep Generative Models for Exploring the Single Cell Landscape · ResearchWorks

Tried and failed

zero-shot transfer learning with pretrained audio embeddings applied to idiosyncratic atypical vocalization classification. Outcome: no signal. Reason: generic pretrained representations failed to capture idiosyncratic acoustic patterns of atypical vocalizations

Foundations of Cognitive, Affective, and Communicative Systems for Neurodiverse Individuals · MIT

Tried and failed

self-supervised speech embeddings with dynamic time warping applied to pathological speech intelligibility and impairment classification. Outcome: no signal. Reason: pre-trained representations failed to capture non-linguistic acoustic deviations and breathing impairments, yielding zero sensitivity and negative correlation

Novel Methods For Detection And Analysis Of Atypical Aspects In Speech · EPFL

Contrastive loss objectives suffer from inappropriate weighting, symmetry, and similarity metrics

10 theses · 5 institutions

Contrastive formulations degraded model performance when loss weights were set to extremes, when symmetric objectives were forced onto hierarchical data, or when continuous distributions collapsed under conventional losses. Replacing Pearson correlation with cosine similarity, averaging positive pairs inside logarithms, or applying uniform attraction without conditional matching also led to performance drops.

Tried and failed

extreme weighting in multi-task contrastive learning applied to medical image classification. Outcome: worse than baseline. Reason: setting the contrastive loss weight either too low or too high degraded target task performance

Integrating domain knowledge and deep learning for enhanced chest X-ray diagnosis and localization · UT Austin

Tried and failed

non-negative representation learning without contrastive structured loss applied to knowledge graph completion. Outcome: worse than baseline. Reason: cannot sufficiently distinguish similar yet distinct entities without negative samples or structured relation constraints

Integrating structural and semantic understanding for robust knowledge graph construction: From knowledge graph completion to zero-shot entity linking · Iowa State

Tried and failed

contrastive slow feature learning with scalar modulation applied to deep representation learning. Reason: did not yield significant improvements over standard representation learning baselines

Biologically plausible unsupervised learning in shallow and deep neural networks · EPFL

Tried and failed

symmetric contrastive loss in self-supervised learning applied to visual perceptual similarity feature extraction. Outcome: worse than baseline. Reason: None

Advancements in perceptual quality assessment for interactive media : from mobile cloud gaming to human avatar videos and facial expressions · UT Austin

Considered and rejected

Considered and rejected: Rejected cosine similarity in the contrastive loss function in favor of exponentiated Pearson correlation similarity due to better validation r2.

Improving Microestimates of Poverty from Satellite Images · Harvard

Considered and rejected

Considered and rejected: Rejected multiplicative normalization in Supervised Contrastive Loss (SCL), substituting an additive dis-alignment factor in SSCL to accommodate continuous similarity metrics.

Enhancing next generation networks with security, sensing and management · UT Austin

Considered and rejected

Considered and rejected: Decided against symmetric bidirectional contrastive loss in favor of an asymmetric text-to-image contrastive loss reflecting document-region hierarchy.

Language-Centric Medical Image Understanding · MIT

Considered and rejected

Considered and rejected: Rejected averaging positive pairs inside the logarithm of the contrastive loss function due to generic performance degradation.

Learning Simplicity and Structure in Quantum Field Theory · Harvard

Tried and failed

Uniform attraction and repulsion contrastive loss applied to representation learning for classification. Outcome: worse than baseline. Reason: uniformly pulling positives and pushing negatives without conditional distribution matching degraded classification accuracy

Implicit distributional matching at high dimensionality · UT Austin

Tried and failed

conventional contrastive losses applied to continuous dynamical parameter distributions. Reason: caused dimensional collapse or severe performance degradation on continuous distributions

On the Certification of Deep Learning-based Dynamical System Identification · MIT

Global pooling and poorly aligned view augmentations discard task-critical information

10 theses · 7 institutions

Generating contrastive views with low mutual information, overlapping spatial crops, or unsupervised attention masks led to representation collapse and lost task-relevant signals. Furthermore, relying on global vision embeddings or CLS tokens diluted fine-grained local context compared to localized features and token mean pooling.

Tried and failed

asymmetric multi-resolution contrastive learning with attention masking applied to unsupervised image representation learning. Outcome: worse than baseline. Reason: unsupervised attention masking degraded representation quality without supervision

Attention based contrastive learning with non-Siamese architecture · Iowa State

Tried and failed

contrastive learning across low mutual information views applied to visual representation learning. Outcome: did not generalise. Reason: insufficient shared information between views removes task-relevant signals needed for downstream transfer

Towards General-purpose Vision via Multiview Contrastive Learning · MIT

Tried and failed

supervised contrastive loss on overlapping spatial patches applied to continual representation learning from image streams. Outcome: worse than baseline. Reason: treating overlapping spatial crops as positive class pairs harmed representation quality compared to batch-splitting them

Unsupervised Progressive Learning for Agents in Online and Dynamic Environments · Georgia Tech

Tried and failed

conditioning policy on global vision embeddings directly applied to robot policy cost correction. Outcome: worse than baseline. Reason: lacks fine-grained spatial representations needed for object-referential tasks

Discovering and Engineering the Computation Underlying Large Intelligent Agents · MIT

Tried and failed

CLS token embedding for sequence representation applied to molecular representation learning. Outcome: worse than baseline. Reason: Mean pooling across all token embeddings captured contextual representation significantly better than CLS token.

Contextual representations of the chemical space for task agnostic machine learning methods · Imperial

Tried and failed

concatenating transformer cls embedding with cnn features applied to clinical relation extraction. Outcome: worse than baseline. Reason: sentence-level representations diluted or conflicted with localized convolutional feature representations

Learning to Improve Clinical Decisions and AI Safety by Leveraging Structure · MIT

Considered and rejected

Considered and rejected: Rejected Barlow Twins contrastive learning framework for trajectory representations as its performance degraded (0.737 F1 vs 0.881 F1 for SimCLR) due to representations collapsing on simple augmentations.

Geospatial representation learning: from road networks to trajectories · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected using the [CLS] token for sentence representation in BiGS downstream classification; used mean pooling over non-padding token embeddings because it yielded superior performance.

Designing for Inference in Future Generative Models · Cornell

Considered and rejected

Considered and rejected: Concatenating speaker embedding vectors directly to input acoustic features for anchored ASR; rejected due to lack of discrimination capacity.

WAKE WORD DETECTION AND ITS APPLICATIONS · JScholarship

Tried and failed

standard contrastive learning for sequential state representations applied to subgoal discovery in reinforcement learning. Outcome: did not generalise. Reason: produces ambiguous subgoal ordering and suffers distraction from neighboring non-key states without temporal geometric sampling

Reinforcement learning with temporal logic and causal constraints · Georgia Tech

Joint training objectives and improper multimodal fusion cause representation collapse and interference

9 theses · 8 institutions

Pre-fusion contrastive learning and spatial feature cross-talk diluted raw features, while joint training with masked token prediction or dual objectives triggered convergence failures. Omitting predictor heads in non-contrastive setups or freezing unimodal encoders similarly caused representation collapse or limited multimodal alignment.

Tried and failed

pre-fusion contrastive learning on raw multimodal representations applied to separable spatial-temporal multimodal fusion. Outcome: worse than baseline. Reason: un-fused raw feature contrastive objectives degraded downstream performance compared to post-fusion feature learning

Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech

Considered and rejected

Considered and rejected: Rejected training embeddings using Word-Level (WL) image-to-word regression without context because it failed to preserve textual information and degraded word similarity

Language Grounding in Vision · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Rejected end-to-end joint multimodal pretraining, opting instead to disentangle unsupervised contrastive alignment from downstream supervised adaptation.

Multimodal Representation Learning for Mental Health: Transfer Learning, PEFT, and Contrastive Learning · DalSpace

Considered and rejected

Considered and rejected: Adding a 2-layer projection/predictor network on embeddings in EmbeddingEarth was rejected as it did not improve performance.

Deep learning methods for satellite image timeseries · Imperial

Considered and rejected

Considered and rejected: Rejected standard contrastive/instance-level InfoNCE loss between multimodal concepts in SimZSS because identical object concepts co-occur multiple times across a batch; used semantic classification cross-entropy instead.

Exploiting Representation Similarities in Self-Supervised Learning for Vision Tasks · EPFL

Tried and failed

non-contrastive self-supervised learning without predictor network applied to cross-modal representation learning. Reason: removing negative pairs without a predictor head caused complete representation collapse

Learning from talking faces: from deepfake detection to speech recognition · Imperial

Tried and failed

freezing pretrained text encoder in contrastive learning applied to multimodal contrastive representation learning. Outcome: worse than baseline. Reason: limits joint alignment between modalities during contrastive representation learning

Visual Representation Learning from Synthetic Data · MIT

Considered and rejected

Considered and rejected: Rejected contrastive learning on audio features derived from spatial fusion (Cross Contr) because fusing each audio token with 64 visual tokens dilutes raw audio features.

Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech

Considered and rejected

Considered and rejected: Rejected joint pre-training of masked audio token prediction and contrastive loss simultaneously because dual objectives caused convergence failures.

Towards Integrated Audio-Visual Learning: From Vision-to-Audio Generation to a Unified Audio-Visual Framework · ResearchWorks

Lost to a baseline

On CommonsenseQA IHdev set, ConKGP with perturbations + contrastive learning (76.74%) was beaten by contrastive learning alone (77.64%)

Knowledge-Enhanced Language Models: Integrating Structured Knowledge Graphs in Large Language Model for Multiple-Choice Question-Answering Tasks · Carleton University Institutional Repository

Faulty negative sampling strategies and class imbalance introduce noise and computational bottlenecks

10 theses · 7 institutions

Using unaligned background negatives, self-comparisons, or in-batch random negatives introduced severe label bias and failed to teach fine distinctions. Severe class imbalances and periodic physiological waveforms produced corrupted false negatives, while excessive batch size requirements created intractable computational costs.

Tried and failed

supervised contrastive learning with positive-unlabeled data applied to positive-unlabeled classification with high class imbalance. Reason: gradient bias and false negative assumptions in low supervision and high class-prior regimes caused severe performance collapse

Robust and efficient learning in high dimensions from noisy data · UT Austin

Tried and failed

contrastive learning using self-comparisons as negative pairs applied to genomic coverage profile representations. Outcome: worse than baseline. Reason: comparing identical coverage vectors degenerated representation quality and increased cross-chromosome variance versus using biological replicates

Unsupervised learning with high-throughput sequencing data · Iowa State

Considered and rejected

Considered and rejected: Using trivial negative samples with random backgrounds unaligned to the anchor's class in contrastive learning was rejected as it degraded debiasing performance compared to matched dictionary queues.

Towards Robust Vision Models · EPFL

Considered and rejected

Considered and rejected: Rejected relying purely on in-batch random negatives for contrastive training because they only teach coarse distinctions.

On Translation as a Symmetry of Language · Harvard

Tried and failed

joint detection and contrastive identification training applied to sparse audio event speaker separation. Outcome: worse than baseline. Reason: dummy negative labels from severe background class imbalance corrupted contrastive loss learning

Decoding the Depths: Developing a Click Separator for Predictive Speaker Recognition in Sperm Whale Conversations Using Machine Learning · MIT

Tried and failed

false negative filtering in contrastive learning applied to multimodal vision-language representation learning. Outcome: did not generalise. Reason: filtering false negatives improved dominant class classification but yielded negligible benefit for grounding or retrieval

Language-Centric Medical Image Understanding · MIT

Tried and failed

contrastive mutual information estimation with negative sampling applied to representation learning. Outcome: unstable. Reason: very small batch sizes fail to generate sufficient negative pairs via latent variable permutation

Novel approaches for learning representations · UT Austin

Considered and rejected

Considered and rejected: Rejected standard negative-pair contrastive time-series frameworks (SimCLR, MoCo, SwAV) due to erroneous negative sampling on periodic physiological waveforms.

Toward accurate health monitoring through large-scale Photoplethysmography signal from wearable devices · Georgia Tech

Tried and failed

exact inclusion-exclusion decomposition for negative sampling applied to contrastive representation learning. Outcome: infeasible cost. Reason: requires at least N positive samples and becomes computationally intractable for large batch sizes

Robust Learning from Uncurated Data · MIT

Considered and rejected

Considered and rejected: Rejected self-supervised contrastive learning due to excessive batch size and compute requirements

Efficient and Robust Machine Learning Methods for Challenging Traffic Video Sensing Applications · ResearchWorks

Complex graph neural network embeddings fail to improve over simple adjacency structures

9 theses · 7 institutions

Incorporating graph embeddings or joint GNN training into reinforcement learning and classification yielded no significant advantage over simple adjacency matrices or one-hot vectors. Graph structure embeddings also injected noise into sparse graphs, required external supervision labels, or distorted structural explanations by over-prioritizing node features.

Tried and failed

joint training of GNN embeddings with RL applied to reinforcement learning policy and value networks. Outcome: worse than baseline. Reason: yielded no significant performance gain compared to simple one-hot encoding

Learning Simplicity and Structure in Quantum Field Theory · Harvard

Tried and failed

random walk graph representation learning applied to biological network node classification. Outcome: worse than baseline. Reason: Low-dimensional embeddings failed to outperform simple adjacency matrix representations.

Benchmarking Methods For Predicting Phenotype Gene Associations · Virginia Tech

Tried and failed

contrastive learning and early-fusion graph convolutional networks applied to heterogeneous dense and sparse biological graphs. Outcome: worse than baseline. Reason: None

Towards Network-Guided Large-Scale Foundation Models on Single-Cell Transcriptomics · Virginia Tech

Tried and failed

incorporating initial node and relevance score embeddings applied to graph neural network question answering reasoning. Outcome: no signal. Reason: the added embeddings were completely dispensable and provided no performance benefit

Knowledge Reasoning with Graph Neural Networks · Georgia Tech

Tried and failed

combining language models with graph structure embeddings applied to link prediction on sparse graphs. Outcome: worse than baseline. Reason: graph structure information introduced noise on sparse graphs, degrading prediction performance

Temporal link prediction in the wild · Imperial

Considered and rejected

Considered and rejected: Rejected standard cluster/community-based multi-aspect graph embeddings because they fail to capture diverse higher-order relational semantics in heterogeneous networks.

Automatic Question Answering and Knowledge Discovery from Electronic Health Records · Virginia Tech

Considered and rejected

Considered and rejected: Rejected GNN graph embeddings and spectral/diffusion methods because GNNs require supervised labels and spectral methods cannot handle node attributes.

Optimal transport methods for analyzing highly multiplexed spatial proteomic data · Oxford

Considered and rejected

Considered and rejected: Rejected DiffPool for subgraph embeddings because subgraph selection is learned end-to-end and cannot generate embeddings for arbitrary subgraphs without a separate learning process

Graph Embedding for Retrieval · EPFL

Considered and rejected

Considered and rejected: Rejected aligning explanation embeddings with dataset average embeddings (as in GNNInterpreter) because it prioritizes node features over structural details and creates disconnected explanation graphs.

Motif-based Graph Neural Networks: Representation, interpretation, and extraction · Iowa State

Left open by the authors

Problems the authors named and did not get to.

Left open

Benchmark contrastive learning directly against multitask learning for integrating radiomics features with deep learning on chest X-rays. Blocker: None

Integrating domain knowledge and deep learning for enhanced chest X-ray diagnosis and localization · UT Austin

Left open

Integrate contrastive learning, reinforcement learning from user feedback, and GNNs into the DPK-GLM biomedical retrieval framework. Blocker: Vague high-level directions without specific architecture, loss formulation, or feedback dataset specified

Enhancing General Language Models for Biomedical Test Retrieval via Diversified Prior Knowledge · YorkSpace

Left open

Develop domain-specific data augmentations for self-supervised contrastive learning on healthcare time series and modalities. Blocker: The direction is very broad without specific datasets, target modalities, or concrete augmentation formulations specified

Towards generalization of deep learning in pervasive human motion analysis · Imperial

Left open

Develop a theoretical framework analyzing how contrastive learning affects feature distribution alignment in domain adaptation. Blocker: Lacks specific mathematical framework, formulation, or concrete hypotheses to test

Exploring Deep Representation Learning on Vision and Language Intelligence · DukeSpace

Left open

Develop graph neural network embeddings or structure-agnostic representations for chemical structures in Bayesian optimization. Blocker: Lacks specific target problem, chemical dataset, and concrete baseline metrics

Bayesian optimisation in chemical problems · Imperial

Left open

Train a graph convolutional network to learn AST representations and token embeddings directly for performance modeling without handcrafted features. Blocker: None

Exploring synergies between program synthesis and machine learning · UT Austin

Left open

Implement a Graph Neural Network to replace the embedding prediction network in KD-EMD for spatial and relational knowledge distillation. Blocker: None

From Symbolic Reasoning to Object Embeddings: Advanced Approaches of Knowledge Distillation in Compacted Neural Networks · Texas Tech

Left open

Formulate contrastive representation learning as a multi-label classification objective using binary cross-entropy loss. Blocker: None

Understanding and Improving Representational Robustness of Machine Learning Models · MIT

Left open

Add a contrastive loss term alongside unlikelihood loss during Possibility Exploration Fine-Tuning to refine pairwise response similarity. Blocker: None

Towards Human-like Dialogue Systems: Enhancing Sensibility, Diversity, Proactiveness, and Naturalness in Spoken Open-domain Chatbots · Research Repository UCD

Left open

Design and evaluate alternative training objectives beyond pairwise text-text contrastive loss to improve temporal reasoning in cross-modal audio-text retrieval. Blocker: None

Searching audiovisual media with natural language queries · Oxford

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.