Chapter Four · failure evidence

What Large Language Models & Prompt Engineering got wrong, from 50 dissertations

Across diverse tasks, large language models and prompt engineering techniques frequently underperformed simpler supervised classifiers, struggled with domain-specific extraction, and failed on complex reasoning or code generation. Iterative prompt modifications often yielded negligible gains over baseline prompts, while fine-tuning and advanced prompting structures encountered high implementation overhead, capacity bottlenecks, and behavioral degradation. These records come from PhD theses at 17 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Zero-shot prompting fails on specialized domain tasks and complex extraction

11 theses · 6 institutions

Zero-shot prompts achieved near-random accuracy or low recall when applied to domain-specific extraction, legal statutory interpretation, and nuanced subjective annotation. Models lacked the necessary domain context, structured reasoning, or fine-tuning to outperform simple baselines like rule-based sentiment analyzers.

Tried and failed

direct zero-shot prompting of large language models applied to clinical note information extraction. Outcome: did not generalise. Reason: failed to adapt and performed poorly on advanced reasoning models without structured reasoning or fine-tuning

Machine Learning Approaches for Drug Combination Discovery · Cornell

Tried and failed

prompt-based zero-shot large language model classification applied to student text behavior classification. Outcome: worse than baseline. Reason: produced substantially lower AUC and more than double the false positives compared to embedding-based classifiers

MEASURING AND UNDERSTANDING STUDENTS’ SELF-REGULATED LEARNING IN TEXTUAL DATA IN COMPUTER-BASED LEARNING ENVIRONMENTS · Penn

Tried and failed

zero-shot large language model prompting applied to commonsense logic rule generation. Outcome: did not generalise. Reason: models fail to consistently output valid compositional rules without task-specific fine-tuning

Neurosymbolic Programming in Scallop: Design, Implementation, and Applications · Penn

Tried and failed

zero-shot prompting with basic statutory definitions applied to confidential information redaction in text. Outcome: did not generalise. Reason: Abstract legal definitions lacked sufficient context to identify complex, non-standard sensitive categories accurately.

LLM-Assisted Detecting and Redacting Confidential Information for Government Information Disclosure · Virginia Tech

Tried and failed

large language model zero-shot text classification applied to moral framing text annotation. Outcome: no signal. Reason: LLM predictions failed to align with validated domain-specific dictionary benchmarks across categories

COMPUTATIONAL METHODS ON POLITICAL MORAL LANGUAGE IN ONLINE SPACES · Cornell

Tried and failed

zero-shot large language model generation without exemplars applied to generating software vulnerability exploits. Outcome: no signal. Reason: the model lacked domain-specific context and test patterns needed to synthesize functional exploit payloads

Secure Coding Practice in Java: Automatic Detection, Repair, and Vulnerability Demonstration · Virginia Tech

Tried and failed

zero-shot large language model reasoning applied to temporal knowledge graph link prediction. Outcome: did not generalise. Reason: models fail to understand structured graph entities and temporal relations without specialized fine-tuning

Temporal link prediction in the wild · Imperial

Tried and failed

zero-shot prompting of general-purpose LLMs applied to extracting viral mutations from text. Outcome: no signal. Reason: general-purpose models achieved extremely low recall and F1 scores on domain-specific extraction

Discovering Viral Hosts, Mutations, and Diseases using Machine Learning · Virginia Tech

Tried and failed

zero-shot large language model prompting applied to nuanced text annotation tasks. Outcome: no signal. Reason: models achieved near-random performance on subjective constructs

More Than a Sum of Its Parts: Teamwork Through a Multidimensional Lens · Penn

Tried and failed

direct LLM zero-shot prompting without retrieval applied to biological perturbation effect prediction. Outcome: no signal. Reason: models lacked external experimental context and structured reasoning to predict perturbation outcomes above chance

Practical Algorithms for Modeling Causality to Accelerate Scientific Discovery · MIT

Lost to a baseline

Zero-shot prompting on several LLMs (including Llama3 8B and Mistral 7B) achieved lower accuracy than the rule-based VADER sentiment analysis baseline

Leveraging social media data and large language models for understanding public health behaviours in the context of infectious diseases and vaccination · EPFL

Prompted generative models lose to supervised baselines and specialized encoders

6 theses · 5 institutions

Few-shot and in-context prompts underperformed specialized encoder architectures like BERT and simple supervised classifiers on domain text classification and sequence extraction. In several instances, prompted generative models failed to beat naive majority-class baselines or simpler language modeling alternatives.

Tried and failed

few-shot prompting of pretrained language models applied to domain-specific text classification. Outcome: worse than baseline. Reason: general pretraining lacks alignment with specialized, nuanced coding rubrics compared to supervised simple classifiers

Evaluating language models applied to student thinking about experiments · Cornell

Tried and failed

prompting pretrained generative large language models applied to complex domain-specific text classification. Outcome: worse than baseline. Reason: zero-shot and few-shot prompting failed to capture complex domain constructs compared to fine-tuned domain-specific encoders

Modeling Legal Constructs · Cornell

Lost to a baseline

Most LLM prompts for stance extraction across all 1,000 tweets had lower overall accuracy than a naive majority-class baseline

Leveraging social media data and large language models for understanding public health behaviours in the context of infectious diseases and vaccination · EPFL

Lost to a baseline

FIST with only prompt input yielded higher perplexity (30.2 word PPL / 25.9 BPE) and lower ROUGE-1 F1 (0.181) than PSA (31.6 word / 21.3 BPE / 0.265 F1) on WritingPrompts.

Towards Effective and Controllable Neural Text Generation · DSpace at SUNY Buffalo

Considered and rejected

Considered and rejected: Rejected relying solely on in-context prompting (zero-shot, few-shot, and Chain-of-Thought) for specialized legal construct classification due to poor performance and severe class imbalance errors.

Modeling Legal Constructs · Cornell

Tried and failed

fine-tuned large language models for sequence extraction applied to named entity recognition. Outcome: worse than baseline. Reason: Generative LLMs underperformed specialized smaller encoder models like BERT on token-level span classification tasks.

Consistency-aware and LLM-assisted methods for named entity recognition · Iowa State

Considered and rejected

Considered and rejected: Rejected Large Language Models (LLMs) in favor of LDA topic models and dictionary-based content analysis.

The Rise of Counterpopulism: a Comparative Analysis of Social Movements Resisting Right-Wing Populism in Power in Italy, the United Kingdom and the United States of America · IRIS - SNS - prod

Prompt engineering and instruction refinement yield negligible performance gains

7 theses · 5 institutions

Manual prompt engineering, perspective augmentations, and formatting prompts with human annotation guidelines failed to provide meaningful improvements over simple baseline prompts. Re-prompting failed to correct persistent model biases, and complex prompt instructions increased overfitting risks without reliably controlling output complexity.

Tried and failed

manual iterative prompt engineering applied to large language model text annotation. Outcome: no signal. Reason: manual refinements yielded only marginal performance gains over baseline prompts

Human-Centered AI in Computational Social Science: Evaluating Automated Annotation with Large Language Models · Penn

Tried and failed

using human annotation guidelines as LLM prompts applied to text classification and rationale extraction. Outcome: worse than baseline. Reason: instructions optimized for human annotators did not translate effectively to LLM reasoning and yielded no improvement

Towards Bridging and Governing Decentralized Communities · MIT

Tried and failed

First-person perspective text augmentation applied to Clinical text classification. Outcome: no signal. Reason: Perspective shift provided no meaningful performance gain over baseline prompts

Transforming SDOH Screening: Towards a General Framework for Transformer-based Prediction of Social Determinants of Health · Virginia Tech

Tried and failed

prompt engineering on large language models applied to zero-shot tone classification of transcripts. Reason: re-prompting failed to correct the model's conservative classification bias, causing poor agreement with human annotators

Investigating the Structural Evolution of Ad Design: Computational Analysis of Information Content and Interactivity in Ads · MIT

Considered and rejected

Considered and rejected: Rejected complex prompt instructions for closed-book QA, sticking to 'Q: <question> A:' because sophisticated prompts did not improve accuracy enough to justify model overfitting risks.

Beyond Scaling: Frontiers of Retrieval-Augmented LMs · ResearchWorks

Considered and rejected

Considered and rejected: Rejected relying solely on grade-level prompt engineering to vary summary complexity because lower grade prompts often generated complex language.

Language as Design: Adapting Language to Different Online Audiences · ResearchWorks

Considered and rejected

Considered and rejected: Rejected using zero-shot LLM prompts with detailed definitions and implicit instructions, choosing concise prompts without added explanations.

Machine Learning Approaches to Understanding Legislative Economic Discourse · Research Repository UCD

Code repair and generation struggle with complex logic and architectural context

4 theses · 3 institutions

Prompting models for code conflict resolution and algorithmic problem solving resulted in meaningless outputs, rule overgeneralization, and failure on multi-step state transitions. Practitioners rejected language models for critical code development and multi-file refactoring because debugging machine-generated logic errors was inefficient and models lacked broader architectural context.

Tried and failed

large language models for code repair applied to software build conflict resolution. Reason: generated meaningless outputs and achieved low resolution accuracy on complex merge conflict tasks

Automating The Detection and Resolution of Build Conflicts in Software Merge for Java Programs · Virginia Tech

Tried and failed

prompting LLMs with explicit domain-specific rules applied to automated code conflict resolution. Outcome: worse than baseline. Reason: models overgeneralized explicit rules, incorrectly misapplying patterns to unrelated code contexts

Automating The Detection and Resolution of Build Conflicts in Software Merge for Java Programs · Virginia Tech

Tried and failed

prompting large language models for algorithmic generation applied to complex dynamic programming problems. Outcome: did not generalise. Reason: models fail on complex algorithmic reasoning requiring multi-step state transitions and optimal substructure

Are Bigger LLMs Always Better? A Study of Open and Closed-Source Models in Code Generation and Translation · Virginia Tech

Considered and rejected

Considered and rejected: Rejected using Large Language Models (ChatGPT) for writing critical code because debugging machine-generated logic errors takes more time than writing manually and models cannot answer contextual 'why' questions.

Code, communication, and culture : an ethnography of open source quantum software communities · UT Austin

Considered and rejected

Considered and rejected: Rejected open-ended optimization prompts (e.g., 'make my code faster') and multi-file refactoring with ChatGPT due to lack of extensive non-code and architectural context

Studying Developers’ Engagement with ChatGPT for Issue Resolution · HARVEST

Supervised fine-tuning causes behavioral degradation and imposes high training overhead

5 theses · 5 institutions

Supervised fine-tuning collapsed citation grounding, degraded output format compliance, and lacked the self-revision mechanisms available in prompting. Researchers also rejected fine-tuning for prototyping and specialized detection tasks due to extensive annotation requirements, behavioral instability, and prohibitive computational costs.

Tried and failed

supervised fine-tuning for structured reasoning applied to large language models for clinical grounding. Outcome: worse than baseline. Reason: optimizing for diagnostic correctness collapsed citation grounding and degraded structured output format compliance

Causal Inference and Evidence-Grounded Language Models for Trustworthy Personalized Clinical Decision Support · Georgia Tech

Tried and failed

uninformed magnitude pruning with weight regeneration applied to large language models. Outcome: worse than baseline. Reason: causes irreparable knowledge damage on complex tasks that weight regeneration fine-tuning cannot recover

Chasing efficiency in the era of large language models · UT Austin

Considered and rejected

Considered and rejected: Fine-tuning large language models rejected for prototyping due to extensive annotation/training requirements and difficulty maintaining consistent behavioral parameters.

Balancing Performance and Social Considerations for Autonomous Agents Interacting with Humans · Cornell

Considered and rejected

Considered and rejected: Rejected fine-tuning large language models (GPT-4/Gemini) for affective state detection due to prohibitive computational costs and limited annotated data, opting for instruction-based zero-shot prompting.

Advancing Ransomware Detection with Machine Learning: Human Factors, Evolution, and Detection · Research Repository UCD

Tried and failed

Supervised fine-tuning on positive reasoning paths only applied to Mathematical question answering. Outcome: worse than baseline. Reason: Lacked error-correction and self-revision mechanisms present in prompting

Knowledge-Centric Multimodal Intelligence: Understanding, Reasoning, and Generation · Virginia Tech

Mathematical and multi-step reasoning degrades under direct or uninformative prompting

4 theses · 4 institutions

Models struggled with explicit symbolic equations and multi-step mathematical problems, leading to degraded accuracy and higher error rates. Direct prompting for intermediate answers reduced answer consistency compared to resampling techniques, while empty noise exclusion hints worsened downstream reasoning performance.

Tried and failed

prompting language models with explicit symbolic equations applied to synthetic tabular data generation. Outcome: worse than baseline. Reason: LLMs struggle with complex mathematical reasoning, degrading generation accuracy and increasing MSE

New Approaches to Synthetic Tabular Data Generation · Virginia Tech

Lost to a baseline

UnifiedQA zero-shot answering on StrategyQA scored 58.95% accuracy, heavily underperforming GPT-3 zero-shot prompting (58.08% / 63.32% few-shot) and other reasoning baselines when not given gold supporting facts.

Incidental Supervision for Natural Language Understanding · Penn

Tried and failed

direct prompting for intermediate answers applied to multi-step mathematical reasoning. Outcome: worse than baseline. Reason: yields lower answer consistency across reasoning benchmarks than forward-reverse resampling

In the Blink of an Eye: A Unified Theory for Feature Emergence in Generative Models · Harvard

Tried and failed

prompting with an empty noise exclusion hint applied to mathematical reasoning under noisy input. Outcome: worse than baseline. Reason: providing an empty exclusion hint degraded downstream problem-solving performance compared to explicit noise identification

Quantitative reasoning ability of large language models under noisy data · Iowa State

Complex prompting structures overwhelm the capacity of smaller language models

2 theses · 2 institutions

Applying complex reasoning prompts such as tree structures or chain-of-thought to lower-capacity models led to poor performance. The representational overhead caused model confusion, high labeling variability, and frequent classification errors.

Tried and failed

tree-structured prompting applied to small language model reasoning. Outcome: worse than baseline. Reason: representational overhead overwhelmed the capacity of smaller models

Backward Reasoning in LLMs: A Strategy for Identifying Irrelevant Context in Mathematical Word Problems · Virginia Tech

Tried and failed

chain-of-thought prompting with lower-capacity LLM applied to complex multi-class text annotation. Outcome: unstable. Reason: Model exhibited confusion, high labeling variability, and frequent classification errors compared to larger models

Value Expressions in Patents: Their Relationship to Patent Valuation and Technological Orientation · Georgia Tech

Automated evaluation prompts generate generic feedback rather than pinpointing errors

2 theses · 2 institutions

When evaluating student solutions or writing, models produced vague, generic, and repetitive critiques rather than identifying specific error locations. These automated assessments failed to provide the precise diagnostic feedback offered by human instructors.

Tried and failed

large language models for automated grading applied to student engineering derivations and solutions. Reason: Outputted full solutions and generic suggestions rather than pinpointing specific error locations in student work

Developing an Automated Practice Environment and Feedback Engine for Guided Instruction in the Deformable Bodies Course · Virginia Tech

Tried and failed

prompting large language models for critique generation applied to evaluating complex student writing. Outcome: worse than baseline. Reason: generated critiques were vague, generic, and repetitive compared to human instructor assessments

Addressing teachers’ needs for integrating generative AI into a pedagogical feedback model · Iowa State

Left open by the authors

Problems the authors named and did not get to.

Left open

Evaluate zero-shot, one-shot, and few-shot prompting on LLM annotation performance for pedagogical discourse moves in classroom transcripts. Blocker: Access to the annotated elementary classroom transcript dataset.

Comparing human and generative AI annotation performance: Understanding linguistic features of teacher talk in classroom discussions · Iowa State

Left open

Apply the PARSE framework to evaluate LLM perturbation bias on prompts in healthcare, legal guidance, and job matching domains. Blocker: None

PARSE: A Framework for Evaluating Linguistic and Typographical Perturbation Bias in Large Language Models · YorkSpace

Left open

Benchmark general-purpose large language models against specialized predictive models on reaction condition tasks such as solvent prediction. Blocker: None

Integrating AI across the Chemistry Discovery Cycle: Advancing Sustainable Chemistry through Digital Methods · EPFL

Left open

Compare empirical neural scaling statistics in large language models against predictions from prior analytical scaling models. Blocker: None

Decomposing Deep Neural Network Minds into Parts · MIT

Left open

Implement and evaluate early-fusion n-gram LM prediction prompting or conditioning for neural LLM text generation. Blocker: None

Navigating the Ocean of Language Model Training Data · ResearchWorks

Left open

Adapt summarization evaluation prompting methods to handle documents exceeding context limits and structured tabular inputs. Blocker: None

Fine-grained evaluation for text summarization · UT Austin

Left open

Apply advanced large language models and hybrid quantum-classical algorithms to source code representations for software defect prediction. Blocker: The unfinished work lacks concrete specifications, specific model architectures, or target benchmarks

Enhancing Software Defect Prediction: Investigating Diverse Representations of Source Code as Feature Values in Classical and Quantum Machine Learning Approaches · HARVEST

Left open

Apply lifelong learning fine-tuning techniques to large language models for generating automated code review comments across evolving codebases. Blocker: None

Towards Sustainable AI for Continuous Integration Quality Gates · Queens University Institutional Repository

Left open

Apply large language models to resolve grounding and policy errors in compositional decision-making tasks. Blocker: None

Enabling Compositional Generalization of AI Systems · MIT

Left open

Finetune large language models directly on large annotated hardware design codebases instead of using demonstration libraries. Blocker: None

Harnessing Large Language Models Towards More Accessible Hardware Accelerator Design · Georgia Tech

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.