ResearchResearch paperLarge Language Models1 source · Oct 7, 2026

How Do LLMs Change Predictions Under Negation?

Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative.

Key points

  • Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it.
  • We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?").
  • Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris").
  • To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines.

Sources (1)

  • [1]How Do LLMs Change Predictions Under Negation?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:12 AM
    Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative.
    Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 7, 2026Q-Learning with Scalar Adjoint Matching
  3. Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
  4. Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
  5. Sep 28, 2026openai/openai-python v3.20.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related