ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Latent space bias directions in LLMs capture confidence, not fairness

Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models.

Key points

  • We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks.
  • We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias.
  • Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics.
  • Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care.

Sources (1)

  • [1]Latent space bias directions in LLMs capture confidence, not fairness
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:42 PM
    Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models.
    We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks.

Extractive summary: sentences quoted from the sources.

Related