Latent space bias directions in LLMs capture confidence, not fairness
Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models.
ProofPaper ↗
Key points
- We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks.
- We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias.
- Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics.
- Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care.
Sources (1)
- [1]Latent space bias directions in LLMs capture confidence, not fairnessarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:42 PM
Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models.
We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks.
Extractive summary: sentences quoted from the sources.