AION
Research paperReinforcement Learning · Agents & Tool Use · Large Language Models1 source · Oct 7, 2026

Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation

Language-model agents favor their own group because they have watched their members favor each other.

Key points

  • We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls.
  • First, the large effect of a bare group label reported in earlier work appears only when giving others points costs the agent nothing; once the agent can keep points for itself, that effect collapses on every model that shows it.
  • Second, under a stake, interaction history becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5-8 points out of 10 on 11, grows with the number of rounds played, and extends to labeled strangers the agent has never met.
  • Group favoritism is thus conformity to observed group behavior, carried to strangers by the label and overridden by individual reputation.

Sources (1)

  • [1]Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:46 PM
    Language-model agents favor their own group because they have watched their members favor each other.
    We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls.

Extractive summary: sentences quoted from the sources.