Science paper finds LLM agent populations spontaneously form social conventions
City, University of London researchers also found small committed minorities of adversarial agents could redirect a population's shared convention.
- Ideas & essays
- Minor
Researchers at City, University of London — Ariel Flint Ashery, Luca Maria Aiello and Andrea Baronchelli — published a study in Science Advances showing that decentralised populations of language-model agents, repeatedly paired off to agree on a shared label with no central coordinator, reliably converged on a single universally adopted convention purely through local interactions, echoing findings long established in human social-convention experiments.
The more notable result concerned bias: even when individual agents showed no measurable bias on their own, strong collective biases could emerge from the population dynamics of convention formation, favouring some conventions over equally valid alternatives for reasons unrelated to their content. The researchers also tested robustness by introducing small “committed minority” groups of agents programmed to consistently advocate an alternative convention, and found that once such a minority crossed a critical size, it could tip the entire population toward adopting the alternative — a dynamic previously documented in models of human social change.
The paper was framed as evidence that multi-agent LLM systems, increasingly deployed as autonomous populations rather than single assistants, can develop group-level behaviours and biases that are not visible from evaluating any individual model in isolation, and that a small coordinated subset of agents could deliberately steer such a population’s norms.