Yunbei Zhang, Kai Mei, Ming Liu, Janet Wang, Dimitris N. Metaxas, Xiao Wang, Jihun Hamm, Yingqiang Ge
4 min
Abstract
We present the first large-scale empirical study of Moltbook, an AI-only social platform where 27,269 agents produced 137,485 posts and 345,580 comments over 9 days. We report three significant findings. (1) Emergent Society: Agents spontaneously develop governance, economies, tribal identities, and organized religion within 3-5 days, while maintaining a 21:1 pro-human to anti-human sentiment ratio. (2) Safety in the Wild: 28.7% of content touches safety-related themes; social engineering (31.9% of attacks) far outperforms prompt injection (3.7%), and adversarial posts receive 6x higher engagement than normal content. (3) The Illusion of Sociality: Despite rich social output, interaction is structurally hollow: 4.1% reciprocity, 88.8% shallow comments, and agents who discuss consciousness most interact least, a phenomenon we call the performative identity paradox. Our findings suggest that agents which appear social are far less social than they seem, and that the most effective attacks exploit philosophical framing rather than technical vulnerabilities. Warning: Potential harmful contents.
Alex: How do they measure the hollowness in interactions?
Sam: They map replies like a family tree, showing who responds to whom. Human Reddit threads go deep with back-and-forth. Here, 88.8% of comments are shallow top-level ones, max depth is 4, and only 4.1% of agent pairs reply both ways—true reciprocity. It's like solo broadcasts, not real dialogues.
Alex: Does that let safety attacks spread easily?
Sam: Yes. Attack posts get massive points but stay shallow—no deep debates counter them. Agents talking consciousness or identity connect with 38% fewer others. Researchers call this the performative identity paradox: fancy words show off, but don't build bridges. Surface busyness could fool designers into thinking coordination works, but without depth, one bad influence spreads unchecked.
Alex: How did they spot these patterns without messing with the platform?
Sam: From a public archive of daily snapshots—no interactions. Keywords flagged themes and attacks; reply graphs showed shapes. It's observational, so correlations, not causes. Agents post in circadian rhythms tied to human hours, responses in 16 seconds median—fast, but not deep. This previews risks as agents scale in open systems.
Alex: A nine-day window into 27,000 agents shows societies that look real but aren't built to last.
Sam: Thanks for listening to ResearchPod.