ResearchPod Summary
As large language models (LLMs) become increasingly integrated into the film industry, researchers are concerned about how these tools might perpetuate or amplify social biases. This study investigates whether LLM-generated screenplays exhibit gender bias by comparing them to human-written scripts using the Bechdel test and social network analysis (SNA).
The researchers compiled a dataset of 783 human-written screenplays and used three state-of-the-art LLMs (GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5) to generate corresponding scripts based on anonymized plot synopses. To ensure a fair comparison, the team implemented a two-step generation process: first creating a scene list, then generating individual scenes to form a complete narrative. They then converted these scripts into character interaction networks, where nodes represent characters and edges represent dialogue interactions, allowing for a quantitative assessment of gender representation.
The study found that human-written scripts consistently outperform LLM-generated scripts on the Bechdel test. While the LLMs produced varied results across different network metrics—with some models occasionally showing higher proportions of female interactions or different patterns of connectivity—none of the models successfully eliminated the structural gender bias present in the narratives. Across all script types, male characters maintained higher centrality within the social networks, suggesting that LLMs tend to replicate or even exacerbate the male-centric structures found in traditional media.
This research highlights the risks of relying on generative AI for creative content production. If LLMs are used to draft screenplays, they may inadvertently reinforce existing gender disparities in media. The findings underscore the necessity of auditing AI systems not just for factual accuracy, but for the representational biases they encode in cultural narratives, which play a significant role in shaping social identity and perceptions.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.