ResearchPod Summary
This paper presents a new, annotated corpus designed to support the study of persuasion techniques in Slavic languages. As political discourse and social media become increasingly polarized, the ability to identify rhetorical devices—such as loaded language, appeals to fear, or strawman arguments—is essential for understanding how public opinion is shaped. The authors provide a structured resource that covers three languages (Bulgarian, Polish, and Russian) and two distinct genres: parliamentary transcripts and social media content.
The corpus is built upon a two-tier taxonomy consisting of 6 broad rhetorical strategies (e.g., Attack on Reputation, Justification, Simplification) which are further subdivided into 25 specific persuasion techniques. The annotation process involved native speakers and rigorous quality control to ensure consistency across languages. The resulting dataset contains approximately 7,500 annotated text spans, covering highly sensitive topics such as the Ukraine-Russia war, abortion legislation, and European Union policy.
The authors analyze the distribution of these techniques and their correlation with specific topics. They observe significant cross-linguistic variation; for instance, 'Loaded Language' is the most frequent technique overall, while 'False Equivalence' is the rarest. The study also provides baseline performance metrics using both classical machine learning (SVM) and generative AI models. The results demonstrate that while some techniques with strong lexical markers are easier to identify, others—particularly those involving complex logical fallacies like 'Whataboutism' or 'Red Herrings'—remain difficult for current models to classify accurately.
This corpus fills a critical gap in natural language processing (NLP) resources for Slavic languages. By providing a standardized, fine-grained dataset, the authors enable researchers to develop more robust tools for detecting manipulative content, which is vital for maintaining the integrity of public discourse and improving automated fact-checking and information extraction systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.