ResearchPod Summary
This study investigates the challenges of annotating evaluative language—specifically the Appraisal theory categories of Affect, Judgement, and Appreciation—within popular science discourse. The authors compare the performance of linguists in training against a trained expert linguist and several large language models (LLMs). By using a corpus of English TED talk transcripts, the researchers aim to determine if LLMs can effectively resolve complex, subjective linguistic annotation tasks that typically yield low inter-annotator agreement among humans.
The researchers compiled the 'EmotionalizTED' corpus, consisting of 190 TED talk transcripts across eleven domains. Twenty-four students (linguists in training) performed manual annotations on a sentence level to avoid the difficulties associated with identifying the exact span of evaluative expressions. A senior linguist provided a gold-standard reference set for comparison. The authors then tested three LLMs using various prompts, ultimately finetuning the best-performing model to classify the Attitude categories.
The study reveals that while linguists in training struggle to reach high agreement scores due to the subjective and context-dependent nature of Appraisal theory, LLMs perform significantly better. The models achieved an F1-score of 0.77, showing a high degree of alignment with the trained expert linguist. The findings suggest that LLMs are capable of handling complex, hierarchical linguistic theories, potentially reducing the need for extensive manual labor in digital humanities and computational social science research.
Manual annotation of complex linguistic constructs is time-consuming and often inconsistent. This paper demonstrates that LLMs can serve as reliable tools for automating the classification of evaluative language, opening new pathways for large-scale discourse analysis. By validating LLM performance against expert human benchmarks, the study provides a framework for researchers to integrate machine-supported annotation into their workflows, even for highly subjective theoretical constructs.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.