ResearchPod Summary
As audiobook catalogs grow, platforms face the challenge of matching listeners with the right content and the right narrator. While content-based features (genre, author, topic) are well-studied, the specific acoustic and paralinguistic qualities of narration—such as tone, pace, and vocal texture—remain under-explored. This study investigates whether these narration qualities influence listener appeal and whether they can be used to improve personalization and narrator casting.
The researchers analyzed a large dataset of 8,854 English-language audiobooks from LibriVox, featuring 1,206 unique narrators across 65 genres. They extracted 129 acoustic and vocal features—including frequency, energy, spectral properties, and tempo—using pre-trained audio models. To measure appeal, they used a 'view-rate' (views per day since publication) as a proxy. The study employed generalized linear models (GLMs) to assess global effects, genre-specific models to capture nuances, and linear mixed-effects (LME) models to isolate narration quality from title-specific biases. They further validated these findings using more granular 'return-rate' data from a proprietary Spotify dataset.
The study demonstrates that acoustic features have a consistent, non-trivial association with audiobook appeal, explaining a measurable portion of variance even when content-level factors are controlled. Key findings include:
These results provide the first systematic computational evidence that narration style is a distinct, measurable driver of listener engagement. By demonstrating that acoustic features can predict appeal independently of the book's content, the study offers a foundation for developing more sophisticated audiobook recommendation systems and data-driven narrator casting tools that prioritize the 'voice' of the story.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.