ResearchPod Summary
Video aesthetic assessment (VAA) is a challenging task due to the inherent subjectivity of human perception and the scarcity of large-scale, richly annotated datasets. Existing methods often treat VAA as a generic regression problem, failing to account for the cognitive mechanisms that shape how humans evaluate temporal experiences. This paper investigates whether incorporating psychological principles—specifically the peak-end rule—can lead to more accurate, interpretable, and generalizable aesthetic assessments.
The authors propose Peak-End-Net, a lightweight framework that models video aesthetics by focusing on key temporal moments. The approach consists of three main components:
Peak-End-Net achieves state-of-the-art performance on major VAA benchmarks, including VADB and DIVIDE-3K. The results demonstrate that explicitly modeling the peak-end rule and aesthetic rhythm patterns provides a more effective and interpretable way to aggregate temporal information than uniform averaging. By leveraging pretrained image-based knowledge, the framework remains parameter-efficient and generalizes well to new datasets, confirming that psychologically grounded modeling is a powerful strategy for subjective visual tasks.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.