ResearchPod Summary
Traditional fully-supervised Salient Object Detection (SOD) models require expensive, pixel-level manual annotations. This paper addresses the challenge of achieving high-performance RGB-D SOD using only sparse, cost-effective scribble annotations, which typically lack the structural detail necessary for accurate saliency map generation.
The authors introduce a two-stage framework:
The proposed method effectively bridges the gap between weakly-supervised and fully-supervised performance. By utilizing SAM to generate dense pseudo-labels, the model overcomes the limitations of previous scribble-supervised approaches that relied on rougher, model-predicted labels. Experimental results across seven standard datasets demonstrate that the combination of SAM-PAG and S²Diff outperforms existing scribble-supervised methods and achieves competitive results compared to fully-supervised benchmarks.
This research provides a scalable solution for salient object detection by significantly reducing the annotation burden while maintaining high accuracy. By integrating foundation models like SAM with advanced generative architectures like conditional diffusion models, the study offers a robust pipeline for tasks where pixel-level ground truth is unavailable or prohibitively expensive to obtain.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.