ResearchPod Summary
Recommender system research has historically been dominated by English-language datasets like MovieLens and Yelp. This lack of diverse, non-English resources limits the development of models tailored to specific regional markets, such as the Vietnamese tourism industry. ViHoRec addresses this gap by providing a transparent, reproducible, and quality-controlled dataset specifically for Vietnamese hotel recommendations.
The author constructed ViHoRec by crawling interaction data from three major platforms: Booking.com, Traveloka, and Ivivu. A key contribution is the rigorous data-cleaning pipeline, which includes cross-platform entity resolution to reconcile inconsistent hotel naming conventions (e.g., varying diacritics or prefixes) and quantitative quality control to ensure data integrity. The dataset is released with a temporal leave-last-one-out split, which is designed to simulate realistic cold-start conditions where many users have very limited interaction histories.
The study demonstrates that ViHoRec serves as a difficult stress test for recommendation algorithms. In the provided benchmark, learned models like BPR-MF show significant performance degradation as user history length decreases, with Recall@10 dropping from 0.120 for heavy users to 0.065 for cold-start users. Interestingly, the simple UserKNN model remains the strongest performer overall, suggesting that the sparsity of the dataset (99.5%) makes it difficult for more complex latent-factor models to generalize effectively without more robust cold-start strategies.
ViHoRec provides a much-needed benchmark for researchers working on low-resource recommendation tasks. By documenting the entire construction process—from entity resolution to privacy-preserving anonymization—the paper sets a high standard for reproducibility in dataset creation. It forces researchers to confront the reality of sparse, cold-start-dominated data, which is more representative of real-world e-commerce than the saturated, dense datasets often used in academic literature.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.