Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence
For NLP practitioners using ratings as weak labels, this study provides evidence that ratings and textual sentiment often disagree, highlighting the need for validation.
This study investigates rating-sentiment incongruence in Sri Lankan tourism reviews, finding that 18.6% of 16,156 reviews show mismatch between star ratings and textual sentiment, with Conservative Rater and Obligatory 5-Star behaviors as main drivers. It demonstrates that star ratings are not interchangeable with textual sentiment and should be validated before use as ground-truth labels.
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak labels for textual sentiment, yet whether the two actually agree is rarely questioned. This study investigates sentiment-rating incongruence, where the sentiment expressed in review text differs from the sentiment implied by the assigned star rating, in Sri Lankan tourism attraction reviews. A dataset of 16,156 reviews from 2010 to 2023 is analyzed using a transformer-based sentiment pipeline that derives textual sentiment independently of assigned ratings. Incongruence occurs in 18.6% of reviews and falls into six directional patterns, with Conservative Rater and Obligatory 5-Star behaviors accounting for the majority of mismatches. Prevalence also varies across venue types, with museums showing the highest rates. Statistical tests, logistic regression, Random Forest, and SHAP analysis identify venue type, reviewer expertise, review length, and temporal factors as contributors to rating-text divergence. Overall, this study demonstrates that star ratings are not interchangeable with textual sentiment and should be validated before being treated as ground-truth labels in NLP.