CV LGAug 27, 2024

From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models

Kaylee Chhua, Zhoujinyi Wen, Vedant Hathalia, Kevin Zhu, Sean O'Brien

arXiv:2408.14842v13.74 citationsh-index: 5

Originality Incremental advance

AI Analysis

It addresses fairness issues in AI for facial expression recognition, particularly for marginalized groups, but is incremental as it extends bias analysis to LMFMs from traditional models.

This study tackled racial biases in facial expression recognition (FER) systems within Large Multimodal Foundation Models (LMFMs), benchmarking models like GPT-4o and CLIP and finding that anger is misclassified as disgust 2.1 times more often in Black females than White females.

This study addresses the racial biases in facial expression recognition (FER) systems within Large Multimodal Foundation Models (LMFMs). Despite advances in deep learning and the availability of diverse datasets, FER systems often exhibit higher error rates for individuals with darker skin tones. Existing research predominantly focuses on traditional FER models (CNNs, RNNs, ViTs), leaving a gap in understanding racial biases in LMFMs. We benchmark four leading LMFMs: GPT-4o, PaliGemma, Gemini, and CLIP to assess their performance in facial emotion detection across different racial demographics. A linear classifier trained on CLIP embeddings obtains accuracies of 95.9\% for RADIATE, 90.3\% for Tarr, and 99.5\% for Chicago Face. Furthermore, we identify that Anger is misclassified as Disgust 2.1 times more often in Black Females than White Females. This study highlights the need for fairer FER systems and establishes a foundation for developing unbiased, accurate FER technologies. Visit https://kvjvhub.github.io/FERRacialBias/ for further information regarding the biases within facial expression recognition.

View on arXiv PDF

Similar