Abstract
Background: Advances in generative artificial intelligence (GAI) has prompted interest in its application in sensitive fields, including mental health. Yet the validity of its nonverbal social-cognitive abilities remains unclear. This study addresses two issues that probe GAI’s social intelligence: its capacity to interpret nonverbal channels (bodily gestures and vocal tone) with accuracy comparable to humans, and whether accuracy of emotion recognition is different by emotional valence (positive vs. negative).Methods Emotion recognition accuracy of three Gemini GAI models (Pro 1.5, Pro 2 (Experimental) and Flash 2) was evaluated using the EU-Emotion Stimulus Set of bodily gestures and vocal tone depicting basic emotions. Each stimulus was presented 20 times to each model, yielding 9,360 responses. Model accuracy was compared to validated human benchmarks.Results: Advanced GAI models achieved near-human accuracy in recognizing emotions from vocal tone (e.g., Gemini Pro 2 vs. humans: z = 0.81, p = .416). However, their performance was significantly lower than that of humans in interpreting bodily gestures (e.g., Gemini Pro 2 vs. humans: z = -3.22, p = .001). GAI models recognized positive emotions more accurately than negative ones. This bias was significant across all models for bodily gestures and in the advanced models for vocal tone.Conclusion: GAI models exhibit a non-humanlike social-cognitive profile, excelling with positive emotions but struggling with the negative content. This finding carries profound clinical and theoretical relevance, highlighting the risks of integrating GAI into sensitive fields and underscoring the need for continued validation of its limitations and biases.
| Original language | English |
|---|---|
| Journal | IJHCS-D-25-01207 |
| DOIs | |
| State | E-pub ahead of print - 2025 |
Keywords
- Generative Artificial Intelligence (GAI)
- Emotion recognition
- Social Cognition
- Nonverbal Communication
- Algorithmic Bias
- Gemini Models
Fingerprint
Dive into the research topics of 'When the Body Speaks and the Tone Echoes: Exploring Gai Multimodal Emotion Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver