Enhancing Emotion Recognition with MTAR
3724 10.21437/Interspeech.2024-2350
feature fusion strategy at the feature-level, originally integrat- 2.2.2. Speech Representations
ing a multimodal fusion module to enhance the acquisition of
modality-specific and modality-invariant representations. To obtain a comprehensive understanding of acoustic features,
we employ a pretrained SSL model, WavLM [22], as our speech
Our contributions can be summarized as follows: self-supervised encoder. WavLM utilizes a hybrid architec-
• We propose a novel feature extraction approach for learning ture that includes convolutional neural network (CNN) layers
multimodal emotion information by leveraging both the SSR and a transformer encoder to effectively capture speech fea-
and handcrafted features extracted from MTAR. To the best tures and contextual information. It has been fine-tuned for
of our knowledge, this is the first systematic attempt to in- various downstream speech tasks [23]. We denote HS =
tegrate music theory-inspired acoustics with SSR to comple- (h1S , h2S , · · · , hm
S ) ∈ R
m∗d
to represent the speech SSR, where
ment emotion-related information. m denotes the number of frames extracted from an utterance, d
• We propose a multimodal fusion module that comprehen- is the dimension of hidden representations.
sively integrates modality-specific and modality-invariant
2.2.3. Text Representations
emotional information from both speech and text modali-
ties. This module is designed to fully exploit emotional cues To acquire comprehensive information regarding lexical fea-
present in diverse modalities, ensuring a holistic understand- tures, we leverage a pretrained SSL model, RoBERTa [24],
ing of emotional content across different sources. as our text encoder. RoBERTa is an extension of the bidi-
• The experimental results demonstrate that the proposed ap- rectional encoder representations from the transformer (BERT)
proach adeptly addresses the MER tasks, surpassing the per- model, a widely-used model in natural language processing
formance of existing feature extraction and fusion methods. (NLP), which is specifically designed to address challenges re-
lated to long-range dependencies and is finely tuned for various
NLP tasks. Pretrained on extensive corpora, including a diverse
2. Proposed Method range of texts from various sources, such as a dataset compris-
This section details our proposed MER system, which is based ing 58 million tweets, RoBERTa exhibits exceptional contextual
on the multimodal fusion method leveraging the MTAR and understanding, thereby enhancing text-related tasks. We denote
SSR. As illustrated in Fig. 1, the network consists of three HT = (h1T , h2T , · · · , hn
T) ∈ R
n∗d
to represent the text SSR,
main components: an embedding module for encoding MTAR where n denotes the number of tokens extracted from an utter-
and SSR, a multimodal fusion module for integrating modality- ance, d is the dimension of hidden representations.
specific and modality-invariant emotional information, and an
2.3. Multimodal Fusion (MF) Module
emotion prediction module for predicting the emotion label.
On the basis of previous study [25], our MF is composed of
2.1. Model Description eight multi-level fusion (MLF) blocks, and two collaborative
As illustrated, raw audio utterances are fed into dedicated en- fusion (CF) blocks. The objective is to facilitate the learning of
coder networks designed to extract MTAR, including VEE- modality-specific representations and modality-invariant repre-
derived MTAR, AEP-derived MTAR, and SSR. Simultaneously, sentations.
transcripts undergo a self-supervised encoder to extract text In this section, we offer an in-depth explanation of the op-
SSR. These distinct modalities of information are then harmo- erations of the MLF and CF blocks.
nized utilizing the proposed multimodal fusion method, thereby MLF Block adheres to the structure of a standard trans-
facilitating the final emotion recognition process. former layer, incorporating a cross-attention module, residual
2.2. Embedding Module connections, and a BiGRU module. Initially, we employ four
MLF blocks to derive AEP-aware and VEE-aware text SSR
2.2.1. Music Theory-inspired Acoustic Representations (HTA , HTV ) ∈ Rn∗d , as well as AEP-aware and VEE-aware
speech SSR (HSA , HSV ) ∈ Rm∗d . This is achieved by utiliz-
We represent the speech sample using MTAR features, as ing HV (or HA ) as queries and HS (or HT ) as keys and values
demonstrated by Li et al. [16]. The MTAR features can be within each MLF block.
categorized into two groups: a 41-dimensional MTAR derived
from VEE and a 39-dimensional MTAR derived from AEP pro- Q = HV (or HA ), K = HS (or HT ), V = HS (or HT ). (1)
cesses. More specifically, the VEE-derived MTAR draws inspi-
ration from five music theory subgroups and includes descrip- HSA (orHSV , HTA , HTV ) = BiGRU(Cross-Attention(Q, K, V )).
tors related to ten MIDI notes, three music dynamics, five music (2)
main intervals, nine microtonal music attributes, and 114 de-
CF Block is employed to extract complementary informa-
scriptors associated with the syntactic structure of music. Addi-
tion from the AEP and VEE in MTAR:
tionally, the AEP-derived MTAR predominantly captures musi-
cal interval information, including melodic and harmonic in- First, we generate HSMTAR ∈ Rm∗m , HTMTAR ∈ Rn∗n
tervals. We adopt HA = (h1A , h2A , · · · , hm A) ∈ R
m∗d
and from AEP features HA and VEE features HV using a combi-
1 2 m
HV = (hV , hV , · · · , hV ) ∈ R m∗d
to symbolize the AEP and nation of BiGRU and fully-connected (FC) layers, as follows:
VEE-derived MTAR, respectively, where m denotes the num-
ber of frames extracted from an utterance, d is the dimension of HSMTAR (or HTMTAR ) = FC(BiGRU(HV ⊕ HA )). (3)
hidden representations. Then, the speech SSR HS or text SSR HT are multiplied by the
In the VEE encoder and AEP encoder, VEE-derived MTAR collaborative representation HSMTAR or HTMTAR to obtain the
and AEP-derived MTAR are processed by a Bidirectional Gated weighted collaborative SSR (HS′ ∈ Rm∗d , HT′ ∈ Rn∗d ).
Recurrent Unit (BiGRU) [21] with tanh as the activation func-
tion and a dropout rate of 0.5. HS′ (or HT′ ) = HS (or HT ) · HSMTAR (or HTMTAR ). (4)
3725
Figure 1: The overall architecture of our proposed method.
Next, we employ four additional MLF blocks to obtain both consistent with experimental protocols used in many previous
′ ′
MTAR-aware speech SSR (HSA , HSV ) ∈ Rm∗d and MTAR- studies [8, 16]: neutral, happiness, sadness, and anger. Addi-
′ ′
aware text SSR (HTA , HTV ) ∈ Rn∗d . This is achieved by em- tionally, we merge happiness and excitement into one category.
ploying HS′ (or HT′ ) as queries and (HSA , HSV , HTA , HTV ) as
keys and values within each MLF block. 3.2. Experimental Procedure
′ ′ ′ ′
We conduct three experiments in this study. In Experiment 1,
HSA (orHSV , HTA , HTV ) = BiGRU(Cross-Attention(Q, K, V )). we examine the effect of MTAR on SSR for MER, compar-
(5) ing it with a Mel Spectrogram. In Experiment 2, we explore
Finally, we adopt attentive statistics pooling [26] to obtain the effect of fusion methods on MER, using concatenation and
a 1-dimension vector for each output. The final MTAR-aware co-attention [8] as benchmarks from previous works. In Exper-
′ ′ ′ ′
speech and text representations (HSA , HSV , HTA , HTV ) are iment 3, we further investigate the effect of MTAR on com-
concatenated and written as follows: mon speech and text-based SSR (Wav2vec 2.0 [28], Hubert
′ ′ ′ ′
[29], WavLM [22], BERT [30], DeBERTa [31], RoBERTa [24])
MTAR
HST = HSA ⊕ HSV ⊕ HTA ⊕ HTV . (6) within the domain of emotion recognition tasks.
3726
3.4. Evaluation hancement in both single and multimodal emotion recognition
model performance. When compared with the commonly uti-
In evaluating our results on IEMOCAP, for which a standard
lized concatenate method, the observed improvements in UAR
train, dev, and test split is lacking, we adopt a common strat-
are 1.62% for the speech modality 2.28% and 2.11% for multi-
egy: leave-one-speaker-out cross-validation, as also employed
modal approaches. These results signify that the adoption of the
in previous works [32, 33]. The categorical MER performance
MF method enhances the performance of MER. Furthermore, in
is assessed using the Unweighted Average Recall (UAR) and F1
comparison with previous work [8], the observed improvements
scores, which have been widely utilized in experiments with un-
in UAR are 0.86% for the speech modality 1.61% and 1.51% for
balanced data to evaluate performance [34, 35], across the four
multimodal approaches. These findings suggest a new feasibil-
distinct emotional labels.
ity for feature fusion techniques, offering a promising path for
advancing emotion recognition systems.
4. Results and Discussion
Table 3: The effectiveness of MTAR on speech and text-based
To analyze the effect of MTAR on SSR for emotion recognition, SSR for emotion recognition.
we compare the prediction performance characteristics of single
and multimodal emotion recognition, as shown in Table 1.
Modality Model UAR (%) F1 (%)
Table 1: Comparison of MER performance obtained by different
Wav2vec 2.0 68.43 65.91
SSR that integrated the proposed MTAR and Mel Spectrogram. Hubert 68.47 67.72
WavLM 70.46 70.24
Single modal
Modality Model UAR (%) F1 (%)
Wav2vec 2.0 + MTAR 70.06 68.08
MTAR [16] 61.92 61.47 Hubert + MTAR 72.05 71.52
WavLM 70.46 70.24
Single modal WavLM + MTAR 73.47 73.01
WavLM + Mel Spectrogram 71.70 71.15
WavLM + MTAR 73.47 73.01
BERT 70.74 70.69
RoBERTa 71.33 71.08 DeBERTa 71.45 70.96
RoBERTa + Mel Spectrogram 72.12 72.08 RoBERTa 71.33 71.08
RoBERTa + MTAR 75.11 74.79 Multimodal
Multimodal
RoBERTa + WavLM 76.34 76.15 BERT + MTAR 73.91 73.52
RoBERTa + WavLM + Mel Spectrogram 77.08 76.42 DeBERTa + MTAR 74.87 74.51
RoBERTa + WavLM + MTAR 79.89 79.35 RoBERTa + MTAR 75.11 74.79
The results show that the utilization of MTAR yields su- In Table 3, we delve deeper into the effect of MTAR on
perior performance compared with relying solely on SSR. The common speech and text-based SSR. Our results reveal that
observed improvements in UAR are 3.01% for single modal ap- MTAR yields notable enhancements in the efficacy of com-
proaches, 3.78% for utilizing only text SSR, and 3.55% for em- mon SSR for emotion recognition. Specifically, for speech-
ploying both speech and text SSR in a multimodal approach. based SSR, the observed improvements in UAR are as fol-
These findings suggest that incorporating MTAR for emo- lows: Wav2vec 2.0 demonstrates a noteworthy enhancement of
tional recognition outperforms relying solely on SSR. More- 1.63%, Hubert exhibits a substantial improvement of 3.58%,
over, when compared with considering MTAR and speech SSR and WavLM shows a marked increase of 3.01%. Additionally,
or MTAR and text SSR, integrating of MTAR, speech and text for text-based SSR, BERT showcases a commendable enhance-
SSR results in notable enhancements, with UAR improvements ment of 3.17%, while DoBERTa achieves a notable improve-
of 6.42% and 4.78%, respectively. These results underscore the ment of 3.42%. RoBERTa demonstrates a remarkable increase
significance of integrating multiple modalities of information of 3.78%. Collectively, these results underscore the substantial
to further enhance the performance of emotion recognition sys- performance gains observed in SSR through the incorporation
tems. Noting that, compared with the Mel spectrogram, MTAR of MTAR.
achieved improvements of 1.77% in single modal, 2.99%, and
2.81% in multimodal approaches, respectively. 5. Conclusions and Future Work
To evaluate the effect of fusion methods on emotion recog-
nition, we contrast the predictive performance attributes of sin- In this paper, we integrated handcrafted MTAR and SSR for
gle and multimodal emotion recognition, delineated in Table 2. MER tasks. Our findings revealed superior performance when
incorporating handcrafted MTAR compared to merely integrat-
Table 2: Comparison of MER performance obtained by the pro- ing individual modality SSR, demonstrating the effectiveness of
posed MF and baseline fusion methods. MTAR in enhancing SSR in emotion recognition. Additionally,
we introduced a novel MF method for integrating modality-
Modality Model UAR (%) F1 (%) specific and modality-invariant emotional information, consis-
WavLM + MTAR (Concat) 71.85 71.12 tently achieving higher accuracy in MER. For future research,
Single modal WavLM + MTAR [8] 72.61 71.92
WavLM + MTAR (MF) 73.47 73.01
we advocate for further exploration of the intricate relationship
between music and speech to devise innovative fusion methods.
RoBERTa + MTAR (Concat) 72.83 72.24
RoBERTa + MTAR [8] 73.50 73.35
Multimodal
RoBERTa + MTAR (MF) 75.11 74.79 6. Acknowledgements
RoBERTa + WavLM + MTAR (Concat) 77.78 77.29
RoBERTa + WavLM + MTAR [8] 78.38 77.83 This work was financially supported by JST SPRING, Grant
RoBERTa + WavLM + MTAR (MF) 79.89 79.35 Number JPMJSP2125, and in part by JST CREST Grant Num-
ber JPMJCR19A3, Japan, and JSPS KAKENHI Grant Number
Evidently, the proposed fusion method demonstrates an en- 21H05054.
3727
7. References [19] L. Pepino, P. Riera, L. Ferrer, and A. Gravano, “Fusion ap-
proaches for emotion recognition from speech using acoustic
[1] P. Singh, R. Srivastava, K. Rana, and V. Kumar, “A multimodal and text-based features,” in ICASSP 2020-2020 IEEE Interna-
hierarchical approach to speech emotion recognition from audio tional Conference on Acoustics, Speech and Signal Processing
and text,” Knowledge-Based Systems, vol. 229, p. 107316, 2021. (ICASSP). IEEE, 2020, pp. 6484–6488.
[2] T. Deschamps-Berger, L. Lamel, and L. Devillers, “End-to-end [20] T. Mittal, U. Bhattacharya, R. Chandra, A. Bera, and D. Manocha,
speech emotion recognition: challenges of real-life emergency “M3er: Multiplicative multimodal emotion recognition using fa-
call centers data recordings,” in 2021 9th International Confer- cial, textual, and speech cues,” in Proceedings of the AAAI con-
ence on Affective Computing and Intelligent Interaction (ACII). ference on artificial intelligence, vol. 34, no. 02, 2020, pp. 1359–
IEEE, 2021, pp. 1–8. 1367.
[3] S. Zepf, J. Hernandez, A. Schmitt, W. Minker, and R. W. Picard, [21] J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evalu-
“Driver emotion recognition for intelligent vehicles: A survey,” ation of gated recurrent neural networks on sequence modeling,”
ACM Computing Surveys (CSUR), vol. 53, no. 3, pp. 1–30, 2020. arXiv preprint arXiv:1412.3555, 2014.
[4] M. Akagi, X. Han, R. Elbarougy, Y. Hamada, and J. Li, “To- [22] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li,
ward affective speech-to-speech translation: Strategy for emo- N. Kanda, T. Yoshioka, X. Xiao et al., “Wavlm: Large-scale self-
tional speech recognition and synthesis in multiple languages,” supervised pre-training for full stack speech processing,” IEEE
in Signal and Information Processing Association Annual Summit Journal of Selected Topics in Signal Processing, vol. 16, no. 6,
and Conference (APSIPA), 2014 Asia-Pacific. IEEE, 2014, pp. pp. 1505–1518, 2022.
1–10.
[23] S. Dang, T. Matsumoto, Y. Takeuchi, and H. Kudo, “Using Semi-
[5] J. Tian, D. Hu, X. Shi, J. He, X. Li, Y. Gao, T. Toda, X. Xu, and supervised Learning for Monaural Time-domain Speech Separa-
X. Hu, “Semi-supervised multimodal emotion recognition with tion with a Self-supervised Learning-based SI-SNR Estimator,” in
consensus decision-making and label correction,” in Proceedings Proc. INTERSPEECH 2023, 2023, pp. 3759–3763.
of the 1st International Workshop on Multimodal and Responsible [24] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy,
Affective Computing, 2023, pp. 67–73. M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A
[6] L.-W. Chen and A. Rudnicky, “Exploring wav2vec 2.0 fine tuning robustly optimized bert pretraining approach,” arXiv preprint
for improved speech emotion recognition,” in ICASSP 2023-2023 arXiv:1907.11692, 2019.
IEEE International Conference on Acoustics, Speech and Signal [25] J. He, X. Shi, X. Li, and T. Toda, “Mf-aed-aec: Speech emotion
Processing (ICASSP). IEEE, 2023, pp. 1–5. recognition by leveraging multimodal fusion, asr error detection,
[7] N. Naderi and B. Nasersharif, “Cross corpus speech emotion and asr error correction,” in ICASSP 2024-2024 IEEE Interna-
recognition using transfer learning and attention-based fusion of tional Conference on Acoustics, Speech and Signal Processing
wav2vec2 and prosody features,” Knowledge-Based Systems, vol. (ICASSP). IEEE, 2024, pp. 11 066–11 070.
277, p. 110814, 2023. [26] K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statis-
[8] H. Zou, Y. Si, C. Chen, D. Rajan, and E. S. Chng, “Speech emo- tics pooling for deep speaker embedding,” arXiv preprint
tion recognition with co-attention based multi-level acoustic infor- arXiv:1803.10963, 2018.
mation,” in ICASSP 2022-2022 IEEE International Conference on [27] C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower,
Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap:
pp. 7367–7371. Interactive emotional dyadic motion capture database,” Language
[9] S. Padi, S. O. Sadjadi, D. Manocha, and R. D. Sriram, resources and evaluation, vol. 42, pp. 335–359, 2008.
“Multimodal emotion recognition using transfer learning from [28] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec
speaker recognition and bert-based models,” arXiv preprint 2.0: A framework for self-supervised learning of speech repre-
arXiv:2202.08974, 2022. sentations,” Advances in neural information processing systems,
[10] K. R. Scherer, “Vocal affect expression: a review and a model vol. 33, pp. 12 449–12 460, 2020.
for future research.” Psychological bulletin, vol. 99, no. 2, p. 143, [29] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhut-
1986. dinov, and A. Mohamed, “Hubert: Self-supervised speech rep-
resentation learning by masked prediction of hidden units,”
[11] A. Gabrielsson and P. N. Juslin, Emotional expression in music.
IEEE/ACM Transactions on Audio, Speech, and Language Pro-
Oxford University Press, 2003.
cessing, vol. 29, pp. 3451–3460, 2021.
[12] A. D. Patel, “Language, music, syntax and the brain,” Nature neu- [30] X. Qin, Z. Wu, T. Zhang, Y. Li, J. Luan, B. Wang, L. Wang, and
roscience, vol. 6, no. 7, pp. 674–681, 2003. J. Cui, “Bert-erc: Fine-tuning bert is enough for emotion recogni-
[13] L. Jäncke, “The relationship between music and language,” p. tion in conversation,” in Proceedings of the AAAI Conference on
123, 2012. Artificial Intelligence, vol. 37, no. 11, 2023, pp. 13 492–13 500.
[14] T. Fujisawa, K. Takami, and N. D. Cook, “On the role of pitch [31] P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-
intervals in the perception of emotional speech,” in ISCA & IEEE enhanced bert with disentangled attention,” arXiv preprint
Workshop on Spontaneous Speech Processing and Recognition, arXiv:2006.03654, 2020.
2003. [32] Y. Gao, H. Shi, C. Chu, and T. Kawahara, “Enhancing two-
[15] B. Yang and M. Lugger, “Emotion recognition from speech sig- stage finetuning for speech emotion recognition using adapters,”
nals using new harmony features,” Signal processing, vol. 90, in ICASSP 2024-2024 IEEE International Conference on Acous-
no. 5, pp. 1415–1423, 2010. tics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp.
11 316–11 320.
[16] X. Li, X. Shi, D. Hu, Y. Li, Q. Zhang, Z. Wang, M. Unoki,
and M. Akagi, “Music theory-inspired acoustic representation for [33] Y. Gao, C. Chu, and T. Kawahara, “Two-stage finetuning of
speech emotion recognition,” IEEE/ACM Transactions on Audio, wav2vec 2.0 for speech emotion recognition with asr and gender
Speech, and Language Processing, 2023. pretraining,” in Proc. Interspeech, 2023.
[17] P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli, [34] X. Shi, S. Li, and J. Dang, “Dimensional emotion prediction based
“Multimodal fusion for multimedia analysis: a survey,” Multime- on interactive context in conversation.” in INTERSPEECH, 2020,
dia systems, vol. 16, pp. 345–379, 2010. pp. 4193–4197.
[35] X. Shi, X. Li, and T. Toda, “Emotion awareness in multi-utterance
[18] S. Yoon, S. Byun, and K. Jung, “Multimodal speech emotion
turn for improving emotion prediction in multi-speaker conversa-
recognition using audio and text,” in 2018 IEEE Spoken Language
tion,” in Proc. Interspeech, vol. 2023, 2023, pp. 765–769.
Technology Workshop (SLT). IEEE, 2018, pp. 112–118.
3728