Enhancing Emotion Recognition with MTAR

0% found this document useful (0 votes)
7 views5 pages

Uploaded by

김민석
License
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Interspeech 2024

1-5 September 2024, Kos, Greece

Multimodal Fusion of Music Theory-Inspired and Self-Supervised


Representations for Improved Emotion Recognition

Xiaohan Shi1 , Xingfeng Li2 , Tomoki Toda1


1
Nagoya University, Japan
2
Hainan University, China
[Link]@[Link], lixingfeng@[Link],
tomoki@[Link]

Abstract evident in providing complementary information to SSR [7].


For example, Zou et al. incorporated multiple acoustic fea-
Multimodal emotion recognition (MER) is a rapidly evolving tures, including Mel-frequency cepstral coefficients, spectro-
field aimed at integrating information from various modali- grams, and Wav2vec 2.0 embeddings, for categorical emotion
ties, such as speech and text, to deepen our understanding of recognition. Their approach yielded an absolute improvement
emotions. However, challenges in feature extraction and fu- of 7.03% compared with using Wav2vec 2.0 embeddings alone
sion hinder further advancements in MER performance. To [8]. Moreover, Padi et al. presented a multimodal emotion
address these challenges, we propose a MER method using recognition (MER) framework using Mel spectrogram and fine-
self-supervised representations and handcrafted music theory- tuning of pre-trained BERT models, offering complementary
inspired representations across different modalities to compre- emotional insights derived from both speech and text modali-
hensively capture emotional information. Additionally, we in- ties [9].
troduce a novel multimodal fusion method to explore modality-
specific and modality-invariant relationships, thereby reducing Motivated by the findings of these studies, this paper fo-
distribution gaps between different modalities in MER. Exten- cuses on the feature extraction of multimodal emotions by in-
sive experimental validation underscores the effectiveness of tegrating SSR and handcrafted features, particularly exploring
our approach, with state-of-the-art results showing a 3.55% im- handcrafted features of music theory-inspired acoustic repre-
provement compared with the baseline. These results validate sentation (MTAR) to enhance the effectiveness of SSR in emo-
the effectiveness of our proposed method, signifying a notable tion recognition. Historically and today, speech and music ex-
enhancement in MER performance. pression and perception have dominated the nonverbal commu-
Index Terms: multimodal emotion recognition, multimodal fu- nication of emotion [10, 11]. There is widespread evidence
sion, self-supervised learning supporting the recognition of emotion in both stimuli, which
has been proven to develop in parallel [12, 13]. Fujisawa et
al. studied the relationship between emotion perception and
1. Introduction F0 on the basis of music intervals, suggesting that the interval
Emotion recognition has been a growing area of focus within structure is a fruitful means for determining the emotional va-
affective computing for interpreting a human subject’s feelings, lence of speech [14]. In addition, Yang et al. applied harmony
thoughts and behavioral responses. With the widespread use perception known from music to improve emotion recognition
of social networks, more and more people tend to express their performance and confirmed that musical interval-inspired rep-
feelings by sharing speech on the internet. Therefore, a consid- resentations associated with F0 counters are important cues for
erable amount of work has recently focused on multimodal ap- affective speech content [15]. Li et al. further solidified the
proaches that utilize both speech and text information to explore link between music and speech by designing an acoustic repre-
acoustic and lexical cues in emotional states [1]. Its applications sentation of emotional speech regarding vocal emotion expres-
have extended across diverse domains of human-computer in- sion (VEE) and auditory emotion perception (AEP) processes
teraction, including call-center interfaces [2], in-vehicle dash- via investigating music theory contents and showing promis-
board systems [3], speech-to-speech translation platforms [4], ing performance compared with acoustic traditions of speech
and others. Despite the considerable progress achieved by these in emotion analysis [16]. In line with these advancements, our
approaches, there still exist two fundamental issues for this task: study first introduces a robust feature extraction method applied
1) how to extract acoustic and lexical features that are most ef- to MER by integrating the handcrafted MTAR and SSR.
fective in distinguishing between emotions; and 2) how to de- The fusion strategy is another key aspect of MER. Two fun-
sign suitable fusion methods for integrating multiple modalities damental fusion strategies for MER are feature and decision-
for emotion recognition. This study addresses both issues to level fusion [17]. For instance, Yoon et al. proposed a deep
model the recognition process of emotion in a multimodal sce- dual recurrent neural network (RNN) for encoding audio-text
nario. sequences, subsequently concatenating their outputs to pre-
Most features employed in the speech and text modali- dict the final emotion. It achieves an accuracy of 71.8% on
ties for emotion recognition are typically categorized into two IEMOCAP [18]. Moreover, Pepino et al. designed multiple
main groups: self-supervised representations (SSR) and hand- dual RNNs to represent audio-text sequences, and conducted a
crafted features [5]. Within its areas of interest, SSR have re- comparative analysis between feature and decision-level fusion
cently gained considerable attention due to their advances in approaches, highlighting their comparable performance [19].
capturing multiple emotional characteristics comprehensively Most researchers believe that decision-level fusion is performed
and reliably [6]. Incidentally, handcrafted features also play more easily but ignore the relevance among representations of
a significant role in emotion recognition, and are increasingly different modalities [20]. In contrast, our study introduces a

3724 10.21437/Interspeech.2024-2350
feature fusion strategy at the feature-level, originally integrat- 2.2.2. Speech Representations
ing a multimodal fusion module to enhance the acquisition of
modality-specific and modality-invariant representations. To obtain a comprehensive understanding of acoustic features,
we employ a pretrained SSL model, WavLM [22], as our speech
Our contributions can be summarized as follows: self-supervised encoder. WavLM utilizes a hybrid architec-
• We propose a novel feature extraction approach for learning ture that includes convolutional neural network (CNN) layers
multimodal emotion information by leveraging both the SSR and a transformer encoder to effectively capture speech fea-
and handcrafted features extracted from MTAR. To the best tures and contextual information. It has been fine-tuned for
of our knowledge, this is the first systematic attempt to in- various downstream speech tasks [23]. We denote HS =
tegrate music theory-inspired acoustics with SSR to comple- (h1S , h2S , · · · , hm
S ) ∈ R
m∗d
to represent the speech SSR, where
ment emotion-related information. m denotes the number of frames extracted from an utterance, d
• We propose a multimodal fusion module that comprehen- is the dimension of hidden representations.
sively integrates modality-specific and modality-invariant
2.2.3. Text Representations
emotional information from both speech and text modali-
ties. This module is designed to fully exploit emotional cues To acquire comprehensive information regarding lexical fea-
present in diverse modalities, ensuring a holistic understand- tures, we leverage a pretrained SSL model, RoBERTa [24],
ing of emotional content across different sources. as our text encoder. RoBERTa is an extension of the bidi-
• The experimental results demonstrate that the proposed ap- rectional encoder representations from the transformer (BERT)
proach adeptly addresses the MER tasks, surpassing the per- model, a widely-used model in natural language processing
formance of existing feature extraction and fusion methods. (NLP), which is specifically designed to address challenges re-
lated to long-range dependencies and is finely tuned for various
NLP tasks. Pretrained on extensive corpora, including a diverse
2. Proposed Method range of texts from various sources, such as a dataset compris-
This section details our proposed MER system, which is based ing 58 million tweets, RoBERTa exhibits exceptional contextual
on the multimodal fusion method leveraging the MTAR and understanding, thereby enhancing text-related tasks. We denote
SSR. As illustrated in Fig. 1, the network consists of three HT = (h1T , h2T , · · · , hn
T) ∈ R
n∗d
to represent the text SSR,
main components: an embedding module for encoding MTAR where n denotes the number of tokens extracted from an utter-
and SSR, a multimodal fusion module for integrating modality- ance, d is the dimension of hidden representations.
specific and modality-invariant emotional information, and an
2.3. Multimodal Fusion (MF) Module
emotion prediction module for predicting the emotion label.
On the basis of previous study [25], our MF is composed of
2.1. Model Description eight multi-level fusion (MLF) blocks, and two collaborative
As illustrated, raw audio utterances are fed into dedicated en- fusion (CF) blocks. The objective is to facilitate the learning of
coder networks designed to extract MTAR, including VEE- modality-specific representations and modality-invariant repre-
derived MTAR, AEP-derived MTAR, and SSR. Simultaneously, sentations.
transcripts undergo a self-supervised encoder to extract text In this section, we offer an in-depth explanation of the op-
SSR. These distinct modalities of information are then harmo- erations of the MLF and CF blocks.
nized utilizing the proposed multimodal fusion method, thereby MLF Block adheres to the structure of a standard trans-
facilitating the final emotion recognition process. former layer, incorporating a cross-attention module, residual
2.2. Embedding Module connections, and a BiGRU module. Initially, we employ four
MLF blocks to derive AEP-aware and VEE-aware text SSR
2.2.1. Music Theory-inspired Acoustic Representations (HTA , HTV ) ∈ Rn∗d , as well as AEP-aware and VEE-aware
speech SSR (HSA , HSV ) ∈ Rm∗d . This is achieved by utiliz-
We represent the speech sample using MTAR features, as ing HV (or HA ) as queries and HS (or HT ) as keys and values
demonstrated by Li et al. [16]. The MTAR features can be within each MLF block.
categorized into two groups: a 41-dimensional MTAR derived
from VEE and a 39-dimensional MTAR derived from AEP pro- Q = HV (or HA ), K = HS (or HT ), V = HS (or HT ). (1)
cesses. More specifically, the VEE-derived MTAR draws inspi-
ration from five music theory subgroups and includes descrip- HSA (orHSV , HTA , HTV ) = BiGRU(Cross-Attention(Q, K, V )).
tors related to ten MIDI notes, three music dynamics, five music (2)
main intervals, nine microtonal music attributes, and 114 de-
CF Block is employed to extract complementary informa-
scriptors associated with the syntactic structure of music. Addi-
tion from the AEP and VEE in MTAR:
tionally, the AEP-derived MTAR predominantly captures musi-
cal interval information, including melodic and harmonic in- First, we generate HSMTAR ∈ Rm∗m , HTMTAR ∈ Rn∗n
tervals. We adopt HA = (h1A , h2A , · · · , hm A) ∈ R
m∗d
and from AEP features HA and VEE features HV using a combi-
1 2 m
HV = (hV , hV , · · · , hV ) ∈ R m∗d
to symbolize the AEP and nation of BiGRU and fully-connected (FC) layers, as follows:
VEE-derived MTAR, respectively, where m denotes the num-
ber of frames extracted from an utterance, d is the dimension of HSMTAR (or HTMTAR ) = FC(BiGRU(HV ⊕ HA )). (3)
hidden representations. Then, the speech SSR HS or text SSR HT are multiplied by the
In the VEE encoder and AEP encoder, VEE-derived MTAR collaborative representation HSMTAR or HTMTAR to obtain the
and AEP-derived MTAR are processed by a Bidirectional Gated weighted collaborative SSR (HS′ ∈ Rm∗d , HT′ ∈ Rn∗d ).
Recurrent Unit (BiGRU) [21] with tanh as the activation func-
tion and a dropout rate of 0.5. HS′ (or HT′ ) = HS (or HT ) · HSMTAR (or HTMTAR ). (4)

3725
Figure 1: The overall architecture of our proposed method.

Next, we employ four additional MLF blocks to obtain both consistent with experimental protocols used in many previous
′ ′
MTAR-aware speech SSR (HSA , HSV ) ∈ Rm∗d and MTAR- studies [8, 16]: neutral, happiness, sadness, and anger. Addi-
′ ′
aware text SSR (HTA , HTV ) ∈ Rn∗d . This is achieved by em- tionally, we merge happiness and excitement into one category.
ploying HS′ (or HT′ ) as queries and (HSA , HSV , HTA , HTV ) as
keys and values within each MLF block. 3.2. Experimental Procedure

′ ′ ′ ′
We conduct three experiments in this study. In Experiment 1,
HSA (orHSV , HTA , HTV ) = BiGRU(Cross-Attention(Q, K, V )). we examine the effect of MTAR on SSR for MER, compar-
(5) ing it with a Mel Spectrogram. In Experiment 2, we explore
Finally, we adopt attentive statistics pooling [26] to obtain the effect of fusion methods on MER, using concatenation and
a 1-dimension vector for each output. The final MTAR-aware co-attention [8] as benchmarks from previous works. In Exper-
′ ′ ′ ′
speech and text representations (HSA , HSV , HTA , HTV ) are iment 3, we further investigate the effect of MTAR on com-
concatenated and written as follows: mon speech and text-based SSR (Wav2vec 2.0 [28], Hubert
′ ′ ′ ′
[29], WavLM [22], BERT [30], DeBERTa [31], RoBERTa [24])
MTAR
HST = HSA ⊕ HSV ⊕ HTA ⊕ HTV . (6) within the domain of emotion recognition tasks.

2.4. Emotion Classification Module 3.3. Implementation


Emotion classification is conducted using the output representa- Our deep learning models were developed using Python 3.7 and
MTAR
tions HST of the MF module, which is subsequently passed PyTorch 1.11.0. The model was trained and evaluated on a
through a FC layer and a SoftMax activation function. computer with Intel(R) Xeon(R) Gold 6248 CPU @ 2.50GHz,
32GB RAM, and one NVIDIA Tesla V100 GPU.
MTAR MTAR
P (yemo |HST ) = SoftMax(FC(HST )). (7) For the speech modality, the self-supervised encoder was
initialized using the WavLM model1 , resulting in speech SSR
where yemo is the predicted emotion classification. with a dimensionality of 1024. For the text modality, we utilized
ground truth text from IEMOCAP as the source for the self-
3. Experiments supervised encoder. Specifically, we employed the RoBERTa
model2 , with a hidden size of 768, 12 attention layers, and 12
3.1. Dataset
attention heads. Both the WavLM and RoBERTa models un-
Interactive Emotional Dyadic Motion Capture (IEMOCAP) derwent fine-tuning during the training process. Most specifi-
database is a widely used corpus in affective computing [27]. cally, all speech fine-tuned SSL models utilize the large-sized
It contains approximately 12 hours of audio-visual recordings model, while all text fine-tuned SSL models utilize the base-
and is designed for two-person dialogs. Each dialog in IEMO- sized model.
CAP has been segmented into utterances with continuous labels
in the Valence-Arousal dimension and category labels for spe- 1 [Link]

cific emotional states. We consider four categorical emotions 2 [Link]

3726
3.4. Evaluation hancement in both single and multimodal emotion recognition
model performance. When compared with the commonly uti-
In evaluating our results on IEMOCAP, for which a standard
lized concatenate method, the observed improvements in UAR
train, dev, and test split is lacking, we adopt a common strat-
are 1.62% for the speech modality 2.28% and 2.11% for multi-
egy: leave-one-speaker-out cross-validation, as also employed
modal approaches. These results signify that the adoption of the
in previous works [32, 33]. The categorical MER performance
MF method enhances the performance of MER. Furthermore, in
is assessed using the Unweighted Average Recall (UAR) and F1
comparison with previous work [8], the observed improvements
scores, which have been widely utilized in experiments with un-
in UAR are 0.86% for the speech modality 1.61% and 1.51% for
balanced data to evaluate performance [34, 35], across the four
multimodal approaches. These findings suggest a new feasibil-
distinct emotional labels.
ity for feature fusion techniques, offering a promising path for
advancing emotion recognition systems.
4. Results and Discussion
Table 3: The effectiveness of MTAR on speech and text-based
To analyze the effect of MTAR on SSR for emotion recognition, SSR for emotion recognition.
we compare the prediction performance characteristics of single
and multimodal emotion recognition, as shown in Table 1.
Modality Model UAR (%) F1 (%)
Table 1: Comparison of MER performance obtained by different
Wav2vec 2.0 68.43 65.91
SSR that integrated the proposed MTAR and Mel Spectrogram. Hubert 68.47 67.72
WavLM 70.46 70.24
Single modal
Modality Model UAR (%) F1 (%)
Wav2vec 2.0 + MTAR 70.06 68.08
MTAR [16] 61.92 61.47 Hubert + MTAR 72.05 71.52
WavLM 70.46 70.24
Single modal WavLM + MTAR 73.47 73.01
WavLM + Mel Spectrogram 71.70 71.15
WavLM + MTAR 73.47 73.01
BERT 70.74 70.69
RoBERTa 71.33 71.08 DeBERTa 71.45 70.96
RoBERTa + Mel Spectrogram 72.12 72.08 RoBERTa 71.33 71.08
RoBERTa + MTAR 75.11 74.79 Multimodal
Multimodal
RoBERTa + WavLM 76.34 76.15 BERT + MTAR 73.91 73.52
RoBERTa + WavLM + Mel Spectrogram 77.08 76.42 DeBERTa + MTAR 74.87 74.51
RoBERTa + WavLM + MTAR 79.89 79.35 RoBERTa + MTAR 75.11 74.79

The results show that the utilization of MTAR yields su- In Table 3, we delve deeper into the effect of MTAR on
perior performance compared with relying solely on SSR. The common speech and text-based SSR. Our results reveal that
observed improvements in UAR are 3.01% for single modal ap- MTAR yields notable enhancements in the efficacy of com-
proaches, 3.78% for utilizing only text SSR, and 3.55% for em- mon SSR for emotion recognition. Specifically, for speech-
ploying both speech and text SSR in a multimodal approach. based SSR, the observed improvements in UAR are as fol-
These findings suggest that incorporating MTAR for emo- lows: Wav2vec 2.0 demonstrates a noteworthy enhancement of
tional recognition outperforms relying solely on SSR. More- 1.63%, Hubert exhibits a substantial improvement of 3.58%,
over, when compared with considering MTAR and speech SSR and WavLM shows a marked increase of 3.01%. Additionally,
or MTAR and text SSR, integrating of MTAR, speech and text for text-based SSR, BERT showcases a commendable enhance-
SSR results in notable enhancements, with UAR improvements ment of 3.17%, while DoBERTa achieves a notable improve-
of 6.42% and 4.78%, respectively. These results underscore the ment of 3.42%. RoBERTa demonstrates a remarkable increase
significance of integrating multiple modalities of information of 3.78%. Collectively, these results underscore the substantial
to further enhance the performance of emotion recognition sys- performance gains observed in SSR through the incorporation
tems. Noting that, compared with the Mel spectrogram, MTAR of MTAR.
achieved improvements of 1.77% in single modal, 2.99%, and
2.81% in multimodal approaches, respectively. 5. Conclusions and Future Work
To evaluate the effect of fusion methods on emotion recog-
nition, we contrast the predictive performance attributes of sin- In this paper, we integrated handcrafted MTAR and SSR for
gle and multimodal emotion recognition, delineated in Table 2. MER tasks. Our findings revealed superior performance when
incorporating handcrafted MTAR compared to merely integrat-
Table 2: Comparison of MER performance obtained by the pro- ing individual modality SSR, demonstrating the effectiveness of
posed MF and baseline fusion methods. MTAR in enhancing SSR in emotion recognition. Additionally,
we introduced a novel MF method for integrating modality-
Modality Model UAR (%) F1 (%) specific and modality-invariant emotional information, consis-
WavLM + MTAR (Concat) 71.85 71.12 tently achieving higher accuracy in MER. For future research,
Single modal WavLM + MTAR [8] 72.61 71.92
WavLM + MTAR (MF) 73.47 73.01
we advocate for further exploration of the intricate relationship
between music and speech to devise innovative fusion methods.
RoBERTa + MTAR (Concat) 72.83 72.24
RoBERTa + MTAR [8] 73.50 73.35
Multimodal
RoBERTa + MTAR (MF) 75.11 74.79 6. Acknowledgements
RoBERTa + WavLM + MTAR (Concat) 77.78 77.29
RoBERTa + WavLM + MTAR [8] 78.38 77.83 This work was financially supported by JST SPRING, Grant
RoBERTa + WavLM + MTAR (MF) 79.89 79.35 Number JPMJSP2125, and in part by JST CREST Grant Num-
ber JPMJCR19A3, Japan, and JSPS KAKENHI Grant Number
Evidently, the proposed fusion method demonstrates an en- 21H05054.

3727
7. References [19] L. Pepino, P. Riera, L. Ferrer, and A. Gravano, “Fusion ap-
proaches for emotion recognition from speech using acoustic
[1] P. Singh, R. Srivastava, K. Rana, and V. Kumar, “A multimodal and text-based features,” in ICASSP 2020-2020 IEEE Interna-
hierarchical approach to speech emotion recognition from audio tional Conference on Acoustics, Speech and Signal Processing
and text,” Knowledge-Based Systems, vol. 229, p. 107316, 2021. (ICASSP). IEEE, 2020, pp. 6484–6488.
[2] T. Deschamps-Berger, L. Lamel, and L. Devillers, “End-to-end [20] T. Mittal, U. Bhattacharya, R. Chandra, A. Bera, and D. Manocha,
speech emotion recognition: challenges of real-life emergency “M3er: Multiplicative multimodal emotion recognition using fa-
call centers data recordings,” in 2021 9th International Confer- cial, textual, and speech cues,” in Proceedings of the AAAI con-
ence on Affective Computing and Intelligent Interaction (ACII). ference on artificial intelligence, vol. 34, no. 02, 2020, pp. 1359–
IEEE, 2021, pp. 1–8. 1367.
[3] S. Zepf, J. Hernandez, A. Schmitt, W. Minker, and R. W. Picard, [21] J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evalu-
“Driver emotion recognition for intelligent vehicles: A survey,” ation of gated recurrent neural networks on sequence modeling,”
ACM Computing Surveys (CSUR), vol. 53, no. 3, pp. 1–30, 2020. arXiv preprint arXiv:1412.3555, 2014.
[4] M. Akagi, X. Han, R. Elbarougy, Y. Hamada, and J. Li, “To- [22] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li,
ward affective speech-to-speech translation: Strategy for emo- N. Kanda, T. Yoshioka, X. Xiao et al., “Wavlm: Large-scale self-
tional speech recognition and synthesis in multiple languages,” supervised pre-training for full stack speech processing,” IEEE
in Signal and Information Processing Association Annual Summit Journal of Selected Topics in Signal Processing, vol. 16, no. 6,
and Conference (APSIPA), 2014 Asia-Pacific. IEEE, 2014, pp. pp. 1505–1518, 2022.
1–10.
[23] S. Dang, T. Matsumoto, Y. Takeuchi, and H. Kudo, “Using Semi-
[5] J. Tian, D. Hu, X. Shi, J. He, X. Li, Y. Gao, T. Toda, X. Xu, and supervised Learning for Monaural Time-domain Speech Separa-
X. Hu, “Semi-supervised multimodal emotion recognition with tion with a Self-supervised Learning-based SI-SNR Estimator,” in
consensus decision-making and label correction,” in Proceedings Proc. INTERSPEECH 2023, 2023, pp. 3759–3763.
of the 1st International Workshop on Multimodal and Responsible [24] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy,
Affective Computing, 2023, pp. 67–73. M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A
[6] L.-W. Chen and A. Rudnicky, “Exploring wav2vec 2.0 fine tuning robustly optimized bert pretraining approach,” arXiv preprint
for improved speech emotion recognition,” in ICASSP 2023-2023 arXiv:1907.11692, 2019.
IEEE International Conference on Acoustics, Speech and Signal [25] J. He, X. Shi, X. Li, and T. Toda, “Mf-aed-aec: Speech emotion
Processing (ICASSP). IEEE, 2023, pp. 1–5. recognition by leveraging multimodal fusion, asr error detection,
[7] N. Naderi and B. Nasersharif, “Cross corpus speech emotion and asr error correction,” in ICASSP 2024-2024 IEEE Interna-
recognition using transfer learning and attention-based fusion of tional Conference on Acoustics, Speech and Signal Processing
wav2vec2 and prosody features,” Knowledge-Based Systems, vol. (ICASSP). IEEE, 2024, pp. 11 066–11 070.
277, p. 110814, 2023. [26] K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statis-
[8] H. Zou, Y. Si, C. Chen, D. Rajan, and E. S. Chng, “Speech emo- tics pooling for deep speaker embedding,” arXiv preprint
tion recognition with co-attention based multi-level acoustic infor- arXiv:1803.10963, 2018.
mation,” in ICASSP 2022-2022 IEEE International Conference on [27] C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower,
Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap:
pp. 7367–7371. Interactive emotional dyadic motion capture database,” Language
[9] S. Padi, S. O. Sadjadi, D. Manocha, and R. D. Sriram, resources and evaluation, vol. 42, pp. 335–359, 2008.
“Multimodal emotion recognition using transfer learning from [28] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec
speaker recognition and bert-based models,” arXiv preprint 2.0: A framework for self-supervised learning of speech repre-
arXiv:2202.08974, 2022. sentations,” Advances in neural information processing systems,
[10] K. R. Scherer, “Vocal affect expression: a review and a model vol. 33, pp. 12 449–12 460, 2020.
for future research.” Psychological bulletin, vol. 99, no. 2, p. 143, [29] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhut-
1986. dinov, and A. Mohamed, “Hubert: Self-supervised speech rep-
resentation learning by masked prediction of hidden units,”
[11] A. Gabrielsson and P. N. Juslin, Emotional expression in music.
IEEE/ACM Transactions on Audio, Speech, and Language Pro-
Oxford University Press, 2003.
cessing, vol. 29, pp. 3451–3460, 2021.
[12] A. D. Patel, “Language, music, syntax and the brain,” Nature neu- [30] X. Qin, Z. Wu, T. Zhang, Y. Li, J. Luan, B. Wang, L. Wang, and
roscience, vol. 6, no. 7, pp. 674–681, 2003. J. Cui, “Bert-erc: Fine-tuning bert is enough for emotion recogni-
[13] L. Jäncke, “The relationship between music and language,” p. tion in conversation,” in Proceedings of the AAAI Conference on
123, 2012. Artificial Intelligence, vol. 37, no. 11, 2023, pp. 13 492–13 500.
[14] T. Fujisawa, K. Takami, and N. D. Cook, “On the role of pitch [31] P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-
intervals in the perception of emotional speech,” in ISCA & IEEE enhanced bert with disentangled attention,” arXiv preprint
Workshop on Spontaneous Speech Processing and Recognition, arXiv:2006.03654, 2020.
2003. [32] Y. Gao, H. Shi, C. Chu, and T. Kawahara, “Enhancing two-
[15] B. Yang and M. Lugger, “Emotion recognition from speech sig- stage finetuning for speech emotion recognition using adapters,”
nals using new harmony features,” Signal processing, vol. 90, in ICASSP 2024-2024 IEEE International Conference on Acous-
no. 5, pp. 1415–1423, 2010. tics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp.
11 316–11 320.
[16] X. Li, X. Shi, D. Hu, Y. Li, Q. Zhang, Z. Wang, M. Unoki,
and M. Akagi, “Music theory-inspired acoustic representation for [33] Y. Gao, C. Chu, and T. Kawahara, “Two-stage finetuning of
speech emotion recognition,” IEEE/ACM Transactions on Audio, wav2vec 2.0 for speech emotion recognition with asr and gender
Speech, and Language Processing, 2023. pretraining,” in Proc. Interspeech, 2023.
[17] P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli, [34] X. Shi, S. Li, and J. Dang, “Dimensional emotion prediction based
“Multimodal fusion for multimedia analysis: a survey,” Multime- on interactive context in conversation.” in INTERSPEECH, 2020,
dia systems, vol. 16, pp. 345–379, 2010. pp. 4193–4197.
[35] X. Shi, X. Li, and T. Toda, “Emotion awareness in multi-utterance
[18] S. Yoon, S. Byun, and K. Jung, “Multimodal speech emotion
turn for improving emotion prediction in multi-speaker conversa-
recognition using audio and text,” in 2018 IEEE Spoken Language
tion,” in Proc. Interspeech, vol. 2023, 2023, pp. 765–769.
Technology Workshop (SLT). IEEE, 2018, pp. 112–118.

3728

You might also like