0% found this document useful (0 votes)
13 views12 pages

Multi-Label Emotion Detection Framework

The paper presents a Multi-label Emotion Detection Architecture (MEDA) designed to detect multiple emotions in textual expressions by utilizing a Multi-Channel Emotion-Specified Feature Extractor (MC-ESFE) and an Emotion Correlation Learner (ECorL). MEDA enhances emotion detection by separately encoding features for each emotion and learning their correlations, leading to improved performance on emotional datasets. Experimental results demonstrate that MEDA outperforms existing state-of-the-art methods in multi-label emotion detection tasks.

Uploaded by

asrrriyasriyas8
License
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views12 pages

Multi-Label Emotion Detection Framework

The paper presents a Multi-label Emotion Detection Architecture (MEDA) designed to detect multiple emotions in textual expressions by utilizing a Multi-Channel Emotion-Specified Feature Extractor (MC-ESFE) and an Emotion Correlation Learner (ECorL). MEDA enhances emotion detection by separately encoding features for each emotion and learning their correlations, leading to improved performance on emotional datasets. Experimental results demonstrate that MEDA outperforms existing state-of-the-art methods in multi-label emotion detection tasks.

Uploaded by

asrrriyasriyas8
License
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO.

1, JANUARY-MARCH 2023 475

Multi-Label Emotion Detection via


Emotion-Specified Feature Extraction
and Emotion Correlation Learning
Jiawen Deng and Fuji Ren , Senior Member, IEEE

Abstract—Textual emotion detection is an attractive task while previous studies mainly focused on polarity or single-emotion
classification. However, human expressions are complex, and multiple emotions often co-occur with non-negligible emotion
correlations. In this paper, a Multi-label Emotion Detection Architecture (MEDA) is proposed to detect all associated emotions
expressed in a given piece of text. MEDA is mainly composed of two modules: Multi-Channel Emotion-Specified Feature Extractor (MC-
ESFE) and Emotion Correlation Learner (ECorL). MEDA captures underlying emotion-specified features through MC-ESFE module,
which is composed of multiple channel-wise ESFE networks. Each channel in MC-ESFE is devoted to the feature extraction of a
specified emotion from sentence-level to context-level through a hierarchical structure. With underlying features, emotion correlation
learning is implemented through an emotion sequence predictor in ECorL. Furthermore, we define a new loss function: multi-label focal
loss. With this loss function, the model can focus more on misclassified positive-negative emotion pairs and improve the overall
performance by balancing the prediction of positive and negative emotions. The evaluation of proposed MEDA architecture is carried
out on emotional corpus: RenCECps and NLPCC2018 datasets. The experimental results indicate that the proposed method can
achieve better performance than state-of-the-art methods in this task.

Index Terms—Multi-label, emotion detection, emotion correlation, multi-label focal loss

1 INTRODUCTION label emotion detection has gained burgeoning attention


because of its vast potential applications.
the rapid development of social media platforms,
A S
such as microblogs and twitter, it is convenient for
users to share their attitudes about any topic. Essentially,
Multi-label emotion detection task aims to recognize all
possible emotions in a piece of textual expression [10]. In
conventional emotion detection networks, textual informa-
understanding the latent emotions expressed in such user-
tion is often encoded together into a representation vector
generated content has gained much attention because of its
and then directly fed into the classifier[11], [12]. However, in
vast potential applications [1], [2], such as emotional chat-
a textual expression with multiple emotions, there may be
bots [3], emotional text-to-speech synthesizer [4], and some emotions with relatively weaker intensity. If informa-
patient emotion monitoring [5]. tion of each emotion is mixed and encoded together into a
As a fundamental task in sentiment analysis, emotion shared vector, the weaker emotions with subtle features
detection has been deeply studied in the literature. Not only could be covered by stronger emotions and be challenging to
the basic emotional polarity classification [6], but emotion recognize. To accurately recognize the emotions expressed,
detection has also delved into more granular analysis [7], the quality of underlying emotional feature representation
[8], such as love, hate, angry, and surprise. While many has an important influence on the final prediction.
such kinds of researches have been implemented, most of In most previous researches, multi-label emotion detec-
them are conducted in the single-emotion environment [9]. tion task is often narrowed down into multiple binary classi-
They are based on the assumption that certain textual data fications [13], in which each emotion is detected respectively
is associated with only one emotion. However, in real-world without considering their correlations. However, emotion
conditions, people often hold multiple complex emotions correlation information provides non-ignorable features and
simultaneously, and a textual expression is often associated is useful for improving the performance of emotion detec-
with multiple emotions simultaneously. Therefore, multi- tion. The definition of emotion correlation can be illustrated
based on Plutchik’s work. In an emotional expression, emo-
tion correlation mainly refers to positive or negative emo-
 The authors are with the Graduate School of Advanced Technology and
Science, Tokushima University, Tokushima 2-1, Minami Josanjima-cho,
tional correlation. Positively correlated emotions are similar
Tokushima 770-8506, Japan. E-mail: c501847002@[Link], to each other and often appearing together but with different
ren@[Link]. intensities. Such as the emotion pair ‘Joy’ and ‘Love’ tend to
Manuscript received 21 May 2020; revised 6 October 2020; accepted 23 October appear simultaneously. Negatively associated emotions are
2020. Date of publication 27 October 2020; date of current version 28 February often opposite to each other and rarely appear together, such
2023. as ‘Love’ and ‘Sorrow’. Emotion correlation can be utilized
(Corresponding author: Jiawen Deng.)
Recommended for acceptance by C. Strapparava. to facilitate more in-depth emotion analysis in multi-label
Digital Object Identifier no. 10.1109/TAFFC.2020.3034215 emotion recognition task.
1949-3045 © 2020 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.
See ht_tps://[Link]/publications/rights/[Link] for more information.
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
476 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

In this paper, a Multi-label Emotion Detection Architec- The rest of this paper is organized as follows: Section 2
ture (MEDA) is proposed to address the above challenges. presents a brief overview of related work on multi-label emo-
MEDA is mainly composed of two modules: Multi-Channel tion detection task. Details of the proposed MEDA are given
Emotion-Specified Feature Extractor (MC-ESFE) and Emo- in Section 3. The experimental setting and details are pre-
tion Correlation Learner (ECorL). MC-ESFE consists of mul- sented in Section 4. The performance of MEDA is discussed
tiple channels by which the features of each emotion are in Section 5. Finally, conclusions are drawn in Section 6.
separately encoded. Each channel is devoted to the underly-
ing feature representation of a specified emotion from both
sentence-level and context-level. Furthermore, an external 2 RELATED WORK
emotion lexicon is introduced as prior knowledge to inte- As an important task in natural language processing, textual
grate more detailed emotional information. ECorL module sentiment analysis is an emerging research field. Tradition-
is devoted to learning emotion correlation based on ally, sentiment analysis is applied to predict a polarity label
extracted emotion-specified features from MC-ESFE. In or star level of product reviews [14], [15], or film reviews
ECorL, multi-label emotion detection task is transformed as [16]. With the delving of related researches, emotion classifi-
an emotion sequence prediction task. Bidirectional GRU cation is becoming more granular [17]. There are many dif-
network is taken as the emotion sequence predictor, and the ferent taxonomies [18], such as Paul Ekman’s six basic
emotions are sequentially predicted in a fixed path. In the emotions [19], and Fuji Ren’s eight basic emotions [20].
hidden state of each step, the emotion correlations of cur- As one of the most obvious clues of sentiment analysis,
rent emotion are learned by information interaction with emotional lexical resources directly encode sentimental
the context of other emotions flowed from both forward knowledge [1], [21], and are widely used, such as WordNet-
and backward directions. Considering that the proposed affect [22], NRC emotion lexicon [23], and Hownet [24].
MEDA network extracts emotional information from sen- These emotional lexicons are the basis for early hand-
tence level, context level, and emotion correlation level, an crafted features based emotional analysis [21], [25]. Given
ensemble model called MEDA-FS is proposed to integrate the recent success of deep learning models, various neural
emotional information from different levels. MEDA-FS can network models have been proposed and have achieved
realize the maximization of information retention and avoid highly competitive performance in sentiment analysis.
information loss during bottom-up learning. During the LSTM networks [26] have shown its superior performance
training, positive-negative emotion correlation is incorpo- in context information encoding [27], [28], and CNN net-
rated into the proposed multi-label focal loss function. By works [29], [30] are often utilized to extract local informa-
introducing a weighting factor, our loss will focus more on tion within a sentence. Multi-task ensemble network
misclassified emotion pairs and balance the prediction performs well in emotional information integration and can
between positive and negative emotions. alleviate the impact of insufficient corpus [31], [32]. To
Compared with existing multi-label emotion detection achieve robust emotion feature representation, Emo2Vec is
methods, the proposed MEDA architecture extracts both emo- trained in [33] to encode emotional semantics by multi-task
tion-specified features and emotion correlations. The perfor- learning six different emotion-related tasks. Different
mance of the proposed MEDA is verified in Chinese emotional modalities, such as text, emoji, and images, can also be com-
corpus: RenCECps. Experimental results show that the pro- bined [34] to express emotion and complement each other
posed architecture achieves state-of-the-art performance on for emotion classification. To further model the effects of
RenCECps and demonstrates the effectiveness of MEDA. sentimental relations, a modified GRNN is proposed [6] to
The major contributions of our paper can be summarized encode the sentiment polarity and sentiment modifier con-
as follows. text separately.
In most previous studies, the complexity of emotion
1. MEDA architecture composed of MC-ESFE and detection task is often narrowed down by focusing on single
ECorL modules is proposed for the textual multi- emotion classification. However, human emotion is com-
label emotion detection task. MC-ESFE can encode plex in reality, and the textual expression often contains
emotion-specified features in the corresponding multiple emotions simultaneously. To address this problem,
channel respectively, which strengthens the underly- the multi-label emotion detection task can be viewed as a
ing feature representation of each emotion. ECorL is special multi-label classification [10], [35], in which emo-
proposed to learn emotion correlations by transform- tions are the multiple labels. Attention-based multi-label
ing multi-label emotion detection task into emotion sentence classifier is proposed in [36] to imitate how
sequence prediction task. humans comprehend and classify emotions. A dual atten-
2. MEDA-FS is proposed to fuse the information at sen- tion-based transfer learning model is proposed in [37] to
tence-level, context-level, and emotion correlation extract both general sentiment words and other emotion-
level, which can realize the maximization of informa- specific words. Linguistic characteristics are explored in
tion retention during bottom-up learning. [12] to reduce the limitation of lexicon coverage and size.
3. Multi-label focal loss function considering emotion However, most of the above models do not take multi-
correlation information is proposed for multi-label label correlations into account, and they assume that the
learning. This loss function contributes to model multiple labels are independent of each other. Some stud-
training by focusing on misclassified emotion pairs ies try to explore the label correlations. To reduce the effect
and balancing the prediction of positive and nega- of irrelevant labels, prior knowledge of co-occur label rela-
tive emotions. tionships are incorporated [38] as a constraint for emotion
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 477

Fig. 1. (a) The illustration of the proposed MEDA: Multi-label Emotion Detection Architecture. MEDA mainly composes two modules: Multi-channel
emotion-specified feature extractor (MC-ESFE) and emotion correlation learner (ECorL). (b) The illustration of lth ESFE channel of emotion el in
MC-ESFE module. (c) The illustration of sentence-level encoder of lth ESFE channel in MC-ESFE module.

prediction and ranking. Some approaches attempt to Multi-label emotion detection task aims to detect all pos-
implicitly estimate label correlations by the modification of sible emotions from the pre-defined emotional label set: E ¼
loss function. The label-correlation sensitive loss function ½e1 ; e2 ; . . . eL . Considering the important influence of contex-
is first proposed in [39] with the BP-MLL algorithm. Joint tual information on this task, the previous k sentences
binary cross-entropy (JBCE) loss is proposed by Huihui He occurred before current sentence s are taken as the context
[14] in his joint binary neural network (JBNN) to capture sentences: scxt ¼ ½sk ; . . . s2 ; s1 . Given a sentence s ¼
label relations. Multi-label classification can also be trans- ½w1 ; w2... wn  and its context scxt , our proposed multi-label
formed into a sequence generation problem [40], [41] to emotion recognition model MEDA is trained to output the
capture label correlations. To reduce the computational predicted probability distribution PML ¼ ½p1 ; p2 ; . . . pL  of
complexity, partial label dependence can also contribute to each emotion, denoted as:
this task, which is demonstrated in [42]. Deep canonical
correlation analysis (DCCA) performs well in feature- PML ¼ fMEDA ðs; scxt Þ: (1)
aware label embedding and label-correlation aware predic-
tion [43], [44]. A semisupervised multi-label method is pro- MC-ESFE module is composed of L parallel channel-
posed in [45] while label correlations are incorporated by wise ESFE. In each channel, ESFE extracts emotion-specified
modifying the loss function. Multi-Label classification and features from sentence-level to context-level through a hier-
label correlations learning can also be realized in a joint archical structure. The output of each channel is combined
cxt
learning framework [46], [47]. into an emotion-specified feature matrix: XES ¼ ½xcxt
ES1 ;
In contrast with most current methods, we focus on emo- xcxt
ES2 ; . . . ; x cxt
ESL . In ECorL module, emotion correlations are
cxt
tion-specified feature extraction and emotion correlation. further learned from XES and multi-label emotions are pre-
There are mainly two fundamental differences: dicted. Specifically, MEDA architecture is very flexible, and
the algorithm applied in each module can be replaced by
1. The information of each emotion is encoded sepa-
other state-of-the-art algorithms.
rately, which can concentrate more on underlying
emotion-specified feature extraction. This implemen-
3.1 MC_ESFE: Multi-Channel Emotion Specified
tation contributes to further emotion correlation
Feature Extractor
learning in emotion sequence predictor.
In this paper, a Multi-Channel Emotion-Specified Feature
2. The proposed multi-label focal loss function consid-
Extractor (MC-ESFE) is proposed for underlying fundamen-
ers emotion correlation information. It can pay more
tal feature extraction. MC-ESFE is composed of L channel-
attention to misclassified emotion pairs and balance
wise ESFE, and L is equal to the number of emotions. Each
the prediction of positive and negative emotions.
channel focuses on the feature extraction of a specified emo-
tion, and each emotion’s information is separately encoded
3 PROPOSED METHOD in each channel. In this way, more details of each emotion
To comprehensively obtain emotional information of texts, could be summarized, and the features of weak emotions
Multi-label Emotion Detection Architecture (MEDA) is pro- are prevented from being covered by strong emotions to
posed in this paper. It mainly composes two modules: some extent.
Multi-Channel Emotion-Specified Feature Extractor (MC- Fig. 1 shows the hierarchical structure of lth ESFE-chan-
ESFE) and Emotion Correlation Learner (ECorL). The nel corresponding to emotion el , l 2 ½1; L. Each channel
framework of MEDA is shown as Fig. 1. contains a sentence-level encoder and a context-level
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
478 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

encoder, which focus on feature extraction of emotion el on el . More attention weight will be assigned to words related
both sentence-level and context-level. to emotion el in the current lth ESFE channel. Attention
weight ai and weighted emotional feature vector xew are
3.1.1 Sentence Level Encoder defined as follows:
In lth ESFE channel, given a sentence s, the sentence-level   
l ei ¼ W2T s W1T  hi þ b1 þ b2 (3Þ
encoder fSEn projects input sentence s to emotion-specified
s
feature xESl : expðei Þ
ai ¼ P n (4Þ
k¼1 expðek Þ
xsESl ¼ fSEn
l
ðsÞ; l 2 ½1; L: (2) xew ¼ ½a1 h1 : a2 h2 : . . . ai hi : . . .; (5Þ

in which s indicates the sigmoid activation function,


l
In sentence-level encoder fSEn , as shown in Fig. 1c, two w1 ; b1 ; w2 ; b2 indicate the model parameters, and [:] indi-
parallel architectures with different embedding methods cates the concatenation operation.
are employed to generate: (1) emotional feature representa- Finally, emotional feature vector xew and general embed-
tion xew , (2) general sentence representation xBERT . They are ding xBERT is integrated, and emotion-specified sentence-
integrated into emotion-specified sentence-level representa- level representation xsESl is generated as follows:
tion xsESl for further context-level learning.
General Sentence Representation. Inspired by the pre- xsESl ¼ tanhðWe  xew þ WB  xBERT þ bs Þ; (6)
trained language model learning approach and transfer
learning techniques, pre-trained Chinese BERT model [48] in which We , WB and bs indicate the model parameters.
is applied to yield general sentence representation in this
paper. BERT stands for Bidirectional Encoder Representa- 3.1.2 Context Level Encoder
tions from Transformers. Chinese BERT is designed to pre- Context level encoding is channel-wise implemented as
train deep bidirectional representations from unlabeled well. As shown in Fig. 1b, in lth ESFE channel correspond-
Chinese text by jointly conditioning on both left and right ing to emotion el , given a sentence s and its context scxt ¼
context in all layers. It remedies the limitation of insuffi- l
½sk ; . . . s2 ; s1 , context-level encoder fCEn gives contex-
cient training corpora and contributes to syntactic and cxt
tual emotional feature xESl . GRU network is utilized to learn
semantic sentence representation. Given a sentence s, the contextual information from previous k sentences, and the
general sentence representation xBERT is generated from output of final step is captured as the context-level repre-
Chinese BERT. sentation, which is denoted as:
Emotional Feature Representation. Arguably, it is accepted
 sk s1 
that general sentence representation generated by pre-trained xcxt l s
ESl ¼ fCEn xESl ; . . . ; xESl ; xESl ; l 2 ½1; L
language model does not contain specific emotional features,  sk s1  (7)
¼ fGRU xESl ; . . . ; xESl ; xsESl
as no emotion-related knowledge has been included in the
training process. To generate emotional sentence representa-
s l
tion, emotional features are further extracted based on an xESl
i
¼ fSEn ðsi Þ; i 2 ½1; k: (8)
external n-dimensional emotion lexicon.
With the input sentence s ¼ ½w1 ; w2 ; . . ., emotional Contextual emotion-specified features xcxtESl learned from
words we ¼ ½we1 ; we2 ; . . . occurred in s are first extracted by each channel in MC-ESFE is output and combined as the
cxt
matching the emotion lexicon. The embedding of emotional emotional feature matrix: XES ¼ ½xcxt cxt cxt
ES1 ; xES2 ; . . . ; xESL .
cxt
words consists of two parts. The first is general word XES is flowed into ECorL module for further emotion corre-
embedding, which is realized by mapping the pre-trained lation learning.
Word2vec word embedding matrix. Each word is embed-
ded as vw2v 2 R1D , in which D is the embedding dimen- 3.2 ECorL: Emotion Correlation Learner
sion. The second is emotional word embedding, which is Emotion correlations are indispensable in multi-label emo-
realized based on n-dimensional emotion lexicon. Each tion detection task. In this paper, ECorL (Emotion Correla-
word is embedded as vemo 2 R1n , in which n is the number tion Learner) is proposed to give emotion prediction based
of emotions annotated in emotion lexicon and the value on emotion correlation learning.
means the intensity of corresponding emotion. Finally, emo- MC-ESFE module project inputs into a sequence of con-
tional word embedding is represented as E ¼ ½ve1 ; ve2 ; . . ., in tinuous emotional representations XES cxt
¼ ½xcxt cxt
ES1 ; xES2 ; . . ..
which vei 2 R1ðDþnÞ . cxt
ECorL module takes XES as input. In ECorL module, multi-
Considered the polysemy of emotional words in differ- label emotion detection task is transformed as emotion
ent contexts, BiGRU (Bidirectional Gated Recurrent Neural sequence prediction task, and emotions are predicted in a
Networks) [49] and attention network [50] are utilized to fixed path. Refer to the previous work [40], the order of
make the network pay more attention to significant emo- emotion sequence is set according to its occurred cumula-
tional words. Take emotional embedding E as input, the tive number in the corpus. BiGRU is taken as the emotional
output
* (
of the hidden
*
state
(
of BiGRU in each step is hi ¼ sequence predictor. The operation is formulated as follows:
½h : h, in which h and h are the output of hidden states  
i i i i
from forward and backward directions, respectively. The He ¼ fBiGRU xcxt cxt cxt
ES1 ; xES2 ; . . . ; xESL (9Þ
attention mechanism considers the contributions of differ-
ent emotional words to the prediction of specified-emotion PML ¼ tanhðWECor  He þ bECor Þ (10Þ
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 479

in which He ¼ ½he1 ; . . . hel ; . . . heL ; are the hidden states of learning, MEDA-FS is proposed to fuse the information
each step, WECor and bECor are the learned weight and from different levels. MEDA-FS consists of three sub-mod-
biases, and PML is the predicted probability of each emotion. els: S-MC-ESFE, C-MC-ESFE, and MEDA, which give emo-
In lth step of BiGRU, the learning of hidden state hel can be tion predictions on sentence-level, context-level, and
viewed as the feature extraction of a specified emotion el . emotion correlation level, respectively.
Invalid information of current input xcxt ESl can be filtered S-MC-ESFE, gives sentence-level predictions PES s
¼
s s
because of the gating mechanism. With the bidirectional ½pES1 ; . . . ; pESL . It is obtained during the pre-training step of
network, emotional feature hel is learned based on the infor- sentence-level encoder in MC-ESFE, which is detailed in
s
mation of other emotions flowed from both forward and Section 3.3. PES represents the prediction based on the
backward hidden state. In this way, emotional information underlying information, without considering the emotion
interaction is realized. The hidden states of BiGRU are out- correlations and contextual information.
cxt
put and fed into emotion interaction layer. This layer is a C-MC-ESFE, gives context-level predictions PES based on
fully-connected layer and aimed to realize further emotional sentence-level predictions of current sentence s, denoted as
s
information interaction. In this way, the final emotion pre- PES , and sentence-level predictions of its context scxt , denoted
s s
diction is obtained: PML ¼ ½p1 ; p2 ; ::pL . as ½PESk ; . . . ; PES1 . GRU network is utilized to learn contex-
tual information and its final output is taken as the prediction:
3.3 Network Pre-Training in MC-ESFE  s 
cxt s s
Each channel in MC-ESFE is dedicated to obtaining corre- PES ¼ fGRU PESk ; . . . PES1 ; PES : (14)
sponding emotional information, which belongs to the
MEDA, gives prediction PML ¼ ½p1 ; p2 ; ::pL  by consider-
underlying feature extraction in the MEDA framework. The
ing both contextual information and emotion correlation.
quality of feature representation has a direct impact on the
MEDA-FS, gives final predictions by comprehensively
performance of upper-level emotion predictions. To improve
fuse the information from above three level, denoted as:
the underlying feature representation, network-based trans-
fer learning is employed to pre-train the sentence-level s
P ¼ ws  PES cxt
þ wc  PES þ wML  PML : (15)
encoder in each channel. During transfer learning, a predic-
tion layer is added to emotional sentence representation xsESl in which ws , wc and wML are the weight parameters of each
for single-emotion prediction: level’s information.
 
psESl ¼ s wl  xsESl þ bl ; (11) 3.5 Definition of Multi-Label Focal Loss
Multi-label (ML) loss function [52] is one of the most com-
in which xsESl is the sentence-level representation of input monly used loss functions in multi-label learning. Instead of
sentence s, and psESl indicates the predicted probability of concentrating on individual label discrimination like tradi-
emotion el expressed in sentence s. The first-step pre-train- tional cross-entropy loss function, ML-loss focused on con-
ing is implemented on positive-negative annotated emo- sidering the correlations between the different labels.
tional datasets. The second-step is fine-tuning. During fine- Inspired by [51], we rewrite ML-loss and called it multi-
tuning, the multi-emotion annotation fs; ½y1 ; y2 ; . . . yL g of label focal loss. Multi-label focal loss not only considers
each sentence s in dataset D is transformed into multiple emotion correlation but also focus more on misclassified
single-emotion annotations: fs; y1 g, fs; y2 g. . .fs; yL g. In this emotion pairs. Besides, it introduces a harmonic parameter
way, we reconstructed multiple binary-dataset: D ^¼
to reduce the influence of the imbalance prediction of posi-
^ ^ ^ ^
fD1 ; D2 ; . . . DL g. For each binary-dataset Dl , l 2 ½1; L, sen- tive and negative emotions. The definition of multi-label
tence s 2 D ^ l is fed into lth ESFE
^ l is annotated as fs; yl g. D
focal loss is defined as follows:
channel to fine-tune the sentence-level parameters. During
pre-training, binary focal loss [51] is utilized: XN 1 X  
EMLFL ¼   aikl  expð pik  pil (16Þ

jYi j Yi
i¼1  ð k;lÞ2Y i Y i
EFL ¼  at ð1  pt Þr logðpt Þ (12Þ i
 
i r
 i r
akl ¼ w  1  pk þ ð1  wÞ  pl ; (17Þ

psESl if yl ¼ 1
pt ¼ ; (13Þ in which Yi denotes the set of positive emotions expressed
1  psESl otherwise
in ith instance si , and Yi denotes the negative emotion set. pik
in which r is a modulating factor, and it aimed to reduce the and pil are the predicted probability of positive emotion ek
relative loss of well-classified examples. a 2 ½0; 1; is a and negative emotion el respectively. Therefore, the training
weighting factor to address the problem of class imbalance. with above loss function is equivalent to maximizing the dif-
at ¼ a for positive label and at ¼ 1  a for negative label. ference of negatively related emotion pair of (pik  pil ). This
leads the system to output a higher probability for positive
3.4 MEDA-FS: Multi-Level Information Fusion emotion while a lower probability for negative emotion. In
The proposed MEDA architectural learns emotional infor- this way, the emotion correlation of negatively related emo-
mation from sentence-level to context-level, from single- tion pairs can be taken into consideration.
emotion level in MC-ESFE to multi-emotion level in ECorL aikl is a weighting factor and mainly affected by two
module. Each layer in MEDA network learns different lev- parameters: w 2 ð0; 1Þ is a harmonic factor aimed to balance
els of information. To realize the maximization of informa- the prediction between positive and negative emotions, and
tion retention and avoid information loss during bottom-up r > 0 is a modulating factor aimed to make the loss put
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
480 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

TABLE 1 In terms of the embedding of emotional words, it mainly


Cumulative Number of Each Emotion in Ren-CECps and consists of two parts. The first part is the general word
NLPCC2018 Datasets embedding. It is initialized by 300-dimensional Word2vec
Ren-CECps NLPCC2018 word embedding, which is trained on Chinese microblog
data [54]. The second part is emotional embedding by map-
Love 11909 Hate 3533 Happiness 2534 ping an external n-dimensional emotion lexicon. Existing
Anxiety 10099 Anger 2236 Sadness 1502
Sorrow 8184 Surprise 1121 Surprise 811 emotion lexicons are very rare due to the subjective and
Joy 6223 Neutral 2488 Anger 765 inconsistent annotation. In our experiments, a dimensional
Expect 4633 - - Fear 770 emotion lexicon is manually built based on word-level
annotation on Ren-CECps. In our lexicon, each emotion
word is annotated as an 8-dimensional vector v. Each
more focus on hard and misclassified examples during dimension corresponding an emotion in [‘Love’, ‘Anxiety’,
training. Significantly, while r ¼ 0, the proposed multi-label ‘Sorrow’, ‘Joy’, ‘Expect’, ‘Hate’, ‘Anger’, ‘Surprise’], and the
focal loss is equivalent to the multi-label loss function. value represents the emotion intensity. For example, the
For well-classified positive-negative emotion pairs emotional word ’不幸’ (‘Unfortunately’ in English) is repre-
(ek ; el ), predicted probability pik tends to 1 while pil tends to sented as [0., 0.23, 0.62, 0., 0., 0., 0., 0.], which means that
0. In this case, the difference of (pik  pil ) tends to the maxi- this word expresses stronger emotion of ‘Sorrow’ and
mum, which means the minimum of expððpik  pil Þ, and weaker emotion of ‘Anxiety’, and the intensities are 0.62,
the weighting factor aikl tends to 0. Thereby the loss of well- and 0.23 respectively. Specially, an extra token named
classified positive-negative emotion pairs is minimized. ‘[EMO_PAD]’ is added to emotion lexicon, and its embed-
Conversely, for hard-classified emotion pairs, the difference ding vector is initialized by zeros. This token will be treated
of (pik  pil ) tends to the minimum, which could be caused as emotional word if the current sentence does not contain
by pik tending to 0 or pil tending to 1. In response to the above any other emotional words.
two cases, w  ð1  pik Þr and ð1  wÞ  ðpil Þr are introduced to For RenCECps, because of the non-coexistence of
give more focus on misclassified pik and pil respectively. ‘Neutral’ label with other emotion labels, the final predic-
tion is subject to the condition: only if the prediction of
4 EXPERIMENTS SETUP ‘Neutral’ label obtained the highest probability among all
labels, the sentence is predicted as ‘Neutral’. Otherwise, it is
4.1 Datasets
predicted as emotions contained.
We employ two different datasets to evaluate the proposed In terms of the number of context sentences k, we set k ¼
architecture, which are listed below: 3, which means that the previous 3 sentences are taken as
Ren-CECps Dataset is an annotated emotional corpus with
the contextual information. Particularly, replication pad-
Chinese blog texts [21]. The corpus is annotated in the docu-
ding with the last sentence is utilized while the number of
ment, paragraph, and sentence level. Each level is annotated
contextual sentences is less than 3.
with eight emotional categories (‘Joy’, ‘Hate’, ‘Love’,
We set the dropout as 0.2 in EcorL module to avoid over-
‘Sorrow’, ‘Anxiety’, ‘Surprise’, ‘Anger’, and ‘Expect’) and cor-
fitting. The hidden size of BiGRU in sentence-level encoder
responding discrete emotional intensity value from 0.0 to 1.0. is 64 in each direction. For the binary focal loss utilized dur-
In our experiments, those emotions with an intensity greater ing the network pre-training, modulating factor r is set to 2,
than 0.0 are labeled as 1, otherwise 0. ‘Neutral’ is regarded as and the weighting factor a is set to 0.75. In multi-label focal
the 9th emotion label in case the sentence holds no emotion. loss, we set the modulating factor r to 2 and harmonic factor
After pre-processing, there is a total of 27091 sentences in w to 0:4. Adam optimization method is applied to train the
training data and 7681 sentences in testing data. The average model by minimizing the proposed multi-label focal loss.
number of emotions expressed in a sentence is 1.4468.
NLPCC2018 Dataset consists of code-switching texts in 4.3 Metrics
Chinese, and concerns another language English on a small
In multi-label emotion detection task, the evaluation is more
scale [53]. There are total 5 emotions annotated: ‘Happiness’,
complicated than traditional single-label emotion classifica-
‘Sadness’, ‘Anger’, ‘Fear’, and ‘Surprise’. After pre-process-
tion. In this paper, some popular evaluation measures typi-
ing, there is 4611 texts in training data and 955 texts in testing cally utilized in this task are utilized to measure the
data. The average number of emotions expressed in a sen- performance of proposed methods [52].
tence is 1.1466. Micro F1-score and Macro F1-score are utilized as the
The cumulative number of each emotion ei on Ren- main metrics to evaluate the global performance of each
CECps and NLPCC Datasets is calculated: model. F1 score is the harmonic mean of precision and
X
N recall. Micro F1-score gives each sample the same impor-
CNi ¼ ðyn;i ¼ 1Þ; (18) tance, while Macro F1-score takes all classes as equally
n¼1
important. Hamming Loss (HL) is the fraction of labels that
in which yn;i is the annotation of emotion ei in nth sample. are incorrectly predicted. Average precision (AP) evaluates
The statistical results are shown in Table 1. the average fraction of labels ranked above a particular label
y : y 2 Yi are actually in Yi , in which Yi is positive emotion
4.2 Experimental Details set of sentence. Coverage evaluates how far it is needed to
In this section, we illustrate the experimental details during go down the ranked emotion list to cover all the relevant
the model training. emotions in the instance. One Error (OE) evaluates the
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 481

TABLE 2
Comparison Results of Proposed Model and Baselines on RenCECps Dataset

Micro F1: % (") Macro F1: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
BR 46.40 34.79 63.69 0.2464 2.8313 0.5221 0.1789
CC 46.97 33.62 63.16 0.2282 2.9721 0.5234 0.1965
LP 45.15 42.51 62.62 0.2069 2.9117 0.5275 0.1861
BP-MLL 48.89 38.13 55.45 0.2241 3.1272 0.4625 0.3234
DPCNN 49.99 35.47 65.43 0.1583 3.0555 0.4834 0.1993
HANs 54.54 41.36 70.65 0.1504 2.4631 0.4520 0.1362
SGM 55.60 - - 0.1758 - - -
DATN - 45.70 73.20 - - 0.4150 -
SGM-IFC 58.60 - - 0.1613 - - -
S-MC-ESFE 59.24 47.73 75.19 0.1367 2.3170 0.3760 0.1163
C-MC-ESFE 55.30 34.34 74.76 0.1213 2.2765 0.3915 0.1134
MEDA 59.71 47.25 75.76 0.1378 2.2369 0.3763 0.1084
MEDA-FS 60.76 48.31 76.51 0.1249 2.2226 0.3618 0.1062

fraction of sentences whose top-ranked emotion is not in the CC and LP, we take pre-trained BERT model as sentence
relevant emotion set. Ranking Loss (RL) evaluates the aver- encoder and Gaussian Naive Bayes as the classifier, and all
age fraction of label pairs that are reversely ordered for experiments are implemented based on Scikit-multilearn
instance. library. The results of baselines BP-MLL, SGM, DATN, and
SGM-IFC on RenCECps dataset are adopted from the pub-
4.4 Baseline Models lished papers [38], [58], [59]. For others, the comparison
To demonstrate the performance of the proposed MEDA experiments are implemented based on the open-source
model, some baseline methods are compared in our codes shared on GitHub.
experiments:
BR [55], Binary Relevance, based on the label indepen- 5 EXPERIMENTAL RESULTS AND DISCUSSION
dence assumption, transforms a multi-label classification
problem into multiple binary classification problems. Experimental results of the proposed method and baseline
CC [56], Classifier Chains, a multi-label model that models are reported in Section 5.1. The discussions are
arranges binary classifiers into a classifier chain to capture organized into two sections. In section 5.2, we analyzed
the label correlations. the contribution of multi-level information from each sub-
LP, LabelPowerset, creates one multi-class classifier for model. In section 5.3, we evaluate the effectiveness of emo-
every label combination attested in the training set. tional features by ablation experiments. In Section 5.4, we
BP-MLL [52], is derived from the backpropagation algo- explore the effectiveness of proposed multi-label focal loss
rithm by employing a novel error function to capture the on this task.
characteristics of multi-label learning.
DPCNN [57], a low-complexity word-level deep pyra- 5.1 Experimental Results
mid CNN network that can efficiently capture global repre- Experimental results of the proposed methods against base-
sentations of text. lines are shown in Tables 2 and 3, the best two results on
HANs [50], hierarchical attention networks that mirror each metric are in bold and in bold italics, respectively.
the hierarchical structure of documents. HANs can find the As the results shown in Table 2, the proposed model sig-
essential words and sentences in a document while taking nificantly outperforms baseline models and achieves state-
the contextual information into consideration. of-the-art performance on RenCECps. Compared with
SGM [40], transfers multi-label classification task to a SGM-IFC [59], which has previously achieved the state-of-
sequence generation problem and can capture the correla- the-art performances, proposed MEDA-FS has improved
tions between labels. micro-F1 score from 58.60 to 60.76 percent and reduced
In previous studies, several emotion classification meth- hamming loss from 0.1613 to 0.1249. Compared with
ods have been implemented in RenCECps datasets and DATN, the proposed MEDA-FS has improved macro-F1
achieved the previous state-of-the-art performances. There- score from 45.70 to 48.31 percent, improved average preci-
fore, we take them as baselines to verify the performance of sion from 73.20 to 76.51 percent, and reduced one error
our method in RenCECps, which includes: from 0.4150 to 0. 3618. Besides, our model outperforms
DATN [58], divides the sentence representation into two other deep learning methods and commonly used machine
different feature spaces, which aims to capture the general learning methods to a great extent, such as BR algorithm
sentiment words and the other critical emotion-specific and SGM model.
words via a dual attention mechanism. Table 3 shows the experimental results of proposed
SGM-IFC [59], utilizes the attention-based Seq2Seq model and baselines on NLPCC2018 dataset. Our proposed
model to solve the multi-label problem. An initialized fully model achieved excellent results on almost all metrics
connection layer is employed to capture the correlation except hamming loss. The hamming loss of proposed
between any two different labels. For the baselines of BR, MEDA-FS is 0.1728, while the best is 0.1617 (achieved by
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
482 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

TABLE 3
Comparison Results of Proposed Model and Baselines on NLPCC2018 Dataset

Micro F1: % (") Macro F: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
BR 48.92 41.07 67.74 0.2975 2.1645 0.4958 0.2771
CC 49.92 40.51 68.63 0.2790 2.1221 0.4883 0.2668
LP 47.67 36.81 67.04 0.2456 2.1592 0.5159 0.2758
BP-MLL 55.66 41.65 74.78 0.2584 1.8896 0.4002 0.2066
DPCNN 46.07 34.25 64.22 0.2420 2.3482 0.5414 0.3231
HANs 55.69 42.78 76.92 0.2805 1.7930 0.3758 0.1835
SGM 57.11 36.28 64.24 0.1843 2.7813 0.4395 0.4267
S-MC-ESFE 63.32 49.23 77.19 0.1849 1.7340 0.3780 0.1694
C-MC-ESFE 60.59 46.90 76.43 0.1719 1.7592 0.3895 0.1749
MEDA 61.21 47.70 75.90 0.1696 1.7665 0.4021 0.1775
MEDA-FS 63.02 49.42 77.12 0.1728 1.7288 0.3812 0.1681

LP). HL is the fraction of wrong labels to the total number of specified features from both sentence-level and context-
labels and penalizes only the individual labels. There are level in each channel. This feature matrix is extracted from
mainly two reasons for the higher hamming loss. One rea- the under-layer and each dimension focused on a certain
son is that weak emotions are difficult to predict accurately. emotion, which could conclude more detailed emotion-
MC-ESFE module can prevent the features of weak emo- specified information. Another is ECorL module, which
tions from being covered by strong emotions to some extent, learns more global semantic information and emotion corre-
but not completely. Their emotional features are not notice- lations based on above emotion-specified features. These
able and are difficult to recognize. The classifier tends to two modules enable MEDA to give emotion predictions
conservatively predict them as negative emotions to ensure based on context and emotion correlation information.
the whole performance among all emotion labels. Another To verify whether the emotion correlation information is
reason is that the data distribution is imbalanced. It is hard learned in MEDA, we visualize the emotional correlation
to guarantee the performance of low-source emotion catego- coefficients matrix. It is calculated with Pearson product-
ries. In future work, more attention will be paid to the detec- moment correlation coefficients, which indicates the level to
tion of weak and low- source emotions. In addition to which two emotions vary together:
hamming loss, the global performance of proposed method
can also be reflected by other multi-label metrics, such as Rij ¼ covðEi ; Ej Þ=sEi  sEj ; (19)
micro-F1, macro-F1, and average precision, on which the
proposed method has achieved satisfying performance. Where Ei ¼ ½E1i ; E2i; . . . ENi  and Eni is the emotional
intensity of emotion ei in the nth sample. covðEi ; Ej Þ is the
covariance of ei and ej , and s is the standard deviation.
5.2 Discussion of Sub-Models Figs. 2 and 3 show the comparison of the actual correlation
MEDA-FS is composed of 3 sub-models: MEDA, S-MC- coefficients matrix on Ren-CECps and the predicted correla-
ESFE and C-MC-ESFE. These sub-models are devoted to tion coefficients matrix in MEDA model. We can observe
learning information from different levels and contributing that the distribution of positively/negatively related emo-
to a more comprehensive ensemble model. To further tion pairs predicted in MEDA is similar to the real distribu-
explore the contribution of each sub-model, we further ana- tion on Ren-CECps. Taking ‘Love’ as an example. Fig. 3
lyze their performance on RenCECps in this section. The shows that in actual distribution, the most positively related
comparison results are shown in Table 4.
MEDA. As the global performance shown in Table 2.
MEDA (micro-F1 ¼ 59.71 percent, HL ¼ 0.1378) outper-
forms the previous state-of-the-art model SGM-IFC (micro-
F1 ¼ 58.60 percent, HL ¼ 0.1613), and outperforms another
two sub-models on micro-F1, AP and ranking loss. MEDA
network consists of two modules. The first is MC-ESFE,
which is a hierarchical network and extracted emotion-

TABLE 4
Comparison Results of Sub-Models on RenCECps

Micro Macro
P R F1 P R F1
S-MC-ESFE 52.16, 68.55, 59.24 42.48 56.72 47.73
C-MC-ESFE 59.34, 51.77, 55.30 43.15 32.01 34.34
MEDA 51.81, 70.46, 59.71 41.44 57.54 47.25
MEDA-FS 55.77, 66.72, 60.76 46.10 52.21 48.31
Fig. 2. Emotional correlation coefficients matrix in RenCECp.
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 483

TABLE 5
Ablation Study on RenCECps Dataset

MEDA MEDA-FS
With Without With Without
Micro F1: % 59.71 56.22 60.76 57.33
Macro F1: % 47.25 42.91 48.31 44.16
AP: % 75.76 73.22 76.51 74.11
Hamming Loss 0.1378 0.1342 0.1249 0.1313
Coverage 2.2369 2.3703 2.2226 2.3322
One Error 0.3763 0.4101 0.3618 0.3959
Ranking Loss 0.1084 0.1239 0.1062 0.1189

‘With’ and ‘without’ denote with and without emotional features.

From Table 4, we can see that the lower F1-score of C-


MC-ESFE mainly because of the lower recall during predic-
Fig. 3. Emotional correlation coefficients matrix learned by MEDA. tion. Its micro recall is 51.77 percent while S-MC-ESFE is
68.55 percent. Although its recall is lower, it can ensure that
the prediction is more accurate: micro-precision of C-ESFE
emotion with ‘Love’ is ‘Joy’ (þ0.20) while the most nega- is 59.34 percent while S-MC-ESFE is 52.16 percent. This
tively related emotion is ‘Anxiety’ (-0.38). This means that means that the prediction given by C-MC-ESFE model is
emotions ‘Love’ and ‘Joy’ often occur together while ‘Love’ more rigorous. Therefore, with higher precision, the C-MC-
and ‘Anxiety’ rarely appear together. The above emotion ESFE model improves the confidence for the final prediction
correlation information can also be learned by MEDA: cor- of the ensemble model.
relation coefficient of ‘Love’ and ‘Joy’ is þ0.51 while ‘Love’ S-MC-ESFE, C-MC-ESFE, and MEDA mean different lev-
and ‘Anxiety’ is -0.65. Besides, there are some emotion pairs els of information from sentence-level, context-level, and
with emotion correlation that have been learned, such as emotion correlation level. They are integrated into MEDA-
‘Love’ and ‘Sorrow’ (-0.39), ‘Anxiety’ and ‘Joy’ (-0.43), ‘Hate’ FS and contribute to more accurate and stable predictions in
and ‘Anger’ (þ0.40), etc. The results demonstrate the ability emotion detection task.
of emotion correlation learning in proposed MEDA.
S-MC-ESFE. Results in Table 2 indicate that the prediction 5.3 Ablation Experiments
of S-MC-ESFE is better than baselines on most metrics. Com-
In the sentence-level embedding, we extract the emotional
pared with MEDA, S-MC-ESFE achieves a higher macro-F1
features based on the external emotional lexicon. To evalu-
value (47.73 while 47.25 percent). Although this gap is small,
ate the effect of emotional features on experimental results,
it can reflect the average level of emotion detection of each
we train the model without this feature on RenCECps data-
emotion category in S-MC-ESFE. The higher macro-F1 of S-
set. The experimental ablation results are shown in Table 5.
MC-ESFE suggests that for some sparse-resources emotion
From Table 5, both MEDA model and MEDA-FS model
categories, it could give more accurate prediction than MEDA
with emotional features outperform the models without
model. S-MC-ESFE is an intermediate model derived from
emotional features on almost all metrics. It is revealed that
MC-ESFE during sentence-level pre-training. In S-MC-ESFE,
considering emotional features can make contributions to
each channel is trained channel-wise and can be considered
the classification improvement. In deep emotion recognition
as multiple binary emotion classifier. During the training of
models, low-resource emotional datasets have been chal-
each classifier, only the parameters of the corresponding
lenging, and effectively incorporating existing emotional
channel are updated, which could make the model focus
resources is the key to improving performance. In proposed
more on the feature extraction of a specified emotion. Take a
MEDA, external emotional lexicon works as prior knowl-
sentence as an example: ‘For a long time, I write about funny
edge and is directly incorporated in sentence-level encoding.
things in my blog, but this time, my heart is heavy.’ In the
This method implements external knowledge supplementa-
channel of ‘Joy’, feature extraction will pay more attention to
tion in the simplest way and contributes to the effective
the words ‘funny things’, while ‘Sorrow’ channel focused
extraction of emotional features.
more on ‘heart is heavy’. Therefore, in each channel, the pre-
diction of whether the sentence contains the corresponding
emotion will be more accurate. 5.4 Discussion of Multi-Label Focal-Loss
C-MC-ESFE. In C-MC-ESFE model, contextual informa- In this section, we discuss the effectiveness of proposed
tion is further considered compared with S-MC-ESFE multi-label focal loss (ML-FL) on emotion detection results.
model. The results in Table 2 shows that both macro-F1 and In the definition of multi-label focal loss, w is a harmonic
micro-F1 are inferior to S-MC-ESFE. However, C-MC-ESFE factor aimed to balance the prediction of positive and neg-
achieves a better hamming loss (HL ¼ 0.1213) than MEDA ative labels. In this way, it has an effect on balancing the
(HL ¼ 0.1378). To further explore the role of C-MC-ESFE in results of precision and recall, thus obtains an optimal F1
MEDA-FS model, we further compared the micro/macro value. To verify the influence of w in emotion detection,
precision and recall of each sub-model. The results are we vary the value of w from 0. to 1. and compared it with
shown in Table 4. two other commonly used loss functions: binary cross-
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
484 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

prediction difference between positive-negative emotion


pairs. The weighting factor a is dedicated to balancing the
prediction of positive and negative labels, which aimed to
recognize as many emotions as possible while ensuring the
accuracy of prediction. The weight a is consists of two parts
to control the prediction loss of positive and negative emo-
tions, respectively:
 r  r
aipos ¼ w  1  pik ; aineg ¼ ð1  wÞ  pil : (20)

Considering the limit case, if harmonic factor w gradually


increases to the maximum w ¼ 1:
 r
aipos  1  pik ; aineg  0: (21)

In this case, as long as the model predicts all the emotion


as pi ¼ 1, it is possible to minimize aipos ¼ 0, thereby mini-
mizing the loss. In this way, the prediction gap between
positive and negative emotion pairs expððpik  pil Þ can only
play a weak role. Therefore, the results show a higher recall
while precision is difficult to be guaranteed: while w ¼ 1:0,
the micro-precision, recall, and F1-score are 39.22, 85.29,
and 53.74 percent, respectively.
Conversely, as w gradually decreases, aineg gradually
Fig. 4. The comparison results for cross-entropy(CE), multi-label loss
function(ML), and proposed multi-label focal loss for different values of increases. In this way, the prediction error of the negative
weight. label will bring greater losses. To reduce the loss, the model
predicts the positive label more conservatively, and thus the
entropy loss function(CE-loss) and multi-label loss function recall decreased and precision could be guaranteed to some
(ML-loss). The comparison experiments are implemented extent. Therefore, it can be assumed that by choosing an
on RenCECps dataset, and the results are shown in Fig. 4, appropriate value of w, it is possible to reach a balance
Table 6. between precision and recall, and then achieve satisfactory
The results of CE-loss and ML-loss both show a higher results. As the results in Fig. 4, take F1-score to measure the
recall (81.06 and 80.60 percent in micro-recall) while lower overall performance, while w 2 ½0:3; 0:5, proposed multi-
precision (41.95 and 44.39 percent in micro-precision). Preci- label focal loss outperforms cross-entropy and multi-label
sion is the average probability of relevant retrieval, while loss function. To be specific, while w ¼ 0:4, its micro-precision
recall is the average probability of complete retrieval. They is 51.81 percent, micro- recall is 70.46 percent, micro-F1-score
are two metrics restrain mutually [60]. In this emotion is 59.71 percent. Compared with ML-loss, although the recall
detection task, we hope to recognize as many emotions as drops, its micro-precision improved 7.42 percent and micro-
possible, based on the premise of ensuring precision. It can F1 improved 2.46 percent. Table 6 shows the comparison
be seen from the trend of the curve in Fig. 4: a proper w can results of different loss functions on multi-label metrics,
modulate the value between recall and precision, thus which demonstrate that multi-label loss function outperforms
achieve both higher precision and F1-score to alleviate the others.
above problems.
We will analyze the role of the parameter w in the curve 6 CONCLUSIONS
change. From the tendency of the curve in Fig. 4, we can see In this paper, a Multiple-label Emotion Detection Architec-
that as the weight w increases, precision shows a downward ture (MEDA) was proposed for the textual multi-label emo-
trend, recall shows an upward trend while the overall trend tion detection task. MEDA was composed of two modules,
of F1-score is to rise first and then fall. The loss function pro- its key idea was to capture emotion-specified features by
posed in this paper is committed to maximizing the MC-ESFE module in advance, and then learn emotion

TABLE 6
Results Comparison of MEDA Model With Different Loss Functions

Micro % (") Macro: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
P R F1 P R F1
CE-loss 41.95 81.06 55.29 37.26 64.71 42.79 70.64 0.1900 2.4496 0.4598 0.1349
ML-loss 44.39 80.60 57.25 35.49 69.84 46.07 72.27 0.1745 2.3652 0.4421 0.1251
ML-FL (w ¼ 0.4) 51.81 70.46 59.71 41.44 57.54 47.25 75.76 0.1378 2.2369 0.3763 0.1084

Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 485

correlations based on above features in ECorL module. In [9] A. Bandhakavi, N. Wiratunga, and D. Padmanabhan, “Lexicon
based feature extraction for emotion text classification,” Pattern
MC-ESFE module, information of each emotion reflected in Recognit. Lett., vol. 93, no. 1, pp. 133–142, Jul. 2017.
the text was separately encoded from sentence-level to con- [10] S. M. Liu and J. H. Chen, “A multi-label classification based
text-level, which contributed a lot to underlying fundamen- approach for sentiment classification,” Expert Syst. Appl., vol. 42,
tal feature extraction. In ECorL module, bidirectional-GRU no. 3, pp. 1083–1093, Feb. 2015.
[11] N. Colneri^c and J. Demsar, “Emotion recognition on twitter: Com-
network was utilized as emotion sequence predictor and parative study and training a unison model,” IEEE Trans. Affective
emotion correlation learning was implemented among emo- Comput., vol. 11, no. 3, pp. 433–446, Third Quarter 2018.
tion-specified features. MEDA-FS integrated three sub- [12] D. Pan and J. Nie, “Mutux at semeval-2018 task 1: Exploring
impacts of context information on emotion detection,” in Proc.
models derived from MEDA, and realized information 12th Int. Workshop Semantic Eval., 2018, pp. 345–349.
fusion from sentence-level, context-level, and emotion cor- [13] H. Huihui and R. Xia, “Joint binary neural network for Multi-label
relation level. Furthermore, to incorporate emotion correla- learning with applications to emotion classification,” in Proc. Int.
tion information into model training, multi-label focal loss Conf. Nature Lang. Process. Chin. Compu., 2018, pp. 250–259.
[14] P. D. Turney, “Thumbs up or thumbs down? Semantic orientation
was proposed for multi-label learning. The proposed model applied to unsupervised classification of reviews,” in Proc. Assoc.
achieved satisfactory performance and outperformed state- Comput. Linguist., 2002, pp. 417–424.
of-the-art models on both RenCECps and NLPCC2018 data- [15] F. Ren and Y. Wu, “Predicting user-topic opinions in twitter with
sets, which demonstrated the effectiveness of the proposed social and topical context,” IEEE Trans. Affective Comput., vol. 4,
no. 4, pp. 412–424, Fourth Quarter 2013.
method for multi-label emotion detection. [16] B. Pang, L. Lee, and S. Vaithyanathan, “Thumbs up? Sentiment
There is still much space for improvements in our works. classification using machine learning techniques,” in Proc. Conf.
Discernible feature representation of the weak emotion cate- Empir. Methods Natural Lang. Process., 2012, pp. 79–86.
[17] S. Mohammad and S. Kiritchenko. “Understanding emotions: A
gory is a critical problem in multi-label emotion detection dataset of tweets to study interactions between affect categories,”
task. Our proposed MC-ESFE module can prevent the fea- in Proc. 11th Int. Conf. Lang. Resour. Eval., 2018, pp. 198–209.
tures of weak emotions from being covered by strong emo- [18] N. H. Frijda, “The laws of emotion,” Amer. Psychol., vol. 43, no. 5,
tions to some extent, but not completely. In future work, we pp. 349–358, May 1988.
[19] P. Ekman, “An argument for basic emotions,” Cogn. Emotion,
will try to explore more effective methods to recognize vol. 6, no. 3-4, pp. 169–200, 1992.
weak emotions more accurately. During the emotional fea- [20] Q. Changqin and F. Ren. “Sentence emotion analysis and recogni-
ture extraction in our model, an external emotion lexicon tion based on emotion words using Ren-cecps,” Int. J. Adv. Intell.,
vol. 2, no. 1, pp. 105–117, Jul. 2010.
was severed as prior knowledge to enhance emotional fea- [21] J. Li and F. Ren, “Creating a chinese emotion lexicon based on cor-
ture representation. Abundant resources are the basis of pus Ren-cecps,” in Proc. IEEE Int. Conf. Cloud Comp. Intell. Syst.,
neural network training. In future work, more emotional 2011, pp. 80–84.
data will be incorporated for better emotion understanding. [22] C. Strapparava and A. Valitutti, “Wordnet affect: An affective
extension of wordnet,” in Proc. 4th Int. Conf. Lang. Resour. Eval.,
2004, vol. 4, pp. 1083–1086.
ACKNOWLEDGMENTS [23] S. Mohammad, “Portable features for classifying emotional text,”
in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguist.: Hum.
This work was supported in part by the Research Clusters pro- Lang. Technol., 2012, pp. 587–591.
gram of Tokushima University under Grant No. 2003002. This [24] Q. Liu and S. Li, “Word semantic similarity computation based
research has been partially supported by NSFC-Shenzhen on hownet,” Comput. Linguist. Chin. Lang. Process., vol. 7, no. 2,
Joint Foundation (Key Project) (Grant No. U1613217). pp. 59–76, 2002.
[25] S. M. Mohammad and P. D. Turney, “Crowdsourcing a word–
emotion association lexicon,” Comput. Intell., vol. 29, no. 3,
REFERENCES pp. 436–465, Sep. 2012.
[26] F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget:
[1] W. Liang, H. Xie, Y. Rao, R. Y. K. Lau, and F. L. Wang, “Universal Continual prediction with LSTM,” in Proc. 9th Int. Conf. Artif. Neu-
affective model for readers’ emotion classification over short ral Netw., 1999, pp. 850–855.
texts,” Expert Syst. Appl., vol. 114, pp. 322–333, Dec. 2018. [27] S. Poria, E. Cambria, D. Hazarika, and N. Majumder, “Context-
[2] F. Ren and K. Matsumoto, “Semi-Automatic creation of youth slang dependent sentiment analysis in user-generated videos,” in Proc.
corpus and its application to affective computing,” IEEE Trans. 55th Annu. Meeting Assoc. Comput. Linguist., 2017, pp. 873–883.
Affect. Comput., vol. vol.7, no. 2, pp. 176–189, Second Quarter 2016. [28] D. Tang, B. Qin, and T. Liu, “Aspect level sentiment classification
[3] H. Y. Shum, X. He, and D. Li, “From eliza to xiaoice: Challenges with deep memory network,” in Proc. Conf. Empir. Methods Natural
and opportunities with social chatbots,” Front. Inf. Technol. Elec- Lang. Process., 2016, pp. 214–224.
tron. Eng., vol. 19, no. 1, pp. 10–26, Jan. 2018. [29] Y. Kim, “Convolutional neural networks for sentence classi-
[4] R. Jayakrishnan, G. N. Gopal, and M. S. Santhikrishna, “Multi- fication,” in Proc. Conf. Empir. Methods Natural Lang. Process., 2014,
class emotion detection and annotation in malayalam novels,” in pp. 1746–1751.
Proc. Int. Conf. Comput. Commun. Inform., 2018, pp. 1–5. [30] T. Rao, X. Li, H. Zhang, and M. Xu, “Multi-level region-based con-
[5] T. Kowatsch, M. Nißen, C. H. I. Shih and D. R€ uegger, “Text-based volutional neural network for image emotion classification,” Neu-
healthcare chatbots supporting patient and health professional rocomputing, vol. 333, no. 14, pp. 429–439, Mar. 2019.
teams: Preliminary results of a randomized controlled trial on [31] M. S. Akhtar, D. Ghosal, and A. Ekbal, “A Multi-task ensemble
childhood obesity,” in Proc. Persuasive Embodied Agents Behav. framework for emotion, sentiment and intensity prediction,”
Chang., 2017, pp. 1–10. 2018, arXiv:1808.01216.
[6] C. Chen, R. Zhuo, and J. Ren, “Gated recurrent neural network [32] M. Chae, T. H. Kim, Y. H. Shin, J. W. Kim, and S. Y. Lee, “End-to-
with sentimental relations for sentiment classification,” Inf. Sci., end multimodal emotion and gender recognition with dynamic
vol. 502, pp. 268–278, Oct. 2019. weights of joint loss,” 2018, arXiv: 1809.00758.
[7] D. Sznycer and A. W. Lukaszewski, “The emotion–valuation [33] P. Xu, A. Madotto, C. S. Wu, J. H. Park, and P. Fung, “Emo2vec:
constellation: Multiple emotions are governed by a common Learning generalized emotion representation by multi-task train-
grammar of social valuation,” Evol. Hum. Behav., vol. 40, no. 4, ing,” in Proc. 9th Workshop Comput. Approaches Subjectivity, Senti-
pp. 395–404, Jul. 2019. ment Social Media Anal., 2018, pp. 292–298.
[8] X. Kang, F. Ren, and Y. Wu, “Exploring latent semantic informa- [34] A. Illendula and A. Sheth, “Multimodal emotion classification,” in
tion for textual emotion recognition in blog articles,” IEEE/CAA J. Proc. World Wide Web Conf., 2019, pp. 439–449.
Automatica Sinica, vol. 5, no. 1, pp. 204–216, Jan. 2018.

Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
486 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023

[35] L. Buitinck, J. V. Amerongen, and E. Tan, “Multi-emotion detec- [55] O. Luaces, J. Dıez, J. Barranquero, and J. D. Coz, “Binary relevance
tion in user-generated reviews,” in Proc. Eur. Conf. Inf. Retrieval, efficacy for multilabel classification,” Program. Artif. Intell., vol. 1,
2015, pp. 43–48. no. 4, pp. 303–313, 2012.
[36] Y. Kim, H. Lee, and K. Jung, “AttnConvnet at semeval-2018 task 1: [56] J. Read, B. Pfahringer, G. Holmes, and E. Frank, “Classifier
Attention-based convolutional neural networks for multi-label chains for multi-label classification,” Mach. Learn., vol. 85,
emotion classification,” in Proc. 12th Int. Workshop Semantic Eval., no. 3, pp. 333–359, 2011.
2018, pp. 141–145. [57] R. Johnson and T. Zhang, “Deep pyramid convolutional neural
[37] J. Yu, L. Marujo, J. Jiang, and P. Karuturi, “Improving multi-label networks for text categorization,” in Proc. 55th Annu. Meeting
emotion classification via sentiment classification with dual atten- Assoc. Comput. Linguist., 2017, vol. 1, pp. 562–570.
tion transfer network,” in Proc. Conf. Empir. Methods Natural Lang. [58] J. Yu, L. Marujo, J. Jiang, P. Karuturi, and W. Brendel, “Improving
Process., 2018, pp. 1097–1102. multi-label emotion classification via sentiment classification with
[38] D. Zhou, Y. Yang, and Y. He, “Relevant emotion ranking from text dual attention transfer network,” in Proc. Assoc. Comput. Linguis-
constrained with emotion relationships,” in Proc. Conf. North tics, 2018, pp. 1097–1102.
Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol., 2018, [59] W. Liao, Y. Wang, Y. Yin, X. Zhang, and P. Ma, “Improved
vol. 1, pp. 561–571. sequence generation model for multi-label classification via CNN
[39] M. L. Zhang and Z. H. Zhou, “Multi-label neural networks with and initialized fully connection,” Neurocomputing, vol. 382, no. 21,
applications to functional genomics and text categorization,” IEEE pp. 188–195, Mar. 2020.
Trans. Knowl. Data Eng., vol. 18, no. 10, pp. 1338–1351, Oct. 2006. [60] D. M. Powers, “Evaluation: From precision, recall and F-measure
[40] P. Yang, X. Sun, W. Li, S. Ma, W. Wu, and H. Wang, “SGM: to ROC, informedness, markedness and correlation,” J. Mach.
Sequence generation model for multi-label classification,” in Proc. Learn. Technol., vol. 2, no. 1, pp. 37–63, Dec. 2011.
27th Int. Conf. Comput. Linguistics, 2018, pp. 3915–3926.
[41] B. Zhao, X. Li, X. Lu, and Z. Wang, “A CNN–RNN architecture for Jiawen Deng received the double master’s degree
multi-label weather recognition,” Neurocomputing, vol. 322, no. 17, in advanced technology and science from Tokush-
pp. 47–57, Dec. 2018. ima University, Japan, and mechanical engineering
[42] S. Lian, J. Liu, R. Lu, and X. Luo, “Captured multi-label relations from Nantong University, China. She is currently
via joint deep supervised autoencoder,” Appl. Soft. Comput., working toward the PhD degree with Tokushima
vol. 74, pp. 709–728, Jan. 2019. University. Her research interests include natural
[43] C. K. Yeh, W. C. Wu, W. J. Ko, and Y. C. F. Wang, “Learning deep language processing and affective computing.
latent space for multi-label classification,” in Proc. 31th AAAI Conf.
Artif. Intell., 2017, pp. 2838–2844.
[44] K. Wang, M. Yang and W. Yang, “Deep correlation structure pre-
served label space embedding for Multi-label classification,” in
Proc. Asian Conf. Mach. Learn., 2018, vol. 95, pp. 1–16.
[45] D. A. Phan, Y. Matsumoto, and H. Shindo, “Autoencoder for Fuji Ren (Senior Member, IEEE) received the
semisupervised multiple emotion detection of conversation tran- PhD degree from the Faculty of Engineering, Hok-
scripts,” IEEE Trans. Affect. Comput., to be published, doi: kaido University, Japan, in 1991. From 1991
10.1109/TAFFC.2018.2885304. to1994, he worked at CSK as a chief researcher.
[46] Z. F. He, M. Yang, Y. Gao, H. D. Liu, and Y. Yin, “Joint multi-label In 1994, he joined the Faculty of Information Sci-
classification and label correlations with missing labels and fea- ences, Hiroshima City University, as an associate
ture selection,” Knowl. Based Syst., vol. 163, pp. 145–158, Jan. 2019. professor. Since 2001, he has been a professor of
[47] M. Rei and A. Søgaard, “Jointly learning to label sentences and the Faculty of Engineering, Tokushima University.
tokens,” in Proc. 33th AAAI Conf. Artif. Intell., 2019, pp. 6916–6923. His current research interests include natural lan-
[48] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- guage processing, artificial intelligence, affective
training of deep bidirectional transformers for language under- computing, emotional robot. He is the academi-
standing,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. cian of The Engineering Academy of Japan and EU Academy of Scien-
[49] J. Chung, C. Gulcehre, K. H. Cho, and Y. Bengio, “Empirical eval- ces. He is an editor-in-chief of the International Journal of Advanced
uation of gated recurrent neural networks on sequence mod- Intelligence, a vice president of CAAI, and a fellow of The Japan Federa-
eling,” in NIPS Workshop Deep Learn., 2014, arXiv:1412.3555. tion of Engineering Societies, a fellow of IEICE, a fellow of CAAI. He is the
[50] Z. Yang, D. Yang, C. Dyer, X. He, and A. Smola, “Hierarchical president of International Advanced Information Institute, Japan.
attention networks for document classification,” in Proc. Conf.
North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Tech-
nol., 2016, pp. 1480–1489. " For more information on this or any other computing topic,
[51] T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss please visit our Digital Library at [Link]/csdl.
for dense object detection,” in Proc. IEEE Int. Conf. Comput. Vis.,
2017, pp. 2980–2988.
[52] M. L. Zhang and Z. H. Zhou, “Multilabel neural networks with
applications to functional genomics and text categorization,” IEEE
Trans. Knowl. Data Eng., vol. 18, no. 10, pp. 1338–1351, Oct. 2006.
[53] Z. Wang, S. Li, F. Wu, Q. Sun, and G. Zhou, “Overview of NLPCC
2018 shared task 1: Emotion detection in code-switching text,” in
Proc. CCF Int. Conf. Natural Lang. Process. Chin. Comput., 2018,
pp. 429–433.
[54] S. Li, Z. Zhao, R. Hu, W. Li, T. Liu, and X. Du, “Analogical reason-
ing on chinese morphological and semantic relations,” in Proc.
56th Annu. Meeting Assoc. Comput. Linguist., 2018, pp. 138–143.

Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.

You might also like