Multi-Label Emotion Detection Framework
Multi-Label Emotion Detection Framework
Abstract—Textual emotion detection is an attractive task while previous studies mainly focused on polarity or single-emotion
classification. However, human expressions are complex, and multiple emotions often co-occur with non-negligible emotion
correlations. In this paper, a Multi-label Emotion Detection Architecture (MEDA) is proposed to detect all associated emotions
expressed in a given piece of text. MEDA is mainly composed of two modules: Multi-Channel Emotion-Specified Feature Extractor (MC-
ESFE) and Emotion Correlation Learner (ECorL). MEDA captures underlying emotion-specified features through MC-ESFE module,
which is composed of multiple channel-wise ESFE networks. Each channel in MC-ESFE is devoted to the feature extraction of a
specified emotion from sentence-level to context-level through a hierarchical structure. With underlying features, emotion correlation
learning is implemented through an emotion sequence predictor in ECorL. Furthermore, we define a new loss function: multi-label focal
loss. With this loss function, the model can focus more on misclassified positive-negative emotion pairs and improve the overall
performance by balancing the prediction of positive and negative emotions. The evaluation of proposed MEDA architecture is carried
out on emotional corpus: RenCECps and NLPCC2018 datasets. The experimental results indicate that the proposed method can
achieve better performance than state-of-the-art methods in this task.
In this paper, a Multi-label Emotion Detection Architec- The rest of this paper is organized as follows: Section 2
ture (MEDA) is proposed to address the above challenges. presents a brief overview of related work on multi-label emo-
MEDA is mainly composed of two modules: Multi-Channel tion detection task. Details of the proposed MEDA are given
Emotion-Specified Feature Extractor (MC-ESFE) and Emo- in Section 3. The experimental setting and details are pre-
tion Correlation Learner (ECorL). MC-ESFE consists of mul- sented in Section 4. The performance of MEDA is discussed
tiple channels by which the features of each emotion are in Section 5. Finally, conclusions are drawn in Section 6.
separately encoded. Each channel is devoted to the underly-
ing feature representation of a specified emotion from both
sentence-level and context-level. Furthermore, an external 2 RELATED WORK
emotion lexicon is introduced as prior knowledge to inte- As an important task in natural language processing, textual
grate more detailed emotional information. ECorL module sentiment analysis is an emerging research field. Tradition-
is devoted to learning emotion correlation based on ally, sentiment analysis is applied to predict a polarity label
extracted emotion-specified features from MC-ESFE. In or star level of product reviews [14], [15], or film reviews
ECorL, multi-label emotion detection task is transformed as [16]. With the delving of related researches, emotion classifi-
an emotion sequence prediction task. Bidirectional GRU cation is becoming more granular [17]. There are many dif-
network is taken as the emotion sequence predictor, and the ferent taxonomies [18], such as Paul Ekman’s six basic
emotions are sequentially predicted in a fixed path. In the emotions [19], and Fuji Ren’s eight basic emotions [20].
hidden state of each step, the emotion correlations of cur- As one of the most obvious clues of sentiment analysis,
rent emotion are learned by information interaction with emotional lexical resources directly encode sentimental
the context of other emotions flowed from both forward knowledge [1], [21], and are widely used, such as WordNet-
and backward directions. Considering that the proposed affect [22], NRC emotion lexicon [23], and Hownet [24].
MEDA network extracts emotional information from sen- These emotional lexicons are the basis for early hand-
tence level, context level, and emotion correlation level, an crafted features based emotional analysis [21], [25]. Given
ensemble model called MEDA-FS is proposed to integrate the recent success of deep learning models, various neural
emotional information from different levels. MEDA-FS can network models have been proposed and have achieved
realize the maximization of information retention and avoid highly competitive performance in sentiment analysis.
information loss during bottom-up learning. During the LSTM networks [26] have shown its superior performance
training, positive-negative emotion correlation is incorpo- in context information encoding [27], [28], and CNN net-
rated into the proposed multi-label focal loss function. By works [29], [30] are often utilized to extract local informa-
introducing a weighting factor, our loss will focus more on tion within a sentence. Multi-task ensemble network
misclassified emotion pairs and balance the prediction performs well in emotional information integration and can
between positive and negative emotions. alleviate the impact of insufficient corpus [31], [32]. To
Compared with existing multi-label emotion detection achieve robust emotion feature representation, Emo2Vec is
methods, the proposed MEDA architecture extracts both emo- trained in [33] to encode emotional semantics by multi-task
tion-specified features and emotion correlations. The perfor- learning six different emotion-related tasks. Different
mance of the proposed MEDA is verified in Chinese emotional modalities, such as text, emoji, and images, can also be com-
corpus: RenCECps. Experimental results show that the pro- bined [34] to express emotion and complement each other
posed architecture achieves state-of-the-art performance on for emotion classification. To further model the effects of
RenCECps and demonstrates the effectiveness of MEDA. sentimental relations, a modified GRNN is proposed [6] to
The major contributions of our paper can be summarized encode the sentiment polarity and sentiment modifier con-
as follows. text separately.
In most previous studies, the complexity of emotion
1. MEDA architecture composed of MC-ESFE and detection task is often narrowed down by focusing on single
ECorL modules is proposed for the textual multi- emotion classification. However, human emotion is com-
label emotion detection task. MC-ESFE can encode plex in reality, and the textual expression often contains
emotion-specified features in the corresponding multiple emotions simultaneously. To address this problem,
channel respectively, which strengthens the underly- the multi-label emotion detection task can be viewed as a
ing feature representation of each emotion. ECorL is special multi-label classification [10], [35], in which emo-
proposed to learn emotion correlations by transform- tions are the multiple labels. Attention-based multi-label
ing multi-label emotion detection task into emotion sentence classifier is proposed in [36] to imitate how
sequence prediction task. humans comprehend and classify emotions. A dual atten-
2. MEDA-FS is proposed to fuse the information at sen- tion-based transfer learning model is proposed in [37] to
tence-level, context-level, and emotion correlation extract both general sentiment words and other emotion-
level, which can realize the maximization of informa- specific words. Linguistic characteristics are explored in
tion retention during bottom-up learning. [12] to reduce the limitation of lexicon coverage and size.
3. Multi-label focal loss function considering emotion However, most of the above models do not take multi-
correlation information is proposed for multi-label label correlations into account, and they assume that the
learning. This loss function contributes to model multiple labels are independent of each other. Some stud-
training by focusing on misclassified emotion pairs ies try to explore the label correlations. To reduce the effect
and balancing the prediction of positive and nega- of irrelevant labels, prior knowledge of co-occur label rela-
tive emotions. tionships are incorporated [38] as a constraint for emotion
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 477
Fig. 1. (a) The illustration of the proposed MEDA: Multi-label Emotion Detection Architecture. MEDA mainly composes two modules: Multi-channel
emotion-specified feature extractor (MC-ESFE) and emotion correlation learner (ECorL). (b) The illustration of lth ESFE channel of emotion el in
MC-ESFE module. (c) The illustration of sentence-level encoder of lth ESFE channel in MC-ESFE module.
prediction and ranking. Some approaches attempt to Multi-label emotion detection task aims to detect all pos-
implicitly estimate label correlations by the modification of sible emotions from the pre-defined emotional label set: E ¼
loss function. The label-correlation sensitive loss function ½e1 ; e2 ; . . . eL . Considering the important influence of contex-
is first proposed in [39] with the BP-MLL algorithm. Joint tual information on this task, the previous k sentences
binary cross-entropy (JBCE) loss is proposed by Huihui He occurred before current sentence s are taken as the context
[14] in his joint binary neural network (JBNN) to capture sentences: scxt ¼ ½sk ; . . . s2 ; s1 . Given a sentence s ¼
label relations. Multi-label classification can also be trans- ½w1 ; w2... wn and its context scxt , our proposed multi-label
formed into a sequence generation problem [40], [41] to emotion recognition model MEDA is trained to output the
capture label correlations. To reduce the computational predicted probability distribution PML ¼ ½p1 ; p2 ; . . . pL of
complexity, partial label dependence can also contribute to each emotion, denoted as:
this task, which is demonstrated in [42]. Deep canonical
correlation analysis (DCCA) performs well in feature- PML ¼ fMEDA ðs; scxt Þ: (1)
aware label embedding and label-correlation aware predic-
tion [43], [44]. A semisupervised multi-label method is pro- MC-ESFE module is composed of L parallel channel-
posed in [45] while label correlations are incorporated by wise ESFE. In each channel, ESFE extracts emotion-specified
modifying the loss function. Multi-Label classification and features from sentence-level to context-level through a hier-
label correlations learning can also be realized in a joint archical structure. The output of each channel is combined
cxt
learning framework [46], [47]. into an emotion-specified feature matrix: XES ¼ ½xcxt
ES1 ;
In contrast with most current methods, we focus on emo- xcxt
ES2 ; . . . ; x cxt
ESL . In ECorL module, emotion correlations are
cxt
tion-specified feature extraction and emotion correlation. further learned from XES and multi-label emotions are pre-
There are mainly two fundamental differences: dicted. Specifically, MEDA architecture is very flexible, and
the algorithm applied in each module can be replaced by
1. The information of each emotion is encoded sepa-
other state-of-the-art algorithms.
rately, which can concentrate more on underlying
emotion-specified feature extraction. This implemen-
3.1 MC_ESFE: Multi-Channel Emotion Specified
tation contributes to further emotion correlation
Feature Extractor
learning in emotion sequence predictor.
In this paper, a Multi-Channel Emotion-Specified Feature
2. The proposed multi-label focal loss function consid-
Extractor (MC-ESFE) is proposed for underlying fundamen-
ers emotion correlation information. It can pay more
tal feature extraction. MC-ESFE is composed of L channel-
attention to misclassified emotion pairs and balance
wise ESFE, and L is equal to the number of emotions. Each
the prediction of positive and negative emotions.
channel focuses on the feature extraction of a specified emo-
tion, and each emotion’s information is separately encoded
3 PROPOSED METHOD in each channel. In this way, more details of each emotion
To comprehensively obtain emotional information of texts, could be summarized, and the features of weak emotions
Multi-label Emotion Detection Architecture (MEDA) is pro- are prevented from being covered by strong emotions to
posed in this paper. It mainly composes two modules: some extent.
Multi-Channel Emotion-Specified Feature Extractor (MC- Fig. 1 shows the hierarchical structure of lth ESFE-chan-
ESFE) and Emotion Correlation Learner (ECorL). The nel corresponding to emotion el , l 2 ½1; L. Each channel
framework of MEDA is shown as Fig. 1. contains a sentence-level encoder and a context-level
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
478 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023
encoder, which focus on feature extraction of emotion el on el . More attention weight will be assigned to words related
both sentence-level and context-level. to emotion el in the current lth ESFE channel. Attention
weight ai and weighted emotional feature vector xew are
3.1.1 Sentence Level Encoder defined as follows:
In lth ESFE channel, given a sentence s, the sentence-level
l ei ¼ W2T s W1T hi þ b1 þ b2 (3Þ
encoder fSEn projects input sentence s to emotion-specified
s
feature xESl : expðei Þ
ai ¼ P n (4Þ
k¼1 expðek Þ
xsESl ¼ fSEn
l
ðsÞ; l 2 ½1; L: (2) xew ¼ ½a1 h1 : a2 h2 : . . . ai hi : . . .; (5Þ
in which He ¼ ½he1 ; . . . hel ; . . . heL ; are the hidden states of learning, MEDA-FS is proposed to fuse the information
each step, WECor and bECor are the learned weight and from different levels. MEDA-FS consists of three sub-mod-
biases, and PML is the predicted probability of each emotion. els: S-MC-ESFE, C-MC-ESFE, and MEDA, which give emo-
In lth step of BiGRU, the learning of hidden state hel can be tion predictions on sentence-level, context-level, and
viewed as the feature extraction of a specified emotion el . emotion correlation level, respectively.
Invalid information of current input xcxt ESl can be filtered S-MC-ESFE, gives sentence-level predictions PES s
¼
s s
because of the gating mechanism. With the bidirectional ½pES1 ; . . . ; pESL . It is obtained during the pre-training step of
network, emotional feature hel is learned based on the infor- sentence-level encoder in MC-ESFE, which is detailed in
s
mation of other emotions flowed from both forward and Section 3.3. PES represents the prediction based on the
backward hidden state. In this way, emotional information underlying information, without considering the emotion
interaction is realized. The hidden states of BiGRU are out- correlations and contextual information.
cxt
put and fed into emotion interaction layer. This layer is a C-MC-ESFE, gives context-level predictions PES based on
fully-connected layer and aimed to realize further emotional sentence-level predictions of current sentence s, denoted as
s
information interaction. In this way, the final emotion pre- PES , and sentence-level predictions of its context scxt , denoted
s s
diction is obtained: PML ¼ ½p1 ; p2 ; ::pL . as ½PESk ; . . . ; PES1 . GRU network is utilized to learn contex-
tual information and its final output is taken as the prediction:
3.3 Network Pre-Training in MC-ESFE s
cxt s s
Each channel in MC-ESFE is dedicated to obtaining corre- PES ¼ fGRU PESk ; . . . PES1 ; PES : (14)
sponding emotional information, which belongs to the
MEDA, gives prediction PML ¼ ½p1 ; p2 ; ::pL by consider-
underlying feature extraction in the MEDA framework. The
ing both contextual information and emotion correlation.
quality of feature representation has a direct impact on the
MEDA-FS, gives final predictions by comprehensively
performance of upper-level emotion predictions. To improve
fuse the information from above three level, denoted as:
the underlying feature representation, network-based trans-
fer learning is employed to pre-train the sentence-level s
P ¼ ws PES cxt
þ wc PES þ wML PML : (15)
encoder in each channel. During transfer learning, a predic-
tion layer is added to emotional sentence representation xsESl in which ws , wc and wML are the weight parameters of each
for single-emotion prediction: level’s information.
psESl ¼ s wl xsESl þ bl ; (11) 3.5 Definition of Multi-Label Focal Loss
Multi-label (ML) loss function [52] is one of the most com-
in which xsESl is the sentence-level representation of input monly used loss functions in multi-label learning. Instead of
sentence s, and psESl indicates the predicted probability of concentrating on individual label discrimination like tradi-
emotion el expressed in sentence s. The first-step pre-train- tional cross-entropy loss function, ML-loss focused on con-
ing is implemented on positive-negative annotated emo- sidering the correlations between the different labels.
tional datasets. The second-step is fine-tuning. During fine- Inspired by [51], we rewrite ML-loss and called it multi-
tuning, the multi-emotion annotation fs; ½y1 ; y2 ; . . . yL g of label focal loss. Multi-label focal loss not only considers
each sentence s in dataset D is transformed into multiple emotion correlation but also focus more on misclassified
single-emotion annotations: fs; y1 g, fs; y2 g. . .fs; yL g. In this emotion pairs. Besides, it introduces a harmonic parameter
way, we reconstructed multiple binary-dataset: D ^¼
to reduce the influence of the imbalance prediction of posi-
^ ^ ^ ^
fD1 ; D2 ; . . . DL g. For each binary-dataset Dl , l 2 ½1; L, sen- tive and negative emotions. The definition of multi-label
tence s 2 D ^ l is fed into lth ESFE
^ l is annotated as fs; yl g. D
focal loss is defined as follows:
channel to fine-tune the sentence-level parameters. During
pre-training, binary focal loss [51] is utilized: XN 1 X
EMLFL ¼ aikl expð pik pil (16Þ
jYi j Yi
i¼1 ð k;lÞ2Y i Y i
EFL ¼ at ð1 pt Þr logðpt Þ (12Þ i
i r
i r
akl ¼ w 1 pk þ ð1 wÞ pl ; (17Þ
psESl if yl ¼ 1
pt ¼ ; (13Þ in which Yi denotes the set of positive emotions expressed
1 psESl otherwise
in ith instance si , and Yi denotes the negative emotion set. pik
in which r is a modulating factor, and it aimed to reduce the and pil are the predicted probability of positive emotion ek
relative loss of well-classified examples. a 2 ½0; 1; is a and negative emotion el respectively. Therefore, the training
weighting factor to address the problem of class imbalance. with above loss function is equivalent to maximizing the dif-
at ¼ a for positive label and at ¼ 1 a for negative label. ference of negatively related emotion pair of (pik pil ). This
leads the system to output a higher probability for positive
3.4 MEDA-FS: Multi-Level Information Fusion emotion while a lower probability for negative emotion. In
The proposed MEDA architectural learns emotional infor- this way, the emotion correlation of negatively related emo-
mation from sentence-level to context-level, from single- tion pairs can be taken into consideration.
emotion level in MC-ESFE to multi-emotion level in ECorL aikl is a weighting factor and mainly affected by two
module. Each layer in MEDA network learns different lev- parameters: w 2 ð0; 1Þ is a harmonic factor aimed to balance
els of information. To realize the maximization of informa- the prediction between positive and negative emotions, and
tion retention and avoid information loss during bottom-up r > 0 is a modulating factor aimed to make the loss put
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
480 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023
TABLE 2
Comparison Results of Proposed Model and Baselines on RenCECps Dataset
Micro F1: % (") Macro F1: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
BR 46.40 34.79 63.69 0.2464 2.8313 0.5221 0.1789
CC 46.97 33.62 63.16 0.2282 2.9721 0.5234 0.1965
LP 45.15 42.51 62.62 0.2069 2.9117 0.5275 0.1861
BP-MLL 48.89 38.13 55.45 0.2241 3.1272 0.4625 0.3234
DPCNN 49.99 35.47 65.43 0.1583 3.0555 0.4834 0.1993
HANs 54.54 41.36 70.65 0.1504 2.4631 0.4520 0.1362
SGM 55.60 - - 0.1758 - - -
DATN - 45.70 73.20 - - 0.4150 -
SGM-IFC 58.60 - - 0.1613 - - -
S-MC-ESFE 59.24 47.73 75.19 0.1367 2.3170 0.3760 0.1163
C-MC-ESFE 55.30 34.34 74.76 0.1213 2.2765 0.3915 0.1134
MEDA 59.71 47.25 75.76 0.1378 2.2369 0.3763 0.1084
MEDA-FS 60.76 48.31 76.51 0.1249 2.2226 0.3618 0.1062
fraction of sentences whose top-ranked emotion is not in the CC and LP, we take pre-trained BERT model as sentence
relevant emotion set. Ranking Loss (RL) evaluates the aver- encoder and Gaussian Naive Bayes as the classifier, and all
age fraction of label pairs that are reversely ordered for experiments are implemented based on Scikit-multilearn
instance. library. The results of baselines BP-MLL, SGM, DATN, and
SGM-IFC on RenCECps dataset are adopted from the pub-
4.4 Baseline Models lished papers [38], [58], [59]. For others, the comparison
To demonstrate the performance of the proposed MEDA experiments are implemented based on the open-source
model, some baseline methods are compared in our codes shared on GitHub.
experiments:
BR [55], Binary Relevance, based on the label indepen- 5 EXPERIMENTAL RESULTS AND DISCUSSION
dence assumption, transforms a multi-label classification
problem into multiple binary classification problems. Experimental results of the proposed method and baseline
CC [56], Classifier Chains, a multi-label model that models are reported in Section 5.1. The discussions are
arranges binary classifiers into a classifier chain to capture organized into two sections. In section 5.2, we analyzed
the label correlations. the contribution of multi-level information from each sub-
LP, LabelPowerset, creates one multi-class classifier for model. In section 5.3, we evaluate the effectiveness of emo-
every label combination attested in the training set. tional features by ablation experiments. In Section 5.4, we
BP-MLL [52], is derived from the backpropagation algo- explore the effectiveness of proposed multi-label focal loss
rithm by employing a novel error function to capture the on this task.
characteristics of multi-label learning.
DPCNN [57], a low-complexity word-level deep pyra- 5.1 Experimental Results
mid CNN network that can efficiently capture global repre- Experimental results of the proposed methods against base-
sentations of text. lines are shown in Tables 2 and 3, the best two results on
HANs [50], hierarchical attention networks that mirror each metric are in bold and in bold italics, respectively.
the hierarchical structure of documents. HANs can find the As the results shown in Table 2, the proposed model sig-
essential words and sentences in a document while taking nificantly outperforms baseline models and achieves state-
the contextual information into consideration. of-the-art performance on RenCECps. Compared with
SGM [40], transfers multi-label classification task to a SGM-IFC [59], which has previously achieved the state-of-
sequence generation problem and can capture the correla- the-art performances, proposed MEDA-FS has improved
tions between labels. micro-F1 score from 58.60 to 60.76 percent and reduced
In previous studies, several emotion classification meth- hamming loss from 0.1613 to 0.1249. Compared with
ods have been implemented in RenCECps datasets and DATN, the proposed MEDA-FS has improved macro-F1
achieved the previous state-of-the-art performances. There- score from 45.70 to 48.31 percent, improved average preci-
fore, we take them as baselines to verify the performance of sion from 73.20 to 76.51 percent, and reduced one error
our method in RenCECps, which includes: from 0.4150 to 0. 3618. Besides, our model outperforms
DATN [58], divides the sentence representation into two other deep learning methods and commonly used machine
different feature spaces, which aims to capture the general learning methods to a great extent, such as BR algorithm
sentiment words and the other critical emotion-specific and SGM model.
words via a dual attention mechanism. Table 3 shows the experimental results of proposed
SGM-IFC [59], utilizes the attention-based Seq2Seq model and baselines on NLPCC2018 dataset. Our proposed
model to solve the multi-label problem. An initialized fully model achieved excellent results on almost all metrics
connection layer is employed to capture the correlation except hamming loss. The hamming loss of proposed
between any two different labels. For the baselines of BR, MEDA-FS is 0.1728, while the best is 0.1617 (achieved by
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
482 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023
TABLE 3
Comparison Results of Proposed Model and Baselines on NLPCC2018 Dataset
Micro F1: % (") Macro F: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
BR 48.92 41.07 67.74 0.2975 2.1645 0.4958 0.2771
CC 49.92 40.51 68.63 0.2790 2.1221 0.4883 0.2668
LP 47.67 36.81 67.04 0.2456 2.1592 0.5159 0.2758
BP-MLL 55.66 41.65 74.78 0.2584 1.8896 0.4002 0.2066
DPCNN 46.07 34.25 64.22 0.2420 2.3482 0.5414 0.3231
HANs 55.69 42.78 76.92 0.2805 1.7930 0.3758 0.1835
SGM 57.11 36.28 64.24 0.1843 2.7813 0.4395 0.4267
S-MC-ESFE 63.32 49.23 77.19 0.1849 1.7340 0.3780 0.1694
C-MC-ESFE 60.59 46.90 76.43 0.1719 1.7592 0.3895 0.1749
MEDA 61.21 47.70 75.90 0.1696 1.7665 0.4021 0.1775
MEDA-FS 63.02 49.42 77.12 0.1728 1.7288 0.3812 0.1681
LP). HL is the fraction of wrong labels to the total number of specified features from both sentence-level and context-
labels and penalizes only the individual labels. There are level in each channel. This feature matrix is extracted from
mainly two reasons for the higher hamming loss. One rea- the under-layer and each dimension focused on a certain
son is that weak emotions are difficult to predict accurately. emotion, which could conclude more detailed emotion-
MC-ESFE module can prevent the features of weak emo- specified information. Another is ECorL module, which
tions from being covered by strong emotions to some extent, learns more global semantic information and emotion corre-
but not completely. Their emotional features are not notice- lations based on above emotion-specified features. These
able and are difficult to recognize. The classifier tends to two modules enable MEDA to give emotion predictions
conservatively predict them as negative emotions to ensure based on context and emotion correlation information.
the whole performance among all emotion labels. Another To verify whether the emotion correlation information is
reason is that the data distribution is imbalanced. It is hard learned in MEDA, we visualize the emotional correlation
to guarantee the performance of low-source emotion catego- coefficients matrix. It is calculated with Pearson product-
ries. In future work, more attention will be paid to the detec- moment correlation coefficients, which indicates the level to
tion of weak and low- source emotions. In addition to which two emotions vary together:
hamming loss, the global performance of proposed method
can also be reflected by other multi-label metrics, such as Rij ¼ covðEi ; Ej Þ=sEi sEj ; (19)
micro-F1, macro-F1, and average precision, on which the
proposed method has achieved satisfying performance. Where Ei ¼ ½E1i ; E2i; . . . ENi and Eni is the emotional
intensity of emotion ei in the nth sample. covðEi ; Ej Þ is the
covariance of ei and ej , and s is the standard deviation.
5.2 Discussion of Sub-Models Figs. 2 and 3 show the comparison of the actual correlation
MEDA-FS is composed of 3 sub-models: MEDA, S-MC- coefficients matrix on Ren-CECps and the predicted correla-
ESFE and C-MC-ESFE. These sub-models are devoted to tion coefficients matrix in MEDA model. We can observe
learning information from different levels and contributing that the distribution of positively/negatively related emo-
to a more comprehensive ensemble model. To further tion pairs predicted in MEDA is similar to the real distribu-
explore the contribution of each sub-model, we further ana- tion on Ren-CECps. Taking ‘Love’ as an example. Fig. 3
lyze their performance on RenCECps in this section. The shows that in actual distribution, the most positively related
comparison results are shown in Table 4.
MEDA. As the global performance shown in Table 2.
MEDA (micro-F1 ¼ 59.71 percent, HL ¼ 0.1378) outper-
forms the previous state-of-the-art model SGM-IFC (micro-
F1 ¼ 58.60 percent, HL ¼ 0.1613), and outperforms another
two sub-models on micro-F1, AP and ranking loss. MEDA
network consists of two modules. The first is MC-ESFE,
which is a hierarchical network and extracted emotion-
TABLE 4
Comparison Results of Sub-Models on RenCECps
Micro Macro
P R F1 P R F1
S-MC-ESFE 52.16, 68.55, 59.24 42.48 56.72 47.73
C-MC-ESFE 59.34, 51.77, 55.30 43.15 32.01 34.34
MEDA 51.81, 70.46, 59.71 41.44 57.54 47.25
MEDA-FS 55.77, 66.72, 60.76 46.10 52.21 48.31
Fig. 2. Emotional correlation coefficients matrix in RenCECp.
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 483
TABLE 5
Ablation Study on RenCECps Dataset
MEDA MEDA-FS
With Without With Without
Micro F1: % 59.71 56.22 60.76 57.33
Macro F1: % 47.25 42.91 48.31 44.16
AP: % 75.76 73.22 76.51 74.11
Hamming Loss 0.1378 0.1342 0.1249 0.1313
Coverage 2.2369 2.3703 2.2226 2.3322
One Error 0.3763 0.4101 0.3618 0.3959
Ranking Loss 0.1084 0.1239 0.1062 0.1189
TABLE 6
Results Comparison of MEDA Model With Different Loss Functions
Micro % (") Macro: % (") AP: % (") HL (#) Coverage (#) OE (#) RL (#)
P R F1 P R F1
CE-loss 41.95 81.06 55.29 37.26 64.71 42.79 70.64 0.1900 2.4496 0.4598 0.1349
ML-loss 44.39 80.60 57.25 35.49 69.84 46.07 72.27 0.1745 2.3652 0.4421 0.1251
ML-FL (w ¼ 0.4) 51.81 70.46 59.71 41.44 57.54 47.25 75.76 0.1378 2.2369 0.3763 0.1084
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
DENG AND REN: MULTI-LABEL EMOTION DETECTION VIA EMOTION-SPECIFIED FEATURE EXTRACTION AND EMOTION... 485
correlations based on above features in ECorL module. In [9] A. Bandhakavi, N. Wiratunga, and D. Padmanabhan, “Lexicon
based feature extraction for emotion text classification,” Pattern
MC-ESFE module, information of each emotion reflected in Recognit. Lett., vol. 93, no. 1, pp. 133–142, Jul. 2017.
the text was separately encoded from sentence-level to con- [10] S. M. Liu and J. H. Chen, “A multi-label classification based
text-level, which contributed a lot to underlying fundamen- approach for sentiment classification,” Expert Syst. Appl., vol. 42,
tal feature extraction. In ECorL module, bidirectional-GRU no. 3, pp. 1083–1093, Feb. 2015.
[11] N. Colneri^c and J. Demsar, “Emotion recognition on twitter: Com-
network was utilized as emotion sequence predictor and parative study and training a unison model,” IEEE Trans. Affective
emotion correlation learning was implemented among emo- Comput., vol. 11, no. 3, pp. 433–446, Third Quarter 2018.
tion-specified features. MEDA-FS integrated three sub- [12] D. Pan and J. Nie, “Mutux at semeval-2018 task 1: Exploring
impacts of context information on emotion detection,” in Proc.
models derived from MEDA, and realized information 12th Int. Workshop Semantic Eval., 2018, pp. 345–349.
fusion from sentence-level, context-level, and emotion cor- [13] H. Huihui and R. Xia, “Joint binary neural network for Multi-label
relation level. Furthermore, to incorporate emotion correla- learning with applications to emotion classification,” in Proc. Int.
tion information into model training, multi-label focal loss Conf. Nature Lang. Process. Chin. Compu., 2018, pp. 250–259.
[14] P. D. Turney, “Thumbs up or thumbs down? Semantic orientation
was proposed for multi-label learning. The proposed model applied to unsupervised classification of reviews,” in Proc. Assoc.
achieved satisfactory performance and outperformed state- Comput. Linguist., 2002, pp. 417–424.
of-the-art models on both RenCECps and NLPCC2018 data- [15] F. Ren and Y. Wu, “Predicting user-topic opinions in twitter with
sets, which demonstrated the effectiveness of the proposed social and topical context,” IEEE Trans. Affective Comput., vol. 4,
no. 4, pp. 412–424, Fourth Quarter 2013.
method for multi-label emotion detection. [16] B. Pang, L. Lee, and S. Vaithyanathan, “Thumbs up? Sentiment
There is still much space for improvements in our works. classification using machine learning techniques,” in Proc. Conf.
Discernible feature representation of the weak emotion cate- Empir. Methods Natural Lang. Process., 2012, pp. 79–86.
[17] S. Mohammad and S. Kiritchenko. “Understanding emotions: A
gory is a critical problem in multi-label emotion detection dataset of tweets to study interactions between affect categories,”
task. Our proposed MC-ESFE module can prevent the fea- in Proc. 11th Int. Conf. Lang. Resour. Eval., 2018, pp. 198–209.
tures of weak emotions from being covered by strong emo- [18] N. H. Frijda, “The laws of emotion,” Amer. Psychol., vol. 43, no. 5,
tions to some extent, but not completely. In future work, we pp. 349–358, May 1988.
[19] P. Ekman, “An argument for basic emotions,” Cogn. Emotion,
will try to explore more effective methods to recognize vol. 6, no. 3-4, pp. 169–200, 1992.
weak emotions more accurately. During the emotional fea- [20] Q. Changqin and F. Ren. “Sentence emotion analysis and recogni-
ture extraction in our model, an external emotion lexicon tion based on emotion words using Ren-cecps,” Int. J. Adv. Intell.,
vol. 2, no. 1, pp. 105–117, Jul. 2010.
was severed as prior knowledge to enhance emotional fea- [21] J. Li and F. Ren, “Creating a chinese emotion lexicon based on cor-
ture representation. Abundant resources are the basis of pus Ren-cecps,” in Proc. IEEE Int. Conf. Cloud Comp. Intell. Syst.,
neural network training. In future work, more emotional 2011, pp. 80–84.
data will be incorporated for better emotion understanding. [22] C. Strapparava and A. Valitutti, “Wordnet affect: An affective
extension of wordnet,” in Proc. 4th Int. Conf. Lang. Resour. Eval.,
2004, vol. 4, pp. 1083–1086.
ACKNOWLEDGMENTS [23] S. Mohammad, “Portable features for classifying emotional text,”
in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguist.: Hum.
This work was supported in part by the Research Clusters pro- Lang. Technol., 2012, pp. 587–591.
gram of Tokushima University under Grant No. 2003002. This [24] Q. Liu and S. Li, “Word semantic similarity computation based
research has been partially supported by NSFC-Shenzhen on hownet,” Comput. Linguist. Chin. Lang. Process., vol. 7, no. 2,
Joint Foundation (Key Project) (Grant No. U1613217). pp. 59–76, 2002.
[25] S. M. Mohammad and P. D. Turney, “Crowdsourcing a word–
emotion association lexicon,” Comput. Intell., vol. 29, no. 3,
REFERENCES pp. 436–465, Sep. 2012.
[26] F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget:
[1] W. Liang, H. Xie, Y. Rao, R. Y. K. Lau, and F. L. Wang, “Universal Continual prediction with LSTM,” in Proc. 9th Int. Conf. Artif. Neu-
affective model for readers’ emotion classification over short ral Netw., 1999, pp. 850–855.
texts,” Expert Syst. Appl., vol. 114, pp. 322–333, Dec. 2018. [27] S. Poria, E. Cambria, D. Hazarika, and N. Majumder, “Context-
[2] F. Ren and K. Matsumoto, “Semi-Automatic creation of youth slang dependent sentiment analysis in user-generated videos,” in Proc.
corpus and its application to affective computing,” IEEE Trans. 55th Annu. Meeting Assoc. Comput. Linguist., 2017, pp. 873–883.
Affect. Comput., vol. vol.7, no. 2, pp. 176–189, Second Quarter 2016. [28] D. Tang, B. Qin, and T. Liu, “Aspect level sentiment classification
[3] H. Y. Shum, X. He, and D. Li, “From eliza to xiaoice: Challenges with deep memory network,” in Proc. Conf. Empir. Methods Natural
and opportunities with social chatbots,” Front. Inf. Technol. Elec- Lang. Process., 2016, pp. 214–224.
tron. Eng., vol. 19, no. 1, pp. 10–26, Jan. 2018. [29] Y. Kim, “Convolutional neural networks for sentence classi-
[4] R. Jayakrishnan, G. N. Gopal, and M. S. Santhikrishna, “Multi- fication,” in Proc. Conf. Empir. Methods Natural Lang. Process., 2014,
class emotion detection and annotation in malayalam novels,” in pp. 1746–1751.
Proc. Int. Conf. Comput. Commun. Inform., 2018, pp. 1–5. [30] T. Rao, X. Li, H. Zhang, and M. Xu, “Multi-level region-based con-
[5] T. Kowatsch, M. Nißen, C. H. I. Shih and D. R€ uegger, “Text-based volutional neural network for image emotion classification,” Neu-
healthcare chatbots supporting patient and health professional rocomputing, vol. 333, no. 14, pp. 429–439, Mar. 2019.
teams: Preliminary results of a randomized controlled trial on [31] M. S. Akhtar, D. Ghosal, and A. Ekbal, “A Multi-task ensemble
childhood obesity,” in Proc. Persuasive Embodied Agents Behav. framework for emotion, sentiment and intensity prediction,”
Chang., 2017, pp. 1–10. 2018, arXiv:1808.01216.
[6] C. Chen, R. Zhuo, and J. Ren, “Gated recurrent neural network [32] M. Chae, T. H. Kim, Y. H. Shin, J. W. Kim, and S. Y. Lee, “End-to-
with sentimental relations for sentiment classification,” Inf. Sci., end multimodal emotion and gender recognition with dynamic
vol. 502, pp. 268–278, Oct. 2019. weights of joint loss,” 2018, arXiv: 1809.00758.
[7] D. Sznycer and A. W. Lukaszewski, “The emotion–valuation [33] P. Xu, A. Madotto, C. S. Wu, J. H. Park, and P. Fung, “Emo2vec:
constellation: Multiple emotions are governed by a common Learning generalized emotion representation by multi-task train-
grammar of social valuation,” Evol. Hum. Behav., vol. 40, no. 4, ing,” in Proc. 9th Workshop Comput. Approaches Subjectivity, Senti-
pp. 395–404, Jul. 2019. ment Social Media Anal., 2018, pp. 292–298.
[8] X. Kang, F. Ren, and Y. Wu, “Exploring latent semantic informa- [34] A. Illendula and A. Sheth, “Multimodal emotion classification,” in
tion for textual emotion recognition in blog articles,” IEEE/CAA J. Proc. World Wide Web Conf., 2019, pp. 439–449.
Automatica Sinica, vol. 5, no. 1, pp. 204–216, Jan. 2018.
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.
486 IEEE TRANSACTIONS ON AFFECTIVE COMPUTING, VOL. 14, NO. 1, JANUARY-MARCH 2023
[35] L. Buitinck, J. V. Amerongen, and E. Tan, “Multi-emotion detec- [55] O. Luaces, J. Dıez, J. Barranquero, and J. D. Coz, “Binary relevance
tion in user-generated reviews,” in Proc. Eur. Conf. Inf. Retrieval, efficacy for multilabel classification,” Program. Artif. Intell., vol. 1,
2015, pp. 43–48. no. 4, pp. 303–313, 2012.
[36] Y. Kim, H. Lee, and K. Jung, “AttnConvnet at semeval-2018 task 1: [56] J. Read, B. Pfahringer, G. Holmes, and E. Frank, “Classifier
Attention-based convolutional neural networks for multi-label chains for multi-label classification,” Mach. Learn., vol. 85,
emotion classification,” in Proc. 12th Int. Workshop Semantic Eval., no. 3, pp. 333–359, 2011.
2018, pp. 141–145. [57] R. Johnson and T. Zhang, “Deep pyramid convolutional neural
[37] J. Yu, L. Marujo, J. Jiang, and P. Karuturi, “Improving multi-label networks for text categorization,” in Proc. 55th Annu. Meeting
emotion classification via sentiment classification with dual atten- Assoc. Comput. Linguist., 2017, vol. 1, pp. 562–570.
tion transfer network,” in Proc. Conf. Empir. Methods Natural Lang. [58] J. Yu, L. Marujo, J. Jiang, P. Karuturi, and W. Brendel, “Improving
Process., 2018, pp. 1097–1102. multi-label emotion classification via sentiment classification with
[38] D. Zhou, Y. Yang, and Y. He, “Relevant emotion ranking from text dual attention transfer network,” in Proc. Assoc. Comput. Linguis-
constrained with emotion relationships,” in Proc. Conf. North tics, 2018, pp. 1097–1102.
Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol., 2018, [59] W. Liao, Y. Wang, Y. Yin, X. Zhang, and P. Ma, “Improved
vol. 1, pp. 561–571. sequence generation model for multi-label classification via CNN
[39] M. L. Zhang and Z. H. Zhou, “Multi-label neural networks with and initialized fully connection,” Neurocomputing, vol. 382, no. 21,
applications to functional genomics and text categorization,” IEEE pp. 188–195, Mar. 2020.
Trans. Knowl. Data Eng., vol. 18, no. 10, pp. 1338–1351, Oct. 2006. [60] D. M. Powers, “Evaluation: From precision, recall and F-measure
[40] P. Yang, X. Sun, W. Li, S. Ma, W. Wu, and H. Wang, “SGM: to ROC, informedness, markedness and correlation,” J. Mach.
Sequence generation model for multi-label classification,” in Proc. Learn. Technol., vol. 2, no. 1, pp. 37–63, Dec. 2011.
27th Int. Conf. Comput. Linguistics, 2018, pp. 3915–3926.
[41] B. Zhao, X. Li, X. Lu, and Z. Wang, “A CNN–RNN architecture for Jiawen Deng received the double master’s degree
multi-label weather recognition,” Neurocomputing, vol. 322, no. 17, in advanced technology and science from Tokush-
pp. 47–57, Dec. 2018. ima University, Japan, and mechanical engineering
[42] S. Lian, J. Liu, R. Lu, and X. Luo, “Captured multi-label relations from Nantong University, China. She is currently
via joint deep supervised autoencoder,” Appl. Soft. Comput., working toward the PhD degree with Tokushima
vol. 74, pp. 709–728, Jan. 2019. University. Her research interests include natural
[43] C. K. Yeh, W. C. Wu, W. J. Ko, and Y. C. F. Wang, “Learning deep language processing and affective computing.
latent space for multi-label classification,” in Proc. 31th AAAI Conf.
Artif. Intell., 2017, pp. 2838–2844.
[44] K. Wang, M. Yang and W. Yang, “Deep correlation structure pre-
served label space embedding for Multi-label classification,” in
Proc. Asian Conf. Mach. Learn., 2018, vol. 95, pp. 1–16.
[45] D. A. Phan, Y. Matsumoto, and H. Shindo, “Autoencoder for Fuji Ren (Senior Member, IEEE) received the
semisupervised multiple emotion detection of conversation tran- PhD degree from the Faculty of Engineering, Hok-
scripts,” IEEE Trans. Affect. Comput., to be published, doi: kaido University, Japan, in 1991. From 1991
10.1109/TAFFC.2018.2885304. to1994, he worked at CSK as a chief researcher.
[46] Z. F. He, M. Yang, Y. Gao, H. D. Liu, and Y. Yin, “Joint multi-label In 1994, he joined the Faculty of Information Sci-
classification and label correlations with missing labels and fea- ences, Hiroshima City University, as an associate
ture selection,” Knowl. Based Syst., vol. 163, pp. 145–158, Jan. 2019. professor. Since 2001, he has been a professor of
[47] M. Rei and A. Søgaard, “Jointly learning to label sentences and the Faculty of Engineering, Tokushima University.
tokens,” in Proc. 33th AAAI Conf. Artif. Intell., 2019, pp. 6916–6923. His current research interests include natural lan-
[48] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- guage processing, artificial intelligence, affective
training of deep bidirectional transformers for language under- computing, emotional robot. He is the academi-
standing,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. cian of The Engineering Academy of Japan and EU Academy of Scien-
[49] J. Chung, C. Gulcehre, K. H. Cho, and Y. Bengio, “Empirical eval- ces. He is an editor-in-chief of the International Journal of Advanced
uation of gated recurrent neural networks on sequence mod- Intelligence, a vice president of CAAI, and a fellow of The Japan Federa-
eling,” in NIPS Workshop Deep Learn., 2014, arXiv:1412.3555. tion of Engineering Societies, a fellow of IEICE, a fellow of CAAI. He is the
[50] Z. Yang, D. Yang, C. Dyer, X. He, and A. Smola, “Hierarchical president of International Advanced Information Institute, Japan.
attention networks for document classification,” in Proc. Conf.
North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Tech-
nol., 2016, pp. 1480–1489. " For more information on this or any other computing topic,
[51] T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss please visit our Digital Library at [Link]/csdl.
for dense object detection,” in Proc. IEEE Int. Conf. Comput. Vis.,
2017, pp. 2980–2988.
[52] M. L. Zhang and Z. H. Zhou, “Multilabel neural networks with
applications to functional genomics and text categorization,” IEEE
Trans. Knowl. Data Eng., vol. 18, no. 10, pp. 1338–1351, Oct. 2006.
[53] Z. Wang, S. Li, F. Wu, Q. Sun, and G. Zhou, “Overview of NLPCC
2018 shared task 1: Emotion detection in code-switching text,” in
Proc. CCF Int. Conf. Natural Lang. Process. Chin. Comput., 2018,
pp. 429–433.
[54] S. Li, Z. Zhao, R. Hu, W. Li, T. Liu, and X. Du, “Analogical reason-
ing on chinese morphological and semantic relations,” in Proc.
56th Annu. Meeting Assoc. Comput. Linguist., 2018, pp. 138–143.
Authorized licensed use limited to: VIT University- Chennai Campus. Downloaded on July 07,2025 at 09:08:51 UTC from IEEE Xplore. Restrictions apply.