0% found this document useful (0 votes)
10 views10 pages

Mapping Emotion Categories in AI Models

Uploaded by

Medy
License
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views10 pages

Mapping Emotion Categories in AI Models

Uploaded by

Medy
License
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A Mapping on Current Classifying Categories of Emotions Used in

Multimodal Models for Emotion Recognition


Ziwei Gong Xinyi Hu Muyin Yao
Columbia University Boston University Tufts University
zg2272@[Link] xhu07@[Link] yaomuyin@[Link]

Xiaoning Zhu Julia Hirschberg


JYLlink Co., Ltd. Columbia University
zhuxiaoning@[Link] julia@[Link]

Abstract els of intensity (Plutchik, 2001) based on the com-


munication function of emotions; Barrett’s (Wilson-
In Emotion Detection within Natural Language
Processing and related multimodal research,
Mendenhall et al., 2011) biological approach stud-
the growth of datasets and models has led to ies brain responses to emotions through intepreting
a challenge: disparities in emotion classifica- EEGs and physiological changes (Hess, 2017); and
tion methods. The lack of commonly agreed emphasizing cultural influence, the construction-
upon conventions on the classification of emo- ist theory adds social and linguistic elements to
tions creates boundaries for model compar- emotion understanding (Wilson-Mendenhall et al.,
isons and dataset adaptation. In this paper, 2011). These theories are independent but some-
we compare the current classification methods
times interconnected, providing a foundation for
in recent models and datasets and propose a
valid method to combine different emotion cat- potential integration. Different theories are mostly
egories. Our proposal arises from experiments considered to be independent theories of emotion,
across models, psychological theories, and hu- yet these classification approaches are often inter-
man evaluations, and we examined the effect connected and sometimes built upon each other,
of proposed mapping on models. providing a basis to connect them. However, few
studies explore ways to connect or combine these
1 Introduction
different categorizations.
Emotion recognition, as an essential ability for In computer science, researchers face the chal-
good interpersonal relations (Mancini et al., 2018), lenge of choosing an emotion theory when build-
has long been a major subject in psychology, and ing datasets for emotion detection. Recent work
for the last two decades has received increasing in emotion classification has shifted towards using
attention from the field of computer science, espe- multimodal data sources like audio, video, and text
cially artificial intelligence (De Silva et al., 1997; (Poria et al., 2019; Shen et al., 2020), and some
Gong et al., 2023). Yet in this process a divergence even explore incorporating additional factors like
has emerged from newly published datasets and personality and social connections to leverage more
models — the misalignment between different cat- information for deep learning models (Kahou et al.,
egories of emotions. To resolve such disparity be- 2015). Due to varying annotation methods and
tween emotion datasets, we propose a psychology- mismatch in the set of labels, a model typically
based solution for computer scientists to solve the selects a single dataset for experiments, although
problem of misalignment in emotion classification more data could improve its performance. A sig-
datasets, which is caused by the independent nature nificant issue arises from the lack of alignment in
of emotion classification theories. labeling schemas across datasets, making it chal-
In the field of psychology, there are many differ- lenging for models to leverage multiple datasets
ent theories on how to classify emotions focus- in supervised learning (Bostan and Klinger, 2018).
ing on different aspects. Various theories clas- This disparity results in a lack of cohesion in the
sify emotions based on different factors: Ekman’s literature, hinders direct performance comparisons,
theory focuses on universal facial expressions and complicates dataset combination and training.
(Ekman, 1992), comparing the facial expressions Since annotating such datasets is costly and time-
of westerners and Aboriginal residents of New consuming, a mapping method that can unify ex-
Guinea; Plutchik’s evolutionary perspective catego- isting datasets could benefit the community. Cur-
rizes emotions into 8 primary emotions with 3 lev- rently, little research in both psychology and com-
19
Proceedings of The 18th Linguistic Annotation Workshop (LAW-XVIII), pages 19–28
March 22, 2024 ©2024 Association for Computational Linguistics
puter science explores the relationship between dif- tion of that emotion. So it it possible that in the
ferent emotion categories. While there are studies annotation process annotators sometimes use com-
mapping categorical emotions onto dimensional mon sense understanding of emotions to annotate
models (Hoffmann et al., 2012) and recent work and only use the definitions provided as references.
inproviding more grounded emotion categories in Given these considerations and results, we decide
Dutch (De Bruyne et al., 2020), the mapping be- not to modify emotions common to both categories.
tween multiple categorical emotions, which cre- Higher-Level Emotions: Emotions exclusive to
ates misalignment in emotion datasets for machine higher-level categories, such as anticipation and
learning, remains unstudied. surprise, are mapped based on past literature, often
This paper aims to establish a valid mapping of considering valence and arousal of various emo-
emotion categories based on psychological theories tions. Valence measures the positiveness or neg-
and validated through machine learning models. ativity of an emotional stimulus (De Silva et al.,
We select the five most commonly used emotion 1997), and emotions with similar valence are pre-
classification methods in large emotion datasets, sumed to be more closely related. Arousal level,
propose a valid mapping method rooted in psycho- measuring the intensity of emotion, is also a cue to
logical theory, verify it through human evaluation, the similarity of emotions. Emotions with compa-
and assess its impact on emotion recognition mod- rable arousal and valence levels are more likely to
els. Our mapping method is an initial effort to be paired, contrasting with emotions that differ in
create a continuous mapping approach connecting these aspects.
these discrete emotion classification methods. Human Evaluations: When faced with tied
choices, we conduct human evaluations on each
2 Methods theory to determine the best mapping choice in the
situation of a tie. Detailed evaluations are carried
2.1 Datasets
out for each theory. We illustrate our mapping
We choose 4 diverse datasets, each employing dis- choice for the emotion "surprise" as an example of
tinct modalities and emotion classification methods. our decision-making process.
We include both datasets that reflect real-life scenar-
ios such as MEmoR (Shen et al., 2020) and MELD 2.3 The Classification for Surprise as
(Poria et al., 2019), and those focusing on facial Example
features like IEMOCAP (Busso et al., 2008). Ad- Surprise characterizes the feeling of shock due to
ditionally, we include the FER-2013 (Goodfellow perceiving things or experience out of expectation.
et al., 2013) computer vision dataset to investigate To map surprise onto a 6-emotion classification
our mapping method’s impact on a single-modality (neutral, sadness, joy, disgust, anger, and fear), we
dataset. These datasets span various classification employed a bipolar model integrating valence and
methods: MEmoR employs Plutchik’s Wheel of arousal dimensions. Russell introduced this model
Emotion (14 emotions), MELD and IEMOCAP in 1977 (Russell and Mehrabian, 1977), with mo-
adopt Ekman’s basic 6 emotions, and FER-2013 tivation as an initial component. Surprise may be
features 7 common emotions as labels. considered a negative emotion, since previous stud-
ies associate surprise with a negative valence (No-
2.2 Mapping Method ordewier and Breugelmans, 2013) and high arousal
Our approach to developing a mapping theory be- levels (Russell and Mehrabian, 1977). Based on
tween emotion classification methods follows the Liu et al.’s research, high-arousal, low-valence
following procedure. emotions are akin to anger (Liu et al., 2010). How-
Common Emotions: Emotions shared by both ever, the potential for positive valence-associated
categories remain unaltered. Although these emo- surprise introduces ambiguity in conversion, possi-
tions might have different definitions across theo- bly favoring mapping to neutral.
ries, our sample annotation process suggests anno- We leverage biological distinctions between
tators seldom find them non-transferable. Consid- emotions as a reference. A recent study utilizing
ering the annotation process of large datasets, it is biomarkers to analyze EEG profiles across brain
common that their annotators are asked to choose regions offers valuable findings. Among surprise-
an emotion that best describes the current scene or combined emotions, the spectral biomarker’s mean
utterance rather than strictly following the defini- differences (0.114) and the temporal biomarker’s
20
14 fine- 6 emo- 3 senti- 3 Mapping Between Different Emotion
9 primary 7 basic
grained tions ments
anticipation Categories
interest anticipation
neutral neutral neutral 3.1 Mapping Results
neutral neutral
fear fear fear fear
disgust Table 1 shows the resulting unique mapping ta-
boredom disgust disgust disgust ble between the 5 most popular emotion classifica-
sadness sadness sadness sadness tion methods, ranging from 14 categories to 3 cate-
anger negative
annoyance anger anger gories. To validate our mapping, a re-annotation of
surprise anger randomly sampled emotions mapped to their cate-
distraction surprise surprise gories achieves an accuracy of 0.96 (Annotator 1)
joy
serenity joy and 0.917 (Annotator 2), with a fair inter-annotator
joy joy positive
trust trust agreement of 0.318 (Cohen’s Kappa). Thus, this
mapping method has proved to have fairly high ac-
Table 1: Mapping results. This table demonstrates how
14 fine-grained emotions, listed on the leftmost column,
curacy when used to reconstruct datasets. We con-
are mapped onto 9 primary emotions, Ekman’s basic clude that it is possible to map emotion categories
emotions, 6 emotions, and the 3 sentiments. onto each other with relatively high accuracy. The
proposed mapping method is one directional, from
more categories to fewer categories. Mapping data
from fewer categories to more categories is possible
mean differences (0.058) are lowest for the neutral- but requires additional annotation to determine the
surprise pairing (Mancini et al., 2018). Hence, both resulting co-domain labels. Additionally, this map-
anger and neutral are considered possible mappings ping method can be used by future researchers with
for surprise. To test this hypothesis, we imple- more fine grained labeling methods when creating
mented a program to convert surprise into anger datasets, since mapping from more fine grained
and neutral. These converted emotions were mixed labeling to less fine grained labeling requires no
with randomly selected samples of other emotions. additional information.
Annotators, at least two per data point, participated
3.2 Map analysis
in the evaluation. All annotators were English-
speaking college students, with half of them fa- The main contribution of our work is that we are
miliar with the TV show "The Big Bang Theory." the first to propose a mapping method for numer-
Annotation materials included clips, scripts, and ous emotion categorization methods from psycho-
emotion definitions per category. Evaluation re- logical theories and have validated it with human
sults favored the surprise-to-anger conversion, as it evaluation and experiments. Analyzing the final
achieved higher accuracy. Hence, we chose to map mapping produced, we found that across all cate-
surprise to anger based on annotation outcomes. gorization methods, the categories in negative emo-
tions are more fine-grained than either positive or
neutral emotions, given the number of emotions
that are mapped into negative emotions. For exam-
2.4 The Annotation Process ple, from the 14-categories, there are 8 emotions
that were mapped into “negative”, 3 mapped into
“positive” and 3 mapped into “neutral”. This imbal-
At least two annotators are asked to annotate one ance could be caused by both biases in the dataset
data point. All annotators are college students and underlying psychological mechanisms. Since
studying in a university where English is the first the data for the datasets are collected from TV
language, since the datasets are all in English. The shows or other commercialized media, it could be
students age between 18 to 22. The annotators are that a dataset may not necessarily contain emotion
provided with clips and scripts during the annota- proportions that are reflective of actual human emo-
tion, and half of the annotators are familiar with tional expressions. The underlying psychological
the TV show, the Big Bang Theory. The emotions mechanisms would also be an aspect to discuss for
and definitions of each emotion in each category other researchers.
are also provided to the annotators to help interpre- Moreover, while several emotions seem more
tation. difficult to be mapped into other categories, such
21
as surprise and trust, in the experiment we found Emotion Category 3 6 7 9 14
MEmoR Accuracy 0.924 0.867 0.884 0.869 0.864
it still has an acceptable evaluation score. For ex- CNN Accuracy 81.78 65.39 65.28 - -
ample, it is difficult to determine whether ‘surprise’
is a good surprise or a bad surprise in real life, but Table 2: Experimental results from the MEmoR model
in our mapping, ‘surprise’ is mapped into anger and the CNN model. This table shows the overall ac-
with a high agreement in human evaluation. One curacy of the models trained and tested on datasets re-
constructed based on each 3 classification method. The
possible reason for this is that the current cate-
highest achieved is bolded. The MEmoR model uses
gories make humans, the annotators, more likely to visual, audio, textual features. In the CNN model, only
choose negative surprise as “surprise” and consider visual information is used.
taking positive surprise as “joy” or “hopeful”. We
attribute this alignment to the disparity among emo-
tion classification theories and their unique aspects
in understanding human emotions. Nevertheless,
our mapping method establishes a consistent stan-
dard grounded in existing datasets.
Although the same emotion categories may
have different definitions for different classification
methods, each of the emotions are still mapped into
the corresponding emotion with the same name Figure 1: Contrast in attention heat maps across 9 ran-
dom images: a CNN model trained on a 7-category
in our mapping. Although we acknowledge the
dataset (left) vs. the same dataset categorized into 3
slight difference in meaning, for the purpose of groups (right). Regions of high attention are shown in
mapping, emotions still prove to be more simi- red.
lar to corresponding emotion with the same name
despite the different interpretations. Our current
and 3 sentiments respectively.
mapping method sucessfully proposes a uniform
Multimodality MEmoR Model (Shen et al.,
standard, yet its accuracy is limited in datasets that
2020) is a fusion multi-modal model is provided
are largely different from the existing datasets in
by (Shen et al., 2020). The model extracts repre-
terms of domain, conversation style, etc. Further-
sentative multimodal features, including audio fea-
more, since we are the first to propose a mapping
tures, video features, and text features, personality
for different emotion classification theories from a
features, and uses an attention-based multimodal
psychological perspective, there are a limited num-
reasoning method. In the experiment we use the
ber of existing studies that we could compare to.
MEmoR dataset reconstructed based on our map-
We hope our proposal, as a first attempt to solve
ping, which has 5 groups of labels. Each model
this disparity, could also serve as a start point for
will be trained tested on each classification method.
others who seek to solve the problem.
4.2 Results
4 Mapping effects on ML Models
Results of the experiments on the MEmoR model
To analysis the effect of the proposed mapping on and CNN model are shown in Table 2. From these
machine learning models, we set up an experiment experiments, we have found that models generally
to check the accuracy of emotion re-categorization perform better when there are fewer emotion cat-
after applying the mapping method in Table 1 to egories, meaning that more fine-grained emotions
both the MEmoR and the CNN dataset. We se- are more difficult for models to differentiate, re-
lected two models to study the effect of the map- gardless of which modality or which combination
ping methods on emotion detection models. of modalities is used. This finding validates that
our mapping is accurate, as it is the general un-
4.1 Models
derstanding in the machine learning community
Vision CNN is commonly used in recognition and that using fewer classification categories, when cor-
classification tasks (Albawi et al., 2017; Suryani rectly applied, leads to higher accuracy since the
et al., 2016). We reconstructed the FER-2013 complexity of the task is reduced. However, the ex-
Dataset (7 basic emotions) based on our mapping perimental results for the MEmoR model show that
in Table 1 to recreated the dataset with 6 emotions training and testing on 7 categories does achieve
22
Figure 2: Confusion matrices generated by three CNN models trained on a dataset, all learning from the same set of
pictures but with labels categorized into 7 (left), 6 (middle) and 3 categories (left). Columns represent the predicted
label and rows represent the true label.

higher accuracy than 6 categories, while still lower vide a study of the effect of using different emo-
than results on 3 categories. However, on the CNN tion classification methods when training models.
model, we see a higher accuracy on 6 categories We are the first group of researchers attempting
compared to the 7 categories. By looking closely at to bridge the different psychological emotion the-
the confusion matrices (Figure 2) of CNN models, ories and lend them consistency in the computer
we see that the improvement was mainly on the ad- science world. Moreover, using our mapping al-
justed category, and the accuracy of the categories lows researchers to obtain a larger and more flexi-
that remain untouched from the transition remains ble dataset for training and testing and to analyze
in the same range. A possible reason for this is the model’s ability to differentiate emotions using
that classifying emotion into 7 categories is derived different emotion categories, as well as identify the
from Ekman’s basic emotion theory, which is based best model across all datasets.
on facial expression. Thus it is possible that such a
categorization method is easier for models to learn Acknowledgements
through facial expression recognition. However, to We thank [Link] Brotherton, Bingpu Zhao, and
determine the cause, there should be more research Jingyi Song for their support and valuable feedback.
on separated models and modalities. We encourage This research is supported in part by the Defense
future researchers to look into this question. Advanced Research Projects Agency (DARPA),
Visualization of the CNN model’s attention is via the The Computational Cultural Understanding
shown in Table 1. We observe that the attention of (CCU) program contract HR001122C0034. The
the model trained with more fine-grained emotions views, opinions and/or findings expressed are those
is more spread out through the face, with some of the author and should not be interpreted as rep-
stress around the eye and mouth area. In com- resenting the official views or policies of the De-
parison, the attention of the model trained on senti- partment of Defense or the U.S. Government.
ments is more focused on specific areas and created
red dots on the heat map. The difference indicates
that there are more subtle cues to distinguish fine-
grained emotion on the face, requiring the model
to learn to predict based on more information from
different areas, compared to sentiments that are
simpler and distinguishable through some key area
like the mouth (smiling or not, for example).

5 Conclusion

In this paper, we propose the first complete map-


ping that connects different emotion categories for
multimodal emotion recognition studies, and pro-
23
Limitations
A limitation of our mapping is that it proposes a
unified standard within a set range of 3 to 14 cat-
egories. Yet for some particular tasks, creating a
recognizer that is sensitive to a particular facial
expression or emotion that is not included in our
proposed method may be necessary. We encourage
future researchers to expand on top of our classi-
fication method using similar methods. However,
we hope that providing a unified standard would
benefit the community by decreasing deviance and
making it easier for scholars who wish to adopt an
existing dataset for a particular task.
Moreover, while several emotions seem harder to
be mapped into other categories, we found accept-
able evaluation score for the mapping, but there are
limitations. Similarly to the mapping of "surprise",
whether the emotion “trust” was a neutral emotion
or a positive emotion is hard to decide. In our clas-
sification, we followed the steps described in our
“Methods” section to determine which classifica-
tion gives better accuracy and thus determines the
mapping. Although our current mapping method
proposes a uniform standard, its accuracy is limited
in datasets that are largely different from the exist-
ing datasets in terms of domain, conversation style,
etc. we also acknowledge potential difficulties in
mapping certain emotions, and we anticipate revi-
sions and improvements to our current mapping
method after the construction of larger datasets in
the future to better bridge the differences between
various data sets.
Furthermore, since we are the first to propose
a mapping for different emotion classification the-
ories from a psychological perspective, there is a
limited number of existing studies that we could
compare to. We hope our proposal, as a first at-
tempt to solve this disparity, could also serve as a
start point for others who seek to solve the problem.

24
References Holger Hoffmann, Andreas Scheck, Timo Schuster, Stef-
fen Walter, Kerstin Limbrecht, Harald C. Traue, and
Saad Albawi, Tareq Abed Mohammed, and Saad Al- Henrik Kessler. 2012. Mapping discrete emotions
Zawi. 2017. Understanding of a convolutional neural into the dimensional space: An empirical approach.
network. In 2017 International Conference on Engi- In 2012 IEEE International Conference on Systems,
neering and Technology (ICET), pages 1–6. Man, and Cybernetics (SMC), pages 3316–3320.
Laura-Ana-Maria Bostan and Roman Klinger. 2018. Samira Ebrahimi Kahou, Xavier Bouthillier, Pas-
An analysis of annotated corpora for emotion clas- cal Lamblin, Caglar Gulcehre, Vincent Michalski,
sification in text. In Proceedings of the 27th Inter- Kishore Konda, Sébastien Jean, Pierre Froumenty,
national Conference on Computational Linguistics, Yann Dauphin, Nicolas Boulanger-Lewandowski,
pages 2104–2119, Santa Fe, New Mexico, USA. As- Raul Chandias Ferrari, Mehdi Mirza, David Warde-
sociation for Computational Linguistics. Farley, Aaron Courville, Pascal Vincent, Roland
Memisevic, Christopher Pal, and Yoshua Bengio.
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe 2015. Emonets: Multimodal deep learning ap-
Kazemzadeh, Emily Mower Provost, Samuel Kim, proaches for emotion recognition in video.
Jeannette Chang, Sungbok Lee, and Shrikanth
Yisi Liu, Olga Sourina, and Minh Khoa Nguyen. 2010.
Narayanan. 2008. Iemocap: Interactive emotional
Real-time eeg-based human emotion recognition and
dyadic motion capture database. Language Re-
visualization. In 2010 International Conference on
sources and Evaluation, 42:335–359.
Cyberworlds, pages 262–269.
Luna De Bruyne, Orphee De Clercq, and Veronique Giacomo Mancini, Roberta Biolcati, Sergio Agnoli, Fed-
Hoste. 2020. An emotional mess! deciding on a erica Andrei, and Elena Trombini. 2018. Recognition
framework for building a Dutch emotion-annotated of facial emotional expressions among italian pre-
corpus. In Proceedings of the Twelfth Language adolescents, and their affective reactions. Frontiers
Resources and Evaluation Conference, pages 1643– in psychology, 9:1303.
1651, Marseille, France. European Language Re-
sources Association. Marret K. Noordewier and Seger M. Breugelmans. 2013.
On the valence of surprise. Cognition and Emotion,
L.C. De Silva, T. Miyasato, and R. Nakatsu. 1997. Fa- 27(7):1326–1334. PMID: 23560688.
cial emotion recognition using multi-modal informa- Robert Plutchik. 2001. The nature of emotions: Human
tion. In Proceedings of ICICS, 1997 International emotions have deep evolutionary roots, a fact that
Conference on Information, Communications and may explain their complexity and provide tools for
Signal Processing. Theme: Trends in Information clinical practice. American Scientist, 89(4):344–350.
Systems Engineering and Wireless Multimedia Com-
munications (Cat., volume 1, pages 397–401 vol.1. Soujanya Poria, Devamanyu Hazarika, Navonil Ma-
jumder, Gautam Naik, Erik Cambria, and Rada Mi-
Paul Ekman. 1992. Are there basic emotions? Psycho- halcea. 2019. MELD: A multimodal multi-party
logical review, 99 (3). dataset for emotion recognition in conversations. In
Proceedings of the 57th Annual Meeting of the As-
Ziwei Gong, Qingkai Min, and Yue Zhang. 2023. Elic- sociation for Computational Linguistics, pages 527–
iting rich positive emotions in dialogue generation. 536, Florence, Italy. Association for Computational
In Proceedings of the First Workshop on Social In- Linguistics.
fluence in Conversations (SICon 2023), pages 1–8,
Toronto, Canada. Association for Computational Lin- James A Russell and Albert Mehrabian. 1977. Evidence
guistics. for a three-factor theory of emotions. Journal of
Research in Personality, 11(3):273–294.
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Guangyao Shen, Xin Wang, Xuguang Duan, Hongzhi
Aaron Courville, Mehdi Mirza, Ben Hamner, Will Li, and Wenwu Zhu. 2020. Memor: A dataset for
Cukierski, Yichuan Tang, David Thaler, Dong-Hyun multimodal emotion reasoning in videos. In Proceed-
Lee, Yingbo Zhou, Chetan Ramaiah, Fangxiang Feng, ings of the 28th ACM International Conference on
Ruifan Li, Xiaojie Wang, Dimitris Athanasakis, John Multimedia, MM ’20, page 493–502, New York, NY,
Shawe-Taylor, Maxim Milakov, John Park, Radu USA. Association for Computing Machinery.
Ionescu, Marius Popescu, Cristian Grozea, James
Bergstra, Jingjing Xie, Lukasz Romaszko, Bing Xu, Dewi Suryani, Patrick Doetsch, and Hermann Ney. 2016.
Zhang Chuang, and Yoshua Bengio. 2013. Chal- On the benefits of convolutional neural network com-
lenges in representation learning: A report on three binations in offline handwriting recognition. In 2016
machine learning contests. 15th International Conference on Frontiers in Hand-
writing Recognition (ICFHR), pages 193–198.
Ursula Hess. 2017. Chapter 5 - emotion categorization.
Christine D. Wilson-Mendenhall, Lisa Feldman Bar-
In Henri Cohen and Claire Lefebvre, editors, Hand-
rett, W. Kyle Simmons, and Lawrence W. Barsalou.
book of Categorization in Cognitive Science (Second
2011. Grounding emotion in situated conceptualiza-
Edition), second edition edition, pages 107–126. El-
tion. Neuropsychologia, 49(5):1105–1127.
sevier, San Diego.
25
A Appendix
A.1 Experiment Design for CNN Model
To explore the effects of our mapping method on
CNN models, we built a simple CNN model with
three convolutional layers, feeds into a fully con-
nected layer, and outputs from a softmax layer. The
model is trained on unimodal (visual) information
on the FER-2013 (Goodfellow et al., 2013) dataset
for emotion classification. The CNN model is se-
lected to study the effect of the mapping methods
on unimodal [Link] model was trained us-
ing batch size=256 for 60 epoches on single GPU.
We reconstructed the FER-2013 (Goodfellow et al.,
2013) Dataset based on our mapping. Since the
dataset is originally classified labeled with 7 basic
emotions, we recreated the dataset with 6 emotions
and 3 sentiments classification methods respec-
tively (Table 3 (Appendix)). The mapping method
is shown in Figure 3 (Appendix). Each CNN model
will be tested on all 3 classification methods using
the same hyper-parameter and trained for 60 epochs
in two stages on the same hardware. All three mod-
els are trained to convergence before stopping at
epochs 60.

A.2 Experiment Design for MEmoR Model


MEmoR Model (Shen et al., 2020) is a fusion
multi-modal model is provided by (Shen et al.,
2020). The model extracts representative multi-
modal features, including audio features, video
features, and text features, personality features,
and uses an attention-based multimodal reasoning
method. The experiment use the MEmoR (Shen
et al., 2020) dataset reconstructed based on our
mapping. The reconstructed dataset has 5 groups
of labels, following the 5 most popular emotion
classification theories. Each model will be tested
on all 5 classification methods and each modality
(visual, textual, audio) in order to explore the ef-
fect of our mapping on models. For simplicity, we
choose the default parameters and model structure
given in the MEmoR model, except to revise the
model to fit the change in the size of the label. All 5
classification methods experimented with are listed
in Table 3 (Appendix). The mapping method is
shown in Figure 3 (Appendix).

26
14 fine-grained emotions 9 primary emotions 7 basic emotions 6 emotions 3 sentiments
joy,

anger,

disgust,

sadness, joy,
joy,
surprise, anger, joy,
anger,
fear, disgust, anger,
disgust, positive,
anticipation, sadness, disgust,
sadness, negative
trust, surprise, sadness,
fear, neutral
serenity, fear, fear,
surprise,
interest, anticipation, trust, neutral
neutral
annoyance, neutral

boredom,

distraction,

neutral

Table 3: Emotion Categories

27
Figure 3: Mapping method in graph. This graph demonstrates how 14 fine-grained emotions, listed on the leftmost
column, are mapped onto 9 primary emotions, Ekman’s basic emotions, 6 emotions, and the 3 sentiments.

28

You might also like