Impact of Emotion Models on Text Detection
ABSTRACT Text emotion detection is a pivotal aspect of natural language processing, with wide-ranging
applications involving human-computer interactions. Machine learning agents have been trained with
supervised methods, thus relying on labeled datasets. However, the arbitrary selection of emotion models
while labeling such datasets poses significant challenges in the performance and generalizability of the
produced machine learning predictors, primarily when evaluated against unseen data, as it effectively
introduces bias to the process. This study investigates the impact of emotion model selection on the
efficacy of machine learning systems for text emotion detection. Eight labeled datasets were employed
to train linear regression, feedforward neural network, and BERT-based deep learning models. Results
demonstrated a notable decrease in accuracy when models trained on one dataset were tested on others,
underscoring the inherent incompatibilities in labeling across datasets. To prove that the emotion model
significantly impacts predictors’ performance, we propose a standardized emotion label mapping utilizing
James Russell’s circumplex model of affect that turns the emotion model into a parameter rather than a
fixed element. Cross-dataset testing with this shared emotion mapping yielded significant, non-negligible
changes in accuracy (both improvement and degradation). This fact highlights the impact of the emotion
model (traditionally arbitrarily selected) during machine learning training and performance, arguing that
improvements in accuracy reported in related research literature might be due to differences in the used
emotion model rather than the new algorithms introduced.
INDEX TERMS Affective computing, natural language processing, sentiment analysis, text emotion
detection, text emotion recognition.
2024 The Authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.
VOLUME 12, 2024 For more information, see [Link] 70489
A. D. L. Languré, M. Zareei: Evaluating the Effect of Emotion Models
Furthermore, current research often introduces new arousal (from low to high) [10]. Klaus Scherer’s three-
machine learning algorithms purportedly enhancing accuracy dimensional model of emotion incorporates valence, arousal,
but typically sidesteps cross-dataset testing. Consequently, and a third dimension denoting the perceived degree of
it can be argued that these studies only demonstrate the new control over emotion [11]. James Gross’s component process
algorithm’s heightened overfitting to the used dataset rather model proposes that emotions stem from basic component
than validating its generalization capabilities. processes, including affective responses, cognitive appraisals,
This research aims to highlight the influence of selecting and physiological changes [12]. Influenced by contextual
an arbitrary emotion model on machine learning algorithms’ factors, the appraisal component shapes individuals’ emo-
performance and generalization capabilities in text emotion tional experiences based on their interactions with their
detection. We propose a novel emotion labeling approach environment and past encounters. Additionally, the pleasure-
centered around a universally recognized framework such as arousal-dominance (PAD) model, developed by Mehrabian
James Russell’s circumplex model of affect. We demonstrate and Russell, augments the circumplex model by introducing
significant changes in ML models’ performance by manip- a third dimension, dominance, representing the perceived
ulating the emotion model while keeping all other elements degree of control or power associated with an emotion [13].
the same. This underscores the novelty of our approach and
the importance of considering emotion models as variable
B. EMOTIONS AS NUMERICAL VALUES
hyperparameters, an element often overlooked in current
In Russell’s circumplex model of affect, valence represents an
literature.
emotion’s positive or negative nature, ranging from pleasant
to unpleasant. Emotions with positive valence are typically
II. PRELIMINARIES associated with happiness, joy, and contentment, while those
A. SENTIMENT AND EMOTIONS IN TEXT with negative valence include sadness, anger, and fear.
Two fundamental domains within NLP for text encom- Arousal, on the other hand, reflects an emotion’s intensity or
pass sentiment analysis (SA) and text emotion detection activation level, ranging from low to high. Emotions with low
arousal are calm and relaxed, while those with high arousal
(TED) [5]. SA focuses on discerning whether data conveys
are intense and stimulating.
positivity, negativity, or neutrality, which can be evaluated
This model represents emotions as points in a two-
along a single dimension. In contrast, TED delves into iden-
dimensional space defined by valence and arousal axes. They
tifying a spectrum of human emotion categories, including
can, therefore, be assigned numerical pairs of values analo-
but not limited to anger, happiness, or sadness, necessitating a
gous to cartesian coordinates. This representation allows for a
more intricate analytical approach. Central to TED is using an
comprehensive understanding of various emotional states by
emotion model (EM) as a theoretical framework or guideline
categorizing them based on their positions within this space.
for understanding and classifying emotions within textual
For example, emotions like excitement and euphoria would
data. An EM encompasses a structured representation of
be located in the high arousal, positive valence quadrant,
emotion categories, facilitating the classification of emotions
while emotions like depression and fatigue would be in
in text data. By leveraging predefined emotion categories, the
the low arousal, negative valence quadrant. By mapping
EM enables the accurate identification of emotional states
emotions this way, Russell’s model provides a structured
conveyed by the text, thereby enhancing the interpretability
framework for analyzing and interpreting emotional experi-
and utility of TED systems.
ences, facilitating research in psychology, neuroscience, and
EMs can be divided into two primary types: categorical
TED, among others.
and dimensional [6]. Categorical models offer discrete
classifications or labels for emotions. For instance, Robert
Plutchik’s wheel of emotions delineates eight primary C. MACHINE LEARNING
emotions (joy, trust, fear, surprise, sadness, disgust, anger, Machine learning (ML) is a subfield of artificial intelli-
and anticipation), which can be further combined to generate gence (AI) that develops algorithms and models to enable
secondary emotions [7]. Similarly, Carroll Izard’s differential computers to learn from data, make predictions, and make
emotions theory posits ten basic emotions, including interest, decisions [14]. Unlike traditional rule-based programming,
joy, surprise, anger, sadness, disgust, contempt, shame, guilt, where explicit instructions are provided for solving problems,
and fear [8]. Meanwhile, Paul Ekman’s six basic emotions ML systems learn patterns and relationships directly from
model identifies anger, disgust, fear, happiness, sadness, and data by identifying underlying structures and adjusting their
surprise [9]. parameters.
In contrast, dimensional EMs do not lend themselves ML algorithms can be broadly categorized into super-
to discrete representation. Instead, they are positioned vised, unsupervised, and reinforcement learning paradigms.
along a continuum defined by axes. Dimensional EMs, In supervised learning, the algorithm is trained on labeled
such as the affective circumplex model by James Russell, data, where each example is associated with a target
conceptualize emotions as points on a two-dimensional space output. Unsupervised learning involves discovering patterns
defined by valence (ranging from positive to negative) and or structures in unlabeled data. In contrast, reinforcement
learning relies on learning through trial and error, with the neurons, with information flowing strictly in one direction,
algorithm receiving feedback as rewards or penalties based from the input layer through one or more hidden layers to the
on its actions. output layer [16].
One of ML’s critical strengths is its ability to generalize In the context of TED, an FNN is employed to learn the
from training data to make predictions or decisions on complex mappings between textual features extracted from
new, unseen data. This generalization capability enables ML the input data and the corresponding emotion labels. The
models to adapt and perform well in diverse and complex architecture of an FNN typically includes an input layer, one
scenarios. or more hidden layers, and an output layer. Each neuron
One application of ML is TED, where algorithms are in the input layer represents a feature of the input data,
trained to recognize and classify emotions conveyed in text, such as word frequencies or TF-IDF scores. In contrast,
such as those found in social media networks, internet neurons in the hidden layers perform computations and
blogs, review sites, or online newspapers. In this context, extract higher-level representations of the input data. The
ML leverages an EM, which serves as the framework for output layer produces predictions or classifications based on
categorizing emotions within the text. The process typically the learned representations, with each neuron corresponding
involves training the ML model using labeled datasets, where to a different emotion category.
each text sample is associated with one or more emotion The training process of an FNN involves iteratively
labels based on the chosen EM. During training, the algorithm adjusting the weights and biases of the network’s con-
learns to identify patterns and relationships between the nections to minimize a predefined loss function, typically
features extracted from the text data and the corresponding using an optimization algorithm such as stochastic gradient
emotion labels. This process often involves adjusting the descent (SGD) or Adam. A FNN can be trained using
model’s parameters to optimize its performance in predicting a backpropagation algorithm, where gradients of the loss
the correct emotion labels for new, unseen text samples. function concerning the network parameters are computed
Once the model is trained, it undergoes testing and and used to update the parameters in the opposite direction
validation to assess its performance and generalization of the gradient. During training, the FNN learns to map
capabilities. Testing involves evaluating the model’s accuracy the input text features to the correct emotion labels,
and effectiveness in classifying emotions on a separate enabling it to make accurate predictions on new, unseen text
dataset not used during training. Validation ensures the model samples.
performs reliably across different datasets and scenarios,
helping identify and address potential biases or limitations.
The present research will discuss three ML algorithms: 3) BERT
linear regression, simple feedforward neural networks, and In TED, using BERT (Bidirectional Encoder Representa-
BERT. tions from Transformers) leverages pre-trained transformer-
based language models. BERT, a deep learning model
developed by Google, captures contextual information from
1) LINEAR REGRESSION
textual data by considering both left and right contexts
In TED, linear regression is a predictive modeling technique
bidirectionally. Unlike traditional models, BERT does not
to infer the emotional states associated with textual data [15].
require handcrafted features or prior feature engineering,
It establishes a linear relationship between the features
as it learns contextual representations directly from the text.
extracted from the text, such as word frequencies or TF-IDF
BERT models can be fine-tuned for sequence classification
scores, and the corresponding emotion labels. Through
tasks, where they learn to predict emotion labels from input
the learning process, the linear regression model attempts
text data [17]. The training begins with initializing the
to discern underlying patterns and associations within the
BERT tokenizer and model, utilizing a pre-trained model
textual features to predict the emotional responses conveyed
for text tokenization and sequence classification. Texts and
by the text. Upon completing the training phase, the model
corresponding emotion labels are extracted from the dataset,
can generate predictions for new text samples based on their
and labels are encoded for numerical representation. Texts
feature representations, assigning probabilities to different
are then tokenized and encoded using the BERT tokenizer,
emotion categories.
incorporating special tokens and truncation/padding to ensure
uniform input length.
2) FEEDFORWARD NEURAL NETWORKS During the training loop, a BERT model can be improved
A feedforward neural network (FNN) is a fundamental type of using the Adam optimizer with a cross-entropy loss function
artificial neural network (ANN) commonly used for various to minimize the discrepancy between predicted and actual
ML tasks, including TED. Unlike other neural network emotion labels. The model is trained over multiple epochs,
architectures, such as recurrent neural networks (RNNs) with gradients calculated and updated using backpropagation.
or convolutional neural networks (CNNs), FNNs do not Evaluation of the validation set involves making predictions
have feedback loops or cyclic connections between neurons. using the trained model and computing evaluation metrics
Instead, they consist of multiple layers of interconnected such as accuracy, precision, recall, and F1 score.
D. PYTHON IMPLEMENTATION essential for obtaining the best possible performance from an
PyTorch and scikit-learn are two prominent Python libraries ML model.
widely used in ML and data science. PyTorch is an In linear regression, common hyperparameters include
open-source ML framework primarily known for its flexi- the learning rate (for gradient descent-based optimization
bility and dynamic computational graph construction [18]. algorithms), the regularization parameter (e.g., L1 or L2
It provides a platform for building and training NNs, offering regularization strength), and the choice of optimization
extensive support for deep learning tasks such as image algorithm. In FNN, hyperparameters include the number of
classification, NLP, and reinforcement learning. PyTorch’s hidden layers, the number of neurons in each layer, the
popularity stems from its intuitive interface, which allows activation function used in each layer, the learning rate,
researchers and practitioners to experiment with complex and the batch size. For BERT-based classifiers or other
models and algorithms efficiently. transformer models, hyperparameters include the learning
On the other hand, scikit-learn is a comprehensive rate, the number of layers, the number of attention heads, the
ML library built on top of Python’s scientific computing sequence length, the batch size, and the choice of pretraining
stack [19], including NumPy, SciPy, and Matplotlib. Unlike checkpoint.
PyTorch, scikit-learn focuses on traditional ML algorithms,
providing implementations for various supervised and
unsupervised learning techniques such as classification, III. EXPERIMENTS
regression, clustering, and dimensionality reduction. A. PROBLEM STATEMENT
The experiments for this research paper were implemented Current research on TED has demonstrated the remarkable
in Python using the PyTorch and scikit-learn libraries. capability of state-of-the-art algorithms to achieve prediction
accuracies exceeding 90% [27]. However, while considerable
emphasis has been placed on algorithm selection and
E. MEASURING PERFORMANCE IN ML ALGORITHMS
hyperparameter tuning, other critical factors, such as the
TED, fundamentally a classification task, prioritizes pre- choice of EM, have often been overlooked. In a significant
diction and accuracy as success metrics [20], [21]. While portion of literature and survey papers, approximately 77%
precision holds value, this research does not center on fail to consider the influence of EMs when enumerating
constructing a novel, finely-tuned ML algorithm. Instead, factors affecting model performance [28].
we investigate the impact of EM selection on prediction From the ML research perspective, TED is typically
accuracy. Therefore, precision takes a second plane. To iso- approached as a classification task, with emotions serving
late this effect, we intentionally employed three established as the classifiers based on the chosen EM. Various EMs
algorithms (linear regression, FNN, and BERT) without have been utilized as the foundation for TED research
extensive optimization. without explicit justification for their selection. Moreover,
Previous research in TED demonstrates that peer-accepted the proliferation of classifiers within these EMs contributes
performance improvements typically range from 5% upwards to algorithmic complexity, potentially obscuring the true
[22], [23], [24], [25]. While lacking a universally stan- determinants of model performance. Thus, disparities in
dardized definition of ‘‘significant change’’ within this performance across trained models may be attributed not
field, our study aligns with the benchmarks established by only to computational limitations during training but also to
current research and state-of-the-art findings. By adhering challenges in generalization owing to the differing number of
to these recognized improvement measures, we provide a classifiers [29].
contextualized basis for evaluating the impact of our research. Given these considerations, the absence of a standardized
number of classifiers, stemming from the utilization of
F. HYPERPARAMETERS diverse EMs, hinders the establishment of baseline compu-
In ML, hyperparameters are preconditions or configurations tational requirements necessary for optimal model fitting.
set before the learning process begins and not directly learned Consequently, discrepancies in accuracy performance may
from the data during training. They control aspects of the partly result from disparities in hardware resources. More-
learning process and model architecture, such as the learning over, existing research predominantly focuses on training
rate, the number of hidden layers in a NN, or the choice of and validating model performance within closed datasets,
kernel in a support vector machine. Unlike model parameters, neglecting cross-dataset testing on unseen data. The rationale
which are learned from the data, hyperparameters must be often cited for this omission is the incorporation of
chosen beforehand by the practitioner based on intuition, k-fold cross-validation during training, deeming external
experimentation, or domain knowledge [26]. dataset testing unnecessary [30], and research has been
Selecting appropriate hyperparameters can improve model proven to contribute to the body of knowledge using this
performance, generalization, and faster convergence during methodology [31].
training. Conversely, poorly chosen hyperparameters can Furthermore, the incomparability of datasets in terms
result in suboptimal performance, overfitting, or underfitting of quality, distribution, and emotional categories poses a
of the model. Thus, effectively tuning hyperparameters is significant challenge. Each dataset is tailored to specific
C. OBJECTIVES E. METHODOLOGY
This study investigates the impact of EM selection on the 1) DATA COLLECTION
performance and generalizability of ML models for TED. Eight publicly available text datasets, each annotated with
A method is proposed to establish equivalence between emo- emotion labels according to distinct EMs, were obtained
tion tags across models to enable meaningful comparisons for this study, as illustrated in Table 1. During the initial
between models trained on datasets with different emotion acquisition phase, datasets were downloaded, decompressed,
label schemes to establish equivalence between emotion tags and consolidated into a shared directory. A Python program
across models. ML models will then be trained on each was developed to systematically read, parse, and format the
dataset using standardized labels derived from the original disparate datasets into a unified SQLite3 database to facilitate
EM. The performance of these dataset-specific models subsequent analysis. This centralized database stored all
will be compared to models trained on optimized, shared collected text samples within a single table, streamlining data
emotion labels mapped across datasets. Model performance access and enabling cross-dataset comparisons.
will be evaluated both on the dataset used for training The selection criteria for these datasets include general
and by cross-testing on datasets with different underlying public availability, previous usage of peer-accepted research
EMs. This cross-dataset testing is intended to assess the papers, and compatibility with the tools we used during
generalization ability of models trained under different experimentation.
conditions. By comparing performance within and across At this point, we applied only basic text pre-processing,
datasets, the study seeks to reveal patterns in how EM as shown in the following piece of code, and the regex
choice affects model accuracy and generalizability in TED responsible for text cleaning.
tasks. The goal is to gain insight into how the selection
def clean_text(text):
of EMs impacts ML model development for real-world text = [Link](r’\s+’, ’ ’, text)
applications. return [Link](r’[^\w\s@#]’, ’’, text)
The table containing all the datasets was defined with the
following columns:
• id: Just a numerical consecutive.
• source_id: The id assigned by the dataset. Some of the
datasets provided an ID for the text. For example, those
datasets containing tweets provided the tweet ID.
• dataset_id: Each dataset was named from one to eight,
as ‘‘DatasetFourLoader’’ for instance.
• text: The actual text as labeled example.
• source_emotion: The emotion or label as assigned by the
source.
• shared_emotion: An emotion label assigned per the first
EM mapping dictionary defined in this research.
• quadrant_emotion: An emotion label assigned per the
second EM mapping dictionary defined in this research.
The initial hypothesis posits that arbitrarily selecting an
EM significantly influences a TED model’s performance.
If valid, this suggests that EMs should be considered
hyperparameters and externalized to allow optimization
independently of the ML algorithm. We have implemented
this EM optimization process as EM mappings, making it
possible to enrich the original EM and produce a new one.
This study introduces two EM mappings: ‘‘shared emotions’’
and ‘‘quadrant emotions.’’
Since each dataset was labeled independently, there is no
formal methodology to create a unified set of labels for
comparison despite some coincidental overlaps. As depicted
in Table 1, the number of distinct labels varies from 4 to
28 across datasets. The first mapping, ‘‘shared emotions,’’
aims to mitigate label disparities and reduce their overall
count, thus minimizing the number of classifiers required.
It is imperative to highlight that this ‘‘shared emotions’’
mapping, like any hyperparameter, remains arbitrary but
can be tailored and fine-tuned to enhance performance. The number of labels for each dataset in their original EM,
In essence, this mapping was not hardcoded in the exper- shared, and quadrant mappings can be seen in Table 4.
iments, allowing practitioners the flexibility to modify it We wrote a Python program to implement both emotion
according to their needs. For this research, the ‘‘shared emo- mappings and store the results in the same database.
tions’’ mapping reduced the possible labels to 10 emotions An excerpt of the resulting database can be seen in Table 5.
and a ‘‘neutral’’ one, as presented in Table 2. Both emotion mappings are provided as Python dictionaries.
The second mapping is denoted as ‘‘quadrant emotion,’’ The data collection and processing flow is represented in
which draws from Russell’s circumplex model of affect, Figure 1.
delineating quadrants based on valence and arousal. Each
distinct emotion tag, as labeled independently by respective 3) MODEL TRAINING AND EVALUATION
datasets, is assigned arbitrary valence and arousal values We selected three distinct algorithms for our experimental
within the range of -1 to 1, as presented in Table 3. setup: linear regression, FNN, and BERT. Each algorithm was
Consequently, this mapping yields a Cartesian plane wherein employed to train models corresponding to different datasets
emotions are positioned within one of four quadrants (or and emotion mapping combinations. Specifically, for dataset
quadrant 0 in the case of neutral emotion). The designation for one, three models were trained using linear regression: one
the ‘‘quadrant emotion’’ corresponds to the specific quadrant with the original emotion labels, another with the shared
on the Cartesian plane where the emotion is situated (Q0, Q1, emotion mapping, and a third with the quadrant emotion
Q2, Q3, Q4). mapping. This exact procedure was repeated for dataset one
Once more, as with the first mapping, this equivalence across the FNN and BERT implementations, and this process
is provided as a Python dictionary and considered a was repeated with all datasets individually.
hyperparameter. Therefore, practitioners can change and Consequently, our experimentation trained 72 distinct
tune these values and see how performance might improve. models (including all eight datasets, three algorithms,
TABLE 3. ‘‘Quadrant emotions’’ mapping. TABLE 4. Comparison of labels after emotion mapping, ordered by
dataset id.
TABLE 5. Excerpt of the resulting database and the implemented emotion mappings.
TABLE 6. Accuracy comparison of linear regression models, ordered by considerations. Each dataset exhibited distinct characteris-
dataset id, expressed in percentages.
tics, including unique label sets (EM), variations in class rep-
resentation (class imbalances), and diverse domains. Based
on their origins, it can be inferred that the individuals involved
in labeling and the original authors of the texts represent
disparate demographics, each with distinct perspectives on
emotional expression.
It is essential to keep in mind that the objective of
these experiments was not to propose a better algorithm for
TED (with better performance) but to observe the changes
in performance, as low as it might be initially, derived
TABLE 7. Accuracy comparison of FNN models, ordered by dataset id,
expressed in percentages.
from changing only the selected EM for training while
maintaining the same algorithm, same dataset, and same
hyper-parameters.
Regarding the initial emotion mapping, referred to as
shared emotions, no significant changes (less than 5%) in
prediction accuracy or performance were observed while
maintaining consistent datasets and algorithms, as can be
seen in the first delta column in tables 6, 7, and 8. The most
significant performance variations were observed in models
trained with dataset 5, specifically Google GoEmotions, but
TABLE 8. Accuracy comparison of BERT models, ordered by dataset id,
still below the 5% mark. This discrepancy can be attributed
expressed in percentages. to significant alterations in this dataset labels, resulting from
the emotion mapping and label reduction process, reducing
the label count from 28 to only eleven shared emotions
(table 4).
An important observation stemming from the minimal
impact of the initial emotion mapping adjustment is the
consistency it implies in the algorithms’ behavior. This
consistency indicates that the algorithms remained stable
without significantly introducing imbalance or skewness
resulting from the modification in emotion mapping.
In the extreme scenario of maximum label reduction (as the
model’s original accuracy, the second column its accuracy quadrant emotion only allows for up to 5 labels), the second
during cross-dataset testing, and the third column showing delta column in tables 6, 7, and 8 demonstrates a remarkable
the corresponding difference in performance. Subsequent increase in prediction accuracy. The available labels were
columns in each table display analogous data but utilize minimized under this second emotion mapping, while the
shared and quadrant emotion mapping. algorithm and datasets remained unchanged.
Table 9 pertains to the outcomes of cross-dataset testing The findings underscore the sensitivity of ML models to
conducted with the linear regression model. In contrast, label modifications, suggesting that label mapping serves as
Table 10 delineates the results obtained from the FNN a critical hyperparameter in model optimization. However,
model. Finally, Table 11 presents and compares the outcomes it is essential to exercise caution when interpreting results
associated with the BERT models. in the context of extreme label customization, as those in
the second delta column from tables 6, 7, and 8, since
IV. DISCUSSION they may lead to exaggerated or unrealistic performance
The eight datasets utilized in this study were developed outcomes. Effectively, as with any other hyper-parameter,
independently by various researchers, lacking compatibility EM manipulation can lead to model overfitting.
TABLE 9. Accuracy comparison of cross-dataset testing for linear regression models, ordered by dataset id, expressed in percentages.
TABLE 10. Accuracy comparison of cross-dataset testing for FNN models, ordered by dataset id, expressed in percentages.
TABLE 11. Accuracy comparison of cross-dataset testing for BERT models, ordered by dataset id, expressed in percentages.
While label reduction can yield valuable insights into from cross-dataset testing with different emotion mappings.
model behavior and enhance computational efficiency, Delta columns show the differences in performance. The
researchers must remain cognizant of the potential trade-offs, first delta column quantifies the change in accuracy when
such as reduced model expressiveness and generalization trained models were presented with unseen data from other
capabilities. datasets without any emotion mapping equivalences. The
Research in TED aims to develop models capable of con- second delta column represents the change in accuracy when
sistently predicting emotions in diverse text samples, thereby the models used for cross-dataset testing used the shared
demonstrating consistent performance across unseen data emotions mapping. Finally, the third delta column shows the
encountered ‘‘in the wild.’’ Cross-dataset testing is a crucial changes in accuracy when models trained with the quadrant
evaluation metric for assessing the model’s performance emotion mappings were used during cross-dataset testing.
under real-world conditions. In this context, the subsequent
phase of the experiments focused on cross-dataset testing as 1) CROSS-DATASET TESTING IN LINEAR REGRESSION
the authentic benchmark of model performance. MODELS
The observed variations in cross-dataset testing performance
A. CROSS-DATASET TESTING underscore potential limitations arising from using linear
The performance of trained models for predicting emotions regression models in the context of NLP and TED, with
in unseen data varies considerably. When tested on dif- a particular sensitivity to the number of classifiers (labels
ferent datasets, the models showed significant changes in in the EM). Linear regression’s core assumption of a
performance (over 5% in either increase or decrease) as linear relationship between features and target variables may
depicted in the delta columns from Tables 9, 10, and 11. become increasingly strained as the number of emotion labels
These tables illustrate the changes in performance resulting varies. More distinct emotion categories introduce a greater
need to model complex, potentially non-linear boundaries emotion signals that are less susceptible to variations across
between emotions within the feature space. datasets.
The subjective labeling of emotional content and linguistic The improvement in cross-dataset testing outcomes sug-
variation between different text corpora further exacerbates gests the existence of shared underlying representations of
this effect. Datasets relying on fine-grained emotion labels emotion within the datasets. It is plausible that the complexity
may have subtle linguistic cues and patterns distinguishing of multiple labels led models to overfit specific dataset char-
these emotions, which a linear model struggles to represent acteristics, while label reduction inadvertently emphasized
adequately. Similarly, when a model has learned decision transferable patterns of emotional expression. Ultimately,
boundaries based on a specific set of labels, it can be NNs appear adept at identifying core linguistic markers of
ill-equipped to deal with a differing label set found in emotion regardless of nuanced labeling distinctions.
cross-dataset testing. In other words, these results suggest that having more
For the present study, results for the cross-dataset testing labels does not necessarily mean that the dataset is of better
in the linear regression models (table 9) show an average of quality, and therefore, the trained model accuracy might be
20.45% accuracy while using the original EM, 21.69% with compromised.
the shared emotions mapping and 54.04% for the quadrant For the FNN models (table 10), the most significant change
emotions. These measurements contrast with the original was presented by dataset 5, comparing its performance using
dataset’s accuracy average of 64.69% its own EM against the rest of the unseen datasets, with a
The most significant performance change is found on 59.38% increase in accuracy.
dataset id 7 using its original EM, demonstrating an 80.4% Out of the 24 different cross-dataset testing scenarios,
decrease in accuracy when evaluating other datasets. When 4 reported non-significant changes in accuracy (less
using the shared emotions mapping, the most significant than 5%), while the remaining 20 did show changes
change is from the model trained with dataset id 3, with a above 5%.
73.18% decrease in accuracy. Finally, the most considerable
variation for the quadrant emotions mapping is present in the 3) CROSS-DATASET TESTING IN BERT-BASED MODELS
model trained with dataset id 8, with a decrease in accuracy Integrating BERT-based classifiers in the experiments exhib-
of 36.78%. 23 of the 24 cross-dataset testing measurements ited significant enhancements in baseline accuracy, further
present a significant (above 5%) change in performance augmented by notable advancements during cross-dataset
for the models with linear regression algorithms. The only testing. This outcome, particularly when coupled with label
non-significant difference (less than 5%) was reported by the reduction, underscores the intricate relationship between
model trained with dataset id 6 with the quadrant emotions BERT’s attributes and the manipulation of emotion labels.
mapping, with a variation of only a 1.47% increase. BERT models derive their efficacy from ample pre-training
on extensive text corpora utilizing masked language modeling
2) CROSS-DATASET TESTING IN FNN MODELS and next-sentence prediction methodologies. This fosters a
In contrast to the findings observed with linear regression profound contextual comprehension of language, resulting
models, using FNNs revealed some cases of a notable in the generation of nuanced word representations. It is
increase in accuracy during cross-dataset testing with a plausible that these representations empower a BERT clas-
reduction in emotion labels. This outcome underscores the sifier to grasp the underlying semantic connections among
impact of label manipulation and the inherent characteristics diverse emotion labels across datasets, even in the presence
of NNs on model performance. of misaligned labeling schemes.
FNNs inherently excel in modeling non-linear relation- Moreover, BERT’s attention mechanisms facilitate sen-
ships, a crucial aspect for effectively capturing subtle sitivity to context-specific word interactions, aligning well
linguistic nuances associated with emotions. Their layered with the expression of emotions through intricate linguistic
architecture facilitates intricate feature transformations, combinations, thereby potentially bridging the disparities
potentially making them less susceptible to dataset-specific introduced by disparate emotion labels.
variations than linear models. Moreover, implicit regulariza- It is imperative to underscore that adjusting the number
tion mechanisms during training can help mitigate overfitting of labels should not be considered a trivial modification.
concerns inherent in smaller datasets. These experiments substantially impact performance while
The significant enhancement in cross-dataset accuracy upholding consistent datasets, preprocessing techniques, and
observed with label reduction suggests potential inconsis- algorithmic selections. Reducing label sets may compel
tencies or semantic discrepancies across the original sets of BERT models to attend to broader contextual emotional
emotion labels. By leveraging their enhanced representational expression patterns that rely less on dataset-specific labeling
capabilities, NNs may identify shared emotional concepts peculiarities. However, excessive abstraction in label reduc-
despite diverse labeling schemes. Furthermore, aligning label tion poses the risk of forfeiting discriminative information,
reduction with established psychological constructs, such thereby diminishing the capacity to discern pertinent nuances
as optimized emotion categories (and, in this way, tuning of emotion. Conversely, employing an excessive number
the used EM), can encourage models to learn fundamental of labels, particularly those that are overly specific or
granular, can introduce variability and noise during model The implications of these findings extend beyond academic
training. This can lead to difficulties for BERT classifiers circles, impacting developers and practitioners working on
in identifying meaningful patterns and associations amidst affective-aware applications. As the demand for sophisticated
many fine-grained emotional distinctions. As a result, the text-based emotional analysis grows (be it in customer service
model may struggle to generalize effectively to unseen data, bots, mental health assessments, or personalized content
thereby reducing overall accuracy and robustness. delivery), so does the necessity for robust and adaptable
The augmentation in cross-dataset accuracy implies that TED systems. This research provides a clear directive for
BERT’s robust linguistic pre-training and contextual empha- future developments by integrating standardized emotion
sis enable it to circumvent certain limitations associated with labels and fostering a more systematic approach to EM
strict overfitting to individual dataset characteristics. These selection. This, in turn, paves the way for more reliable
findings underscore the pivotal role of nuanced semantic and nuanced emotional analysis tools, which are crucial
comprehension in navigating the intrinsic subjectivity and for advancing human-computer interaction in increasingly
variability of emotion labeling in textual data. digital landscapes.
For the BERT model (table 11) dataset 5 and its comparison In summary, this research delivers valuable insights for
using its own EM against unseen data is also the biggest practitioners seeking to enhance the efficacy and adaptability
change in performance, with an increase of 50.17%. Out of of ML models operating within emotionally nuanced
the 24 accuracy results, only three of them showed changes text datasets. Furthermore, it prompts pertinent inquiries
of less than 5% variation. The remaining 21 results showed regarding the necessity for unified theoretical frameworks
significant, above 5% variations. governing the modeling and labeling of emotions within the
During all the cross-dataset testings, we found that out computational linguistics domain.
of 72 experiments (24 for each of linear regression, FNN,
and BERT), in 8 cases, the variance in accuracy was non- VI. SCOPE AND LIMITATIONS
significant (less than 5%), while in the remaining 64 cases, The experimentation encompassed three distinct models:
the change was significant. This means that for around 88% linear regression, FNN, and BERT. While the datasets utilized
of the experiments with all elements equal, the introduced were publicly accessible, it is essential to note that they do
emotion mappings significantly impacted performance and not represent an exhaustive compilation. The source code
generalizability. implementation of all algorithms has been disclosed along
with this paper; however, it is conceivable that practition-
ers may discover alternative, more optimized approaches.
V. CONCLUSION Nonetheless, it is essential to clarify that the primary objective
For the research question: How does the arbitrary selection of of this research was not to identify an optimal implementation
an emotion model affect the performance and generalizability nor to delve into aspects such as text preprocessing, model
of machine learning models for text emotion detection, and architecture, programming methodology, or other ancillary
how can emotion labels from different models be stan- factors. Instead, the focus remained on demonstrating the
dardized to improve compatibility and generalization across impact of alterations in the emotion model (EM) while
diverse datasets? This investigation provides compelling maintaining all other variables constant and adhering to a
evidence of the significant influence exerted by the arbitrary straightforward methodology.
selection of EMs on the effectiveness and adaptability of All datasets employed in the study were exclusively in
ML models employed in TED. When all other experimental English, as were the corresponding labels. The experimenta-
variables remained constant, alterations in emotion mappings tion used Google Colab Pro, leveraging prioritized access to
consistently induced model accuracy and performance high-performance NVIDIA A100 GPUs. While the specific
fluctuations across different datasets. Notably, in 88% of specifications of the allocated compute instance may vary
the conducted experiments, these fluctuations surpassed within Google’s cloud environment, it is noteworthy that
the predetermined threshold of 5%, thus substantiating the A100 GPUs consistently offered substantial acceleration and
rejection of the null hypothesis. memory capacity conducive to model training and evaluation.
We propose that integrating standardized emotion labels Nevertheless, it is pertinent to acknowledge that variations in
as a hyper-parameter for TED research is pivotal in computational resources may influence the reported findings,
fostering compatibility across datasets and enhancing overall as previously stated in this document.
model generalization. Models trained with standardized
labels consistently performed better than those trained with
dataset-specific tags. These outcomes underscore the critical REFERENCES
importance of deliberate and informed decisions regarding [1] I Gartner. (2019). 2021 Strategic Roadmap for Enterprise AI:
the choice of EMs and the implementation of standardized Natural Language Architecture. [Online]. Available: [Link]
[Link]/en/documents/3994504
labeling practices in developing robust TED systems that
[2] H. Li, B. X. B. Yu, G. Li, and H. Gao, ‘‘Restaurant survival prediction using
would perform consistently in both controlled environments customer-generated content: An aspect-based sentiment analysis of online
and ‘‘on the wild’’ with unseen text. reviews,’’ Tourism Manage., vol. 96, Jun. 2023, Art. no. 104707.
[3] I Gartner. (2021). Hype Cycle for Natural Language Technologies, 2021. [28] A. D. L. Languré and M. Zareei, ‘‘Breaking barriers in sentiment analysis
[Online]. Available: [Link] and text emotion detection: Toward a unified assessment framework,’’
[4] P. Vyas, M. Reisslein, B. P. Rimal, G. Vyas, G. P. Basyal, and P. Muzumdar, IEEE Access, vol. 11, pp. 125698–125715, 2023.
‘‘Automated classification of societal sentiments on Twitter with machine [29] C. Cortes, L. D. Jackel, and W.-P. Chiang, ‘‘Limits on learning machine
learning,’’ IEEE Trans. Technol. Soc., vol. 3, no. 2, pp. 100–110, Jun. 2022. accuracy imposed by data quality,’’ in Proc. Knowl. Discovery Data
[5] P. Nandwani and R. Verma, ‘‘A review on sentiment analysis and emotion Mining, 1995, pp. 1–8.
detection from text,’’ Social Netw. Anal. Mining, vol. 11, no. 1, p. 81, [30] M. Stone, ‘‘Cross-validatory choice and assessment of statistical predic-
Aug. 2021. tions,’’ J. Roy. Stat. Soc., Ser. B Methodol., vol. 36, no. 2, pp. 111–133,
[6] P. S. Sreeja and G. Mahalakshmi, ‘‘Emotion models: A review,’’ Int. J. 1974.
Control Theory Appl., vol. 10, no. 8, pp. 651–657, 2017. [31] G. M. Shafiq, T. Hamza, M. F. Alrahmawy, and R. El-Deeb, ‘‘Enhancing
[7] R. Plutchik, Emotion: A Psychoevolutionary Synthesis. Manhattan, NY, Arabic aspect-based sentiment analysis using end-to-end model,’’ IEEE
USA: Harper & Row, 1980. Access, vol. 11, pp. 142062–142076, 2023.
[8] C. E. Izard, Human Emotions. New York, NY, USA: Springer, 1977. [32] CrowdFlower. (2016). Emotion Detection From Text. [Online]. Available:
[9] P. Ekman, ‘‘An argument for basic emotions,’’ Cognition Emotion, vol. 6, [Link]
nos. 3–4, pp. 169–200, 1992. [33] P. Govindaraj. (2020). Emotions Dataset for NLP. Kaggle. [Online]. Avail-
[10] J. A. Russell, ‘‘A circumplex model of affect,’’ J. Personality Social able: [Link]
Psychol., vol. 39, no. 6, p. 1161, 1980. nlp/data
[11] K. R. Scherer, ‘‘Emotions: Definition, nature, and general issues,’’ in [34] E. Saravia, H.-C. T. Liu, Y.-H. Huang, J. Wu, and Y.-S. Chen, ‘‘CARER:
Handbook of Emotions. New York, NY, USA: Guilford Press, 2000, Contextualized affect representations for emotion recognition,’’ in Proc.
pp. 137–154. Conf. Empirical Methods Natural Lang. Process. Brussels, Belgium:
[12] J. J. Gross, ‘‘The component process model of emotion regulation,’’ Association for Computational Linguistics, Nov. 2018, pp. 3687–3697.
Psychol. Inquiry, vol. 9, no. 2, pp. 80–99, 1998. [Online]. Available: [Link]
[13] A. Mehrabian, ‘‘Pleasure-arousal-dominance: A general framework for [35] E. S. Dan-Glauser and K. R. Scherer, ‘‘The difficulties in emotion
describing and measuring individual differences in temperament,’’ in regulation scale (DERS): Factor structure and consistency of a French
Current Psychology. New York, NY, USA: Springer, 1996. translation,’’ Swiss J. Psychol., vol. 72, no. 1, pp. 5–11, Jan. 2013, doi:
10.1024/1421-0185/a000093.
[14] T. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997.
[36] D. Demszky, D. Movshovitz-Attias, J. Ko, A. Cowen, G. Nemade, and
[15] D. Maulud and A. M. Abdulazeez, ‘‘A review on linear regression
S. Ravi, ‘‘GoEmotions: A dataset of fine-grained emotions,’’ in Proc. 58th
comprehensive in machine learning,’’ J. Appl. Sci. Technol. Trends, vol. 1,
Annu. Meeting Assoc. Comput. Linguistics (ACL), 2020, pp. 4040–4054.
no. 44, pp. 140–147, Dec. 2020.
[37] C. Strapparava and R. Mihalcea, ‘‘Semeval-2007 task 14: Affective
[16] B. Duncan and Y. Zhang, ‘‘Neural networks for sentiment analysis
text,’’ in Proc. 4th Int. Workshop Semantic Evaluations (SemEval), 2007,
on Twitter,’’ in Proc. IEEE 14th Int. Conf. Cognit. Informat. Cog-
pp. 70–74.
nit. Comput. (ICCI*CC), Jul. 2015, pp. 275–278. [Online]. Available:
[38] D. Ghazi, D. Inkpen, and S. Szpakowicz, ‘‘Detecting emotion stimuli in
[Link]
emotion-bearing sentences,’’ in Proc. 16th Int. Conf., Cairo, Egypt. Cham,
[17] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, ‘‘BERT: Pre-training
Switzerland: Springer, 2015, pp. 152–165.
of deep bidirectional transformers for language understanding,’’ in Proc.
[39] S. M. Mohammad and F. Bravo-Marquez, ‘‘WASSA-2017 shared task on
Conf. North Amer. Chapter Assoc. Comput. Linguistics, Human Lang.
emotion intensity,’’ in Proc. Workshop Comput. Approaches Subjectivity,
Technol., vol. 1, J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis,
Sentiment Social Media Anal. (WASSA), Copenhagen, Denmark, 2017,
MN, USA: Association for Computational Linguistics, Jun. 2019,
pp. 65–77.
pp. 4171–4186. [Online]. Available: [Link]
[18] A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin,
A. Desmaison, L. Antiga, and A. Lerer, ‘‘Automatic differentiation in
Pytorch,’’ in Proc. NIPS-W, 2017, pp. 1–4.
[19] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel,
M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, ALEJANDRO DE LEÓN LANGURÉ received the
A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, bachelor’s degree in computer science from the
‘‘Scikit-learn: Machine learning in Python,’’ J. Mach. Learn. Res., vol. 12, School of Engineering and Sciences, Tecnológico
pp. 2825–2830, Jul. 2011. de Monterrey, and the master’s degree from the
[20] S. M. Basha and D. Rajput, Survey on Evaluating the Performance of School of Economics and Business, Anahuac
Machine Learning Algorithms: Past Contributions and Future Roadmap. University. He is currently pursuing the Ph.D.
Cambridge, MA, USA: Academic, pp. 153–164, Jan. 2019. degree in computer sciences with the School
[21] A. Alslaity and R. Orji, ‘‘Machine learning techniques for emotion of Engineering and Sciences, Tecnológico de
detection and sentiment analysis: Current state, challenges, and future Monterrey, with a focus on machine learning for
directions,’’ Behaviour Inf. Technol., vol. 43, pp. 1–26, Dec. 2022. affective computing and text emotion detection.
[22] J. Herzig, M. Shmueli-Scheuer, and D. Konopnicki, Emotion Detection
From Text via Ensemble Classification Using Word Embeddings. New
York, NY, USA: ACM, Oct. 2017.
[23] W. Witon, P. Colombo, A. Modi, and M. Kapadia, ‘‘Disney at IEST 2018:
Predicting emotions using an ensemble,’’ in Proc. 9th Workshop Comput. MAHDI ZAREEI (Senior Member, IEEE) received
Approaches Subjectivity, Sentiment Social Media Anal., A. Balahur, the [Link]. degree in computer networks from the
S. M. Mohammad, V. Hoste, and R. Klinger, Eds. Brussels, Belgium: University of Science Malaysia, in 2011, and the
Association for Computational Linguistics, Oct. 2018, pp. 248–253.
Ph.D. degree from the Communication Systems
[Online]. Available: [Link]
and Networks Research Group, Malaysia-Japan
[24] B. Kratzwald, S. Ilic, M. Kraus, S. Feuerriegel, and H. Prendinger,
International Institute of Technology, Universiti
‘‘Decision support with text-based emotion recognition: Deep learning for
affective computing,’’ Decis. Support Syst., vol. 115, pp. 24–35, Mar. 2018. Teknologi Malaysia, Malaysia, in 2016. In 2017,
[25] F. Anzum and M. L. Gavrilova, ‘‘Emotion detection from micro- he joined the School of Engineering and Sciences,
blogs using novel input representation,’’ IEEE Access, vol. 11, Tecnológico de Monterrey, as a Postdoctoral Fel-
pp. 19512–19522, 2023. low, where he was a Research Professor, in 2019.
[26] P. Probst, A.-L. Boulesteix, and B. Bischl, ‘‘Tunability: Importance of His research interests include wireless sensor and ad hoc networks, energy
hyperparameters of machine learning algorithms,’’ J. Mach. Learn. Res., harvesting sensors, information security, and machine learning. He is a
vol. 20, no. 1, Jan. 2019, Art. no. 19341965. member of Mexican National Researchers System (Level I). He is an
[27] P. Lin, X. Luo, and Y. Fan, ‘‘A survey of sentiment analysis based on deep Associate Editor of IEEE ACCESS and PLOS One.
learning,’’ Int. J. Comput. Inf. Eng., vol. 14, no. 12, pp. 473–485, 2020.