Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi)
Not a member yet
1071 research outputs found
Sort by
The Accuracy Comparison Between Word2Vec and FastText On Sentiment Analysis of Hotel Reviews
Word embedding vectorization is more efficient than Bag-of-Word in word vector size. Word embedding also overcomes the loss of information related to sentence context, word order, and semantic relationships between words in sentences. Several kinds of Word Embedding are often considered for sentiment analysis, such as Word2Vec and FastText. Fast Text works on N-Gram, while Word2Vec is based on the word. This research aims to compare the accuracy of the sentiment analysis model using Word2Vec and FastText. Both models are tested in the sentiment analysis of Indonesian hotel reviews using the dataset from TripAdvisor.Word2Vec and FastText use the Skip-gram model. Both methods use the same parameters: number of features, minimum word count, number of parallel threads, and the context window size. Those vectorizers are combined by ensemble learning: Random Forest, Extra Tree, and AdaBoost. The Decision Tree is used as a baseline for measuring the performance of both models. The results showed that both FastText and Word2Vec well-to-do increase accuracy on Random Forest and Extra Tree. FastText reached higher accuracy than Word2Vec when using Extra Tree and Random Forest as classifiers. FastText leverage accuracy 8% (baseline: Decision Tree 85%), it is proofed by the accuracy of 93%, with 100 estimators.
Word embedding vectorization is more efficient than Bag-of-Word in word vector size. Word embedding also overcomes the loss of information related to sentence context, word order, and semantic relationships between words in sentences. Several kinds of Word Embedding are often considered for sentiment analysis, such as Word2Vec and FastText. Fast Text works on N-Gram, while Word2Vec is based on the word. This research aims to compare the accuracy of the sentiment analysis model using Word2Vec and FastText. Both models are tested in the sentiment analysis of Indonesian hotel reviews using the dataset from TripAdvisor.Word2Vec and FastText use the Skip-gram model. Both methods use the same parameters: number of features, minimum word count, number of parallel threads, and the context window size. Those vectorizers are combined by ensemble learning: Random Forest, Extra Tree, and AdaBoost. The Decision Tree is used as a baseline for measuring the performance of both models. The results showed that both FastText and Word2Vec well-to-do increase accuracy on Random Forest and Extra Tree. FastText reached higher accuracy than Word2Vec when using Extra Tree and Random Forest as classifiers. FastText leverage accuracy 8% (baseline: Decision Tree 85%), it is proofed by the accuracy of 93%, with 100 estimators
Implementation of Ensemble Method in Schizophrenia Identification Based on Microarray Data
Schizophrenia is a chronic mental illness that leads the patient to hallucinations and delusions with a prevalence of 0.4% worldwide. The importance early detection of Schizophrenia is tracking the pre-syndrome of Schizophrenia during the active phase, and could reduce psychosis symptomatic. However, the method sometimes cannot detect the symptoms accurately. As an alternative, machine learning can be implemented on microarray data for early detection. This study aimed to implement three ensemble methods, i.e., Random Forest (RF), Adaptive Boosting (AdaBoost), and Extreme Gradient Boosting (XGBoost) to identify Schizophrenia. Hyperparameter tuning was performed to improve the performance of the models. Based on the results, we found that the model 6, which is developed by the XGBoost method, performs better than other models with the value of accuracy and F1-score are 0.87 and 0.87, respectively.Schizophrenia is a chronic mental illness that leads the patient to hallucinations and delusions with a prevalence of 0.4% worldwide. The importance early detection of Schizophrenia is tracking the pre-syndrome of Schizophrenia during the active phase, and could reduce psychosis symptomatic. However, the method sometimes cannot detect the symptoms accurately. As an alternative, machine learning can be implemented on microarray data for early detection. This study aimed to implement three ensemble methods, i.e., Random Forest (RF), Adaptive Boosting (AdaBoost), and Extreme Gradient Boosting (XGBoost) to identify Schizophrenia. Hyperparameter tuning was performed to improve the performance of the models. Based on the results, we found that the model 6, which is developed by the XGBoost method, performs better than other models with the value of accuracy and F1-score are 0.87 and 0.87, respectively
Critical Section Overhead Reduction for OpenMP Program by Nesting a Serial Loop to Increase Task Granularity of Parallel Loop
This paper presents a simple method to reduce performance loss due to a parallel program's massive critical sections of parallel numerical integration. The method transforms a fine-grain parallel loop into a coarse grain parallel loop that nests a sequential loop. The coarse grain parallel loop is by nesting a loop block to make task granularities coarser than that naive one. In addition to the overhead reduction, the method makes the parallel work fraction significantly more significant than the serial fraction. As a result, nesting a serial loop within a parallel loop improves the parallel program's performance. Compared to the naïve method, which does not scale the performance of a parallel program of numerical integration, the nesting serial loop method scales a parallel program up to 3.26 times a fold relative to its sequential program on a quad-core processor. This result shows that the proposed method makes the parallel program much faster than the naïve method.
 
Educational Data Mining Using Cluster Analysis Methods and Decision Trees based on Log Mining
Educational Data Mining (EDM) often appears to be applied in big data processing in the education sector. One of the educational data that can be further processed with EDM is activity log data from an e-learning system used in teaching and learning activities. The log activity can be further processed more specifically by using log mining. The purpose of this study was to process log data from the Sebelas Maret University Online Learning System (SPADA UNS) to determine student learning behavior patterns and their relationship to the final results obtained. The data mining method applied in this research is cluster analysis with the K-means Clustering and Decision Tree algorithms. The clustering process is used to find groups of students who have similar learning patterns. While the decision tree is used to model the results of the clustering in order to enable the analysis and decision-making processes. Processing of 11,139 SPADA UNS log data resulted in 3 clusters with a Davies Bouldin Index (DBI) value of 0.229. The results of these three clusters are modeled by using a Decision Tree. The decision tree model in cluster 0 represents a group of students who have a low tendency of learning behavior patterns with the highest frequency of access to course viewing activities obtained accuracy of 74.42% . In cluster 1, which contains groups of students with high learning behavior patterns, have a high frequency of access to viewing discussion activities obtained accuracy of 76.47%. While cluster 2 is a group of students who have a pattern of learning behavior that is having a high frequency of access to the activity of sending assignments obtained accuracy of 90.00%.
Educational Data Mining (EDM) often appears to be applied in big data processing in the education sector. One of the educational data that can be further processed with EDM is activity log data from an e-learning system used in teaching and learning activities. The log activity can be further processed more specifically by using log mining. The purpose of this study was to process log data from the Sebelas Maret University Online Learning System (SPADA UNS) to determine student learning behavior patterns and their relationship to the final results obtained. The data mining method applied in this research is cluster analysis with the K-means Clustering and Decision Tree algorithms. The clustering process is used to find groups of students who have similar learning patterns. While the decision tree is used to model the results of the clustering in order to enable the analysis and decision-making processes. Processing of 11,139 SPADA UNS log data resulted in 3 clusters with a Davies Bouldin Index (DBI) value of 0.229. The results of these three clusters are modeled by using a Decision Tree. The decision tree model in cluster 0 represents a group of students who have a low tendency of learning behavior patterns with the highest frequency of access to course viewing activities obtained accuracy of 74.42% . In cluster 1, which contains groups of students with high learning behavior patterns, have a high frequency of access to viewing discussion activities obtained accuracy of 76.47%. While cluster 2 is a group of students who have a pattern of learning behavior that is having a high frequency of access to the activity of sending assignments obtained accuracy of 90.00%.
 
Development of Mastoid Air Cell System Extraction Method on Temporal CT-scan Image
Mastoiditis is disease that to infection of the mastoid bone cavity that affects the size of the air cell system of the temporal bone. Visually, the information temporal CT image mastoid bone has can assist medical experts in viewing the mastoid air cell system (MACS), but the fact that medical personnel are experiencing difficulties in determining the size MACS is due to the many different characteristics and objects overlap, so that in the measurement of the area, precise and accurate results have not been obtained. This study aims to separate the object of the MACS with the development of extraction. The proposed method uses Morphology and Regionprops operations. The dataset used in the testing process is 347 of 5 patients indicated for Mastoiditis. The results obtained can calculate the area of MACS for each test image. Based on image testing, the area of the smallest MACS in this study was 0.589 cm2 and the largest was 6.183 cm2. This, the smaller the size of the MACS indicates the severity of infection, so this study can help medical personnel make decisions and take appropriate treatment actions.
Mastoiditis is disease that to infection of the mastoid bone cavity that affects the size of the air cell system of the temporal bone. Visually, the information temporal CT image mastoid bone has can assist medical experts in viewing the mastoid air cell system (MACS), but the fact that medical personnel are experiencing difficulties in determining the size MACS is due to the many different characteristics and objects overlap, so that in the measurement of the area, precise and accurate results have not been obtained. This study aims to separate the object of the MACS with the development of extraction. The proposed method uses Morphology and Regionprops operations. The dataset used in the testing process is 347 of 5 patients indicated for Mastoiditis. The results obtained can calculate the area of MACS for each test image. Based on image testing, the area of the smallest MACS in this study was 0.589 cm2 and the largest was 6.183 cm2. This, the smaller the size of the MACS indicates the severity of infection, so this study can help medical personnel make decisions and take appropriate treatment actions
Word2Vec on Sentiment Analysis with Synthetic Minority Oversampling Technique and Boosting Algorithm
Customer opinion is an important aspect in determining the success of a company or service provider. By determining the sentiment of the existing opinion, the company can use it as an evaluation material to improve the quality of the service or product provided. Sentiment analysis can be used as a measure of opinion sentiment with input data in the form of a corpus which will be classified into positive or negative classes to obtain the level of customer satisfaction with a product or service. Aspect-based sentiment analysis can be used by companies to analyze more specifically and find out what aspects need to be improved. In this research, an aspect-based sentiment analysis was conducted on Telkomsel users on Twitter. The data used is 16,992 tweets from users who discuss several aspects such as Telkomsel's services and signals in Twitter. In this research Word2Vec was used for feature expansion to minimize vocabulary mismatch caused by limited words in tweets. The results showed that Word2Vec, Synthetic Minority Oversampling Technique (SMOTE), and Boosting algorithm combination with Logistic Regression classifier achieve highest accuracy of 95.10% for signal aspect and using hyperparameters makes the service aspect get the highest accuracy of 93.34%.
Customer opinion is an important aspect in determining the success of a company or service provider. By determining the sentiment of the existing opinion, the company can use it as an evaluation material to improve the quality of the service or product provided. Sentiment analysis can be used as a measure of opinion sentiment with input data in the form of a corpus which will be classified into positive or negative classes to obtain the level of customer satisfaction with a product or service. Aspect-based sentiment analysis can be used by companies to analyze more specifically and find out what aspects need to be improved. In this research, an aspect-based sentiment analysis was conducted on Telkomsel users on Twitter. The data used is 16,992 tweets from users who discuss several aspects such as Telkomsel's services and signals in Twitter. In this research Word2Vec was used for feature expansion to minimize vocabulary mismatch caused by limited words in tweets. The results showed that Word2Vec, Synthetic Minority Oversampling Technique (SMOTE), and Boosting algorithm combination with Logistic Regression classifier achieve highest accuracy of 95.10% for signal aspect and using hyperparameters makes the service aspect get the highest accuracy of 93.34%
Aspect Based Sentiment Analysis with FastText Feature Expansion and Support Vector Machine Method on Twitter
Social media such as Twitter has now become very close to society. Twitter users can express current issues, their opinions, product reviews, and many other things both positive and negative. Twitter is also used by companies to monitor the assessment of their products among the public as insight that will be used to evaluate what aspects of their products need to be further developed. Twitter with its limitation of only allowing users to post a maximum tweet of 280 characters will make a lot of abbreviated and difficult to understand words used, so it will allow vocabulary mismatch problems to occur. Therefore, in this paper, research conducted on aspect-based sentiment analysis of Telkomsel’s products from the aspects of signal and service by applying feature expansion using Fasttext word embedding to overcome vocabulary mismatch problem and classification with the Support Vector Machine (SVM) method. Sampling technique with Synthetic Minority Oversampling Technique (SMOTE) used to overcome data imbalance. The experimental results show that feature expansion can increase the performance of model. The final results obtained F1-Score value of the model for the signal aspect increased by 27.91% with F1-Score 95.93%, and for the service aspect increased by 42.36% with F1-Score 94.53%.Social media such as Twitter has now become very close to society. Twitter users can express current issues, their opinions, product reviews, and many other things both positive and negative. Twitter is also used by companies to monitor the assessment of their products among the public as insight that will be used to evaluate what aspects of their products need to be further developed. Twitter with its limitation of only allowing users to post a maximum tweet of 280 characters will make a lot of abbreviated and difficult to understand words used, so it will allow vocabulary mismatch problems to occur. Therefore, in this paper, research conducted on aspect-based sentiment analysis of Telkomsel’s products from the aspects of signal and service by applying feature expansion using Fasttext word embedding to overcome vocabulary mismatch problem and classification with the Support Vector Machine (SVM) method. Sampling technique with Synthetic Minority Oversampling Technique (SMOTE) used to overcome data imbalance. The experimental results show that feature expansion can increase the performance of model. The final results obtained F1-Score value of the model for the signal aspect increased by 27.91% with F1-Score 95.93%, and for the service aspect increased by 42.36% with F1-Score 94.53%
Analysis of the Quality of Natural Dyes in Weaving Exposed to Sunlight Using MSE and PSNR Parameters
It is widely assumed that natural dyes in weaving degrade in quality when exposed to sunlight for an extended period. This indication is visible to the naked eye. There is currently no standard for evaluating the quality of natural dyes. The Boti tribe's weaving on Timor Island, East Nusa Tenggara Province, is one type of weaving that uses natural dyes. The dye is made from corn flour and a combination of "nobah" leaves and the bark of the "bauk ulu" tree (from the local language). White (from corn flour) and blue-black are the colors produced by dyeing the yarn. The purpose of this research is to examine the image quality of the Boti tribe's woven fabric. The parameters used were Means Square Error (MSE), Peak Signal to Noise (PSNR), and RGB values. The image of the weaving used as a reference is compared to the image of the sun-dried weaving. The image capture distance was 30 cm, and the cropped RGB image size was 423x623x3. The experimental method was used in the research. The drying time was one hour, and it was repeated every hour between 10:00 and 15:00 local time. The sun-dried images were photographed, and parameter comparisons were performed for analysis. The results demonstrated that the MSE and PSNR methods were effective in measuring the image quality of weaving dyed with natural dyes. The average value has changed by 8.42% for the R-value, 8.58% for the G value, and 9.68% for the B value. The average PSNR for RGB images is 9.44288 dB, and the MSE is 7477.52. For grayscale images, the average PSNR is 10.52 dB and the average MSE is 5832.06.It is widely assumed that natural dyes in weaving degrade in quality when exposed to sunlight for an extended period of time. This indication is clearly visible to the naked eye. There is currently no standard for evaluating the quality of natural dyes. The Boti tribe's weaving on Timor Island, East Nusa Tenggara Province, is one type of weaving that uses natural dyes. The dye is made from corn flour and a combination of "nobah" leaves and the bark of the "bauk ulu" tree (from the local language). White (from corn flour) and blue-black are the colors produced by dyeing the yarn. The purpose of this research is to examine the image quality of the Boti tribe's woven fabric. The parameters used were Means Square Error (MSE), Peak Signal to Noise (PSNR), and RGB values. The image of the weaving used as a reference is compared to the image of the sun-dried weaving. The image capture distance was 30 cm, and the cropped RGB image size was 423x623x3. The experimental method was used in the research. The drying time was one hour, and it was repeated every one hour between 10:00 and 15:00 local time. The sun-dried images were photographed, and parameter comparisons were performed for analysis. The results demonstrated that the MSE and PSNR methods were effective in measuring the image quality of weaving dyed with natural dyes. The average value has changed by 8.42% for the R value, 8.58% for the G value, and 9.68% for the B value. The average PSNR for RGB images is 9.44288 dB, and the MSE is 7477.52. For grayscale images, the average PSNR is 10.52 dB and the average MSE is 5832.06
Big Five Personality Assessment Using KNN method with RoBERTA
Personality is the general way a person responds to and interacts with others. Personality is also often defined as the quality that distinguishes individuals. Social media was created to help people communicate remotely and easily. These personalities fall into five categories known as the Big Five personality traits, namely Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism (OCEAN). The use of K-Nearest Neighbour (KNN) is a method of classifying objects based on the training data closest to them. To overcome the data imbalance during training data, we use K-Means SMOTE (Synthetic Minority Oversampling Technique). Other features such as LIWC (Linguistic Inquiry Word Count), Information Gain, Robustly Optimized BERT Approach (RoBERTa), and hyperparameter tuning can improve the performance of the systems we build. The focus of this study is to present an analysis of Twitter user behavior that can be used to predict the personality of the Big Five Personality using the KNN method. The Important aspect to consider when using this method, namely accuracy in classifying the Big Five Personalities. The experimental results show that the accuracy of the KNN method is 72.09%, which is 95.28% gain above the specified baseline.Personality is the general way a person responds to and interacts with others. Personality is also often defined as the quality that distinguishes individuals. Social media was created to help people communicate remotely and easily. These personalities fall into five categories known as the Big Five personality traits, namely Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism (OCEAN). The use of K-Nearest Neighbour (KNN) is a method of classifying objects based on the training data closest to them. To overcome the data imbalance during training data, we use K-Means SMOTE (Synthetic Minority Oversampling Technique). Other features such as LIWC (Linguistic Inquiry Word Count), Information Gain, Robustly Optimized BERT Approach (RoBERTa), and hyperparameter tuning can improve the performance of the systems we build. The focus of this study is to present an analysis of Twitter user behavior that can be used to predict the personality of the Big Five Personality using the KNN method. The Important aspect to consider when using this method, namely accuracy in classifying the Big Five Personalities. The experimental results show that the accuracy of the KNN method is 72.09%, which is 95.28% gain above the specified baseline
Implementation of Maggot Cage Temperature and Humidity Control Using ESP8266 Based On the Internet of Things
Black Soldier Fly (BSF) is a fly that can produce a maggot or larvae that are useful for human life, like a decomposer waste in the form of composting, animal feed, animal oil production, source of chitin, and for the economic incomes of the society. This study aims to develop a device that can be used to control the maggot cage temperature and humidity using the ESP8266 microcontroller based on the Internet of Things (IoT). The benefit of this study is the utilization of nozzle-based water spraying that can be used to maintain the maggot cage temperature and humidity to improve the quality of maggot cultivation results. In this study, the sensor used to read the temperature and humidity on the maggot cage is DHT11, then used a water spraying method to handle the temperature and humidity controlled by using the ESP8266 and online based on the Blynk IoT platform. This study result shows that the device built in this study can maintain the maggot cage temperature between 28 to 30oC and humidity over 60%.Black Soldier Fly (BSF) is a fly that can produce a maggot or larvae that are useful for human life, like a decomposer waste in the form of composting, animal feed, animal oil production, source of chitin, and for the economic incomes of the society. This study aims to develop a device that can be used to control the maggot cage temperature and humidity using the ESP8266 microcontroller based on the Internet of Things (IoT). The benefit of this study is the utilization of nozzle-based water spraying that can be used to maintain the maggot cage temperature and humidity to improve the quality of maggot cultivation results. In this study, the sensor used to read the temperature and humidity on the maggot cage is DHT11, then used a water spraying method to handle the temperature and humidity controlled by using the ESP8266 and online based on the Blynk IoT platform. This study result shows that the device built in this study can maintain the maggot cage temperature between 28 to 30oC and humidity over 60%