JOIV : International Journal on Informatics Visualization
Not a member yet
786 research outputs found
Sort by
A Framework of Forensic Analysis and Visualization: Using WhatsApp Chat Data as a Case Study
Digital forensic analysis involves studying and analyzing acquired evidence artifacts using methodical approaches. However, unstructured data could be time-consuming and difficult in the forensic examination phase. Automation in digital forensic processes has recently been seen as a potential solution to improve analysis processes. Therefore, we propose a forensic analysis and visualization framework via exploratory data analysis (EDA) using WhatsApp chat datasets as a case study. Univariate and multivariate EDA visualization models were applied to the datasets. The framework's utility was demonstrated through forensic analysis simulation scenarios: linkage (interaction) and attribution (who was responsible). origination (evaluation of source), and sequencing (timeline). It was conducted in a controlled experiment environment using Python scripting. The aim is to test the extent to which EDA visualization models can visualize complete and accurate artifacts based on the scenarios. Our evidence-based findings demonstrated the suitability of specific univariate and multivariate in visualizing complete and accurate data. The framework was able to visualize key metadata such as incoming and outgoing chats, sender identification, communication timeline, and shared media. The findings suggested that the EDA approach aligns with forensic analysis, as it helps describe investigative clues by analyzing data patterns. Additionally, an expert review was conducted, in which the experts confirmed the adequacy of the simulation scenarios and the usefulness of the forensic visualization. Furthermore, the results of this study could aid in presenting evidence in a court of law
A Survey on Forms of Visualization and Tools Used in Topic Modelling
In this paper, we surveyed recent publications on topic modeling and analyzed the forms of visualizations and tools used. Expectedly, this information will help Natural Language Processing (NLP) researchers to make better decisions about which types of visualization are appropriate for them and which tools can help them. This could also spark further development of existing visualizations or the emergence of new visualizations if a gap is present. Topic modeling is an NLP technique used to identify topics hidden in a collection of documents. Visualizing these topics permits a faster understanding of the underlying subject matter in terms of its domain. This survey covered publications from 2017 to early 2022. The PRISMA methodology was used to review the publications. One hundred articles were collected, and 42 were found eligible for this study after filtration. Two research questions were formulated. The first question asks, "What are the different forms of visualizations used to display the result of topic modeling?" and the second question is "What visualization software or API is used? From our results, we discovered that different forms of visualizations meet different purposes of their display. We categorized them as maps, networks, evolution-based charts, and others. We also discovered that LDAvis is the most frequently used software/API, followed by the R language packages and D3.js. The primary limitation of this survey is it is not exhaustive. Hence, some eligible publications may not be included
Cheating Detection for Online Examination Using Clustering Based Approach
Online exams have become increasingly popular due to their convenience in eliminating the need for physical exams and allowing students to take exams from remote locations. However, one of the drawbacks of online exams is that they make cheating easier, and it can be difficult for online proctoring to detect subtle movements by the students. This could lead to doubts about students' exam results' value and overall credibility. To address this pressing issue, we present a cheating detection method using a CCTV camera to monitor students' faces, eyes, and devices to determine whether they cheat during exams. If suspicious behavior indicative of cheating is detected, a warning is raised to alert the students. A custom dataset was developed to train the model. The dataset consisted of recordings of pre-determined cheating behavior by 50 participants. These videos captured various poses and behaviors encoded and analyzed using a clustering approach. The encoded clustering method continuously tracks the students' faces, eyes, and body gestures throughout an exam. Experimental results show that the proposed approach effectively detects cheating behavior with a favorable accuracy of 83%. The proposed method offers a promising solution to the growing concern about cheating in online exams. This approach can significantly enhance the integrity and reliability of online assessment processes, fostering trust among educational institutions and stakeholders
Evaluation of Cryptocurrency Price Prediction Using LSTM and CNNs Models
Cryptocurrencies created by Nakamoto in 2009 have gained significant interest due to their potential for high returns. However, the cryptocurrency market's unpredictability makes it challenging to forecast prices accurately. To tackle this issue, a deep learning model has been developed that utilizes Long Short-Term Memory (LSTM) neural networks and Convolutional Neural Networks (CNNs) to predict cryptocurrency prices. LSTMs, a type of recurrent neural network, are well-suited for analyzing time series data and have been successful in various prediction applications. Additionally, CNNs, primarily used for image analysis tasks, can be employed to extract relevant patterns and characteristics from input data in Bitcoin price prediction applications. This study contributes to the existing related works on cryptocurrency price prediction by exploring various predictive models and techniques, which involve a machine learning model, deep learning model, time series analysis, and as well as a hybrid model that combines deep learning methods to predict cryptocurrency prices as well as enhance the accuracy and reliability of the price predictions. To ensure accurate predictions in this study, a trustworthy dataset from investing.com was sought. The dataset, sourced from investing.com, consists of 1826 time series data samples. The dataset covers the time frame from January 1, 2018, to December 31, 2022, providing data for a period of 5 years. Subsequently, pre-processing was conducted on the dataset to guarantee the quality of the input. As a result of absent values and concerns regarding the dataset's obsolescence, an alternative dataset was sourced to avoid these issues. The performance of the LSTM and CNN models was evaluated using root mean squared error (RMSE), mean squared error (MSE), mean absolute error (MAE) and R-squared (R2). It was observed that they outperformed each other to a certain degree in short-term forecasts compared to long-term predictions, where the R2Â values for LSTM range from 0.973 to 0.986, while for CNNs, they range from 0.972 to 0.988 for 1 day, 3 days and 7 days windows length. Nevertheless, the LSTM model demonstrated the most favorable performance with the lowest error rate. The RMSE values for the LSTM model ranged from 1203.97 to 1645.36, whereas the RMSE values for the CNNs model ranged from 1107.77 to 1670.93. As a result, the LSTM model exhibited a lower error rate in RMSE and achieved the highest accuracy in R2Â compared to the CNNs model. Considering these comparative outcomes, the LSTM model can be deemed as the most suitable model for this specific cas
Identification of Coffee Types Using an Electronic Nose with the Backpropagation Artificial Neural Network
Coffee is one of the famous plants’ commodities in the world. There are some coffee powders such as Arabica dan Robusta. This study aimed to identify two various coffee powders, Arabica and Robusta based on the blended aroma profiles, employing the backpropagation Artificial Neural Network (ANN). Four taste sensors were employed, namely TGS 2602, 2610, 2611, and 2620, to capture the diverse coffee aroma. These detectors were combined with the aroma sensors having transducers integrated with signal amplifiers or processors, which featured a load of 10 KΩ resistance. Three aroma types were investigated, namely Arabica coffee, Robusta coffee, and without coffee beans. The neural network architecture consisted of four inputs from all sensors, with one hidden layer housing eight neurons. Two neuron outputs were employed for classification, with 70 samples used for training ANN for each type. During the training phase, the developed neural network showed an impressive accuracy rate of 91.90%. TGS 2602 and 2611 sensors showed the most significant differences among the three aroma types. When analyzing ground Robusta coffee, TGS 2602 and 2611 sensors recorded 2.967 volts and 1.263 volts, with a gas concentration of 17.92 ppm and 2441.8 ppm. Similarly, the sensors for ground Arabica coffee displayed 3.384 volts and 1.582 volts with a gas concentration of 20.445 ppm and 3058.5 ppm in both TGS 2602 and 2611, respectively. The implemented ANN with aroma sensor as input successfully identify the coffee powders
Modified LeNet-5 Architecture to Classify High Variety of Tourism Object: A Case Study of Tourism Object for Education in Tinalah Village
This research aims to modify a CNN (Convolutional Neural Network) based on LeNet-5 to reduce overfitting in a Tinalah Tourism Village dataset object detection. Tinalah Tourism Village has many objects that can be identified for tourism education and enhanced tourist experience. While these objects, spread across the different sites of Tinalah do vary, some share similarities in their histogram patterns. Visually, if the size of a picture is reduced in the LeNet-5 ‘preferred size’ feature, it will inevitably lose some of its information, making pictures too similar reducing accuracy. In order to learn and classify objects, this research performs a modification on LeNet-5 architecture to provide a better performance geared toward larger input imaging. The previous state-of-the-art architecture showed an overfitting performance where the training accuracy performed too much better than the testing accuracy in our dataset. We brought in a dropout layer to reduce overfitting, increase the dense layer's size, and add a convolution layer. We then compared the modified LeNet-5 with other state-of-the art architecture, such as LeNet-5 and AlexNet. Results showed that a modified LeNet-5 outperformed other architectures, especially in performing accuracy for testing the Tinalah dataset, reaching 0.913 or (91,3 %). This research discusses the dataset, the modified LeNet-5 architecture, and performance comparison between state-of-the-art CNN architecture. Our CNN architecture can be developed by involving a transfer learning mechanism to provide greater accuracy for further research
Chest X-ray Image Classification to Identify Lung Diseases Using Convolutional Neural Network and Convolutional Block Attention Module
Image classification, the process of categorizing and labeling groups of pixels or vectors within an image based on specific rules, is continuously developed by many researchers in the world to solve many problems. One of those problems is x-ray image classification to determine lung diseases. This research tries to solve the problem of classifying COVID-19, pneumonia, and healthy lungs using x-ray images. The image datasets were collected from several sources. This research aims to build a reliable and robust Convolutional Neural Network (CNN) enhanced with Convolutional Block Attention Module (CBAM) mechanism. CNN is used to do the feature extraction and the classification, whereas CBAM is used to improve the performance of the CNN by focusing on the important features in given data. Research methods are done through extensive data selection, preprocessing, and parameter tuning to achieve a well-performing model. While there is still a lack of research on x-ray classification using the attention mechanism, this research proposes it as the main method. This research also does a further experiment on the effect of the imbalanced dataset on the model. The evaluation is done using a cross-validation method. This research results reach 97.74% of accuracy, precision, recall, and f1-score. This research concludes that CBAM increases the performance of a CNN module. Using a larger dataset can be beneficial in this kind of research as well as evaluation by radiologists
A Detection and Response Architecture for Stealthy Attacks on Cyber-Physical Systems
There has been an increased reliance on interconnected Cyber-Physical Systems (CPS) applications. This reliance has caused tremendous growth in high assurance challenges. Due to the functional interdependence between the internal systems of CPS applications, the utilities' ability to reliably provide services could be disrupted if security threats are not addressed. To address this challenge, we propose a multi-level, multi-agent detection and response architecture built on the formalisms of Hidden Markov Models (HMM) and Markov Decision Processes (MDP). We have evaluated the performance of the proposed architecture on one of the critical smart grid applications, Advanced Metering Infrastructure (AMI). This paper utilizes a simulation tool called SecAMI for performance evaluation. A Stealthy attack scenario contains multiple distinct multi-stage attacks deployed concurrently in a network to compromise the system and stop several critical services in a CPS. The results show that the proposed architecture effectively detects and responds to stealthy attack scenarios against Cyber-Physical Systems. In particular, the simulation results show that the proposed system can preserve the availability of more than 93% of the AMI network under stealthy attacks. A future study may evaluate the effectiveness of various stealthy attack strategies and detection and response systems. The high availability of any AMI should be protected against new attack techniques. The proposed system will also determine a distributed IDS's efficient placement for intrusion detection sensors and response nodes within an AMI
Multi-Temporal Factors to Analyze Indonesian Government Policies regarding Restrictions on Community Activities during COVID-19 Pandemic
Concerning the implementation of the government policy regarding the Restriction of Community Activities (PPKM) during the COVID-19 pandemic era, there are still discrepancies in the economic sector and population mobility. This issue emerges due to irrelevant data and information in one region of Indonesia. The data differences should be carefully solved when implementing the PPKM policy. Besides, the PPKM must also pay attention to some specific factors related to the real conditions of a region, such as the data on the epidemiology of COVID-19, economic situations, and population mobility. These three are called Multi Factors. Then, based on the data, COVID-19 has a specific spreading period that cannot be repeated and thus is called temporal. Therefore, using the Multi-Temporal Factors approach to identify their correlation with the PPKM policy by applying Machine Learning, such as the Multiple Linear Regression model and Dynamic Factors, is essential. This research aims to analyze the characteristics and correlations of the COVID-19 pandemic data and the effectiveness of the government's policy on community activities (PPKM) based on the data quality. The results show that the accuracy of the multiple linear regression models is 84%. The Dynamic Factor shows that the five most important factors are idr_close, positive, retail_recreation, station, and healing. Based on the ANOVA test, all independent variables significantly influence the dependent one. The linear multiple regression models do not display any symptoms of heteroscedasticity. Thus, based on the data quality, the implementation of PPKM by the government has a practical impact
Mapping User Experience Information Overload Problems Across Disciplines
User Experience (UX) has been increasing linearly with the systems and digital media. UX concept describes a human factor as an experience with the life cycle of digital technology. UX increases the usability of the product in the industry more than functionality. Interest in UX has produced a huge amount of product and research articles. Moreover, this interdisciplinary topic becomes increased significantly because of the wider applications. However, this benefit become a problem due to the number of publications. The information overload problem is the result of the increasing UX topic. Several researchers solved this problem with qualitative analysis, but it cannot solve the overload problem. In this paper, we purposed bibliometric analysis and research profiling to interpret UX information on the map, with the publications from 1998-2022, a dataset compiled in RIS format to provide article metadata. As a result, the UX information map from the topic with the five clusters. Therefore, to provide information on the topic's coherence, we propose a coupling network. A related topic is shown as a link; a direct link means high coherence between topics. The analysis was carried out using the 5W1H approach (what, where, who, when, why, and how). The results show that UX is indeed an interdisciplinary field, especially with a design approach and user experience. In addition, to determine novice researchers, determining the focus of research can be done by taking into account previous research goals and maps