5975 research outputs found
Sort by
Ensemble Method for Indonesian Twitter Hate Speech Detection
Due to the massive increase of user-generated web content, in particular on social media networks where anyone can give a statement freely without any limitations, the amount of hateful activities is also increasing. Social media and microblogging web services, such as Twitter, allowing to read and analyze user tweets in near real time. Twitter is a logical source of data for hate speech analysis since users of twitter are more likely to express their emotions of an event by posting some tweet. This analysis can help for early identification of hate speech so it can be prevented to be spread widely. The manual way of classifying out hateful contents in twitter is costly and not scalable. Therefore, the automatic way of hate speech detection is needed to be developed for tweets in Indonesian language. In this study, we used ensemble method for hate speech detection in Indonesian language. We employed five stand-alone classification algorithms, including Naïve Bayes, K-Nearest Neighbours, Maximum Entropy, Random Forest, and Support Vector Machines, and two ensemble methods, hard voting and soft voting, on Twitter hate speech dataset. The experiment results showed that using ensemble method can improve the classification performance. The best result is achieved when using soft voting with F1 measure 79.8% on unbalance dataset and 84.7% on balanced dataset. Although the improvement is not truly remarkable, using ensemble method can reduce the jeopardy of choosing a poor classifier to be used for detecting new tweets as hate speech or not
Classification of The NTEV Problems on The Commercial Building
Neutral to Earth Voltage (NTEV) is one of power quality (PQ) problems in the commercial building that need to be resolved. The classification of the NTEV problems is a method to identify the source types of disturbance in alleviating the problems. This paper presents the classification of NTEV source in the commercial building which is known as the harmonic, loose termination, and lightning. The Euclidean, City block, and Chebyshev variables for K-Nearest Neighbor (K-NN) classifying are being utilized in order to identify the best performance for classifying the NTEV problems. Then, S-Transform (ST) is applied as a pre-processing signal to extract the desired features of NTEV problem for classifier input. Furthermore, the performance of K-NN variables is validated by using the confusion matrix and linear regression. The classification results show that all the K-NN variables capable to identify the NTEV problems. While the K-NN results show that the Euclidean and City block variables are well performed rather than the Chebyshev variable. However, the Chebyshev variable is still reliable as the confusion matrix shows minor misclassification. Then, the linear regression outperformed the percentage close to a perfect value which is hundred percent
IoT: Electrocardiogram (ECG) Monitoring System
Internet of Things (IoT) has many applications in the medical field. With remote-information gathering, healthcare professionals can evaluate, diagnose and treat patients in remote locations using telecommunications technology. This study aimed to develop a small-scale electrocardiogram (ECG) monitoring device that will measure heart rates and waveforms and send the data in a database and a web server. An ECG acquisition device was developed using a single-lead heart rate monitor sensor and an Arduino microcontroller. A program, which will process, analyze and upload the ECG data is coded using MATLAB and C# programs. The collected information is viewed in a Graphical User Interface (GUI) display, coded using C# and in a webpage. Rapid Application Technology (RAD) was used in the methodology, which began with a quick design of the system. The hardware and software systems underwent a prototyping cycle for development. Once finished, the integration of the system is conducted to construct a complete IoT-based ECG monitoring system. For testing using t-test, a sample size of 18 and a a= 0.05 is used. Testing resulted into t-test values that lie in the non-critical zone for all ECG parameters, denoting that there is no significant difference between the gathered data. The device’s percent reliability in detecting ECG conditions such as normal sinus rhythm, sinus tachycardia, sinus bradycardia and flatline, is 83.33%. The percent difference for the heart rate is 0.35 %, which falls within the acceptable medical standard of 99% accuracy. The device was deemed functional and reliable
Performance Investigation of Coaxial Cable with Transmission Line Parameters Based on Lossy Dielectric Medium
This paper presents the analysis of high performance for coaxial cable with transmission line parameters. The modeling for performance of coaxial cable contains many parameters, in this paper will discuss the more effective parameter is the type of dielectric mediums. This analysis of the performance related to dielectric mediums with respect to dielectric losses and its effect upon cable properties, dielectrics versus characteristic impedance, and the attenuation in the coaxial line for different dielectrics. The analysis depends on a simple mathematical model for coaxial cables to test the influence of the insulators (Dielectrics) performance
An Efficient Schema of a Special Permutation Inside of Each Pixel of an Image for its Encryption
The developments of communications and digital transmissions have pushed the data encryption to grow quickly to protect the information, against any hacking or digital plagiarisms. Many encryption algorithms are available on the Internet, but it's still illegal to use a number of them. Therefore, the search for new the encryption algorithms is still current. In this work, we will provide a preprocessing of the securisation of the data, which will significantly enhance the crypto-systems. Firstly, we divide the pixel into two blocks of 4 bits, a left block that contains the most significant bit and another a right block which contains the least significant bits and to permute them mutually. Then make another permutation for each of group. This pretreatment is very effective, it is fast and is easy to implement and, only consumes little resource
Measuring the Road Traffic Intensity using Neural Network with Computer Vision
Traffic congestion plagues all driver around the world. To solve this problem computer vision can be used as a tool to develop alternative routes and eliminate traffic congestions. In the current generation with increasing number of cameras on the streets and lower cost for Internet of Things(IoT) this solution will have a greater impact on current systems. In this paper, the Macroscopic Urban Traffic model is used using computer vision as its source and traffic intensity monitoring system is implemented. The input of this program is extracted from a traffic surveillance camera and another program running a neural network classification which can classify and distinguish the vehicle type is on the road. The neural network toolbox is trained with positive and negative input to increase accuracy. The accuracy of the program is compared to other related works done and the trends of the traffic intensity from a road is also calculated
Educational Data Mining and Analysis of Students’ Academic Performance Using WEKA
In this competitive scenario of the educational system, the higher education institutes use data mining tools and techniques for academic improvement of the student performance and to prevent drop out. The authors collected data from three colleges of Assam, India. The data consists of socio-economic, demographic as well as academic information of three hundred students with twenty-four attributes. Four classification methods, the J48, PART, Random Forest and Bayes Network Classifiers were used. The data mining tool used was WEKA. The high influential attributes were selected using the tool. The internal assessment attribute in the continuous evaluation process makes the highest impact in the final semester results of the students in our dataset. The results showed that random forest outperforms the other classifiers based on accuracy and classifier errors. Apriori algorithm was also used to find the association rule mining among all the attributes and the best rules were also displayed
To Improve Feature Extraction and Opinion Classification Issues in Customer Product Reviews Utilizing an Efficient Feature Extraction and Classification (EFEC) Algorithm
Currently, customer's product review opinion plays an essential role in deciding the purchasing of the online product. A customer prefers to acquire the opinion of other customers by viewing their opinion during online products' reviews, blogs and social networking sites, etc. The majority of the product reviews including huge words. A few users provide the opinion; it is tough to analysis and understands the meaning of reviews. To improve user fulfillment and shopping experience, it has become a general practice for online sellers to allow their users to review or to communicate opinions of the products that they have sold. The major goal of the paper is to solve feature extraction problem and opinion classification problem from customers utilized product reviews which extract the feature words and opinion words from product reviews. To propose an Efficient Feature Extraction and Classification (EFEC) algorithm is implementing to extracts a feature from opinion words. The reviewer usually marks both positive and negative parts of the reviewed product, despite the fact that their general opinion on the product may be positive or negative. An EFEC algorithm is utilized to predict the number of positive and negative opinion in reviews. Based on Experimental evaluations, proposed algorithm improves accuracy 15.05%, precision 13.7%, recall 15.59% and F-measure 15.07% of the proposed system compared than existing methodologie
Electromagnetic Radiation (EMR) of Human Body Before and After Jogging
The research is on the electromagnetic radiation of human body before and after jogging. 30 healthy students from UiTM with an age range of 23-25 years old volunteered. The seven locations of chakra points were measured. The body frequency (in MHz) is captured using frequency detector by taking the reading of the frequency 5 times at each point at the same location; hence, the average value is calculated for data analysis. This frequency measurement is recorded two times which is before and after jogging with a consistence protocol for all participants. The data in terms of frequency (Hertz) is converted into 15 colours of bio-energies representing the health level. The finding shows that 63.3% of participants’ health level improved after jogging. While 33.3% of participants had decrement in their health level. The results also indicate improvement in bio-energies score for five out of seven chakra points after jogging
Solving N-Queens Problem Using Subproblems based on Genetic Algorithm
Nowadays, permutation problems with large state spaces and the path to solution is irrelevant such as N-Queens problem has the same general property for many important applications such as integrated-circuit design, factory-floor layout, job-shop scheduling, automatic programming, telecommunications network optimization, vehicle routing, and portfolio management. Therefore, methods which are able to find a solution are very important. Genetic algorithm (GA) is one the most well-known methods for solving N-Queens problem and applicable to a wide range of permutation problems. In the absence of specialized solution for a particular problem, genetic algorithm would be efficient. But holism and random choices cause problem for genetic algorithm in searching large state spaces. So, the efficiency of this algorithm would be demoted when the size of state space of the problem grows exponentially. In this paper, the subproblems used based on genetic algorithm to cover this weakness. This proposed method is trying to provide partial view for genetic algorithm by locally searching the state space. This method works to take shorter steps toward the solution. To find the first solution and other solutions in N-Queens problem using proposed method: dividing N-Queens problem into subproblems, which configuring initial population of genetic algorithm. The proposed method is evaluated and compares it with two similar methods that indicate the amount of performance improvement