24 research outputs found
Classification of three pathological voices based on specific features groups using support vector machine
Determining and classifying pathological human sounds are still an interesting area of research in the field of speech processing. This paper explores different methods of voice features extraction, namely: Mel frequency cepstral coefficients (MFCCs), zero-crossing rate (ZCR) and discrete wavelet transform (DWT). A comparison is made between these methods in order to identify their ability in classifying any input sound as a normal or pathological voices using support vector machine (SVM). Firstly, the voice signal is processed and filtered, then vocal features are extracted using the proposed methods and finally six groups of features are used to classify the voice data as healthy, hyperkinetic dysphonia, hypokinetic dysphonia, or reflux laryngitis using separate classification processes. The classification results reach 100% accuracy using the MFCC and kurtosis feature group. While the other classification accuracies range between~60% to~97%. The Wavelet features provide very good classification results in comparison with other common voice features like MFCC and ZCR features. This paper aims to improve the diagnosis of voice disorders without the need for surgical interventions and endoscopic procedures which consumes time and burden the patients. Also, the comparison between the proposed feature extraction methods offers a good reference for further researches in the voice classification area
Assessing the effectiveness of data mining tools in classifying and predicting road traffic congestion
Traffic congestion is a significant issue in cities, impacting the environment, commuters, and the economy. Predicting congestion is crucial for efficient network operation, but high-quality data and computational techniques are challenging for scientists and engineers. The revolution of data mining and machine learning has enabled the development of effective prediction methods. Machine learning (ML) approaches have shown potential in predicting traffic congestion, with classification being a key area of study. Open-source software tools WEKA and Orange are used to predict and classify traffic congestion. However, there is no single best strategy for every situation. This study compared the effectiveness of both data mining tools for predicting congestion in one of the areas of the capital of the Hashemite Kingdom of Jordan, Amman, by testing several classifiers including support vector machine (SVM), K-nearest neighbors (KNN), logistic regression (LR), and random forest (RF) classifications. The results showed that the Orange mining tool was superior in predicting traffic congestion, with a prediction accuracy of 100% for Random forest, logistic regression, and 99.8% for KNN. On the other hand, results were better in WEKA for the SVM classifier with an accuracy of 99.7%
Enhancing stroke prediction using the waikato environment for knowledge analysis
State-of-the-art data mining tools incorporate advanced machine learning (ML) and artificial intelligence (AI) models, and it is widely used in classification, association rules, clustering, prediction, and sequential models. Data mining is important for the process of diagnosing and predicting diseases in the early stages, and this contributes greatly to the development of the health services sector. This study utilized classification to predict the stroke of a sample of the patient dataset that was taken from Kaggle. The classification model was created using the data mining program waikato environment for knowledge analysis (WEKA). This data mining tool helped identify individuals most at risk of stroke based on analysis of features extracted from the patient’s dataset. These features were used in classification processes according to the naive Bayes (NB), random forest (RF), support vector machine (SVM), and multi-layer perceptron (MLP) algorithms. Analysis of the classification results of the previous algorithms showed that the SVM outperformed other algorithms in terms of accuracy (94.4%), sensitivity (100%), and F-measure (97.1%). However, the NB algorithm had the best performance in terms of precision (95.7%)
Crack detection based on mel-frequency cepstral coefficients features using multiple classifiers
Crack detection plays an essential role in evaluating the strength of structures. In recent years, the use of machine learning and deep learning techniques combined with computer vision has emerged to assess the strength of structures and detect cracks. This research aims to use machine learning (ML) to create a crack detection model based on a dataset consisting of 2432 images of different surfaces that were divided into two groups: 70% of the training dataset and 30% of the testing dataset. The Orange3 data mining tool was used to build a crack detection model, where the support vector machine (SVM), gradient boosting (GB), naive Bayes (NB), and artificial neural network (ANN) were trained and verified based on 3 sets of features, mel-frequency cepstral coefficients (MFCC), delta MFCC (DMFCC), and delta-delta MFCC (DDMFCC) were extracted using MATLAB. The experimental results showed the superiority of SVM with a classification accuracy of (100%), while for NB the accuracy reached (93.9%-99.9%), and (99.9%) for ANN, and finally in GB the accuracy reached (99.8%)
Enhancing internet of things security: evaluating machine learning classifiers for attack prediction
The internet of things (IoT) has contributed to improving the quality of service and operational efficiency in many areas, such as smart cities, but this technology has faced a major dilemma: the problem of cyber-attacks of various types. In this study, we relied on the use of machine learning (ML) and deep learning (DL) techniques to present a proposed model of an intrusion detection system (IDS) for detecting different types of IoT attacks that include ARP_poisoning, DOS_SYN_Hping, MQTT_Publish, NMAP_FIN_SCAN, NMAP_OS_DETECTION, and Thing_Speak. However, the proposed model is built using Orange3 data mining tools. The model consists of random forest (RF), artificial neural network (ANN), logistic regression (LR), and support vector machine (SVM) classifiers. On the other hand, the data set that is used was obtained from the Kaggle platform's real-time IoT infrastructure data set, called RT-IoT2022. The data set consists of a huge number of records, which are processed and then reduced to 7,481 records using linear discriminant analysis. In the next stage, the data set is fed to the Orange3 data mining tool, which is divided into 70% of the training dataset and 30% of the test dataset, in addition to using fold-cross validation to increase accuracy and avoid overfitting. Thus, the experimental results showed the superiority of RF with a classification accuracy of (99.9%), while the accuracy in ANN reached (99.8%), (97.8%) in LR, and finally, for SVM, the accuracy reached (92.9%)
Driving behavior analytics: an intelligent system based on machine learning and data mining techniques
One of the most common causes of road accidents is driver behavior. To reduce abnormal driver behavior, it must be detected early on. Previous research has demonstrated that behavioral and physiological indicators affect drivers' performance. The goal of this study is to consider the feasibility of classifying driver behavior as either aggressive (sudden left or right turns, accelerating and braking), normal (average driving events) or slow (keeping a lower-than-average speed). Innovation in data mining and machine learning (ML) has allowed for the creation of powerful prediction tools. ML techniques have shown potential in predicting driver behavior, with classification being a critical study area. The data set was gathered using the Kaggle platform. This study classifies driver behavior using Orange3 data mining tools and tests several classifiers, including AdaBoost, CN2 rule inducer, and random forest (RF) classifiers. The results showed that AdaBoost was superior in predicting driver behavior, with 100% accuracy, while the classification accuracy in CN2 rule inducer and RF was 99.8% and 95.4%, respectively. These results demonstrate the possibility of early and highly accurate driver behavior prediction and use it to create a ML-based driver behavior detection system
An automated system for classifying types of cerebral hemorrhage based on image processing techniques
The brain is one of the most important vital organs in the human body. It is responsible for most of the body’s basic activities, such as breathing, heartbeat, thinking, remembering, speaking, and others. It also controls the central nervous system. Cerebral hemorrhage is considered one of the most dangerous diseases that a person may be exposed to during his life. Therefore, the correct and rapid diagnosis of the hemorrhage type is an important medical issue. The innovation in this work lies in extracting a huge number of effective features from computed tomography (CT) images of the brain using the Orange3 data mining technique, as the number of features extracted from each CT image reached (1,000). The proposed system then uses the extracted features in the classification process through logistic regression (LR), support vector machine (SVM), k-nearest neighbor algorithm (KNN), and convolutional neural networks (CNN), which classify cerebral hemorrhage into four main types: epidural hemorrhage, subdural hemorrhage, intraventricular hemorrhage, and intraparenchymal hemorrhage. A total of (1,156) CT images were tested to verify the validity of the proposed model, and the results showed that the accuracy reached the required success level with an average of (97.1%)
Miniaturized L-Shaped and U-Shaped Resonator-Based 8-Bit and 12-Bit Chipless RFID Tag
Chipless Radio Frequency Identification (RFID) technology is a wireless technology that uses radio frequency signals (RF) to identify objects automatically. Chipless RFID tags promise a low-cost and printable solution for item tracking, authentication, and sensing in various applications, including supply chain management and the Internet of Things (IoT). In this work, two compact-sized prototypes of 8-bit chipless radio frequency identification (RFID) tags are modeled as L-shaped and U-shaped resonators. The proposed tags are printed on (14.5mm 14.5mm) Rogers RT5880 substrate with a dielectric constant, and a thickness, . The simulated RCS responses for both the L-shaped and the U-shaped 8-bit chipless RFID tags are presented. A prototype of 12 back-to-back L-shaped resonators with different lengths corresponding to resonant frequencies between 5 and 8.5 GHz is proposed to increase the chipless RFID tag capacity. The proposed 12 back-to-back L-shaped resonators chipless RFID tag is printed on (15mm 25mm) Rogers RT5880 substrate. The simulated RCS response for the compact 12-bit L-shaped resonator-based chipless RFID tag is calculated.
ABSTRAK: Teknologi Pengenalan Frekuensi Radio Tanpa Cip (Chipless RFID) merupakan teknologi tanpa wayar yang menggunakan isyarat frekuensi radio (RF) bagi mengenal pasti objek secara automatik. Tag RFID tanpa wayar menawarkan penyelesaian kos rendah yang berpotensi bagi menjejak item, pengesahan, dan pengesanan dalam pelbagai aplikasi termasuk pengurusan rantaian bekalan dan Internet Benda (IoT). Kajian ini membentangkan dua prototaip bersaiz kompak; iaitu tag RFID 8-bit tanpa wayar yang dimodelkan sebagai resonator berbentuk-L dan berbentuk-U. Tag yang dicadangkan ini dicetak pada substrat Rogers RT5880 (14.5 mm × 14.5 mm) dengan pemalar dielektrik, , dan ketebalan, h = 1.575 mm. Respons simulasi RCS bagi kedua-dua tag 8-bit berbentuk-L dan berbentuk-U dibentangkan dalam kajian ini. Bagi meningkatkan kapasiti tag RFID tanpa wayar, satu prototaip yang terdiri daripada 12 resonator berturutan berbentuk-L dengan panjang berbeza dan berfrekuensi resonan antara 5 hingga 8.5 GHz telah dicadangkan. Tag RFID tanpa wayar ini dicetak pada substrat Rogers RT5880 (15 mm × 25 mm). Respons simulasi RCS bagi tag kompak RFID 12-bit tanpa wayar turut dikira
Detection and classification of pneumonia using the Orange3 data mining tool
A chest X-ray can convey a lot about a patient's condition. However, it requires a specialized and skilled doctor to determine the type of lung disease with high accuracy. Here comes the role of deep learning techniques (DL) and artificial intelligence (AI) in accelerating the process of detecting lung diseases and classifying them with high precision, which saves time and effort for the patient and the doctor alike. This work presents a proposed model for a machine learning (ML) and AI system to analyze chest X-ray images and categorize them into four cases normal, viral pneumonia, bacterial pneumonia, and coronavirus disease 2019 (COVID-19). The system relies on extracting Mel frequency cepstral coefficient (MFCC) features from a dataset consisting of 4,800 chest X-ray images, and then these features are used to train four basic classifiers based on the data mining tool Orange3, which are adaptive boosting (AdaBoost), decision trees (DTs), gradient boosting (GB), and random forest (RF). The model was tested and evaluated, where the AdaBoost classifier excelled with an accuracy of 100%, followed by RF with an accuracy of 99.5%. Finally, GB and DTs came with a classification accuracy of 98.5%, and 97.2%, respectively
Artificial intelligence-powered smart roads: leveraging orange3 for traffic signs recognition
Traffic sign recognition systems are an important concern of advance driver assistance systems (ADAS) and intelligent autonomous vehicles. Recently, many studies have emerged that aim to employ artificial intelligence (AI) and machine learning (ML) to detect and classify traffic signs to improve a system that can be embedded in vehicles to increase efficiency and safety. This work's primary goal is to address traffic sign identification and recognition utilizing a 2,339-image open-source dataset from Kaggle. Our detection model for extracting and classifying traffic sign suggestions is built using Orange3 data mining tools, based on four classifiers random forest (RF), k-nearest neighbors (KNN), decision tree (DT), and adaptive boosting (AdaBoost). Signs are classified into eight categories: don't go signs, go signs, horn signs, roundabout signs, danger signs, crossing signs, speed limit sign, and unallowed signs. The results of examining and evaluating the proposed model based on the performance evaluation metrics showed that RF outperformed with an accuracy rate of 99.8%, followed by AdaBoost with a classification accuracy of 99.2%, and the classification accuracy of DT and KNN was 98.3% and 94.9%, respectively
