32 research outputs found

    Feature Fusion of Deep Spatial Features and Handcrafted Spatiotemporal Features for Human Action Recognition

    No full text
    Human action recognition plays a significant part in the research community due to its emerging applications. A variety of approaches have been proposed to resolve this problem, however, several issues still need to be addressed. In action recognition, effectively extracting and aggregating the spatial-temporal information plays a vital role to describe a video. In this research, we propose a novel approach to recognize human actions by considering both deep spatial features and handcrafted spatiotemporal features. Firstly, we extract the deep spatial features by employing a state-of-the-art deep convolutional network, namely Inception-Resnet-v2. Secondly, we introduce a novel handcrafted feature descriptor, namely Weber’s law based Volume Local Gradient Ternary Pattern (WVLGTP), which brings out the spatiotemporal features. It also considers the shape information by using gradient operation. Furthermore, Weber’s law based threshold value and the ternary pattern based on an adaptive local threshold is presented to effectively handle the noisy center pixel value. Besides, a multi-resolution approach for WVLGTP based on an averaging scheme is also presented. Afterward, both these extracted features are concatenated and feed to the Support Vector Machine to perform the classification. Lastly, the extensive experimental analysis shows that our proposed method outperforms state-of-the-art approaches in terms of accuracy

    A guided random forest based feature selection approach for activity recognition

    No full text
    Selection of relevant and non-redundant features for human activity classification is a crucial task in human activity recognition studies, where researchers try to identify the minimum number of features that can still achieve good classification performance. In this paper, we propose a feature selection method based on guided random forest in the context of human activity recognition. In order to select small set of important features using the guided random forest, we first train an ordinary random forests on the dataset for collecting the feature importance scores, and then, inject the collected importance scores to influence the feature selection process in the guided random forest. The proposed guided random forest has a number of key advantages such as trees in the guided random forest are completely independent from one-another, allows parallel computing, low computational cost as it grows only two ensembles, and can select a very small set of high quality features without losing classification accuracy. Using five benchmark datasets, we show that the guided random forest can select compact feature subsets compared to previously proposed methods, while preserving recognition accuracy. We also apply random forests classifier to evaluate the effectiveness of the feature selection methods and demonstrate that random forests with feature subset selected by guided random forest performs comparatively better than feature subsets selected by other feature selection methods.</p

    Human activity recognition from wearable sensors using extremely randomized trees

    No full text
    Learning and recognizing the physical activities of human based on wearable sensor has a wide range of applications in many fields such as assistive healthcare and security surveillance. In this paper, we propose an activity recognition framework based on extremely randomized trees and guided random forest to recognize both simple and complex activities from wearable sensor. In order to recognize different activities using extremely randomized trees, we first select most important features from all available features applying the feature selection method, namely, guided random forest; then, selected features are used to build the classifier for classifying activities. The proposed framework is extremely efficient in terms of recognition performance and computational time as it can recognize both small and large set of activities very accurately with different number of features in different sensor settings, while it needs fairly small amount of time for training and classification. The evaluation results of the experiments conducted on four benchmark data sets indicate that the proposed technique performs better than the classic activity recognition systems with respect to recognition accuracy and computational time; the proposed approach yielded the maximum recognition rate of 99.6%.</p

    Dynamic Facial Emotion Recognition Using Deep Spatial Feature and Handcrafted Spatiotemporal Feature on Spark

    No full text
    One important challenge of dynamic facial emotion recognition is to effectively obtain the spatial and dynamic change of face structure from videos. Besides, there is an increasing demand for distributed computing of videos, as a result of the speedy production of videos from numerous multimedia sources. To address the above issues, in this work, we propose a novel method for dynamic facial emotion recognition on top of Spark. Furthermore, we introduce an effective dynamic feature descriptor namely, Volume Symmetric Local Graph Structure (VSLGS), which extracts the spatiotemporal features. We also utilize the convolutional neural network (CNN) to obtain deep spatial features. Lastly, these obtained features are concatenated and fed to Spark MLlib Multilayer Perceptron (MLP) classifier to recognize the dynamic facial emotions. An extensive experimental investigation is performed to prove the effectiveness of our method over state-of-the-art methods. Furthermore, we also showed the scalability of the proposed method experimentally.</p

    A multimodal multitask deep learning framework for vibrotactile feedback and sound rendering

    No full text
    Data-driven approaches are often utilized to model and generate vibrotactile feedback and sounds for rigid stylus-based interaction. Nevertheless, in prior research, these two modalities were typically addressed separately due to challenges related to synchronization and design complexity. To this end, we introduce a novel multimodal multitask deep learning framework. In this paper, we developed a comprehensive end-to-end data-driven system that encompasses the capture of contact acceleration signals and sound data from various texture surfaces. This framework introduces novel encoder-decoder networks for modeling and rendering vibrotactile feedback through an actuator while routing sound to headphones. The proposed encoder-decoder networks incorporate stacked transformers with convolutional layers to capture both local variability and overall trends within the data. To the best of our knowledge, this is the first attempt to apply transformer-based data-driven approach for modeling and rendering of vibrotactile signals as well as sounds during tool-surface interactions. In numerical evaluations, the proposed framework demonstrates a lower RMS error compared to state-of-the-art models for both vibrotactile signals and sound data. Additionally, subjective similarity evaluation also confirm the superiority of proposed method over state-of-the-art

    A Distributed Automatic Video Annotation Platform

    No full text
    In the era of digital devices and the Internet, thousands of videos are taken and share through the Internet. Similarly, CCTV cameras in the digital city produce a large amount of video data that carry essential information. To handle the increased video data and generate knowledge, there is an increasing demand for distributed video annotation. Therefore, in this paper, we propose a novel distributed video annotation platform that explores the spatial information and temporal information. Afterward, we provide higher-level semantic information. The proposed framework is divided into two parts: spatial annotation and spatiotemporal annotation. Therefore, we propose a spatiotemporal descriptor, namely, volume local directional ternary pattern-three orthogonal planes (VLDTP&ndash;TOP) in a distributed manner using Spark. Moreover, we developed several state-of-the-art appearance-based and spatiotemporal-based feature descriptors on top of Spark. We also provide the distributed video annotation services for the end-users so that they can easily use the video annotation and APIs for development to produce new video annotation algorithms. Due to the lack of a spatiotemporal video annotation dataset that provides ground truth for both spatial and temporal information, we introduce a video annotation dataset, namely, STAD which provides ground truth for spatial and temporal information. An extensive experimental analysis was performed in order to validate the performance and scalability of the proposed feature descriptors, which proved the excellence of our proposed approach

    Face Sketch Image Generation from Facial Attributes Using StyleGAN2

    No full text
    During forensic investigations, composite sketches are frequently used to track down suspects when photographic evidence is unavailable. Descriptions from victims or eyewitnesses are used to create these sketches. Forensic artists are essential due to the sketches’ investigative and prosecutorial applications, but their work is expensive and time-consuming. Existing systems for generating face sketches from facial attributes have shown reasonable performance but struggle with limited training samples. This work uses data augmentation, specifically adaptive discriminator augmentation (ADA), to mitigate these challenges. Our framework uses a pre-trained text encoder, bidirectional LSTM, to encode descriptions into sentence embeddings, and a style-based generative adversarial network, StyleGAN2, fine-tuned for sketch generation. Extensive experiments demonstrate that this approach outperforms existing state-of-the-art models.</p

    Fusion of Deep Learned and Handcrafted Features for Paddy Disease Recognition

    No full text
    Rice is a fundamental food grain worldwide, playing a vital role in both agriculture and public health. However, rice leaf diseases significantly threaten cultivation, affecting farmers globally. Early identification and effective management of these diseases are critical for ensuring healthy rice crops and sufficient food supply for the growing population. Traditional manual diagnosis of paddy diseases remains prevalent but is often inefficient, time-consuming, and susceptible to errors. To address this, our study introduces a novel end-to-end framework for accurately diagnosing paddy diseases through advanced image analysis of paddy leaves. This approach combines deep learning and handcrafted feature extraction techniques. The InceptionResNetV2 pre-trained network is employed to extract deep features from each image, while the Local Neighborhood Encoded Pattern (LNEP) captures texture features. These combined features are then used to identify discriminative patterns, which are fed into a multi-scale 1D Convolutional Neural Network (CNN) classifier. Extensive investigations performed on the Paddy Doctor dataset reveal that the proposed method exhibits promising performance in comparison to state-of-the-art methods

    An integrated approach to classify gender and ethnicity

    No full text
    Faces express many social indications, including gender, ethnicity, age, expression and identity, most of them have drawn thriving attention from various research communities, for instance neuroscience, computer science and psychology. In this paper, we propose a new approach to classify gender and ethnicity by merging both texture and shape features extracted from face images. Gabor filter is used to extract the texture features and histogram of oriented gradients (HOG) is used to extract the shape features from face images. In order to achieve higher performance we combined both texture and shape features. After combining, the size of feature vector obtained is in a high dimension, to decrease the dimensionality Kernel PCA has been implemented. Finally, to classify the gender and ethnicity we used Support Vector Machine. The experimental result shows the effectiveness of proposed framework.</p

    Hand sign language recognition for Bangla alphabet using Support Vector Machine

    No full text
    The sign language considered as the main language for deaf and dumb people. So, a translator is needed when a normal person wants to talk with a deaf or dumb person. In this paper, we present a framework for recognizing Bangla Sign Language (BSL) using Support Vector Machine. The Bangla hand sign alphabets for both vowels and consonants have been used to train and test the recognition system. Bangla sign alphabets are recognized by analyzing its shape and comparing its features that differentiates each sign. In proposed system, hand signs are first converted to HSV color space from RGB image. Then Gabor filters are used to acquire desired hand sign features. Since feature vector obtained using Gabor filter is in a high dimension, to reduce the dimensionality a nonlinear dimensionality reduction technique that is Kernel PCA has been used. Lastly, Support Vector Machine (SVM) is employed for classification of candidate features. The experimental results show that our proposed method outperforms the existing work on Bengali hand sign recognition.</p
    corecore