1,721,129 research outputs found

    Nonlinear Projection Trick in Kernel Methods: An Alternative to the Kernel Trick

    No full text
    In kernel methods such as kernel principal component analysis (PCA) and support vector machines, the so called kernel trick is used to avoid direct calculations in a high (virtually infinite) dimensional kernel space. In this brief, based on the fact that the effective dimensionality of a kernel space is less than the number of training samples, we propose an alternative to the kernel trick that explicitly maps the input data into a reduced dimensional kernel space. This is easily obtained by the eigenvalue decomposition of the kernel matrix. The proposed method is named as the nonlinear projection trick in contrast to the kernel trick. With this technique, the applicability of the kernel methods is widened to arbitrary algorithms that do not use the dot product. The equivalence between the kernel trick and the nonlinear projection trick is shown for several conventional kernel methods. In addition, we extend PCA-L1, which uses L-1-norm instead of L-2-norm (or dot product), into a kernel version and show the effectiveness of the proposed approach.Y

    Principal Component Analysis by L-p-Norm Maximization

    No full text
    This paper proposes several principal component analysis (PCA) methods based on L-p-norm optimization techniques. In doing so, the objective function is defined using the L-p-norm with an arbitrary p value, and the gradient of the objective function is computed on the basis of the fact that the number of training samples is finite. In the first part, an easier problem of extracting only one feature is dealt with. In this case, principal components are searched for either by a gradient ascent method or by a Lagrangian multiplier method. When more than one feature is needed, features can be extracted one by one greedily, based on the proposed method. Second, a more difficult problem is tackled that simultaneously extracts more than one feature. The proposed methods are shown to find a local optimal solution. In addition, they are easy to implement without significantly increasing computational complexity. Finally, the proposed methods are applied to several datasets with different values of p and their performances are compared with those of conventional PCA methods.N

    Implementing Kernel Methods Incrementally by Incremental Nonlinear Projection Trick

    No full text
    Recently, the nonlinear projection trick (NPT) was introduced enabling direct computation of coordinates of samples in a reproducing kernel Hilbert space. With NPT, any machine learning algorithm can be extended to a kernel version without relying on the so called kernel trick. However, NPT is inherently difficult to be implemented incrementally because an ever increasing kernel matrix should be treated as additional training samples are introduced. In this paper, an incremental version of the NPT (INPT) is proposed based on the observation that the centerization step in NPT is unnecessary. Because the proposed INPT does not change the coordinates of the old data, the coordinates obtained by INPT can directly be used in any incremental methods to implement a kernel version of the incremental methods. The effectiveness of the INPT is shown by applying it to implement incremental versions of kernel methods such as, kernel singular value decomposition, kernel principal component analysis, and kernel discriminant analysis which are utilized for problems of kernel matrix reconstruction, letter classification, and face image retrieval, respectively.Y

    Feature-Level Ensemble Knowledge Distillation for Aggregating Knowledge from Multiple Networks

    No full text
    Knowledge Distillation (KD) aims to transfer knowledge in a teacher-student framework, by providing the predictions of the teacher network to the student network in the training stage to help the student network generalize better. It can use either a teacher with high capacity or an ensemble of multiple teachers. However, the latter is not convenient when one wants to use feature-map-based distillation methods. In this paper, we empirically show that using several non-linear transformation layer cope well with multiple-teacher setting compared to other kinds of feature-map-level distillation methods. Comprehensively, this paper proposes a versatile and powerful training algorithm named FEature-level Ensemble knowledge Distillation (FEED), which aims to transfer the ensemble knowledge using multiple teacher networks. In this study, we introduce a couple of training algorithms that transfer ensemble knowledge to the student at the feature-map-level. Among the feature-map-level distillation methods, using several non-linear transformations in parallel for transferring the knowledge of the multiple teachers helps the student find more generalized solutions. We name this method as parallel FEED, and experimental results on CIFAR-100 and ImageNet show that our method has clear performance enhancements, without introducing any additional parameters or computations at test time. We also show the experimental results of sequentially feeding teacher's information to the student, hence the name sequential FEED, and discuss the lessons obtained. Additionally, the empirical results on measuring the reconstruction errors at the feature map give hints for the enhancements.N

    Backdoor Attacks in Federated Learning by Rare Embeddings and Gradient Ensembling

    No full text
    Recent advances in federated learning have demonstrated its promising capability to learn on decentralized datasets. However, a considerable amount of work has raised concerns due to the potential risks of adversaries participating in the framework to poison the global model for an adversarial purpose. This paper investigates the feasibility of model poisoning for backdoor attacks through rare word embeddings of NLP models. In text classification, less than 1% of adversary clients suffices to manipulate the model output without any drop in the performance on clean sentences. For a less complex dataset, a mere 0.1% of adversary clients is enough to poison the global model effectively. We also propose a technique specialized in the federated learning scheme called Gradient Ensemble, which enhances the backdoor performance in all our experimental settings.N

    Principal component analysis based on L1-norm maximization

    No full text
    A method of principal component analysis (PCA) based on a new L1-norm optimization technique is proposed. Unlike conventional PCA, which is based on L2-norm, the proposed method is robust to outliers because it utilizes the L1-norm, which is less sensitive to outliers. It is invariant to rotations as well. The proposed L1-norm optimization technique is intuitive, simple, and easy to implement. It is also proven to find a locally maximal solution. The proposed method is applied to several data sets and the performances are compared with those of other conventional methods.N

    Cultural Event Recognition by Subregion Classification with Convolutional Neural Network

    No full text
    In this paper, a novel cultural event classification algorithm based on convolutional neural networks is proposed. The proposed method firstly extracts regions that contain meaningful information.. Then, convolutional neural networks are trained to classify the extracted regions. The final classification of a scene is performed by combining the classification results of each extracted region of the scene probabilistically. Compared to the state-of-the-art methods for classifying Chalearn Looking at People cultural event recognition database, the proposed methods shows competitive results.Y

    Human detection by neural networks using a low-cost short-range Doppler radar sensor

    No full text
    In this paper, we propose the human detection technique using Neural Networks to effectively classify the Doppler signals caused by human walking along with the background noise sources. The frequency or phase feature vectors converted from the given input signal are directly used as the input of Neural Networks. In addition, Gaussian noise is added in the input nodes of Neural Network in order to prevent the overfitting problem. We developed the low-cost & short-range K-band Doppler radar for the experiment. The proposed technique was examined with human walking data accompanied with the background noises caused by the fan, rain, snow, and other outdoor environmental factors. The trained Neural Network detection technique can detect human walking with 95.2% of the true positive rate and it has 4.6% of the false positive rate.N

    Feature Extraction with Weighted Samples Based on Independent Component Analysis

    No full text
    This study investigates a new method of feature extraction for classification problems with a considerable amount of outliers. The method is a weighted version of our previous work based on the independent component analysis (ICA). In our previous work, ICA was applied to feature extraction for classification problems by including class information in the training. The resulting features contain much information on the class labels producing good classification performances. However, in many real world classification problems, it is hard to get a clean dataset and inherently, there may exist outliers or dubious data to complicate the learning process resulting in higher rates of misclassification. In addition, it is not unusual to find the samples with the same inputs to have different class labels. In this paper, Parzen window is used to estimate the correctness of the class information of a sample and the resulting class information is used for feature extraction.Y

    Time-Domain Measurement Data Accumulation for Slow Moving Point Target Detection in Heavily Cluttered Environments Using CNN

    No full text
    In modern radars, the target detection probability is increased by lowering the detection threshold via signal processing to detect a point target with a small radar cross-section value. However, a lower threshold increases the number of false targets. In the conventional tracking method, which uses a general tracking filter, the measurement data between scans should be compared. Therefore, for a large amount of acquired measurement data, the computational complexity can be reduced by accumulating the acquired measurement data over time, recognizing the target movement as a pattern, and training a convolutional neural network (CNN) model. Here, we propose a method to create a desired target scenario by transfer learning and estimate the target position using the activation map of a binary detector CNN model. The model can detect a target using the actual acquired radar data, and the processing time remains constant, regardless of the number of false alarms.Y
    corecore