581 research outputs found

    Do unbalanced data have a negative effect on LDA?

    Get PDF
    For two-class discrimination, Xie and Qiu [The effect of imbalanced data sets on LDA: a theoretical and empirical analysis, Pattern Recognition 40 (2) (2007) 557–562] claimed that, when covariance matrices of the two classes were unequal, a (class) unbalanced data set had a negative effect on the performance of linear discriminant analysis (LDA). Through re-balancing 10 real-world data sets, Xie and Qiu [The effect of imbalanced data sets on LDA: a theoretical and empirical analysis, Pattern Recognition 40 (2) (2007) 557–562] provided empirical evidence to support the claim using AUC (Area Under the receiver operating characteristic Curve) as the performance metric. We suggest that such a claim is vague if not misleading, there is no solid theoretical analysis presented in Xie and Qiu [The effect of imbalanced data sets on LDA: a theoretical and empirical analysis, Pattern Recognition 40 (2) (2007) 557–562], and AUC can lead to a quite different conclusion from that led to by misclassification error rate (ER) on the discrimination performance of LDA for unbalanced data sets. Our empirical and simulation studies suggest that, for LDA, the increase of the median of AUC (and thus the improvement of performance of LDA) from re-balancing is relatively small, while, in contrast, the increase of the median of ER (and thus the decline in performance of LDA) from re-balancing is relatively large. Therefore, from our study, there is no reliable empirical evidence to support the claim that a (class) unbalanced data set has a negative effect on the performance of LDA. In addition, re-balancing affects the performance of LDA for data sets with either equal or unequal covariance matrices, indicating that having unequal covariance matrices is not a key reason for the difference in performance between original and re-balanced data

    High performance latent dirichlet allocation for text mining

    Get PDF
    This thesis was submitted for the degree of Doctor of Philosophy and awarded by Brunel University.Latent Dirichlet Allocation (LDA), a total probability generative model, is a three-tier Bayesian model. LDA computes the latent topic structure of the data and obtains the significant information of documents. However, traditional LDA has several limitations in practical applications. LDA cannot be directly used in classification because it is a non-supervised learning model. It needs to be embedded into appropriate classification algorithms. LDA is a generative model as it normally generates the latent topics in the categories where the target documents do not belong to, producing the deviation in computation and reducing the classification accuracy. The number of topics in LDA influences the learning process of model parameters greatly. Noise samples in the training data also affect the final text classification result. And, the quality of LDA based classifiers depends on the quality of the training samples to a great extent. Although parallel LDA algorithms are proposed to deal with huge amounts of data, balancing computing loads in a computer cluster poses another challenge. This thesis presents a text classification method which combines the LDA model and Support Vector Machine (SVM) classification algorithm for an improved accuracy in classification when reducing the dimension of datasets. Based on Density-Based Spatial Clustering of Applications with Noise (DBSCAN), the algorithm automatically optimizes the number of topics to be selected which reduces the number of iterations in computation. Furthermore, this thesis presents a noise data reduction scheme to process noise data. When the noise ratio is large in the training data set, the noise reduction scheme can always produce a high level of accuracy in classification. Finally, the thesis parallelizes LDA using the MapReduce model which is the de facto computing standard in supporting data intensive applications. A genetic algorithm based load balancing algorithm is designed to balance the workloads among computers in a heterogeneous MapReduce cluster where the computers have a variety of computing resources in terms of CPU speed, memory space and hard disk space

    LDA aan stijgende luchtbellen

    No full text
    Laser Doppler Anemometrie (LDA) is een meetmethode, die de stroming niet beïnvloedt. Met LDA wordt informatie over de snelheid verkregen uit een gemeten Dopplerfrequentie. Dit verslag behandelt LDA aan luchtbellen in een forward LDA opsteUing met referentiebundel. De diameter van de bellen is groot in verhouding tot de diameter van de laserbundels, waardoor de metingen niet beschreven worden door de klassieke LDA-theorie. In dit onderzoek zijn metingen gedaan aan stijgende bellen in verschillende vloeistoffen. Bij een meting is de equivalente diameter constant. Met verschillende luchtuitlaten zijn bellen verkregen met een diameter van 2 tot 5 mm. Aan de bellen zijn LDA-signalen gemeten door reflectie en door refractie. Bij reflectie spiegelt één laserstraal aan het beloppervlak, terwijl de ander langs de bel gaat. Bij refractie breken beide laserstralen, terwijl ze de luchtbel passeren. De Dopplerfrequenties veranderen sterk met de hoogte boven de luchtuitlaat. Op basis van de klassieke LDA-theorie zou dit overeen komen met snelheden van 20cm/s tot 40cm/s. Het frequentieverloop heeft een zekere periodiciteit. De signalen van reflectie en refractie verschillen wezenlijk van elkaar. Bij de verklaring van de Dopplersignalen is de theorie van bellenstromingen gebruikt. De experimenten leiden met de theorie van lichtbreking tot de conclusie, dat trillingen van het beloppervlak variaties in de Dopplersignalen geven. De hoogte boven de luchtuitlaat is om te zetten in de tijd na loslaten van de bel. Met deze tijd is de 'frequentie van de Dopplerfrequentie-oscillaties' te bepalen. Bij reflectie is deze frequentie voor verschillende belgroottes, oppervlaktespanningen en viscositeiten vergeleken met de frequentie, die de theorie van Lamb (1932) geeft voor trillingen van bolvormige bellen. Kwantitatief hebben veranderingen dezelfde invloed op beide frequenties, kwalitatief verschillen de frequenties maximaal een factor 1.6.Kramers Laboratorium voor Fysische TechnologieApplied Science

    Application of DPLS-Based LDA in Corn Qualitative Near Infrared Spectroscopy Analysis

    No full text
    NIR technology is a rapid, nondestructive and user-friendly method ideally suited for Qualitative analysis. In this paper the authors present the use of discriminant partial least Squares (DPLS)-based linear discriminant analysis (LDA) in corn qualitative near infrared spectroscopy analysis. Firstly, a training set including 30 corn varieties (each variety has 20 samples) was used to build the DPLS regression model, and 28 principal components (DPLS-PCs) were obtained from original spectrum. Secondly, the DPLS-PCs scores of the training set were extracted as DPLS features. Thirdly, LDA was applied to the DPLS features, determining 26 principal components (LDA-PCs). A test sample was first projected onto the DPLS-PCs and then onto the LDA-PCs, and finally 26 DPLS+LDA features were obtained. The recognition results were obtained by minimum distance classifier. DPLS+LDA method achieved 96. 18% recognition rate, while traditional DPLS regression method and DPLS feature extraction method only achieved 85. 38% and 95. 76% recognition rate respectively. The experiment results indicated that DPLS +LDA method is with better generalization ability compared with traditional DPLS regression method and NIRS analysis by DPLS+LDA method is an efficient way to discriminate corn species

    Spectral analysis of individual realization LDA data

    No full text
    The estimation of the autocorrelation function (act) or the spectral density function (sdt) from LDA data poses unique data-processing problems. The random sampling times in LDA preclude the use of the spectral methods for equi-spaced samples. As a consequence, special data-processing algorithms are used to process the LDA data. However, the random sampling causes an additional statistical variability of the spectral estimates that obscures the behaviour of the sdf in the high frequency range. The maximum frequency at which reliable estimates can be made is usually less than the mean data rate. For LDA measurements in gas flows the mean data rate is often small compared to the highest frequencies of the velocity fluctuations. As a consequence, the small scales of the turbulent fluctuations cannot be studied from the estimated sdf's with the presently available data-processing methods. It is the objective of the present study to modify an existing data-processing method such that information on the spectral density can be revealed at much higher frequencies. The modification consists of two elements. First, a locally sealed autocorrelation function is computed. This modification of the conventional slotting technique results in a much lower statistical variance at small lag times. Next, the locally scaled acf is cosine-transformed using a lag window whose width is varied with frequency. The modified estimator is applied to two types of stimulated data to illustrate its performance. It is shown that the modified slotting technique in conjunction with a variable window forms a powerful spectral estimator for low data density flows.Aerospace Engineerin

    LDA experiments on a mixing layer

    No full text
    Investigation on plane turbulent two phase mixing layers serves to get insight in the mutual interaction between the behaviour of the injected gas bubbles and the turbulence of the liquid phase. An experimental setup for investigations on such a mixing layer has been built. The measuring section is 20 cm in depth, 40 cm in width, and 150 cm in height. Measurements have been done on the liquid phase (water) with use of Laser Doppler Anemometry. These measurements mainly serve to determine the quality of the setup. The LDA measurements relate, first, to averaged velocities and their profiles, and second, to turbulent quantities, viz. rms values and ttu-stresses. From the turbulent quantities, ID-spectra and autocorrelation functions have been determined. As a direct result of the measurements, the air distribution was improved. The experimental results indicate that a leakage between the two sections exists. This should be repaired, along with the obliquity of the splitter plate. To get more reliable spectra, the data rate of the LDA measurements has to be increased. Another requirement for future measurements is the use of an accurate and more stable traverse system for the LDA probe.Kramers Laboratorium voor Fysische TechnologieApplied Science

    A Systematic Comparison of Search-Based Approaches for LDA Hyperparameter Tuning

    No full text
    Context: Latent Dirichlet Allocation (LDA) has been successfully used in the literature to extract topics from software documents and support developers in various software engineering tasks. While LDA has been mostly used with default settings, previous studies showed that default hyperparameter values generate sub-optimal topics from software documents. Objective: Recent studies applied meta-heuristic search (mostly evolutionary algorithms) to configure LDA in an unsupervised and automated fashion. However, previous work advocated for different meta-heuristics and surrogate metrics to optimize. The objective of this paper is to shed light on the influence of these two factors when tuning LDA for SE tasks. Method: We empirically evaluated and compared seven state-of-the-art meta-heuristics and three alternative surrogate metrics (i.e., fitness functions) to solve the problem of identifying duplicate bug reports with LDA. The benchmark consists of ten real-world and open-source projects from the Bench4BL} dataset. Results: Our results indicate that (1) meta-heuristics are mostly comparable to one another (except for random search and CMA-ES), and (2) the choice of the surrogate metric impacts the quality of the generated topics and the tuning overhead. Furthermore, calibrating LDA helps identify twice as many duplicates than untuned LDA when inspecting the top five past similar reports. Conclusion: No meta-heuristic and/or fitness function outperforms all the others, as advocated in prior studies. However, we can make recommendations for some combinations of meta-heuristics and fitness functions over others for practical use. Future work should focus on improving the surrogate metrics used to calibrate/tune LDA in an unsupervised fashionSoftware Engineerin
    corecore