1,720,960 research outputs found
Study of Single and Ensemble Classifiers of Classification Tree and Support Vector Machine
A classifier is such a rule that can be used to group an object into predetermined group or classs based on its attributes. There are two types of approach to develop a classifier rules are a parametric and a nonparametric. Parametric method requires certain assumptions to obtain the best classification but not all assumptions are met so that makes it difficult for researchers. The violation of the assumptions might lead to the lack of the effectiveness and the validity results. Recently, people pay more attention to non parametric classifiers such as Support Vector Machine (SVM) and Classification Tree (CT) to overcome the violation of the assumptions of parametric method. Some resent research figured out that an ensemble of classifiers could be an effective way to improve the classification accuracy and reduce the prediction variation of a single classifier (Valentini dan Dietterich 2000). The ensemble method is combining the class predictions resulted by a set of single classifiers into a single prediction by applying a majority vote rule. Among some popular techniques a method of bagging (bootstrap agregating) by Breiman (1996) is the simplest but powerful technique. The data used in this research are simulation data and real-life data. Simulation data are used to assess and compare the performance of single and ensemble classifiers of classification tree and SVM in three different data structures: (1) a situation where the members of different classes are perfectly linear separable, (2) a situation where the members of different classes are linerseparable but not perfect and (3) a situation where the members of different classes could not be separated by a linear function. Single and ensemble classifiers of classification trees and SVM will be applied to classify the successful study of postgraduate IPB students in Statistics department enrollment 2000-2010. Our research revealed that SVM resulted better classifier compared to Classification Tree. It is valid for all three data structure under consideration. Moreover, ensemble treatment to the classifier succeeded in improving the classification performance, especiality when radial kernel function is embedded in the procedure. Ensemble SVM in real-life data with a radial kernel function has the best performance compared to other methods and is the most appropriate method to classify the successful study of postgraduate IPB students in Statistics department enrollment 2000-2010
MODIFIKASI DALAM PENAKSIR RASIO MENGGUNAKAN RANK SET SAMPLING
Penelitian ini bertujuan untuk memodifikasi penaksir rasio dengan menggunakan Rank Set Sampling mrssRˆ . Varians dari penaksir ini akan diperoleh dan dibandingkan dengan varians rasio Rank Set Sampling klasik rssRˆ . Dengan kondisi perbandingan ini menghasilkan penaksir yang dimodifikasi lebih efisien daripada penaksir klasik
Penerapan Metode Boosting Pada Cart Untuk Mengklasifikasikan Korban Kecelakaan Lalu Lintas Di Kota Palu
Kota Palu sebagai ibu kota Provinsi Sulawesi Tengah dengan kecelakaan lalu lintas yang cukup tinggi yang setiap tahunnya memiliki kematian sekitar 365 jiwa. Kecelakaan lalu lintas dipengaruhi oleh beberapa faktor, diantaranya jenis pelanggaran, jenis kecelakaan, dan lain-lain. Tujuan yang akan dicapai dalam penelitian ini adalah untuk menentukan ketepatan klasifikasi pada korban kecelakaan lalu lintas di Kota Palu dengan menggunakan metode boosting serta faktor-faktor yang mempengaruhinya. Hasil dari penelitian ini menunjukkan bahwa ketepatan klasifikasi metode boosting sebesar 82% dan metode CART sebesar 77,9%. Hasil tersebut menunjukkan bahwa metode boosting dapat meningkatan tingkat akurasi. Sedangkan faktor-faktor yang mempengaruhi korban kecelakaan lalu lintas di Kota Palu adalah faktor jenis kecelakaan (X1), peran korban dalam kecelakaan (X4), jenis pelanggaran (X7) dan usia (X3) korban kecelakaan lalu lintas di Kota Palu
A Gaussian Mixture Model Approach to Profiling Stunting Risk Across Indonesian Provinces
Stunting is still a major health problem in Indonesia, with notable differences between provinces. Although the national rate has decreased over time, regional gaps continue, emphasizing the role of data in helping to explain what contributes to the issue. This study aims to segment 38 provinces in Indonesia based on maternal and child health indicators associated with stunting prevalence. The variables used include the percentage of low birth weight (LBW) infants, the percentage of infants born short, the percentage of pregnant women with chronic energy deficiency (CED), exclusive breastfeeding (EBF) coverage, prevalence of diarrhea in toddlers, and prevalence of acute respiratory infections (ARI) in toddlers. The clustering analysis was performed using the Gaussian Mixture Model (GMM) with the number of clusters varied from 2 to 7. Model selection was based on the Bayesian Information Criterion (BIC), where the lowest value indicated the optimal model. The results show that the model with two clusters was selected, with a BIC value of 1358.24, which indicates the best balance between model fit and complexity. This clustering reveals that provinces are grouped based on similarities in maternal and child health profiles, not on geographic proximity, meaning that the GMM method does not rely on spatial location to form clusters
Analisis Regresi Logistik Biner Untuk Mengklasifikasi Penderita Hipertensi Berdasarkan Kebiasaan Merokok Di RSU Mokopido Toli-Toli
Hipertensi adalah suatu penyakit yang disebabkan oleh peningkatan abnormal tekanan darah, baik tekanan darah sistolik maupun tekanan darah diastolik. Salah satu faktor resiko terjadinya hipertensi adalah kebiasaan merokok. Penelitian ini bertujuan untuk mendapatkan klasifikasi penderita hipertensi di RSU Mokopido Toli-toli dengan menggunakan regresi logistik biner. Variabel respon yang digunakan dalam penelitian yaitu hipertensi (Y) yang dikategorikan jadi dua kategori yaitu hipertensi dan tidak hipertensi. Variabel prediktor yang diteliti adalah kebiasaan merokok yang berupa lama merokok (X1), jumlah rokok yang dihisap (X2), jenis rokok yang dihisap (X3), cara menghisap rokok (X4), jenis kelamin (X5) dan usia (X6). Dari hasil penelitian menunjukkan bahwa faktor-faktor yang mempengaruhi penderita hipertensi adalah lama merokok (X1), jenis rokok yang dihisap (X3) dan cara menghisap rokok (X4). Persamaan regresi logistik biner dengan fungsi logit yang dihasilkan adalah g(x) = 0,981+ 2,478X1- 2,364X3+1,549X4 . Nilai ketepatan klasifikasi penderita hipertensi dengan menggunakan analisis regresi logistik biner yaitu sebesar 77,6
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
