1,720,982 research outputs found

    Segmental Analysis-Based Authorship Discrimination between the Holy Quran and Prophet’s Statements

    No full text
    Stylometry has got a lot of interest during these recent years because it solved many authorship problems and disputes that were difficult to handle. Author discrimination consists in checking whether two texts are written by the same author or not. In this investigation, we try to make an author discrimination between the Quran (The holy words and statements of God in the Islamic religion) and the Hadith (statements said by the prophet Muhammad) in a segmental form. In fact, 14 text segments are extracted from the Quran book and 11 text segments are extracted from the Bukhari Hadith. These segments have more or less the same size in terms of words and the medium size is about 2080 words per text segment. The Quran is taken in its entirety, whereas for the Prophet’s statements, we chose only the certified texts of the Bukhari book. That is, four series of experiments are done and commented. The first series of experiments concerns several experiments of authorship attribution using different state of the art features and classifiers, the second series of experiments analyses the different texts by using a new parameter called COST, the third series of experiments consists in an authorship discrimination using the frequency of a particular word (“الذين ” meaning those/who in English) and the fourth series of experiments performs a hierarchical clustering on the 25 text segments, in order to assess the real number of clusters (author styles) and to see if the hypothesis of a unique author is possible. This investigation, which represents the continuation of a previous research work on the same topic [Sayoud 2012-a], has further clarified an old enigma, which was impossible to solve for fourteen centuries: all the results of this investigation show unanimously that the two books should have two different authors.   La stylométrie a montré beaucoup d'intérêts durant ces dernières années car elle a résolu beaucoup de problèmes et disputes qui étaient difficiles à manipuler. La discrimination d'auteur consiste à vérifier si deux textes sont écrits par le même auteur ou non. Dans cette étude, nous essayons d'exécuter une discrimination d'auteur entre le Coran (Les mots saints et déclarations de Dieu dans la religion Islamique) et le Hadith (déclarations prononcées par le Prophète Muhammad) sous une forme segmentale. En fait, 14 segments de texte sont extraits du Coran et 11 segments sont extraits du Hadith de Bukhari. Ces segments ont plus ou moins la même taille en terme de mots et la taille moyenne est d'environ 2080 mots par segment. Le Coran est pris entièrement, tandis que pour les déclarations du Prophète, nous avons choisi seulement les textes certifiés du livre de Bukhari. Ainsi, quatre séries d'expériences sont faites et commentées. La première série d'expériences concerne plusieurs expériences d'identification d'auteur utilisant différents classifieurs et caractéristiques de l'état de l'art ; la seconde série d'expériences analyse les différents textes en utilisant un nouveau paramètre appelé 'COST'; la troisième série d'expériences consiste en une discrimination d'auteurs utilisant la fréquence d'un mot particulier ("الذين " signifiant Ceux/ Qui en Français) et la quatrième série d'expériences exécute un regroupement hiérarchique sur les 25 segments de texte, dans le but d'estimer le nombre réel de clusters (Styles d'auteurs) et de voir si l'hypothèse d'un auteur unique est possible. Cette étude, qui représente la suite d'un travail de recherche précédent sur le même sujet (Sayoud 2012a), a clarifié d'avantages une ancienne énigme, qui était impossible de résoudre durant quatorze siècles : en fait, tous les résultats de cette étude montrent que les deux livres devraient avoir deux auteurs différents

    AUTHORSHIP IDENTIFICATION OF SEVEN ARABIC RELIGIOUS BOOKS -A FUSION APPROACH-

    No full text
    In this paper, we conduct an investigation of automatic authorship attribution on seven Arabic religious books, namely: the holy Quran, Hadith and five other books written by five religious scholars. The Arabic styles are almost the same (i.e. Standard Arabic) for the seven books. The genre is the same and the topics of the different books are also the same (i.e. Religion). The authorship characterization is based on four different features: character trigrams, character tetragrams, word unigrams and word bigrams. The task of authorship identification is ensured by four conventional classifiers: Manhattan distance, Multi-Layer Perceptron, Support Vector Machines and Linear Regression. Furthermore, a fusion approach has been proposed to enhance the performances of authorship attribution, with two fusion techniques. The novelty of this research work lies in the following points: the proposal of a new type of fusion and the proposal of a new optimal rule dealing with unbalanced text documents. A particular application is dedicated to the authorship discrimination between the Quran and Hadith, in order to see if the two books could have the same author or not. Results show good authorship attribution performances with an overall score ranging from 96% and 99% of correct attribution by using the conventional classifiers. This score reaches 100% of correct attribution by using the proposed fusion techniques. Concerning the application of discrimination, results have revealed that the Quran and Hadith books are stylistically different and should belong to two different authors

    AUTHORSHIP IDENTIFICATION OF SEVEN ARABIC RELIGIOUS BOOKS -A FUSION APPROACH-

    No full text
    In this paper, we conduct an investigation of automatic authorship attribution on seven Arabic religious books, namely: the holy Quran, Hadith and five other books written by five religious scholars. The Arabic styles are almost the same (i.e. Standard Arabic) for the seven books. The genre is the same and the topics of the different books are also the same (i.e. Religion). The authorship characterization is based on four different features: character trigrams, character tetragrams, word unigrams and word bigrams. The task of authorship identification is ensured by four conventional classifiers: Manhattan distance, Multi-Layer Perceptron, Support Vector Machines and Linear Regression. Furthermore, a fusion approach has been proposed to enhance the performances of authorship attribution, with two fusion techniques. The novelty of this research work lies in the following points: the proposal of a new type of fusion and the proposal of a new optimal rule dealing with unbalanced text documents. A particular application is dedicated to the authorship discrimination between the Quran and Hadith, in order to see if the two books could have the same author or not. Results show good authorship attribution performances with an overall score ranging from 96% and 99% of correct attribution by using the conventional classifiers. This score reaches 100% of correct attribution by using the proposed fusion techniques. Concerning the application of discrimination, results have revealed that the Quran and Hadith books are stylistically different and should belong to two different authors

    Biometrics

    No full text
    The term biometrics is derived from the Greek words: bio (life) and metrics (to measure). “Biometric technologies” are defined as automated methods of verifying or recognizing the identity of a living person based on a physiological or behavioral characteristic. Several techniques and features were used over time to recognize human beings several years before the birth of Christ. Today, this research field has become very employed in many applications such as security applications, multimedia applications and banking applications. Also, many methods have been developed to strengthen the biometric accuracy and reduce the imposture errors by using several features such as face, speech, iris, finger vein, etc. From a security purpose and economic point of view, biometrics has brought a great benefit and has become an important tool for governments and institutions. However, citizens are expressing their thorough worry, which is due to the freedom limitations and loss of privacy. This paper briefly presents some new technologies that have recently been proposed in biometrics with their levels of reliability, and discusses the different social and ethic problems that may result from the abusive use of these technologies.</p

    Biometrics

    No full text
    corecore