1,721,012 research outputs found
A doubly projected analysis for Lexical Tables
This paper aims at showing how external information contributes in analysing a lexical table by enriching the readability of factorial maps. The theoretical frame is given by Principal Component Analysis onto a Reference Subspace, a method based on the orthogonal projection of a correlation structure on the space spanned by an external set of explanatory variables. In previous papers the idea of a projected lexical analysis has been introduced by using a single reference space for terms. Here we consider a double projection strategy by involving external informative structures both on documents and terms, i.e. on rows and columns of a lexical table
Rotated canonical correlation analysis for multilingual corpora
This paper aims at proposing the joint use of Canonical Correlation Analysis and Procrustes Rotations (RCA), when we deal with a text and its translation into another language. The basic idea is representing words in the two different natural languages on a common reference space. The main characteristic of this space is to be lan-guage independent, although Procrustes Rotation is performed transforming the lexical table derived from trans-lation by minimizing its distance from the lexical table belonging to the original corpus, while the subsequent Canonical Correlation Analysis treats symmetrically the two word sets. The most interesting RCA feature is building a unique reference space for representing the correlation structure in the data, inducing the two systems of canonical factors to lie on the same space. These graphical representations enables us to read distances be-tween corresponding points in terms of different way of translating the same word in relation with the general context defined by the canonical variates. Trying to understand the distances between matched points could rep-resent an useful tool for enriching lexical resources in a translation procedure. In this paper we propose the com-parison of the most frequent content bearing words in the two languages, analyzing one year (2003) of Le Monde Diplomatique and its Italian edition
Extracting and Classifying Keywords in Textual Data Analysis
In this paper, we consider a peculiar lexical table having as general term the number of times the forms of two different vocabularies, collected on the same units, are simultaneously present. On this peculiar matrix, we first of all apply a factorial data analysis method for visualizing and extracting keywords and successively, by means of a co-clustering technique, we identify classes of keywords for the two different corpora. The main results of this strategy are shown by an application on two corpora defined by the language used by a set of firms on their official web sites for describing their core mission and the language they use in searching new employers
Il candidato ideale: Analisi delle Professionalità e delle Competenze
Obiettivo del lavoro è proporre l'utilizzo di strumenti che consentano di sintetizzare l'informazione testuale contenuta in un corpus, visualizzandola in modo opportuno. In particolare si propone di utilizzare gli oggetti simbolici per comprimere l'informazione contenuta in database testuali di grosse dimensioni e di rappresentare l'informazione con nuovi strumenti quali le 3D zoom sta
An Integrated Approach for the Analysis of the TF/IDF Matrix
In this paper we propose to analyse a puculiar data matrix, known as term frequency/inverse document frequency (TF/IDF) matrix, with a strategy based on the jointly use of a Factorial Data Analysis and a Classification Method
Relazioni non simmetriche tra Corpora
In this paper the language used by firms for searching new employers by web is studied. Particularly, we are in-teresting in evaluating the dependence between two corpora, e.g. one defined by the forms used for describing the skills of the candidates for jobs and the other defined by the forms used by firms for describing their mission. The method used is textual data analysis, more precisely, a non symmetrical correspondence analysis on a pecu-liar lexical table forms/forms with an ad hoc weighting system on the explanatory variables. Furthermore, the main results of an application on a sample of firms are showed in terms of friendly readable graphical represen-tations
A doubly projected analysis for lexical tables
This paper aims at showing how external information contributes in analysing a lexical table by enriching the readability of factorial maps. The theoretical frame is given by Principal Component Analysis onto a Reference Subspace, a method
based on the orthogonal projection of a correlation structure on the space spanned by an external set of explanatory variables. In previous papers the idea of a projected lexical analysis has been introduced by using a single reference space for terms. Here we consider a double projection strategy by involving external informative structures both on documents and terms, i.e. on rows and columns of a lexical table
The Ideal Candidate. Analysis of Professional Competences through Text Mining of Job Offers
The aim of this paper is to propose analytical tools for identifying peculiar aspects of job market for graduates. We propose a strategy for dealing with daa tat have different source and nature
CO.ME. T.A. – COvid-19 MEdia Textual Analysis: una piattaforma per il monitoraggio dei media}
- …
