1,721,108 research outputs found
An efficient Particle Swarm Optimization approach to cluster short texts
Short texts such as evaluations of commercial products, news, FAQ’s and scientific abstracts are important resources on the Web due to the constant requirements of people to use this on line information in real life. In this context, the clustering of short texts is a significant analysis task and a discrete Particle Swarm Optimization (PSO) algorithm named CLUDIPSO has recently shown a promising performance in this type of problems. CLUDIPSO obtained high quality results with small corpora although, with larger corpora, a significant deterioration of performance was observed. This article presents CLUDIPSO★, an improved version of CLUDIPSO, which includes a different representation of particles, a more efficient evaluation of the function to be optimized and some modifications in the mutation operator. Experimental results with corpora containing scientific abstracts, news and short legal documents obtained from the Web, show that CLUDIPSO★ is an effective clustering method for short-text corpora of small and medium size.Fil: Cagnina, Leticia Cecilia. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Errecalde, Marcelo Luis. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Ingaramo, Diego. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Rosso, Paolo. Universidad Politécnica de Valencia; Españ
Silhouette + Attraction: A Simple and Effective Method for Text Clustering
This article presents Sil-Att, a simple and effective method for text clustering, which is based on two main concepts: the silhouette coefficient and the idea of attraction. The combination of both principles allows to obtain a general technique that can be used either as a boosting method, which improves results of other clustering algorithms, or as an independent clustering algorithm. The experimental work shows that Sil-Att is able to obtain high quality results on text corpora with very different characteristics. Furthermore, its stable performance on all the considered corpora is indicative that it is a very robust method. This is a very interesting positive aspect of Sil-Att with respect to the other algorithms used in the experiments, whose performances heavily depend on specific characteristics of the corpora being considered.Fil: Errecalde, Marcelo L.. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo En Inteligencia Computacional; ArgentinaFil: Cagnina, Leticia Cecilia. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo En Inteligencia Computacional; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Rosso, Paolo. Universidad Politecnica de Valencia; Españ
Knowledge discovery applying text mining techniques in Psychology
La extracción de conocimiento en bases de datos es un proceso complejo que en última instancia busca darle sentido a los datos. La minería de datos sólo constituye una etapa de este proceso cuyo objetivo consiste en la obtención de patrones y modelos aplicando métodos estadísticos y técnicas de aprendizaje automático. El presente artículo de revisión examina cómo pueden aplicarse las técnicas de minería de textos en el campo de la psicología. En este contexto, se describen los dos grandes propósitos de las técnicas de minería de textos: la descripción y la predicción. Finalmente, se destaca que la aplicación de técnicas de minería de textos en nuestra disciplina hace posible la medición o evaluación de distintos constructos psicológicos, a diferencia de la utilización de los tradicionales cuestionarios o encuestas.The knowledge discovery in databases (KDD) is concerned with the non-trivial process of making sense of data. Data mining is only a step in the KDD process that consists in pattern recognition using statistics and machine learning techniques. This literature review focuses on how text mining techniques can be applied in Psychology. In this context, the two main purposes of text mining techniques will be introduced: description and prediction. Finally, this paper highlights the use of text mining techniques as a psychological assessment tool, which differs from the use of standard questionnaires or scales.Fil: Mariñelarena-dondena, Luciana. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional de San Luis. Facultad de Psicología; ArgentinaFil: Errecalde, Marcelo Luis. Universidad Nacional de San Luis. Facultad de Psicología; ArgentinaFil: Castro Solano, Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad de Palermo. Facultad de Ciencias Sociales. Departamento de Psicología. Centro de Investigación y Posgrados; Argentin
A Possibilistic Defeasible Logic Programming Approach to Argumentation-Based Decision Making
The development of symbolic approaches to decision-making has become an evergrowing research line in artificial intelligence; argumentation has contributed to that with its unique strengths. Following this trend, this article proposes a general-purpose decision framework based on argumentation. Given a set of alternatives posed to the decisionmaker, the framework represents the agent’s preferences and knowledge by an epistemic component developed using possibilistic defeasible logic programming. The reasons by which a particular alternative is deemed better than another are explicitly considered in the argumentation process involved in warranting information from the epistemic component. The information warranted by the dialectical process is then used in decision rules that implement the agent’s general decision-making policy. Essentially, decision rules establish patterns of behaviour of the agent specifying under which conditions a set of alternatives will be considered acceptable; moreover, a methodology for programming the agent’s epistemic component is defined. It is demonstrated that programming the agent’s epistemic component following this methodology exhibits some interesting properties with respect to the selected alternatives; also, when all the relevantinformation regarding the agent’s preferences is specified, its choice behaviour coincides with respect to the optimum preference derived from a rational preference relation.Fil: Ferretti, Edgardo. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Errecalde, Marcelo . Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Garcia, Alejandro Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Bahia Blanca; Argentina. Universidad Nacional del Sur; ArgentinaFil: Simari, Guillermo Ricardo. Universidad Nacional del Sur; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Bahia Blanca; Argentin
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Analysis of language traits with natural language processing techniques in early detection of depression
El desarrollo de métodos computacionales que utilizan información de la Web para la detección temprana de riesgos es un área de investigación socialmente relevante, científicamente atractiva y actualmente en pleno crecimiento. La depresión es uno de los trastornos mentales más frecuentes a nivel mundial y con alta incidencia de suicidio en los casos más severos. Por lo tanto, su detección temprana podría derivar en un tratamiento a tiempo e incluso salvar vidas. En este trabajo, se analiza la relación que existe entre los modelos computacionales que permiten la detección automática de depresión y las propiedades lingüísticas del texto escrito por personas que experimentan la enfermedad. Se utilizan representaciones textuales que forman parte del estado del arte en clasificación de documentos y que cubren aspectos lingüísticos, sintácticos y semánticos. Los resultados obtenidos con clasificadores estándares indican que las incrustaciones de palabras capturan información precisa para detectar indicios de depresión de forma rápida y segura.The development of computational methods using information from the Web for early detection of risks is a socially relevant, scientifically attractive and currently a growing area of research. Depression is one of the most frequent mental disorders in the world and with high incidence of suicide in the most severe cases. Therefore, early detection of this illness could lead to a timely treatment and to save lives. This paper analyzes the relationship between computational models that allow the automatic detection of depression and the linguistic properties of the text written by people who experience the disease. State-of-the-art text representations in document classification are used, covering linguistic, syntactic and semantic aspects. The results obtained with standard classifiers indicate that word embeddings capture precise information to detect quickly and safely signs of depression.Fil: Garciarena Ucelay, María José. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Cagnina, Leticia Cecilia. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - San Luis; Argentina. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; ArgentinaFil: Errecalde, Marcelo Luis. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; Argentina. Universidad Nacional de la Patagonia Austral; Argentina. Universidad Nacional de La Plata; Argentin
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
A simple method for recommending specialized specifications for diabetes monitoring
Under glycemic variability, a characterization of the desired blood glucose (BG) behavior is needed to assess if a given artificial pancreas (AP) respects its specification. The specification is an essential element to detect any deviation from an adequate insulin policy. Specializing the monitoring specification is therefore of utmost importance as existing guidelines for diabetes management are general and do not take into account how the personal factors and lifestyle affect the glycemic behavior. Surely, recommending personalized monitoring specifications may provide flexible and appropriate treatment goals to be attained by diabetic patients in order to account for their actual treatment needs. In this work, we use machine learning models to characterize glycemic behavior in synthetic healthy individuals. To account for the day-by-day fluctuation in BG levels, we use a stochastic process superimposed on a deterministic model of the glucose-insulin dynamics. The obtained characterization of the glycemic behavior in healthy individuals is then used as the target class to predict, and thus recommend, personalized monitoring specifications to diabetic patients. Results show that the approach stands as a feasible strategy to recommending appropriate and realistic monitoring goals for diabetic patients based on healthy individuals who share a similar glycemic behavior. Eventually, the incorporation of a recommender approach on an intelligent monitoring system for the AP will allow on-line adaptation of the treatment requirements for each patient.Fil: Avila, Luis Omar. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - San Luis; ArgentinaFil: Errecalde, Marcelo Luis. Universidad Nacional de San Luis. Facultad de Ciencias Físico Matemáticas y Naturales. Departamento de Informática. Laboratorio Investigación y Desarrollo en Inteligencia Computacional; Argentin
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
