1,721,051 research outputs found
MoDeSuS: A machine learning tool for selection of molecular descriptors in qsar studies applied to molecular informatics
The selection of the most relevant molecular descriptors to describe a target variable in the context of QSAR (Quantitative Structure-Activity Relationship) modelling is a challenging combinatorial optimization problem. In this paper, a novel software tool for addressing this task in the context of regression and classification modelling is presented. The methodology that implements the tool is organized into two phases. The first phase uses a multiobjective evolutionary technique to perform the selection of subsets of descriptors. The second phase performs an external validation of the chosen descriptors subsets in order to improve reliability. The tool functionalities have been illustrated through a case study for the estimation of the ready biodegradation property as an example of classification QSAR modelling. The results obtained show the usefulness and potential of this novel software tool that aims to reduce the time and costs of development in the drug discovery process.Fil: Martínez, María Jimena. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; ArgentinaFil: Razuc, Marina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; Argentina. Provincia de Buenos Aires. Gobernación. Comisión de Investigaciones Científicas; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; Argentin
On Artificial Gene Regulatory Networks
Gene regulatory networks (GRNs) represent dependencies between genes and their products during protein synthesis at the molecular level. At the present there exist many inference methods that infer GRNs form observed data. However, gene expression data sets have in general considerable noise that make understanding and learning even simple regulatory patterns difficult. Also, there is no well-known method to test the accuracy of inferred GRNs. Given these drawbacks, characterizing the effectiveness of different techniques to uncover gene networks remains a challenge. The development of artificial GRNs with known biological features of expression complexity, diversity and interconnectivities provides a more controlled means of investigating the appropriateness of those techniques. In this work we introduce this problem in terms of machine learning and present a review of the main formalisms that have been used to build artificial GRNs.Fil: Carballido, Jessica Andrea. Universidad Nacional del Sur; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentina. Universidad Nacional del Sur; Argentin
Direct Method for Structural Observability Analysis
A noncombinatorial method for structural observability analysis is presented in this paper. The technique rearranges the process occurrence matrix to a specific block lower-triangular pattern by means of bigraphs and digraphs in two consecutive stages. The algorithmic core is constituted of a new node classification that leads to suitable maximum-matching decompositions even for structurally singular matrices. A three-step strategy for the identification and analysis of forbidden subsets was also designed to take into account the additional numeric constraints that guarantee further solvability of the final pattern. In contrast with other structural techniques, the proposed method treats complex nonlinear models in a remarkably efficient way. Its performance was compared with existing structural observability techniques for three industrial problems. The final results revealed that the direct method is extremely robust and efficient in computing times, becoming more efficacious as problems grow in size and complexity.Fil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; ArgentinaFil: Sanchez, Mabel Cristina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; ArgentinaFil: Brignole, Nélida Beatriz. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentin
Filtering non-balanced data using an evolutionary approach
Matrices that cannot be handled using conventional clustering, regression or classification methods are often found in every big data research area. In particular, datasets with thousands or millions of rows and less than a hundred columns regularly appear in biological so-called omic problems. The effectiveness of conventional data analysis approaches is hampered by this matrix structure, which necessitates some means of reduction. An evolutionary method called PreCLAS is presented in this article. Its main objective is to find a submatrix with fewer rows that exhibits some group structure. Three stages of experiments were performed. First, a benchmark dataset was used to assess the correct functionality of the method for clustering purposes. Then, a microarray gene expression data matrix was used to analyze the method’s performance in a simple classification scenario, where differential expression was carried out. Finally, several classification methods were compared in terms of classification accuracy using an RNA-seq gene expression dataset. Experiments showed that the new evolutionary technique significantly reduces the number of rows in the matrix and intelligently performs unsupervised row selection, improving classification and clustering methods.Fil: Carballido, Jessica Andrea. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; ArgentinaFil: Cecchini, Rocío Luján. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; Argentin
CGD-GA: A graph-based genetic algorithm for sensor network design
The foundations and implementation of a genetic algorithm (GA) for instrumentation purposes are presented in this paper. The GA constitutes an initialization module of a decision support system for sensor network design. The method development entailed the definition of the individual's representation as well as the design of a graph-based fitness function, along with the formulation of several other ad hoc implemented features. The performance and effectiveness of the GA were assessed by initializing the instrumentation design of an ammonia synthesis plant. The initialization provided by the GA succeeded in accelerating the sensor network design procedures. It also accomplished a great improvement in the overall quality of the resulting instrument configuration. Therefore, the GA constitutes a valuable tool for the treatment of real industrial problems.Fil: Carballido, Jessica Andrea. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Laboratorio de Investigación y Desarrollo en Computación Científica; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; ArgentinaFil: Brignole, Nélida Beatriz. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; Argentin
Can we gain insight about the ductile behavior of materials by using polymer informatics?
Predicting the ductile behavior of thermoplastic materials is a significant challenge for both industry and research. In this study, we present a predictive model that classifies polymers based on their ductility degree, which is a relationship created in the present work to approach this problem. It comes from relating two critical and measurable properties from the tensile test. This target was discretized into three classes, more ductile, intermediate, and less ductile. The feature selection process employed for finding the most relevant molecular descriptors for the predictive model used two approaches: a classical one and an expert-guided one. A new metric, called relaxed %CC, was presented to prioritize the models that reduce the misclassification between the extremes of the ductile scale, which is considered more important than confusion with the intermediate class. Our final model was able to successfully classify polymers, achieving a precision rate of 0.91, an %CC of 89.47 % (traditional accuracy), and a relaxed %CC of 100 % (no extremes confusion). This approach has the potential to help both the industry and R&D by selecting polymers with suitable ductile properties for specific applications during the design stage before their synthesis.Fil: Cravero, Fiorella. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; ArgentinaFil: Ponzoni, Ignacio. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Instituto de Ciencias e Ingeniería de la Computación. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación. Instituto de Ciencias e Ingeniería de la Computación; ArgentinaFil: Diaz, Monica Fatima. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentina. Universidad Nacional del Sur. Departamento de Ingeniería Química; Argentin
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
On Evolutionary Algorithms for Biclustering of Gene Expression Data
Past decades have seen the rapid development of microarray technologies making available large amounts of gene expression data. Hence, has become increasingly important to count with reliable methods that interpret this information in order to discover new biological knowledge. In this review paper we aim to describe the main existing evolutionary methods that analyze microarray gene expression data by means of biclustering techniques. Strategies will be classified according to the evaluation metric they use to quantify the quality of the biclusters. In this context, the main evaluation measures namely mean squared residue, virtual error and transposed virtual error are first presented. Then, the main evolutionary algorithms, which find biclusters in gene expression data matrices using those metrics, are described and compared.Fil: Carballido, Jessica Andrea. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; ArgentinaFil: Gallo, Cristian Andrés. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; ArgentinaFil: Dussaut, Julieta Sol. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional del Sur. Departamento de Ciencias e Ingeniería de la Computación; Argentin
BiHEA: A Hybrid Evolutionary Approach for Microarray Biclustering
In this paper a new hybrid approach that integrates an evolutionary algorithm with local search for microarray biclustering is presented. The novelty of this proposal is constituted by the incorporation of two mechanisms: the first one avoids loss of good solutions through generations and overcomes the high degree of overlap in the final population; and the other one preserves an adequate level of genotypic diversity. The performance of the memetic strategy was compared with the results of several salient biclustering algorithms over synthetic data with different overlap degrees and noise levels. In this regard, our proposal achieves results that outperform the ones obtained by the referential methods. Finally, a study on real data was performed in order to demonstrate the biological relevance of the results of our approach. © 2009 Springer Berlin Heidelberg.Fil: Gallo, Cristian Andrés. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional del Sur; ArgentinaFil: Carballido, Jessica Andrea. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina. Universidad Nacional del Sur; ArgentinaFil: Ponzoni, Ignacio. Consejo Nacional de Investigaciones Científicas y Técnicas. Centro Científico Tecnológico Conicet - Bahía Blanca. Planta Piloto de Ingeniería Química. Universidad Nacional del Sur. Planta Piloto de Ingeniería Química; Argentina. Universidad Nacional del Sur; Argentin
- …
