530 research outputs found
Graphs and Attributes used for the attribute-structure correlation pattern mining
## SCPM: An implementation of an algorithm for structural correlation pattern mining.
The structural correlation measures how a set of attributes induces dense subgraphs in an attributed graph. A structural correlation pattern is a dense subgraph induced by a particular attribute set. Structural correlation pattern mining is useful to analyze how different attribute sets are correlated to dense subgraphs in several real-life attributed graphs.
**Relevant Publications**
* Arlei Silva, Wagner Meira, Jr., and Mohammed J. Zaki. Structural correlation pattern mining for large graphs. In Proceedings of the Eighth Workshop on Mining and Learning with Graphs (MLG '10).
* Arlei Silva, Wagner Meira, Jr., and Mohammed J. Zaki. Mining Attribute-structure Correlated Patterns in Large Attributed Graphs. In Proceedings of the VLDB Endowment (PVLDB '12).
* Arlei Silva. Structural correlation pattern mining for large graphs. M.Sc Thesis, Computer Science Department, Universidade Federal de Minas Gerais, 2011.
* Arlei Silva, Wagner Meira Jr. Structural correlation pattern mining for large graphs. Thesis and Dissertation Contest of the Brazilian Computer Society (CTD'12).
## HOW TO
cd to trunk and run make
see README in trunk
## Datasets:
### Description:
#### ATTRIBUTE FILE:
Format: Lists the attributes of each vertex from the graph.
,,...,
Example:
1,A,C
2,A
3,A,C,D
4,A,D
5,A,E
6,A,B,C
7,A,B,E
8,A,B
9,A,B
10,A,B,D
11,A,B
#### GRAPH FILE:
Format: Lists the neighbors of each vertex from the graph (adjacency list). Although the graph is undirected, each edge must be included in both directions.
,,...,
Example:
1,4
2,3
3,2,4,5,6,7
4,1,3,5,6
5,3,4,6
6,3,4,5,7,8,9,10
7,3,6,8,11
8,6,7,9,10,11
9,6,8,10,11
10,6,8,9,11
11,7,8,9,10
### REAL DATASETS
Lastfm:
attributes: attrLastFm.csv.tar.bz2
network: graphLastFm.csv.tar.gz
DBLP:
attributes: newAttrDBLP.csv.tar.bz2
network: newGraphDBLP.csv.tar.bz2
CITESEER:
attributes: attrCiteseer.csv.tar.bz2
network: graphCiteseer.csv.tar.bz2</p
Data mining and analysis : fundamental concepts and algorithms / Mohammed J. Zaki, Rensselaer Polytechnic Institute, Troy, New York, Wagner Meira, Jr., Universidade Federal de Minas Gerais, Brazil.
computer bookfair2016Includes bibliographical references and index.xi, 593 pages :A comprehensive overview of data mining from an algorithmic perspective, integrating related concepts from machine learning and statistics
A visual demonstration of supramolecular chemistry: observable fluorescence enhancement upon host-guest inclusion
PT: J; CR: BOZZELLI JW, 1982, J CHEM EDUC, V59, P787 BRESLOW R, 1998, J CHEM EDUC, V75, P705 BUCCIGROSS JM, 1996, J CHEM EDUC, V73, P275 BURROWS HD, 1983, J CHEM EDUC, V60, P228 CATENA GC, 1989, ANAL CHEM, V61, P905 CONN MM, 1997, CHEM REV, V97, P1647 CONRADI S, 1997, J CHEM EDUC, V74, P1122 CRAM DJ, 1992, NATURE, V356, P29 CRAMER F, 1967, J AM CHEM SOC, V89, P14 DIEDERICH F, 1990, J CHEM EDUC, V67, P813 EBBESEN TW, 1989, J PHYS CHEM-US, V93, P7139 FEMIA RA, 1985, ENVIRON SCI TECHNOL, V19, P155 FYFE MCT, 1997, ACCOUNTS CHEM RES, V30, P393 HAMILTON AD, 1990, J CHEM EDUC, V67, P821 KONDO J, 1976, J BIOCH, V79, P393 LACKOWICZ JR, 1983, PRINCIPLES FLUORESCE, CH7 LEHN JM, 1988, ANGEW CHEM INT EDIT, V27, P89 LERNER DA, 1989, ANAL CHIM ACTA, V227, P297 LI S, 1992, CHEM REV, V92, P1457 WAGNER BD, 1998, J PHOTOCH PHOTOBIO A, V114, P151 ZARZYCKI PK, 1996, J CHEM EDUC, V73, P459 ZIESSEL RF, 1997, J CHEM EDUC, V74, P673; NR: 22; TC: 6; J9: J CHEM EDUC; PG: 4; GA: 274KQSource type: Electronic(1
Modeling Performance of Parallel Programs
The actual performance of parallel programs is often disappointing, especially in comparison to the peak performance offered by the underlying hardware. There are many sources of performance degradation and understanding these sources is necessary to improve application performance. In this paper we discuss performance modeling, an approach to understanding the performance of parallel systems. We present a survey of current approaches to modeling (both analytical modeling based on system parameters, and structural modeling based on the structure of the program), and propose a combination of these two approaches as a promising direction for new work. This combination is explored by evaluating and proposing improvements to lost cycles analysis, which already contains features from both approaches, and also combines measurement and modeling. Supported by CNPq, Brazil, Grant No. 200.862-93/6 1 Introduction One disappointing contrast in parallel systems is between the nominal (e.g., peak..
Modeling Performance of Parallel Programs
The actual performance of parallel programs is often disappointing, especially in comparison to the peak performance offered by the underlying hardware. There are many sources of performance degradation and understanding these sources is necessary to improve application performance. In this paper we discuss performance modeling, an approach to understanding the performance of parallel systems. We present a survey of current approaches to modeling (both analytical modeling based on system parameters, and structural modeling based on the structure of the program), and propose a combination of these two approaches as a promising direction for new work. This combination is explored by evaluating and proposing improvements to lost cycles analysis, which already contains features from both approaches, and also combines measurement and modeling
Efficient Data Mining for Frequent Itemsets in Dynamic and Distributed Databases
Data Mining is one of the central activities associated with understanding and exploiting the world of digital data. It is the mechanized process of modeling large databases by means of discovering useful patterns. A frequent itemset is a pattern describing a relevant subset of the data, and a collection of frequent itemsets is particularly useful because it is an extremely compact model of the database. Discovering frequent itemsets in large databases is usually a hard computational task, which can be even harder when data is dynamic and distributed. Applying traditional algorithms in such data results in high communication overhead, excessive wastage of CPU and I/O resources, privacy violations, and often does not meet the stringent rapid response times, to essentially an interactive process of exploiting the data. Hence, there is an urgent need for non-trivial algorithms that can effectively mine frequent itemsets in dynamic and distributed databases. Such algorithms are presented in this master thesis
Using linear algebra for protein structural comparison and classification
In this article, we describe a novel methodology to extract semantic characteristics from protein structures using linear algebra in order to compose structural signature vectors which may be used efficiently to compare and classify protein structures into fold families. These signatures are built from the pattern of hydrophobic intrachain interactions using Singular Value Decomposition (SVD) and Latent Semantic Indexing (LSI) techniques. Considering proteins as documents and contacts as terms, we have built a retrieval system which is able to find conserved contacts in samples of myoglobin fold family and to retrieve these proteins among proteins of varied folds with precision of up to 80%. The classifier is a web tool available at our laboratory website. Users can search for similar chains from a specific PDB, view and compare their contact maps and browse their structures using a JMol plug-in
Medical Imaging Processing Architecture on ATMOSPHERE Federated Platform
[EN] This paper describes the development of applications in the frame of the ATMOSPHERE platform. ATMOSPHERE provides means for developing container-based applications over a federated cloud offering measurin he trustworthiness of the applications. In this paper we show the design of a transcontinental application in the frame of medical imaging that keeps the data at one end and uses the processing capabilities of the resources available at the other end. The applications are described using TOSCA blueprints and the federation
of IaaS resources is performed by the Fogbow middleware. Privacy guarantees are provided by means of SCONE and intensive computing resources are integrated through the use of GPUs directly mounted on the containers.The work in this article has been co-funded by
project ATMOSPHERE, funded jointly by the European Commission under the Cooperation Programme, Horizon 2020 grant agreement No 777154
and the Brazilian Ministerio de Ci ¿ encia, Tecnologia e ¿
Inovac¿ao (MCTI), number 51119. ¿
The authors also want to acknowledge the research grant from the regional government of the Comunitat Valenciana (Spain), co-funded by the European Union ERDF funds (European Regional Development Fund) of the Comunitat Valenciana 2014-
2020, with reference IDIFEDER/2018/032 (HighPerformance Algorithms for the Modelling, Simulation and early Detection of diseases in Personalized
Medicine).Blanquer Espert, I.; Alberich-Bayarri, Á.; García-Castro, F.; Teodoro, G.; Meirelles, A.; Nascimento, B.; Meira Jr., W.... (2019). Medical Imaging Processing Architecture on ATMOSPHERE Federated Platform. ScitePress. 589-594. https://riunet.upv.es/handle/10251/181077S58959
Um modelo baseado na evolução temporal de consumo e sua aplicação em domínios de recomendação
Em domínios de recomendação o gosto dos usuários, bem como o próprio domínio, varia ao longo do tempo. Porém, essa evolução é pouco compreendida na literatura. Um maior entendimento deste processo permitiria melhorar as recomendações. Neste trabalho, modelamos a evolução temporal de forma a identificar itens potencialmente relevantes para recomendação. Utilizamos o conceito de transições evolutivas representativas, transições entre quaisquer par de itens ao longo do tempo, e extraímos tais transições através da modelagem proposta. Além disso, realizamos experimentos para validar nossa premissa de que transições evolutivas representativas existem e são: mensuráveis, relevantes, úteis e não óbvias
- …
