252 research outputs found
Not Available
Version: 1.0.0
Imports: utils, minimalRSD, stats
Published:2017-03-21
Author: Shwetank Lall [aut, cre], Arpan Bhowmik [ctb], Eldho Varghese [aut], Seema Jaggi [ctb], Cini Varghese [ctb]
Maintainer: Shwetank Lall
License: GPL-2 | GPL-3 [expanded from: GPL (≥ 2)]
NeedsCompilation: no
Citation: FMC citation info
In views: ExperimentalDesignAn R package to generate cost effective minimally changed run sequences for symmetrical as well as asymmetrical factorial designsNot Availabl
Recommended from our members
Learning from aggregated data
Data aggregation is ubiquitous in modern life. Due to various reasons like privacy, scalability, robustness, etc., ground truth data is often subjected to aggregation before being released to the public, or utilised by researchers and analysts. Learning from aggregated data is a challenging problem that requires significant algorithmic innovation, since naive application of standard techniques to aggregated data is vulnerable to the ecological fallacy. In this work, we explore three different versions of this setting.
First, we tackle the problem of using generalised linear models when features/covariates are fully observed but the targets are only available as histograms- a common scenario in the healthcare domain where many datasets contain both non-sensitive attributes like age, sex, zip-code, etc., as well as privacy sensitive attributes like healthcare records. We introduce an efficient algorithm that uses alternating data imputation and GLM estimation steps to learn predictive models in this setting.
Next, we look at the problem of learning sparse linear models when both features and targets are in aggregated form, specified as empirical estimates of group-wise means computed over different sub-groups of the population. We show that if the true sub-populations are heterogeneous enough, the optimal sparse parameter can be recovered within an arbitrarily small tolerance even in the presence of noise, provided the empirical estimates are obtained from a sufficiently large number of observations.
Third, we tackle the scenario of predictive modelling with data that is subjected to spatio-temporal aggregation. We show that by formulating the problem in the frequency domain, we can bypass the mathematical and representational challenges that arise due to non-uniform aggregation, misaligned sampling periods and aliasing. We introduce a novel algorithm that uses restricted Fourier transforms to estimate a linear model which, when applied to spatio-temporally aggregated data, has a generalisation error that is provably close to the optimal performance by the best possible linear model that can be learned from the non-aggregated data set.
We then focus our attention on the complementary problem that involves designing aggregation strategies that can allow learning, as well as developing algorithmic techniques that can use only the aggregates to train a model that works on individual samples. We motivate our methods by using the example of Gaussian regression, and subsequently extend our techniques to subsume binary classifiers and generalised linear models. We deonstrate the effectiveness of our techniques with empirical evaluation on data from healthcare and telecommunication.
Finally, we present a concrete example of our methods applied to a real life practical problem. Specifically, we consider an application in the domain of online advertising where the complexity of bidding strategies require accurate estimates of most probable cost-per-click or CPC incurred by advertisers, but the data used for training these CPC prediction models are only available as aggregated invoices supplied by an ad publisher on a daily or hourly basis. We introduce a novel learning framework that can use aggregates computed at varying levels of granularity for building individual-level predictive models. We generalise our modelling and algorithmic framework to handle data from diverse domains, and extend our techniques to cover arbitrary aggregation paradigms like sliding windows and overlapping/non-uniform aggregation. We show empirical evidence for the efficacy of our techniques with experiments on both synthetic data and real data from the online advertising domain as well as healthcare to demonstrate the wider applicability of our framework.Electrical and Computer Engineerin
Performance Evaluation of Polybenzimidazole for Potential Aerospace Applications
With the increasing use of polymer based composite materials, there is an increasing demand of polymeric resins with high glass transition temperature, high thermal stability and excellent mechanical properties at high temperature. Polybenzimidazole (PBI) is a recently emerged high performance polymer. It has the highest glass transition temperature of any commercially available organic polymer, high decomposition temperature, good oxidation resistance and it maintains excellent strength at cryogenic temperatures. Due to its excellent thermal and mechanical properties, PBI has great potential to be used for many high temperature applications. The present work has contributed to the knowledge and understanding of several aspects of the PBI polymer. Different problems related to the processing of unfilled and carbon nano-fibers reinforced PBI in the form of film, coating and adhesive are highlighted. Performance of PBI after exposure to various environmental conditions is evaluated. PBI has shown great potential to be used as fire resistant coating in aircraft. It also has revealed its potential to be used for different space applications.Structure Integrity and Composite groupAerospace Engineerin
A Harmony Search Based Wrapper Feature Selection Method for Holistic Bangla Word Recognition
AbstractA lot of search approaches have been explored for the selection of features in pattern classification domain in order to discover significant subset of the features which produces better accuracy. In this paper, we introduced a Harmony Search (HS) algorithm based feature selection method for feature dimensionality reduction in handwritten Bangla word recognition problem. This algorithm has been implemented to reduce the feature dimensionality of a technique described in one of our previous papers by Bhowmik et al.1. In the said paper, a set of 65 elliptical features were computed for handwritten Bangla word recognition purpose and a recognition accuracy of 81.37% was achieved using Multi Layer Perceptron (MLP) classifier. In the present work, a subset containing 48 features (approximately 75% of said feature vector) has been selected by HS based wrapper feature selection method which produces an accuracy rate of 90.29%. Reasonable outcomes also validates that the introduced algorithm utilizes optimal number of features while showing higher classification accuracies when compared to two standard evolutionary algorithms like Genetic Algorithm (GA), Particle Swarm Optimization (PSO) and statistical feature dimensionality reduction technique like Principal Component Analysis (PCA). This confirms the suitability of HS algorithm to the holistic handwritten word recognition problem
Computational intelligence based lossless regeneration (CILR) of blocked gingivitis intraoral image transportation
This paper presented that an intraoral image has been wrapped during wireless transportation with an encryption tool with an added essence of lossless regeneration property. Threshold based cryptographic transportation has provided the construction of reliable and robust medical data communication system. The accumulation of threshold shares only would result to the formation of the intraoral gingivitis image at the receivers’ end. The proposed technique dealt with the generation of n number of partial shares by creating a unique frame structure by the dentist / physician. Additional feature has been proposed on the computational lossless transportation.The existing techniques cause a high computational complexity. The proposed technique ensured the lossless regeneration property while blocked gingivitis image sharing. Filling of bits have been incorporated to ensure the static sized homogeneous blocks of intraoral gingivitis image. A graphical masking method had been deployed, followed by successive decryption procedure on minimum threshold shares that ensure lossless data regeneration. This can guide the dental treatment with enhanced accuracy. Different types of statistical testing like entropy analysis and histogram analysis confirms the exhibition of authenticity, confidentiality, and integrity of our proposed technique
Recurrence relation and DNA sequence: A state-of-art technique for secret sharing
During the transmission over the Internet, protection of data and information is an important issue. Efficient cryptographic techniques are used for protection but everything depends on the encryption key and robustness of encryption algorithm. Threshold cryptography provides the development of reliable and strong encryption and key management machine which can reconstruct the message even in the case of destruction of some particular numbers of shares and at the opposite the data cannot be reconstructed unless an allowable set of shares are been gathered. The earlier techniques available in literature result in high computational complexity in the course of both sharing and reconstructing of message. Our method employs a brand new easy protecting technique based totally on unit matrix. The simple AND operation is used for percentage generation and reconstruction can be finished by way of easy ORing the stocks with threshold cost. We are proposing a sharing approach in conjunction with conventional cryptography technique for key control to make the key greater sturdy and for encryption we have used a session key the use of the idea of recurrence relation and DNA series Different types of experimental results confirm authenticity, confidentiality, integrity and acceptance of our technique
Privileged authenticity in reconstruction of digital encrypted shares
Efficient message reconstruction mechanism depends on the entire partial shares received in random manner. This paper proposed a technique to ensure the authenticated accumulation of shares based on the privileged share. Threshold number of received shares inclusive of the privileged share, were being accumulated together to validate the original message. Although attaining threshold number of shares or more excluding the privileged share, it would not be possible to reconstruct the original message. Encryptional procedure has been put into the desired partial shares to confuse the evaesdroppers. Decisive parameter termed as hash tag has been extracted from the cumulative shares and bitwise checking procedure has been carried out. In appearance of first mismatch, rests of the checking bits were ignored, as test case put under failure transaction. Different statistical tests namely floating frequency, entropy value have proved the robustness of the proposed technique. Thus, extensive experiments were conducted to evaluate the security and efficiency with better productivity
- …
