1,720,977 research outputs found

    Principled machine learning

    Get PDF
    We introduce the underlying concepts which give rise to some of the commonly used machine learning methods, excluding deep-learning machines and neural networks. We point to their advantages, limitations and potential use in various areas of photonics. The main methods covered include parametric and non-parametric regression and classification techniques, kernel-based methods and support vector machines, decision trees, probabilistic models, Bayesian graphs, mixture models, Gaussian processes, message passing methods and visual informatics

    Dataset supporting the conference paper entitled "Real-time Room Occupancy Estimation with Bayesian Machine Learning using a Single PIR Sensor and Microcontroller"

    No full text
    This dataset supports the paper entitled &quot;Real-time Room Occupancy Estimation with Bayesian Machine Learning using a Single PIR Sensor and Microcontroller&quot; for the Sensors Applications Symposium 2017. Assigned DOI: 10.5258/SOTON/405795</span

    A generalizable and open-source algorithm for real-life monitoring of tremor in Parkinson’s disease

    Get PDF
    Wearable sensors can objectively and continuously monitor daily-life tremor in Parkinson’s Disease (PD). We developed an open-source algorithm for real-life monitoring of PD tremor which achieves generalizable performance across different wrist-worn devices. We achieved this using a unique combination of two independent, complementary datasets. The first was a small, but extensively video-labeled gyroscope dataset collected during unscripted activities at home (n = 24 PD; n = 24 controls). We used this to train and validate a logistic regression tremor detector based on cepstral coefficients. The second was a large, unsupervised dataset (n = 517 PD; n = 50 controls, data collected for 2 weeks with a different device), used to externally validate the algorithm. Results show that our algorithm can reliably quantify real-life PD tremor (sensitivity of 0.61 (0.20) and specificity of 0.97 (0.05)). Weekly aggregated tremor time and power showed excellent test-retest reliability and moderate correlation to MDS-UPDRS rest tremor scores. This opens possibilities to support clinical trials and individual tremor management with wearable technology

    Complementary MR measures of white matter and their relation to cardiovascular health and cognition

    Get PDF
    The microstructural and macrostructural integrity of white matter (WM) underpins efficient brain function, and is known to decline with age and vascular burden. Key aspects of WM health include axonal fibre density, myelination, free-water content, and the presence of tissue damage or lesions. Magnetic Resonance Imaging (MRI) offers multiple complementary sequences to non-invasively estimate these properties in vivo. For example, diffusion-weighted imaging (DWI) provides sensitive measures of microstructure, while T1-weighted and T2-weighted MRI can estimate total WM volume and hyper-intensities, and magnetisation transfer imaging (MT) and T1:T2 ratios can indicate myelin content. In this study, we leveraged all of these MRI-derived measures in a large population-based cohort (Cam-CAN) to identify latent WM factors and test how these factors relate to cardiovascular health and cognitive performance. Among 11 commonly-used WM metrics [Fractional Anisotropy (FA); Mean Signal Diffusion (MSD); Mean Signal Kurtosis (MSK); Neurite Density Index (NDI); fibre Orientation Dispersion Index (ODI); Free water volume faction (Fiso); spread of Mean Signal Diffusivity values (MSDvar); Magnetisation Transfer Ratio (MTR); T1:T2 ratio; volume of White Matter Hyper-Intensities (WMHI); White Matter Volume (WMV)], latent factor analysis showed that four factors were needed to explain 89% of the variance, which we interpreted in terms of (1) fibre density/myelination, (2) free-water / tissue damage, (3) fibre-crossing complexity and (4) microstructural complexity. These factors showed distinct effects of age and sex. To test the validity of these factors, we related them to measures of cardiovascular health and cognitive performance. Specifically, we ran path analyses linking (1) cardiovascular factors to the WM factors, and (2) the WM factors to cognitive measures. Even after adjusting for age and sex, we found that a vascular factor related to pulse pressure predicted the WM factor capturing free-water/tissue damage, and that several WM factors made unique predictions for fluid intelligence and processing speed. Our results show that there is both complementary and redundant information across common MR measures of WM, and their underlying latent factors may be useful for pinpointing the differential causes and contributions of white matter health in aging

    Real-time room occupancy estimation with Bayesian machine learning using a single PIR sensor and microcontroller

    No full text
    This paper presents the implementation and deployment of a compute/memory intensive non-parametric Bayesian machine learning algorithm on a microcontroller unit (MCU) to estimate room occupancy in a Smart Room using a single analogue PIR sensor. We envisage an IoT device consisting of a resource-constrained MCU, PIR sensor and a battery running the occupancy estimation algorithm and operating over days or months without recharging or replacing the battery. Both hardware-independent and hardware-dependent optimizations are performed to reduce memory footprint and yet provide acceptable real-time performance while consuming less energy. We show a significant reduction in the on-chip memory usage in the MCUs by the algorithm through optimisation of the machine learning models and of the static memory footprint and dynamic memory usage. We also show that a low-end MCU does not meet the real-time requirements of the application without causing high average power consumption. However, a moderately high-performance MCU with a higher clock frequency and hardware floating-point unit provides 19x improvement in the execution time of the algorithm, better meeting the real-time specification of the application and reducing power consumption. Further, we estimate the battery lifetime of the IoT device if it operates continuously in a Smart Room. With a typical size battery, an IoT device consisting of a Cortex-M4F MCU and PIR sensor can operate for more than a month without replacement or recharging of the battery while running the compute-intensive Bayesian machine learning algorithm

    Quantifying arm swing in Parkinson’s disease: a method accounting for arm activities during free-living gait

    Get PDF
    Background: Accurately measuring hypokinetic arm swing during free-living gait in Parkinson’s disease (PD) is challenging due to other concurrent arm activities. We developed a method to isolate gait segments without these arm activities. Methods: Wrist accelerometer and gyroscope data were collected from 25 individuals with PD and 25 age-matched controls while performing unscripted activities in their home environment. This was done after overnight withdrawal of dopaminergic medication (‘pre-medication’) and approximately one hour after intake (‘post-medication’). Using video annotations as ground truth, we trained and evaluated two classifiers: one for detecting gait and one for detecting gait segments without other arm activities. Based on the filtered gait segments, arm swing was quantified using the median and 95th percentile range of motion (RoM). These arm swing parameters were evaluated in three ways: (1) the agreement between predicted and video-annotated gait segments without other arm activities, (2) the sensitivity to differences between PD and controls, and (3) the sensitivity to the effects of dopaminergic medication. Results: On the most affected side, the mean (SD) balanced accuracy for detecting gait without other arm activities was 0.84 (0.10) pre-medication and 0.88 (0.09) post-medication. The agreement between arm swing parameters of predicted and video-annotated gait segments without other arm activities was high irrespective of medication state (intra-class correlation coefficients: median RoM: 0.99; 95th percentile RoM: 0.97). Both the median and 95th percentile RoM were smaller in PD pre-medication compared to controls (median: Δ=-18.80∘, 95% CI [-30.63, -10.60], p < 0.001; 95th percentile: Δ=-28.34∘, 95% CI [-38.26, -18.18], p < 0.001), and smaller in pre- compared to post-medication (median: Δ=-12.31∘, 95% CI [-21.35, -5.59], p < 0.001; 95th percentile: Δ=-19.04∘, 95% CI [-28.48, -11.14], p < 0.001). The differences in RoM between pre- and post-medication were larger after filtering gait for the median (p < 0.01) and 95th percentile RoM (p = 0.01). Conclusions: Filtering out gait segments with other concurrent arm activities is feasible and increases the change in arm swing parameters following dopaminergic medication in free-living conditions. This approach may be used to monitor treatment effect and disease progression in daily life

    Causal distributional effects for functional data

    Get PDF
    LAUREA MAGISTRALENella nostra ricerca proponiamo un modello per l'inferenza causale, un'area di ricerca indirizzata a molteplici applicazioni. L'obiettivo è fornire un metodo per l'analisi di effetti in distribuzione indotti da una variabile binaria, piuttosto che effetti in media. Il modello proposto si compone di due fasi. La prima, in linea con la grande maggioranza della ricerca accademica, riguarda la costruzione di una risposta potenziale. Proponiamo infatti un modello Bayesiano per dati funzionali sotto i due trattamenti binari. La seconda rappresenta la misurazione dell'effetto causale a partire dai due insiemi di curve, uno per ogni trattamento. A tal fine, abbiamo sviluppato un quadro teorico per trattare la funzione di distribuzione cumulativa empirica (CDF) di dati ad alta dimensionalità, che non è una procedura banale in contesti con dati funzionali. Di conseguenza, forniamo un metodo pratico per calcolare le dissimilarità tra CDF, utilizzando metodologie di trasporto ottimale. Abbiamo testato il modello risultante sia su dataset sintetici che su dati reali provenienti da studi sul Morbo di Parkinson. Inoltre, confrontiamo le sue prestazioni con metodi consolidati nel campo, come gli approcci basati su kernel e i modelli per funzioni di distribuzione, per valutare l'efficacia e la robustezza della nostra proposta.In this research we propose a model for causal inference, a critical area of investigation for many domains of application. The goal is to provide a tool for analyzing the distributional effects induced by a binary treatment variable, rather than average effects. The proposed model consists of two phases. The first one, coherently with the vast majority of academic research, is related to the construction of a Bayesian potential outcome for functional observations under the two binary treatments. The second one represents the measurement of the causal effect starting from the two sets of curves, one for each treatment. To address this challenge, we developed a theoretical framework for dealing with empirical cumulative distribution function (CDF) of high-dimensional data, which is a popular challenge in functional data analysis. Consequently, we provide a practical method for computing dissimilarities between CDFs, by means of optimal transport methodologies. We evaluate the model using both synthetic datasets and real-world data from Parkinson’s Disease studies. We also compare its performance with well-established methods in the field, such as kernel-based approaches and distribution function models, to assess the effectiveness and robustness of our proposal

    Dirichlet process based models for functional data: a Bayesian approach to random effects estimation

    No full text
    LAUREA MAGISTRALEIn diverse discipline, come economia, scienze ambientali e in medicina, i dati sono raccolti sotto forma di curve con valori in un dominio temporale. In questi settori, i modelli lineari e generalizzati misti lineari migliorano la flessibilità dei modelli lineari tradizionali incorporando effetti casuali. Tuttavia, in applicazioni reali, le ipotesi parametriche convenzionali per tali effetti possono risultare troppo limitanti. Per ovviare a questa limitazione, in questo studio esploriamo l'uso di misture di processi di Dirichlet (DPM) come alternativa non parametrica per modellare gli effetti casuali in contesti di regressione funzionale. Le DPM non solo offrono una maggiore flessibilità ma inducono anche un effetto clustering sui dati. In questa tesi viene sviluppato e valutato un modello Function-on-Scalar che incorpora le DPM, considerando prima gli effetti casuali come variabili scalari e poi espandendo il modello per includere curve casuali funzionali. Diverse alternative del modello sono studiate e presentate al fine di migliorare le prestazioni e mitigare i problemi di dimensionalità. Attraverso l'impiego di dati sintetici, dimostriamo che il nostro modello performa con successo la task di regressione e contemporaneamente scopre cluster latenti. Concludiamo aggiungendo un ulteriore livello di complessità e considerando una variabile binaria per includere un effetto trattamento. Testiamo quest'ultimo modello su dati sintetici e su un set di dati reali di individui affetti dal morbo di Parkinson. I risultati mostrano un buon fit per i coefficienti funzionali e gli effetti casuali nel contesto sintetico, mentre sui dati reali ci permettono di fare inferenza significativa sugli effetti stimati.In several disciplines, like economics, environmental and health science, observations are represented as curves over a time domain. In these fields, linear and generalized linear mixed models enhance the flexibility of traditional linear models by incorporating random effects. However, in real world application, the conventional parametric assumptions for such effects may be too limiting. To address this limitation, in this study we explore the use of Dirichlet Process Mixtures (DPMs) as a nonparametric alternative for modeling random effects in functional regression settings. DPMs not only provide greater flexibility but also induce a natural clustering effect. We develop and evaluate a Function-on-Scalar model that incorporates DPMs by considering first scalar random effects and then we expand the model to include functional random curves. We study different alternatives of the model in order to improve the performance and to mitigate dimensionality issues. Through the employment of synthetic data, we show that our model successfully performs regression while simultaneously uncovering latent clusters. We conclude by adding another level of complexity and considering a binary treatment effect variable. We test the latter model on synthetic data and on real world dataset of individuals affected by Parkinson's Disease. The results show a strong fit for the functional coefficients and random effects in the synthetic setting, while on the real dataset they allow us to make meaningful inferences about the estimated effects

    A deterministic inference framework for discrete nonparametric latent variable models:learning complex probabilistic models with simple algorithms

    Get PDF
    Latent variable models provide a powerful framework for describing complex data by capturing its structure with a combination of more compact unobserved variables. The Bayesian approach to statistical latent models additionally provides a consistent and principled framework for dealing with uncertainty inherent in the data described with our model. However, in most Bayesian latent variable models we face the limitation that the number of unobserved variables has to be specied a priori. With the increasingly larger and more complex data problems such parametric models fail to make most out of the data available. Any increase in data passed into the model only affects the accuracy of the inferred posteriors and models fail to adapt to adequately capture new arising structure. Flexible Bayesian nonparametric models can mitigate such challenges and allow the learn arbitrarily complex representations given enough data is provided. However,their applications are restricted to applications in which computational resources are plentiful because of the exhaustive sampling methods they require for inference. At the same time we see that in practice despite the large variety of exible models available, simple algorithms such as K-means or Viterbi algorithm remain the preferred tool for most real world applications.This has motivated us in this thesis to borrow the exibility provided by Bayesian nonparametric models,but to derive easy to use, scalable techniques which can be applied to large data problems and can be ran on resource constraint embedded hardware. We propose nonparametric model-based clustering algorithms nearly as simple as K-means which overcome most of its challenges and can infer the number of clusters from the data. Their potential is demonstrated for many different scenarios and applications such as phenotyping Parkinson and Parkisonism related conditions in an unsupervised way. With few simple steps we derive a related approach for nonparametric analysis on longitudinal data which converges few orders of magnitude faster than current available sampling methods. The framework is extended to effcient inference in nonparametric sequential models where example applications can be behaviour extraction and DNA sequencing. We demonstrate that our methods could be easily extended to allow for exible online learning in a realistic setup using severely limited computational resources. We develop a system capable of inferring online nonparametric hidden Markov models from streaming data using only embedded hardware. This allowed us to develop occupancy estimation technology using only a simple motion sensor
    corecore