39936 research outputs found
Sort by
Solution blow spun polymeric nanofibres embedding cyclodextrin complexes of miltefosine: An approach to the production of sprayable dressings for the treatment of cutaneous leishmaniasis
Leishmaniasis is a neglected tropical disease caused by Leishmania genus protozoa. Treating this disease effectively and safely remains a significant challenge. Herein, hydroxypropyl-beta-cyclodextrin (HPßCD) and miltefosine (MF), an alkylphospholipid currently used for the treatment of leishmaniasis, were incorporated into nonwoven mats made of nanofibres of polyvinylpyrrolidone and the amphiphilic block copolymer Tetronic® 1307. The mats were produced in straightforward manner by solution blow spinning (SBS), after the optimisation of the experimental setup for the in-situ production. Scanning electron microscopy, FTIR spectroscopy, X-ray diffraction and differential scanning calorimetry were used to fully characterize the fibres morphology and structure. Both MF and HPßCD were embedded into the fibres at proportions adequate for the therapeutic action of MF, without affecting their global morphology. The release kinetics was controlled by the fast dissolution of the hygroscopic polymeric matrix. HPßCD-MF-loaded fibres demonstrated active against Leishmania promastigotes, displaying higher activity than MF, in addition to a reduced cytotoxicity in macrophages. The functionalised fibres affected the expression levels of parasite genes related to proliferation, differentiation, and drug response. This work demonstrates the potential of SBS for the in-situ delivery of drugs in the form of sprayable dressings, highlighting the use of CD complexes of antileishmanial agents.The authors gratefully acknowledge the financial support provided by the Ministerio de Ciencia e Innovación from Spain (MICIU/AEI/10.13039/501100011033, projects PID2020-112713RB-C21 -C22, and PID2023-147765OB-C21, -C22), Fundación La Caixa (LCF/PR/PR13/51080005) and Fundación Roviralta (Chair “Maria Francisca de Roviralta of Molecular Parasitology, Leishmaniasis and One Health”). Z.D. acknowledges the Asociación de Amigos de la Universidad de Navarra for her doctoral grant. P.G. thanks the Government of Navarra as well as the Ministerio de Educación Cultura y Deporte from Spain (FPU23/01618) for his doctoral grant
Combined model-based and data-driven approach for the control of a soft robotic neck
This paper delves into the potential of integrating model-based and data-driven techniques for controlling the performance of a soft robotic neck. Artificial intelligence (AI) methods, such as machine learning and deep learning, have shown their applicability in modelling and controlling robotic systems with complex nonlinear behaviours. However, model-based approaches have also proven to be effective analytical alternatives, even if they rely on simplified approximations of the robot model. The control system proposed in this work combines the closed loop analytical model of the soft robotic neck with a Multi-Layer Perceptron (MLP) network trained to minimise the neck pose error. The MLP undergoes training with three different data treatments, and the results are compared to determine the most effective one. The experimental results obtained demonstrate the robustness of the proposed technique and its potential as an alternative to classical solutions, whether purely based on analytical models or data-driven models.Nicole A. Continelli, Luis F. Nagua and Concepción A. Monje were supported by the project ADAPTA, with reference PLEC2023-010218, funded by MICIU /AEI /10.13039/501100011033, and project SIROCO, with reference PID2023-147343OB-I00, funded by MICIU /AEI/10.13039/501100011033 and FEDER, UE . Pablo M. Olmos was supported by the Comunidad de Madrid IDEA-CM project (TEC-2024/COM-89), the ELLIS Unit Madrid (European Laboratory for Learning and Intelligent Systems), the 2024 Leonardo Grant for Scientific Research and Cultural Creation from the BBVA Foundation, and AEI/FEDER/UE (PID2021-123182OB-I00 EPiCENTER; PID2024-157856NB-I00 CARTESIAN)
Essays on the Quantitative Analysis of Climate Heterogeneity
This PhD thesis advocates the use of quantile-based approaches as flexible tools to capture different forms of climate heterogeneity. Chapter 1, “Heterogeneous Predictive Association of CO2 with Global Warming”, applies the Quantile Factor Model of Chen et al. (2021) to study how detrended CO2-growth predicts the quantile factors of temperature deviations from a linear trend, thereby describing short-run predictability in climate dynamics. Chapter 2, “Quantitative Analysis of Climate Heterogeneity via an Unconditional Quantile Vector Error-Correction Model”, introduces an Unconditional Quantile Vector Error Correction Model to model and forecast the long-run relationships between the unconditional quantiles of temperature and the radiative forcing of several GHGs. Chapter 3, “Macroeconomic Effects of Temperature Distributional Shocks”, extends recent research on the macroeconomic impact of shocks to the global-average temperature by incorporating the effects of shocks at different part of the global or local temperature distributions. Finally, Chapter 4, “High-Frequency Density Nowcasts of U.S. State-Level Carbon Dioxide Emissions”, employs quantile panel regression methods to generate density nowcasts of state-level CO2 emissions in the United States.Programa de Doctorado en Economía por la Universidad Carlos III de MadridPresidente: Mikkel Bennedsen.- Secretario: Álvaro Escribano Sáez.- Vocal: Marina Friedric
Novel computational techniques for decision support through medical imaging
Mención Internacional en el título de doctorClinical experts have to daily confront complex decisions such as disease diagnosis, grading, or treatment determination. To assist them, the computational progress has boosted the creation of Clinical Decision Support Systems (CDSSs). Nowadays, many of these systems are designed with artificial intelligence (AI), giving them the ability to work with all kinds of clinical data, as for example images. Medical imaging has major importance in clinical and pharmacological processes, nevertheless, they are data with high complexity that require intense computation to analyze them. The automation of the image analysis, classification, and examination processes promises to speed up and increase the accuracy of the daily work of the clinical experts, being a current issue. Particularly in lung infection diseases, which are a major concern for the health systems globally, due to the high mortality rates and the spread capability of the infector microorganisms. This propagation produces high infection rates, as in the COVID-19 pandemic, which proved the major importance of fast treatment and isolation of the patient to control incidence rates. Another current concern is the emergence of antimicrobial resistance (AMR), which poses a major challenge in fighting infectious diseases, like tuberculosis, where the AMRs have no response to current treatment schemes. Currently, all these challenges are being addressed thanks to novel computational techniques. This thesis focuses on the proposal of new techniques and analysis for the current CDSSs, exploring the use-case of clinical images for lung infection diseases. This dissertation addresses two main computational streams; image classification systems and automatic image assessment techniques, applying an experimental comparative approach. Previous studies have proved the capability of the CDSSs to support the detection and assessment of diseases with medical images. Despite that, there are many issues to address, especially in a high-precision area like medicine where the systems have to be reliable, robust, and precise. During the dissertation, we addressed different open questions as the decision of the imaging modality during the design of a CDSS, and the implications of applying parallel and distributed learning techniques. In this document, a whole classification system is also proposed. The system is able to differentiate four lung infection diseases and it is researched to maximize the accuracy with deep learning ensemble techniques. This dissertation also presents image assessment techniques. The assessment systems analyze images to provide information without (or with a minimum) human direct interaction. In this area a novel diagnosis system for tuberculosis is proposed, the technique is able to find bacteria in the sputum microscopy images automatically, allowing for increasing the precision of this diagnosis and assessment procedure. Also, a semi-supervised approach to evaluate time-lapse microscopy (TLM) sequences is presented. The system speeds up the evaluation of TLM and the labeling procedure, opening new paths for microbial deep learning future research. To summarize, this thesis includes original contributions to the CDSSs area, proposing ad-hoc techniques, new labeling procedures, and novel analyses to fill literature gaps, increasing the reliability and speed of clinical procedures. The techniques proposed beyond the use-case of lung infection disease are open to being used in other image modalities and diseases, opening new research and application paths.Los expertos clínicos deben enfrentar diariamente decisiones complejas como el diagnóstico de enfermedades, la clasificación o la determinación del tratamiento. Para asistirlos, el progreso computacional ha impulsado la creación de Sistemas de Soporte a la Decisión Clínica (SSDC). Hoy en día, muchos de estos sistemas se diseñan con inteligencia artificial (IA), dándoles la capacidad de trabajar con todo tipo de datos clínicos, como por ejemplo imágenes. La imagenología médica tiene una gran importancia en los procesos clínicos y farmacológicos. Sin embargo, se trata de datos de alta complejidad que requieren una intensa computación para su análisis. La automatización de los procesos de análisis, clasificación y examen de imágenes promete acelerar y aumentar la precisión del trabajo diario de los expertos clínicos, siendo un tema de actualidad.
Particularmente, las enfermedades de infección pulmonar son una gran preocupación para los sistemas de salud a nivel mundial, debido a las altas tasas de mortalidad y la capacidad de propagación de los microorganismos infecciosos. Esta propagación produce altas tasas de infección, como en la pandemia de COVID-19, que demostró la gran importancia del tratamiento rápido y el aislamiento del paciente para controlar las tasas de incidencia. Otra preocupación actual es la aparición de la resistencia a los antimicrobianos (RAM), que plantea un gran desafío en la lucha contra enfermedades infecciosas, como la tuberculosis, donde la RAM no responde a los esquemas de tratamiento actuales.
Actualmente, todos estos desafíos se están abordando gracias a nuevas técnicas computacionales. Esta tesis se centra en la propuesta de nuevas técnicas y análisis para los SSDC actuales, explorando el caso de uso de imágenes clínicas para enfermedades de infección pulmonar. Esta disertación aborda dos corrientes computacionales principales: sistemas de clasificación de imágenes y técnicas de evaluación automática de imágenes, aplicando un enfoque comparativo experimental. Estudios previos han demostrado la capacidad de los SSDC para apoyar la detección y evaluación de enfermedades con imágenes médicas. A pesar de esto, quedan muchos problemas por abordar, especialmente en un área de alta precisión como la medicina, donde los sistemas deben ser fiables, robustos y precisos.
Durante la disertación, abordamos diferentes preguntas abiertas como la decisión de la modalidad de imagenología durante el diseño de un SSDC y las implicaciones de aplicar técnicas de aprendizaje paralelo y distribuido. En este documento, también se propone un sistema de clasificación completo. El sistema es capaz de diferenciar cuatro enfermedades de infección pulmonar y se investiga para maximizar la precisión con técnicas de *ensemble* de aprendizaje profundo. Esta disertación también presenta técnicas de evaluación de imágenes. Los sistemas de evaluación analizan imágenes para proporcionar información sin (o con una mínima) interacción humana directa. En esta área se propone un novedoso sistema de diagnóstico para la tuberculosis; la técnica es capaz de encontrar bacterias en las imágenes de microscopía de esputo automáticamente, lo que permite aumentar la precisión de este procedimiento de diagnóstico y evaluación. Además, se presenta un enfoque semi-supervisado para evaluar secuencias de microscopía de lapso de tiempo (TLM). El sistema acelera la evaluación de TLM y el procedimiento de etiquetado, abriendo nuevos caminos para futuras investigaciones de aprendizaje profundo microbiano. En resumen, esta tesis incluye contribuciones originales al área de los SSDC, proponiendo técnicas *ad-hoc*, nuevos procedimientos de etiquetado y análisis novedosos para llenar vacíos en la literatura, aumentando la fiabilidad y la velocidad de los procedimientos clínicos. Las técnicas propuestas, más allá del caso de uso de las enfermedades de infección pulmonar, están abiertas a ser utilizadas en otras modalidades de imagen y enfermedades, abriendo nuevas vías de investigación y aplicación.Programa de Doctorado en Ciencia y Tecnología Informática por la Universidad Carlos III de MadridPresidente: Pedro Ángel Cuenca Castillo.- Secretario: Félix García Carballeira.- Vocal: Fabrizio Marozz
Data-driven chance-constrained optimization based on Gaussian Mixture Models
In this work, we present a novel approach based on the combination of Gaussian Mixture Models (GMM) and Chance-Constrained Optimization (CCO). The method proposed deals with the often difficult task of deriving exact estimates of the individual constraint quantiles in the presence of uncertainty for some or all of the optimization problem parameters. Although COO solution methods have been extensively studied when uncertainty is assumed to be normally distributed, only approximate solutions are considered when the uncertainty is not normal. We propose a reformulation of the COO problem that provides exact and tractable solutions. The performance of the method has been studied under different GMM parameterizations in simulated and real-world applications.This work is part of the R&D project PID2023-151013NB-I00, funded byMCIN/AEI/10.13039/501100011033 and by ERDF/EU
On Using Curved Mirrors to Decrease Shadowing in VLC
Proceedings of 2024 IEEE Global Communications Conference (GLOBECOM), 8-12 Dec. 2024, Cape Town, South AfricaVisible light communication (VLC) complements radio frequency in indoor environments with large wireless data traffic. However, VLC is hindered by dramatic path losses when an opaque object is interposed between the transmitter and the receiver. Prior works propose the use of plane mirrors as optical reconfigurable intelligent surfaces (ORISs) to enhance communications through non-line-of-sight links. Plane mirrors rely on their orientation to forward the light to the target user location, which is challenging to implement in practice. This paper studies the potential of curved mirrors as static reflective surfaces to provide a broadening specular reflection that increases the signal coverage in mirror-assisted VLC scenarios. We study the behavior of paraboloid and semi-spherical mirrors and derive the irradiance equations. We provide extensive numerical and analytical results and show that curved mirrors, when developed with proper dimensions, may reduce the shadowing probability to zero, while static plane mirrors of the same size have shadowing probabilities larger than 65%. Furthermore, the signal-to-noise ratio offered by curved mirrors may suffice to provide connectivity to users deployed in the room even when a line-of-sight link blockage occurs.Borja Genoves Guzman has received funding from the European Union under the Marie Skłodowska-Curie grant agreement No 101061853
Information on income concentration and redistribution preferences: The case of the top 1% in the United States
[EN] In recent decades, income inequality has soared in the United States, but few studies utilize experimental methodologies to assess whether having reliable information on income concentration affects the formation of redistribution attitudes. We advance this emerging literature through an analysis of information on the income owned by the top 1% of the income distribution ¿ a group that is socio-politically meaningful and axial to the recent rise of inequality. The empirical evidence draws on a novel split-sample survey experiment (N=4,000). The results indicate that having the real value of income concentration does not have an average, significant effect on redistribution attitudes. Neither does it moderate the positive association between prior perceptions of concentration and these attitudes. However, the analysis suggests that intense cognitive effort is related to preference revision. This is because, information reduces pro-redistribution attitudes among extreme overestimators who process the information more intensely.[FR] Au cours des dernières décennies, les inégalités de revenus ont explosé aux États-Unis. Pourtant, peu d'études utilisent des méthodologies expérimentales pour évaluer si le fait de disposer d'informations fiables sur la concentration des revenus a une influence sur les opinions concernant la redistribution. Cet article contribue à l'étude émergente de cette question en analysant les informations sur les revenus détenus par le 1% supérieur de la distribution des revenus - un groupe qui est important au plan socio-politique et qui joue un rôle central dans l'augmentation récente des inégalités. Les données empiriques s'appuient sur une nouvelle expérience d'enquête à échantillon partagé (N=4.000). Les résultats indiquent que le fait de connaître la valeur réelle de la concentration des revenus n'a pas d'effet moyen significatif sur les attitudes à l'égard de la redistribution. Cela ne modère pas non plus l'association positive entre les perceptions antérieures de la concentration des revenus et ces attitudes. L'étude révèle cependant qu'un effort cognitif intense est lié à la révision des préférences. Cela s'explique par le fait que l'information réduit les attitudes favorables à la redistribution chez les surestimateurs extrêmes qui traitent l'information plus intensément.[ES] En las últimas décadas, la desigualdad de ingresos se ha disparado en Estados Unidos, pero pocos estudios han utilizado metodologías experimentales para evaluar si el hecho de disponer de información fiable sobre la concentración de ingresos afecta a la formación de actitudes hacia la redistribución. En este artículo se contribuye a esta literatura emergente mediante un análisis de la información sobre los ingresos que obtiene el 1% más alto de la distribución de ingresos, un grupo sociopolíticamente significativo y crucial para el reciente aumento de la desigualdad. La evidencia empírica se basa en un experimento de encuesta original de muestra partida (N=4.000). Los resultados indican que disponer del valor real de la concentración de ingresos no tiene un efecto medio significativo sobre las actitudes hacia la redistribución. Tampoco modera la asociación positiva entre las percepciones previas del grado de concentración y tales actitudes. Sin embargo, el análisis sugiere que un esfuerzo cognitivo intenso está relacionado con la revisión de las preferencias. Esto se debe a que la información reduce las actitudes pro-redistributivas entre los sobreestimadores extremos que procesan la información con más intensidad
Use of transfer learning for affordable in-context fake review generation
This work has been submitted to the IEEE for possible publication Copyright may be transferred without notice, after which this version may no longer be accessibleFake reviews are a threat to the trust of online shopping platforms. To produce them, attackers may use assorted techniques based on machine learning. Transfer learning may enable them to leverage already trained models, thus reducing the training requirements. However, the feasibility of these techniques for generating in-context targeted fake reviews has not been explored yet. To address this issue, this paper analyses the suitability of transfer learning using existing models (TS and BART) on different domains (restaurants and technological products). Our results show that: (1) Our work reaches better realism and diversity than previous proposals using artificial intelligence techniques; (2) Our reviews produced with TS can be spot by an automatic detector with a precision of 49% at the highest; (3) Human detection is only 6% better than random guessing, at the highest; and (4) only 1 hour of training and Sk real reviews are needed to produce realistic fake reviews.This work was supported by the Madrid Government (Comunidad de Madrid-Spain),
in part by the Multiannual Agreement with UC3M (“Fostering Young Doctors Research”, DEPROFAKE-CM-UC3M), in part by the V PRICIT (Research
and Technological Innovation Regional Programme); in part by the Excellence Program for University Researchers; in part by the INCIBE grant APAMciber within the framework of the Recovery, Transformation and Resilience Plan funds, in part by the European Union (Next Generation). Jose Maria de Fuentes and Lorena Gonzalez have also received support from UC3M’s Requalification programme, in part by the Spanish Ministerio de Ciencia, Innovacion y Universidades with EU recovery funds (Convocatoria de la Universidad Carlos III de
Madrid de Ayudas para la recualificación del sistema universitario español para 2021-2023, de 1 de julio de 2021). This project has received funding from the European Union’s Horizon 2020 research and innovation programme under Grant 101021377 (TRUSTaWARE), and in part by the Funding for APC: Universidad Carlos III de Madrid under Grant CRUE-Madroño 2024
Data‑driven analysis of anomalous transport and three‑wave‑coupling effects in an E × B plasma
The collisionless cross-field electron transport in an E x B plasma configuration, representative of a Hall thruster, is studied using bispectral analysis on the data of a fully-kinetic simulation. The nonlinear, in-phase interaction of the oscillations of the azimuthal electric field and the electron density, both tied to the fundamental electron cyclotron drift instability (ECDI) mode, is found to be the main driver of electron transport. Higher-wavenumber ECDI modes do not drive anomalous transport directly; however, they are nonlinearly coupled with each other and with the fundamental ECDI mode. In addition, there is a smaller contribution from a lower-wavenumber mode, not predicted by linear ECDI theory. A reduced model obtained by sparse regression of the data suggests the existence of an inverse energy cascade from the higher ECDI modes to the fundamental one, which would mean that these modes do contribute to transport, albeit indirectly.This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (ZARATHUSTRA project, grant agreement No 950466). Additionally, during part of this period Bayón-Buján enjoyed a grant from the Consejería de Educación, Universidades, Ciencia y Portavocía of the Community of Madrid (grant PEJ-2021-AI/TIC-23158)
Bayesian Learning in High Dimensional Medical Data
Mención Internacional en el título de doctorThe increasing availability of diagnostic methods and new medical technologies provides healthcare professionals with access to a huge variety of data from various sources and of different types This enables a more in-depth analysis of complex diseases and disorders allowing for more precise diagnoses Additionally many medical tests such as medical imaging and genetic tests generate highly detailed data For instance in medical imaging resolutions can reach hundreds of thousands of voxels However this level of precision which is directly associated with a high dimensionality combined with the multi-modality of medical data introduces new challenges for effective data analysis In fact this variety of information is impossible to be fully analyzed by a human with healthcare professionals struggling to examine all relationships between different data sources and thoroughly process high-dimensional data This is where Machine Learning (ML) algorithms are emerging as powerful tools for analyzing complex clinical data settings Although ML algorithms can efficiently process such data and assist in diagnostic decision-making they face a major limitation: the challenge of working in wide-data environments This concept refers to the constraints that ML algorithms face when dealing with datasets that have much more features than samples often leading to poor generalization (overfitting) and computational limitations This scenario far from being present in isolated cases is common across various medical fields such as psychiatry and neurology In these cases the complexity of the diseases requires the collection of data from multiple sources for accurate diagnoses while patient recruitment for clinical studies is often challenging resulting in datasets with numerous features (with one or more data views) but few useful patients Therefore this thesis focuses on defining ML algorithms capable of working in different wide-data scenarios Additionally as these models are intended to work in clinical environments they must also be interpretable This is not only crucial for explaining how diagnostic decisions are made but also for identifying clinically useful biomarkers that can assist in future diagnoses and research This thesis begins by proposing a simple model inspired by the well-known Sparse Bayesian Linear Regression (SBLR) which is capable of efficiently work within wide-data scenarios while performing a compact selection of the most relevant variables To do so the proposed model the Dual Bayesian Linear Regression with Feature Selection (DBL-FS) operates directly over the dual space while imposing sparsity on the features dimension This allows for the generation of a sparse model in terms of features and subsequently enables it to perform a Feature Selection (FS) process while efficiently training it by means of a dual formulation This model has been tested on two real medical datasets one related to schizophrenia and the other to bacterial outbreaks and in both cases it outperformed all proposed baselines Additionally with the assistance of our clinical collaborators we have validated the biomarkers identified by the model providing insights into its clinical utility Next we propose a more complex model inspired on the Sparse Relevance Vector Machine (SRVM) model the Bayesian sparse formulation of the Support Vector Machines (SVM) The proposed algorithm the Relevance Feature and VectorMachine (RFVM) operates in the dual space like DBL-FS but additionally produces sparse solutions in both features and samples Thus the model can efficiently handle wide-data environments using kernels while providing a final solution that is as compact as possible by retaining only the minimum number of features and samples needed for the final diagnostic decision Although this approach may seem counterintuitive in settings with limited data availability we show that controlled removal of noisy or outlier samples can potentially improve the final FS and enhance the generalization of the model Moreover by reducing the dataset the linear formulation of the model enhances interpretability by keeping just the information needed for the diagnosis This model has been tested on several wide genetic datasets outperforming state-of-the-art models under study It has also demonstrated remarkable compression capabilities generating the most compact solutions particularly in terms of sample reduction also enhancing the FS Finally we validated the algorithm’s clinical utility through the identified biomarkers and introduced a new approach for interpreting the model’s decision-making process by analyzing the final selected samples Finally we propose a multi-source model based on the Sparse Bayesian Partial Least Squares (SBPLS) model which we have named BAyesian Latent Data Unified Representation (BALDUR) This algorithm projects and combines multiple data sources into a common latent space with the objective of solving a specific diagnostic task studying the differences and similarities of the different data sources by means of the relations established in the latent space Moreover this model can work simultaneously in both the dual and primal spaces allowing it to combine both wide and non-wide views The model also has the ability to discard non-informative features or even entire views keeping only the necessary information to carry out a diagnosis Furthermore this last point not only enhances the model interpretability but also improves its performance by avoiding the introduction of noisy information into the latent space generation The model has been tested on two real datasets related to neurodegenerative diseases one for Parkinson’s disease and the other for Alzheimer’s In both cases BALDUR not only achieved the best performance among all benchmark models but also provided highly interpretable results After conducting an extensive review of the state of the art and with the assistance of our clinical collaborators we were able to validate the biomarkers identified by the model demonstrating its potential as a clinical tool In summary in this thesis we aim to provide insights into the potential of Bayesian algorithms for the development of tools for clinical practice That is the flexibility offered by Bayesian modeling allows researchers to adapt the models to specific and recent needs such as the wide-data and multi-modality nature of some datasets Hence the proposed models stand out when dealing with these types of scenarios by effectively capturing the underlying information of the databases and carrying out precise and reliable diagnosesLa aparición de nuevos métodos de diagnóstico y nuevas tecnologías médicas ha permitido a la comunidad clínica disponer de una creciente cantidad de datos provenientes de diferentes fuentes y de diversa naturaleza Esto ha posibilitado llevar a cabo análisis más exhaustivos y complejos de diferentes enfermedades y trastornos para obtener así diagnósticos más precisos Además muchas pruebas clínicas tales como la imagen médica o los test genéticos permiten la generación de datos de alta precisión Por ejemplo en el caso de la imagen médica podemos contar con imágenes de resoluciones del orden de cientos de miles de vóxeles Sin embargo esta precisión asociada a una alta dimensionalidad junto con la multi-modalidad de este tipo de datos supone nuevos retos para su análisis De hecho su análisis a menudo resulta inviable que los profesionales clínicos analicen todas las relaciones entre las diferentes fuentes de información Es por ello que los algoritmos de Aprendizaje Máquina (Machine Learning ML) están empezando a popularizarse por su potencial para el análisis de datos clínicos de gran complejidad Sin embargo pese a que estos son capaces de procesar de forma muy eficiente diferentes tipos de datos y proporcionar apoyo a la toma de decisiones diagnosticas presentan limitaciones en escenarios con una baja cantidad de muestras y una elevada dimensionalidad lo que se conoce como wide-data En estos escenarios los algoritmos de ML trabajan con muchas mas variables que muestras lo que suele dar lugar a una falta de generalización en las soluciones (sobreentrenamiento) junto con problemas de computo Este lejos de ser un caso aislado es muy común a lo largo de ciertas aplicaciones médicas como la psiquiatría o la neurología En estos casos debido a la complejidad de las enfermedad que requiere de la recogida de datos de diferentes fuentes para realizar diagnósticos precisos unido a la dificultad de reclutar pacientes para estudios clínicos se generan bases de datos con muchas variables (de una o varias vistas) y pocos pacientes o muestras Tratando de dar solución a estas limitaciones en esta tesis nos vamos a centrar en el diseño de diferentes algoritmos de ML capaces de trabajar en escenarios de wide-data Además ya que estos modelos están pensados para trabajar en entornos clínicos deberán ser interpretables Este último punto no solo es importante para ser capaces de explicar como se ha realizado el diagnostico sino para la búsqueda de biomarcadores de utilidad clínica los cuales puedan ayudar en futuros diagnósticos e investigaciones Esta tesis empieza proponiendo un primer modelo inspirándose en el conocido Sparse Bayesian Linear Regression (SBLR) En concreto extiende este modelo para permitirle trabajar de forma eficiente en entornos de wide-data así como de realizar una selección compacta de las variables más relevantes Además debido a su formulación lineal el algoritmo es interpretable Así el modelo propuesto denominado Dual Bayesian Linear Regression with Feature Selection (DBL-FS) es capaz de trabajar directamente sobre el espacio dual mientras impone dispersión sobre las variables del modelo proporcionando así un modelo disperso entrenado mediante el uso de métodos núcleo Este algoritmo ha sido evaluado sobre dos problemas reales dentro del ámbito clínico uno relacionado con esquizofrenia y otro con brotes bacterianos proporcionando en ambos casos prestaciones por encima del resto de modelos de referencia analizados Además con ayuda de nuestros colaboradores médicos hemos sido capaces de validar los biomarcadores encontrados corroborando su utilidad clínica A continuación se presenta la segunda propuesta de esta tesis la cual se inspira en las conocidas Sparse Relevance Vector Machines (SRVM) la formulación bayesiana y dispersa de las Support Vector Machines (SVM) El algoritmo propuesto denominado Relevance Feature and Vector Machine (RFVM) es capaz de trabajar en el espacio dual como el DBL-FS generando además soluciones dispersas tanto en variables como muestras Es decir el modelo es capaz de trabajar de forma eficiente en entornos de wide-data mediante el empleo de métodos núcleo proporcionando una solución final lo más compacta posible manteniendo únicamente las variables y muestras necesarias para realizar la decisión Así aunque pueda sonar contra intuitivo en escenarios donde contamos con pocos datos la eliminación controlada de las muestras ruidosas o atípicas nos permite reducir la selección de características final y mejorar la capacidad de generalización del modelo Además la compacidad de la solución junto a la formulación lineal del algoritmo potencia la interpretabilidad del mismo Este método ha sido evaluado sobre diferentes bases de datos genéticas en las cuales ha obtenido las mejores prestaciones frente al resto de modelos de referencia estudiados Además este algoritmo ha destacado por su gran capacidad compresiva generando las soluciones altamente compactas especialmente en cuanto al numero de muestras lo cual a su vez ayuda a mejorar la selección de características Por otro lado hemos podido validar la utilidad clínica del modelo mediante los biomarcadores encontrados así como de ayudar a interpretar la toma de decisiones del modelo mediante el análisis de las muestras finales seleccionadas Finalmente se ha propuesto un algoritmo multi-modal basado en el conocido Sparse Bayesian Partial Least Squares (SBPLS) el cual hemos llamado BAyesian Latent Data Unified Representation (BALDUR) Este método permite proyectar y combinar eficientemente diferentes tipos de datos médicos dentro de un espacio latente común con el objetivo de resolver una tarea diagnostica concreta Además este modelo es capaz de trabajar de forma simultanea tanto en el espacio dual como primal siendo capaz de combinar tanto vistas wide como no-wide Por otro lado es capaz de eliminar aquellas variables o incluso modalidades que no son informativas para resolver la tarea diagnostica Este último punto no solo ayuda a generar un modelo más interpretable eliminando la información irrelevante sino que mejora las prestaciones del modelo evitando la introducción de ruido en el espacio latente Este método ha sido evaluado sobre dos bases de datos reales relacionadas con enfermedades neurodegenerativas una de Parkinson y otra de Alzheimer En ambos casos no sólo ha dado las mejores prestaciones de entre todos los modelos estudiados sino que ha proporcionado resultados altamente interpretables Además mediante una revisión del estado del arte y con la ayuda de nuestros colaboradores clínicos hemos podido validar los biomarcadores encontrados por el modelo y así demostrar su potencial uso clínico En resumen en esta tesis destacamos el potencial de los modelos bayesianos para el desarrollo de herramientas de uso clínico Esto se debe a la flexibilidad que aporta el modelado bayesiano que nos permite adaptar los algoritmos a diferentes escenarios como por ejemplo al manejo de wide-data y bases de datos multi-modales Así los modelos propuestos en esta tesis son capaces de trabajar de forma eficiente dentro de estos escenarios capturando la información subyacente de los datos y llevando a cabo diagnósticos precisos y fiablesPrograma de Doctorado en Tratamiento de Señales e Ingeniería de las Comunicaciones por la Universidad Carlos III de MadridPresidenta: Inmaculada Mora Jiménez.- Secretaria: Ascensión Gallardo Antolín.- Vocal: Raúl Santos Rodrígue