INRIA a CCSD electronic archive server
Not a member yet
122212 research outputs found
Sort by
Learning-Guided Force-Feedback Model Predictive Control with Obstacle Avoidance for Robotic Deburring
Model Predictive Control (MPC) is widely used for torque-controlled robots, but classical formulations often neglect real-time force feedback and struggle with contact-rich industrial tasks under collision constraints. Deburring in particular requires precise tool insertion, stable force regulation, and collision-free circular motions in challenging configurations, which exceeds the capability of standard MPC pipelines. We propose a framework that integrates force-feedback MPC with diffusion-based motion priors to address these challenges. The diffusion model serves as a memory of motion strategies, providing robust initialization and adaptation across multiple task instances, while MPC ensures safe execution with explicit force tracking, torque feasibility, and collision avoidance. We validate our approach on a torque-controlled manipulator performing industrial deburring tasks. Experiments demonstrate reliable tool insertion, accurate normal force tracking, and circular deburring motions even in hard-to-reach configurations and under obstacle constraints. To our knowledge, this is the first integration of diffusion motion priors with force-feedback MPC for collision-aware, contact-rich industrial tasks
Fairness in Generative AI is Understudied, Underachieved, Undervalued
Despite groundbreaking advancements in generative models during the last decade, concerns about their fairness remain underexplored. Behind their impressive capabilities, these models perpetuate and amplify biases present in their training data, reinforcing societal inequalities and harming marginalized groups. Yet, fairness in generative models has been overshadowed by a rush towards efficiency and novelty in the research community, effectively treating it as a side problem. This paper argues that fairness necessitates to be wholly integrated into generative AI design as a performance-critical component. Supporting our argumentation, we highlight the complex, context-dependent nature of fairness, the challenges of quantifying it as well as the potential trade-offs it incurs. We deduce several research recommendations to ultimately urge the community to shift focus and integrate fairness more deeply into the development of generative models.</div
Quantifying cell traction forces at the single-fiber scale in 3D: An approach based on deformable photopolymerized fiber arrays
International audienceThe forces exerted by cells upon the fibers of the extracellular matrix play a decisive role in cell motility in physiopathology. How the local physical properties of the matrix (density, stiffness, orientation) affect cellular forces remains, however, poorly understood. Existing approaches to measure cell three-dimensional (3D) traction forces within fibrous substrates lack control over the local properties and rely on continuum approaches, not suited for measuring forces at the scale of individual fibers. Herein, an approach is proposed to fabricate multilayer arrays of suspended deformable fibers spanning a wide range of fine-tunable geometrical and mechanical properties using two-photon polymerization. Atomic Force Microscopy is used to thoroughly investigate the properties of individual fibers, including Young’s modulus and stiffness. This approach is combined with a reference-free method for measuring traction forces in 3D, which relies on automated segmentation of the fibers coupled with finite element modeling. The force measurement pipeline is applied to study forces exerted by endothelial cells, fibroblasts, or macrophages, and reveals how these forces are influenced by fiber density and stiffness. Additionally, coupling to fast volumetric imaging with lattice light-sheet microscopy enables the measurement of the low-intensity and short-lived tractions exerted by amoeboid cells, such as dendritic cells. Our technology will be instrumental for monitoring and studying cell behavior at the single-fiber level at extracellular matrix density interfaces, which play a crucial role in both physiological and pathological contexts, such as tumor boundaries
Reconstruction de la séquence d’activation du coeur par mesure non invasive
Cardiac arrhythmias are among the leading causes of death worldwide. While some can be detected by clinical tests such as electrocardiograms (ECGs), others occur suddenly in patients considered healthy and can sometimes lead to sudden death. A key question is therefore whether there are any warning signs of these events in the electrical signals from the torso surface and, if so, how to identify them. In particular, during heart contractions, the cells are synchronized by an electrical activation wave that travels throughout the heart. The activation sequence of the heart cells can then provide cardiologists with valuable information about the origin of possible rhythm disorders. The inverse problem of cardiac electrophysiology, or electrocardiographic imaging (ECGi), aims to reconstruct the electrical activity of the heart, and in particular the activation sequence, from non-invasive measurements of electrical potential on the surface of the torso. This problem is mathematically ill-posed in the sense of Hadamard, which makes it particularly difficult to solve. Studying the properties of this inverse problem and developing new methods for solving it is a major challenge for improving and systematizing the detection of cardiac arrhythmias. In this context, this thesis proposes new approaches to addressing the problem of electrocardiographic imaging. Classically, the ECGi problem is formulated as a constrained optimization problem, requiring the satisfaction of a partial mathematical model of the heart’s electrical activity, called the ”source model”. Constraint terms, called ”regularization”, are also added to ensure the invertibility and stability of the minimization problem. The choice of source model, as well as the constraints imposed on the electrical potentials and voltages on the heart modify significantly the properties of this inverse problem. In this thesis, we first focused on describing the inherent difficulties of the inverse problem in cardiac electrophysiology. We presented several formulations of the problem, from source models describing the underlying electrophysiology to the mathematical and numerical methods used for its resolution. We then developed a new source model for ECGi. The ”depth-averaged” model was first established in two dimensions for simplified geometries, then heuristically extended to three dimensions under the name ”epicardial model”. This model couples a surface transmembrane voltage and a surface extracellular potential with a volume extracardiac potential in the torso, while accounting for flux exchanges between the heart and the torso. It thus offers more diverse modelling possibilities than the models commonly used in ECGi. Certain properties of the epicardial model were used to numerically study the inverse problem in three dimensions. The presence of surface conductivities made it possible to evaluate the impact of the a priori integration of cardiac fibres. This framework also allowed for the introduction of a more physiologically realistic regularisation, based on the total variation of the transmembrane voltage, and for the comparison of different methods for extracting activation sequences from the reconstructed data. In a final exploratory study, we proposed a new method for solving the inverse problem, based on Bayesian theory, combining a simplified model of cardiac electrical activation with a particle filtering technique. This approach makes it possible to reconstruct the activation sequence while explicitly incorporating a degree of uncertainty, thus paving the way for a probabilistic interpretation of electrocardiographic imaging results.Les arythmies cardiaques figurent parmi les principales causes de mortalité dans le monde. Si certaines peuvent être détectées par des examens cliniques tels que l’électrocardiogramme (ECG), d’autres surviennent de manière soudaine chez des patients considérés comme sains, et peuvent parfois mener à une mort subite. Une question centrale est donc de déterminer si des signes précurseurs de ces événements sont présents dans les signaux électriques du torse et, le cas échéant, comment les identifier. En particulier, lors des contractions cardiaques, les cellules sont synchronisées par une onde d’activation électrique, qui parcourt tout le cœur. La séquence d’activation des cellules du cœur peut alors fournir aux cardiologues des indications précieuses sur l’origine d’éventuelles pathologies du rythme. Le problème inverse de l’électrophysiologie cardiaque, ou imagerie électrocardiographique (ECGi), vise à reconstruire l’activité électrique du cœur, et notamment la séquence d’activation, à partir de mesures non invasives de potentiel électrique à la surface du torse. Ce problème est mathématiquement mal posé au sens de Hadamard, ce qui le rend particulièrement difficile à résoudre. étudier les propriétés de ce problème inverse et développer de nouvelles méthodes de résolution constitue un enjeu majeur pour améliorer et systématiser la détection des troubles du rythme cardiaque. Dans ce cadre, cette thèse propose de nouvelles approches pour aborder le problème de l’imagerie électrocardiographique. Classiquement, le problème de l’ECGi est formulé comme une optimisation sous contrainte de satisfaire un modèle mathématique partiel de l’activité électrique du cœur, appelé ”modèle source”. Des termes de contrainte, appelés ”régularisation” sont également ajoutés pour assurer l’inversibilité et la stabilité du problème de minimisation. Le choix du modèle source, ainsi que des contraintes imposées aux potentiels et tensions électriques sur le cœur jouent un rôle déterminant dans les propriétés de ce problème inverse. Dans ce travail, nous nous sommes tout d’abord attachés à décrire les difficultés inhérentes au problème inverse d’électrophysiologie cardiaque. Nous avons présenté plusieurs formulations de ce problème, des modèles source décrivant l’électrophysiologie sous-jacente aux méthodes mathématiques et numériques pour le résoudre. Nous avons ensuite élaboré un nouveau modèle source pour l’ECGi. Le modèle ”moyenné en épaisseur”, a été établi en deux dimensions dans des géométries simplifiées, avant d’être étendu de fa¸con heuristique en trois dimensions, sous le nom de ”modèle épicardique”. Ce modèle couple une tension transmembranaire et un potentiel extracellulaire de surface avec des potentiels extracardiaques dans le torse, tout en prenant en compte les échanges de flux entre le cœur et le torse. Il offre ainsi une plus grande richesse de modélisation que les modèles classiquement utilisés dans le cadre de l’ECGi. Certaines propriétés du modèle épicardique ont été exploitées pour étudier numériquement le problème inverse en trois dimensions. La présence de conductivités surfaciques a permis d’évaluer l’impact de l’intégration a priori de fibres cardiaques. Ce modèle a également rendu possible l’introduction d’une régularisation plus physiologique, fondée sur la variation totale de la tension transmembranaire, ainsi que la comparaison de différentes méthodes d’extraction des séquences d’activation à partir des potentiels reconstruits. Enfin, dans un travail exploratoire, nous avons proposé une nouvelle méthode de résolution du problème inverse, fondée sur la théorie bayésienne, combinant un modèle simplifié de l’activation électrique cardiaque avec une technique de filtrage particulaire. Cette approche permet de reconstruire la séquence d’activation tout en intégrant explicitement une part d’incertitude, ouvrant ainsi la voie à une interprétation probabiliste des résultats de l’imagerie électrocardiographique
Evaluating the effects of preprocessing, method selection, and hyperparameter tuning on SAR-based flood mapping and water depth estimation
International audienceFlood mapping and water depth estimation from Synthetic Aperture Radar (SAR) imagery are crucial for calibrating and validating hydraulic models. This study uses SAR imagery to evaluate various preprocessing (especially speckle noise reduction), flood mapping, and water depth estimation methods. The impact of the choice of method at different steps and its hyperparameters is studied by considering an ensemble of preprocessed images, flood maps, and water depth fields.The evaluation is conducted for two flood events on the Garonne River (France) in 2019 and 2021, using hydrodynamic simulations and in-situ observations as reference data. Results show that the choice of speckle filter alters flood extent estimations with variations of several square kilometers. Furthermore, the selection and tuning of flood mapping methods also affect performance. While supervised methods outperformed unsupervised ones, tuned unsupervised approaches (such as local thresholding or change detection) can achieve comparable results. The compounded uncertainty from preprocessing and flood mapping steps also introduces high variability in the water depth field estimates.This study highlights the importance of considering the entire processing pipeline, encompassing preprocessing, flood mapping, and water depth estimation methods and their associated hyperparameters. Rather than relying on a single configuration, adopting an ensemble approach and accounting for methodological uncertainty should be privileged. For flood mapping, the method choice has the most influence. For water depth estimation, the most influential processing step was the flood map input resulting from the flood mapping step and the hyperparameters of the methods
Prédiction de mouvement avec prise en compte de l'incertitude via des grilles d'occupation dynamiques
Scene understanding and situational awareness remain fundamental challenges in autonomous driving, as accurate perception and prediction of the environment are critical for ensuring safety. These challenges arise from various factors such as partial or unreliable sensing, unfamiliar environments, complex behavior of agents, and possibility of diverse future behaviors. To address these issues, this thesis explores the potential of Bayesian Dynamic Occupancy Grid Maps (DOGMs) to model and predict the scene evolution through deep learning techniques. The original contribution to knowledge of this research is the development of a multi-task motion prediction framework that integrates DOGMs with semantics to bridge agent-specific behavior prediction and generalized scene-level motion forecasting. Within this framework, occupancy flow predictions are leveraged to represent dynamic behavior, enabling the model to anticipate future occupancy states of both recognized and unrecognized agents.In the first phase, the thesis examines the multi-step prediction of agent-agnostic occupancy grid maps as a video prediction problem. To address the inherent spatiotemporal challenges of scene dissolution and fading artifacts, DOGM prediction with a fixed reference frame is proposed and compared against the conventional ego-centric DOGM approach. Experiments with various video prediction networks demonstrate that an allo-centric DOGM representation has superior ability to predict the same urban driving scene. This constitutes the first contribution of the thesis.Building on this, the second contribution explores a new research direction in DOGM-based future prediction through vehicle motion forecasting. A novel framework is developed that incorporates semantic labels and map information with DOGMs to predict future vehicle segmentation grids.Moving beyond video prediction networks, this approach studies deep learning-based spatiotemporal modeling and conditional variational method to anticipate diverse vehicle behaviors. Experiments highlight the potential of incorporating vehicle semantic labels into the input, alongside predicting a sequence of vehicle-specific grids. This approach effectively captures complex motion patterns in relation to surrounding agents as well as scene structures.The third contribution extends these insights by formulating a multi-task framework that leverages occupancy flow prediction to integrate agent-agnostic and vehicle-specific motion forecasting.The framework jointly predicts sequence of vehicle grids, occupancy flow, and occupancy state grids iii iv across the entire driving scene. Prediction of occupancy flow grids significantly improves performance, particularly in uncertain environments and when learning multi-modal future trajectories in complex driving scenarios. Each grid provides complementary and redundant information, collectively refining prediction capabilities.Overall, this thesis makes several contributions to motion prediction in autonomous driving by leveraging DOGMs for scene evolution modeling. All algorithms were validated on real-world datasets, demonstrating the effectiveness of the proposed methods. The findings have significant implications for safer navigation and collision avoidance, offering improvements in interaction-aware planning and enabling better generalization to unseen scenarios.La compréhension de la scène et la perception situationnelle restent des défis fondamentaux en conduite autonome, car une perception et une prédiction précises de l'environnement sont essentielles pour garantir la sécurité. Ces défis résultent de divers facteurs, tels que des capteurs partiellement fiables, des environnements inconnus, le comportement complexe des agents et la diversité des évolutions futures possibles. Pour répondre à ces problématiques, cette thèse explore le potentiel des grilles d'occupations dynamiques (DOGMs) bayésiennes pour modéliser et prédire l'évolution de la scène à l'aide de techniques d'apprentissage profond. La contribution originale de cette recherche réside dans le développement d'un modèle multi-tâches de prédiction de mouvement intégrant les DOGMs avec les informations sémantiques afin de combiner la prédiction des comportements spécifiques aux agents et la prévision généralisée des dynamiques de la scène. Dans ce modèle, la prédiction du flux d'occupation est exploitée pour représenter les dynamiques de mouvement, permettant ainsi au modèle d'anticiper les futurs états d'occupation des agents reconnus et non reconnus.Dans un premier temps, la thèse examine la prédiction multi-temporelle des grilles d'occupation indépendantes des agents comme un problème de prédiction vidéo. Afin de résoudre les problèmes spatiotemporels inhérents, tels que le flou et la disparition d'éléments de la scène, une prédiction de DOGM avec un référentiel fixe est proposée et comparée à l'approche conventionnelle en référentiel égo-centrique. Des expériences menées avec différents réseaux de prédiction vidéo démontrent que la représentation allo-centrique des DOGMs offre une meilleure capacité à prédire la même scène de conduite urbaine. Cette approche constitue la première contribution de la thèse.Dans un second temps, la thèse explore une nouvelle direction de recherche en matière de prédiction future basée sur les DOGMs, en se concentrant sur la prévision des mouvements des véhicules. Un modèle innovant est développé, intégrant les labels sémantiques et les informations cartographiques aux DOGMs afin de prédire les futures grilles de segmentation des véhicules. Dépassant les réseaux de prédiction vidéo, cette approche exploite des modèles spatiotemporels basés sur l'apprentissage profond et une méthode variationnelle conditionnelle pour anticiper la diversité des comportements des véhicules. Les expériences mettent en évidence le potentiel de l'intégration des labels sémantiques des véhicules en entrée, ainsi que la prédiction d'une séquence de grilles spécifiques aux véhicules. Cette approche permet de capturer efficacement les motifs complexes de déplacement en lien avec les agents environnants et la structure de la scène.La troisième contribution étend ces travaux en formulant un modèle multi-tâches exploitant la prédiction du flux d'occupation afin d'intégrer la prévision des mouvements à la fois indépendants des agents et spécifiques aux véhicules. Le modèle prédit conjointement une séquence de grilles de véhicules, le flux d'occupation et les états d'occupation sur l'ensemble de la scène de conduite. La prédiction du flux d'occupation améliore significativement les performances, en particulier dans des environnements incertains et pour l'apprentissage de trajectoires futures multi-modales dans des scénarios de conduite complexes. Chaque grille fournit des informations complémentaires et redondantes, affinant ainsi la précision des prédictions.Dans l'ensemble, cette thèse apporte plusieurs contributions à la prédiction du mouvement en conduite autonome en exploitant les DOGMs pour la modélisation de l'évolution de la scène. Tous les algorithmes ont été validés sur des ensembles de données réelles, démontrant l'efficacité des méthodes proposées. Ces résultats ont des implications majeures pour une navigation plus sûre et une meilleure prévention des collisions, en améliorant la planification tenant compte des interactions et en permettant une meilleure généralisation aux scénarios précédement non rencontrés
Towards Interpretable Time Series Foundation Models
International audienceIn this paper, we investigate the distillation of time series reasoning capabilities into small, instruction-tuned language models as a step toward building interpretable time series foundation models. Leveraging a synthetic dataset of meanreverting time series with systematically varied trends and noise levels, we generate natural language annotations using a large multimodal model and use these to supervise the fine-tuning of compact Qwen models. We introduce evaluation metrics that assess the quality of the distilled reasoning-focusing on trend direction, noise intensity, and extremum localization-and show that the post-trained models acquire meaningful interpretive capabilities. Our results highlight the feasibility of compressing time series understanding into lightweight, language-capable models suitable for on-device or privacy-sensitive deployment. This work contributes a concrete foundation toward developing small, interpretable models that explain temporal patterns in natural language.</div
18th International Conference on Similarity Search and Applications: 18th International Conference, SISAP 2025, Reykjavik, Iceland, October 1–3, 2025, Proceedings
International audienceThis book constitutes the refereed proceedings of the 18th International Conference on Similarity Search and Applications, SISAP 2025, held in Reykjavik, Iceland, during October 2025.The 18 full papers and 9 short papers included in the proceedings were carefully reviewed and selected from 58 submissions. The book also contains 2 Doctoral Symposium papers, 3 demonstration papers, and 5 indexing challenge papers. The contributions were organized in the following topical sections: Applications; Foundations; Search and Retrieval; Similarity Measures; Clustering; Indexing; Learning-Based Methods; Doctoral Symposium; Demonstration Papers; and Indexing Challenge
Does 3D Gaussian Splatting Need Accurate Volumetric Rendering?
International audienceAbstract Since its introduction, 3D Gaussian Splatting (3DGS) has become an important reference method for learning 3D representations of a captured scene, allowing real‐time novel‐view synthesis with high visual quality and fast training times. Neural Radiance Fields (NeRFs), which preceded 3DGS, are based on a principled ray‐marching approach for volumetric rendering. In contrast, while sharing a similar image formation model with NeRF, 3DGS uses a hybrid rendering solution that builds on the strengths of volume rendering and primitive rasterization. A crucial benefit of 3DGS is its performance, achieved through a set of approximations, in many cases with respect to volumetric rendering theory. A naturally arising question is whether replacing these approximations with more principled volumetric rendering solutions can improve the quality of 3DGS. In this paper, we present an in‐depth analysis of the various approximations and assumptions used by the original 3DGS solution. We demonstrate that, while more accurate volumetric rendering can help for low numbers of primitives, the power of efficient optimization and the large number of Gaussians allows 3DGS to outperform volumetric rendering despite its approximations