Portail HAL des publications du LIRMM
Not a member yet
13279 research outputs found
Sort by
Méthodes et Modèles pour l'élaboration automatisée de Graphes de Connaissances dans le domaine juridique : Application aux Ressources Juridiques et Juridico-Pratiques des Collectivités Locales et Territoriales
This thesis examines the construction of knowledge graphs from texts, focusing primarily on information extraction from unstructured texts. The objective of this work is to explore various aspects of information extraction from specialized corpora in the legal domain. To this end, we divide this study into two subtasks: terminology extraction and relation extraction.Terminology extraction aims to automatically identify relevant terms in a given corpus of texts. Subsequently, from these terms, we extract the relations that link them. For relation extraction, two approaches are conceivable: either by determining in advance the types of relations that structure the terms, or by using the context, notably the verbs or other actions that link these terms (in the field of OpenIE). Thus, we address three main sub-problems.We introduce the terminology extraction system InfoGlean KeyTerms, composed of three modules: one for named entity recognition (NER), one for the extraction of relevant terms/text segments (KPE), and a final one for the extraction of legal entities. Expert annotations were provided in addition to this system to constitute a terminological base.After building this terminological base, we implemented two relation extraction systems: Relational Embeddings Model (REM) and GOREX. REM identifies typed relations between the extracted terms using the lexical network rezoJDM. REM represents the relation pairs using a Word2Vec model, then classifies the types of relations. GOREX, on the other hand, exploits the principle of OpenIE by focusing on the verbs or action terms in the local context of the terms. GOREX uses LLMs to perform this task.The analysis of the results revealed promising research avenues to be explored in future work for all systems. More specifically, the implementation of a hybrid relation extraction system could be an interesting path to explore.Cette thèse examine la construction de graphes de connaissances à partir de textes, en se concentrant principalement sur l'extraction d'informations à partir de textes non structurés. L'objectif de ce travail est d'explorer différents aspects de l'extraction d'informations à partir de corpus spécialisés dans le domaine juridique. Pour ce faire, nous divisons cette étude en deux sous-tâches : l'extraction terminologique et l'extraction de relations.L'extraction terminologique vise à identifier automatiquement les termes pertinents dans un corpus de textes donné. Ensuite, à partir de ces termes, nous extrayons les relations qui les lient. Pour l'extraction de relations, deux approches sont envisageables : soit en déterminant au préalable des types de relations qui structurent les termes, soit en utilisant le contexte, notamment les verbes ou autres actions qui relient ces termes (dans le domaine de l'OpenIE). Nous abordons donc trois sous-problématiques principales.Nous introduisons le système d'extraction terminologique InfoGlean KeyTerms, composé de trois modules : un pour la reconnaissance d'entités nommées (NER), un pour l'extraction de termes/segments textuels pertinents (KPE) et un dernier pour l'extraction d'entités juridiques. Des annotations d'experts ont été fournies en complément à ce système pour constituer une base terminologique.Après la construction de cette base terminologique, nous avons mis en place deux systèmes d'extraction de relations : Relational Embeddings Model (REM) et GOREX. REM identifie des relations typées entre les termes extraits en utilisant le réseau lexical rezoJDM. REM représente les paires de relations à l'aide d'un modèle Word2Vec, puis classe les types de relations. GOREX, quant à lui, exploite le principe de l'OpenIE en se concentrant sur les verbes ou termes d'action dans le contexte local des termes. GOREX utilise les LLMs pour réaliser cette tâche.L'analyse des résultats a révélé des pistes de recherche prometteuses à approfondir dans les travaux futurs pour l'ensemble des systèmes. Plus spécifiquement, la mise en place d'un système hybride d'extraction de relations pourrait être un chemin intéressant à explorer
RCAviz
RCAviz is an online visualization tool created for exploring knowledge organized into interconnected classifications using Relational Concept Analysis [Rouane Hacène 2013]. This tool has been motivated by the Knomana project, which aims to collect and explore knowledge on pesticidal and antimicrobial plants, in order to replace synthetic products in human, animal and plant health. A detailed description is available in [Muller 2022]
A SHAP-based controversy analysis through communities on Twitter
International audienceControversy encompasses content that draws diverse perspectives, along with positive and negative feedback on a specific event, resulting in the formation of distinct user communities. we explore the explainability of controversy through the lens of SHAP (SHapley Additive exPlanations) method, aiming to provide a fair assessment of the individual contributions of different text features of tweets to controversy detection. We conduct an analysis of topic discussions on Twitter from a community perspective, investigating the role of text in accurately classifying tweets into their respective communities. To achieve this, we introduce a SHAP-based pipeline designed to quantify the influence of impactful text features on the predictions of three tweet classifiers. Text content alone offers interesting controversy detection accuracy. It can contain predictive features for controversy detection. For instance, negative connotations, pejorative tendencies and positive qualifying adjectives tend to impact the controversy model detection
Evaluating Human Trajectory Prediction with Metamorphic Testing
International audienceThe prediction of human trajectories is important for planning in autonomous systems that act in the real world, e.g. automated driving or mobile robots. Human trajectory prediction is a noisy process, and no prediction does precisely match any future trajectory. It is therefore approached as a stochastic problem, where the goal is to minimise the error between the true and the predicted trajectory. In this work, we explore the application of metamorphic testing for human trajectory prediction. Metamorphic testing is designed to handle unclear or missing test oracles. It is well-designed for human trajectory prediction, where there is no clear criterion of correct or incorrect human behaviour. Metamorphic relations rely on transformations over source test cases and exploit invariants. A setting well-designed for human trajectory prediction where there are many symmetries of expected human behaviour under variations of the input, e.g. mirroring and rescaling of the input data. We discuss how metamorphic testing can be applied to stochastic human trajectory prediction and introduce the Wasserstein Violation Criterion to statistically assess whether a follow-up test case violates a label-preserving metamorphic relation
Comparing two bootstrapped regions in images: the D-test
International audienceObjectives: Many molecular imaging diagnoses involve comparing two regions of interest (ROIs) in the image or different images. Since the images are obtained by measuring a random phenomenon, such comparisons should be based on a statistical test to ensure reliability. Recent studies have shown that use of the bootstrap approach provides access to the statistical variability of reconstructed values in molecular images. However, although there is general agreement that this increase in information should make diagnosis based on molecular images more reliable, no approach has been proposed in the relevant literature to use bootstrap replicates to enhance the reliability of comparisons of two ROIs. In this paper, we propose to fill this gap by introducing the first statistical test that allows us to compare two sets of pixels/voxels for which bootstrap replicates are available. Material and methods : After presenting the theoretical basis of this non-parametric statistical test, this article describes how to calculate it in practice. Finally, it proposes two experiments based on quantitative comparisons and expert judgment to assess its relevance. Results : The results obtained are consistent with expert diagnosis on synthetic data. This validates the relevance of the D-test. Conclusion : This paper presents the first statistical test to compare two ROIs in reconstructed images for which the statistical variability information is accessible
Low-Resource Fully-Digital BPSK Demodulation Technique for Intra-Body Wireless Sensor Networks
International audienceIn the context of intra-body communication, gal-vanic coupling has gained significant attention. This communication technique leverages conductivity of biological tissues to use them as a medium for modulated electrical signals. This paper introduces a novel fully-digital demodulation technique for use in galvanic coupling transmission with BPSK-modulated signals at low frequencies (below MHz). The technique relies on an original digital processing algorithm applied on samples collected by an analogue-to-digital converter. The algorithm employs a voting principle among several demodulated data candidates and implement an error detection scheme to offer high robustness and reliability. Moreover, it has been specifically designed to minimise the computational operations and resources required to enable a compact and energy-efficient implementation. Simulations and experiments demonstrate good performance, with a perfect transmission success rate of 100% for signals with a Signal-to-Noise Ratio (SNR) of -3 dB or higher, and a success rate that remains higher than 97% for signals with an SNR down to -7dB
Don't bite the hand that feeds you: Meta food webs help in the face of the Eltonian shortfall
International audienceThis article is a Response to the Letter by Brimacombe et al, https://doi.org/10.1111/gcb.17360, which was related to the paper of Botella et al., https://doi.org/10.1111/gcb.17167
A Structural Testing Approach for SRAMs using Cell-Aware Methodology
International audienc
Nonvolatile and SEU-Recoverable Latch Based on FeFET and CMOS for Energy-Harvesting Devices
International audienceNonvolatile memories are widely used in emerging energy-harvesting Internet-of-Things (IoT) applications, and nonvolatile memories constructed from FeFET devices hold great promise. This paper presents a nonvolatile and single-event-upset (SEU)-recoverable latch based on FeFET and CMOS for energyharvesting devices. The latch uses n-type FeFET devices to provide nonvolatility without any additional control signals. Moreover, since the soft error problem has become increasingly severe, radiation hardening by design gains a great attention as a promising approach to mitigate the reliability issue. The latch uses feedback interlocked loops with n-type FeFETs and C-elements, enabling it to provide nonvolatility and SEU-recovery simultaneously. Simulation results with Candence Virtuoso verifies that the proposed latch design has correct functioning with excellent performance compared to the state-of-the-art designs
Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models
International audienceAs a promising paradigm to collaboratively train models with decentralized data, Federated Learning (FL) can be exploited to fine-tune Large Language Models (LLMs). While LLMs correspond to huge size, the scale of the training data significantly increases, which leads to tremendous amounts of computation and communication costs. The training data is generally non-Independent and Identically Distributed (non-IID), which requires adaptive data pro- cessing within each device. Although Low-Rank Adaptation (LoRA) can significantly reduce the scale of parameters to update in the fine-tuning process, it still takes unaffordable time to transfer the low-rank parameters of all the layers in LLMs. In this paper, we propose a Fisher Information-based Efficient Curriculum Federated Learning framework (FibecFed) with two novel methods, i.e., adaptive federated curriculum learning and efficient sparse parameter update. First, we propose a fisher information- based method to adaptively sample data within each device to improve the effectiveness of the FL fine-tuning process. Second, we dynamically select the proper layers for global aggregation and sparse parameters for local update with LoRA so as to improve the efficiency of the FL fine-tuning process. Extensive experimental results based on 10 datasets demonstrate that FibecFed yields excellent performance (up to 45.35% in terms of accuracy) and superb fine-tuning speed (up to 98.61% faster) com- pared with 17 baseline approaches). Our code will be publicly available