HAL Portal UTC Université de Technologie de Compiègne
Not a member yet
    11652 research outputs found

    Des IA au service de l'espace littéraire du XIXe siècle : évaluation et analyse des outils de reconnaissance d'entités nommées spatiales

    No full text
    In fields related to the acquisition and enhancement of spatial data, geographic Named Entity Recognition (NER) can become an entry point for literary analysis. When NER is applied to literary text corpora, its results can be expressed, among other formats, as maps. The task of transcribing images (such as PDFs) into text, which is the first step in the processing chain for acquiring textual data and creating digital corpora, can be done manually or by using Optical Character Recognition (OCR) systems. While the use of an OCR system allows for much faster task execution compared to long and tedious manual entry, it is undeniable that OCR introduces transcription errors. These errors can involve noise or silence (addition, substitution, and deletion of characters). The same applies to manual transcriptions, where noise and silence can also be semantic in nature, with one term being replaced by another.Machine-generated errors can often be repetitive, which means that they can be modeled and automatically corrected in the outputs. However, when they are unique, implementing an effective automatic correction routine remains challenging. To overcome the issue of OCR transcription quality and obtain higher-quality texts, users implement cleaning strategies that are costly in both time and resources. Nevertheless, one might ask to what extent errors in spatial named entity recognition are truly attributable to data noise, and what impact this noise has on the subsequent uses of this data. In a user survey conducted among researchers in Digital Humanities and Natural Language Processing, we observed that many of them use uncorrected OCR transcriptions to “probe” their corpus and gain an initial overview of the major trends to analyze. However, do these corpora allow the examination of attested phenomena that can lead to viable scientific applications? We present an evaluation of various NER tools on raw OCR transcriptions. To assess the robustness of these NER tools, we compare the results on ten books, each comprising between 100 and 400 pages, (in their reference and OCR versions) from the French, English, and Portuguese collections of ELTeC - European Literary Text Collection. We demonstrate that a significant proportion of the errors are attributable to the models and not to the data, and, quite surprisingly, that certain entities are detected in noisy data even when they are misspelled. Thus, these interferences in NER, arising from silences and noise in corpus processing — where some named entities are not recognized, resulting in “silence,” and others are detected in error, introducing “noise” — are not merely inherent to the quality of the input data.While noise can be overcome (through various strategies such as correction, careful reading, etc.), silence results in the loss of information and remains an open issue regarding its impact on many potential future uses. This research aims to establish methods for evaluating NER systems on real-world data, meaning data as it is acquired by literary researchers, and focuses on finding strategies to assist users in the extraction and exploration of this data.Dans les domaines liés à l'acquisition et à la valorisation des données spatiales, la reconnaissance d'entités nommées (REN ou en anglais NER pour Named Entity Recognition) géographiques peu devenir un point d'entrée à l'analyse littéraire. La tâche de REN appliquée à des corpus de texte littéraires, voit ses résultats s'exprimer, entre autres, sous forme de cartes. La tâche de transcription de l'image (PDF par exemple) en texte, qui est le premier maillon de la chaîne de traitement pour l'acquisition des données textuelles et la constitution des corpus numériques, peut être assurée manuellement ou par utilisation des systèmes de reconnaissance optique de caractère (OCR).Si l'utilisation d'un système OCR permet une plus grande rapidité d'exécution de la tâche, comparativement à une saisie manuelle longue et fastidieuse, force est de constater qu'elle est à l'origine d'erreurs de transcription. Il peut s'agir de bruit ou de silence (ajout, substitution et suppression de caractères). Il en va de même pour les transcriptions manuelles, le bruit et le silence pouvant aussi être de nature sémantique, un terme est remplacé par un autre. Ces erreurs produites par la machine peuvent être répétitives et donc modélisables et corrigeables automatiquement dans les sorties, mais lorsqu'elles sont singulières une routine de correction automatique efficace reste difficile à mettre en place. Pour pallier le problème lié à la qualité des transcriptions OCR et obtenir des textes de meilleure qualité, les utilisateurs mettent en place des stratégies, de nettoyage, coûteuses en temps et en moyen. Néanmoins, on peut se demander dans quelle mesure les erreurs dans la reconnaissance des entités nommées spatiales sont réellement imputables au bruitage des données, et quel est l'impact de ce bruit sur les usages consécutifs de ses données. Nous avons constaté au cours d'une enquête utilisateur menée auprès de chercheurs en Humanités Numériques ou en Traitement Automatique des Langues que nombre d'entre eux utilisent des transcriptions OCR non corrigées pour « sonder » leur corpus et avoir une première vision des grandes tendances à analyser. Mais, ces corpus permettent-ils d'examiner des phénomènes attestés qui donneront lieu à des applications scientifiques viables ? Nous présentons une évaluation de différents outils de REN sur des transcriptions OCR brutes. Afin d'évaluer la robustesse des outils de REN, nous comparons les résultats sur dix œuvres (dans leur version de référence et leur version OCR) issues des collections française, anglaise et portugaise d'ELTeC - European Literary Text Collection, comptant entre 100 et 400 pages chacune. Nous montrons qu'une proportion non négligeable des erreurs est imputable aux modèles, et non aux données, et que, de façon très étonnante, certaines entités sont détectées dans les données bruitées, même lorsqu'elles sont mal orthographiées. Ainsi, ces interférences dans la REN, qui découlent des silences et bruits dans les exploitations des corpus -- certaines entités nommées ne sont pas reconnues, amenant donc du « silence », d'autres sont détectées par erreur, apportant du « bruit » -- ne sont pas simplement inhérentes à la qualité des données en entrée. Si le bruit peut être dépassé (selon plusieurs stratégies : correction, lecture attentive, etc.), le silence fait disparaître de l'information et reste un problème ouvert quant à son impact sur bon nombre d'utilisations postérieures potentielles. Ces travaux de recherches s'appliquent à établir des méthodes pour l'évaluation des systèmes de REN sur des données du monde réel, c'est à dire des données telles que les chercheurs en littérature les acquiers, et s'attache à trouver des stratégies pour assister l'utilisateur i ce dans l'extraction et l'exploration de ces données

    Adversarial Semi-Supervised Domain Adaptation for Semantic Segmentation: A New Role for Labeled Target Samples

    No full text
    International audienceAdversarial learning baselines for domain adaptation (DA) approaches in the context of semantic segmentation are under explored in semi-supervised framework. These baselines involve solely the available labeled target samples in the supervision loss. In this work, we propose to enhance their usefulness on both semantic segmentation and the single domain classifier neural networks. We design new training objective losses for cases when labeled target data behave as source samples or as real target samples. The underlying rationale is that considering the set of labeled target samples as part of source domain helps reducing the domain discrepancy and, hence, improves the contribution of the adversarial loss. To support our approach, we consider a complementary method that mixes source and labeled target data, then applies the same adaptation process. We further propose an unsupervised selection procedure using entropy to optimize the choice of labeled target samples for adaptation. We illustrate our findings through extensive experiments on the benchmarks GTA5, SYNTHIA, and Cityscapes. The empirical evaluation highlights competitive performance of our proposed approach

    Bounds in Wasserstein Distance for Locally Stationary Functional Time Series

    No full text
    Functional time series (FTS) extend traditional methodologies to accommodate data observed as functions/curves. A significant challenge in FTS consists of accurately capturing the time-dependence structure, especially with the presence of time-varying covariates. When analyzing time series with time-varying statistical properties, locally stationary time series (LSTS) provide a robust framework that allows smooth changes in mean and variance over time. This work investigates Nadaraya-Watson (NW) estimation procedure for the conditional distribution of locally stationary functional time series (LSFTS), where the covariates reside in a semi-metric space endowed with a semi-metric. Under small ball probability and mixing condition, we establish convergence rates of NW estimator for LSFTS with respect to Wasserstein distance. The finite-sample performances of the model and the estimation method are illustrated through extensive numerical experiments both on functional simulated and real data

    Enhancing Recommender Systems Using Textual Embeddings from Pre-trained Language Models

    No full text
    International audienceRecent advancements in language models and pre-trained language models like BERT and RoBERTa have revolutionized natural language processing, enabling a deeper understanding of human-like language. In this paper, we explore enhancing recommender systems using textual embeddings from pre-trained language models to address the limitations of traditional recommender systems that rely solely on explicit features from users, items, and user-item interactions. By transforming structured data into natural language representations, we generate high-dimensional embeddings that capture deeper semantic relationships between users, items, and contexts. Our experiments demonstrate that this approach significantly improves recommendation accuracy and relevance, resulting in more personalized and context-aware recommendations. The findings underscore the potential of PLMs to enhance the effectiveness of recommender systems

    Prediction of temperature and strain rate dependent flow behaviors for AA6061-T4 sheet using phenomenology and machine learning-based approaches

    No full text
    International audienceThe plastic flow behaviors of AA6061-T4 sheets at different temperatures (21-300 degrees C) and strain rates (0.002-4 s-1) were studied. Significant nonlinear effects of temperature and strain rate on flow behaviors were revealed, as well as underlying micromechanical factors. Phenomenology and machine learning-based constitutive models were developed. Both models were formulated in the framework of a temperature-dependent linear combination regulated by a transition function to capture the evolution of strain-hardening behavior with increasing temperature. Novel mathematical functions for describing temperature and strain rate sensitivities were formulated for the phenomenological constitutive model. The threshold temperature related to microstructure evolution was considered in the modeling. A data-enrichment strategy based on extrapolating experimental data via classical strain hardening laws was adopted to improve neural network training. An efficient inverse identification strategy, focusing solely on the transition function, was proposed to enhance the prediction accuracy of post-necking deformation by both constitutive models

    Multi-Stage Microwave-Assisted Extraction of Phenolic Compounds from Tunisian Walnut (Juglans regia L.) Bark

    No full text
    International audienceThis study aimed to optimize the extraction of total phenolic compounds (TPC) from Tunisian walnut bark using microwave treatment. Initially, a preliminary investigation was conducted to establish optimal levels for ethanol concentration, liquid–solid ratio, temperature, and time, which were then applied in subsequent conventional solvent extraction (CSE) experiments. To enhance the extraction yield, multi-stage microwave-assisted extraction (MS MAE) was evaluated using three microwave power settings: 100, 200, and 300 W. The results showed a statistically significant (p < 0.05) effect of microwave irradiation combined with multiple solvent extraction stages. The optimized MS MAE protocol, employing 300 W power, six stages of 10 min each, and a liquid–solid ratio of 10 mL/g, achieved an 86% recovery of TPC. In contrast, extraction involving 10 stages of 30 min each without microwave irradiation recovered only 79% of TPC. UHPLC–MS analysis revealed that the phenolic profile of the extracts was dominated by gallic acid, vanillic acid, and quercetin, and that microwave treatment did not significantly alter the qualitative or quantitative composition of these major phenolic compounds compared to conventional extraction. These findings demonstrate that MS MAE is a time-saving, energy-saving, solvent-reducing, and highly efficient extraction technology for producing bioactive extracts from walnut bark

    Experimental study of biocompatible polycaprolactone composite membranes blended with short oligomers of biosourced triarylmethane-based polyaryletherketone

    No full text
    International audienceOur first original goal was to synthesize new poly(aryl ether ketone)s by using successively one of three different triarylmethane monomers. These monomers were respectively produced from benzaldehyde, vanillin, and veratraldehyde by reaction with two molecules of phenol in excess of sulfuric acid. Thus, each monomer was allowed to react with one equivalent of 4,4′-difluorobenzophenone through a typical SNAR polycondensation in dimethylformamide without the help of toluene to afford only short oligomers of each desired poly(aryl ether ketone). After realizing their physico-chemical characterizations, the oligomers were employed as fillers added at 25 wt% to produce original polycaprolactone composite membranes via a classic solvent casting technique using chloroform, followed by the regeneration of used membranes in 2-methylfuran at room temperature. These impermeable membranes were tested for their resistances and exhibited Young's modulus values varying from 80 to 180 mPa. According to the surface wetting characterization, the surfaces of these membranes were clearly hydrophilic proven by values evolving between 63 and 76°. Less hydrophobic than pristine PCL and impermeable to water solution, the properties of our PCL composite membranes convince us to perform biocompatibility tests with cell viability found in a range of value between 71 and 87 % and a cell adhesion control dependent on the type of oligomer filler

    0

    full texts

    11,652

    metadata records
    Updated in last 30 days.
    HAL Portal UTC Université de Technologie de Compiègne
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇