Portail HAL des publications du LIRMM
Not a member yet
13279 research outputs found
Sort by
Jargon : Une suite de modèles de langues et de référentiels d'évaluation pour les domaines spécialisés du français
JEP-TALN-RECITAL : 35esJournées d'Études sur la Parole (JEP 2024) 31e Conférence sur le Traitement Automatique des Langues Naturelles (TALN 2024) 26e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RECITAL 2024)National audiencePretrained Masked Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important applications : the domains of transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations : continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets.Les modèles de langue préentraînés (PLM) constituent aujourd’hui de facto l’épine dorsale de la plupart des systèmes de traitement automatique des langues. Dans cet article, nous présentons Jargon, une famille de PLMs pour des domaines spécialisés du français, en nous focalisant sur trois domaines : la parole transcrite, le domaine clinique / biomédical, et le domaine juridique. Nous utilisons une architecture de transformeur basée sur des méthodes computationnellement efficaces(LinFormer) puisque ces domaines impliquent souvent le traitement de longs documents. Nous évaluons et comparons nos modèles à des modèles de l’état de l’art sur un ensemble varié de tâches et de corpus d’évaluation, dont certains sont introduits dans notre article. Nous rassemblons les jeux de données dans un nouveau référentiel d’évaluation en langue française pour ces trois domaines. Nous comparons également diverses configurations d’entraînement : préentraînement prolongé en apprentissage autosupervisé sur les données spécialisées, préentraînement à partir de zéro, ainsi que préentraînement mono et multi-domaines. Nos expérimentations approfondies dans des domaines spécialisés montrent qu’il est possible d’atteindre des performances compétitives en aval, même lors d’un préentraînement avec le mécanisme d’attention approximatif de LinFormer. Pour une reproductibilité totale, nous publions les modèles et les données de préentraînement, ainsi que les corpus utilisés
Database Repairing with Soft Functional Dependencies
International audienceA common interpretation of soft constraints penalizes the database for every violation of every constraint, where the penalty is the cost (weight) of the constraint. A computational challenge is that of finding an optimal subset: a collection of database tuples that minimizes the total penalty when each tuple has a cost of being excluded. When the constraints are strict (i.e., have an infinite cost), this subset is a “cardinality repair” of an inconsistent database; in soft interpretations, this subset corresponds to a “most probable world” of a probabilistic database, a “most likely intention” of a probabilistic unclean database, and so on. Within the class of functional dependencies, the complexity of finding a cardinality repair is thoroughly understood. Yet, very little is known about the complexity of finding an optimal subset for the more general soft semantics. The work described in this manuscript makes significant progress in that direction. In addition to general insights about the hardness and approximability of the problem, we present algorithms for two special cases (and some generalizations thereof): a single functional dependency, and a bipartite matching. The latter is the problem of finding an optimal “almost matching” of a bipartite graph where a penalty is paid for every lost edge and every violation of monogamy. For these special cases, we also investigate the complexity of additional computational tasks that arise when the soft constraints are used as a means to represent a probabilistic database in the case of a probabilistic unclean database
Herbivorous fish feeding dynamics and energy expenditure on a coral reef: Insights from stereo‐video and AI‐driven 3D tracking
International audienceUnveiling the intricate relationships between animal movement ecology, feeding behavior, and internal energy budgeting is crucial for a comprehensive understanding of ecosystem functioning, especially on coral reefs under significant anthropogenic stress. Here, herbivorous fishes play a vital role as mediators between algae growth and coral recruitment. Our research examines the feeding preferences, bite rates, inter‐bite distances, and foraging energy expenditure of the Brown surgeonfish ( Acanthurus nigrofuscus ) and the Yellowtail tang ( Zebrasoma xanthurum ) within the fish community on a Red Sea coral reef. To this end, we used advanced methods such as remote underwater stereo‐video, AI‐driven object recognition, species classification, and 3D tracking. Despite their comparatively low biomass, the two surgeonfish species significantly influence grazing pressure on the studied coral reef. A. nigrofuscus exhibits specialized feeding preferences and Z. xanthurum a more generalist approach, highlighting niche differentiation and their importance in maintaining reef ecosystem balance. Despite these differences in their foraging strategies, on a population level, both species achieve a similar level of energy efficiency. This study highlights the transformative potential of cutting‐edge technologies in revealing the functional feeding traits and energy utilization of keystone species. It facilitates the detailed mapping of energy seascapes, guiding targeted conservation efforts to enhance ecosystem health and biodiversity
peerannot: A framework for label aggregation in crowdsourced datasets
National audienceThis work presents peerannot, an image data classification library of image data whose labels are generated by crowdsourcing. It is written in Python and allows a comparison of aggregation classification methods with other reference libraries
Modeling compliant bistable mechanisms: An energy method based on the high-order smooth curvature model
International audienceThis paper applies an energy method based on the high-order smooth curvature model to address the challenges of kinetostatically modeling the complicated nonlinear post-buckling behavior of inclined compliant beams. The pro- posed energy method is grounded in the principle of minimum strain energy, implying that the total strain energy is minimized at the equilibrium configuration. In this work, the high-order smooth curvature model is adopted to accu- rately model the bending strain energy of large deformation beams. Subsequently, the Lagrange multiplier method is employed to ascertain the minimum of strain energy while concurrently determining the corresponding tip loads. Additionally, the deformation shape and maximum stress are determined via the smooth curvature model. The pro- posed method is introduced for the first time in modeling compliant bistable mechanisms, and it is proven that the method can be used for modeling compliant beams with inclined angles ranging from 0 to 90 degrees. Following the modeling, finite element analysis and experimental tests are conducted to verify the accuracy of the proposed energy method. A comprehensive comparative analysis between proposed method and existing methods is conducted. The comparison results prove that the model is more computationally efficient without compromising modeling accuracy The proposed modeling method can not only be used for modeling compliant bistable mechanisms but also has ex- tendable applications in modeling initially-curved compliant beams, contact-aid design problems, and distributed load problems
M4.3 - Specification of semantic artefact description
Semantic artefacts (SA) are key for the description of data and for making data FAIR (findable, accessible, interoperable and reusable) [1]. SA is a broader term to include ontologies, terminologies, taxonomies, thesauri, vocabularies, metadata schemas and semantic standards. Describing SAs is fundamental to make them FAIR themselves. The Metadata for Ontology Description and Publication (MOD) was developed to provide the vocabulary required to describe ontologies, and Semantic Artefacts in general. The Data Catalogue Vocabulary (DCAT) [2] was designed to describe datasets and resources that can be catalogued. This milestone presents a DCAT-based standard for description of Semantic Artefacts and their catalogues, building on the MOD vocabulary as well as recommendations and outcomes from the FAIRsFAIR project and the Research Data Alliance Vocabulary and Semantic Services Interest Group (RDA VSSIG). This milestone also makes a distinction between the MOD specification and a series of mappings with other vocabularies, presented in a machine-actionable way. Last but not least, we describe MOD profiles and how to formalise them in a machine-actionable and composable way. The next step (deliverable D4.3) related to this milestone will be to specify a common Application Programming Interface (API) for interoperability of SA catalogues in the European Open Science Cloud (EOSC) ecosystem and beyond, building from MOD descriptions of SAs. The API for SA-catalogues will enable interoperability and unified access to their content, enabling seamless querying and use by stakeholders independent of domain. The API will be adopted by FAIR-IMPACT T4.2’s use case SA-catalogues and it will be publicly available for other catalogues to deploy; via the API, other registries could consume content from multiple SA-catalogues. The implementation of this API will be the topic of an upcoming FAIR-IMPACT Open Call
Material Scrunching Enables Working Channels in Miniaturized Vine-Inspired Robots
International audienceA new subclass of soft robot, known as tip-extending or "vine" robots, consists of long inflatable devices that move through the environment by extending from the tip. A key requirement for many applications of these robots is a working channel-a hollow tube through the core of the robot for passing tools, sensors, fluids, etc. While working channels have been proposed in a few vine robots, it remains an open challenge to create miniaturized vine robots (diameter < 1 cm) with working channels that enable continuous access through the core. In this paper, we analyze the growth models of current vine robot designs and show that the working channel greatly increases required pressure to grow at small scales due to internal friction. Based on this insight, we propose the concept of storing scrunched material at the tip of the vine robot to circumvent this frictional force. We validate our models and demonstrate this concept via prototypes down to diameters of 2.3 mm. Overall, this work enables the creation of miniaturized vine robots with working channels, which significantly enhances their practicality and potential for impact in applications such as minimally invasive surgery
Normalisation automatique de variables issues de bases de données en agroécologie
International audienceThe objective of this work is to propose and evaluate methods of matching between sourcevariables and candidate variables from the agroecology domain (in English). The aim of ourapproach is to support the experts to standardize the databases, as well as to link the sourcevariables to the candidate variables of the dictionaries of the models used in agroecology (i.e.AEGIS).L'objectif de ces travaux est de proposer et évaluer des méthodes de mise en correspondance entre des variables sources et des variables candidates du domaine de l'agroécologie (en anglais). Le but de notre démarche est d'aider l'expert à normaliser les bases de données, ainsi qu'à lier les variables sources aux variables candidates des dictionnaires des modèles utilisés en agroécologie (AEGIS)
Zonal statistics datasets of climate indicators for Brazilian municipalities
International audienceClimate trends and weather indicators are used in several research fields due to their importance in statistical modeling, frequently used as covariates. Usually, climate indicators are available as grid files with different spatial and time resolutions. The availability of a time series of climate indicators compatible with administrative boundaries is scattered in Brazil, not fully available for several years, and produced with diverse methodologies. In this paper, we propose time series of climate indicators for the Brazilian municipalities produced using zonal statistics derived from the ERA5-Land reanalysis indicators. As a result, we present datasets with zonal statistics of climate indicators with daily data, covering the period from 1950 to 2022