ARPHA Preprints
Not a member yet
    49208 research outputs found

    ChatGPT as a Semantic Engineering Assistant: Lessons from Ontology Design in the Agricultural Biodiversity Domain

    No full text
    Modeling species names in biodiversity ontologies is particularly difficult in multilingual contexts, where semantic conflation often occurs. A good example is the common name "pimenta." In Brazilian Portuguese, experts usually refer to Capsicum spp. (chili peppers), while its direct translation “pepper” in English often denotes Piper nigrum (black pepper) (Soares et al. 2025a). In Brazilian markets, however, Piper nigrum is more accurately associated with “pimenta-do-reino" (“pimenta-negra”). This issue was observed on Wikipedia, when translating the Portuguese page for “pimenta” into English, the entry switches from Capsicum spp. to black pepper (Piper nigrum), showing how easily semantic drift can appear in multilingual data modeling. The correct association between common names used in the agricultural market with species would be a way to avoid the misunderstanding of these cultural differences. However, another challenge in vocabulary management emerges, which is how to manage species names in ontologies to keep them updated as the taxonomy itself updates. Some agriculturally controlled vocabularies, such as Agrotermos (Telles et al. 2024) lack automated mechanisms for updating taxonomic classifications. For example, Prochilodus cearensis, Prochilodus scrofa, and Prochilodus margravii are all listed in Agrotermos as preferred terms, i.e., the authorized, standard term selected to represent a concept in a controlled vocabulary, while according to the Global Biodiversity Information Facility (GBIF) Backbone Taxonomy (GBIF Secretariat 2023) these names are synonyms, as shown in Table 1. When developing the Agricultural Product Types Ontology (APTO), which was designed to represent products traded in Brazilian agricultural markets based on Agrotermos and AGROVOC, we proposed two approaches using generative AI, specifically OpenAI's ChatGPT-4, as a semantic engineering assistant to automate the inclusion of scientific names in the ontology:Prompt-based queries with a plugin accessing the GBIF APIA ChatGPT-generated Python script that converted GBIF taxonomy data into Web Ontology Language (OWL) formatThese AI-supported methods automated the construction of APTO’s “Organism” module, integrating taxonomic hierarchies and managing synonyms. ChatGPT effectively identified synonymy (e.g., see Table 1) and reduced manual labor in ontology development. The first approach is no longer reproducible since OpenAI has replaced plugins by GPTs. As such, we are currently developing a GPT named Taxonomy OWLizer 2.0*1, which is an evolution of the first approach described in that paper. Concerns about scalability, reproducibility, and hallucinations (false, made-up information) remain, highlighting the need for expert oversight throughout the process. When ChatGPT was used without API access, hallucinations appeared more frequently. For instance, when asked to check a list of plant species names for typos, it incorrectly suggested that Euterpe edulis was a synonym of Euterpe oleracea, even though both are recognized as distinct species in widely used catalogues such as the GBIF Backbone Taxonomy (Soares et al. 2025a).This case study demonstrates that generative AI can support but not yet replace human-led ontology development. It also emphasizes AI’s potential contribution to biodiversity informatics, particularly for managing evolving and multilingual vocabularies. All tools and source code related to our work are archived on Zenodo (Soares et al. 2025b). Detailed protocols are provided in Soares et al. 2025a

    A comparative analysis of the management of thoracolumbar burst fractures: short-segment posterior stabilization versus long-segment posterior instrumentation

    No full text
    Abstract Introduction: The lumbar and thoracic spine sustain 90% of all spinal fractures. There is ongoing discussion regarding the most effective treatment for thoracolumbar burst fractures. Clinical results of long-segment and short-segment posterior stabilization are contrasted in this research. Materials and methods: There were thirty patients in total; fifteen underwent short-segment stabilization, and fifteen underwent long-segment stabilization. All patients were assessed and treated according to the advanced trauma life support (ATLS) protocol. Mobilization was started as tolerated. Postoperative X-rays were taken, and follow-up occurred monthly for six months, then every two months thereafter up to one year. Functional outcomes were determined by employing the ASIA impairment scale and VAS scores. Results: Long-segment group operating times were 136.1±11.31 and 79.4±11.7 minutes, respectively, substantially longer than those of short-segment groups (p<0.05). In comparison to short-segment groups, long-segment groups experienced a statistically significant greater blood loss of 1263.3±151.74 and 876.7±189.8 ml, respectively (p<0.05). The ASIA impairment scale measurement, change in Beck’s index, and kyphotic angle were statistically insignificant among both groups. In the long-segment group, most patients had a Denis pain scale score of P2 (40%) and a Denis work scale score of W3 (46.7%), but in the short-segment group, most patients had scores of P3 (60%) and W4 (46.7%). We encountered no major complications. Conclusion: While long-segment fixation provides greater stability, short-segment fixation results in less operative time and blood loss without compromising clinical outcomes. Longer follow-up studies with larger sample sizes are recommended

    Our experience in the management of Fournier’s gangrene – a single-center retrospective study

    No full text
    Abstract Introduction: Fournier’s gangrene (FG) is a rare and potentially life-threatening infection that leads to necrosis of soft tissue. This condition constitutes a medical emergency, necessitating prompt surgical intervention to mitigate the potential consequences. Aim: To analyze the demographic and clinical characteristics of a small cohort of patients with FG. Patients and methods: The present retrospective study included 31 patients with Fournier’s gangrene who were hospitalized in the Department of General Surgery from January 2020 to December 2023. A comprehensive examination of the patients’ demographic characteristics, comorbidities, presence of diabetes mellitus, microbial agents involved, and methods used for wound management was conducted. Results: The study found that men, particularly those over the age of 55, were more commonly affected than women. Escherichia coli was identified as the predominant microbial agent. The prevalence of diabetes mellitus was found to be higher among female patients. All patients received prompt surgery according to established protocols. Enzyme proteolysis was our method of choice for wound management. Ten patients underwent adjunctive surgery while seven patients had reconstructive procedures. The mortality rate registered was 25.8%. The mean length of hospital stay was 12.8 days. Conclusion: Fournier’s gangrene has a high mortality and complication rate despite the current treatment options. Wound management with enzyme proteolysis yielded promising results

    A study on the expression of EZH2, Bcl-2 and Ber-EP in BCC

    No full text
    Basal cell carcinoma is the most common malignant tumour in humans. In cases with indistinct morphology on H&E-stained slides, immunohistochemistry may help distinguish basal cell carcinoma from other similar-appearing lesions. Our study aimed to investigate the expression of a marker panel comprising EZH2, Bcl-2, and Ber-EP4 in morphologically diagnosed, CK20-verified cutaneous basal cell carcinomas.Materials and methods: A cross-sectional study of 50 histologically confirmed cases of basal cell carcinoma was conducted. Immunohistochemical staining was performed using the following markers: EZH2, Bcl-2, Ber-EP4, and CK20. Due to the lack of a standardised method for evaluating markers, we adopted and modified the staining index (SI), which semi-quantitatively combines staining intensity and the percentage of positive cells. The results were systematised and interpreted using IBM SPSS.Results: All 50 examined tumours tested negative for CK20 (100%), thereby excluding mimics. All 50 tumours stained positive for EZH2 and Bcl-2 (100%), and only one stained negative for Ber-EP4 (98% positive). We found no association between histological type and EZH2 (p = 0.376), Bcl-2 (p = 0.376), and Ber-EP4 (p = 0.318), respectively, or their co-expression (p = 0.258). High co-expression of two of the three markers was observed in 33 of the 50 examined cases (66%), and a low co-expression in 4 cases (8%).Conclusion: The marker panel demonstrates co-expression of the three markers in the context of negative CK20 in over 90% of the cases. In challenging cases, it is important to consider clinical, morphological, and immunohistochemical features together

    New records of two riparian species of Labiduridae (Insecta, Dermaptera) from Guizhou Province, China

    No full text
    We report new provincial records of two riparian earwigs, Forcipula decolyi de Bormans, 1900 and Nala lividipes (Dufour, 1820), from Fanjingshan National Nature Reserve, Tongren City, Guizhou Province, China. The specimens conform to previously published diagnostic characters of these species, notably forceps and genital morphology. These records extend the known distributions of F. decolyi and N. lividipes within China and underscore the incompleteness of current faunal inventories. We also provide distribution maps for the two species

    Relationship between sprint and vertical jump force-time metrics in elite male professional futsal players

    No full text
    While sprinting capabilities are crucial for success in futsal, there is a lack of research on how they relate to countermovement vertical jump (CMJ) force-time metrics during both the eccentric and concentric phases of the movement. Thus, the purpose of this study was to examine the relationship between short-distance sprint speed and acceleration capabilities over 5m and 10m sprint distances and CMJ performance within a group of professional athletes. Twenty-two male futsal players competing in the top-tier national league volunteered to participate in this study. Following completion of the warm-up protocol, athletes stepped onto a uni-axial force plate and performed two non-consecutive CMJs with no arm swing, followed by two 10m sprints. The body mass-dependent force-time metrics were analyzed in both absolute and relative terms, while sprint analysis included 5m and 10m sprint speed and average acceleration over 0-5m and 5-10m sprint distances. Pearson product-moment correlation coeffi  cients (r) were used to examine the strength of the relationship between performance parameters of interest (p<.05). The results revealed that 10m sprint speed was positively associated with CMJ concentric peak velocity (r=.455; p=.032) and jump height (r=.457; p=.033). Besides providing sports practitioners working with this specific group of athletes with referent values about sprint and jump performance characteristics, the findings of this study suggest that athletes capable of generating greater CMJ concentric peak velocity and jump height tend to attain greater 10m sprint velocities

    Untangling Attribution in Biodiversity Data Records

    No full text
    The exactness, fitness-for-purpose (FFP) and reliability of primary biodiversity data can be enhanced by additional data beyond the basic taxon-location-date triad (Hill et al. 2010). Often, the only available data are the labels in legacy specimen collections. The digitization process is most efficient if all available information can be collected at once in a single event of specimen handling, rather than in separate phases. There is, however, a compromise between producing a faster catalogue for immediate use and housekeeping, and an accurate, wider-FFP database where all data have been thoroughly checked.Recognizing potential sources of error at digitization time may help making choices. During a dataset integration procedure, the quality and reliability of the data capture was analyzed. The dataset consisted of transcribed label data of over 58K pinned insects of agriculturally-relevant groups in XXth-century collections at six institutions in Spain, that resulted in almost 6000 collector strings. But collector names could beunidentified;misread;ambiguous,duplicated under variants; ormisplaced or misattributed to/from another entity, e.g. a location.This resulted in a high entropy level where one collector could be databased in multiple ways, artificially inflating the corresponding catalogues. The entropy was much higher in collections where collectors contributed few specimens, which is the case for university-based collections.By using simple indexing and cross-referencing techniques, the roster of names was significantly reduced, but full disambiguation of collectors required mining ancillary sources and consulting with people with long-standing knowledge of the collections. Overall, 42% of collector names were in error, resulting in excess entropy. METHODSSpecimen data were collated from the TETTRIS INC-STEP Project of Spanish pollinators complemented by some agriculturally-relevant groups in six collections deposited at five academic and research institutions in Madrid, Barcelona, Valencia and Pamplona (see Suppl. material 1 for full details). MCNB, MNCN, MUVHN were chiefly historical and created by researchers over long careers, while MZNA-R, MZNA-Z, UCME came mainly from academic coursework activities.Our procedures agreed with the overall strategy of Groom et al. (2022), while devising a specific workflow. Disambiguation and attribution included a number of steps (Suppl. material 1). First, names were normalized to facilitate grouping and sorting (7 steps). Then, ambiguous names were (whenever possible) attributed to actual persons by internal checks (clustering of names, matching localities and dates), consultation with external references (e.g. student lists), or consultation with collection curators (5 steps). Names were finally given a four-level identity qualification resulting from the disambiguation exercise (see Suppl. material 1). For further analyses, “very low” and “low” levels were considered a poor attribution, while “medium” and “full” levels were deemed good attribution.RESULTSThe disambiguation exercise reduced collector names from 5939 to 3473 (a 41.5% decrease). The total entropy, measured as Shannon’s H’, decreased by 8% while Simpson’s dominance D increased by 29%, as expected (merging names created larger collections for some collectors). However, there were differences among collections. The three academic collections had both higher diversity and higher diversity reduction than the historical collections (Fig. 1). Historical collections tended to be much more concentrated (30 specimens per collector, 25% single-specimen collectors) than the academic collections (8 and 61%, respectively).These differences can be tracked to how the collections were formed. Historical collections tended to be well documented and created by researchers spanning longer careers, while academic collections were often linked to works created during coursework. Fig. 2 shows the career spans of collectors. The abrupt end of collection at the turn of the century for academic collections can be linked to the introduction of restrictive legislation in Spain (a mandatory reduction of course loads, and unaffordable permission requirements). Full disambiguation often requires manual, time-consuming verifications against a variety of external sources. Automated disambiguation procedures elevated good identification from an initial 21% to 47% in MZNA-R, but it was not possible to use manual checking against sources for this dataset. In contrast, the similar MZNA-Z collection could be checked against coursework lists, which resulted in a 80% good attribution rate (Fig. 3). Overall, 57% of the collector names across collections had a good attribution.CONCLUSIONDoing disambiguation exercises by people with contemporary knowledge of the collections appears to be critical to avoid losing much information from the collections. Curators have an invaluable knowledge of the collections, could locate relevant documents about them, and are the ones able to reduce collection entropy by at least 32%. Their early retirement would negate a significant way to disambiguate names that left little or no trace in the literature record

    Working across Scales: Shared Building-Blocks of a Sustainable Infrastructure Environment for Bio/Geodiversity Data

    No full text
    The bio- and geodiversity communities aim to build digital infrastructures capable of delivering high-quality, well-structured, and richly interlinked data at a scale that enables powerful analyses to support the conservation and management of Earth’s diversity. Informaticians, scientists, and resource managers face the challenge of globally accelerating and upscaling open data availability. Achieving this requires sustainable, resilient, and bi-directional socio-technical connections that promote interactions and feedback both ways across the data landscape. Such interactions will connect global and national repositories to local providers, data stewards, and data users in ways that are context-sensitive and adapted to partners' location, time and communities.Although communities ranging from individual local experts to the governing bodies of global platform consortia are highly motivated and productive, they struggle to manage the rapidly growing volume of data, keep pace with technical innovation, and adhere to shared standards and best practices. Persistent challenges include insufficient recognition and visibility for contributors; functional and operational limitations in infrastructure; and inadequate long-term funding. Although the importance of the life cycles of data, infrastructures, and services is well-recognized, insufficient expertise in assessing the value that they represent and generate, and integrating those benefits into global, whole-community finance strategies presents an additional challenge. However, methodologies developed by global initiatives such as the United Nations System of Environmental-Economic Accounting (SEEA; United Nations et al. 2025) and UN Biodiversity Finance (BIOFIN; Cruz-Trinidad et al. 2024) provide innovative examples for addressing these. Without mechanisms that return generated value to data providers, stewards, and infrastructure developers, the community cannot initiate the self-reinforcing cycle of social and technical capacity growth needed for global upscaling. Instead, the current resource-limited landscape threatens the robustness of community networks and the long-term sustainability of existing infrastructures.The insights from comprehensive community outreach and infrastructure reviews conducted by the U.S.-based BIOFAIR (Building an Integrated, Open, Findable, Accessible, Interoperable, and Reusable) Data Network form the foundation for a living roadmap designed to support effective solutions and sustainable models for global biodiversity data infrastructures (Kunkel et al. 2025). Informed by these findings, the International Partners for the Digital Extended Specimen (IPDES) network are discussing shared and unique strengths and gaps across partners at different levels in this ecosystem. Adding theory of change (Rice et al. 2020), as well as socioeconomic and biodiversity finance considerations to the BIOFAIR roadmap, members of IPDES developed an extended model of infrastructure sustainability and a project proposal. Key elements include:involving economists in the process of integrating data, infrastructure, and services in information-economy value chains into natural capital accounting frameworks (e.g., SEEA), and developing comprehensive finance plans for bio/geodiversity data (e.g. adapting the methodology of BIOFIN);building governance and organizational frameworks that support transparent, community-centered decision-making;engaging the global community through coordinated participation and shared infrastructures;conducting pilot implementations that apply accounting methods to key use cases; andperforming risk assessments of, and establishing safeguards for, the arising feedback loops in a socioeconomic system of returning value to data providers, stewards and users.Together, these actions outline a path toward community-wide finance strategies that can support sustainable, scalable, and globally coherent biodiversity data infrastructures

    MfN DataHub – a Centralized Service for Automated Biodiversity Data Integration at the Museum for Natural History Berlin

    No full text
    The Museum für Naturkunde Berlin (MfN) DataHub*1 is an open-source web service and workflow engine developed to execute automated data-integration and migration workflows in continuous and parallel scenarios. Data migration and integration remain major challenges in publishing biodiversity data that follow international standards. To overcome these, a centralized service was created to coordinate and concentrate the computational power required for large-scale data transformation. Deployed at the Museum für Naturkunde Berlin, it now serves as the core of the institution’s scientific data-management infrastructure.This service supports digitization pipelines by iterating through datasets and files to integrate them into designated target systems. Following the ETL (Extract, Transform, Load) (Moreau 2015) principle, it extracts data from internal databases and shared storages via secure protocols such as SMB*2 and SFTP*3, transforms and validates them, and loads the compliant outputs through target-system API endpoints. The overall architecture and data flow are shown in Fig. 1.Operations are controlled through a web dashboard that allows execution and monitoring of pipelines either manually, automatically, or with AI-agent assistance. Implemented using the Django Web Framework, the service runs modular Python scripts and exposes all functions through RESTful APIs. A MCP*4 server provides an AI-readable interface, enabling both human- and machine-driven operations.Connected to the museum’s storage systems, the DataHub validates, enriches, and transforms data with a dedicated validator ensuring each record meets predefined structural and semantic rules. A persistent integration pipeline imports datasets into the museum’s Specify collection management system and its digital catalog, making data accessible for research and public use. It also prepares standardized packages for external partners such as GBIF (Global Biodiversity Information Facility), following the Darwin Core (Wieczorek et al. 2012) format and specific project requirements. Integration with field-data applications like ODK (Open Data Kit) ensures mobile data collection can enter the same pipeline. All operational steps, including validation, transformation, and API transactions are fully logged for transparency and reproducibility of the operation.A key innovation is the AI-integration layer, linking through the MCP*4 server to an AI agent built with LangChain library and Qwen3 LLM*6 model, executed locally via Ollama platform. This component assists with workflow orchestration by optimizing tasks order, resource allocation, error recovery and live reports thereby reducing manual supervision.By combining ETL*7 pipelines with AI-assisted orchestration, the DataHub provides a flexible, scalable engine adaptable to different collection domains. Its modular and open-source design promotes reproducibility and extension to new systems. The centralized architecture enhances data quality and FAIR (Findability, Accessibility, Interoperability, and Reusability) Wilkinson et al. 2016 compliance, offering a practical and scalable solution that can be deployed in natural-history institutions aiming to modernize their digital-collection infrastructures and ensure the continuous availability of reliable, high-quality biodiversity data. The open-source code and technical documentation of the MfN DataHub are available on the museum's GitHub repository.*

    First records of Opius and Apodesmia (Hymenoptera, Braconidae, Opiinae) from South Korea, with descriptions of newly-recorded species

    No full text
    The subfamily Opiinae comprises more than 2,000 valid species worldwide. Members of this subfamily are koinobiont endoparasitoids, with parasitism generally culminating in the eventual death of the host. Several species of Opiinae have been utilised for biological control of agricultural pests. The genus Opius is the largest genus within Opiinae, with more than 1,000 valid species worldwide. It is divided into several subgenera, classification of which remains under active discussion. The genus Apodesmia was formerly regarded as a subgenus of Opius, but was elevated to genus level, based on differences in the form of the occipital carina.Opius youi Li & van Achterberg, 2013 is recorded for the first time from South Korea, representing the first record of the species outside China. Apodesmia incisula Fischer, 1963 is also newly recorded from South Korea, constituting the first record of the species outside Europe, where it was previously known from Germany and the Netherlands. For each species, detailed morphological descriptions are provided, accompanied by diagnostic characters illustrated with photographs of the relevant body structures. The barcode region of mitochondrial cytochrome c oxidase I (COI) was also analysed for the species

    0

    full texts

    49,208

    metadata records
    Updated in last 30 days.
    ARPHA Preprints
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇