Nara Institute of Science and Technology

naistar NAIST Academic Repository
Not a member yet
    13197 research outputs found

    MultiMSD: A Corpus for Multilingual Medical Text Simplification from Online Medical References

    No full text
    We release a parallel corpus for medical text simplification, which paraphrases medical terms into expressions easily understood by patients. Medical texts written by medical practitioners contain a lot of technical terms, and patients who are non-experts are often unable to use the information effectively. Therefore, there is a strong social demand for medical text simplification that paraphrases input sentences without using medical terms. However, this task has not been sufficiently studied in non-English languages. We therefore developed parallel corpora for medical text simplification in nine languages: German, English, Spanish, French, Italian, Japanese, Portuguese, Russian, and Chinese, each with 10,000 sentence pairs, by automatic sentence alignment to online medical references for professionals and consumers. We also propose a method for training text simplification models to actively paraphrase complex expressions, including medical terms. Experimental results show that the proposed method improves the performance of medical text simplification. In addition, we confirmed that training with a multilingual dataset is more effective than training with a monolingual dataset.conference pape

    Exploring LLM Annotation for Adaptation of Clinical Information Extraction Models under Data-sharing Restrictions

    No full text
    In-hospital text data contains valuable clinical information, yet deploying fine-tuned small language models (SLMs) for information extraction remains challenging due to differences in formatting and vocabulary across institutions. Since access to the original in-hospital data (source domain) is often restricted, annotated data from the target hospital (target domain) is crucial for domain adaptation. However, clinical annotation is notoriously expensive and time-consuming, as it demands clinical and linguistic expertise. To address this issue, we leverage large language models (LLMs) to annotate the target domain data for the adaptation. We conduct experiments on four clinical information extraction tasks, including eight target domain data. Experimental results show that LLM-annotated data consistently enhances SLM performance and, with a larger number of annotated data, outperforms manual annotation in three out of four tasks.conference pape

    A Study on Surface Passivation of Perovskite Solar Cells with First-Principles Insights

    Get PDF
    奈良先端科学技術大学院大学博士(工学)doctoral thesi

    Understanding the characteristics of LLMs in detecting textually dissimilar duplicate bug

    Get PDF
    奈良先端科学技術大学院大学修士(工学)master thesi

    Round Outcome Prediction in VALORANT Using Tactical Features from Video Analysis

    No full text
    Recently, research on predicting match outcomes in esports has been actively conducted, but much of it is based on match log data and statistical information. This research targets the FPS game VALORANT, which requires complex strategies, and aims to build a round outcome prediction model by analyzing minimap information in match footage. Specifically, based on the video recognition model TimeSformer, we attempt to improve prediction accuracy by incorporating detailed tactical features extracted from minimap information, such as character position information and other in-game events. This paper reports preliminary results showing that a model trained on a dataset augmented with such tactical event labels achieved approximately 81% prediction accuracy, especially from the middle phases of a round onward, significantly outperforming a model trained on a dataset with the minimap information itself. This suggests that leveraging tactical features from match footage is highly effective for predicting round outcomes in VALORANT.conference pape

    Genomic insights into the carbohydrate-active enzymes of Trichoderma asperellum USM SD4 for potential efficient biomass degradation

    No full text
    Aims: This study aimed to analyze the Trichoderma asperellum USM SD4 genome to identify carbohydrate-active enzymes (CAZymes) involved in lignocellulosic biomass degradation. These enzymes, particularly glycoside hydrolases (GHs), are crucial for converting biomass into simple sugars. Methodology and results: The genomic DNA of T. asperellum USM SD4 was successfully sequenced, resulting in an assembled genome size of 35.83 Mb containing 10,208 protein-coding sequences. The CAZyme analysis revealed a total of 482 annotated polypeptides, with the highest 54% classified as GHs. The 180 proteins were lignocellulosic CAZymes, with 44% targeting hemicellulose hydrolysis, 42% for cellulose degradation, and 14% for lignin degradation. The present study identifies 14 xylanases and 23 cellulases protein-coding genes in the T. asperellum USM SD4 genome. The secretome analysis revealed 770 proteins, with 182 proteins associated with CAZymes. Among these, one putative xylanase and one putative cellulase were selected for the in vitro characterization. Conclusion, significance and impact of study: This study highlighted the potential of T. asperellum USM SD4 for industrial applications that require efficient biomass degradation.journal articl

    Scalable Anonymous Authentication Scheme Based on Zero-Knowledge Set-Membership Proof

    No full text
    In this article, we propose zero-knowledge named proof, a stateless replay attack prevention strategy that ensures the user’s anonymity against malicious administrators. We begin with adopting the zero-knowledge set-membership proof into an authentication setting in which users would delegate their requests to an agent that obstructs the user’s identity from the administrator. This anonymous agent carries the guarantee of authenticity, which the administrator through the set-membership proof can confirm. Next, we prevent replay attacks from other parties by binding the agent’s identity to the authentication proof verifiable by the administrators. By leveraging these properties, a scalable blockchain-based authentication scheme is then built. We quantitatively evaluate the security and measure the time and monetary cost of our scheme under both ideal and realistic environments. On top of it, we provide a third-party authorization scheme derived from our authentication framework to demonstrate its real-world applicability.journal articl

    Utilization of Healthcare Information Extracted from Medical Documents and Social Media Texts

    Get PDF
    奈良先端科学技術大学院大学博士(工学)doctoral thesi

    Tris[N,N-bis(trimethylsilyl)amide] lanthanum (Ⅲ) as a highly active catalyst for polymerization of six-membered cyclic carbonates: achievements of high molecular weight and property evaluations

    Get PDF
    奈良先端科学技術大学院大学博士(工学)doctoral thesi

    Quantitative Evaluation of Subjective Experience in Virtual Reality

    Get PDF
    奈良先端科学技術大学院大学博士(理学)doctoral thesi

    11,774

    full texts

    13,197

    metadata records
    Updated in last 30 days.
    naistar NAIST Academic Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇