NII Repository (National Institute of Informatics)
Not a member yet
2035 research outputs found
Sort by
Real-MedNLP: Overview of REAL document-based MEDical Natural Language Processing Task
A standard dataset collection is essential for the development of information science. Particularly in the medical field, in which privacy protection is a critical issue, the importance of the dataset is significant. To discuss the validness of various methods, we build the clinical text dataset, Real-MedNLP, for multiple medical tasks. The goal of Real-MedNLP is threefold: (1) Real datasets: Previous medical shared tasks, MedNLP, MedNLP2, and MedNLPDoc, were based on the pseudo dataset, which was built from medical textbooks or dummy clinical texts. This task prepares real radiology and case reports. (2) Bilingual capability: Both English and Japanese data are handled. (3) Practicality: Both fundamental (named entity recognition) and applied practical tasks are handled. This study introduces the task setting of Real-MedNLP and submitted systems. The methods mostly share the common paradigm, which is based on a fundamental language model, such as BERT, aiming to separate the resource problems. Based on their results, this study discusses the feasibility of their approaches to bring us the future direction of medical NLP. Note that the Real-MedNLP is a shared task that handles real Japanese medical texts.conference pape
AMI Team at the NTCIR-16 Real-MedNLP Task
The AMI team participated in subtasks 1 and 2 of the NTCIR-16 Real-MedNLP Task. In this paper, we report our systems employed for subtasks 1 and 2. In subtask 1, the organizer provides a small amount of training data. In recent years, the approach based on BERT has achieved excellent results for such a low-resource situation. We construct two systems based on the BERT model pretrained on biomedical documents (UTH-BERT). We construct the ensemble method with hidden vectors from multiple layers of UTH-BERT and the fine-tuning method with the CRF layer. In subtask 2, participants construct their methods based on the annotation guideline. We construct a multistage method to identify named entities. The system consists of three stages: a candidate extraction stage, an identification stage, and a tag correction stage. We discuss the effectiveness of our systems on the basis of our preliminary experiments and the results in the formal run.conference pape
FRDC at the NTCIR-16 Real-MedNLP Task
In this paper, we describe the approaches of FRDC team for the Real-MedNLP task. Specially, the FRDC team participated in three sub-tasks including Subtask1-CR-EN, Subtask3-CR-EN (ADE), and Subtask3-RR-EN (CI). The Real-MedNLP task aims to promote approaches for supporting real medical services under constrained training resources. We applied pre-trained language models (PTLMs) such as BERT and BioBERT to learn sentence and document representations. For each sub-task, we designed different networks based on PTLMs. Various effective methods such data augmentation were adopted in each sub-task. In the official run, we achieved the best score for the CI sub-task, and ranked 2nd in the ADE sub-task.conference pape
Approach for Named Entity Recognition and Case Identification Implemented by ZuKyo-JA Sub-team at the NTCIR-16 Real-MedNLP Task
We describe our submissions to NTCIR-16 Real-MedNLP shared task. This paper presents the approach of the ZuKyo-JA subteam for solving the Japanese part of Subtask1 and Subtask3 (Subtask1-CR-JA, Subtask1-RR-JA, Subtask3-RR-JA) based on a sliding-window approach using Japanese BERT pre-trained masked-language model. A lot of methods used for these subtasks share in common, regardless of the difference in the task. We also show a method to aggressively use medical knowledge for data labeling, data augmentation, and the same class identification for the subtask3-RR-JA.conference pape
Overview of the NTCIR-16 Session Search (SS) Task
This is an overview of the NTCIR-16 Session Search (SS) task. The task features the Fully Observed Session Search subtask (FOSS) and the Partially Observed Session Search subtask (POSS). This year, we received 28 runs from 6 teams in total. This paper will describe the task background, data, subtasks, evaluation measures, and the evaluation results, respectively.conference pape
ãªãŒãã³ãµã€ãšã³ã¹ã®ããã®ããŒã¿ç®¡çåºç€ãã³ããã㯠ïœåŠè¡ç ç©¶è ã®ããã® âå人æ å ±â ã® åæ±ãæ¹ã«ã€ããŠïœ
technical repor
ãã€ãªãªãœãŒã¹æ€çŽ¢ã·ã¹ãã ã®èª²é¡
2022幎6æ6æ¥ JAPAN OPEN SCIENCE SUMMIT 2022, ã»ãã·ã§ã³E1ãç ç©¶ããŒã¿ã®ãæ°ããèŠã€ãæ¹ããèãããconference objec
é»åãžã£ãŒãã«åé¡ã®åãæã®äžã€ãšããŠã®ã転æå¥çŽã
äŒè°åïŒç¬¬24å倧åŠå³æžé€šãšåœç«æ
å ±åŠç ç©¶æãšã®é£æºã»ååæšé²äŒè°
éå¬å ŽæïŒãªã³ã©ã€ã³
æ¥æïŒ2022幎6æ29æ¥ïŒæ°ŽïŒ14:00ïœ15:45conference objec
æµ·å€äž»èŠïŒåŠè¡æ å ±åºç€ã®ããã·ã¥ããŒãâœèŒåæ ãŒãªãŒãã³ãµã€ãšã³ã¹ææšã«æ³šâœ¬ããŠãŒ
äŒè°åïŒRAåè°äŒç¬¬8å幎次⌀äŒ
éå¬å°ïŒä»å°åžåœéå±ç€ºã»ã³ã¿ãŒ å±ç€ºæ£ 察é¢ïŒWebé
ä¿¡ãã€ããªãã
äŒæïŒ2022幎8æ30æ¥ - 2022幎8æ31æ¥
äž»å¬ïŒäž»å¬è
äžè¬ç€Ÿå£æ³äººãªãµãŒãã»ã¢ãããã¹ãã¬ãŒã·ã§ã³åè°äŒconference poste
SPARC Japan ã»ãããŒ2021 ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã ç·åèšè«ïŒç¬¬2éšïŒç ç©¶ããŒã¿ã«é¢ããåã¹ããŒã¯ãã«ããŒãšã®è°è«ãçºè¡šè³æ
SPARC Japan ã»ãããŒ2021ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã
éå¬å ŽæïŒãªã³ã©ã€ã³éå¬
æ¥æïŒ2022幎2æ22æ¥ïŒç«ïŒ13:00-16:55conference objec