NII Repository (National Institute of Informatics)
Not a member yet
2035 research outputs found
Sort by
Overview of the NTCIR-16 Data Search 2 Task
NTCIR-16 Data Search 2 is the second round of the Data Search task at NTCIR. The first round of Data Search (NTCIR-15 Data Search) focused on the retrieval of a statistical data collection. This round also addressed the problem of ad-hoc data retrieval (IR subtask) and planned the other subtasks including question answering (QA) subtask and user interface (UI) subtask. This paper introduces the task definition, test collection, and evaluation methodology of the subtasks of NTCIR-16 Data Search 2. The IR subtask attracted seven research groups, from which we received 25 English runs and 23 Japanese runs. The evaluation results of these runs are presented and discussed in this paper.conference pape
TUA1 at the NTCIR-16 DialEval-2 Task
In this paper, we report the work of TUA1 team in the dialogue evaluation (DialEval-2) task of NTCIR-16, which consists of two subtasks: the Dialogue Quality (DQ) subtask and the Nugget Detection (ND) subtask. Our proposed method consists of two parts: a feature extractor and a feedforward network. The feature extractor employs pre-trained Transformer networks to extract the hidden representations of the dialogue utterances and employs a Latent Dirichlet Allocation (LDA) method to extract the topic information of these utterances. The feedforward network then concatenates the hidden representations and the topics extracted by the feature extractor, compresses them through several feedforward layers into a desired dimension, and finally predicts the quality scores and nugget types of the dialogues. Since the DialEval-2 task dataset was composed of the one-to-one translated Chinese and English dialogues, we employ the pre-trained Transformer networks for Chinese and English, respectively, to extract the hidden representations of the dialogues. This makes it possible to process the sub-tasks on the Chinese or English datasets simultaneously. We train the neural network models based on the mean squared error for both dialogue quality prediction and nugget detection subtasks. Our proposed method reaches the best scores for RSNOD and NMD metrics in both Chinese and English dialogue quality subtasks among all participants. The results indicate that the proposed method is promising in learning a dialogue quality prediction system for generating very close predictions to the human annotators.conference pape
NKUST at the NTCIR-16 DialEval-2 Task
It is important to evaluate the quality of dialogues generated by chatbots. Most previous automatic evaluation methods have been based on models (e.g., LSTM ) that are capable of processing time series. This study presents three models for dialogue quality and two nugget detection subtasks, respectively. Specifically, the first model uses a Pegasus model that can transform dialogues into short summaries; the second model uses a Bi-LSTM that merely adjusts the internal model structure; and the third model is a multi-agent model simulating situations in which multiple annotators generate different evaluation results for the same text. The experimental results show that certain opinions may need to be corroborated by more refined experimental design and the testing of more model parameters before they are applicable to this issue.conference pape
THUIR-LL at the NTCIR-16 Lifelog-4 Task: Enhanced Interactive Lifelog Search Engine
With the development of digital information storage technology and portable sensing devices, users are gradually accustomed to recording their personal life~(i.e., lifelog) in various digital ways. Therefore, the retrieval of lifelogging has become a new and essential research topic in related fields. Unlike traditional search engines, in lifelog, text and other data automatically recorded in real-time by sensors bring challenges to data arrangement and search. As the dataset is highly personalized, interactions and feedback from users should also be considered in the search engine. This paper describes our interactive approach for the NTCIR-16 Lifelog-4 Task. The task is to search relevant lifelog images from the users' daily lifelog given an event topic. A significant challenge is how to bridge the semantic gap between lifelog images and event-level topics. We propose a framework to address this problem with a multi-functional and flexible feedback mechanism and result presentation for interaction in a search engine. Besides, we propose a query text parsing procedure that parses the long query text into keywords and fills the fields automatically. We analyzed the interactive lifelog search engine with 12 topics constructed by ourselves according to LSC'18 development topics. Finally, we achieved an official result of 741 at the NTCIR-16 Lifelog-4 task in terms of RelRet score over 48 topics.conference pape
Overview of the NTCIR-16 QA Lab-PoliInfo-3 Task
The goal of the NTCIR-16 QA Lab-PoliInfo-3 task is to develop real-world complex question answering (QA) techniques using Japanese political information such as local assembly minutes and newsletters. QA Lab-PoliInfo-3 consists of four subtasks: QA Alignment, Question Answering, Fact Verification, and Budget Argument Mining. In this paper, we present the data used and the results of the formal run.conference pape
NICTmed at the NCTIR-16 Real-MedNLP Task
This paper describes NICTmed team's challenge to Subtask1-CR-EN, Subtask1-CR-JA, Subtask3-CR-EN, and Subtask3-CR-JA in NTCIR-16 Real-MedNLP. In Real-MedNLP, approximately 100 annotated real clinical reports in both English and Japanese are given to participants. Subtask1-CR-EN/JA and Subtask3-CR-EN/JA are both based on Case Reports, Subtask1 is few-resource Named Entity Recognition (NER) and Subtask3 is information extraction for adverse drug event (ADE). We used multilingual BERT (mBERT) and XLM-RoBERTa (XLM-R) to compare how effective the multilingual pre-trained models work in specific domain downstream tasks English and Japanese. Our experiment used no external data to adjust conditions of English and Japanese experiments. As a result, we confirmed that the multilingual pre-training models provide similar level of accuracy in Japanese as in English, and got rank 3 in Entity F1 of all target entities for Subtask1-CR-JA, top rank in Report-level precision and F1 for Subtask3-CR-JA.conference pape
第32å ããããã®åŠè¡æ å ±ã·ã¹ãã æ§ç¯æ€èšå§å¡äŒ è°äºèŠæš
äŒè°åïŒç¬¬32å ããããã®åŠè¡æ
å ±ã·ã¹ãã æ§ç¯æ€èšå§å¡äŒ
éå¬å ŽæïŒãªã³ã©ã€ã³äŒè°
æ¥æïŒ2022幎1æ26æ¥ïŒæ°ŽïŒ15ïŒ00ïœ17ïŒ00conference outpu
SPARC Japan ã»ãããŒ2021 ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã åŠè¡æ å ±ã€ã³ãã©ãå®çŸããç ç©¶ããŒã¿ã®ç®¡çãšåŸªç°ãçºè¡šè³æ
SPARC Japan ã»ãããŒ2021ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã
éå¬å ŽæïŒãªã³ã©ã€ã³éå¬
æ¥æïŒ2022幎2æ22æ¥ïŒç«ïŒ13:00-16:55conference objec
SPARC Japan ã»ãããŒ2021 ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã ç 究ãå·¡ãæ¿çååããèŠãç ç©¶ããŒã¿ããªã·ãŒã®åœ¹å²ãããã¥ã¡ã³ã
SPARC Japan ã»ãããŒ2021ãç ç©¶ããŒã¿ããªã·ãŒãç®æããã®ãšã¯ã
éå¬å ŽæïŒãªã³ã©ã€ã³éå¬
æ¥æïŒ2022幎2æ22æ¥ïŒç«ïŒ13:00-16:55conference objec