NII Repository (National Institute of Informatics)
Not a member yet
2035 research outputs found
Sort by
Information Retrieval Evaluation as Search Simulation
Due to the empirical nature of the Information Retrieval (IR) task, experimental evaluation of IR methods and systems is essential. Historically, evaluation initiatives such as TREC, CLEF, and NTICR have made significant impacts on IR research and resulted in many test collections that can be reused by researchers to study a wide range of IR tasks in the future. However, despite its great success, the traditional Cranfield evaluation methodology using a test collection has significant limitations, especially for evaluating an interactive IR system, and it remains an open challenge how to evaluate interactive IR systems using reproducible experiments. In this talk, I will discuss how we can address this challenge by framing the problem of IR evaluation more generally as search simulation, i.e., having an IR system interact with simulated users and measuring the performance of the system based on its interaction with the simulated users. I will first present a general formal framework for evaluating IR systems based on search session simulation, discussing how the framework can not only cover the traditional Cranfield evaluation method as a special case but also reveal potential limitations of the traditional IR evaluation measures. I will then review the recent research progress in developing formal models for user simulation and evaluating user simulators. Finally, I will discuss how we may leverage the current IR test collections to support simulation-based evaluation by developing and deploying user simulators based on those existing collections. I will conclude the talk with a brief discussion of important future research directions in simulation-based IR evaluation.conference pape
NYUCIN at the NTCIR-16 Dataset Search 2 Task
In this paper, we describe the approach and results of the NYUCIN team in the NTCIR-16 conference. We participated in the Data Search 2 Task, which is a shared task on ad-hoc retrieval for governmental statistical data composed of multiple subtasks. We report our work on the two subtasks we participated in: the English IR Subtask and the UI Subtask. For the IR Subtask, we explored learning-to-rank approaches based on deep learning models. Given the limited training data available for this task, we employed a transfer learning method to train a deep neural network that learns how to match web tables and news articles using data available on the Web. The official evaluation shows that our approach attained the highest score among all submitted runs across all evaluation metrics. In particular, for the nDCG@5 measure, our score of 0.246 represents a 30% improvement compared to the second-best result in NTCIR-16 Data Search 2 Task. For the experimental UI Subtask, we performed a preliminary user study to evaluate the effectiveness of the user interface of Auctus, a dataset search engine developed by our team.conference pape
fuys at the NTCIR-16 QA Lab-PoliInfo-3 Budget Argument Mining Subtask
This paper reports on the achievements of Budget Argument Mining subtask of the NTCIR-16 QA Lab-PoliInfo-3 task of fuys team. We have assigned ArgumentClass and RelatedID in different ways. ArgumentClass was assigned using BERT. We also thought that the accuracy could be improved by adding the flag indicating whether a speaker is a legislator or not ("giin-flag"). RelatedID was assigned using keyword extraction with TFIDF. The results showed good results for ArgumentClass, but the improvement in accuracy of RelatedID could not be confirmed. Although there was a difference in the results for the "giin-flag", the difference was small, and no advantage was found with or without the "giin-flag".conference pape
Ibrk at the NTCIR-16 QA Lab-PoliInfo-3 Budget Argument Mining Subtask
In this study, we construct a system to predict argument labels for statements in meeting minutes by using sequence labeling methodology and validate the effectiveness of prediction performance for various input methods of utterance data to predict argument labels for money expressions effectively. To evaluate the validation of the system, we will use the Budget Argument Mining task data in NTCIR 16 QA Lab Poliinfo 3. We train an argument label prediction model on the training data that exists in the data, dividing it into two types: data for model training and data for model validation. As the prediction model, we use the Bidirectional LSTM CNNs CRF model to predict argument labels for each word in the input data and output a series of argument labels. In the experiment, we compare the prediction accuracy of models obtained by changing the data input method, such as the range of sentences containing money expressions. As a result of the experiment, we found that the prediction accuracy of argument labels was higher when each sentence was entered into the model rather than when all the statements of the assembly member were entered into the model. Furthermore, we found that the prediction accuracy of the argument labels can be improved by replacing the numbers in the money expression with special tokens.conference pape
SRCB at the NTCIR-16 Real-MedNLP Task
The SRCB participated in subtask1: Few-resource Named Entity Recognition (NER) and subtask3: Adverse Drug Event detection (ADE) in NTCIR-16 Real-MedNLP. This paper reports our approach to solve the problem and discusses the official results. For the Few-resource NER subtask, we developed NER systems based on pretraining model, span-based classification and prompt learning. In addition, data augmentation and model ensemble are used to further improve performance. For ADE subtask, we mainly adopted two methods: multi-class classification and prompt learning. We employed a two-stage training strategy to solve the long tail distribution problem and applied transfer learning to improve performance of model.conference pape
大学図書館の現状と課題
研修名:2022年度大学図書館職員短期研修
開催期間:2022年10月18日(火)~10月21日(金)
主催:東京大学附属図書館、京都大学附属図書館、国立情報学研究所othe