NII Repository (National Institute of Informatics)
Not a member yet
2035 research outputs found
Sort by
OUHCIR at the NTCIR-16 Data Search-2 Task
In this paper, we report our work and discuss the results for NTCIR- 16 DataSearch-2 IR subtask. NTCIR-16 Data Search-2 was organized to improve the present knowledge and promote the concepts of dataset search among IR researchers. In this particular subtask, we tried to perform ad-hoc retrieval for datasets based on given queries. While this task was available in English and Japanese, we decided to only compete for the English subtask. We sought to perform the ranking of datasets by using traditional BM25-based ranking functions and recent language models. During the evaluation experiments, we also explored the impact of metadata features on the performance of dataset ranking. Our best performing submission achieved a score of 0.153, 0.161 and 0.174 in nDCG@10, nERR@10, and Q-measure, respectively. In all the metrics, this run was ranked 13th among 25 submitted runs.conference pape
RSLDE at NTCIR-16 DialEval-2 Task
In this paper, we report the work of the RSLDE team at the dialogue evaluation (DialEval-2) task of NTCIR-16, including the Chinese and English dialogue quality (DQ) and nugget detection (ND) subtasks. We implemented two sentence-level baselines that fine-tune BERT and XLNet along with a linear layer for the ND subtask. In addition, we propose a model based on BERT to capture the structure and context information of a customer-helpdesk dialogue. This dialogue-level model modifies the input and embeddings of BERT. We add a Transformer encoder layer over this model as our third model for the ND subtask and first model for the DQ subtask. The second model for the DQ subtask is the same dialogue-level model but without the Transformer encoder layer. Our XLNet model generated the best run for both English and Chinese ND subtasks. Our dialogue-level model outperformed the baselines for the Chinese DQ subtask but not for the English DQ subtask.conference pape
Overview of the NTCIR-16 FinNum-3 Task: Investor’s and Manager’s Fine-grained Claim Detection
In the FinNum task series, we proposed numeral category understanding (FinNum-1) and numeral attachment (FinNum-2) tasks for better comprehending the numerals in financial narratives. In FinNum-3, we present a novel task, fine-grained claim detection, and further integrate the new task with the previous. There are two subtasks in FinNum-3: (1) Investor's claim detection (Chinese) and (2) Manager's claim detection (English). This round of FinNum attracted ten research teams, and received 10 submissions for Chinese subtask and 15 submissions for English subtask. This paper provides an overview of FinNum-3, including task definition, data annotation, and participants' results.conference pape
OUC at the NTCIR-16 QA Lab-PoliInfo-3 Budget Argument Mining
The OUC team participated in the Budget Argument Mining subtask of NTCIR-16 Question Answering Lab for Political Information 3 (QA Lab-PoliInfo-3). In this paper, we report on our methods for this task and discuss the results. We performed argument classification using a fine-tuned BERT classifier. This method showed the second highest score (0.5716) among the participants in the test data. We also performed linking relatedID using TF-IDF vectorization of documents and calculation of their cosine similarity. This method showed the highest score (0.6596) among the participants in the test data.conference pape
Attempt to Develop An Approach Based on BERT for Task of NTCIR-16 QA Lab-Poliinfo-3 Budget Argument Mining
The SMLAB team participated in the budget argument mining subtask of the NTCIR-16 QALab Poli-info Task. This paper reports our approach to solving this task and discusses the official results. As an reult, our model underperformed other approaches from other teams drastically.conference pape
Overview of the NTCIR-16 RCIR Task
The NTCIR-16 RCIR pilot task aimed to motivate the development of a first generation of personalised retrieval techniques that integrate reading comprehension measures and eye tracker signals as a source of information when ranking text content. The dataset used in the challenge was newly generated by capturing eye movement measures while experimental participants read text passages on a computer screen. The RCIR challenge included two sub-tasks: a) the comprehension-evaluation task (CET) that involved predicting a measure of a reader’s comprehension for text passages and, b) the comprehension-based retrieval task (CRT) that involved retrieving relevant passage texts ranked by comprehension score. The participating teams were ranked using Spearman’s correlation coefficient (rho) for the CET sub-task and normalised Discounted Cumulative Gain (nDCG score) for the CRT sub-task.conference pape
Syapse at the NTCIR-16 RealMed-NLP task
In this paper, we present our approach to subtasks 1, 2, and 3 of the NTCIR-16 RealMed-NLP challenge. For these challenges, the english language corpora (CR-EN and RR-EN) were used. In subtasks 1 and 2, the goal was to create an NLP system which could add tags to case reports (CR) or radiology reports (RR). In subtask 3, two applications of this system were tested: the ability to determine which RRs from a group referred to the same sample, and the ability to determine the probability that a medication caused side effects in a report. Our approach leveraged keyword extraction through a medical metathesaurus (MetaMap), sentence structuring using a SciSpacy model, and word embeddings using a trained BERT model. Using this model, we were able to complete these three subtasks with high levels of accuracy.conference pape
学術流通における検索・発見基盤の在り方
2022年5月31日 学術情報基盤オープンフォーラム2022, コンテンツトラック1「研究データ流通のベストプラクティスと学術基盤の役割」conference objec
ジェネラルな学術基盤から発⾒する研究データ
2022年6月6日 JAPAN OPEN SCIENCE SUMMIT 2022, セッションE1「研究データの「新しい見つけ方」を考える」conference objec
第34回 これからの学術情報システム構築検討委員会 配付資料
会議名:第34回 これからの学術情報システム構築検討委員会
開催場所:オンライン
日時:2022年10月31日(月)13:00~15:00conference outpu