National Institute for Japanese Language and Linguistics

Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリ
Not a member yet
    3427 research outputs found

    Diachronic Domain Adaptation of Word Sense Disambiguation in Corpus of Historical Japanese Using Word Embeddings

    Get PDF
    東京農工大学茨城大学茨城大学Tokyo University of Agriculture and TechnologyIbaraki UniversityIbaraki University語義タグ付きコーパスを用いた現代日本語の語義曖昧性解消の研究は数多い。しかし,入手可能なタグ付きコーパスが少ないため,日本語の古典語の語義曖昧性解消を高性能に行うことは難しい。そのため,現代日本語文を用いて通時的な領域適応を行うことは,古典語の語義曖昧性解消の性能を高めるひとつの解決方法であると考えられる。本研究では,日本語の古典語の語義曖昧性解消において,領域適応手法のひとつである,分散表現のfine-tuningの効果について調べる。現代文の分散表現であるNWJC2vecの古典語によるfine-tuningや,古典語によって作成した分散表現の現代文によるfine-tuningなど,様々なfine-tuningのシナリオを検証した。さらに,NWJC2vecを古典語でfine-tuningする際には,時代順に段階的に分散表現をfine-tuningする手法についても試した。語義曖昧性解消の対象語の前後二語ずつの単語の分散表現を素性とし,Support Vector Machineの分類器に用いて分類を行った。シナリオは(1)現代文のコーパスの全用例と古典語のコーパスの用例8割を訓練事例とし,残りの2割の古典語の用例をテストとして利用する場合,(2)古典語の用例だけを利用して五分割交差検定を行った場合,(3)現代文のコーパスの全用例を訓練事例とし,古典語全用例をテストする場合の三通りを比較した。最高の精度となったのは,(2)古典語の用例だけを利用したシナリオで,古典語によって作成した分散表現に現代文によるfine-tuningを行った場合であった。There have been many studies on word sense disambiguation (WSD) in contemporary Japanese. However, it is difficult to achieve high performance of WSD in historical Japanese because of the lack of sense-tagged corpora. Therefore, diachronic adaptation using contemporary Japanese could be a solution. We investigated the effectiveness of the fine-tuning of word embeddings for WSD in historical Japanese. A variety of fine-tuning scenarios are examined, including the case where the word embeddings of contemporary Japanese (NWJC2vec) are fine-tuned with historical Japanese and the case where the word embeddings trained with historical Japanese are fine-tuned with contemporary Japanese. Moreover, when NWJC2vec was fine-tuned with a historical corpus, the case where the word embeddings were gradually fine-tuned in the order of time was also tested. The word embeddings of two words before and after the target word are used as the features for the support vector machine, which is a classifier of WSD. The following three scenarios are compared: (1) all the examples from the contemporary Japanese corpus and 80% examples from the historical corpus are used as the training data for the test of the remaining 20% examples from the historical corpus, (2) 5-fold cross validation of the examples of the historical Japanese corpus, and (3) all the examples from the contemporary corpus are used as the training data for test examples from the historical corpus. The best accuracy was achieved when we used word embeddings trained from a historical corpus and fine-tuned with a contemporary corpus in the 5-fold cross validation scenario.application/pdfdepartmental bulletin pape

    Possibilities of Meta-research on Japonic Descriptive Linguistics : Evaluation of 40 Years of Lexical Research on Southern Ryukyuan and Its Implications for Future Methodology

    Get PDF
    名桜大学国立国語研究所 研究系 非常勤研究員信州大学Meio UniversityAdjunct Researcher, Research Department, NINJALShinshu University本論文では,過去40年間の南琉球の語彙研究を事例に,従来の調査方法や成果に対して詳細な評価を行い,それに基づき,日琉諸語を対象に多地点で詳細な語彙研究を行うという目的を達成するための効果的な研究方法について論じる。過去40年間に行われた南琉球の語彙研究成果を収集および評価した結果,研究者と母語話者が協力して行うハイブリッド型研究が,質と量の両面から最も有効な研究手段であることが分かった。このため,今後の語彙研究にとってハイブリッド型の研究形態を積極的に活用していくことが大きな可能性を秘めていると主張する。一方で,南琉球諸語は消滅危機言語であるため研究期間に制限があるにもかかわらず,語彙研究が行われている地点には激しい偏りが存在していることも指摘する。この問題に対して,ハイブリッド型研究であっても面接調査のみに頼ることは非現実的であることを明らかにし,語彙収集に関わる作業を一部遠隔化することで解消できると指摘する。以上の結果を踏まえ,我々が実施している事例を参照しながら,作業を細分化・分担し,各自が居ながらにして作業を効率的に行うハイブリッド遠隔型の語彙研究を提案する。In this study, we evaluated the methodology and results of 40 years of lexical research on Southern Ryukyuan, and, based on this evaluation, we explored the feasibility of a more efficient lexical research methodology for the Japonic languages. First, we found that a hybrid approach to lexicography with native speakers and researchers working together was the most efficient, both qualitatively and quantitatively. Therefore, there is a strong possibility of active adoption of this approach in future lexical research. Second, lexical research on Southern Ryukyuan was found to be characterized by an extreme variability in data availability for each doculect. Despite adopting such a hybrid approach, however, it is not possible to resolve this data variability by relying solely on the highly time-consuming methodology of traditional fieldwork, especially given that all the languages under study are endangered. This problem can be resolved by devising novel remote data gathering techniques for lexical research. Therefore, we devised a new methodology that combines both the hybrid and remote approaches and report on its implementation in this article.application/pdfdepartmental bulletin pape

    話者の立場に着目した国会会議録における言語表現の通時的分析

    Get PDF
    application/pdf国立国語研究所会議名: シンポジウム「日常会話コーパス」VII, 開催地: オンライン, 会期: 2022年3月7日, 主催: 国語研共同研究プロジェクト「大規模日常会話コーパスに基づく話し言葉の多角的研究」conference objec

    Examining the Rubric Focusing on the Level Distinction for Japanese Speaking Test: Analysis of the Ratings by Non-Japanese Language Teachers

    No full text
    東京大学東京大学国立国語研究所The University of TokyoThe University of TokyoNational Institute for Japanese Language and Linguistics本研究では,根本ほか(2020)で使用されたスピーキングテストSTAR(Speaking Test of Active Reaction)のルーブリックを改良するために,従前のルーブリックver.1と改訂版ルーブリックver.2を用い,非日本語教師97名による状況対応タスクのレベル判定比較実験を行った。ルーブリックver.2は,ver.1の判定実験でのコメント分析結果をもとに,各評価項目をレベルごとに全て記述するのではなく,複数レベルにわたって記述することで,ルーブリックの記述の簡便化を図り,判定レベルの弁別性を焦点化した「弁別性焦点化ルーブリック」である。実験の結果,ver.2は中級後半から上級レベルの弁別力が不足している可能性があるものの,ver.1に比べ,受験者の音声を最後まで聞かなくても判定できる割合が上がった。また,レベル判定のための音声サンプルよりもver.2のルーブリックの方がわかりやすくなったとの結果が出た。このことから,ルーブリックをver.2にしたことで実用性が上がったと言える。本研究では,さらに,この弁別性を焦点化したルーブリックを状況対応タスクだけでなく,音読・シャドーイング・絵描写・再話・意見述べといった他のタスクにも応用し,これらのテストタスクのルーブリックを作成した。journal articl

    Historical change in conditional expressions

    Get PDF
    application/pdfORIGINAL PAPER 阪倉篤義, 1975, 「条件表現の変遷」, 『文章と表現』第3章第4節, 角川書店, pp. 255-273. SAKAKURA Atsuyoshi, 1975, Jōken Hyōgen no Hensen, Chapter 3, Section 4, Bunshō to hyōgen, Kadokawa Shoten, pp. 255-273. Translated by Stephen Wright HORN Proofed by John H. Haig (University of Hawai'i at Manoa)journal articl

    Construction of a Language Resource Package of the Minutes of the National Diet of Japan for the Full-Text Search System "Himawari"

    Get PDF
    国立国語研究所 研究系 音声言語研究領域Spoken Language Division, Research Department, NINJAL本稿は,『国会会議録検索システム』に収録されている国会会議録のテキストデータに基づき,全文検索システム『ひまわり』用の『国会会議録』パッケージを構築する方法,および,構築結果を報告する。本パッケージには,1947(第1回)~ 2012年(第182回)に開催された衆議院・参議院の本会議,および,予算委員会の会議録11106件(約4.49億字)を収録している。本パッケージは言語表現の経年変化分析を行うために設計され,会議情報,発言者情報,会議録の構造情報がXMLで付与されている。本稿では,まず,XMLタグを設計するとともに,原資料の表記上の手がかりを使って,設計したタグを会議録に自動的にアノテーションする方法を示す。次に,考案した手法に基づいて『国会会議録』パッケージを構築する。また,構築したパッケージに収録した会議録の基礎情報を示す。最後に,『国会会議録』パッケージを使って,(a)経年変化が大きい表現を抽出する方法,(b)抽出された表現に対する経年変化要因を調査する方法を示すことにより,『国会会議録』パッケージの有用性を示す。This paper presents the method whereby a language resource package of the Minutes of the National Diet of Japan was constructed for the Full-Text Search System "Himawari" from text data stored in the Full-Text Database System for the Minutes of the Diet and reports the results of the construction. This package includes 11106 minutes (about 450 million characters) of the 1st (1947) to 182nd (2012) plenary sessions and budget committee meetings in the House of Representatives and the House of Councillors. Information related to the meetings, speakers, and the document structures of the minutes are annotated to the minutes in XML to facilitate the analysis of temporal changes in linguistic expressions. In this paper, I first describe the XML tags and an automatic annotation method created using notational clues in the minutes, then I detailed the application of the annotation method to the original minute data to construct the package and summarized the results. Finally, this paper classifies the usefulness of the package by showing how it can be used (a) to extract expressions showing large temporal changes and (b) to investigate the factors of the changes.application/pdfdepartmental bulletin pape

    Word Sense Disambiguation of Corpus of Historical Japanese Using Japanese BERT Trained with Contemporary Texts

    Get PDF
    application/pdfTokyo University of Agriculture and TechnologyTokyo University of Agriculture and TechnologyNational Institute for Japanese Language and Linguisticshttps://aclanthology.org/2022.paclic-1.49/journal articl

    The Development of the ‘Simple Japanese Test’ for the use in Japanese Language Education and Research : With a focus on the development process and item analysis results

    Get PDF
    application/pdf東京都立大学城西大学国立国語研究所Tokyo Metropolitan UniversityJosai UniversityNational Institute for Japanese Language and LinguisticsIn this study, we developed the ‘Simple Japanese Test’ for the use in Japanese language education and research, which can be used stably in various environments, is easy to operate, and can be administered without placing too much burden on learners. It is a computer-based web test and has two types: ‘Vocabulary’ and ‘Grammar’. Each test is in a four-way multiple-choice format, with a total of 60 questions, divided into four levels. The time limit was 30 minutes for each. In developing these tests, a trial study was conducted with 80 and 72 questions in each test. Cronbach’s alpha coefficient, an indicator of internal consistency, was 0.945 for vocabulary and 0.951 for grammar, both of which were found to be highly reliable. Item analysis was also conducted, and the questions were refined based on item difficulty, item discrimination power index, choice rate and G-P analysis chart to prepare the test for practical use.journal articl

    NINJAL Parsed Corpus of Modern Japanese の構築と公開

    No full text
    東北大学国立国語研究所名古屋大学大学院人文学研究科弘前大学人文社会科学部journal articl

    『現代語の助詞・助動詞』分類語彙表番号付与版

    No full text
    application/zip国立国語研究所National Institute for Japanese Language and Linguisticsdatase

    3,205

    full texts

    3,427

    metadata records
    Updated in last 30 days.
    Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリ
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇