National Institute for Japanese Language and Linguistics
Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリNot a member yet
3427 research outputs found
Sort by
Word Component Database for “Hands-On Medical Terms” (Version 3) Specification
会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センター我々は、医療用語の合成語の語構造および語構成要素とその意味を明らかにすることを目的に、合成語7,087語を分析し『実践医療用語_語構成要素語彙試案表 Ver.2』を作成した。この作成の過程で、(1)医療用語の選定方法、(2)分割単位の曖昧性、(3)語構造の記述方法、(4)意味ラベルの命名と付与方法に課題がみつかった。そこでこれらの課題を検討し、改良版の試案表Ver.3の作成に着手している。本発表では、改訂版Ver.3の公開にむけて、これらの問題点と当面の方針について述べる。application/pdf奈良先端科学技術大学院大学杏林大学大阪大学京都大学関西大学国立国語研究所Nara Institute of Science and TechnologyKyorin UniversityOsaka UniversityKyoto UniversityKansai UniversityNational Institute for Japanese Language and Linguisticsconference pape
Overview of the Web Content "Tsukuba Vocabulary Checker" for Japanese language teachers
会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センター本発表では、現在構築中の「つくば語彙チェッカー(仮)」についての紹介を行う。「つくば語彙チェッカー(仮)」は、web上で形態素解析を実行し、「リーディング チュウ太」、「日本語教育語彙表 Ver1.0」、「日本語文法項目用例文データベース『はごろも』ver.3」、「EDR電子化辞書」といったデータベースをもとに、頻度や語彙レベルを一括して確認することができる日本語教師向けのコンテンツである。また、この「つくば語彙チェッカー(仮)」は、入力したテキストから空欄補充問題などを作成することが可能な機能も搭載している。本発表では、これらの紹介に加えて、「つくば語彙チェッカー(仮)」の使いやすさ向上のためにどのような改修を行っているか、今後どのような機能を実装予定なのかについても紹介する。application/pdf筑波大学筑波大学筑波大学University of TsukubaUniversity of TsukubaUniversity of Tsukubaconference pape
Measuring linguistics of the wokototen chart made inductively by deciphering kunten materials
University of TsukubaNational Institute of Technology(KOSEN), Gifu CollegeUniversity of ToyamaNational Institute for Japanese Language and LinguisticsIn this paper, we focus on wokototen markings, which are a system of kunten annotations used to facilitate the reading of classical Chinese documents by Japanese readers. Using digitized data, we performed basic measurements of wokototen by using a chart that summarizes the wokototen markings of actual kunten materials described by Hiroshi Tsukishima, and we quantitatively clarified their characteristics. Kunten materials are classical Chinese books with annotations, called kunten, on the Chinese text. The wokototen is a type of kunten. In ancient East Asian countries, kunten systems were developed as a way of directly annotating Chinese documents so that they could be read and understood by non-native readers. For this reason, kunten materials and kunten are treated as historical sources for linguistic and historical research. The shape and position of a wokototen marking determines what kind of reading it indicates. The results of our basic survey quantitatively show that almost all the wokototen charts in actual kunten materials contain particles represented by “te”, “ni”, and “wo”, the most common shapes of wokototen are dots and shapes that can be written with a single stroke, such as |, ─, and \, and that the most common places to find these markings are to the right of characters in the horizontal direction and below characters in the vertical direction.journal articl
A study on the usage of Japanese demonstrative “SO” using Japanese Map Task Dialogues Corpus
国立国語研究所千葉大学National Institute for Japanese Language and LinguisticsChiba Universityjournal articl
AB/C Tone Class Information of the Kohama Dialect for the Reconstruction of the Classified Vocabulary of Proto-Yaeyaman
journal articl
Constructing a Collocation Database for the CEFR-J Wordlist
会議名: 言語資源ワークショップ2022, 開催地: オンライン, 会期: 2022年8月30日-31日, 主催: 国立国語研究所 言語資源開発センター外国語学習において、語彙を運用する力を学習者が身につけるためには語彙リストで単語を個別に記憶するだけでなく、使用頻度の高いフレーズで提示するのが重要である。近年、コーパス準拠英語語彙・フレーズリストが発表されているが、対象学習者レベル、受容・産出語彙の区別は明確ではない。本研究では、CEFR準拠英語学習語彙表CEFR-J Wordlistの活用度を上げ、学習に資するコロケーション選択をするための資料となるコロケーション・データセットを整備する。1億語のイギリス英語均衡コーパスであるBritish National CorpusからUniversal Dependenciesに基づいた構文解析を行い、共起フレーム別(例:動詞+名詞、形容詞+名詞)に共起語セットを抽出した。教育的に有用なコロケーションを教師や学習者が選定するための情報源として、単純共起頻度以外に検索語-共起語ペアについて、各語のCEFR-J Wordlistに基づいたCEFRレベル情報、共起統計(MI, MI2, MI3, t-score, z-score ,logDice, log-likelihood, chi-squared)と散布度指標(DP)の情報を付与した。application/pdf東京外国語大学東京外国語大学Tokyo University of Foreign StudiesTokyo University of Foreign Studiesconference pape
The Construction and longitudinal analysis of a learner corpus of Japanese on the elementary, intermediate and advanced level
会議名: 言語資源ワークショップ2022, 開催地: オンライン, 会期: 2022年8月30日-31日, 主催: 国立国語研究所 言語資源開発センター初級、中級、上級レベルのスロベニア人日本語学習者コーパスを構築した。その中の文法誤用の分析を行った。各レベルで最も頻繁な文法誤用を判別し、レベル間の結果を比較することにより、さまざまなタイプの誤用の発生傾向の変化を分析した。その結果、237のテキストと、1778のマークおよび分類された文法誤用を含むコーパスが作成された。習熟度のレベルごとに、誤用の数が最も多いカテゴリーを決定した。コーパスには縦断的なデータも含まれているため、そのデータも分析し、両方の分析の結果を比較した。最後に、誤用が最も多い文法をより詳細に分析した。分析の結果は、スロベニア人日本語学習者のデータを含む初めての公的コーパスが公開された。application/pdfリュブリャーナ大学文学部アジア研究学科University of Ljubljana, Department of Asian Studiesconference pape
Prosodic Patterns as Evidence for Syntactic Structure: A View from the Genitive Subject Construction in Modern Korean
会議名: Evidence-based Linguistics Workshop 2023, 開催地: 国立国語研究所, 会期: 2023/09/14-15, 主催: 国立国語研究所、神戸大学人文学研究科統語論、音韻論、意味論など言語学の各分野においては、それぞれの現象を検討するために、細分化されたそれぞれの分野内のデータが証拠とされることが多い。しかし有効な証拠は分野内に限らず、分野外のデータから得られることもある。本発表では、現代韓国語の属格主語構造を一例として、統語構造に関する仮説の検証に韻律パターンを証拠として使用することの有効性を示す。現代日本語では、「母親が焼いたチジミ/母親の焼いたチジミ」のように連体修飾節中の主格と属格が交替することが可能であるが、現代韓国語/朝鮮語では方言によって可能性が異なることが指摘されている(Sohn, 2004; 金銀姫 2014)。ここで「母親の」のような名詞句が連体修飾節の主語であるという証拠を示すために、従来の研究では修飾語を加えた複雑な文の意味判断を行わせることが多かった。本発表では、例文を各方言の母語話者に音読させた韻律パターンを分析することで、名詞句が連体修飾節の主語であることの明瞭な証拠が得られることを示す。application/pdf帝京大学国立国語研究所名古屋大学早稲田大学Teikyo UniversityNational Institute for Japanese Language and LinguisticsNagoya UniversityWaseda Universityconference pape
文献レビュー5 小柳正司(2010)『リテラシーの地平:読み書き能力の教育哲学』
application/pdf定住外国人のよみかき研究departmental bulletin pape
Digitizing "Bound forms ('zyosi' and 'zyodôsi') in modern Japanese: Uses and examples" : Application in the Analysis of Simile
神戸大学Kobe University『現代語の助詞・助動詞―用法と実例―』を表形式のデータベースとして電子化し,助詞と助動詞の用法の体系的な記述枠組みとして利用できるようにした。本論文では,その電子化の基本方針とデータベースの内容を概説した。著者は,このデータベースの応用例として,『日本語レトリックコーパス』に収録された直喩のすべての用例を対象として,その構文形に含まれる助詞と助動詞の用法のアノテーションを行った。アノテーションの対象となった助詞,助動詞のほとんどすべての用法について,一貫性のある分類を行うことができた。その結果,直喩には,特定の助詞と助動詞が生起しやすく,それぞれの語の用法の頻度にも著しい偏りがあることが分かった。本論文で示した事例分析の成果は,任意の日本語データに含まれる助詞・助動詞の用法分析を行う上で,このデータベースが汎用性の高い記述枠組みとして利用できる可能性を示唆している。We digitized the volume "Bound forms ('zyosi' and 'zyodôsi') in modern Japanese: Uses and examples" (The National Language Research Institute, 1951) as tabular data and made it available for the analysis of Japanese bound forms. In this paper, we describe how we digitized the volume and what content has been included in the database. We also applied the database as an annotation framework for simile analysis, using examples from The Corpus of Japanese Figurative Language (J-FIG). The database enabled us to analyze the data systematically. Consequently, we found that certain usages of bound forms were dominant in the examples of simile and that they formed patterns of grammatical construction. The results suggest that the database might be used as a general-purpose framework for describing how bound forms are used in Japanese sentences.application/pdfdepartmental bulletin pape