National Institute for Japanese Language and Linguistics
Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリNot a member yet
3427 research outputs found
Sort by
Present state analysis and measurement of pronunciation training effectiveness in English acquisition : Relationships between production patterns and English proficiencies
Juntendo UniversityNational Institute for Japanese Language and LinguisticsJuntendo Universityjournal articl
The *Long-C Constraint and Word-Initial Aspirates in Hateruma Yaeyama, a Southern Ryukyuan Language
application/pdf国際基督教大学 / ベンダ大学国立国語研究所International Christian University / University of VendaNational Institute for Japanese Language and LinguisticsHateruma Yaeyaman is an endangered southern Ryukyuan language spoken in the Hateruma island. In Hateruma, there is a limited distribution of strong aspiration in disyllabic words which is argued to be the result of a prosodic condition for a foot: all feet must have at least one heavy syllable. After presenting the distribution of strong aspiration in Hateruma, we propose the *Long-C constraint that is violated when a consonant is phonetically lengthened. The prosodic requirement for a heavy syllable in a foot, coupled with the *Long-C constraint results in a repair strategy that associates an epenthetic mora with an onset consonant that is realized with strong aspiration.波照間方言は,沖縄県八重山郡竹富町に属する波照間島で話されている言語である。波照間方言の特 徴として,語頭の無声阻害音に強い帯気が報告されている。この帯気は,フットに関する韻律条件,す なわちフットには必ず重音節が必要であるという条件を満たすために実現していると考えられる。本論 文では,強い帯気が観察されるパターンを概観し,*Long-C 制約の適用を提案する。*Long-C 制約とは, 音声的に長い区間を持つ子音を禁止する制約である。本制約と韻律条件の結果,代替操作として頭子音 がモーラを担い,強い帯気が実現すると考える。journal articl
The Tolerance of Loanword Verbs among Japanese Native Speakers : From the Survey with I-JAS
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センター「ドラマが/車がヒットする」のような外来語サ変動詞についての研究は、サ変動詞化の基準を始め十分に解明されていない。本研究では、外来語サ変動詞における日本語母語話者の許容状況を明らかにするため、I-JAS(国立国語研究所)に出現した外来語サ変動詞に対する日本語母語話者の許容度調査を行った。I-JASから収集した73語のうち、BCCWJの出現状況や外来語辞典等の用例と照合し、まだ十分に日本語として定着していない外来語サ変動詞57語を選出した。その上で、大学・大学院生89名から容認度判定を得た。調査の結果、外来語名詞と同様に「意味の縮小・特殊化」が外来語サ変動詞の許容度にも大きく影響しており、日本語に借用された際に意味の縮小や特殊化が起こっていると、原語の意味でのサ変動詞の許容度が低下することが分かった。これらの成果は、外来語の動詞化や、「日本語の外来語/外国語」の判断に関する基準の明確化に貢献し得ると思われる。application/pdf金沢大学Kanazawa Universityconference pape
Design of Real-Time MRI Articulatory Movement Database
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センターわれわれは、日本語に関する調音音声学的研究の新しいインフラ提供をめざして、調音運動データベースの構築を進めてきている。声道全体の形状変化を毎秒14 ないし25 フレームのリアルタイムMRI 動画として記録したデータが、現時点で東京方言16 名、近畿方言5名分収録済である。1 名あたりの発話量は25〜30 分である。データには個々の発話の開始時刻と終了時刻のタグが付与されており、他に発話内容と話者に関するメタデータを検索に利用できる。application/pdf国立国語研究所国立国語研究所早稲田大学国立国語研究所ATR-PromotionsATR-Promotions千葉工業大学甲南大学拓殖大学国立国語研究所国立国語研究所早稲田大学早稲田大学ピコラボNational Institute for Japanese Language and LinguisticsNational Institute for Japanese Language and LinguisticsWaseda UniversityNational Institute for Japanese Language and LinguisticsATR-PromotionsATR-PromotionsChiba Institute of TechnologyKonan UniversityTakushoku UniversityNational Institute for Japanese Language and LinguisticsNational Institute for Japanese Language and LinguisticsWaseda UniversityWaseda UniversityPicolabconference pape
Natural Language Processing : Language Resources and Semantic Processing\n
application/pdf東北大学国立国語研究所Tohoku UniversityNational Institute for Japanese Language and Linguisticsjournal articl
A Reexamination of the Silent Characters in the Kunten Materials of the Early Heian Period : A Study of the Language of Kunten Materials Based on Corpora and Digitized Texts
国立国語研究所 研究系 言語変化研究領域 非常勤研究員Adjunct Researcher, Language Change Division, Research Department, NINJAL平安初期訓点資料は,上代語資料と『古今和歌集』以降の平安時代仮名文学の間の日本語を伝える重要な資料群である。従来書籍の形で用いてきたこのような訓点資料群を電子化テキスト,XMLファイル,あるいはコーパスの形で公開していくことで,これまで明らかになっていなかった平安時代語成立の諸相が垣間見えるものと思われる。本研究では,西大寺本『金光明最勝王経』平安初期点の電子化テキスト(作成中)を用いて,平安初期訓点資料における不読字の全体像を計量的に眺め,考察した。
実際に眺めてみると,いわゆる「置き字」と言われる後世の漢文訓読における読まれない字の在り方とは異なる様相が見受けられる。たとえば,後世,格助詞の「の」として読む「之」は,同じく日本語の格助詞や接続助詞としての用法を持つ「而」「於」といった字と同様に,助詞としては読まず,前後の名詞に必要な意味を読み添えて読む傾向がある。このことは,漢字と意味との一対一の対応よりも,助詞は助詞として等しく名詞や節の後に出現するという,素朴な,それゆえに実態に沿った把握がなされていたことを示すように思われる。
また,平安初期に後世とはやや異なった規則で訓読されたことを考慮すると,平安中期以降の記録体(和化漢文)に見られる特徴的な語法が発生した理由を説明することもできる。たとえば「可……之由」の「之」が不読であるにもかかわらず必ず記されることも,記録体が発生した平安初期に「之」は助詞として把握されておらず,句や節の末尾を示すマークとして捉えられていたと考えることで,引用などを読み誤らないように記すために適切な構文の発生や,そのような文体の発生を可能にした背景の説明が可能になるのである。Kunten materials in the early Heian period hold great importance in determining aspects of Old Japanese between the materials from antiquity and the kana literature of the Heian period after the Kokin Waka Collection. The publication of these kunten materials, which have heretofore been published only in books, in new formats such as electronic texts, XML files, or corpora will help reveal how Old Japanese evolved in the Heian period. This study surveyed and analyzed the muted characters in kunten materials in the early Heian period metrically through the use of electronic texts (currently being compiled) of the Golden Light Sutra (Saidaiji version).
This research has revealed different aspects in reading these texts of what is called okiji, Chinese characters unread in later Japanese readings. For example, the character 之 tended to be read not as a postpositional particle but as a word supplying a preceding or following noun with necessary meaning, like such characters as 而 or 於 that were sometimes used as nominative or conjunctive particles in Japanese. This suggests that in the early Heian period people tended not to identify each Chinese character with a particular meaning in Japanese in a one-to-one fashion but to understand that in practice a postpositional particle simply appeared after a noun or a clause.
Moreover, in light of the slight differences in the rules of kundoku in the early Heian period and later periods, it is possible to explain why characteristic wording in kirokutai (Japanization of Chinese writings) in and after the mid-Heian period developed. For instance, we know that 之 in 可之由 was unread but was always written in kirokutai. If we suppose that 之 was not perceived as a postpositional particle but only as a sign marking the end of a phrase or clause in the early Heian period (when kirokutai was created), we can explain the rise of syntactic conventions to avoid misinterpreting quotations and the background that enabled it to arise.application/pdfdepartmental bulletin pape
Design, Evaluation, and Preliminary Analysis of the Monitor Version of the Corpus of Everyday Japanese Conversation
国立国語研究所 音声言語研究領域国立国語研究所 音声言語研究領域 非常勤研究員国立国語研究所 音声言語研究領域 非常勤研究員国立国語研究所 音声言語研究領域 非常勤研究員国立国語研究所 音声言語研究領域国立国語研究所 音声言語研究領域 非常勤研究員国立国語研究所 音声言語研究領域 非常勤研究員千葉大学国立国語研究所 コーパス開発センター 非常勤研究員Spoken Language Division, Research Department, NINJALAdjunct Researcher, Spoken Language Division, Research Department, NINJALAdjunct Researcher, Spoken Language Division, Research Department, NINJALAdjunct Researcher, Spoken Language Division, Research Department, NINJALSpoken Language Division, Research Department, NINJALAdjunct Researcher, Spoken Language Division, Research Department, NINJALAdjunct Researcher, Spoken Language Division, Research Department, NINJALChiba UniversityAdjunct Researcher, Center for Corpus Development, NINJAL国立国語研究所共同研究プロジェクト「大規模日常会話コーパスに基づく話し言葉の多角的研究」では,『日本語日常会話コーパス』(CEJC)の構築を進めている。CEJCは,日常会話の多様性を捉え自然な会話行動が観察できるよう,様々な種類の会話をバランスよく収めることを目標に掲げている。2021年度末に予定している本公開に先立ち,コーパスの利用可能性や問題などを把握するために,目標とする200時間のうち50時間の会話データについて,2018年12月にモニター公開を開始した。本稿ではまず,コーパスの設計について,会話の収録法,データの公開方針,調査協力者の内訳,コーパスの規模や構成などの観点から概観する。次に,収録されているデータが設計通りバランスがとれているかを,話者と会話の両面から検証する。最後に,コーパスを用いた予備的分析を通して,CEJCモニター版を活用した研究の可能性を示す。We have been constructing the Corpus of Everyday Japanese Conversation (CEJC) under the NINJAL collaborative research project since 2016. The CEJC is designed to contain various kinds of everyday conversations in a balanced manner to capture the diversity of everyday conversations and to observe natural conversational behavior. Prior to the publication of the whole corpus, which scheduled for 2022, we published the monitor version of the CEJC in December 2018. In this paper, we first outlined the design of the monitor version of the CEJC, including recording methods, the release policy of the corpus, corpus size, and annotations. Then, we examined whether the speakers and the conversations in the corpus vary in a balanced manner. Finally, we conducted a preliminary analysis on some linguistic aspects of the monitor version of the CEJC, revealing the possible implications of the corpus.application/pdfdepartmental bulletin pape
Design of BCCWJ-EEG : Balanced Corpus with Human Electroencephalography
application/pdfWaseda UniversityNational Institute for Japanese Language and LinguisticsThe past decade has witnessed the happy marriage between natural language processing (NLP) and the cognitive science of language. Moreover, given the historical relationship between biological and artificial neural networks, the advent of deep learning has re-sparked strong interests in the fusion of NLP and the neuroscience of language. Importantly, this inter-fertilization between NLP, on one hand, and the cognitive (neuro)science of language, on the other, has been driven by the language resources annotated with human language processing data. However, there remain several limitations with those language resources on annotations, genres, languages, etc. In this paper, we describe the design of a novel language resource called BCCWJ-EEG, the Balanced Corpus of Contemporary Written Japanese (BCCWJ) experimentally annotated with human electroencephalography (EEG). Specifically, after extensively reviewing the language resources currently available in the literature with special focus on eye-tracking and EEG, we summarize the details concerning (i) participants, (ii) stimuli, (iii) procedure, (iv) data preprocessing, (v) corpus evaluation, (vi) resource release, and (vii) compilation schedule. In addition, potential applications of BCCWJ-EEG to neuroscience and NLP will also be discussed.journal articl
Construction of a Corpus of Assigned School Essays
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センター児童の作文能力を研究するための資料整備を目的として、現在の児童の作文調査や、過去の作文資料の電子化を進めている。この研究の一環として、国語研究所所蔵の1980年代の作文資料(島村1987)を電子化したので、その概要を報告する。この資料は昭和58年に千葉県内の公立小学校2年、4年、6年の児童の作文を調査したもので、「学校」「先生」「ともだち」の3つの課題を含む。原資料は約1440篇ほどの規模の調査と考えられるが、資料の欠落もあり、電子化した資料は1021篇である。資料の概要と電子化作業の詳細について報告し、既に構築済みの「児童・生徒作文コーパス」(2014-2016)、「「手」作文コーパス」(1992, 2016)との違いについて、文字種の構成比を中心に説明する。application/pdf筑波大学富山大学University of TsukubaUniversity of Toyamaconference pape