National Institute for Japanese Language and Linguistics
Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリNot a member yet
3427 research outputs found
Sort by
Proceedings of Language Resources Workshop 2020
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センターapplication/pdfconference pape
Coverpage and Table of Contents
application/pdf会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センターothe
Citizen's Life Before and After the Period of High Economic Growth in Local Cities : That Era and Region in Movies of Shizuoka, Ibaraki, Kanagawa Prefectural Government News Reels
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センター筆者は先行研究として、神奈川県政ニュースのうち川崎市政ニュース映画を題材に、特にナレーション表現に着目して、戦後昭和2,30 年代の都市部の市民生活などについて考察してきた。本研究では、同様に自治体による行政映画である、茨城県と静岡県の県政ニュースを題材に、高度成長期を挟んだ、地方都市における市民生活の変化を、都市部川崎と比較する。茨城、静岡共、昭和20 年代から、広域自治体レベルでの記録映画が残されている。それら行政によるニュース映画を、政策ニュース映画と総称する。
政策ニュース映画は、1 篇が短く、映像自体が、通常の映画に比べて、シンボリックに表現されているという傾向がある。本研究は、昭和30 年代以降に生まれて行った、地方と都市という対比構造の中で、その地の市民生活の小さな差異が、高度成長期に急激に顕在化していく記録として、政策ニュース映画を捉えるものである。application/pdfフェリス女学院大学国立国語研究所 / 神奈川大学Ferris UniversityNational Institute for Japanese Language and Linguistics / Kanagawa Universityconference pape
The Usage of Quotations in Real-life Conversation and Pseudo Conversation : Focusing on "-tte+verb"
会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センター本発表は、「引用標識『ッテ』+動詞」で表される引用文が、現実会話と小説内会話それぞれにおいて、どのように使用されるのかを調査したものである。調査の結果、レジスターが異なるといえども、「ッテ」に後続する動詞は「いう」が最も多いことがわかった。しかし現実会話中では、先行する話の内容の一部を再度引用して新たな表現として引用部に取り込む「ッテ+いう。」引用文(例:いやでもこのままじゃなっていう。)が多く、一方小説内会話では、話し手の感情を表すために使われるような「ッテ+いう」引用文(例:どうしたっていうんだ?)が多かった。
考察として、目の前いる相手の理解度や話の流れによって、再度同じ内容を挟みこみながら進んでいく「現実会話の特性」と、小説の読み手に対しても登場人物の会話に込めた気持ちを表していることになる「小説内会話の特性」の違いが、上記のような結果に影響したのではないかと考えた。application/pdf国際交流基金日本語国際センターThe Japan Foundation Japanese-Language Institute,Urawaconference pape
KOTONOHA Contest 2020 Excellence Award
application/pdf会議名: 言語資源活用ワークショップ2020, 開催地: オンライン, 会期: 2020年9月8日−9日, 主催: 国立国語研究所 コーパス開発センターothe
Machine Learning-based Sentence Boundary Detection for Modern Japanese Texts
首都大学東京首都大学東京国立国語研究所首都大学東京Tokyo Metropolitan UniversityTokyo Metropolitan UniversityThe National Institute for Japanese Language and LinguisticsTokyo Metropolitan UniversityIn this study, we propose a method to detect sentence boundaries for modern Japanese texts using machine learning. For modern Japanese texts, sentence boundaries are not explicitly marked so that human annotation is inevitable, but the annotation process is far from complete due to enormous number of materials. Therefore, we propose a method to detect sentence boundaries using machine learning. The main contribution of this study is that this method can support the annotation task as a primary annotation. We also show that the accuracy of morphological analysis can be improved by performing sentence boundary detection. Moreover, this is the first work to detect sentence boundaries targeting modern Japanese texts by using modern Japanese data for model training and comparing multiple machine learning methods.本稿では,機械学習を用いて近代の歴史的資料に対して文境界を検出する手法を提案する.近代の歴史的資料は明確な文境界が必ずしも存在しないため,これまで人手作業による文境界の付与が行われてきたが,膨大な資料に対してなかなか作業が進んでいない現状がある.そこで我々は機械学習を用いて文境界を検出する手法を提案する.この手法により膨大な量の資料に対して文境界の一次的なアノテーションを施すことができることに加えて,形態素解析の精度を向上させたことが本研究の貢献である.また,モデルの訓練に日本語の近代語のデータを使用して,複数の機械学習手法を比較して近代の歴史的資料を対象とした文境界推定を行うのは本研究が初めてである.journal articl
Word familiarity Rate and Register Type Estimation Using a Bayesian Linear Mixed Model
application/pdf国立国語研究所National Institute for Japanese Language and LinguisticsThis paper presents research on word familiarity rate estimation using the 'Word List by Semantic Principles'. We collected rating information on 96,557 words in the 'Word List by Semantic Principles' via Yahoo! crowdsourcing. We asked 3,392 subject participants to use their introspection to rate the familiarity and register information of words based on the five perspectives of 'KNOW', 'WRITE', 'READ', 'SPEAK', and 'LISTEN', and each word was rated by at least 16 subject participants. We used Bayesian linear mixed models to estimate the word familiarity rates. We also explored the ratings with the semantic labels used in the 'Word List by Semantic Principles'.本論文では『分類語彙表増補改訂版データベース』に対する単語親密度推定手法について述べる。分類語彙表に収録されている96,557項目に対する評定情報をYahoo!クラウドソーシングを用いて収集した。1項目あたり最低16人(異なり3,392人)の研究協力者に,内省に基づいて「知っている」「書く」「読む」「話す」「聞く」の評定情報付与を依頼した。研究協力者の評定情報から単語親密度をベイジアン線形混合モデルにより推定した。また,推定された単語親密度と分類語彙表の語義情報との関連性について調査した。journal articl
Generation and Evaluation of Concept Embeddings Via Fine-Tuning Using Automatically Tagged Corpus
application/pdfIbaraki UniversityNational Institute for Japanese Language and LinguisticsIbaraki Universityjournal articl
Gender categorization by children with normal hearing and children with cochlear implants
国際医療福祉大学国際医療福祉大学国立国語研究所国際医療福祉大学国際医療福祉大学国際医療福祉大学独立行政法人国立病院機構東京医療センターInternational University of Health and WelfareInternational University of Health and WelfareNational Institute for Japanese Language and LinguisticsInternational University of Health and WelfareInternational University of Health and WelfareInternational University of Health and WelfareNational Hospital organization Tokyo Medical Centerjournal articl