National Institute for Japanese Language and Linguistics

Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリ
Not a member yet
    3427 research outputs found

    Developing ChaKi.NET lite

    Get PDF
    九州大学総和技研Kyushu UniversitySowa Research現代において,コーパスは言語研究に欠かせない資源となっている。言語学の分野では検索・閲覧・集計インターフェイスを備えたコーパスの利用が多いが,情報学等の分野で作成されたコーパスには必ずしもインターフェイスが提供されるわけではない。類型論研究での活用が期待されるUniversal Dependencies(UD)ツリーバンクもそのようなコーパスの1つである。そこで本研究では,既存の高機能コーパスツールであるChaKi.NETを情報抽出用に特化し,新規ユーザにも利用しやすい軽量版であるChaKi.NET liteを開発した。ChaKi.NETは高機能であるがゆえに利用者にとっての学習コストが高かったが,ChaKi.NET liteではUDに合わせたインターフェイスを提供し,アノテーション機能を省くことで目的の機能を利用しやすくした。本稿ではChaKi.NET lite開発の背景と機能について紹介する。Corpora are indispensable resources for contemporary linguistic research. While corpora used for linguistics research usually have an interface, those developed for informatics research tend to lack one. Universal Dependencies (UD) Treebank, which has proved useful for linguistic typology studies, also lacks an interface. In this study, we developed a lightweight corpus tool named ChaKi.NET lite for new users, specialized from the existing sophisticated ChaKi.NET for information extraction. While one takes a long time to learn to use ChaKi.NET, ChaKi.NET lite reduces the required learning time by omitting the annotation function and providing an interface tailored to UD. This paper introduces the background of the development of ChaKi.NET lite and the new functions that this corpus tool provides.application/pdfdepartmental bulletin pape

    All-Words Word Sense Disambiguation for Historical Japanese

    Get PDF
    application/pdfTokyo University of Agriculture and TechnologyTokyo University of Agriculture and TechnologyNational Institute for Japanese Language and LinguisticsPacific Asia Conference on Language, Information and Computation (PACLIC 37), Hong Kong, 2023/12/02-2023/12/04,PACLIC 37 (2023) Organizing Committee and PACLIC Steering Committeejournal articl

    Linguistic Forms Linked to Abstract Thinking in Subject Learning : Case study of "To Suru" in Mathematics.

    Get PDF
    会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センター近年、日本語指導を必要とする外国人児童生徒が増加しており、教科学習に必要な学習言語能力の支援が問題となっている。本稿では、教科学習の中で求められる抽象的思考と結びつく言語形式を分析することを目的とし、中学校教科書のテキストを対象として形態素解析を行った。まず、数学教科書と理科教科書の比較から、共通して出現しやすい表現と特定の教科に出現しやすい表現が存在することを指摘し、ケーススタディとして数学に特徴的な文型として「AをBとする」に注目して分析を行う。「とする」の前後文脈、出現しやすい単元に関して分析を行い、数学では「とする」が具体的事象における要素と数式に出現する要素を同定し、立式の際に思考の枠組みを設定する用法で用いられることを指摘する。本稿の分析は、語彙だけでなく、特定の文型が教科学習で求められる抽象的な思考と結びつくことを示す事例として位置付けられる。application/pdf筑波大学筑波大学筑波大学University of TsukubaUniversity of TsukubaUniversity of Tsukubaconference pape

    Coverpage and Table of Contents

    Get PDF
    application/pdf会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センターothe

    文献レビュー2 ましこ・ひでのり(2014)『ことばの政治社会学』

    Get PDF
    application/pdf定住外国人のよみかき研究departmental bulletin pape

    文献レビュー4 菊池久一 (1995)『〈識字〉の構造:思考を抑圧する文字文化』

    Get PDF
    application/pdf定住外国人のよみかき研究departmental bulletin pape

    Genre Attribute-related Annotations on Fiction Samples in the Balanced Corpus of Contemporary Written Japanese

    Get PDF
    目白大学国立国語研究所 研究系Mejiro UniversityResearch Department, NINJAL我々は『現代日本語書き言葉均衡コーパス』の書籍サンプルに含まれるすべての小説サンプルについて,小説の内容に関するジャンルや舞台設定等の分類情報(「推理」「SF」「アドベンチャー」「ロマンス」など)を付与した。分類情報の策定にあたっては,小説サンプルの取得された各書籍について,書店や出版社の分類情報をはじめ,小説の内容を表すと複数作業者が判断した特徴語句を広く収集し,結果を整理した。各小説サンプルには様々な分類項目を重複して付与した。本稿の作業により,これまで分類されていなかった小説の分類情報が付与された。新たに付与された分類情報により,分類別の語彙分布や文体特徴が確認できるようになった。本稿では,作業手順と情報付与結果を報告する。We categorized genres and settings (e.g., "Mystery," "Science," "Adventure," "Romance," and "Historical") for all fiction works in book samples from the Balanced Corpus of Contemporary Written Japanese. To design the descriptive genre attributes, we explored the classification items of bookshops and publishers. We also newly defined the classification items by exploring characteristic words and phrases in the fiction contents. Thus, we annotated the designed classification items of genre attributes in a multi-label classification setting. The work described in this study enabled the assignment of new classification information for fiction samples in the Balanced Corpus of Contemporary Written Japanese. The genre attributes enabled us to confirm the distribution of vocabulary and stylistic features. We reported the annotation procedures and results of the classification items of the genre attributes.application/pdfdepartmental bulletin pape

    Design and Construction of the Kokinshu-Towokagami Curpus

    No full text
    総合研究大学院大学京都府立大学国立国語研究所The Graduate University for Advanced Studies, SOKENDAIKyoto Prefectural UniversityNational Institute for Japanese Language and Linguistics『古今集遠鏡』は,本居宣長による『古今和歌集』の注釈書である.本書の特徴は,平安時代の文語,江戸時代後期の文語,同時代の口語の複数の文体が混在している点にある.本書のコーパス構築にあたり,XML による構造化の際には quotation タグの type 属性によって分類することで,文体の区別を示した.また,『日本語歴史コーパス』への収録に向けた形態素解析の際には,それぞれの文体によって UniDic を分けて利用し,XML の文書構造タグによって複数の辞書を切り替えて MeCabによる解析を行った.このように構築されたコーパスを活用し,本書におけるム,テム,ナムの俚言訳の方針を調査した結果,宣長の他の資料では得ることができなかった助動詞の解釈の相違点を見ることができた.また,本コーパスは,同時代の『洒落本コーパス』の収録作品との比較による江戸時代後期の口語の検討や,平安時代語と近世語の比較といった活用が考えられ,日本語史研究の多様化が期待できる.journal articl

    Polysemes in the "Word List by Semantic Principles"

    Get PDF
    会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センター2004年に刊行された『分類語彙表増補改訂版』(以下、分類語彙表)はその「まえがき」によると、初版とくらべて多義語の処理を改良して、「同じ単語を意味に応じて何箇所にも出すようにした」と記述されている(P.6)。しかし、現代の小型国語辞書に掲載されている多義語と比べると、『分類語彙表』の多義語は掲出されている分類項目が少ないものがある。例えば、「切る」は『三省堂国語辞典』(第八版)では動詞の意味が27個、造語成分としての意味が3個あるが、これら30個の意味を『分類語彙表』と対照させると、単独の見出しがあるものが3個、「スイッチを切る」のように連語として見出しがあるものが6個で、計9個しか対応していなかった。残りの21個は、単独の見出しで掲出できそうなもの15個、連語として掲出できそうなもの6個であった。本発表の目的は、使用頻度の高い多義語を取り上げ、『分類語彙表』に収録されていない意味を拾い上げ、増補の候補とすることである。application/pdf国立国語研究所National Institute for Japanese Language and LinguisticsIn the foreword of the "Word List by Semantic Principles, Expanded and Revised Edition" (WLSP"), published in 2004, thereis emphasis on the treatment of polysemy cpmpared to the first edition. Specifically, the text notes that "the same word can be listed in several places according to its meaning" (p.6). However, when compared to polysemous words in contemporary concise Japanese dictionaries, the WLSP seems to have fewer categorization entries for some polysemous words. For example, the verb "cut." in the "Sanseido Kokugo Jiten" (8th edition), it is represented with 27 verb meanings and 3 meanings for constituents of compound words ("Zogo-seibun"). In contrast, within the WLSP, of these 30 meanings, only 3 have individual headings, while 6 adopt phrasal headings like "switch off," amounting to 9 meanings that align with the "Sanseido Kokugo Jiten." The other meanings comprise 15 that could be categorized under individual headings and 6 under phrasal headings. This study seeks to highlight these discrepancies in commonly used polysemous words, pinpoint meanings not covered in the WLSP, and suggest potential inclusions for the WLSP's next iteration.conference pape

    Diachronic Variation of Illative and Adversative Conjunctions: An Analysis with the Magazine Register of the Showa-Heisei Corpus of Written Japanese

    Get PDF
    会議名: 言語資源ワークショップ2023, 開催地: オンライン, 会期: 2023年8月28日-29日, 主催: 国立国語研究所 言語資源開発センター『昭和・平成書き言葉コーパス』雑誌レジスターを用いて、昭和・平成期の非文芸ジャンルの書き言葉における順接・逆接の接続詞について考察した。まず、コーパスから作成した短単位n-gram を用いて接続詞語形を網羅的に抽出したところ、順接の接続詞は31語形、逆接の接続詞は17語形が確認できた。これらの語形について口語文体(常体多 / 敬体多)割合を分析すると、語形と口語文体割合には対応関係が見られ、それは各語形の書き言葉的・話し言葉的性質の強弱と対応していることが明らかになった。次に、『日本語歴史コーパス明治・大正編Ⅰ雑誌』のデータと合わせて、刊行年別に接続詞の各語形の使用サンプル率を算出し、明治期から平成期までの量的な通時的変化を概観した。話し言葉的性質の強い語形が増加すること、また特に常体多の口語文体では語形の整理が進み、少数の語形が使用されるようになることが明らかになった。application/pdf東京大学The University of Tokyoconference pape

    3,205

    full texts

    3,427

    metadata records
    Updated in last 30 days.
    Academic Repository of the National Institute for Japanese Language and Linguistics / 国立国語研究所学術情報リポジトリ
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇