National Institute of Informatics, Center of Dataset Sharing and Collaborative Research: NII DSC Reference Portal / 国立情報学研究所
Not a member yet
119 research outputs found
Sort by
楽天データセット(コレクション)
楽天グループ株式会社が事業等を通じて取得したデータから構築したデータセット。以下のデータを含む。- 楽天市場データ - 楽天トラベルデータ - 楽天GORAデータ - 楽天レシピデータ - 筑波大学文単位評価極性タグ付きコーパス(TSUKUBAコーパス) - カテゴリラベル付き商品画像データセット - 文字領域アノテーション画像 - 楽天不動産間取り図と壁ラベル - Rakuten France: ユーザ評価・レビュー有効性情報データ(旧 PriceMinisterデータ) - Rakuten France: 書籍情報・著者名情報 - Rakuten France: マルチモーダルプロダクトデータセット - 楽天ブックス著者名の曖昧性解消実験用書誌データ - 楽天トラベルレビュー: アスペクト・センチメントタグ付きコーパス - Chart Digitization Dataset 以下のデータは提供を終了。- 楽天オークションデータ - 楽天Vikiデータ データの説明は各データのアイテムに記載。Collection of datasets compiled by Rakuten Group, Inc. using data obtained through its business, etc. Included datasets are: - Rakuten Ichiba data - Rakuten Travel data - Rakuten GORA data (Rakuten\u27s golf service) - Rakuten Recipe data - Tsukuba sentiment-tagged corpus (TSUKUBA corpus) - Product images dataset with category label - Images with character area - Floor plan from Rakuten Real Estate and pixel-wise wall label - Rakuten France: user review, products reviews interests (former PriceMinister data) - Rakuten France: book and author information - Rakuten France: Multi-modal Product Dataset - Rakuten Books bibliographic information for author name disambiguation test - Rakuten Travel Review: Aspects and Sentiment-tagged corpus - Chart Digitization Dataset Provision of following datasets has been terminated: - Rakuten Auction data - Rakuten Viki data See each dataset\u27s item for details
宇都宮大学 パラ言語情報研究向け音声対話データベース(UUDB)
音声言語に付随して伝達されるパラ言語情報に主眼を置いて設計・構築された自然対話の音声コーパス.課題遂行を通して友人同士のいわゆる「ため口」を収録. - 「4コマまんが並べ替え課題」:対話を行いながら互いの情報を交換し,ばらばらになった4コマまんがのストーリーを作成する(インフォメーションギャップを用いた課題) - 大学生(同学年・友人同士)7ペア,計4,737発話 - パラ言語情報のアノテーションを含むSpontaneous speech through task-oriented dialogue - "4-frame cartoon sorting task": Four cards each containing one frame extracted from a cartoon are shuffled, and each participant has two cards out of the four, and is asked to estimate the original order without looking at the remaining cards. - 7 pairs of college students (same grade, friends), total 4737 utterances - Emotional state labels includes
残響下日本語連続数字 音声認識評価環境(CENSREC-4)
ハンズフリー環境下における遠隔発話音声認識の課題の中で,残響に着目した残響下音声認識の評価環境.基本セットとエクストラセットの2種類のデータ群より構成される.発話内容はCENSREC-1に準じている. ・基本セット : CENSREC-1のクリーン環境で収録された音声にインパルス応答を畳み込んだシミュレーション評価.残響下連続数字発声データと評価ツールより構成される.学習用計8,440発話,テスト用計4,004発話. ・エクストラセット : 基本セットに加算性雑音を重畳したマルチコンディションに対する評価.残響・雑音下連続数字発声データ(基本セットと同数)と実環境データ(計2,536発話)で構成される.Common platform for evaluating independently speech recognition accuracy and speech interval detection under noisy environment.The target evaluation framework of CENSREC-4 is distant-talking speech recognition in various reverberation environments. The data contained in CENSREC-4 are connected digit utterances as in CENSREC-1. - Basic data set: connected digit utterances under reverberant conditions and an evaluation framework that is provided as HTK-based HMM training and recognition scripts. 8440 utterances for training, 4004 utterances for test. - Extra data set: connected digit utterances (the same as in basic data set) under reverberant and noisy conditions, and recording data (2536 utterances) in real environments
電総研 単語音声データベース(ETL-WD)
音韻バランスを考慮した単語リスト1542語の読み上げ音声Read speech data of phonetically-balanced 1542 words
RWCP 実環境音声・音響データベース(RWCP-SSD)
・非音声音のドライソース - 非音声音の無響室測定データ(ドライソース) - 各種の部屋における再現データ(インパルス応答との畳み込み)・マイクロホンアレーによるインパルス応答と音声データ - 固定音源のインパルス応答 - 固定音源(音声)の測定データ - 移動音源のインパルス応答 - 移動音源(音声)の測定データ - 拡散音源、背景雑音の測定データ・マイクロホンアレーによる近傍音場でのインパルス応答・対話ロボットの頭部伝達関数・測定に関わるアルゴリズムの解説,使用したソフトウェア・上記測定データを用いた研究事例の紹介Vol.1: "RWCP Sound Scene Database in Real Acoustical Environments"Vol.2: Speech data of fixed sound sources measured by microphone-arrayVol.3: Speech data of moving sound sources measured by microphone-arra
Yahoo! 知恵袋データ(第1版)
ヤフー株式会社が提供している知識検索サービス「Yahoo!知恵袋」において解決済みとなった質問と回答を,「ヤフー知恵袋」のデータベースから抽出したもの。This data consists of questions and answers extracted from the "Yahoo! Chiebukuro" database, which were posted and solved at "Yahoo! Chiebukuro", a knowledge retrieval service provided by Yahoo Japan Corporation
千葉大 日本語地図課題対話コーパス(MapTask)
地図を用いた課題遂行対話.情報提供者と情報追従者が各々地図を持ち,対話による情報交換をしながら,情報提供者の持つルートを情報追従者の地図に再現する課題.4人組で8対話ずつ収録,全128対話,約23時間.使用した地図画像と専用音声再生ソフトを含む.Task-oriented dialogues using maps, in which two speakers participate, an instruction-giver who has a map with a route and an instruction-follower who has a map without a route. The giver instructs the follower verbally to reconstruct the giver\u27s route on the follower\u27s map.There are 128 dialogues in total of about 23 hours.The corpus contains the images of the maps used, dedicated software for reproducing the recorded dialogues
NTT・東北大 親密度別単語了解度試験用音声データセット2007(FW07)
「NTT・東北大 親密度別単語了解度試験用音声データセット(FW03)」に含まれる単語音声を1セット20単語に再構成した1,600語(20単語×20セット×親密度4ランク)を下記5条件で編集. ・擬似音声雑音4条件を重畳(SN 比で3, 0, -3, -6dB) ・雑音重畳なし検査等の便宜を図るため,本データセットの一部を音楽CD 5枚に収録したものも同梱The 1600 spoken words (20 words times 20 sets times 4 ranks) reconstructed from the "NTT - Tohoku University Familiarity-controlled Word Lists (FW03)" and edited under the following 5 conditions: - Pseudo speech noise is superposed with four conditions (S/N ratio: 3, 0, -3, -6 dB). - No noise superposition.A part of this data set is contained in the 5 audio CD’s for the convenience of the examination, etc
特定領域研究「メディア教育利用」日本人学生による読み上げ英語音声データベース(UME-ERJ)
音素学習を念頭においた読み上げセット ・音素バランス文(460文) ・難音文(32文) ・実際の音素学習において利用された文(100文) ・ミニマル単語対(302単語対) ・音素バランス単語(300単語)韻律学習を念頭においた読み上げセット ・イントネーションに関する文(94文) ・文強勢,文リズムに関する文(120文) ・単語アクセントに関する単語(109単語,句)文セットは1文あたり約12人,単語セットは1単語あたり約20人の話者により読み上げネイティブ英語教師による評定ラベリングありEnglish Speech Database Read by Japanese Students: 1. Sentences for learning phonemic pronunciation - 460 phonetically-balanced sentences - 32 sentences including phoneme sequences difficult for Japanese to pronounce correctly - 100 sentences designed for test set - 302 minimal-pair words - 300 phonemically balanced words 2. Sentences for learning prosody of speech - 94 sentences with various intonation patterns - 120 sentences with various accent and rhythm patterns - 109 words with various accent patternsEach sentence in the set is read by 12 speakers and each word in the set is read by 20 speakers.Grading lists by native English teachers are attached
理研ワープロ操作対話音声コーパス(RIKEN-DLG)
Vol.1:文書作成依頼対話 ・ワープロ操作の専門家がユーザの希望を聞きながらコンピュータを用いて文書作成を行う.(ユーザ ⇔ 秘書 ⇔ オペレーター ⇔ 専門家の対話) ・文書作成画面を録画したビデオを見ながら,専門家が自分の作業について説明する.(専門家の独話) ・9対話, 9独話(1対話あたり最長2時間)Vol.2~4:質問応答対話 ・ユーザが自ら文書を作成しながら,ワープロ操作方法について専門家に質問をする対話.(ユーザ ⇔ 専門家の対話) ・Vol.2,3:各18対話(1対話あたり最長1時間) ・Vol.4:15対話(1対話あたり最長2時間)※Vol.4は音声データなしVol.1: Dialogues of request for document making - A professional word processor operator makes documents using a computer listening to the user\u27s requirements (Dialogues among a user, a secretary, an assistant, and a professional). - A professional explains his work watching the viedo display of the document making recording (Monologues of a professional). - 9 dialogues and 9 monologues; no more than two hours per dialogue.Vol.2-4: Question-answer dialogues - Dialogues of a user who makes documents by himself/herself and asks questions to the professional about operation of word processor (Dialogues between a user and a professional). - Vol.2,3: 18 dialogues each; no more than one hour per dialogue. - Vol.4: 15 dialogues; no more than two hours per dialogue. (*Vol.4 does not contain speech data