本文へ移動

プレプリント ·日本語 ·未確認

A clinical specific BERT developed with huge size of Japanese clinical narrative

Yoshimasa Kawazoe Daisaku Shibata Emiko Shinohara Eiji Aramaki Kazuhiko Ohe

刊行年
2020-07-09
収録
『medRxiv』
言語
英語
OpenAlex
W3041263301
DOI
10.1101/2020.07.07.20148585
MAG
3041263301
URL
https://www.medrxiv.org/content/medrxiv/early/2020/07/09/2020.07.07.20148585.full.pdf

要旨

Abstract Generalized language models that pre-trained with a large corpus have achieved great performance on natural language tasks. While many pre-trained transformers for English are published, few models are available for Japanese text, especially in clinical medicine. In this work, we demonstrate a development of a clinical specific BERT model with a huge size of Japanese clinical narrative and evaluated it on the NTCIR-13 MedWeb that has pseudo-Twitter messages about medical concerns with eight labels. Approximately 120 millions of clinical text stored at the University of Tokyo Hospital were used as dataset. The BERT-base was pre-trained with the entire dataset and a vocabulary including 25,000 tokens. The pre-training was almost saturated at about 4 epochs, and the accuracies of Masked LM and Next Sentence Prediction were 0.773 and 0.975, respectively. The developed BERT tends to show higher performances on the MedWeb task than the other nonspecific BERTs, however, no significant differences were found. The advantage of training on domain-specific texts may become apparent in the more complex tasks on actual clinical text, and such corpus for the evaluation is required to be developed.

主題

この書誌の出所

  • openalex— W3041263301(2026-08-14取得)

引用

Yoshimasa Kawazoe・Daisaku Shibata・Emiko Shinohara・Eiji Aramaki・Kazuhiko Ohe(2020-07-09) A clinical specific BERT developed with huge size of Japanese clinical narrative 『medRxiv』

Kawazoe2020ClinicalSpecificBERT
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON