从from_pretrained的文档中,我知道我不必每次都下载预训练的向量,我可以使用以下语法保存它们并从磁盘加载:
- a path to a `directory` containing vocabulary files required by the tokenizer, for instance saved using the :func:`~transformers.PreTrainedTokenizer.save_pretrained` method, e.g.: ``./my_model_directory/``.
- (not applicable to all derived classes, deprecated) a path or url to a single saved vocabulary file if and only if the tokenizer only requires a single vocabulary file (e.g. Bert, XLNet), e.g.: ``./my_model_directory/vocab.txt``.
Run Code Online (Sandbox Code Playgroud)
所以,我去了模型中心:
我找到了我想要的模型:
我从他们提供给这个存储库的链接下载了它:
使用掩码语言建模 (MLM) 目标的英语语言预训练模型。它是在本文中介绍的,并首次在此存储库中发布。此模型区分大小写:它区分英语和英语。
存储在:
/my/local/models/cased_L-12_H-768_A-12/
Run Code Online (Sandbox Code Playgroud)
其中包含:
./
../
bert_config.json
bert_model.ckpt.data-00000-of-00001
bert_model.ckpt.index
bert_model.ckpt.meta
vocab.txt …Run Code Online (Sandbox Code Playgroud)