我已经训练了一个主题分类模型。然后当我要将新数据转换为向量进行预测时,它出错了。它显示“NotFittedError: CountVectorizer - Vocabulary is not fit”。但是当我通过将训练数据拆分为训练模型中的测试数据来进行预测时,它起作用了。下面是代码:
from sklearn.externals import joblib
from sklearn.feature_extraction.text import CountVectorizer
import pandas as pd
import numpy as np
# read new dataset
testdf = pd.read_csv('C://Users/KW198/Documents/topic_model/training_data/testdata.csv', encoding='cp950')
testdf.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1800 entries, 0 to 1799
Data columns (total 2 columns):
keywords 1800 non-null object
topics 1800 non-null int64
dtypes: int64(1), object(1)
memory usage: 28.2+ KB
# read columns
kw = testdf['keywords']
label = testdf['topics']
# ?????????
vectorizer = CountVectorizer(min_df=1, stop_words='english')
x_testkw_vec = vectorizer.transform(kw)
Run Code Online (Sandbox Code Playgroud)
这是一个错误
--------------------------------------------------------------------------- …Run Code Online (Sandbox Code Playgroud) python machine-learning scikit-learn text-classification countvectorizer