我正在用文字袋来分类文字.它运作良好,但我想知道如何添加一个不是一个单词的功能.
这是我的示例代码.
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.svm import LinearSVC
from sklearn.feature_extraction.text import TfidfTransformer
from sklearn.multiclass import OneVsRestClassifier
X_train = np.array(["new york is a hell of a town",
"new york was originally dutch",
"new york is also called the big apple",
"nyc is nice",
"the capital of great britain is london. london is a huge metropolis which has a great many number of people living in it. london is also a very …Run Code Online (Sandbox Code Playgroud) python classification machine-learning scikit-learn text-classification