Med*_*edo 28 feature-extraction scikit-learn categorical-data
我想在我的数据集中的10个特征中编码3个分类特征.我用preprocessing从sklearn.preprocessing如下面这样做:
from sklearn import preprocessing
cat_features = ['color', 'director_name', 'actor_2_name']
enc = preprocessing.OneHotEncoder(categorical_features=cat_features)
enc.fit(dataset.values)
Run Code Online (Sandbox Code Playgroud)
但是,我无法继续,因为我收到此错误:
array = np.array(array, dtype=dtype, order=order, copy=copy)
ValueError: could not convert string to float: PG
Run Code Online (Sandbox Code Playgroud)
我很惊讶为什么它抱怨字符串,因为它应该转换它!我在这里错过了什么吗?
pim*_*314 44
如果您阅读文档,OneHotEncoder您将看到输入为fit"输入数组类型为int".因此,您需要为一个热编码数据执行两个步骤
from sklearn import preprocessing
cat_features = ['color', 'director_name', 'actor_2_name']
enc = preprocessing.LabelEncoder()
enc.fit(cat_features)
new_cat_features = enc.transform(cat_features)
print new_cat_features # [1 2 0]
new_cat_features = new_cat_features.reshape(-1, 1) # Needs to be the correct shape
ohe = preprocessing.OneHotEncoder(sparse=False) #Easier to read
print ohe.fit_transform(new_cat_features)
Run Code Online (Sandbox Code Playgroud)
输出:
[[ 0. 1. 0.]
[ 0. 0. 1.]
[ 1. 0. 0.]]
Run Code Online (Sandbox Code Playgroud)
编辑
由于0.20这成为更容易一点,不仅因为OneHotEncoder现在处理字符串好听,而且还因为我们可以轻松地将多列使用ColumnTransformer,请参阅下面的例子
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import LabelEncoder, OneHotEncoder
import numpy as np
X = np.array([['apple', 'red', 1, 'round', 0],
['orange', 'orange', 2, 'round', 0.1],
['bannana', 'yellow', 2, 'long', 0],
['apple', 'green', 1, 'round', 0.2]])
ct = ColumnTransformer(
[('oh_enc', OneHotEncoder(sparse=False), [0, 1, 3]),], # the column numbers I want to apply this to
remainder='passthrough' # This leaves the rest of my columns in place
)
print(ct2.fit_transform(X)) # Notice the output is a string
Run Code Online (Sandbox Code Playgroud)
输出:
[['1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '0.0' '0.0' '1.0' '1' '0']
['0.0' '0.0' '1.0' '0.0' '1.0' '0.0' '0.0' '0.0' '1.0' '2' '0.1']
['0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1.0' '0.0' '2' '0']
['1.0' '0.0' '0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1' '0.2']]
Run Code Online (Sandbox Code Playgroud)
小智 14
您可以使用LabelBinarizer类在一次镜头中应用两种转换(从文本类别到整数类别,然后从整数类别到单热矢量):
cat_features = ['color', 'director_name', 'actor_2_name']
encoder = LabelBinarizer()
new_cat_features = encoder.fit_transform(cat_features)
new_cat_features
Run Code Online (Sandbox Code Playgroud)
请注意,默认情况下,这会返回密集的NumPy数组.您可以通过将sparse_output = True传递给LabelBinarizer构造函数来获取稀疏矩阵.
从文档中:
\n\ncategorical_features : \xe2\x80\x9call\xe2\x80\x9d or array of indices or mask\nSpecify what features are treated as categorical.\n\xe2\x80\x98all\xe2\x80\x99 (default): All features are treated as categorical.\narray of indices: Array of categorical feature indices.\nmask: Array of length n_features and with dtype=bool.\nRun Code Online (Sandbox Code Playgroud)\n\npandas 数据框的列名不起作用。如果您的分类特征是列号 0、2 和 6,请使用:
\n\nfrom sklearn import preprocessing\ncat_features = [0, 2, 6]\nenc = preprocessing.OneHotEncoder(categorical_features=cat_features)\nenc.fit(dataset.values)\nRun Code Online (Sandbox Code Playgroud)\n\n还必须注意的是,如果这些分类特征没有经过标签编码,那么LabelEncoder在使用之前需要对这些特征进行使用OneHotEncoder
| 归档时间: |
|
| 查看次数: |
32272 次 |
| 最近记录: |