OneHotEncoder对分类功能的问题

Med*_*edo 28 feature-extraction scikit-learn categorical-data

我想在我的数据集中的10个特征中编码3个分类特征.我用preprocessing从sklearn.preprocessing如下面这样做:

from sklearn import preprocessing
cat_features = ['color', 'director_name', 'actor_2_name']
enc = preprocessing.OneHotEncoder(categorical_features=cat_features)
enc.fit(dataset.values)
Run Code Online (Sandbox Code Playgroud)

但是,我无法继续,因为我收到此错误:

    array = np.array(array, dtype=dtype, order=order, copy=copy)
ValueError: could not convert string to float: PG
Run Code Online (Sandbox Code Playgroud)

我很惊讶为什么它抱怨字符串,因为它应该转换它!我在这里错过了什么吗?

pim*_*314 44

如果您阅读文档,OneHotEncoder您将看到输入为fit"输入数组类型为int".因此,您需要为一个热编码数据执行两个步骤

from sklearn import preprocessing
cat_features = ['color', 'director_name', 'actor_2_name']
enc = preprocessing.LabelEncoder()
enc.fit(cat_features)
new_cat_features = enc.transform(cat_features)
print new_cat_features # [1 2 0]
new_cat_features = new_cat_features.reshape(-1, 1) # Needs to be the correct shape
ohe = preprocessing.OneHotEncoder(sparse=False) #Easier to read
print ohe.fit_transform(new_cat_features)
Run Code Online (Sandbox Code Playgroud)

输出:

[[ 0.  1.  0.]
 [ 0.  0.  1.]
 [ 1.  0.  0.]]
Run Code Online (Sandbox Code Playgroud)

编辑

由于0.20这成为更容易一点,不仅因为OneHotEncoder现在处理字符串好听,而且还因为我们可以轻松地将多列使用ColumnTransformer,请参阅下面的例子

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import LabelEncoder, OneHotEncoder
import numpy as np

X = np.array([['apple', 'red', 1, 'round', 0],
              ['orange', 'orange', 2, 'round', 0.1],
              ['bannana', 'yellow', 2, 'long', 0],
              ['apple', 'green', 1, 'round', 0.2]])
ct = ColumnTransformer(
    [('oh_enc', OneHotEncoder(sparse=False), [0, 1, 3]),],  # the column numbers I want to apply this to
    remainder='passthrough'  # This leaves the rest of my columns in place
)
print(ct2.fit_transform(X)) # Notice the output is a string
Run Code Online (Sandbox Code Playgroud)

输出:

[['1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '0.0' '0.0' '1.0' '1' '0']
 ['0.0' '0.0' '1.0' '0.0' '1.0' '0.0' '0.0' '0.0' '1.0' '2' '0.1']
 ['0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1.0' '0.0' '2' '0']
 ['1.0' '0.0' '0.0' '1.0' '0.0' '0.0' '0.0' '0.0' '1.0' '1' '0.2']]
Run Code Online (Sandbox Code Playgroud)

  • 我根本不明白这个答案。您在哪里将编码器与数据集中的数据拟合?您能否提供来自问题的数据集更详细的示例? (2认同)

小智 14

您可以使用LabelBinarizer类在一次镜头中应用两种转换(从文本类别到整数类别,然后从整数类别到单热矢量):

cat_features = ['color', 'director_name', 'actor_2_name']
encoder = LabelBinarizer()
new_cat_features = encoder.fit_transform(cat_features)
new_cat_features
Run Code Online (Sandbox Code Playgroud)

请注意,默认情况下,这会返回密集的NumPy数组.您可以通过将sparse_output = True传递给LabelBinarizer构造函数来获取稀疏矩阵.

来源动手机器学习与Scikit,学习和TensorFlow


Hap*_*ing 6

如果数据集在熊猫数据框中,则使用

pandas.get_dummies

会更直接。

*已从pandas.get_getdummies更正为pandas.get_dummies


Abh*_*kur 5

从文档中:

\n\n
categorical_features : \xe2\x80\x9call\xe2\x80\x9d or array of indices or mask\nSpecify what features are treated as categorical.\n\xe2\x80\x98all\xe2\x80\x99 (default): All features are treated as categorical.\narray of indices: Array of categorical feature indices.\nmask: Array of length n_features and with dtype=bool.\n
Run Code Online (Sandbox Code Playgroud)\n\n

pandas 数据框的列名不起作用。如果您的分类特征是列号 0、2 和 6,请使用:

\n\n
from sklearn import preprocessing\ncat_features = [0, 2, 6]\nenc = preprocessing.OneHotEncoder(categorical_features=cat_features)\nenc.fit(dataset.values)\n
Run Code Online (Sandbox Code Playgroud)\n\n

还必须注意的是,如果这些分类特征没有经过标签编码,那么LabelEncoder在使用之前需要对这些特征进行使用OneHotEncoder

\n