从熊猫中的假人重建一个分类变量

the*_*rgo 32 python pandas

pd.get_dummies允许将分类变量转换为虚拟变量.除了重建分类变量是微不足道的事实之外,还有一种首选/快速的方法吗?

Nat*_*han 36

这已经有几年了,所以pandas当最初提出这个问题时,这可能不会出现在工具包中,但这种方法对我来说似乎有点容易.idxmax将返回对应于最大元素的索引(即带有a的索引1).我们这样做axis=1是因为我们想要发生的列名称1.

编辑:我没有打扰它是分类而不仅仅是一个字符串,但你可以像@Jeff那样用它包装它pd.Categorical(和pd.Series,如果需要).

In [1]: import pandas as pd

In [2]: s = pd.Series(['a', 'b', 'a', 'c'])

In [3]: s
Out[3]: 
0    a
1    b
2    a
3    c
dtype: object

In [4]: dummies = pd.get_dummies(s)

In [5]: dummies
Out[5]: 
   a  b  c
0  1  0  0
1  0  1  0
2  1  0  0
3  0  0  1

In [6]: s2 = dummies.idxmax(axis=1)

In [7]: s2
Out[7]: 
0    a
1    b
2    a
3    c
dtype: object

In [8]: (s2 == s).all()
Out[8]: True
Run Code Online (Sandbox Code Playgroud)

编辑以回应@ piRSquared的评论:这个解决方案确实假设1每行有一个.我认为这通常是一种格式.pd.get_dummies如果你有drop_first=True或者有NaN值,那么可以返回全0的行和dummy_na=False(默认)(我丢失的任何情况?).一行全零将被视为它是第一列中命名的变量的实例(例如a,在上面的示例中).

如果drop_first=True,你无法从虚拟数据框中知道"第一"变量的名称是什么,那么除非你保留额外的信息,否则操作是不可逆的; 我建议离开drop_first=False(默认).

由于dummy_na=False是默认值,这肯定会导致问题.如果您想使用此解决方案来反转"dummification"并且您的数据包含任何内容,请dummy_na=True在致电时进行设置.pd.get_dummiesNaNs设置dummy_na=True将始终添加"nan"列,即使该列全为0,因此您可能不希望设置此列,除非您实际拥有NaNs.一个很好的方法可能是设置dummies = pd.get_dummies(series, dummy_na=series.isnull().any()).同样好的是,idxmax解决方案将正确地重新生成您的NaNs(不仅仅是一个表示"nan"的字符串).

还值得一提的是,设置drop_first=True和dummy_na=False表示NaNs与第一个变量的实例无法区分,因此如果您的数据集可能包含任何NaN值,则强烈建议不要这样做.

  • 如果行全为零,则会失败.它适用于此示例,并假设每行只存在一个"1"值. (2认同)

Jef*_*eff 19

In [46]: s = Series(list('aaabbbccddefgh')).astype('category')

In [47]: s
Out[47]: 
0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
dtype: category
Categories (8, object): [a < b < c < d < e < f < g < h]

In [48]: df = pd.get_dummies(s)

In [49]: df
Out[49]: 
    a  b  c  d  e  f  g  h
0   1  0  0  0  0  0  0  0
1   1  0  0  0  0  0  0  0
2   1  0  0  0  0  0  0  0
3   0  1  0  0  0  0  0  0
4   0  1  0  0  0  0  0  0
5   0  1  0  0  0  0  0  0
6   0  0  1  0  0  0  0  0
7   0  0  1  0  0  0  0  0
8   0  0  0  1  0  0  0  0
9   0  0  0  1  0  0  0  0
10  0  0  0  0  1  0  0  0
11  0  0  0  0  0  1  0  0
12  0  0  0  0  0  0  1  0
13  0  0  0  0  0  0  0  1

In [50]: x = df.stack()

# I don't think you actually need to specify ALL of the categories here, as by definition
# they are in the dummy matrix to start (and hence the column index)
In [51]: Series(pd.Categorical(x[x!=0].index.get_level_values(1)))
Out[51]: 
0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
Name: level_1, dtype: category
Categories (8, object): [a < b < c < d < e < f < g < h]
Run Code Online (Sandbox Code Playgroud)

所以我认为我们需要一个"做"这个功能,因为它似乎是一个自然的操作.也许get_categories(),看到这里


sac*_*cuL 9

这是一个相当晚的答案,但由于你要求快速做到这一点,我认为你正在寻找最高效的策略.在大型数据帧(例如10000行)上,通过使用np.where而不是idxmax或get_level_values获得相同的结果,可以获得非常显着的速度提升.我们的想法是索引虚拟数据帧不为0的列名:

方法:

使用与@Nathan相同的样本数据:

>>> dummies
   a  b  c
0  1  0  0
1  0  1  0
2  1  0  0
3  0  0  1

s2 = pd.Series(dummies.columns[np.where(dummies!=0)[1]])

>>> s2
0    a
1    b
2    a
3    c
dtype: object
Run Code Online (Sandbox Code Playgroud)

基准测试:

在一个小的虚拟数据帧上,您将看不到性能上的太大差异.但是,测试不同的策略来解决大型系列中的这个问题:

s = pd.Series(np.random.choice(['a','b','c'], 10000))

dummies = pd.get_dummies(s)

def np_method(dummies=dummies):
    return pd.Series(dummies.columns[np.where(dummies!=0)[1]])

def idx_max_method(dummies=dummies):
    return dummies.idxmax(axis=1)

def get_level_values_method(dummies=dummies):
    x = dummies.stack()
    return pd.Series(pd.Categorical(x[x!=0].index.get_level_values(1)))

def dot_method(dummies=dummies):
    return dummies.dot(dummies.columns)

import timeit

# Time each method, 1000 iterations each:

>>> timeit.timeit(np_method, number=1000)
1.0491090340074152

>>> timeit.timeit(idx_max_method, number=1000)
12.119140846014488

>>> timeit.timeit(get_level_values_method, number=1000)
4.109266621991992

>>> timeit.timeit(dot_method, number=1000)
1.6741622970002936
Run Code Online (Sandbox Code Playgroud)

该np.where方法比get_level_values方法快4倍,比方法快11.5倍idxmax!它也打败了(但只是一点点)这个答案中.dot()概述的类似问题的方法

它们都返回相同的结果:

>>> (get_level_values_method() == np_method()).all()
True
>>> (idx_max_method() == np_method()).all()
True
Run Code Online (Sandbox Code Playgroud)


piR*_*red 5

设置

使用@Jeff 的设置

s = Series(list('aaabbbccddefgh')).astype('category')
df = pd.get_dummies(s)
Run Code Online (Sandbox Code Playgroud)

如果列是字符串

1每行只有一个

df.dot(df.columns)

0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
dtype: object
Run Code Online (Sandbox Code Playgroud)

numpy.where

再次!假设1每行只有一个

i, j = np.where(df)
pd.Series(df.columns[j], i)

0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
dtype: category
Categories (8, object): [a, b, c, d, e, f, g, h]
Run Code Online (Sandbox Code Playgroud)

numpy.where

不假设1每行一个

i, j = np.where(df)
pd.Series(dict(zip(zip(i, j), df.columns[j])))

0   0    a
1   0    a
2   0    a
3   1    b
4   1    b
5   1    b
6   2    c
7   2    c
8   3    d
9   3    d
10  4    e
11  5    f
12  6    g
13  7    h
dtype: object
Run Code Online (Sandbox Code Playgroud)

numpy.where

我们不假设1每行一个,我们删除索引

i, j = np.where(df)
pd.Series(dict(zip(zip(i, j), df.columns[j]))).reset_index(-1, drop=True)

0     a
1     a
2     a
3     b
4     b
5     b
6     c
7     c
8     d
9     d
10    e
11    f
12    g
13    h
dtype: object
Run Code Online (Sandbox Code Playgroud)