包含列表的Pandas列的get_dummies

red*_*gem 1 python pandas data-science

鉴于我有一个DataFrame,其列包含字符串列表,如下所示:

    Name    Fruit
0   Curly   [Apple]
1   Moe     [Orange]
2   Larry   [Apple, Banana]
Run Code Online (Sandbox Code Playgroud)

我怎么把它变成这样的东西?

    Name     Fruit_Apple   Fruit_Orange   Fruit_Banana
0   Curly              1              0              0
1   Moe                0              1              0
2   Larry              1              0              1
Run Code Online (Sandbox Code Playgroud)

我有一种感觉,我会以某种方式使用,pandas.get_dummies()但我似乎无法得到它.有帮助吗?

Jar*_*rad 5

import pandas as pd

df = pd.DataFrame({'Name': ['Curly', 'Moe', 'Larry'],
                   'Fruit': [['Apple'], ['Orange'], ['Apple', 'Banana']]},
                  columns=['Name', 'Fruit'])

# a one-liner... that's pretty long    
dummies_df = pd.get_dummies(
  df.join(pd.Series(df['Fruit'].apply(pd.Series).stack().reset_index(1, drop=True),
                    name='Fruit1')).drop('Fruit', axis=1).rename(columns={'Fruit1': 'Fruit'}),
  columns=['Fruit']).groupby('Name', as_index=False).sum()

print(dummies_df)
Run Code Online (Sandbox Code Playgroud)

我会把它分解为几步:

步骤1:

df['Fruit'].apply(pd.Series).stack().reset_index(1, drop=True)

此步骤适用pd.Series于将列表中的每个项目拆分为新列的列表.stack然后将这些列堆叠成一列,同时保持重要的索引信息.该reset_index部分重置索引的级别1并删除它,因为它不需要.你最终得到这个:

0     Apple
1    Orange
2     Apple
2    Banana
dtype: object 
Run Code Online (Sandbox Code Playgroud)

第2步:

你会注意到pd.Series( *Step 1 here*, name='Fruit1')上面的第1步代码,因为我们接下来会将这个系列加入到现有的数据帧中,所以我们需要a name才能做到这一点.

第3步:

df.join(* steps 1 and 2 code *).drop('Fruit', axis=1).rename(columns={'Fruit1': 'Fruit'})
Run Code Online (Sandbox Code Playgroud)

由于我们现在有一个pd.Series带有name(Fruit1)的Fruit1系列,我们将系列连接到原始系列df,然后有三列.然后我们调用drop来删除原始Fruit列.现在我们只有两列Name,Fruit1但我们想要Fruit命名,Fruit所以我们用它重命名rename.

第4步:

pd.get_dummies(* steps 1, 2, and 3 here*, columns=['Fruit'])
Run Code Online (Sandbox Code Playgroud)

在这里,我们最终调用了get_dummies,我们使用它columns=['Fruit']来专门告诉get_dummies只能获取Fruit列的假人.

    Name  Fruit_Apple  Fruit_Banana  Fruit_Orange
0  Curly          1.0           0.0           0.0
1    Moe          0.0           0.0           1.0
2  Larry          1.0           0.0           0.0
2  Larry          0.0           1.0           0.0
Run Code Online (Sandbox Code Playgroud)

第5步:

dummies_df = (*steps 1, 2, 3, and 4*).groupby('Name', as_index=False).sum()
Run Code Online (Sandbox Code Playgroud)

最后,您groupby在Name列上使用a 并指定as_index=False可选地不将其设置Name为索引.然后将结果与之相加.sum()

最终结果:

    Name  Fruit_Apple  Fruit_Banana  Fruit_Orange
0  Curly          1.0           0.0           0.0
1  Larry          1.0           1.0           0.0
2    Moe          0.0           0.0           1.0
Run Code Online (Sandbox Code Playgroud)

  • 您可以使用pd.crosstab(df_unwind ['Name'],df_unwind ['Fruit1'])替换步骤4和5.df_unwind是步骤3之后的df.Ther不是单行而是更短. (2认同)