将行与pandas中的"其他"组合

nt.*_*jin 5 python pandas

我有一个像这样的pandas数据帧:

  character  count
0         a    104
1         b     30
2         c    210
3         d     40
4         e    189
5         f     20
6         g     10
Run Code Online (Sandbox Code Playgroud)

我希望数据框中只有前3个字符,剩下的组合在一起,others因此表格变为:

  character  count
0         c    210
1         e    189
2         a    104
3    others    100
Run Code Online (Sandbox Code Playgroud)

我怎样才能做到这一点?

谢谢.

Max*_*axU 6

我们可以使用Series.nlargest()方法:

In [31]: new = df.nlargest(3, columns='count')

In [32]: new = pd.concat(
    ...:         [new,
    ...:          pd.DataFrame({'character':['others'],
    ...:                        'count':df.drop(new.index)['count'].sum()})
    ...:         ], ignore_index=True)
    ...:

In [33]: new
Out[33]:
  character  count
0         c    210
1         e    189
2         a    104
3    others     60
Run Code Online (Sandbox Code Playgroud)

或者不那么惯用的解决方案:

In [16]: new = df.nlargest(3, columns='count')

In [17]: new.loc[len(new)] = ['others', df.drop(new.index)['count'].sum()]

In [18]: new
Out[18]:
  character  count
2         c    210
4         e    189
0         a    104
3    others    100
Run Code Online (Sandbox Code Playgroud)

  • 只需添加`new.reset_index(inplace = True,drop = True)`即可获得完全匹配:) (2认同)