Ale*_*sev 5 python group-by dataframe pandas
我有以下数据框:
uniq_id value
2016-12-26 11:03:10 001 342
2016-12-26 11:03:13 004 5
2016-12-26 12:03:13 005 14
2016-12-26 12:03:13 008 114
2016-12-27 11:03:10 009 343
2016-12-27 11:03:13 013 5
2016-12-27 12:03:13 016 124
2016-12-27 12:03:13 018 114
Run Code Online (Sandbox Code Playgroud)
而且我需要获得按值排序的每一天的前N条记录。像这样(对于N = 2):
2016-12-26 001 342
008 114
2016-12-27 009 343
016 124
Run Code Online (Sandbox Code Playgroud)
请在熊猫0.19.x中建议正确的方法
不幸的是,目前还没有这样的方法DataFrameGroupBy.nlargest(),它允许我们执行以下操作:
df.groupby(...).nlargest(2, columns=['value'])
Run Code Online (Sandbox Code Playgroud)
所以这是一个有点丑陋但有效的解决方案:
In [73]: df.set_index(df.index.normalize()).reset_index().sort_values(['index','value'], ascending=[1,0]).groupby('index').head(2)
Out[73]:
index uniq_id value
0 2016-12-26 1 342
3 2016-12-26 8 114
4 2016-12-27 9 343
6 2016-12-27 16 124
Run Code Online (Sandbox Code Playgroud)
PS我想一定有更好的...
更新:如果您的 DF 没有重复的索引值,则以下解决方案也应该有效:
In [117]: df
Out[117]:
uniq_id value
2016-12-26 11:03:10 1 342
2016-12-26 11:03:13 4 5
2016-12-26 12:03:13 5 14
2016-12-26 12:33:13 8 114 # <-- i've intentionally changed this index value
2016-12-27 11:03:10 9 343
2016-12-27 11:03:13 13 5
2016-12-27 12:03:13 16 124
2016-12-27 12:33:13 18 114 # <-- i've intentionally changed this index value
In [118]: df.groupby(pd.TimeGrouper('D')).apply(lambda x: x.nlargest(2, 'value')).reset_index(level=1, drop=1)
Out[118]:
uniq_id value
2016-12-26 1 342
2016-12-26 8 114
2016-12-27 9 343
2016-12-27 16 124
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
1875 次 |
| 最近记录: |