iLo*_*eng 5 python data-manipulation dataframe pandas data-science
我有一个pandas数据帧如下:
user_id product_id order_number
1 1 1
1 1 2
1 1 3
1 2 1
1 2 5
2 1 1
2 1 3
2 1 4
2 1 5
3 1 1
3 1 2
3 1 6
Run Code Online (Sandbox Code Playgroud)
我想查询这个df的最长条纹(没有跳过order_number)和最后一条条纹(自上一个order_number以来).
理想的结果如下:
user_id product_id longest_streak last_streak
1 1 3 3
1 2 0 0
2 1 3 3
3 1 2 0
Run Code Online (Sandbox Code Playgroud)
我很欣赏这方面的任何见解.
用一个循环和defaultdict
a = defaultdict(lambda:None)
longest = defaultdict(int)
current = defaultdict(int)
for i, j, k in df.itertuples(index=False):
if a[(i, j)] == k - 1:
current[(i, j)] += 1 if current[(i, j)] else 2
longest[(i, j)] = max(longest[(i, j)], current[(i, j)])
else:
current[(i, j)] = 0
longest[(i, j)] |= 0
a[(i, j)] = k
pd.concat(
[pd.Series(d) for d in [longest, current]],
axis=1, keys=['longest_streak', 'last_streak']
).rename_axis(['user_id', 'product_id']).reset_index()
user_id product_id longest_streak last_streak
0 1 1 3 3
1 1 2 0 0
2 2 1 3 3
3 3 1 2 0
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
425 次 |
| 最近记录: |