use*_*836 1 pandas pandas-groupby
这类似于将计算列附加到现有数据框,但是,当在pandas v0.14中按多个列进行分组时,该解决方案不起作用.
例如:
$ df = pd.DataFrame([
[1, 1, 1],
[1, 2, 1],
[1, 2, 2],
[1, 3, 1],
[2, 1, 1]],
columns=['id', 'country', 'source'])
Run Code Online (Sandbox Code Playgroud)
以下计算有效:
$ df.groupby(['id','country'])['source'].apply(lambda x: x.unique().tolist())
0 [1]
1 [1, 2]
2 [1, 2]
3 [1]
4 [1]
Name: source, dtype: object
Run Code Online (Sandbox Code Playgroud)
但是将输出分配给新列会导致错误:
df['source_list'] = df.groupby(['id','country'])['source'].apply(
lambda x: x.unique().tolist())
Run Code Online (Sandbox Code Playgroud)
TypeError:带有帧索引的插入列的不兼容索引
将分组结果与初始DataFrame合并:
>>> df1 = df.groupby(['id','country'])['source'].apply(
lambda x: x.tolist()).reset_index()
>>> df1
id country source
0 1 1 [1.0]
1 1 2 [1.0, 2.0]
2 1 3 [1.0]
3 2 1 [1.0]
>>> df2 = df[['id', 'country']]
>>> df2
id country
1 1 1
2 1 2
3 1 2
4 1 3
5 2 1
>>> pd.merge(df1, df2, on=['id', 'country'])
id country source
0 1 1 [1.0]
1 1 2 [1.0, 2.0]
2 1 2 [1.0, 2.0]
3 1 3 [1.0]
4 2 1 [1.0]
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
5289 次 |
| 最近记录: |