我有以下名为ttm的数据框:
usersidid clienthostid eventSumTotal LoginDaysSum score
0 12 1 60 3 1728
1 11 1 240 3 1331
3 5 1 5 3 125
4 6 1 16 2 216
2 10 3 270 3 1000
5 8 3 18 2 512
Run Code Online (Sandbox Code Playgroud)
当我做
ttm.groupby(['clienthostid'], as_index=False, sort=False)['LoginDaysSum'].count()
Run Code Online (Sandbox Code Playgroud)
我得到了我的预期(虽然我希望结果在一个名为'ratio'的新标签下):
clienthostid LoginDaysSum
0 1 4
1 3 2
Run Code Online (Sandbox Code Playgroud)
但是,当我这样做
ttm.groupby(['clienthostid'], as_index=False, sort=False)['LoginDaysSum'].apply(lambda x: x.iloc[0] / x.iloc[1])
Run Code Online (Sandbox Code Playgroud)
我明白了:
0 1.0
1 1.5
Run Code Online (Sandbox Code Playgroud)
谢谢,
对于返回DataFrame后groupby有两种可能的解决方案:
参数as_index=False是什么在起作用漂亮count,sum,mean功能
reset_index从更高级别的index解决方案创建新列
df = ttm.groupby(['clienthostid'], as_index=False, sort=False)['LoginDaysSum'].count()
print (df)
clienthostid LoginDaysSum
0 1 4
1 3 2
Run Code Online (Sandbox Code Playgroud)
df = ttm.groupby(['clienthostid'], sort=False)['LoginDaysSum'].count().reset_index()
print (df)
clienthostid LoginDaysSum
0 1 4
1 3 2
Run Code Online (Sandbox Code Playgroud)
对于第二个需要删除as_index=False,而是添加reset_index:
#output is `Series`
a = ttm.groupby(['clienthostid'], sort=False)['LoginDaysSum'] \
.apply(lambda x: x.iloc[0] / x.iloc[1])
print (a)
clienthostid
1 1.0
3 1.5
Name: LoginDaysSum, dtype: float64
print (type(a))
<class 'pandas.core.series.Series'>
print (a.index)
Int64Index([1, 3], dtype='int64', name='clienthostid')
df1 = ttm.groupby(['clienthostid'], sort=False)['LoginDaysSum']
.apply(lambda x: x.iloc[0] / x.iloc[1]).reset_index(name='ratio')
print (df1)
clienthostid ratio
0 1 1.0
1 3 1.5
Run Code Online (Sandbox Code Playgroud)
为什么有些专栏不见了?
我认为可能会有问题自动排除滋扰列:
#convert column to str
ttm.usersidid = ttm.usersidid.astype(str) + 'aa'
print (ttm)
usersidid clienthostid eventSumTotal LoginDaysSum score
0 12aa 1 60 3 1728
1 11aa 1 240 3 1331
3 5aa 1 5 3 125
4 6aa 1 16 2 216
2 10aa 3 270 3 1000
5 8aa 3 18 2 512
#removed str column userid
a = ttm.groupby(['clienthostid'], sort=False).sum()
print (a)
eventSumTotal LoginDaysSum score
clienthostid
1 321 11 3400
3 288 5 1512
Run Code Online (Sandbox Code Playgroud)
count是groupby对象的内置方法,pandas 知道如何处理它。还指定了另外两件事来确定输出的样子。
# For a built in method, when
# you don't want the group column
# as the index, pandas keeps it in
# as a column.
# |----||||----|
ttm.groupby(['clienthostid'], as_index=False, sort=False)['LoginDaysSum'].count()
clienthostid LoginDaysSum
0 1 4
1 3 2
Run Code Online (Sandbox Code Playgroud)
# For a built in method, when
# you do want the group column
# as the index, then...
# |----||||---|
ttm.groupby(['clienthostid'], as_index=True, sort=False)['LoginDaysSum'].count()
# |-----||||-----|
# the single brackets tells
# pandas to operate on a series
# in this case, count the series
clienthostid
1 4
3 2
Name: LoginDaysSum, dtype: int64
Run Code Online (Sandbox Code Playgroud)
ttm.groupby(['clienthostid'], as_index=True, sort=False)[['LoginDaysSum']].count()
# |------||||------|
# the double brackets tells pandas
# to operate on the dataframe
# specified by these columns and will
# return a dataframe
LoginDaysSum
clienthostid
1 4
3 2
Run Code Online (Sandbox Code Playgroud)
当你使用applyPandas 时,不再知道如何处理组列时你说as_index=False. 它必须相信,如果你使用apply你想要返回的正是你所说的返回,所以它只会把它扔掉。此外,您的列周围有单个括号,表示对系列进行操作。相反,用于as_index=True将分组列信息保留在索引中。然后使用 areset_index将其从索引传输回数据帧。在这一点上,您使用单括号无关紧要,因为之后reset_index您将再次拥有一个数据框。
ttm.groupby(['clienthostid'], as_index=True, sort=False)['LoginDaysSum'].apply(lambda x: x.iloc[0] / x.iloc[1])
0 1.0
1 1.5
dtype: float64
Run Code Online (Sandbox Code Playgroud)
ttm.groupby(['clienthostid'], as_index=True, sort=False)['LoginDaysSum'].apply(lambda x: x.iloc[0] / x.iloc[1]).reset_index()
clienthostid LoginDaysSum
0 1 1.0
1 3 1.5
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
11526 次 |
| 最近记录: |