Kia*_*ann 5 python filter dataframe python-3.x pandas
如何过滤掉我不希望第一个字母为“Z”或任何其他字符的一系列数据(在 Pandas dataFrame 中)。
我有以下熊猫数据帧,df,(其中有 > 25,000 行)。
TIME_STAMP Activity Action Quantity EPIC Price Sub-activity Venue
0 2017-08-30 08:00:05.000 Allocation BUY 50 RRS 77.6 CPTY 066
1 2017-08-30 08:00:05.000 Allocation BUY 50 RRS 77.6 CPTY 066
3 2017-08-30 08:00:09.000 Allocation BUY 91 BATS 47.875 CPTY PXINLN
4 2017-08-30 08:00:10.000 Allocation BUY 43 PNN 8.07 CPTY WCAPD
5 2017-08-30 08:00:10.000 Allocation BUY 270 SGE 6.93 CPTY PROBDMAD
Run Code Online (Sandbox Code Playgroud)
我正在尝试删除 Venue 的第一个字母为“Z”的所有行。
例如,我通常的过滤器代码类似于(过滤掉 Venue = '066' 的所有行
df = df[df.Venue != '066']
Run Code Online (Sandbox Code Playgroud)
我可以看到这个过滤器行过滤掉了我需要的数组,但我不确定如何在过滤器上下文中指定它。
[k for k in df.Venue if 'Z' not in k]
Run Code Online (Sandbox Code Playgroud)
jez*_*ael 12
使用str[0]了选择第一个值或使用startswith,contains用正则表达式^的字符串的开始。对于 invertong boolen 掩码使用~:
df1 = df[df.Venue.str[0] != 'Z']
df1 = df[~df.Venue.str.startswith('Z')]
df1 = df[~df.Venue.str.contains('^Z')]
Run Code Online (Sandbox Code Playgroud)
如果没有NaN更快的 s 值,则使用列表理解:
df1 = df[[x[0] != 'Z' for x in df.Venue]]
df1 = df[[not x.startswith('Z') for x in df.Venue]]
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
9855 次 |
| 最近记录: |