零填充熊猫列

Ale*_*lli 0 python numpy pandas

我有以下数据框,其中col_1是整数类型:

print(df)

col_1 
100
200
00153
00164
Run Code Online (Sandbox Code Playgroud)

如果位数等于3,我想添加两个零。

final_col
00100
00200
00153
00164
Run Code Online (Sandbox Code Playgroud)

我尝试过:

df.col_1 = df.col_1.astype(int).astype(str)

df["final_col"] = np.where(len(df["col_1"]) == 3, "00" + df.col_1, df.col_1 )
Run Code Online (Sandbox Code Playgroud)

但是,它不会产生预期的输出(满足条件时,不会将两位数字相加)。

我该如何解决?

Chr*_*s A 7

使用str.zfill:

df['final_col'] = df['col_1'].astype(str).str.zfill(5)
Run Code Online (Sandbox Code Playgroud)

[出去]

   final_col
0      00100
1      00200
2      00153
3      00164
Run Code Online (Sandbox Code Playgroud)

更新,如果您只想准确填充 len 为 3 的位置,请使用感谢@yatu 指出:Series.where

df.col_1.where(df.col_1.str.len().ne(3),
               df.col_1.astype(str).str.zfill(5))
Run Code Online (Sandbox Code Playgroud)


ank*_*_91 5

另一种使用方式series.str.pad():

df.col_1.astype(str).str.pad(5,fillchar='0')
Run Code Online (Sandbox Code Playgroud)
0    00100
1    00200
2    00153
3    00164
Run Code Online (Sandbox Code Playgroud)

您的解决方案应更新为:

(np.where(df["col_1"].astype(str).str.len()==3, 
       "00" + df["col_1"].astype(str),df["col_1"].astype(str)))
Run Code Online (Sandbox Code Playgroud)

但这在字符串的长度小于5且不等于3时不起作用,因此我建议您不要使用它。


vra*_*a95 1

# after converting it to str , you can foolow up list comprehension.\n\ndf=pd.DataFrame({'col':['100','200','00153','00164']})\ndf['col_up']=['00'+x if len(x)==3 else x for x in df.col ]\ndf\n\n###output\n\n    col    col_up\n0   100     00100\n1   200     00200\n2   00153   00153\n3   00164   00164\n\n\n    ### based on the responses in comments \n  %%timeit -n 10000\n df.col.str.pad(5,fillchar='0') \n142 \xc2\xb5s \xc2\xb1 5.47 \xc2\xb5s per loop (mean \xc2\xb1 std. dev. of 7 runs, 10000 loops each)\n\n\n     %%timeit -n 10000\n ['00'+x if len(x)==3 else x for x in df.col ]\n21.1 \xc2\xb5s \xc2\xb1 952 ns per loop (mean \xc2\xb1 std. dev. of 7 runs, 10000 loops each)\n\n  %%timeit -n 10000\n  df.col.astype(str).str.pad(5,fillchar='0')\n243 \xc2\xb5s \xc2\xb1 7.02 \xc2\xb5s per loop (mean \xc2\xb1 std. dev. of 7 runs, 10000 loops each)\n
Run Code Online (Sandbox Code Playgroud)\n

  • :/ 这是实现这一目标的奇怪方法。我认为 str.pad 和 str.zfill 方法是更正常的方法。 (4认同)
  • @vrana95 另外,您的测试并不等效,您已经 astype str 其中一种方法,但您的解决方案假设它们以字符串开头,并且您没有考虑将一系列写入 df 与写入列表之间的区别到 df - 但这里最重要的问题并不是真正的时间问题,而是可读性和风格的问题。我知道这不一定是您自己实现这一目标的方式,并且您编写它是为了提供替代方案,我只是认为提供更糟糕的替代方案不是一个好主意。不过我应该回去工作了! (3认同)
  • 这不具有代表性。尝试使用更大的数据框 (2认同)