cs9*_*s95 4 python replace pandas
通常,我需要对Series或DataFrames列中的数据执行某种替换或替换操作。
例如,给定一系列字符串,
s = pd.Series(['foo', 'another foo bar', 'baz'])
0 foo
1 another foo bar
2 baz
dtype: object
Run Code Online (Sandbox Code Playgroud)
目标是将所有出现的“ foo”替换为“ bar”,以获取
0 bar
1 another bar bar
2 baz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
目前,我通常很困惑,因为可以使用两种方法来解决此问题:replace和str.replace。我不知道哪种方法是正确的,或者它们之间有什么区别(如果有的话),这使我感到困惑。
replace和之间的主要区别是str.replace什么,使用这两种方法的优点/缺点是什么?
cs9*_*s95 13
跳至TLDR;在此答案的底部,简要介绍了差异。
如果您从实用性的角度考虑这两种方法,就很容易理解它们之间的区别。
.str.replace是一种具有非常特定用途的方法-对字符串数据执行字符串或正则表达式替换。
OTOH .replace更像是一种通用的瑞士军刀,可以用任何其他东西代替任何东西(是的,这包括字符串和正则表达式)。
考虑下面的简单DataFrame,这将构成我们即将进行的讨论的基础。
# Setup
df = pd.DataFrame({
'A': ['foo', 'another foo bar', 'baz'],
'B': [0, 1, 0]
})
df
A B
0 foo 0
1 another foo bar 1
2 baz 0
Run Code Online (Sandbox Code Playgroud)
这两个功能之间的主要区别可以归纳为
使用str.replace对一个字符串列串替换,并replace在一个或多个列任何一般更换。
该文档市场str.replace为“简单的字符串替换”的方法,所以在执行字符串/正则表达式替换上熊猫系列或它的列认为是“矢量化”等同于Python的字符串时,这应该是您的第一选择replace()功能(或者re.sub()是更准确的)。
# simple substring replacement
df['A'].str.replace('foo', 'bar', regex=False)
0 bar
1 another bar bar
2 baz
Name: A, dtype: object
# simple regex replacement
df['A'].str.replace('ba.', 'xyz')
0 foo
1 another foo xyz
2 xyz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
replace适用于字符串替换和非字符串替换。而且,它还意味着一次**可以处理多个列(如果您需要在整个DataFrame中替换值,则也可以replace作为DataFrame方法来访问df.replace()。
# DataFrame-wide replacement
df.replace({'foo': 'bar', 1: -1})
A B
0 bar 0
1 another foo bar -1
2 baz 0
Run Code Online (Sandbox Code Playgroud)
str.replace一次可以更换一件东西。replace使您可以执行多次独立替换,即一次替换很多东西。
您只能为指定一个子字符串或正则表达式模式str.replace。repl可以是可调用的(请参见文档),因此使用regex可以发挥创意,可以在某种程度上模拟多个子字符串的替换,但是这些解决方案充其量是不可靠的。
一种常见的pandaic(可紧急,潘多尼克)模式是使用str.replace正则表达式OR管道通过管道分隔子字符串来删除多个不需要的子字符串|,并且替换字符串为''(空字符串)。
replace当您使用repl2 形式的多个独立替换项时,应该首选。有多种方法可以指定独立的替换项(列表,系列,字典等)。请参阅文档。{'pat1': 'repl1', 'pat2':, ...}
为了说明差异,
df['A'].str.replace('foo', 'text1').str.replace('bar', 'text2')
0 text1
1 another text1 text2
2 baz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
更好地表示为
df['A'].replace({'foo': 'text1', 'bar': 'text2'}, regex=True)
0 text1
1 another text1 text2
2 baz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
在字符串操作的上下文中,str.replace默认情况下启用正则表达式替换。replace除非使用此regex=True开关,否则仅执行完全匹配。
您所做的一切str.replace,也可以做到replace。但是,重要的是要注意两种方法的默认行为之间的以下差异。
str.replace将替换每次出现的子字符串,replace默认情况下仅执行整个单词匹配str.replace除非指定,否则将第一个参数解释为正则表达式regex=False。replace恰恰相反。对比之间的区别
df['A'].replace('foo', 'bar')
0 bar
1 another foo bar
2 baz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
和
df['A'].replace('foo', 'bar', regex=True)
0 bar
1 another bar bar
2 baz
Name: A, dtype: object
Run Code Online (Sandbox Code Playgroud)
还值得一提的是,您只能在时执行字符串替换regex=True。因此,例如,df.replace({'foo': 'bar', 1: -1}, regex=True)它将是无效的。
总而言之,主要区别是
目的。使用
str.replace对一个字符串列串替换,并replace在一个或多个列任何一般更换。用法。
str.replace一次可以更换一件东西。replace使您可以执行多次独立替换,即一次替换很多东西。默认行为。
str.replace默认情况下启用正则表达式替换。replace除非使用此regex=True开关,否则仅执行完全匹配。