熊猫合并意外产生后缀

a11*_*a11 5 merge python-3.x pandas

我正在将两个 Pandas DataFrame 合并在一起,并获得“_x”和“_y”后缀。易于复制下面的示例。我尝试添加, suffixes=(False, False)到合并中,但它返回一个错误:ValueError: columns overlap but no suffix specified: Index(['f1', 'f2', 'f3'], dtype='object')。我一定在这里遗漏了一些明显的东西?我明白为什么使用 join 会发生这种情况,但我没想到它会用于合并。

请忽略复制切片错误。我想不通为什么它不扔在10号线这个错误,但它扔在第17行(如果你知道,有一个悬而未决的问题在这里就可以了!)

系统详细信息:Windows 10
conda 4.8.2
Python 3.8.3
pandas 1.0.5 py38he6e81aa_0 conda-forge

import pandas as pd

#### Build an example DataFrame for easy-to-replicate example ####
myid = [1, 1, 1, 2, 2]
myorder = [3, 2, 1, 2, 1]
y = [3642, 3640, 3632, 3628, 3608]
x = [11811, 11812, 11807, 11795, 11795]
df = pd.DataFrame(list(zip(myid, myorder, x, y)), 
                  columns =['myid', 'myorder', 'x', 'y']) 
df.sort_values(by=['myid', 'myorder'], inplace=True) #Line10
df.reset_index(drop=True, inplace=True)
display(df.style.hide_index())

### Typical analysis on existing DataFrame, Error occurs in here ####
for id in df.myid.unique():
    tempdf = df[mygdf.myid == id]
    tempdf.sort_values(by=['myid', 'myorder'], inplace=True) #Line17
    tempdf.reset_index(drop=True, inplace=True)
    for i, r in tempdf.iloc[1:].iterrows():
        ## in reality, calling a more complicated function here
        ## this is just a simple example
        tempdf.loc[i, 'f1'] = tempdf.x[i-1] - tempdf.x[i]
        tempdf.loc[i, 'f2'] = tempdf.y[i-1] - tempdf.y[i]
        tempdf.loc[i, 'f3'] = tempdf.y[i] +2
   
    what_i_care_about = ['myid', 'myorder', 'f1', 'f2', 'f3']

    df = pd.merge(df, tempdf[what_i_care_about], 
                  on=['myid', 'myorder'], how='outer')
    del tempdf

display(df.style.hide_index())
Run Code Online (Sandbox Code Playgroud)

在此处输入图片说明

Chr*_*per 7

您的问题是您没有合并的列对于两个源 DataFrame 都是通用的。Pandas 需要一种方法来说明哪个来自哪里,因此它添加了后缀,默认'_x'位于左侧和'_y'右侧。

如果您对保留列的源数据框有偏好,那么您可以设置后缀并相应地进行过滤,例如,如果您想保留左侧的冲突列:

# Label the two sides, with no suffix on the side you want to keep
df = pd.merge(
    df, 
    tempdf[what_i_care_about], 
    on=['myid', 'myorder'], 
    how='outer',
    suffixes=('', '_delme')  # Left gets no suffix, right gets something identifiable
)
# Discard the columns that acquired a suffix
df = df[[c for c in df.columns if not c.endswith('_delme')]]
Run Code Online (Sandbox Code Playgroud)

或者,您可以在合并之前删除每个冲突列中的一个,然后 Pandas 无需分配后缀。

  • 有道理,我将连接和合并混为一谈。我想我想要的是合并所有列,所以我应该删除 `on=...` 标准。谢谢 (2认同)