PySpark将列中的null替换为其他列中的值

Lui*_*eal 11 python apache-spark pyspark

我想用一个相邻列中的值替换一列中的空值,例如,如果我有

A|B
0,1
2,null
3,null
4,2
Run Code Online (Sandbox Code Playgroud)

我希望它是:

A|B
0,1
2,2
3,3
4,2
Run Code Online (Sandbox Code Playgroud)

试过

df.na.fill(df.A,"B")
Run Code Online (Sandbox Code Playgroud)

但是没有用,它说值应该是float,int,long,string或dict

有任何想法吗?

Lui*_*eal 21

最后找到了另一种选择:

df.withColumn("B",coalesce(df.B,df.A)) 
Run Code Online (Sandbox Code Playgroud)

  • 该解决方案缺少 **from pyspark.sql.functions import coalesce** (3认同)

Rag*_*ags 5

另一个答案。

如果df1您的数据框下方

rd1 = sc.parallelize([(0,1), (2,None), (3,None), (4,2)])
df1 = rd1.toDF(['A', 'B'])

from pyspark.sql.functions import when
df1.select('A',
           when( df1.B.isNull(), df1.A).otherwise(df1.B).alias('B')
          )\
   .show()
Run Code Online (Sandbox Code Playgroud)