删除 pyspark 数据框中的连续重复项

Question

删除 pyspark 数据框中的连续重复项

Qub*_*bix 4 apache-spark apache-spark-sql pyspark

有一个像这样的数据框：

## +---+---+
## | id|num|
## +---+---+
## |  2|3.0|
## |  3|6.0|
## |  3|2.0|
## |  3|1.0|
## |  2|9.0|
## |  4|7.0|
## +---+---+

Run Code Online (Sandbox Code Playgroud)

我想删除连续的重复，并获得：

## +---+---+
## | id|num|
## +---+---+
## |  2|3.0|
## |  3|6.0|
## |  2|9.0|
## |  4|7.0|
## +---+---+

Run Code Online (Sandbox Code Playgroud)

我在 Pandas 中找到了执行此操作的方法，但在 Pyspark 中却没有找到。

Answer 1

gaw*_*gaw 5

The answer should work as you desired, however there might be room for some optimization:

from pyspark.sql.window import Window as W
test_df = spark.createDataFrame([
    (2,3.0),(3,6.0),(3,2.0),(3,1.0),(2,9.0),(4,7.0)
    ], ("id", "num"))
test_df = test_df.withColumn("idx", monotonically_increasing_id())  # create temporary ID because window needs an ordered structure
w = W.orderBy("idx")
get_last= when(lag("id", 1).over(w) == col("id"), False).otherwise(True) # check if the previous row contains the same id

test_df.withColumn("changed",get_last).filter(col("changed")).select("id","num").show() # only select the rows with a changed ID

Run Code Online (Sandbox Code Playgroud)

Output:

+---+---+
| id|num|
+---+---+
|  2|3.0|
|  3|6.0|
|  2|9.0|
|  4|7.0|
+---+---+

Run Code Online (Sandbox Code Playgroud)

归档时间：	7 年，5 月前
查看次数：	1439 次
最近记录：	7 年，5 月前