小编vis*_*ram的帖子

PySpark dataframe to_json()函数

我有一个如下数据框,

>>> df.show(10,False)
+-----+----+---+------+
|id   |name|age|salary|
+-----+----+---+------+
|10001|alex|30 |75000 |
|10002|bob |31 |80000 |
|10003|deb |31 |80000 |
|10004|john|33 |85000 |
|10005|sam |30 |75000 |
+-----+----+---+------+
Run Code Online (Sandbox Code Playgroud)

将df的整个行转换为一个新列“ jsonCol”,

>>> newDf1 = df.withColumn("jsonCol", to_json(struct([df[x] for x in df.columns])))
>>> newDf1.show(10,False)
+-----+----+---+------+--------------------------------------------------------+
|id   |name|age|salary|jsonCol                                                 |
+-----+----+---+------+--------------------------------------------------------+
|10001|alex|30 |75000 |{"id":"10001","name":"alex","age":"30","salary":"75000"}|
|10002|bob |31 |80000 |{"id":"10002","name":"bob","age":"31","salary":"80000"} |
|10003|deb |31 |80000 |{"id":"10003","name":"deb","age":"31","salary":"80000"} |
|10004|john|33 |85000 |{"id":"10004","name":"john","age":"33","salary":"85000"}|
|10005|sam |30 |75000 |{"id":"10005","name":"sam","age":"30","salary":"75000"} |
+-----+----+---+------+--------------------------------------------------------+
Run Code Online (Sandbox Code Playgroud)

我不需要像上述步骤那样将整个行转换为JSON字符串,而是需要一种解决方案,仅根据字段的值选择几列。我在以下命令中提供了示例条件。

但是当我开始使用when函数时,结果JSON字符串的列名(键)消失了。仅按位置获取列名,而不是实际的列名(键)

>>> newDf2 = df.withColumn("jsonCol", to_json(struct([ when(col(x)!="  ",df[x]).otherwise(None) for x in df.columns])))
>>> …
Run Code Online (Sandbox Code Playgroud)

apache-spark apache-spark-sql pyspark

5
推荐指数
1
解决办法
2497
查看次数

标签 统计

apache-spark ×1

apache-spark-sql ×1

pyspark ×1