小编Anm*_*ave的帖子

如何从pyspark中的数组中提取元素

我有一个以下类型的数据框

col1|col2|col3|col4
xxxx|yyyy|zzzz|[1111],[2222]
Run Code Online (Sandbox Code Playgroud)

我希望我的输出是跟随类型

col1|col2|col3|col4|col5
xxxx|yyyy|zzzz|1111|2222
Run Code Online (Sandbox Code Playgroud)

我的col4是一个数组,我想将它转换为一个单独的列.需要做什么?

我在flatmap中看到很多答案,但是他们正在增加一行,我希望只将元组放在另一列但是在同一行中

以下是我的实际架构:

root
 |-- PRIVATE_IP: string (nullable = true)
 |-- PRIVATE_PORT: integer (nullable = true)
 |-- DESTINATION_IP: string (nullable = true)
 |-- DESTINATION_PORT: integer (nullable = true)
 |-- collect_set(TIMESTAMP): array (nullable = true)
 |    |-- element: string (containsNull = true)
Run Code Online (Sandbox Code Playgroud)

也可以请一些人帮我解释数据帧和RDD

python apache-spark rdd pyspark

8
推荐指数
2
解决办法
2万
查看次数

根据 pyspark 中的条件合并 Spark 中的两行

我有以下格式的输入记录: 输入数据格式

在此输入图像描述

我希望数据按照以下格式进行转换: 输出数据格式

在此输入图像描述

我想根据条件类型合并两行。

据我所知,我需要获取 3 个数据字段的组合键,并在它们相等时比较类型字段。

有人可以帮我使用 Python 在 Spark 中实现吗?

编辑:以下是我在 pyspark 中使用 RDD 的尝试

record = spark.read.csv("wasb:///records.csv",header=True).rdd
print("Total records: %d")%record.count()
private_ip = record.map(lambda fields: fields[2]).distinct().count()
private_port = record.map(lambda fields: fields[3]).distinct().count()
destination_ip = record.map(lambda fields: fields[6]).distinct().count()
destination_port = record.map(lambda fields: fields[7]).distinct().count()
print("private_ip:%d, private_port:%d, destination ip:%d, destination_port:%d")%(private_ip,private_port,destination_ip,destination_port)
types = record.map(lambda fields: ((fields[2],fields[3],fields[6],fields[7]),fields[0])).reduceByKey(lambda a,b:a+','+b)
print types.first()
Run Code Online (Sandbox Code Playgroud)

以下是我到目前为止的输出。

((u'100.79.195.101', u'54835', u'58.96.162.33', u'80'), u'22-02-2016 13:11:03,22-02-2016 13:13:53')
Run Code Online (Sandbox Code Playgroud)

python apache-spark pyspark

4
推荐指数
1
解决办法
1万
查看次数

标签 统计

apache-spark ×2

pyspark ×2

python ×2

rdd ×1