dis*_*ame 8 python machine-learning apache-spark pyspark
我正在寻找等效的变压器,例如MultiLabelBinarizerin sklearn。
到目前为止,我发现的只是这Binarizer并不能真正满足我的需要。
我也在看这个文档,但我看不到任何我想要的东西。
我的输入是一个列,其中每个元素都是一个标签列表:
labels
['a', 'b']
['a']
['c', 'b']
['a', 'c']
Run Code Online (Sandbox Code Playgroud)
输出应该是
labels
[1, 1, 0]
[1, 0, 0]
[0, 1, 1]
[1, 0, 1]
Run Code Online (Sandbox Code Playgroud)
PySpark 相当于什么?
以下解决方案可能不是非常优化,但我认为它非常简单并且快速完成工作。
我们基本上创建一个函数来收集列中包含的所有不同值labels,然后为列中遇到的每个值动态创建 0/1 列labels。
import pyspark.sql.functions as F
def multi_label_binarizer(df, labels_col='labels', output_col='new_labels'):
"""
Function that takes as input:
- `df`, pyspark.sql.dataframe
- `labels_col`, string that indicates an array column containing labels
- `output_col`, string that indicates the name of the new labels column
and returns a multi-label binarized column.
"""
# get set of unique labels and sort them
labels_set = df\
.withColumn('exploded', F.explode('labels'))\
.agg(F.collect_set('exploded'))\
.collect()[0][0]
labels_set = sorted(labels_set)
# dynamically create columns for each value in `labels_set`
for i in labels_set:
df = df.withColumn(i, F.when(F.array_contains(labels_col, i), 1).otherwise(0))
# create new, multi-label binarized array column
df = df.withColumn(output_col, F.array(*labels_set))
return df
multi_label_binarizer(df).show()
+------+---+---+---+----------+
|labels| a| b| c|new_labels|
+------+---+---+---+----------+
|[a, b]| 1| 1| 0| [1, 1, 0]|
| [a]| 1| 0| 0| [1, 0, 0]|
|[c, b]| 0| 1| 1| [0, 1, 1]|
|[a, c]| 1| 0| 1| [1, 0, 1]|
+------+---+---+---+----------+
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
564 次 |
| 最近记录: |