在join和reduceByKey中将执行程序激发出内存不足

Question

在join和reduceByKey中将执行程序激发出内存不足

xua*_*uan 5 out-of-memory executor executors apache-spark

在spark2.0中，我有两个数据框，我需要首先将它们加入并做一个reduceByKey来聚合数据。我总是在执行器中有OOM。提前致谢。

数据

d1（1G，5亿行，已缓存，由col id2分区）

Run Code Online (Sandbox Code Playgroud)

d2（160G，200万行，已缓存，由col id2分区，值col包含5000个浮点数的列表）

id2   value
0     [0.1, 0.2, 0.0001, ...]
1     [0.001, 0.7, 0.0002, ...]
...

Run Code Online (Sandbox Code Playgroud)

现在我需要加入两个表以获取d3并使用spark.sql

select d1.id1, d2.value
FROM d1 JOIN d2 ON d1.id2 = d2.id2

Run Code Online (Sandbox Code Playgroud)

然后在d3上执行reduceByKey并汇总表d1中每个id1的值

d4 = d3.rdd.reduceByKey(lambda x, y: numpy.add(x, y)) \
           .mapValues(lambda x: (x / numpy.linalg.norm(x, 1)).toList)\
           .toDF()

Run Code Online (Sandbox Code Playgroud)

我估计d4的大小为340G。现在我在r3.8xlarge机器上使用以运行作业

mem: 244G
cpu: 64
Disk: 640G

Run Code Online (Sandbox Code Playgroud)

问题

我玩了一些配置，但是执行器中总是有OOM。所以，问题是

是否可以在当前类型的机器上运行此作业？或者我应该只使用更大的机器（多大？）。但是我记得我曾经遇到过一些文章/博客，它们说使用相对较小的机器来进行TB级的处理。
我应该做什么样的改善？例如火花配置，代码优化？
是否可以估计每个执行器所需的内存量？

火花配置

我尝试过的一些Spark配置

config1：

--verbose
--conf spark.sql.shuffle.partitions=200
--conf spark.dynamicAllocation.enabled=false
--conf spark.driver.maxResultSize=24G
--conf spark.shuffle.blockTransferService=nio
--conf spark.serializer=org.apache.spark.serializer.KryoSerializer
--conf spark.kryoserializer.buffer.max=2000M
--conf spark.rpc.message.maxSize=800
--conf "spark.executor.extraJavaOptions=-verbose:gc -     XX:+PrintGCDetails -XX:+PrintGCTimeStamps -XX:MetaspaceSize=100M"
--num-executors 4
--executor-memory 48G
--executor-cores 15
--driver-memory 24G
--driver-cores 3

Run Code Online (Sandbox Code Playgroud)

config2：

--verbose
--conf spark.sql.shuffle.partitions=10000
--conf spark.dynamicAllocation.enabled=false
--conf spark.driver.maxResultSize=24G
--conf spark.shuffle.blockTransferService=nio
--conf spark.serializer=org.apache.spark.serializer.KryoSerializer
--conf spark.kryoserializer.buffer.max=2000M
--conf spark.rpc.message.maxSize=800
--conf "spark.executor.extraJavaOptions=-verbose:gc -XX:+PrintGCDetails -XX:+PrintGCTimeStamps -XX:MetaspaceSize=100M"
--num-executors 4
--executor-memory 48G
--executor-cores 15
--driver-memory 24G
--driver-cores 3

Run Code Online (Sandbox Code Playgroud)

配置3：

--verbose
--conf spark.sql.shuffle.partitions=10000
--conf spark.dynamicAllocation.enabled=true
--conf spark.driver.maxResultSize=6G
--conf spark.shuffle.blockTransferService=nio
--conf spark.serializer=org.apache.spark.serializer.KryoSerializer
--conf spark.kryoserializer.buffer.max=2000M
--conf spark.rpc.message.maxSize=800
--conf "spark.executor.extraJavaOptions=-verbose:gc -XX:+PrintGCDetails -XX:+PrintGCTimeStamps -XX:MetaspaceSize=100M"
--executor-memory 6G
--executor-cores 2
--driver-memory 6G
--driver-cores 3

Run Code Online (Sandbox Code Playgroud)

配置4：

--verbose
--conf spark.sql.shuffle.partitions=20000
--conf spark.dynamicAllocation.enabled=false
--conf spark.driver.maxResultSize=6G
--conf spark.shuffle.blockTransferService=nio
--conf spark.serializer=org.apache.spark.serializer.KryoSerializer
--conf spark.kryoserializer.buffer.max=2000M
--conf spark.rpc.message.maxSize=800
--conf "spark.executor.extraJavaOptions=-verbose:gc -XX:+PrintGCDetails -XX:+PrintGCTimeStamps -XX:MetaspaceSize=100M"
--num-executors 13
--executor-memory 15G
--executor-cores 5
--driver-memory 13G
--driver-cores 5

Run Code Online (Sandbox Code Playgroud)

错误

来自执行器的OOM错误1

ExecutorLostFailure (executor 14 exited caused by one of the running  tasks) Reason: Container killed by YARN for exceeding memory limits. 9.1 GB of 9 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead.

Heap
PSYoungGen      total 1830400K, used 1401721K [0x0000000740000000,   0x00000007be900000, 0x00000007c0000000)
eden space 1588736K, 84% used [0x0000000740000000,0x0000000791e86980,0x00000007a0f80000)
from space 241664K, 24% used [0x00000007af600000,0x00000007b3057de8,0x00000007be200000)
to  space 236032K, 0% used [0x00000007a0f80000,0x00000007a0f80000,0x00000007af600000)
ParOldGen      total 4194304K, used 4075884K [0x0000000640000000, 0x0000000740000000, 0x0000000740000000)
object space 4194304K, 97% used [0x0000000640000000,0x0000000738c5b198,0x0000000740000000)
Metaspace      used 59721K, capacity 60782K, committed 61056K,  reserved 1101824K
class space    used 7421K, capacity 7742K, committed 7808K, reserved 1048576K

Run Code Online (Sandbox Code Playgroud)

来自执行器的OOM错误2

ExecutorLostFailure (executor 7 exited caused by one of the running tasks) Reason: Container marked as failed: container_1477662810360_0002_01_000008 on host: ip-172-18-9-130.ec2.internal. Exit status: 52. Diagnostics: Exception from container-launch.

Heap
PSYoungGen      total 1968128K, used 1900544K [0x0000000740000000, 0x00000007c0000000, 0x00000007c0000000)
eden space 1900544K, 100% used [0x0000000740000000,0x00000007b4000000,0x00000007b4000000)
from space 67584K, 0% used [0x00000007b4000000,0x00000007b4000000,0x00000007b8200000)
to  space 103936K, 0% used [0x00000007b9a80000,0x00000007b9a80000,0x00000007c0000000)
ParOldGen      total 4194304K, used 4194183K [0x0000000640000000, 0x0000000740000000, 0x0000000740000000)
object space 4194304K, 99% used [0x0000000640000000,0x000000073ffe1f38,0x0000000740000000)
Metaspace      used 59001K, capacity 59492K, committed 61056K, reserved 1101824K
class space    used 7300K, capacity 7491K, committed 7808K, reserved 1048576K

Run Code Online (Sandbox Code Playgroud)

容器错误

16/10/28 14:33:21 ERROR CoarseGrainedExecutorBackend: RECEIVED SIGNAL TERM
16/10/28 14:33:26 ERROR Utils: Uncaught exception in thread stdout writer for python
java.lang.OutOfMemoryError: Java heap space
    at org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:504)
    at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)
    at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
    at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$doExecute$3$$anon$2.hasNext(WholeStageCodegenExec.scala:386)
    at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.next(SerDeUtil.scala:120)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.next(SerDeUtil.scala:112)
    at scala.collection.Iterator$class.foreach(Iterator.scala:893)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.foreach(SerDeUtil.scala:112)
    at org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)
    at org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)
    at org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1877)
    at org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)
16/10/28 14:33:36 ERROR Utils: Uncaught exception in thread driver-heartbeater
16/10/28 14:33:26 ERROR Utils: Uncaught exception in thread stdout writer for python
java.lang.OutOfMemoryError: GC overhead limit exceeded
    at java.lang.Double.valueOf(Double.java:519)
    at org.apache.spark.sql.catalyst.expressions.UnsafeArrayData.get(UnsafeArrayData.java:138)
    at org.apache.spark.sql.catalyst.util.ArrayData.foreach(ArrayData.scala:135)
    at org.apache.spark.sql.execution.python.EvaluatePython$.toJava(EvaluatePython.scala:64)
    at org.apache.spark.sql.execution.python.EvaluatePython$.toJava(EvaluatePython.scala:57)
    at org.apache.spark.sql.Dataset$$anonfun$55.apply(Dataset.scala:2517)
    at org.apache.spark.sql.Dataset$$anonfun$55.apply(Dataset.scala:2517)
    at scala.collection.Iterator$$anon$11.next(Iterator.scala:409)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.next(SerDeUtil.scala:121)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.next(SerDeUtil.scala:112)
    at scala.collection.Iterator$class.foreach(Iterator.scala:893)
    at org.apache.spark.api.python.SerDeUtil$AutoBatchedPickler.foreach(SerDeUtil.scala:112)
    at org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)
    at org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)
    at org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1877)
    at org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)
16/10/28 14:33:43 ERROR SparkUncaughtExceptionHandler: [Container in shutdown] Uncaught exception in thread Thread[stdout writer for python,5,main]

Run Code Online (Sandbox Code Playgroud)

更新1

如果我按id2分区，数据d1看起来似乎很歪斜。结果，该联接将导致OOM。如果d1像我之前想象的那样均匀分布，则上面的配置应该可以工作。

更新2

如果有人也遇到类似的问题，我发布了解决问题的尝试。

尝试1

我的问题是，如果我用id2对d1进行分区，那么数据将非常不对称。结果，存在一些包含几乎所有id1的分区。因此，与d2的JOIN将导致OOM错误。为了减轻此类问题，我首先s从id2中识别出一个子集，如果按id2进行分区，则可能会导致这种偏斜的数据。然后，我从仅包含d2的d5 s和不包含d2的d6 创建一个s。幸运的是，d5的大小不是太大。因此，我可以广播d1和d5的连接。然后我加入d1和d6。然后，我将两个结果合并并执行reduceByKey。我非常想解决这个问题。我没有继续这样做，因为我的d1以后可能会变得更大。换句话说，这种方法对我而言并不是真正可扩展的

尝试2

幸运的是，对于我来说，d2中的大多数值都很小。根据我的应用程序，我可以安全地删除较小的值并将向量转换为sparseVector，从而显着减小d2的大小。完成此操作后，我将d1划分为id1并广播加入d2（在删除较小值之后）。当然，必须增加驱动程序的内存以允许较大的广播变量。这对我有用，对我的应用程序也可以扩展。

Answer 1

Tim*_*Tim 5

您可以尝试以下方法：将执行程序的大小减小一点。您目前拥有：

--executor-memory 48G
--executor-cores 15

Run Code Online (Sandbox Code Playgroud)

试试看：

--executor-memory 16G
--executor-cores 5

Run Code Online (Sandbox Code Playgroud)

出于各种原因，较小的执行程序大小似乎是最佳的。其中之一是java堆大小大于32G会导致对象引用从4个字节变为8个，并且所有内存需求都爆满。

编辑：问题实际上可能是d4分区太大（尽管其他建议仍然适用！）。您可以通过将d3重新分区为更大数量的分区（大约为d1 * 4），或将其传递给的 numPartitions可选参数来解决reduceByKey。这两个选项都将触发随机播放，但这比崩溃更好。

归档时间：	9 年，4 月前
查看次数：	7628 次
最近记录：	7 年，6 月前