我有一台机器"c3.8xlarge"的EMR集群,在阅读了几个资源之后,我明白我必须允许在堆外使用相当数量的内存,因为我使用的是pyspark,所以我将集群配置如下:
一位执行人:
司机:
当我cache()使用DataFrame时,它需要大约3.6GB的内存.
现在,当我调用collect()或toPandas()在DataFrame上时,进程崩溃了.
我知道我将大量数据带入驱动程序,但我认为它不是那么大,而且我无法弄清楚崩溃的原因.
当我打电话collect()或toPandas()我收到此错误时:
Py4JJavaError: An error occurred while calling o181.collectToPython.
: org.apache.spark.SparkException: Job aborted due to stage failure: Task 5 in stage 6.0 failed 4 times, most recent failure: Lost task 5.3 in stage 6.0 (TID 110, ip-10-0-47-207.prod.eu-west-1.hs.internal, executor 9): ExecutorLostFailure (executor 9 exited caused by one of the running tasks) Reason: Container marked as failed: container_1511879540686_0005_01_000016 on …Run Code Online (Sandbox Code Playgroud) 我在Hadoop的YARN上运行Spark.这种转换如何运作?在转换之前是否会发生collect()?
另外我需要在每个从节点上安装Python和R才能使转换工作?我很难找到这方面的文件.