该文件包含2000000行:每行包含208列,以逗号分隔,如下所示:
0.0863314058048,0.0208767447842,0.03358010485,0.0,1.0,0.0,0.314285714286,0.336293217457,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0
该程序将这个文件读成一个numpy叙述,我预计它将消耗大约(2000000 * 208 * 8B) = 3.2GB内存.但是,当程序读取此文件时,我发现该程序消耗大约20GB的内存.
我很困惑为什么我的程序会消耗如此多的内存而不符合预期?
这是一个由python编写的简单spark。我收到“诊断:根据请求杀死容器。退出代码为 143”错误。这是我运行此代码的代码和脚本。
输入目录的大小大约为 100GB。但是如果我使用一个小的数据文件(3GB),它会工作得很好。
import sys
from pyspark import SparkContext
sc = SparkContext(appName="job_name")
data1 = sc.textFile(sys.argv[1])
d1 = data1.filter(lambda x: "a string" in x)
print d1.count()
sc.stop()
//-----------------------------------------//
input="xxxxx"
output="yyyyy"
hadoop fs -rmr $output
$SPARK_HOME/bin/spark-submit \
--deploy-mode cluster \
--master yarn \
--num-executors 100 \
--executor-cores 2 \
--driver-cores 2 \
--executor-memory 8g \
--driver-memory 4g \
stat.py \
$input \
$output
Run Code Online (Sandbox Code Playgroud)
所以我写了一个 hadoop 流来完成这项工作,我得到了:
Error: Java heap space
Container killed by the ApplicationMaster.
Container killed on request. Exit code …Run Code Online (Sandbox Code Playgroud)