Lou*_*ger 7 amazon-web-services aws-lambda aws-codebuild
我从AWS Codebuild上传了我的lambda函数源。我的Python脚本使用NLTK,因此需要大量数据。我的.zip软件包太大,RequestEntityTooLargeException发生了。我想知道如何增加通过UpdateFunctionCode命令发送的部署包的大小。
我AWS CodeBuild用来将源代码从GitHub存储库转换为AWS Lambda。这是关联的buildspec文件:
version: 0.2
phases:
install:
commands:
- echo "install step"
- apt-get update
- apt-get install zip -y
- apt-get install python3-pip -y
- pip install --upgrade pip
- pip install --upgrade awscli
# Define directories
- export HOME_DIR=`pwd`
- export NLTK_DATA=$HOME_DIR/nltk_data
pre_build:
commands:
- echo "pre_build step"
- cd $HOME_DIR
- virtualenv venv
- . venv/bin/activate
# Install modules
- pip install -U requests
# NLTK download
- pip install -U nltk
- python -m nltk.downloader -d $NLTK_DATA wordnet stopwords punkt
- pip freeze > requirements.txt
build:
commands:
- echo 'build step'
- cd $HOME_DIR
- mv $VIRTUAL_ENV/lib/python3.6/site-packages/* .
- sudo zip -r9 algo.zip .
- aws s3 cp --recursive --acl public-read ./ s3://hilightalgo/
- aws lambda update-function-code --function-name arn:aws:lambda:eu-west-3:671560023774:function:LaunchHilight --zip-file fileb://algo.zip
- aws lambda update-function-configuration --function-name arn:aws:lambda:eu-west-3:671560023774:function:LaunchHilight --environment 'Variables={NLTK_DATA=/var/task/nltk_data}'
post_build:
commands:
- echo "post_build step"
Run Code Online (Sandbox Code Playgroud)
启动管道时,RequestEntityTooLargeException因为.zip包中的数据太多,所以出现了。请参阅下面的构建日志:
[Container] 2019/02/11 10:48:35 Running command aws lambda update-function-code --function-name arn:aws:lambda:eu-west-3:671560023774:function:LaunchHilight --zip-file fileb://algo.zip
An error occurred (RequestEntityTooLargeException) when calling the UpdateFunctionCode operation: Request must be smaller than 69905067 bytes for the UpdateFunctionCode operation
[Container] 2019/02/11 10:48:37 Command did not exit successfully aws lambda update-function-code --function-name arn:aws:lambda:eu-west-3:671560023774:function:LaunchHilight --zip-file fileb://algo.zip exit status 255
[Container] 2019/02/11 10:48:37 Phase complete: BUILD Success: false
[Container] 2019/02/11 10:48:37 Phase context status code: COMMAND_EXECUTION_ERROR Message: Error while executing command: aws lambda update-function-code --function-name arn:aws:lambda:eu-west-3:671560023774:function:LaunchHilight --zip-file fileb://algo.zip. Reason: exit status 255
Run Code Online (Sandbox Code Playgroud)
当我减少要下载的NLTK数据时,一切工作正常(我尝试仅使用软件包stopwords和wordnet。
有谁有解决这个“尺寸限制问题”的想法?
Jac*_*ack 11
以下是 Lambda 的硬限制(将来可能会改变):
解决这个问题的一个明智方法是从 Lambda 挂载 EFS。这不仅对于加载库很有用,而且对于其他存储也很有用。
浏览一下这些博客:
小智 10
如果有人在 2020 年 12 月之后偶然发现这个问题,那么 AWS 已经进行了重大更新,以支持将 Lambda 函数用作容器映像(最大 10GB !!)。更多信息在这里
小智 7
AWS Lambda 函数可以挂载 EFS。您可以使用 EFS 加载大于 AWS Lambda 的 250 MB 包部署大小限制的库或包。
关于如何设置的详细步骤在这里:https : //aws.amazon.com/blogs/aws/new-a-shared-file-system-for-your-lambda-functions/
在高层次上,这些变化包括:
来自AWS文档:
如果您的部署包大于 50 MB,我们建议您将函数代码和依赖项上传到 Amazon S3 存储桶。
您可以创建部署包并将 .zip 文件上传到您要在其中创建 Lambda 函数的 AWS 区域中的 Amazon S3 存储桶。创建 Lambda 函数时,请在 Lambda 控制台上或使用 AWS 命令行界面 (AWS CLI) 指定 S3 存储桶名称和对象键名称。
您可以使用 AWS CLI 部署包,并且可以使用--code参数指定 S3 存储桶中的对象,而不是使用--zip-file参数来传递部署包。前任:
aws lambda create-function --function-name my_function --code S3Bucket=my_bucket,S3Key=my_file
Run Code Online (Sandbox Code Playgroud)
来自 github ( https://github.com/awslabs/aws-data-wrangler/releases )的 aws wrangler zip 文件包含许多其他库,例如 pandas 和 pymysql。就我而言,这是我唯一需要的层,因为它还有很多其他东西。可能对某些人有用。
您不能增加Lambda的部署程序包大小。AWS Lambda开发人员指南中介绍了AWS Lambda限制。有关这些限制如何工作的更多信息,请参见此处。本质上,您解压缩的程序包大小必须小于250MB(262144000字节)。
PS:虽然可以帮助管理和更快的冷启动,但使用图层并不能解决调整大小的问题。包装尺寸包括Lambda层。
一个功能一次最多可以使用5层。功能和所有层的总解压缩大小不能超过250 MB的解压缩部署程序包大小限制。
我自己没有尝试过这个,但Zappa的人描述了一个可能有帮助的技巧。引自https://blog.zappa.io/posts/slim-handler:
Zappa 压缩大型应用程序并将项目 zip 文件发送到 S3。其次,Zappa 创建了一个非常精简的处理程序,它只包含 Zappa 及其依赖项并将其发送到 Lambda。
当在冷启动时调用 slim 处理程序时,它会从 S3 下载大型项目 zip 并将其解压缩到 Lambda 的共享 /tmp 空间中。对该暖 Lambda 的所有后续调用都共享 /tmp 空间并可以访问项目文件;因此,如果 Lambda 保持温暖,文件可能只下载一次。
这样你应该在 /tmp 中得到 500MB。
更新:
我在几个项目的lambdas中使用了以下代码,它基于zappa使用的方法,但可以直接使用。
# Based on the code in https://github.com/Miserlou/Zappa/blob/master/zappa/handler.py
# We need to load the layer from an s3 bucket into tmp, bypassing the normal
# AWS layer mechanism, since it is too large, AWS unzipped lambda function size
# including layers is 250MB.
def load_remote_project_archive(remote_bucket, remote_file, layer_name):
# Puts the project files from S3 in /tmp and adds to path
project_folder = '/tmp/{0!s}'.format(layer_name)
if not os.path.isdir(project_folder):
# The project folder doesn't exist in this cold lambda, get it from S3
boto_session = boto3.Session()
# Download zip file from S3
s3 = boto_session.resource('s3')
archive_on_s3 = s3.Object(remote_bucket, remote_file).get()
# unzip from stream
with io.BytesIO(archive_on_s3["Body"].read()) as zf:
# rewind the file
zf.seek(0)
# Read the file as a zipfile and process the members
with zipfile.ZipFile(zf, mode='r') as zipf:
zipf.extractall(project_folder)
# Add to project path
sys.path.insert(0, project_folder)
return True
Run Code Online (Sandbox Code Playgroud)
然后可以按如下方式调用它(我通过 env 变量将带有层的存储桶传递给 lambda 函数):
load_remote_project_archive(os.environ['MY_ADDITIONAL_LAYERS_BUCKET'], 'lambda_my_extra_layer.zip', 'lambda_my_extra_layer')
Run Code Online (Sandbox Code Playgroud)
写这段代码的时候,tmp也被封顶了,我想是250MB,但是zipf.extractall(project_folder)上面的调用可以换成直接提取到内存中:unzipped_in_memory = {name: zipf.read(name) for name in zipf.namelist()}
我对一些机器学习模型做的,我猜@rahul的答案是尽管如此,它更通用。
| 归档时间: |
|
| 查看次数: |
3319 次 |
| 最近记录: |