Har*_*h J 7 amazon-web-services amazon-kinesis terraform aws-glue amazon-kinesis-firehose
我在 Terraform 中有一个 Kinesis Firehose 配置,它从 JSON 格式的 Kinesis 流读取数据,使用 Glue 将其转换为 Parquet 并写入 S3。数据格式转换出现问题,我收到以下错误(删除了一些详细信息):
{“attemptsMade”:1,“arrivalTimestamp”:1624541721545,“lastErrorCode”:“DataFormatConversion.InvalidSchema”,“lastErrorMessage”:“架构无效。指定的表没有列。”,“attemptEndingTimestamp”:1624542026951,“rawData ":"xx","sequenceNumber":"xx","subSequenceNumber":null,"dataCatalogTable":{"catalogId":null,"databaseName":"db_name","tableName":"table_name","region" :null,"versionId":"最新","roleArn":"xx"}}
我正在使用的 Glue Table 的 Terraform 配置如下:
resource "aws_glue_catalog_table" "stream_format_conversion_table" {
name = "${var.resource_prefix}-parquet-conversion-table"
database_name = aws_glue_catalog_database.stream_format_conversion_db.name
table_type = "EXTERNAL_TABLE"
parameters = {
EXTERNAL = "TRUE"
"parquet.compression" = "SNAPPY"
}
storage_descriptor {
location = "s3://${element(split(":", var.bucket_arn), 5)}/"
input_format = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat"
output_format = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat"
ser_de_info {
name = "my-stream"
serialization_library = "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe"
parameters = {
"serialization.format" = 1
}
}
columns {
name = "metadata"
type = "struct<tenantId:string,env:string,eventType:string,eventTimeStamp:timestamp>"
}
columns {
name = "eventpayload"
type = "struct<operation:string,timestamp:timestamp,user_name:string,user_id:int,user_email:string,batch_id:string,initiator_id:string,initiator_email:string,payload:string>"
}
}
}
Run Code Online (Sandbox Code Playgroud)
这里需要改变什么?
小智 5
我遇到了“架构无效。指定的表没有列”,其组合如下:
事实证明,如果表是从现有架构创建的,KDF 无法读取表的架构。必须从头开始创建表(与“从现有架构添加表”相对)这尚未记录......目前。
| 归档时间: |
|
| 查看次数: |
2178 次 |
| 最近记录: |