将 xml 数据加载到 hive 表中:org.apache.hadoop.hive.ql.metadata.HiveException

shr*_*e11 5 hive xmldataset

我正在尝试将 XML 数据加载到 Hive 中,但出现错误:

java.lang.RuntimeException: org.apache.hadoop.hive.ql.metadata.HiveException: Hive Runtime Error while processing row {"xmldata":""}

我使用的 xml 文件是:

<?xml version="1.0" encoding="UTF-8"?>
<catalog>
<book>
  <id>11</id>
  <genre>Computer</genre>
  <price>44</price>
</book>
<book>
  <id>44</id>
  <genre>Fantasy</genre>
  <price>5</price>
</book>
</catalog>
Run Code Online (Sandbox Code Playgroud)

我使用的配置单元查询是:

1) Create TABLE xmltable(xmldata string) STORED AS TEXTFILE;
LOAD DATA lOCAL INPATH '/home/user/xmlfile.xml' OVERWRITE INTO TABLE xmltable;

2) CREATE VIEW xmlview (id,genre,price)
AS SELECT
xpath(xmldata, '/catalog[1]/book[1]/id'),
xpath(xmldata, '/catalog[1]/book[1]/genre'),
xpath(xmldata, '/catalog[1]/book[1]/price')
FROM xmltable;

3) CREATE TABLE xmlfinal AS SELECT * FROM xmlview;

4) SELECT * FROM xmlfinal WHERE id ='11
Run Code Online (Sandbox Code Playgroud)

直到第二个查询一切正常,但是当我执行第三个查询时,它给了我错误:

错误如下:

java.lang.RuntimeException: org.apache.hadoop.hive.ql.metadata.HiveException: Hive Runtime Error while processing row {"xmldata":"<?xml version=\"1.0\" encoding=\"UTF-8\"?>"}
    at org.apache.hadoop.hive.ql.exec.ExecMapper.map(ExecMapper.java:159)
    at org.apache.hadoop.mapred.MapRunner.run(MapRunner.java:50)
    at org.apache.hadoop.mapred.MapTask.runOldMapper(MapTask.java:417)
    at org.apache.hadoop.mapred.MapTask.run(MapTask.java:332)
    at org.apache.hadoop.mapred.Child$4.run(Child.java:268)
    at java.security.AccessController.doPrivileged(Native Method)
    at javax.security.auth.Subject.doAs(Subject.java:415)
    at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1438)
    at org.apache.hadoop.mapred.Child.main(Child.java:262)
 Caused by: org.apache.hadoop.hive.ql.metadata.HiveException: Hive Runtime Error    while processing row {"xmldata":"<?xml version=\"1.0\" encoding=\"UTF-8\"?>"}
    at org.apache.hadoop.hive.ql.exec.MapOperator.process(MapOperator.java:675)
    at org.apache.hadoop.hive.ql.exec

FAILED: Execution Error, return code 2 from org.apache.hadoop.hive.ql.exec.MapRedTask
Run Code Online (Sandbox Code Playgroud)

那么到底哪里出错了呢?此外,我正在使用正确的 xml 文件。

谢谢,史瑞

vij*_*mar 4

错误原因:

1) case-1:(您的情况)- xml 内容正在逐行输入到 hive。

输入XML:

<?xml version="1.0" encoding="UTF-8"?>
<catalog>
<book>
  <id>11</id>
  <genre>Computer</genre>
  <price>44</price>
</book>
<book>
  <id>44</id>
  <genre>Fantasy</genre>
  <price>5</price>
</book>
</catalog>  
Run Code Online (Sandbox Code Playgroud)

检查蜂巢:

select count(*) from xmltable;  // return 13 rows - means each line in individual row with col xmldata  
Run Code Online (Sandbox Code Playgroud)

错误原因:

XML 被解读为 13 个未统一的部分。如此无效的 XML

2) 情况 2:xml 内容应作为 singleString 提供给 hive - XpathUDF工作 引用语法:所有函数均遵循以下形式:xpath_ (xml_string, xpath_expression_string).* source

输入.xml

<?xml version="1.0" encoding="UTF-8"?><catalog><book><id>11</id><genre>Computer</genre><price>44</price></book><book><id>44</id><genre>Fantasy</genre><price>5</price></book></catalog>
Run Code Online (Sandbox Code Playgroud)

检查蜂巢:

select count(*) from xmltable; // returns 1 row - XML is properly read as complete XML.
Run Code Online (Sandbox Code Playgroud)

方法 :

xmldata   = <?xml version="1.0" encoding="UTF-8"?><catalog><book> ...... </catalog>
Run Code Online (Sandbox Code Playgroud)

然后像这样应用你的 xpathUDF

select xpath(xmldata, 'xpath_expression_string' ) from xmltable
Run Code Online (Sandbox Code Playgroud)