有没有办法getElementsByTagName只在单个节点级别使用而不是递归?
例如,考虑解析pom.xml文件:
<project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/maven-v4_0_0.xsd">
<parent>
<groupId>com.parent</groupId>
<artifactId>parent</artifactId>
<version>1.0-SNAPSHOT</version>
<relativePath>../pom.xml</relativePath>
</parent>
<modelVersion>2.0.0</modelVersion>
<groupId>com.parent.somemodule</groupId>
<artifactId>some_module</artifactId>
<packaging>jar</packaging>
<version>1.0-SNAPSHOT</version>
<name>Some Module</name>
...
Run Code Online (Sandbox Code Playgroud)
如果我想进入groupId顶级(特别是project->groupId,不是project->parent->groupId),我使用:
xmldoc = minidom.parse('pom.xml')
groupId = xmldoc.getElementsByTagName("groupId")[0].childNodes[0].nodeValue
Run Code Online (Sandbox Code Playgroud)
但不幸的是,groupId无论层次结构级别如何,它都会在文件中找到第一个物理事件project->parent->groupId.我实际上只想在特定节点级别进行非递归查找,而不是在其子级内.有没有办法做到这一点xml.dom?
更新: 我切换到BeautifulSoup但仍然有隐式递归遍历的相同问题:使用BeautifulSoup在Python中查找非递归DOM子节点
您可以迭代getElementsByTagName()结果并获取根级别上的第一个元素:
group_id_element = next(element for element in xmldoc.getElementsByTagName("groupId")
if element.parentNode == xmldoc.documentElement)
print group_id_element.childNodes[0].nodeValue
Run Code Online (Sandbox Code Playgroud)
请注意,使用ElementTree做同样的事情会更容易、更短和更快,它也是标准库的一部分。
希望有帮助。