XPath获取特定长度的文本

Anc*_*end 6 xpath

我正在尝试创建一个XPath查询,每次都会获得549个字符的文本.该文本应该是关于相关主题,在下面的例子中是oranges或apples或pears.如果页面上不存在包含这些单词的元素,那么我希望XPath查询更容易在页面上找到目标/不太具体的文本.

所以为了澄清,我试图创建一个XPath查询,找到包含特定类型文本的元素,如果使用下面的查询找到549个或更多字符,那么我们就完成了,如果没有找到或者返回的文本是少于549个字符,我希望XPath查询在段落形式的页面上获取任何文本(除了按钮,链接,菜单等文本之外的任何内容都有效)并返回此文本的549个字符,如果结果string小于549个字符我想用以下内容连接这两个查询:...在中间.

   substring(normalize-space(//*[self::p or self::div][contains(text(),'apples') or contains(text(),'oranges') or contains(text(),'pears')]), 0, 549)
Run Code Online (Sandbox Code Playgroud)

我一直试图解决这个问题很长一段时间,我将不胜感激任何建议!

提前谢谢了!

Pau*_*mer 8

是.string-length()您可以在谓词中使用xpath中的函数:

substring(normalize-space(//*[string-length( text()) > 549 and (... other conditions ...)]),0,549)
Run Code Online (Sandbox Code Playgroud)

有关如何执行条件以确定是否需要添加省略号,请参阅" 是否存在"if -then - else"语句在XPath中? "

改编上述SO问题的例子:

if (fn:string-length(normalize-space(//*[self::p or self::div][contains(text(),'apples']) > 549)
        then (concat( fn:substring(normalize-space(//*[self::p or self::div][contains(text(),'apples']), 0, 5490), "...") )
        else (normalize-space(//*[self::p or self::div][contains(text(),'apples']))
Run Code Online (Sandbox Code Playgroud)

在我看来,这在XPath中真的很复杂.如果你可以使用XQuery,你将拥有更易读的变换:

for $text in normalize-space(//*[self::p or self::div])
where $text[contains(text(),'apples' or ...]
return
    if (string-length( $text) > 549) then
        concat( substring( $text, 0, 549), "...")
    else
        $text
Run Code Online (Sandbox Code Playgroud)

我怀疑这实际上可以进一步优化(为了可读性,维护),使用多个嵌套的for语句来处理您需要的各种结果.

如果使用XSL:

<xsl:template match="//*[self::p or self::div][contains(text(),'apples' or ...]">
    <xsl:variable name="text" select="normalize-space( . )" />
    <xsl:choose>
        <xsl:when test="string-length( $text)">
            <xsl:value-of select="substring( $text, 0, 549)"/>...
        </xsl:when>
        <xsl:otherwise>
            <xsl:value-of select="$text"/>
        </xsl:otherwise>
    </xsl:choose>
</xsl:template>
Run Code Online (Sandbox Code Playgroud)

您还可以使用matches()xpath函数contains(),通过构造正则表达式来避免拥有如此多的谓词:

matches( //*[self::p or self::div][matches(text(),'(apples|oranges|bananas)'])
Run Code Online (Sandbox Code Playgroud)

最后,请注意,在XPath 中使用//和*非常低效,如果您的文档对其有任何影响,您将看到性能影响.我有一个痒告诉我有一种优化方法,但不幸的是我没有时间研究.