从字符串中检测并提取url?

Shi*_*oft 40 java regex url

这是一个简单的问题,但我不明白.我想检测字符串中的url并用缩短的字符串替换它们.

我从stackoverflow中找到了这个表达式,但结果却是 http

Pattern p = Pattern.compile("\\b(https?|ftp|file)://[-a-zA-Z0-9+&@#/%?=~_|!:,.;]*[-a-zA-Z0-9+&@#/%=~_|]",Pattern.CASE_INSENSITIVE);
        Matcher m = p.matcher(str);
        boolean result = m.find();
        while (result) {
            for (int i = 1; i <= m.groupCount(); i++) {
                String url=m.group(i);
                str = str.replace(url, shorten(url));
            }
            result = m.find();
        }
        return html;
Run Code Online (Sandbox Code Playgroud)

有什么好主意吗?

Whi*_*g34 87

让我继续前言并说明我不是复杂案例的正则表达式的大力倡导者.试图为这样的事情写出完美的表达是非常困难的.也就是说,我碰巧有一个用于检测URL,并且它由350行单元测试用例类支持.有人从一个简单的正则表达式开始,多年来我们已经增加了表达式和测试用例来处理我们发现的问题.这绝对不是微不足道的:

// Pattern for recognizing a URL, based off RFC 3986
private static final Pattern urlPattern = Pattern.compile(
        "(?:^|[\\W])((ht|f)tp(s?):\\/\\/|www\\.)"
                + "(([\\w\\-]+\\.){1,}?([\\w\\-.~]+\\/?)*"
                + "[\\p{Alnum}.,%_=?&#\\-+()\\[\\]\\*$~@!:/{};']*)",
        Pattern.CASE_INSENSITIVE | Pattern.MULTILINE | Pattern.DOTALL);
Run Code Online (Sandbox Code Playgroud)

以下是使用它的示例:

Matcher matcher = urlPattern.matcher("foo bar http://example.com baz");
while (matcher.find()) {
    int matchStart = matcher.start(1);
    int matchEnd = matcher.end();
    // now you have the offsets of a URL match
}
Run Code Online (Sandbox Code Playgroud)

  • 不能正确处理文本中的URL.前面的空格处理不正确(吞没了换行符),并在URL后接受冒号,点等. (4认同)
  • 这应该是公认的答案.辉煌. (3认同)
  • 不幸的是,这个也匹配URL后面的一个点. (3认同)
  • 不能使用"<a href="www.google.com">谷歌链接</a>"它返回"www.google.com" (3认同)

Bul*_*aza 40

/**
 * Returns a list with all links contained in the input
 */
public static List<String> extractUrls(String text)
{
    List<String> containedUrls = new ArrayList<String>();
    String urlRegex = "((https?|ftp|gopher|telnet|file):((//)|(\\\\))+[\\w\\d:#@%/;$()~_?\\+-=\\\\\\.&]*)";
    Pattern pattern = Pattern.compile(urlRegex, Pattern.CASE_INSENSITIVE);
    Matcher urlMatcher = pattern.matcher(text);

    while (urlMatcher.find())
    {
        containedUrls.add(text.substring(urlMatcher.start(0),
                urlMatcher.end(0)));
    }

    return containedUrls;
}
Run Code Online (Sandbox Code Playgroud)

例:

List<String> extractedUrls = extractUrls("Welcome to https://stackoverflow.com/ and here is another link http://www.google.com/ \n which is a great search engine");

for (String url : extractedUrls)
{
    System.out.println(url);
}
Run Code Online (Sandbox Code Playgroud)

打印:

https://stackoverflow.com/
http://www.google.com/
Run Code Online (Sandbox Code Playgroud)


M'v*_*'vy 7

m.group(1)为您提供第一个匹配组,即第一个捕获括号.在这里(https?|ftp|file)

您应该尝试查看m.group(0)中是否存在某些内容,或者用括号括起所有模式并再次使用m.group(1).

您需要重复查找函数以匹配下一个函数并使用新的组数组.


ste*_*ema 5

检测 URL 并非易事。如果它足以让您获得以 https?|ftp|file 开头的字符串,那么它可能没问题。您的问题是,您有一个捕获组,()这些仅在第一部分 http ...

我会使用 (?:) 将这部分设为非捕获组,并在整个内容周围加上括号。

"\\b((?:https?|ftp|file)://[-a-zA-Z0-9+&@#/%?=~_|!:,.;]*[-a-zA-Z0-9+&@#/%=~_|])"
Run Code Online (Sandbox Code Playgroud)