如何通过正则表达式删除 URL 的某些部分?

Rom*_*man 1 java regex string

我有这样的完整链接:

http://localhost:8080/suffix/rest/of/link
Run Code Online (Sandbox Code Playgroud)

如何在 Java 中编写正则表达式,它只返回带有后缀的 url 的主要部分:http://localhost/suffix而没有:/rest/of/link?

  • 可能的协议:http、https
  • 可能的端口:多种可能性

我假设我需要在第三次出现'/'标记(包括)后删除整个文本。我想按以下方式进行,但我不太了解正则表达式,请您帮忙如何正确编写正则表达式?

String appUrl = fullRequestUrl.replaceAll("(.*\\/{2})", ""); //this removes 'http://' but this is not my case
Run Code Online (Sandbox Code Playgroud)

Rah*_*thi 5

我不确定您为什么要为此使用 Regex。Java 提供了一个Query URL Objects来为您做同样的事情。

以下是来自同一站点的示例,以展示其工作原理:

import java.net.*;
import java.io.*;

public class ParseURL {
    public static void main(String[] args) throws Exception {

        URL aURL = new URL("http://example.com:80/docs/books/tutorial"
                           + "/index.html?name=networking#DOWNLOADING");

        System.out.println("protocol = " + aURL.getProtocol());
        System.out.println("authority = " + aURL.getAuthority());
        System.out.println("host = " + aURL.getHost());
        System.out.println("port = " + aURL.getPort());
        System.out.println("path = " + aURL.getPath());
        System.out.println("query = " + aURL.getQuery());
        System.out.println("filename = " + aURL.getFile());
        System.out.println("ref = " + aURL.getRef());
    }
}
Run Code Online (Sandbox Code Playgroud)

这是程序显示的输出:

protocol = http
authority = example.com:80
host = example.com
port = 80
path = /docs/books/tutorial/index.html
query = name=networking
filename = /docs/books/tutorial/index.html?name=networking
ref = DOWNLOADING
Run Code Online (Sandbox Code Playgroud)