可以使用单个正则表达式来修改网址并匹配所有部分,我一直在研究一个,到目前为止我提出的是:
(?:(?P<scheme>[a-z]*?)://)?(?:(?P<username>.*?):?(?P<password>.*?)?@)?(?P<hostname>.*?)/(?:(?:(?P<path>.*?)\?)?(?P<file>.*?\.[a-z]{1,6})?(?:(?:(?P<query>.*?)#?)?(?P<fragment>.*?)?)?)?
Run Code Online (Sandbox Code Playgroud)
但是这不起作用,它应该匹配以下所有示例:
http:// username:password@hostname.tld/path?arg = value#anchor
http://www.domain.com/
http://www.doamin.co.uk/
http://www.yahoo.com/
http://www.google.au/
https:// username:password@domain.com/
ftp:// user:password@domain.com/path/
https://www.blah1.subdoamin.doamin.tld /
domain.tld /#anchor
doamin.tld /?query = 123
domain.co.uk/
domain.tld
http://www.domain.tld/index.php?var1 = blah
http://www.domain.tld /path/to/index.ext
mailto://user@unkwndesign.com
并为所有组件提供命名捕获:
计划,例如.http https ftp ftps callto mailto和任何其他未列出的
用户名
密码
主机名,包括子域,域和tld
路径,例如/ images/profile/
filename,例如file.ext
查询字符串.?foo = bar&bar = foo
片段例如.#锚
使用主机名作为唯一的必填字段.
我们可以假设这是来自特定要求网址的表单,并且不会用于在文本中查找链接.
我正在尝试编写一个正则表达式,它将从URL捕获域和路径.我试过了:
https?:\/\/(.+)(\/.*)
Run Code Online (Sandbox Code Playgroud)
Match 1
0. google.com
1. /foo
Run Code Online (Sandbox Code Playgroud)
但不是我对http://example.com/foo/bar的期望:
预期:
Match 1
0. google.com
1. /foo/bar
Run Code Online (Sandbox Code Playgroud)
实际:
Match 1
0. google.com/foo
1. /bar
Run Code Online (Sandbox Code Playgroud)
我究竟做错了什么?
是否有一种简单的符合标准的方法来检查URL字符串是否是有效格式?要么通过特定的URL类型类,要么可能有人可以告诉我如何对它进行正则表达式验证?
如何检查是否有任何URL包含一个或多个参数?
例如,如果URL是
.../PaymentGatewayManager?mode=5015
Run Code Online (Sandbox Code Playgroud)
那么我们应该能够知道URL包含一个参数.或者如果URL是
.../PaymentGatewayManager
Run Code Online (Sandbox Code Playgroud)
那么我们应该能够知道URL不包含任何参数.或者如果URL是
.../PaymentGatewayManager?mode=5015&test=456123&abc=78
Run Code Online (Sandbox Code Playgroud)
然后我们应该能够知道URL包含三个参数,并且我们还应该能够使用Java中的任何正则表达式知道参数名称和值.
url="www.example.com/thedubaimall"
Run Code Online (Sandbox Code Playgroud)
我想保存之后的一切,com/我只需要thedubaimall
我怎样才能用正则表达式做到这一点?
我有这样的完整链接:
http://localhost:8080/suffix/rest/of/link
Run Code Online (Sandbox Code Playgroud)
如何在 Java 中编写正则表达式,它只返回带有后缀的 url 的主要部分:http://localhost/suffix而没有:/rest/of/link?
我假设我需要在第三次出现'/'标记(包括)后删除整个文本。我想按以下方式进行,但我不太了解正则表达式,请您帮忙如何正确编写正则表达式?
String appUrl = fullRequestUrl.replaceAll("(.*\\/{2})", ""); //this removes 'http://' but this is not my case
Run Code Online (Sandbox Code Playgroud) 我有一个URL,例如: http://www.di.fm/calendar/event/40351
我想用正则表达式解析此URL,并在域之后检索该部分.在这种情况下 :calendar/event/40351
哈希之后还可以有其他信息:(例如calendar/event/40351#event-info)
我一直很难使用正则表达式,但我认为现在我有合法的需求,而且我不确定是否可以使用它们来实现它.
我想要一个正则表达式,在执行时,返回给定URL的所有组件.它不是用于验证格式:前提条件是传递正确的URL(通常是它location.href).
期望的组件是:
奖金:
例子:
/regex/.exec('http://www.stackoverflow.com/questions/1/regex-for-getting-url-components-in-javascript') --> ["http://stackoverflow.com/questions/30868359/regex-for-getting-url-components-in-javascript", "http", "www.stackoverflow.com", undefined, "questions/30868359/regex-for-getting-url-components-in-javascript"]
/regex/.exec('https://localhost:8080/?a=1&b=2') --> ["http://www.stackoverflow.com/questions/1/regex-for-getting-url-components-in-javascript", "https", "localhost", "8080", "", "a=1&b=2"]
编辑:
为了澄清,我需要的是一个小代码,它创建一个代表URl的对象.然后我必须能够修改参数,模式等组件,并再次将结果作为字符串.AFAIK,我不能用本地位置对象做这个,但我一定是错的.
代码的大小必须特别小,因为它必须在页面的标题中同步加载.它最终可能会被复制到每个页面而不是作为外部文件包含在内.所以,起初,我更喜欢不依赖外部依赖.
regex ×8
java ×3
python ×2
url ×2
arguments ×1
c# ×1
c++ ×1
javascript ×1
parameters ×1
string ×1
validation ×1