在两个可能的分隔符之一之前找到一个单词

Mai*_*ura 5 regex r

word:12335
anotherword:2323434
totallydifferentword/455
word/32
Run Code Online (Sandbox Code Playgroud)

我需要前抢字符串:或/仅使用基础R功能.我可以使用stringr但不想在我的包中添加另一个依赖项.单词可以具有可变数量的字符,但总是以(一个)分隔符结束.我不需要保留之后的内容.

Tyl*_*ker 3

也许尝试:

x <- c("word:12335", "anotherword:2323434", "totallydifferentword/455", "word/32")
lapply(strsplit(x, ":|/"), function(z) z[[1]]) #as a list
sapply(strsplit(x, ":|/"), function(z) z[[1]]) #as a string
Run Code Online (Sandbox Code Playgroud)

有一些正则表达式解决方案gsub也可以工作,但根据我处理类似问题的经验,strsplit它会不那么雄辩但更快。

我想这个正则表达式也可以工作:

gsub("([a-z]+)([/|:])([0-9]+)", "\\1", x)
Run Code Online (Sandbox Code Playgroud)

在这种情况下 gsub 更快:

Unit: microseconds
        expr    min     lq median     uq     max
1     GSUB() 19.127 21.460 22.392 23.792 106.362
2 STRSPLIT() 46.650 50.849 53.182 54.581 854.162
Run Code Online (Sandbox Code Playgroud)