通过第一个冒号提取字符串

Mar*_*ler 1 regex string r regex-negation

我有一个字符串数据集,并希望提取一个子字符串,包括第一个冒号.之前我在这里发帖询问如何提取第一个冒号之后的部分:在第一个冒号处拆分字符串 下面我列出了一些解决当前问题的尝试.

我知道^[^:]+:匹配我想保留的部分,但我无法弄清楚如何提取该部分.

这是一个示例数据集和所需的结果.

my.data <- "here is: some text
here is some more.
even: more text
still more text
this text keeps: going."

my.data2 <- readLines(textConnection(my.data))

desired.result <- "here is:
0
even:
0
this text keeps:"

desired.result2 <- readLines(textConnection(desired.result))

# Here are some of my attempts

# discards line 2 and 4 but does not extract portion from lines 1,3, and 5.
ifelse( my.data2 == gsub("^[^:]+:", "", my.data2), '', my.data2)

# returns the portion I do not want rather than the portion I do want
sub("^[^:]+:", "\\1", my.data2, perl=TRUE)

# returns an entire line if it contains a colon
grep("^[^:]+:", my.data2, value=TRUE)

# identifies which rows contain a match
regexpr("^[^:]+:", my.data2)

# my attempt at anchoring the right end instead of the left end
regexpr("[^:]+:$", my.data2)
Run Code Online (Sandbox Code Playgroud)

这个早期的问题涉及返回匹配的反面.如果我从上面链接的上一个问题的解决方案开始,我还没有想出如何在R中实现这个解决方案:正则表达式相反

我最近获得了RegexBuddy来学习正则表达式.这就是我知道^[^:]+:匹配我想要的东西.我只是无法使用该信息来提取匹配项.

我知道这个stringr包裹.也许它可以提供帮助,但我更喜欢基础R的解决方案.

谢谢你的任何建议.

42-*_*42- 6

"我知道^ [^:] +:匹配我想要保留的部分,但我无法弄清楚如何提取该部分."

所以只需将parens包裹起来并在末尾添加".+ $"并使用带引用的sub

sub("(^[^:]+:).+$", "\\1", vec)

 step1 <- sub("^([^:]+:).+$", "\\1", my.data2)
 step2 <- ifelse(grepl(":", step1), step1, 0)
 step2
#[1] "here is:"         "0"                "even:"            "0"               
#[5] "this text keeps:"
Run Code Online (Sandbox Code Playgroud)

目前尚不清楚您是否希望将它们作为单独的向量元素与它们粘贴在一起使用换行符:

> step3 <- paste0(step2, collapse="\n")
> step3
[1] "here is:\n0\neven:\n0\nthis text keeps:"
> cat(step3)
here is:
0
even:
0
this text keeps:
Run Code Online (Sandbox Code Playgroud)