使用hive命令更改DF中的字符串,使用sparklyr更改mutate

Lev*_*man 1 hive r gsub apache-spark sparklyr

使用Hive命令regexp_extract我试图更改以下字符串:

201703170455 to 2017-03-17:04:55
Run Code Online (Sandbox Code Playgroud)

来自:

2017031704555675 to 2017-03-17:04:55.0010
Run Code Online (Sandbox Code Playgroud)

我在sparklyr中尝试使用这个与R中的gsub一起使用的代码:

  newdf<-df%>%mutate(Time1 = regexp_extract(Time, "(....)(..)(..)(..)(..)", "\\1-\\2-\\3:\\4:\\5"))
Run Code Online (Sandbox Code Playgroud)

而这段代码:

newdf<-df%>mutate(TimeTrans = regexp_extract("(....)(..)(..)(..)(..)(....)", "\\1-\\2-\\3:\\4:\\5.\\6"))
Run Code Online (Sandbox Code Playgroud)

但根本不起作用.有关如何使用regexp_extract执行此操作的任何建议?

use*_*411 5

Apache Spark使用Java正则表达式方言而不是R,并且应该引用组$.此外,regexp_replace还用于通过数字索引提取单个组.

你可以使用regexp_replace:

df <- data.frame(time = c("201703170455", "2017031704555675"))
sdf <- copy_to(sc, df)

sdf %>% 
  mutate(time1 = regexp_replace(
    time, "^(....)(..)(..)(..)(..)$", "$1-$2-$3 $4:$5" )) %>%
  mutate(time2 = regexp_replace(
    time, "^(....)(..)(..)(..)(..)(....)$", "$1-$2-$3 $4:$5.$6"))
Run Code Online (Sandbox Code Playgroud)

Source:   query [2 x 3]
Database: spark connection master=local[8] app=sparklyr local=TRUE

# A tibble: 2 x 3
              time            time1                 time2
             <chr>            <chr>                 <chr>
1     201703170455 2017-03-17 04:55          201703170455
2 2017031704555675 2017031704555675 2017-03-17 04:55.5675
Run Code Online (Sandbox Code Playgroud)