God*_*del 9 nlp r machine-learning text-mining tm
假设我有基于文本的训练数据和测试数据.更具体地说,我有两个数据集 - 培训和测试 - 并且它们都有一个包含文本的列,并且对于手头的工作很感兴趣.
我在R中使用tm包来处理训练数据集中的文本列.在删除空格,标点符号和停用词之后,我阻止了语料库并最终创建了1克的文档术语矩阵,其中包含每个文档中单词的频率/计数.然后,我采取了预先确定的截止值,比如50,并且仅保留那些计数大于50的术语.
在此之后,我使用DTM和因变量(存在于训练数据中)训练一个GLMNET模型.到目前为止,一切都运行顺畅.
但是,当我想在测试数据上获得/预测模型或未来可能出现的任何新数据时,我该如何处理?
具体来说,我想要找出的是如何在新数据上创建精确的DTM?
如果新数据集没有任何与原始训练数据相似的单词,则所有术语的计数应为零(这很好).但我希望能够在任何新语料库中复制完全相同的DTM(在结构方面).
任何想法/想法?
Dmi*_*nov 14
tm有这么多的陷阱...看到更有效的text2vec和矢量化插图,它完全回答了这个问题.
对于tm这里,可能有一种更简单的方法来重建第二语料库的DTM矩阵:
crude2.dtm <- DocumentTermMatrix(crude2, control = list
(dictionary=Terms(crude1.dtm), wordLengths = c(3,10)) )
Run Code Online (Sandbox Code Playgroud)
如果我理解正确,你已经创建了一个dtm,并且你想从具有与第一个dtm相同的列(即.术语)的新文档中创建一个新的dtm.如果是这种情况,那么应该通过第一个中的术语来设置第二个dtm,可能是这样的:
首先设置一些可重现的数据......
这是你的训练数据......
library(tm)
# make corpus for text mining (data comes from package, for reproducibility)
data("crude")
corpus1 <- Corpus(VectorSource(crude[1:10]))
# process text (your methods may differ)
skipWords <- function(x) removeWords(x, stopwords("english"))
funcs <- list(tolower, removePunctuation, removeNumbers,
stripWhitespace, skipWords)
crude1 <- tm_map(corpus1, FUN = tm_reduce, tmFuns = funcs)
crude1.dtm <- DocumentTermMatrix(crude1, control = list(wordLengths = c(3,10)))
Run Code Online (Sandbox Code Playgroud)
这是你的测试数据......
corpus2 <- Corpus(VectorSource(crude[15:20]))
# process text (your methods may differ)
skipWords <- function(x) removeWords(x, stopwords("english"))
funcs <- list(tolower, removePunctuation, removeNumbers,
stripWhitespace, skipWords)
crude2 <- tm_map(corpus2, FUN = tm_reduce, tmFuns = funcs)
crude2.dtm <- DocumentTermMatrix(crude2, control = list(wordLengths = c(3,10)))
Run Code Online (Sandbox Code Playgroud)
这是你想做的事情:
现在我们只保留训练数据中存在的测试数据中的术语......
# convert to matrices for subsetting
crude1.dtm.mat <- as.matrix(crude1.dtm) # training
crude2.dtm.mat <- as.matrix(crude2.dtm) # testing
# subset testing data by colnames (ie. terms) or training data
xx <- data.frame(crude2.dtm.mat[,intersect(colnames(crude2.dtm.mat),
colnames(crude1.dtm.mat))])
Run Code Online (Sandbox Code Playgroud)
最后在测试数据中添加训练数据中不在测试数据中的术语的所有空列...
# make an empty data frame with the colnames of the training data
yy <- read.table(textConnection(""), col.names = colnames(crude1.dtm.mat),
colClasses = "integer")
# add incols of NAs for terms absent in the
# testing data but present # in the training data
# following SchaunW's suggestion in the comments above
library(plyr)
zz <- rbind.fill(xx, yy)
Run Code Online (Sandbox Code Playgroud)
zz测试文档的数据框也是如此,但其结构与训练文档相同(即相同的列,尽管其中很多都包含NA,如SchaunW所述).
那是你想要的吗?