在 dfm 中如何检测非英语单词并将其删除?
\ndftest <- data.frame(id = 1:3, \n text = c("Holla this is a spanish word", \n "English online here", \n "Bonjour, comment \xc3\xa7a va?"))\nRun Code Online (Sandbox Code Playgroud)\ndfm 的构造示例如下:
\ntestDfm <- dftest$text %>%\n tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE) %>% %>% tokens_wordstem() %>%\n dfm()\nRun Code Online (Sandbox Code Playgroud)\n我发现 textcat 包作为替代解决方案,但在真实数据集中有很多情况,其中整行都是英语,它仅将其识别为另一种语言的字符。是否有其他方法可以使用 quanteda 在 dfm 中的数据帧或标记中查找非英语行?
\n您可以使用所有英语单词的单词列表来完成此操作。存在这种情况的地方之一是在hunspellpacakges 中,它用于拼写检查。
library(quanteda)\n# find the path in which the right dictionary file is stored\nhunspell::dictionary(lang = "en_US")\n#> <hunspell dictionary>\n#> affix: /home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.aff \n#> dictionary: /home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.dic \n#> encoding: UTF-8 \n#> wordchars: \xe2\x80\x99 \n#> added: 0 custom words\n\n# read this into a vector\nenglish_words <- readLines("/home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.dic") %>% \n# the vector contains extra information on the words, which is removed\n gsub("/.+", "", .)\n\n# let's display a sample of the words\nset.seed(1)\nsample(english_words, 50)\n#> [1] "furnace" "steno" "Hadoop" "alumna" \n#> [5] "gonorrheal" "multichannel" "biochemical" "Riverside" \n#> [9] "granddad" "glum" "exasperation" "restorative" \n#> [13] "appropriate" "submarginal" "Nipponese" "hotting" \n#> [17] "solicitation" "pillbox" "mealtime" "thunderbolt" \n#> [21] "chaise" "Milan" "occidental" "hoeing" \n#> [25] "debit" "enlightenment" "coachload" "entreating" \n#> [29] "grownup" "unappreciative" "egret" "barre" \n#> [33] "Queen" "Tammany" "Goodyear" "horseflesh" \n#> [37] "roar" "fictionalization" "births" "mediator" \n#> [41] "resitting" "waiter" "instructive" "Baez" \n#> [45] "Muenster" "sleepless" "motorbike" "airsick" \n#> [49] "leaf" "belie"\nRun Code Online (Sandbox Code Playgroud)\n有了这个向量,理论上应该包含所有英语单词,但只包含英语单词,我们可以删除非英语标记:
\ntestDfm <- dftest$text %>%\n tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE) %>%\n tokens_keep(english_words, valuetype = "fixed") %>% \n tokens_wordstem() %>%\n dfm()\n\ntestDfm\n#> Document-feature matrix of: 3 documents, 9 features (66.7% sparse).\n#> features\n#> docs this a spanish word english onlin here comment va\n#> text1 1 1 1 1 0 0 0 0 0\n#> text2 0 0 0 0 1 1 1 0 0\n#> text3 0 0 0 0 0 0 0 1 1\nRun Code Online (Sandbox Code Playgroud)\n正如您所看到的,这工作得很好,但并不完美。“\xc3\xa7a va”中的“va”与“comment”一样被选为英语单词。因此,您要做的就是找到正确的单词列表和/或清理它。您还可以考虑删除删除了太多单词的文本。
\n