在 dfm 中查找非英语标记并将其删除

rek*_*rek 0 r quanteda

在 dfm 中如何检测非英语单词并将其删除?

\n
dftest <- data.frame(id = 1:3, \n                     text = c("Holla this is a spanish word", \n                              "English online here", \n                              "Bonjour, comment \xc3\xa7a va?"))\n
Run Code Online (Sandbox Code Playgroud)\n

dfm 的构造示例如下:

\n
testDfm <- dftest$text %>%\n             tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE)  %>%  %>% tokens_wordstem() %>%\n             dfm()\n
Run Code Online (Sandbox Code Playgroud)\n

我发现 textcat 包作为替代解决方案,但在真实数据集中有很多情况,其中整行都是英语,它仅将其识别为另一种语言的字符。是否有其他方法可以使用 quanteda 在 dfm 中的数据帧或标记中查找非英语行?

\n

JBG*_*ber 5

您可以使用所有英语单词的单词列表来完成此操作。存在这种情况的地方之一是在hunspellpacakges 中,它用于拼写检查。

\n
library(quanteda)\n# find the path in which the right dictionary file is stored\nhunspell::dictionary(lang = "en_US")\n#> <hunspell dictionary>\n#>  affix: /home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.aff \n#>  dictionary: /home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.dic \n#>  encoding: UTF-8 \n#>  wordchars: \xe2\x80\x99 \n#>  added: 0 custom words\n\n# read this into a vector\nenglish_words <- readLines("/home/johannes/R/x86_64-pc-linux-gnu-library/4.0/hunspell/dict/en_US.dic") %>% \n# the vector contains extra information on the words, which is removed\n  gsub("/.+", "", .)\n\n# let's display a sample of the words\nset.seed(1)\nsample(english_words, 50)\n#>  [1] "furnace"          "steno"            "Hadoop"           "alumna"          \n#>  [5] "gonorrheal"       "multichannel"     "biochemical"      "Riverside"       \n#>  [9] "granddad"         "glum"             "exasperation"     "restorative"     \n#> [13] "appropriate"      "submarginal"      "Nipponese"        "hotting"         \n#> [17] "solicitation"     "pillbox"          "mealtime"         "thunderbolt"     \n#> [21] "chaise"           "Milan"            "occidental"       "hoeing"          \n#> [25] "debit"            "enlightenment"    "coachload"        "entreating"      \n#> [29] "grownup"          "unappreciative"   "egret"            "barre"           \n#> [33] "Queen"            "Tammany"          "Goodyear"         "horseflesh"      \n#> [37] "roar"             "fictionalization" "births"           "mediator"        \n#> [41] "resitting"        "waiter"           "instructive"      "Baez"            \n#> [45] "Muenster"         "sleepless"        "motorbike"        "airsick"         \n#> [49] "leaf"             "belie"\n
Run Code Online (Sandbox Code Playgroud)\n

有了这个向量,理论上应该包含所有英语单词,但只包含英语单词,我们可以删除非英语标记:

\n
testDfm <- dftest$text %>%\n  tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE)  %>%\n  tokens_keep(english_words, valuetype = "fixed") %>% \n  tokens_wordstem() %>%\n  dfm()\n\ntestDfm\n#> Document-feature matrix of: 3 documents, 9 features (66.7% sparse).\n#>        features\n#> docs    this a spanish word english onlin here comment va\n#>   text1    1 1       1    1       0     0    0       0  0\n#>   text2    0 0       0    0       1     1    1       0  0\n#>   text3    0 0       0    0       0     0    0       1  1\n
Run Code Online (Sandbox Code Playgroud)\n

正如您所看到的,这工作得很好,但并不完美。“\xc3\xa7a va”中的“va”与“comment”一样被选为英语单词。因此,您要做的就是找到正确的单词列表和/或清理它。您还可以考虑删除删除了太多单词的文本。

\n