我在R中使用包arules来生成关联规则.我想限制规则,以便在左侧只有一个特定的元素,让我们称之为"potatoe".
如果我这样做:
rules <- apriori(dtm.mat, parameter = list(sup = 0.4, conf =
0.9,target="rules"), appearance = list(lhs = c("potatoe")))
Run Code Online (Sandbox Code Playgroud)
我在lhs上得到"potatoe",但也包括所有其他类型的东西.如何强制规则只包含一个元素?参数maxlen没有做我想要的,因为,据我所知,我不能指定应用于左边元素的maxlen.
假设您已经生成了规则("问题",在您的问题中),这里是如何对其进行子集化.基本上,您必须将数据强制转换为数据框,然后对其进行子集化.
#Here are the original rules generated with some data I created
# categories are "G", "T", "D", and "potatoe"
> inspect(rules);
lhs rhs support confidencelift
1 {} => {T} 0.3333333 0.3333333 1.0000000
2 {} => {G} 0.5000000 0.5000000 1.0000000
3 {} => {potatoe} 0.5000000 0.5000000 1.0000000
4 {} => {D} 0.5000000 0.5000000 1.0000000
5 {T} => {G} 0.1666667 0.5000000 1.0000000
6 {G} => {T} 0.1666667 0.3333333 1.0000000
7 {T} => {D} 0.1666667 0.5000000 1.0000000
8 {D} => {T} 0.1666667 0.3333333 1.0000000
9 {G} => {potatoe} 0.1666667 0.3333333 0.6666667
10 {potatoe} => {G} 0.1666667 0.3333333 0.6666667
11 {potatoe} => {D} 0.3333333 0.6666667 1.3333333
12 {D} => {potatoe} 0.3333333 0.6666667 1.3333333
#Coerce into data frame
as(rules, "data.frame");
#Restrict LHS to only certain value (here, "potatoe")
rules_subset <- subset(rules, (lhs %in% c("potatoe")));
#Check to see subset rules
inspect(rules_subset);
lhs rhs support confidencelift
1 {potatoe} => {G} 0.1666667 0.3333333 0.6666667
2 {potatoe} => {D} 0.3333333 0.6666667 1.3333333
Run Code Online (Sandbox Code Playgroud)
该方法还允许任意多个LHS值,而不仅仅是一个.比我之前提出的答案容易得多.
use*_*675 -6
这是一种方法:
- 使用检查()生成规则列表。
- 将所有规则复制到文本编辑器中。
- 另存为 .txt 文件。
- 在 Excel 中作为固定宽度分隔文件打开。
- 过滤 LHS 以仅包含“土豆”[原文如此]。
可能有更简单的方法,但至少您不必在左侧手动搜索“potatoe”[原文如此]。