使用Java Regex,如何检查字符串是否包含集合中的任何单词?

use*_*116 35 java regex string-matching

我有一套话说 - 苹果,橘子,梨,香蕉,猕猴桃

我想检查一个句子是否包含上面列出的任何单词,如果是,我想找到匹配的单词.我怎样才能在Regex中实现这一目标?

我目前正在为每组单词调用String.indexOf().我假设这不像正则表达式匹配那么有效吗?

Dav*_*ebb 48

TL; DR对于简单的子串contains()最好,但只匹配整个单词正则表达式可能更好.

查看哪种方法更有效的最佳方法是测试它.

您可以使用String.contains()而不是String.indexOf()简化非正则表达式代码.

要搜索不同的单词,正则表达式如下所示:

apple|orange|pear|banana|kiwi
Run Code Online (Sandbox Code Playgroud)

|作为OR正则表达式的作品.

我非常简单的测试代码如下所示:

public class TestContains {

   private static String containsWord(Set<String> words,String sentence) {
     for (String word : words) {
       if (sentence.contains(word)) {
         return word;
       }
     }

     return null;
   }

   private static String matchesPattern(Pattern p,String sentence) {
     Matcher m = p.matcher(sentence);

     if (m.find()) {
       return m.group();
     }

     return null;
   }

   public static void main(String[] args) {
     Set<String> words = new HashSet<String>();
     words.add("apple");
     words.add("orange");
     words.add("pear");
     words.add("banana");
     words.add("kiwi");

     Pattern p = Pattern.compile("apple|orange|pear|banana|kiwi");

     String noMatch = "The quick brown fox jumps over the lazy dog.";
     String startMatch = "An apple is nice";
     String endMatch = "This is a longer sentence with the match for our fruit at the end: kiwi";

     long start = System.currentTimeMillis();
     int iterations = 10000000;

     for (int i = 0; i < iterations; i++) {
       containsWord(words, noMatch);
       containsWord(words, startMatch);
       containsWord(words, endMatch);
     }

     System.out.println("Contains took " + (System.currentTimeMillis() - start) + "ms");
     start = System.currentTimeMillis();

     for (int i = 0; i < iterations; i++) {
       matchesPattern(p,noMatch);
       matchesPattern(p,startMatch);
       matchesPattern(p,endMatch);
     }

     System.out.println("Regular Expression took " + (System.currentTimeMillis() - start) + "ms");
   }
}
Run Code Online (Sandbox Code Playgroud)

我得到的结果如下:

Contains took 5962ms
Regular Expression took 63475ms
Run Code Online (Sandbox Code Playgroud)

显然,时间将根据搜索的单词数量和搜索的字符串而有所不同,但contains()对于像这样的简单搜索,它似乎比正则表达式快10倍.

通过使用正则表达式在另一个字符串中搜索字符串,你正在使用大锤来破解坚果,所以我想我们不应该对它的速度感到惊讶.保存正则表达式,以便在您要查找的模式更复杂时使用.

一种情况可能要使用正则表达式是,如果indexOf()和contains()不会做的工作,因为你只想匹配整个单词,而不仅仅是子,如要匹配pear,但没有spears.正则表达式很好地处理了这种情况,因为它们具有单词边界的概念.

在这种情况下,我们将模式更改为:

\b(apple|orange|pear|banana|kiwi)\b
Run Code Online (Sandbox Code Playgroud)

该\b说只匹配开头或一个字的结束和支架组或表达式在一起.

注意,在代码中定义此模式时,需要使用另一个反斜杠转义反斜杠:

 Pattern p = Pattern.compile("\\b(apple|orange|pear|banana|kiwi)\\b");
Run Code Online (Sandbox Code Playgroud)


Gui*_*let 7

我不认为正则表达式在性能方面会做得更好,但您可以按照以下方式使用它:

Pattern p = Pattern.compile("(apple|orange|pear)");
Matcher m = p.matcher(inputString);
while (m.find()) {
   String matched = m.group(1);
   // Do something
}
Run Code Online (Sandbox Code Playgroud)

  • 你不能读吗?我从来没有说它效率很高. (6认同)
  • @deporter答案的目的是为如何解决问题提供一个很好的暗示,不提供完美,闪亮,世界一流的解决方案.它可以很容易地进行改进,就可读性而言,如果你有200个字符串(另一个原因是不使用regexp),你可以在`StringBuilder`中使用for循环和连接.我认为我的答案提供了足够的味道. (2认同)

小智 6

这是我找到的最简单的解决方案(与通配符匹配):

boolean a = str.matches(".*\\b(wordA|wordB|wordC|wordD|wordE)\\b.*");
Run Code Online (Sandbox Code Playgroud)