如何使用std :: regex匹配多个结果

Ant*_*ron 19 c++ regex

例如.如果我有一个像"第一个第二个第三个"的字符串,我想在一个操作中匹配每个单词,逐个输出.

我只是认为"(\ b\S*\b){0,}"会起作用.但实际上并没有.

我该怎么办?

这是我的代码:

#include<iostream>
#include<string>
using namespace std;
int main()
{
    regex exp("(\\b\\S*\\b)");
    smatch res;
    string str = "first second third forth";
    regex_search(str, res, exp);
    cout << res[0] <<" "<<res[1]<<" "<<res[2]<<" "<<res[3]<< endl;
}   
Run Code Online (Sandbox Code Playgroud)

我期待着你的帮助.:)

St0*_*0fF 25

只需在regex_searching上迭代字符串,如下所示:

{
    regex exp("(\\b\\S*\\b)");
    smatch res;
    string str = "first second third forth";

    string::const_iterator searchStart( str.cbegin() );
    while ( regex_search( searchStart, str.cend(), res, exp ) )
    {
        cout << ( searchStart == str.cbegin() ? "" : " " ) << res[0];  
        searchStart = res.suffix().first;
    }
    cout << endl;
}
Run Code Online (Sandbox Code Playgroud)

  • 这几乎是正确的。您可能已经注意到了“+=”;)这导致了这样一个事实:“res.position()”是相对于搜索的,而不是原始字符串。所以在第一轮循环的情况下你的话是正确的。 (2认同)
  • 到目前为止,这是对我有意义的唯一解释.谢谢! (2认同)
  • 您还可以使用`searchStart = res.suffix().first`将迭代器移动到最后一个匹配后的第一个字母,而不是`searchStart + = res.position()+ res.length()`如果它更清晰一点. (2认同)
  • @TimMB 谢谢你。看起来它还节省了 2 次操作(这些操作仍然会在幕后完成),因此我同意它看起来更清晰。希望可以将您的建议纳入我的答案中吗? (2认同)

her*_*tao 16

这可以做到regex的C++11.

两个方法:

  1. 您可以使用()in regex来定义捕获.

    像这样:

    string var = "first second third forth";
    
    const regex r("(.*) (.*) (.*) (.*)");  
    smatch sm;
    
    if (regex_search(var, sm, r)) {
        for (int i=1; i<sm.size(); i++) {
            cout << sm[i] << endl;
        }
    }
    
    Run Code Online (Sandbox Code Playgroud)

    现场观看:http://coliru.stacked-crooked.com/a/e1447c4cff9ea3e7

  2. 你可以使用sregex_token_iterator():

    string var = "first second third forth";
    
    regex wsaq_re("\\s+"); 
    copy( sregex_token_iterator(var.begin(), var.end(), wsaq_re, -1),
        sregex_token_iterator(),
        ostream_iterator<string>(cout, "\n"));
    
    Run Code Online (Sandbox Code Playgroud)

    现场观看:http://coliru.stacked-crooked.com/a/677aa6f0bb0612f0


小智 10

您可以使用suffix()函数,并再次搜索,直到找不到匹配项:

int main()
{
    regex exp("(\\b\\S*\\b)");
    smatch res;
    string str = "first second third forth";

    while (regex_search(str, res, exp)) {
        cout << res[0] << endl;
        str = res.suffix();
    }
}   
Run Code Online (Sandbox Code Playgroud)

  • 这样你就可以在每个循环中重新分配str.在我看来是浪费时间和堆积碎片. (2认同)

Ste*_*ven 8

sregex_token_iterator似乎是理想、有效的解决方案,但所选答案中给出的示例仍有很多不足之处。相反,我在这里找到了一些很好的例子:http : //www.cplusplus.com/reference/regex/regex_token_iterator/regex_token_iterator/

为方便起见,我复制粘贴了该页面显示的示例代码。我声称代码没有功劳。

// regex_token_iterator example
#include <iostream>
#include <string>
#include <regex>

int main ()
{
  std::string s ("this subject has a submarine as a subsequence");
  std::regex e ("\\b(sub)([^ ]*)");   // matches words beginning by "sub"

  // default constructor = end-of-sequence:
  std::regex_token_iterator<std::string::iterator> rend;

  std::cout << "entire matches:"; 
  std::regex_token_iterator<std::string::iterator> a ( s.begin(), s.end(), e );
  while (a!=rend) std::cout << " [" << *a++ << "]";
  std::cout << std::endl;

  std::cout << "2nd submatches:";
  std::regex_token_iterator<std::string::iterator> b ( s.begin(), s.end(), e, 2 );
  while (b!=rend) std::cout << " [" << *b++ << "]";
  std::cout << std::endl;

  std::cout << "1st and 2nd submatches:";
  int submatches[] = { 1, 2 };
  std::regex_token_iterator<std::string::iterator> c ( s.begin(), s.end(), e, submatches );
  while (c!=rend) std::cout << " [" << *c++ << "]";
  std::cout << std::endl;

  std::cout << "matches as splitters:";
  std::regex_token_iterator<std::string::iterator> d ( s.begin(), s.end(), e, -1 );
  while (d!=rend) std::cout << " [" << *d++ << "]";
  std::cout << std::endl;

  return 0;
}

Output:
entire matches: [subject] [submarine] [subsequence]
2nd submatches: [ject] [marine] [sequence]
1st and 2nd submatches: [sub] [ject] [sub] [marine] [sub] [sequence]
matches as splitters: [this ] [ has a ] [ as a ]
Run Code Online (Sandbox Code Playgroud)


Beh*_*z.M 6

随意使用我的代码.它将捕获所有比赛中的所有组:

vector<vector<string>> U::String::findEx(const string& s, const string& reg_ex, bool case_sensitive)
{
    regex rx(reg_ex, case_sensitive ? regex_constants::icase : 0);
    vector<vector<string>> captured_groups;
    vector<string> captured_subgroups;
    const std::sregex_token_iterator end_i;
    for (std::sregex_token_iterator i(s.cbegin(), s.cend(), rx);
        i != end_i;
        ++i)
    {
        captured_subgroups.clear();
        string group = *i;
        smatch res;
        if(regex_search(group, res, rx))
        {
            for(unsigned i=0; i<res.size() ; i++)
                captured_subgroups.push_back(res[i]);

            if(captured_subgroups.size() > 0)
                captured_groups.push_back(captured_subgroups);
        }

    }
    captured_groups.push_back(captured_subgroups);
    return captured_groups;
}
Run Code Online (Sandbox Code Playgroud)

  • 答案已根据@AxelRietschin 评论进行了更新。 (2认同)

Pet*_*vin 5

我对文档的阅读是regex_search搜索第一个匹配项,并且没有任何函数std::regex按照您的要求进行“扫描”。但是,Boost 库似乎支持这一点,如C++ tokenize a string using a regular expression 中所述