为什么 std::counting_semaphore::acquire() 在这种情况下会遇到死锁?

Sad*_*Kao 5 c++ multithreading semaphore c++20

我正在std::counting_semaphore使用 Windows 10 和 MinGW x64 在 C++20 上进行测试。

正如我从https://en.cppreference.com/w/cpp/thread/counting_semaphore了解到的,std::counting_semaphore是一个原子计数器。我们可以使用release()增加计数器,使用acquire()减少计数器。如果计数器等于 0,则线程等待。

我构建了以下简化示例来展示我的问题。如果我总是在线程release()之前,则内部计数器值(v)应始终保持在 v 和 v+1 之间,并且此代码不应遭受任何阻塞。acquire()std::counting_semaphore

当我运行这个示例代码时,它经常遇到死锁,但有时它可以正确完成。

我尝试使用std::cout消息来了解死锁情况,但是当我使用std::cout. 另一方面,当我使用 时,僵局消失了std::unique_lock。

示例如下:

#include <iostream>
#include <thread>
#include <atomic>
#include <vector>
#include <mutex>
#include <semaphore>

using namespace std::literals;

std::mutex mtx;

const int numOfThr {2};
const int numOfForLoop {1000};
const int max_smph {numOfThr* numOfForLoop *2};

std::counting_semaphore<max_smph> smph {numOfThr+1};

void thrf_TestSmph ( const int iThr )
{
    for ( int i = 0; i < numOfForLoop; ++i )
    {
//        std::unique_lock ul(mtx);
        //unique_lock can stop deadlock.

        smph.release(); //smph counter ++
        smph.acquire(); //smph counter --

//        if ( i % 1000 == 1 ) std::cout << iThr << " : " << i << "\n";
        //print out message can stop deadlock.
    }
}


int main()
{
    std::cout << "Start testing semaphore ..." << "\n\n";

    std::vector<std::thread> thrf_TestSmphVec ( numOfThr );

    for ( int iThr = 0; iThr < numOfThr; ++iThr )
    {
        thrf_TestSmphVec[iThr] = std::thread ( thrf_TestSmph, iThr );
    }

    for ( auto& thr : thrf_TestSmphVec )
    {
        if ( thr.joinable() )
            thr.join();
    }

    std::cout << "Test is done." << "\n";

    return 0;
}


Run Code Online (Sandbox Code Playgroud)

Vai*_*Man 8

更新:找到此错误报告:https://gcc.gnu.org/bugzilla/show_bug.cgi ?id=104928


这并不是一个真正的答案。

当使用 gcc 或 clang 和 libstdc++ 编译时,我可以在我的 M1 macbook air 上重现无限阻塞。打印消息并不能阻止阻塞。当使用clang和libc++编译时,程序正常完成。

include/c++/11/bits/semaphore_base.h我在libstdc++包含的标头中注意到了这段代码和注释:

    _GLIBCXX_ALWAYS_INLINE void
    _M_release(ptrdiff_t __update) noexcept
    {
      if (0 < __atomic_impl::fetch_add(&_M_counter, __update, memory_order_release))
          return;
      if (__update > 1)
          __atomic_notify_address_bare(&_M_counter, true);
      else
          __atomic_notify_address_bare(&_M_counter, true);
// FIXME - Figure out why this does not wake a waiting thread
//  __atomic_notify_address_bare(&_M_counter, false);
    }
Run Code Online (Sandbox Code Playgroud)

然后我将第一个更改return为__atomic_notify_address_bare(&_M_counter, true);,问题似乎消失了。

该评论已在此提交中提交。

    _GLIBCXX_ALWAYS_INLINE void
    _M_release(ptrdiff_t __update) noexcept
    {
      if (0 < __atomic_impl::fetch_add(&_M_counter, __update, memory_order_release))
          return;
      if (__update > 1)
          __atomic_notify_address_bare(&_M_counter, true);
      else
-         __atomic_notify_address_bare(&_M_counter, false);
+         __atomic_notify_address_bare(&_M_counter, true);
+ // FIXME - Figure out why this does not wake a waiting thread
+ //    __atomic_notify_address_bare(&_M_counter, false);
Run Code Online (Sandbox Code Playgroud)

开发团队似乎已经知道了这个问题,但他们的短期解决方案并没有解决问题。