Clang/GCC编译器内在函数没有相应的编译器标志

inn*_*nat 7 c++ gcc sse clang intrinsics

我知道有这种类似的问题,但在编译不同的文件以不同的标志是不是在这里接受的解决办法,因为这将复杂的代码库真正的快.回答"不,这是不可能的"将会做到.


在任何版本的Clang OR GCC中,是否可以为SSE ​​2/3/3S/4.1编译内在函数,同时只允许编译器使用SSE指令集进行优化?

编辑:例如,我想编译器转_mm_load_si128()来movdqa,但在任何其他地方比这内在功能,类似于MSVC编译器是如何工作的编译器不可以做发射该指令.

EDIT2:我有动态调度程序和几个版本的单个函数,使用内在函数编写不同的指令集.使用多个文件会使维护更加困难,因为相同版本的代码将跨越多个文件,并且有很多这种类型的函数.

EDIT3:请求的示例源代码:https://github.com/AviSynth/AviSynthPlus/blob/master/avs_core/filters/resample.cpp或该文件夹中的大多数文件.

Sco*_*ttD 9

这是一种使用gcc的方法,可能是可以接受的.所有源代码都进入单个源文件.单个源文件分为几个部分.一节根据使用的命令行选项生成代码.main()和处理器功能检测等功能在本节中介绍.另一部分根据目标覆盖编译指示生成代码.可以使用目标覆盖值支持的内部函数.只有在处理器功能检测确认存在所需的处理器功能后,才应调用本节中的功能.此示例具有AVX2代码的单个覆盖部分.编写针对多个目标优化的函数时,可以使用多个覆盖部分.

// temporarily switch target so that all x64 intrinsic functions will be available
#pragma GCC push_options
#pragma GCC target ("arch=core-avx2")
#include <intrin.h>
// restore the target selection
#pragma GCC pop_options

//----------------------------------------------------------------------------
// the following functions will be compiled using default code generation
//----------------------------------------------------------------------------

int dummy1 (int a) {return a;}

//----------------------------------------------------------------------------
// the following functions will be compiled using core-avx2 code generation
// all x64 intrinc functions are available
#pragma GCC push_options
#pragma GCC target ("arch=core-avx2")
//----------------------------------------------------------------------------

static __m256i bitShiftLeft256ymm (__m256i *data, int count)
   {
   __m256i innerCarry, carryOut, rotate;

   innerCarry = _mm256_srli_epi64 (*data, 64 - count);                        // carry outs in bit 0 of each qword
   rotate     = _mm256_permute4x64_epi64 (innerCarry, 0x93);                  // rotate ymm left 64 bits
   innerCarry = _mm256_blend_epi32 (_mm256_setzero_si256 (), rotate, 0xFC);   // clear lower qword
   *data    = _mm256_slli_epi64 (*data, count);                               // shift all qwords left
   *data    = _mm256_or_si256 (*data, innerCarry);                            // propagate carrys from low qwords
   carryOut   = _mm256_xor_si256 (innerCarry, rotate);                        // clear all except lower qword
   return carryOut;
   }

//----------------------------------------------------------------------------
// the following functions will be compiled using default code generation
#pragma GCC pop_options
//----------------------------------------------------------------------------

int main (void)
    {
    return 0;
    }

//----------------------------------------------------------------------------
Run Code Online (Sandbox Code Playgroud)