以下GLSL计算着色器只是复制inImage到outImage.它源于更复杂的后处理过程.
在前几行中main(),单个线程将64个像素的数据加载到共享阵列中.然后,在同步之后,64个线程中的每一个将一个像素写入输出图像.
根据我的同步方式,我会得到不同的结果.我原本以为memoryBarrierShared()是正确的调用,但它会产生以下结果:

这与没有同步或使用相同的结果相同memoryBarrier().
如果我使用barrier(),我得到以下(所需)结果:

条带宽度为32像素,如果我将工作组大小更改为小于或等于32的任何值,我会得到正确的结果.
这里发生了什么?我误解了目的memoryBarrierShared()吗?为什么要barrier()工作?
#version 430
#define SIZE 64
layout (local_size_x = SIZE, local_size_y = 1, local_size_z = 1) in;
layout(rgba32f) uniform readonly image2D inImage;
uniform writeonly image2D outImage;
shared vec4 shared_data[SIZE];
void main() {
ivec2 base = ivec2(gl_WorkGroupID.xy * gl_WorkGroupSize.xy);
ivec2 my_index = base + ivec2(gl_LocalInvocationID.x,0);
if (gl_LocalInvocationID.x == 0) {
for (int i = 0; i < SIZE; i++) {
shared_data[i] = imageLoad(inImage, base + ivec2(i,0));
}
}
// with no synchronization: stripes
// memoryBarrier(); // stripes
// memoryBarrierShared(); // stripes
// barrier(); // works
imageStore(outImage, my_index, shared_data[gl_LocalInvocationID.x]);
}
Run Code Online (Sandbox Code Playgroud)
Chr*_*ica 22
图像加载存储和朋友的问题是,实现不再能够确定着色器仅更改其专用输出值的数据(例如片段着色器之后的帧缓冲).这更适用于计算着色器,它没有专用输出,只能通过将数据写入可写存储(如图像,存储缓冲区或原子计数器)来输出内容.这可能需要在各个传递之间进行手动同步,否则尝试访问纹理的片段着色器可能没有将最新数据写入具有前一遍图像存储操作的纹理中,如计算着色器.
So it may be that your compute shader works perfectly, but it is the synchronization with the following display (or whatever) pass (that needs to read this image data somehow) that fails. For this purpose there exists the glMemoryBarrier function. Depending on how you read that image data in the display pass (or more precisely the pass that reads the image after the compute shader pass), you need to give a different flag to this function. If you read it using a texture, use GL_TEXTURE_FETCH_BARRIER_BIT?, if you use an image load again, use GL_SHADER_IMAGE_ACCESS_BARRIER_BIT?, if using glBlitFramebuffer for display, use GL_FRAMEBUFFER_BARRIER_BIT?...
Though I don't have much experience with image load/store and manual memory snynchronization and this is only what I came up with theoretically. So if anyone knows better or you already use a proper glMemoryBarrier, then feel free to correct me. Likewise does this not need to be your only error (if any). But the last two points from the linked Wiki article actually address your use case and IMHO make it clear that you need some kind of glMemoryBarrier:
Data written to image variables in one rendering pass and read by the shader in a later pass need not use
coherentvariables ormemoryBarrier(). CallingglMemoryBarrierwith theSHADER_IMAGE_ACCESS_BARRIER_BIT?set in barriers between passes is necessary.Data written by the shader in one rendering pass and read by another mechanism (e.g., vertex or index buffer pulling) in a later pass need not use
coherentvariables ormemoryBarrier(). CallingglMemoryBarrierwith the appropriate bits set in barriers between passes is necessary.
EDIT: Actually the Wiki article on compute shaders says
Shared variable access uses the rules for incoherent memory access. This means that the user must perform certain synchronization in order to ensure that shared variables are visible.
Shared variables are all implicitly declared
coherent?, so you don't need to (and can't use) that qualifier. However, you still need to provide an appropriate memory barrier.The usual set of memory barriers is available to compute shaders, but they also have access to
memoryBarrierShared()?;this barrier is specifically for shared variable ordering.groupMemoryBarrier() acts likememoryBarrier()?, ordering memory writes for all kinds of variables, but it only orders read/writes for the current work group.While all invocations within a work group are said to execute "in parallel", that doesn't mean that you can assume that all of them are executing in lock-step. If you need to ensure that an invocation has written to some variable so that you can read it, you need to synchronize execution with the invocations, not just issue a memory barrier (you still need the memory barrier though).
To synchronize reads and writes between invocations within a work group, you must employ the
barrier() function. This forces an explicit synchronization between all invocations in the work group. Execution within the work group will not proceed until all other invocations have reach this barrier. Once past thebarrier()?, all shared variables previously written across all invocations in the group will be visible.
So this actually sounds like you need the barrier there and the memoryBarrierShared is not enough (though you don't need both, as the last sentence says). The memory barrier will just synchronize the memory, but it doesn't stop the execution of the threads to cross it. Thus the threads won't read any old cached data from the shared memory if the first thread has already written something, but they can very well reach the point of reading before the first thread has tried to write anything at all.
This actually fits perfectly to the fact that for 32 and below block sizes it works and that the first 32 pixels work. At least on NVIDIA hardware 32 is the warp size and thus the number of threads that operate in perfect lock-step. So the first 32 threads (well, every block of 32 threads) always work exactly parallel (well, conceptually that is) and thus they cannot introduce any race-conditions. This is also the case why you don't actually need any synchronization if you know you work inside a single warp, a common optimization.