use*_*627 10 python bounding-box object-detection computer-vision scikit-image
在大图像中测试物体检测算法时,我们根据为地面实况矩形给出的坐标检查我们检测到的边界框.
根据Pascal VOC挑战,有这样的:
如果预测的边界框与地面实况边界框重叠超过50%,则认为该边界框是正确的,否则边界框被认为是误报检测.多次检测会受到处罚.如果系统预测了几个与单个地面实况边界框重叠的边界框,则只有一个预测被认为是正确的,其他预测被认为是误报.
这意味着我们需要计算重叠的百分比.这是否意味着地面实况框被检测到的边界框覆盖了50%?或者50%的边界框被地面真值箱吸收了?
我已经搜索了但是我没有为此找到一个标准算法 - 这是令人惊讶的,因为我会认为这是计算机视觉中非常常见的东西.(我是新手).我错过了吗?有谁知道这类问题的标准算法是什么?
Mar*_*oma 35
对于轴对齐的边界框,它相对简单:
def get_iou(bb1, bb2):
"""
Calculate the Intersection over Union (IoU) of two bounding boxes.
Parameters
----------
bb1 : dict
Keys: {'x1', 'x2', 'y1', 'y2'}
The (x1, y1) position is at the top left corner,
the (x2, y2) position is at the bottom right corner
bb2 : dict
Keys: {'x1', 'x2', 'y1', 'y2'}
The (x, y) position is at the top left corner,
the (x2, y2) position is at the bottom right corner
Returns
-------
float
in [0, 1]
"""
assert bb1['x1'] < bb1['x2']
assert bb1['y1'] < bb1['y2']
assert bb2['x1'] < bb2['x2']
assert bb2['y1'] < bb2['y2']
# determine the coordinates of the intersection rectangle
x_left = max(bb1['x1'], bb2['x1'])
y_top = max(bb1['y1'], bb2['y1'])
x_right = min(bb1['x2'], bb2['x2'])
y_bottom = min(bb1['y2'], bb2['y2'])
if x_right < x_left or y_bottom < y_top:
return 0.0
# The intersection of two axis-aligned bounding boxes is always an
# axis-aligned bounding box
intersection_area = (x_right - x_left) * (y_bottom - y_top)
# compute the area of both AABBs
bb1_area = (bb1['x2'] - bb1['x1']) * (bb1['y2'] - bb1['y1'])
bb2_area = (bb2['x2'] - bb2['x1']) * (bb2['y2'] - bb2['y1'])
# compute the intersection over union by taking the intersection
# area and dividing it by the sum of prediction + ground-truth
# areas - the interesection area
iou = intersection_area / float(bb1_area + bb2_area - intersection_area)
assert iou >= 0.0
assert iou <= 1.0
return iou
Run Code Online (Sandbox Code Playgroud)

图像来自这个答案
Mit*_*ers 26
如果您使用的是屏幕(像素)坐标,则最高投票的答案存在数学错误!几周前我提交了一个编辑,为所有读者提供了很长的解释,以便他们理解数学。但是那个编辑没有被审稿人理解并被删除,所以我再次提交了相同的编辑,但这次更简要地总结了。(更新:拒绝 2vs1因为它被认为是“实质性的变化”,呵呵)。
所以我将在这个单独的答案中用它的数学完全解释这个大问题。
所以,是的,一般来说,投票最多的答案是正确的,是计算 IoU 的好方法。但是(正如其他人也指出的那样)它的数学计算对于计算机屏幕来说是完全不正确的。你不能只做(x2 - x1) * (y2 - y1),因为这不会产生正确的面积计算。屏幕索引从像素开始,0,0到 结束width-1,height-1。屏幕坐标的范围是inclusive:inclusive(包括两端),所以像素坐标中从0到的范围10实际上是11个像素宽,因为它包括0 1 2 3 4 5 6 7 8 9 10(11项)。因此,计算屏幕坐标的区域,您必须因此+1按钮添加到每个维度,具体如下:(x2 - x1 + 1) * (y2 - y1 + 1)。
如果您正在使用范围不包含在内的其他坐标系(例如inclusive:exclusive,0to10表示“元素 0-9 但不是 10”的系统),则不需要这种额外的数学运算。但最有可能的是,您正在处理基于像素的边界框。好吧,屏幕坐标0,0从那里开始并从那里上升。
甲1920x1080屏幕从索引0(第一像素)到1919(最后一个像素水平地),并从0(第一像素)到1079(最后一个像素垂直地)。
所以如果我们在“像素坐标空间”中有一个矩形,要计算它的面积,我们必须在每个方向上加 1。否则,我们会得到面积计算的错误答案。
想象一下,我们的1920x1080屏幕有一个基于像素坐标的矩形left=0,top=0,right=1919,bottom=1079(覆盖整个屏幕上的所有像素)。
好吧,我们知道1920x1080像素就是2073600像素,它是 1080p 屏幕的正确区域。
但是如果数学错误area = (x_right - x_left) * (y_bottom - y_top),我们会得到:(1919 - 0) * (1079 - 0)= 1919 * 1079=2070601像素!那是错误的!
这就是为什么我们必须添加+1到每个计算,这为我们提供了以下修正数学:area = (x_right - x_left + 1) * (y_bottom - y_top + 1),给我们:(1919 - 0 + 1) * (1079 - 0 + 1)= 1920 * 1080=2073600像素!这确实是正确的答案!
最短的总结是:像素坐标范围是inclusive:inclusive,所以+ 1如果我们想要像素坐标范围的真实面积,我们必须添加到每个轴。
有关为什么+1需要的更多详细信息,请参阅 Jindil 的回答:https ://stackoverflow.com/a/51730512/8874388
以及这篇 pyimagesearch 文章:https ://www.pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/
而这个 GitHub 评论:https : //github.com/AlexeyAB/darknet/issues/3995#issuecomment-535697357
由于修复的数学没有被批准,任何从最高投票答案中复制代码的人都希望看到这个答案,并且能够通过简单地复制下面的错误修复断言和面积计算行来自己修复它固定inclusive:inclusive(像素)坐标范围:
assert bb1['x1'] <= bb1['x2']
assert bb1['y1'] <= bb1['y2']
assert bb2['x1'] <= bb2['x2']
assert bb2['y1'] <= bb2['y2']
................................................
# The intersection of two axis-aligned bounding boxes is always an
# axis-aligned bounding box.
# NOTE: We MUST ALWAYS add +1 to calculate area when working in
# screen coordinates, since 0,0 is the top left pixel, and w-1,h-1
# is the bottom right pixel. If we DON'T add +1, the result is wrong.
intersection_area = (x_right - x_left + 1) * (y_bottom - y_top + 1)
# compute the area of both AABBs
bb1_area = (bb1['x2'] - bb1['x1'] + 1) * (bb1['y2'] - bb1['y1'] + 1)
bb2_area = (bb2['x2'] - bb2['x1'] + 1) * (bb2['y2'] - bb2['y1'] + 1)
Run Code Online (Sandbox Code Playgroud)
您可以按torchvision如下方式计算。bbox 格式为[x1, y1, x2, y2].
import torch
import torchvision.ops.boxes as bops
box1 = torch.tensor([[511, 41, 577, 76]], dtype=torch.float)
box2 = torch.tensor([[544, 59, 610, 94]], dtype=torch.float)
iou = bops.box_iou(box1, box2)
# tensor([[0.1382]])
Run Code Online (Sandbox Code Playgroud)
对于相交距离,我们不应该添加+1以便有
intersection_area = (x_right - x_left + 1) * (y_bottom - y_top + 1)
Run Code Online (Sandbox Code Playgroud)
(与 AABB 相同)
喜欢这个pyimage 搜索帖子
我同意(x_right - x_left) x (y_bottom - y_top)在数学中使用点坐标,但由于我们处理像素,因此我认为不同。
考虑一维示例:
编辑:我最近才知道这是一种“小方块”方法。
但是,如果您将像素视为点样本(即边界框角在 matplotlib 中显然位于像素的中心),那么您不需要 +1。
请参阅此评论和此插图
一种简单的方法
from shapely.geometry import Polygon
def calculate_iou(box_1, box_2):
poly_1 = Polygon(box_1)
poly_2 = Polygon(box_2)
iou = poly_1.intersection(poly_2).area / poly_1.union(poly_2).area
return iou
box_1 = [[511, 41], [577, 41], [577, 76], [511, 76]]
box_2 = [[544, 59], [610, 59], [610, 94], [544, 94]]
print(calculate_iou(box_1, box_2))
Run Code Online (Sandbox Code Playgroud)
其结果将是0.138211...该装置13.82%。
在下面的代码片段中,我沿着第一个框的边缘构建了一个多边形。然后,我使用 Matplotlib 将多边形剪辑到第二个框。生成的多边形包含四个顶点,但我们只对左上角和右下角感兴趣,因此我采用坐标的最大值和最小值来获取边界框,并将其返回给用户。
import numpy as np
from matplotlib import path, transforms
def clip_boxes(box0, box1):
path_coords = np.array([[box0[0, 0], box0[0, 1]],
[box0[1, 0], box0[0, 1]],
[box0[1, 0], box0[1, 1]],
[box0[0, 0], box0[1, 1]]])
poly = path.Path(np.vstack((path_coords[:, 0],
path_coords[:, 1])).T, closed=True)
clip_rect = transforms.Bbox(box1)
poly_clipped = poly.clip_to_bbox(clip_rect).to_polygons()[0]
return np.array([np.min(poly_clipped, axis=0),
np.max(poly_clipped, axis=0)])
box0 = np.array([[0, 0], [1, 1]])
box1 = np.array([[0, 0], [0.5, 0.5]])
print clip_boxes(box0, box1)
Run Code Online (Sandbox Code Playgroud)