相关疑难解决方法(0)

PyTorch CUDA 错误:遇到非法内存访问

使用 CUDA 相对较新。经过一段看似随机的时间后,我不断收到以下错误: RuntimeError: CUDA error: an invalid memory access was Been应付

我见过人们提出诸如使用cuda.set_device()而不是cuda.device()设置之类的建议torch.backends.cudnn.benchmark = False

但我似乎无法让错误消失。这是我的一些代码: torch.cuda.set_device(torch.device('cuda:0')) torch.backends.cudnn.benchmark = False

class LSTM(nn.Module):
    def __init__(self, input_dim, hidden_dim, num_layers, output_dim):
        super(LSTM, self).__init__()
        self.hidden_dim = hidden_dim
        self.num_layers = num_layers

        self.lstm = nn.LSTM(input_dim, hidden_dim, num_layers, batch_first=True, dropout=0.2)

        self.fc = nn.Linear(hidden_dim, output_dim)

    def forward(self, x):
        h0 = torch.zeros(self.num_layers, x.size(0), self.hidden_dim).requires_grad_().cuda()
        c0 = torch.zeros(self.num_layers, x.size(0), self.hidden_dim).requires_grad_().cuda()

        out, (hn, cn) = self.lstm(x, (h0.detach(), c0.detach()))

        out = self.fc(out[:, -1, :]) 
        
        return out

    def …
Run Code Online (Sandbox Code Playgroud)

python pytorch

15
推荐指数
2
解决办法
5万
查看次数

清除对象后,为何仍在使用GPU中的内存?

从零使用开始:

>>> import gc
>>> import GPUtil
>>> import torch
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% |  0% |
|  1 |  0% |  0% |
|  2 |  0% |  0% |
|  3 |  0% |  0% |
Run Code Online (Sandbox Code Playgroud)

然后创建一个足够大的张量并占用内存:

>>> x = torch.rand(10000,300,200).cuda()
>>> GPUtil.showUtilization()
| ID | GPU | MEM |
------------------
|  0 |  0% | 26% |
|  1 |  0% |  0% |
|  2 | …
Run Code Online (Sandbox Code Playgroud)

python garbage-collection memory-leaks gpu pytorch

9
推荐指数
3
解决办法
570
查看次数

CUDA 错误:在 Colab 上触发设备端断言

我正在尝试在启用 GPU 的情况下在 Google Colab 上初始化张量。

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')

t = torch.tensor([1,2], device=device)
Run Code Online (Sandbox Code Playgroud)

但我收到了这个奇怪的错误。
RuntimeError: CUDA error: device-side assert triggered CUDA kernel errors might be asynchronously reported at some other API call,so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1

即使将该环境变量设置为 1 似乎也没有显示任何进一步的细节。
有人遇到过这个问题吗?

pytorch tensor google-colaboratory

7
推荐指数
3
解决办法
1万
查看次数