获取 RuntimeError:预期标量类型 Half 但在 opt6.7B 微调中的 AWS P3 实例中发现 Float

SRC*_*SRC 5 python pytorch huggingface-transformers huggingface

我有一个简单的代码,它采用 opt6.7B 模型并对其进行微调。当我在 Google colab(Tesla T4,16GB)中运行此代码时,它运行没有任何问题。但是当我尝试在 AWS p3-2xlarge 环境(Tesla V100 GPU,16GB)中运行相同的代码时,它给出了错误。

\n
RuntimeError: expected scalar type Half but found Float\n
Run Code Online (Sandbox Code Playgroud)\n

为了能够在单个 GPU 上运行微调,我使用 LORA 和 peft。在这两种情况下,安装方式完全相同(pip install)。我可以使用with torch.autocast("cuda"):,然后该错误就消失了。但是训练的损失变得非常奇怪,这意味着它不会逐渐减少,而是在很大的范围内(0-5)波动(如果我将模型更改为GPT-J,则损失始终保持为0),而损失是逐渐减少的对于 colab 来说减少。所以我不确定使用是否with torch.autocast("cuda"):是一件好事。

\n

Transfromeers 版本是4.28.0.dev0两种情况。Colab 的 Torch 版本显示1.13.1+cu116,而 p3 的 Torch 版本显示 - 1.13.1(这是否意味着它没有 CUDA 支持?我怀疑,除此之外,这样做torch.cuda.is_available()显示 True)

\n

我能看到的唯一大的区别是,对于 colab,bitsandbytes 有以下设置日志

\n
===================================BUG REPORT===================================\nWelcome to bitsandbytes. For bug reports, please submit your error trace to: https://github.com/TimDettmers/bitsandbytes/issues\n================================================================================\nCUDA_SETUP: WARNING! libcudart.so not found in any environmental path. Searching /usr/local/cuda/lib64...\nCUDA SETUP: CUDA runtime path found: /usr/local/cuda/lib64/libcudart.so\nCUDA SETUP: Highest compute capability among GPUs detected: 7.5\nCUDA SETUP: Detected CUDA version 118\n
Run Code Online (Sandbox Code Playgroud)\n

而对于 p3 则如下

\n
===================================BUG REPORT===================================\nWelcome to bitsandbytes. For bug reports, please submit your error trace to: https://github.com/TimDettmers/bitsandbytes/issues\n================================================================================\nCUDA SETUP: CUDA runtime path found: /opt/conda/envs/pytorch/lib/libcudart.so\nCUDA SETUP: Highest compute capability among GPUs detected: 7.0\nCUDA SETUP: Detected CUDA version 117\nCUDA SETUP: Loading binary /opt/conda/envs/pytorch/lib/python3.9/site-packages/bitsandbytes/libbitsandbytes_cuda117_nocublaslt.so...\n
Run Code Online (Sandbox Code Playgroud)\n

我缺少什么?我不会在这里发布代码。但这确实是一个非常基本的版本,它采用 opt-6.7b 并使用 LORA 和 peft 在 alpaca 数据集上对其进行微调。

\n

为什么在colab中可以运行,而在p3中却不能运行?欢迎任何帮助:)

\n

- - - - - - - - - - 编辑

\n

我发布了一个我实际尝试过的最小代码示例

\n
import os\nos.environ["CUDA_VISIBLE_DEVICES"]="0"\nimport torch\nimport torch.nn as nn\nimport bitsandbytes as bnb\nfrom transformers import AutoTokenizer, AutoConfig, AutoModelForCausalLM\n\nmodel = AutoModelForCausalLM.from_pretrained(\n    "facebook/opt-6.7b", \n    load_in_8bit=True, \n    device_map='auto',\n)\n\ntokenizer = AutoTokenizer.from_pretrained("facebook/opt-6.7b")\nfor param in model.parameters():\n  param.requires_grad = False  # freeze the model - train adapters later\n  if param.ndim == 1:\n    # cast the small parameters (e.g. layernorm) to fp32 for stability\n    param.data = param.data.to(torch.float32)\n\nmodel.gradient_checkpointing_enable()  # reduce number of stored activations\nmodel.enable_input_require_grads()\n\nclass CastOutputToFloat(nn.Sequential):\n  def forward(self, x): return super().forward(x).to(torch.float32)\nmodel.lm_head = CastOutputToFloat(model.lm_head)\n\ndef print_trainable_parameters(model):\n    """\n    Prints the number of trainable parameters in the model.\n    """\n    trainable_params = 0\n    all_param = 0\n    for _, param in model.named_parameters():\n        all_param += param.numel()\n        if param.requires_grad:\n            trainable_params += param.numel()\n    print(\n        f"trainable params: {trainable_params} || all params: {all_param} || trainable%: {100 * trainable_params / all_param}"\n    )\n\nfrom peft import LoraConfig, get_peft_model \n\nconfig = LoraConfig(\n    r=16,\n    lora_alpha=32,\n    target_modules=["q_proj", "v_proj"],\n    lora_dropout=0.05,\n    bias="none",\n    task_type="CAUSAL_LM"\n)\n\nmodel = get_peft_model(model, config)\nprint_trainable_parameters(model)\n\nimport transformers\nfrom datasets import load_dataset\n\ntokenizer.pad_token_id = 0\nCUTOFF_LEN = 256\n\ndata = load_dataset("tatsu-lab/alpaca")\n\ndata = data.shuffle().map(\n    lambda data_point: tokenizer(\n        data_point['text'],\n        truncation=True,\n        max_length=CUTOFF_LEN,\n        padding="max_length",\n    ),\n    batched=True\n)\n# data = load_dataset("Abirate/english_quotes")\n# data = data.map(lambda samples: tokenizer(samples['quote']), batched=True)\n\ntrainer = transformers.Trainer(\n    model=model, \n    train_dataset=data['train'],\n    args=transformers.TrainingArguments(\n        per_device_train_batch_size=4, \n        gradient_accumulation_steps=4,\n        warmup_steps=100, \n        max_steps=400, \n        learning_rate=2e-5, \n        fp16=True,\n        logging_steps=1, \n        output_dir='outputs'\n    ),\n    data_collator=transformers.DataCollatorForLanguageModeling(tokenizer, mlm=False)\n)\n\nmodel.config.use_cache = False  # silence the warnings. Please re-enable for inference!\ntrainer.train()\n
Run Code Online (Sandbox Code Playgroud)\n

这是完整的堆栈跟踪

\n
/tmp/ipykernel_24622/2601578793.py:2 in <module>                                                 \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 [Errno 2] No such file or directory: '/tmp/ipykernel_24622/2601578793.py'                        \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:1639 in train        \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   1636 \xe2\x94\x82   \xe2\x94\x82   inner_training_loop = find_executable_batch_size(                                 \xe2\x94\x82\n\xe2\x94\x82   1637 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   self._inner_training_loop, self._train_batch_size, args.auto_find_batch_size  \xe2\x94\x82\n\xe2\x94\x82   1638 \xe2\x94\x82   \xe2\x94\x82   )                                                                                 \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 1639 \xe2\x94\x82   \xe2\x94\x82   return inner_training_loop(                                                       \xe2\x94\x82\n\xe2\x94\x82   1640 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   args=args,                                                                    \xe2\x94\x82\n\xe2\x94\x82   1641 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   resume_from_checkpoint=resume_from_checkpoint,                                \xe2\x94\x82\n\xe2\x94\x82   1642 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   trial=trial,                                                                  \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:1906 in              \xe2\x94\x82\n\xe2\x94\x82 _inner_training_loop                                                                             \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   1903 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   with model.no_sync():                                                 \xe2\x94\x82\n\xe2\x94\x82   1904 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   tr_loss_step = self.training_step(model, inputs)                  \xe2\x94\x82\n\xe2\x94\x82   1905 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   else:                                                                     \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 1906 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   tr_loss_step = self.training_step(model, inputs)                      \xe2\x94\x82\n\xe2\x94\x82   1907 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82                                                                             \xe2\x94\x82\n\xe2\x94\x82   1908 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   if (                                                                      \xe2\x94\x82\n\xe2\x94\x82   1909 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   args.logging_nan_inf_filter                                           \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:2662 in              \xe2\x94\x82\n\xe2\x94\x82 training_step                                                                                    \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   2659 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   loss = loss / self.args.gradient_accumulation_steps                           \xe2\x94\x82\n\xe2\x94\x82   2660 \xe2\x94\x82   \xe2\x94\x82                                                                                     \xe2\x94\x82\n\xe2\x94\x82   2661 \xe2\x94\x82   \xe2\x94\x82   if self.do_grad_scaling:                                                          \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 2662 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   self.scaler.scale(loss).backward()                                            \xe2\x94\x82\n\xe2\x94\x82   2663 \xe2\x94\x82   \xe2\x94\x82   elif self.use_apex:                                                               \xe2\x94\x82\n\xe2\x94\x82   2664 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   with amp.scale_loss(loss, self.optimizer) as scaled_loss:                     \xe2\x94\x82\n\xe2\x94\x82   2665 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   scaled_loss.backward()                                                    \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/_tensor.py:488 in backward             \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82    485 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   create_graph=create_graph,                                                \xe2\x94\x82\n\xe2\x94\x82    486 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   inputs=inputs,                                                            \xe2\x94\x82\n\xe2\x94\x82    487 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   )                                                                             \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1  488 \xe2\x94\x82   \xe2\x94\x82   torch.autograd.backward(                                                          \xe2\x94\x82\n\xe2\x94\x82    489 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   self, gradient, retain_graph, create_graph, inputs=inputs                     \xe2\x94\x82\n\xe2\x94\x82    490 \xe2\x94\x82   \xe2\x94\x82   )                                                                                 \xe2\x94\x82\n\xe2\x94\x82    491                                                                                           \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/__init__.py:197 in backward   \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   194 \xe2\x94\x82   # The reason we repeat same the comment below is that                                  \xe2\x94\x82\n\xe2\x94\x82   195 \xe2\x94\x82   # some Python versions print out the first line of a multi-line function               \xe2\x94\x82\n\xe2\x94\x82   196 \xe2\x94\x82   # calls in the traceback and some print out the last line                              \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 197 \xe2\x94\x82   Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the bac   \xe2\x94\x82\n\xe2\x94\x82   198 \xe2\x94\x82   \xe2\x94\x82   tensors, grad_tensors_, retain_graph, create_graph, inputs,                        \xe2\x94\x82\n\xe2\x94\x82   199 \xe2\x94\x82   \xe2\x94\x82   allow_unreachable=True, accumulate_grad=True)  # Calls into the C++ engine to ru   \xe2\x94\x82\n\xe2\x94\x82   200                                                                                            \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/function.py:267 in apply      \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   264 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82      "Function is not allowed. You should only implement one "   \xe2\x94\x82\n\xe2\x94\x82   265 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82      "of them.")                                                 \xe2\x94\x82\n\xe2\x94\x82   266 \xe2\x94\x82   \xe2\x94\x82   user_fn = vjp_fn if vjp_fn is not Function.vjp else backward_fn                    \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 267 \xe2\x94\x82   \xe2\x94\x82   return user_fn(self, *args)                                                        \xe2\x94\x82\n\xe2\x94\x82   268 \xe2\x94\x82                                                                                          \xe2\x94\x82\n\xe2\x94\x82   269 \xe2\x94\x82   def apply_jvp(self, *args):                                                            \xe2\x94\x82\n\xe2\x94\x82   270 \xe2\x94\x82   \xe2\x94\x82   # _forward_cls is defined by derived class                                         \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/utils/checkpoint.py:157 in backward    \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   154 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   raise RuntimeError(                                                            \xe2\x94\x82\n\xe2\x94\x82   155 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   "none of output has requires_grad=True,"                                   \xe2\x94\x82\n\xe2\x94\x82   156 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   " this checkpoint() is not necessary")                                     \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 157 \xe2\x94\x82   \xe2\x94\x82   torch.autograd.backward(outputs_with_grad, args_with_grad)                         \xe2\x94\x82\n\xe2\x94\x82   158 \xe2\x94\x82   \xe2\x94\x82   grads = tuple(inp.grad if isinstance(inp, torch.Tensor) else None                  \xe2\x94\x82\n\xe2\x94\x82   159 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82     for inp in detached_inputs)                                          \xe2\x94\x82\n\xe2\x94\x82   160                                                                                            \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/__init__.py:197 in backward   \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   194 \xe2\x94\x82   # The reason we repeat same the comment below is that                                  \xe2\x94\x82\n\xe2\x94\x82   195 \xe2\x94\x82   # some Python versions print out the first line of a multi-line function               \xe2\x94\x82\n\xe2\x94\x82   196 \xe2\x94\x82   # calls in the traceback and some print out the last line                              \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 197 \xe2\x94\x82   Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the bac   \xe2\x94\x82\n\xe2\x94\x82   198 \xe2\x94\x82   \xe2\x94\x82   tensors, grad_tensors_, retain_graph, create_graph, inputs,                        \xe2\x94\x82\n\xe2\x94\x82   199 \xe2\x94\x82   \xe2\x94\x82   allow_unreachable=True, accumulate_grad=True)  # Calls into the C++ engine to ru   \xe2\x94\x82\n\xe2\x94\x82   200                                                                                            \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/function.py:267 in apply      \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   264 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82      "Function is not allowed. You should only implement one "   \xe2\x94\x82\n\xe2\x94\x82   265 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82      "of them.")                                                 \xe2\x94\x82\n\xe2\x94\x82   266 \xe2\x94\x82   \xe2\x94\x82   user_fn = vjp_fn if vjp_fn is not Function.vjp else backward_fn                    \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 267 \xe2\x94\x82   \xe2\x94\x82   return user_fn(self, *args)                                                        \xe2\x94\x82\n\xe2\x94\x82   268 \xe2\x94\x82                                                                                          \xe2\x94\x82\n\xe2\x94\x82   269 \xe2\x94\x82   def apply_jvp(self, *args):                                                            \xe2\x94\x82\n\xe2\x94\x82   270 \xe2\x94\x82   \xe2\x94\x82   # _forward_cls is defined by derived class                                         \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/bitsandbytes/autograd/_functions.py:456 in   \xe2\x94\x82\n\xe2\x94\x82 backward                                                                                         \xe2\x94\x82\n\xe2\x94\x82                                                                                                  \xe2\x94\x82\n\xe2\x94\x82   453 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82                                                                                  \xe2\x94\x82\n\xe2\x94\x82   454 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   elif state.CB is not None:                                                     \xe2\x94\x82\n\xe2\x94\x82   455 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   CB = state.CB.to(ctx.dtype_A, copy=True).mul_(state.SCB.unsqueeze(1).mul   \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 456 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   grad_A = torch.matmul(grad_output, CB).view(ctx.grad_shape).to(ctx.dtype   \xe2\x94\x82\n\xe2\x94\x82   457 \xe2\x94\x82   \xe2\x94\x82   \xe2\x94\x82   elif state.CxB is