SRC*_*SRC 5 python pytorch huggingface-transformers huggingface
我有一个简单的代码,它采用 opt6.7B 模型并对其进行微调。当我在 Google colab(Tesla T4,16GB)中运行此代码时,它运行没有任何问题。但是当我尝试在 AWS p3-2xlarge 环境(Tesla V100 GPU,16GB)中运行相同的代码时,它给出了错误。
\nRuntimeError: expected scalar type Half but found Float\nRun Code Online (Sandbox Code Playgroud)\n为了能够在单个 GPU 上运行微调,我使用 LORA 和 peft。在这两种情况下,安装方式完全相同(pip install)。我可以使用with torch.autocast("cuda"):,然后该错误就消失了。但是训练的损失变得非常奇怪,这意味着它不会逐渐减少,而是在很大的范围内(0-5)波动(如果我将模型更改为GPT-J,则损失始终保持为0),而损失是逐渐减少的对于 colab 来说减少。所以我不确定使用是否with torch.autocast("cuda"):是一件好事。
Transfromeers 版本是4.28.0.dev0两种情况。Colab 的 Torch 版本显示1.13.1+cu116,而 p3 的 Torch 版本显示 - 1.13.1(这是否意味着它没有 CUDA 支持?我怀疑,除此之外,这样做torch.cuda.is_available()显示 True)
我能看到的唯一大的区别是,对于 colab,bitsandbytes 有以下设置日志
\n===================================BUG REPORT===================================\nWelcome to bitsandbytes. For bug reports, please submit your error trace to: https://github.com/TimDettmers/bitsandbytes/issues\n================================================================================\nCUDA_SETUP: WARNING! libcudart.so not found in any environmental path. Searching /usr/local/cuda/lib64...\nCUDA SETUP: CUDA runtime path found: /usr/local/cuda/lib64/libcudart.so\nCUDA SETUP: Highest compute capability among GPUs detected: 7.5\nCUDA SETUP: Detected CUDA version 118\nRun Code Online (Sandbox Code Playgroud)\n而对于 p3 则如下
\n===================================BUG REPORT===================================\nWelcome to bitsandbytes. For bug reports, please submit your error trace to: https://github.com/TimDettmers/bitsandbytes/issues\n================================================================================\nCUDA SETUP: CUDA runtime path found: /opt/conda/envs/pytorch/lib/libcudart.so\nCUDA SETUP: Highest compute capability among GPUs detected: 7.0\nCUDA SETUP: Detected CUDA version 117\nCUDA SETUP: Loading binary /opt/conda/envs/pytorch/lib/python3.9/site-packages/bitsandbytes/libbitsandbytes_cuda117_nocublaslt.so...\nRun Code Online (Sandbox Code Playgroud)\n我缺少什么?我不会在这里发布代码。但这确实是一个非常基本的版本,它采用 opt-6.7b 并使用 LORA 和 peft 在 alpaca 数据集上对其进行微调。
\n为什么在colab中可以运行,而在p3中却不能运行?欢迎任何帮助:)
\n- - - - - - - - - - 编辑
\n我发布了一个我实际尝试过的最小代码示例
\nimport os\nos.environ["CUDA_VISIBLE_DEVICES"]="0"\nimport torch\nimport torch.nn as nn\nimport bitsandbytes as bnb\nfrom transformers import AutoTokenizer, AutoConfig, AutoModelForCausalLM\n\nmodel = AutoModelForCausalLM.from_pretrained(\n "facebook/opt-6.7b", \n load_in_8bit=True, \n device_map='auto',\n)\n\ntokenizer = AutoTokenizer.from_pretrained("facebook/opt-6.7b")\nfor param in model.parameters():\n param.requires_grad = False # freeze the model - train adapters later\n if param.ndim == 1:\n # cast the small parameters (e.g. layernorm) to fp32 for stability\n param.data = param.data.to(torch.float32)\n\nmodel.gradient_checkpointing_enable() # reduce number of stored activations\nmodel.enable_input_require_grads()\n\nclass CastOutputToFloat(nn.Sequential):\n def forward(self, x): return super().forward(x).to(torch.float32)\nmodel.lm_head = CastOutputToFloat(model.lm_head)\n\ndef print_trainable_parameters(model):\n """\n Prints the number of trainable parameters in the model.\n """\n trainable_params = 0\n all_param = 0\n for _, param in model.named_parameters():\n all_param += param.numel()\n if param.requires_grad:\n trainable_params += param.numel()\n print(\n f"trainable params: {trainable_params} || all params: {all_param} || trainable%: {100 * trainable_params / all_param}"\n )\n\nfrom peft import LoraConfig, get_peft_model \n\nconfig = LoraConfig(\n r=16,\n lora_alpha=32,\n target_modules=["q_proj", "v_proj"],\n lora_dropout=0.05,\n bias="none",\n task_type="CAUSAL_LM"\n)\n\nmodel = get_peft_model(model, config)\nprint_trainable_parameters(model)\n\nimport transformers\nfrom datasets import load_dataset\n\ntokenizer.pad_token_id = 0\nCUTOFF_LEN = 256\n\ndata = load_dataset("tatsu-lab/alpaca")\n\ndata = data.shuffle().map(\n lambda data_point: tokenizer(\n data_point['text'],\n truncation=True,\n max_length=CUTOFF_LEN,\n padding="max_length",\n ),\n batched=True\n)\n# data = load_dataset("Abirate/english_quotes")\n# data = data.map(lambda samples: tokenizer(samples['quote']), batched=True)\n\ntrainer = transformers.Trainer(\n model=model, \n train_dataset=data['train'],\n args=transformers.TrainingArguments(\n per_device_train_batch_size=4, \n gradient_accumulation_steps=4,\n warmup_steps=100, \n max_steps=400, \n learning_rate=2e-5, \n fp16=True,\n logging_steps=1, \n output_dir='outputs'\n ),\n data_collator=transformers.DataCollatorForLanguageModeling(tokenizer, mlm=False)\n)\n\nmodel.config.use_cache = False # silence the warnings. Please re-enable for inference!\ntrainer.train()\nRun Code Online (Sandbox Code Playgroud)\n这是完整的堆栈跟踪
\n/tmp/ipykernel_24622/2601578793.py:2 in <module> \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 [Errno 2] No such file or directory: '/tmp/ipykernel_24622/2601578793.py' \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:1639 in train \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 1636 \xe2\x94\x82 \xe2\x94\x82 inner_training_loop = find_executable_batch_size( \xe2\x94\x82\n\xe2\x94\x82 1637 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 self._inner_training_loop, self._train_batch_size, args.auto_find_batch_size \xe2\x94\x82\n\xe2\x94\x82 1638 \xe2\x94\x82 \xe2\x94\x82 ) \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 1639 \xe2\x94\x82 \xe2\x94\x82 return inner_training_loop( \xe2\x94\x82\n\xe2\x94\x82 1640 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 args=args, \xe2\x94\x82\n\xe2\x94\x82 1641 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 resume_from_checkpoint=resume_from_checkpoint, \xe2\x94\x82\n\xe2\x94\x82 1642 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 trial=trial, \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:1906 in \xe2\x94\x82\n\xe2\x94\x82 _inner_training_loop \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 1903 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 with model.no_sync(): \xe2\x94\x82\n\xe2\x94\x82 1904 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 tr_loss_step = self.training_step(model, inputs) \xe2\x94\x82\n\xe2\x94\x82 1905 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 else: \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 1906 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 tr_loss_step = self.training_step(model, inputs) \xe2\x94\x82\n\xe2\x94\x82 1907 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 1908 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 if ( \xe2\x94\x82\n\xe2\x94\x82 1909 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 args.logging_nan_inf_filter \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/transformers/trainer.py:2662 in \xe2\x94\x82\n\xe2\x94\x82 training_step \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 2659 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 loss = loss / self.args.gradient_accumulation_steps \xe2\x94\x82\n\xe2\x94\x82 2660 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 2661 \xe2\x94\x82 \xe2\x94\x82 if self.do_grad_scaling: \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 2662 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 self.scaler.scale(loss).backward() \xe2\x94\x82\n\xe2\x94\x82 2663 \xe2\x94\x82 \xe2\x94\x82 elif self.use_apex: \xe2\x94\x82\n\xe2\x94\x82 2664 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 with amp.scale_loss(loss, self.optimizer) as scaled_loss: \xe2\x94\x82\n\xe2\x94\x82 2665 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 scaled_loss.backward() \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/_tensor.py:488 in backward \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 485 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 create_graph=create_graph, \xe2\x94\x82\n\xe2\x94\x82 486 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 inputs=inputs, \xe2\x94\x82\n\xe2\x94\x82 487 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 ) \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 488 \xe2\x94\x82 \xe2\x94\x82 torch.autograd.backward( \xe2\x94\x82\n\xe2\x94\x82 489 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 self, gradient, retain_graph, create_graph, inputs=inputs \xe2\x94\x82\n\xe2\x94\x82 490 \xe2\x94\x82 \xe2\x94\x82 ) \xe2\x94\x82\n\xe2\x94\x82 491 \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/__init__.py:197 in backward \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 194 \xe2\x94\x82 # The reason we repeat same the comment below is that \xe2\x94\x82\n\xe2\x94\x82 195 \xe2\x94\x82 # some Python versions print out the first line of a multi-line function \xe2\x94\x82\n\xe2\x94\x82 196 \xe2\x94\x82 # calls in the traceback and some print out the last line \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 197 \xe2\x94\x82 Variable._execution_engine.run_backward( # Calls into the C++ engine to run the bac \xe2\x94\x82\n\xe2\x94\x82 198 \xe2\x94\x82 \xe2\x94\x82 tensors, grad_tensors_, retain_graph, create_graph, inputs, \xe2\x94\x82\n\xe2\x94\x82 199 \xe2\x94\x82 \xe2\x94\x82 allow_unreachable=True, accumulate_grad=True) # Calls into the C++ engine to ru \xe2\x94\x82\n\xe2\x94\x82 200 \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/function.py:267 in apply \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 264 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 "Function is not allowed. You should only implement one " \xe2\x94\x82\n\xe2\x94\x82 265 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 "of them.") \xe2\x94\x82\n\xe2\x94\x82 266 \xe2\x94\x82 \xe2\x94\x82 user_fn = vjp_fn if vjp_fn is not Function.vjp else backward_fn \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 267 \xe2\x94\x82 \xe2\x94\x82 return user_fn(self, *args) \xe2\x94\x82\n\xe2\x94\x82 268 \xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 269 \xe2\x94\x82 def apply_jvp(self, *args): \xe2\x94\x82\n\xe2\x94\x82 270 \xe2\x94\x82 \xe2\x94\x82 # _forward_cls is defined by derived class \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/utils/checkpoint.py:157 in backward \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 154 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 raise RuntimeError( \xe2\x94\x82\n\xe2\x94\x82 155 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 "none of output has requires_grad=True," \xe2\x94\x82\n\xe2\x94\x82 156 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 " this checkpoint() is not necessary") \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 157 \xe2\x94\x82 \xe2\x94\x82 torch.autograd.backward(outputs_with_grad, args_with_grad) \xe2\x94\x82\n\xe2\x94\x82 158 \xe2\x94\x82 \xe2\x94\x82 grads = tuple(inp.grad if isinstance(inp, torch.Tensor) else None \xe2\x94\x82\n\xe2\x94\x82 159 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 for inp in detached_inputs) \xe2\x94\x82\n\xe2\x94\x82 160 \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/__init__.py:197 in backward \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 194 \xe2\x94\x82 # The reason we repeat same the comment below is that \xe2\x94\x82\n\xe2\x94\x82 195 \xe2\x94\x82 # some Python versions print out the first line of a multi-line function \xe2\x94\x82\n\xe2\x94\x82 196 \xe2\x94\x82 # calls in the traceback and some print out the last line \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 197 \xe2\x94\x82 Variable._execution_engine.run_backward( # Calls into the C++ engine to run the bac \xe2\x94\x82\n\xe2\x94\x82 198 \xe2\x94\x82 \xe2\x94\x82 tensors, grad_tensors_, retain_graph, create_graph, inputs, \xe2\x94\x82\n\xe2\x94\x82 199 \xe2\x94\x82 \xe2\x94\x82 allow_unreachable=True, accumulate_grad=True) # Calls into the C++ engine to ru \xe2\x94\x82\n\xe2\x94\x82 200 \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/torch/autograd/function.py:267 in apply \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 264 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 "Function is not allowed. You should only implement one " \xe2\x94\x82\n\xe2\x94\x82 265 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 "of them.") \xe2\x94\x82\n\xe2\x94\x82 266 \xe2\x94\x82 \xe2\x94\x82 user_fn = vjp_fn if vjp_fn is not Function.vjp else backward_fn \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 267 \xe2\x94\x82 \xe2\x94\x82 return user_fn(self, *args) \xe2\x94\x82\n\xe2\x94\x82 268 \xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 269 \xe2\x94\x82 def apply_jvp(self, *args): \xe2\x94\x82\n\xe2\x94\x82 270 \xe2\x94\x82 \xe2\x94\x82 # _forward_cls is defined by derived class \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 /opt/conda/envs/pytorch/lib/python3.9/site-packages/bitsandbytes/autograd/_functions.py:456 in \xe2\x94\x82\n\xe2\x94\x82 backward \xe2\x94\x82\n\xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 453 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82\n\xe2\x94\x82 454 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 elif state.CB is not None: \xe2\x94\x82\n\xe2\x94\x82 455 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 CB = state.CB.to(ctx.dtype_A, copy=True).mul_(state.SCB.unsqueeze(1).mul \xe2\x94\x82\n\xe2\x94\x82 \xe2\x9d\xb1 456 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 grad_A = torch.matmul(grad_output, CB).view(ctx.grad_shape).to(ctx.dtype \xe2\x94\x82\n\xe2\x94\x82 457 \xe2\x94\x82 \xe2\x94\x82 \xe2\x94\x82 elif state.CxB is
我有同样的错误。在谷歌上搜索后,终于通过在我的火车方法之前添加带有 torch.autocast("cuda"): 的代码得到了解决方案。像这样:
with torch.autocast("cuda"):
trainer.train()
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
3123 次 |
| 最近记录: |