ERROR:torch.distributed.elastic.multiprocessing.api:failed (exitcode: -9) local_rank: 0
Author: cpt9m0Created Apr 25, 2023Updated Jun 15, 2025
I am trying to re-train alpaca on the following machine:

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 525.105.17 Driver Version: 525.105.17 CUDA Version: 12.0 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|===============================+======================+======================|
| 0 Tesla T4 Off | 00000000:00:1B.0 Off | 0 |
| N/A 24C P0 25W / 70W | 2MiB / 15360MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
| 1 Tesla T4 Off | 00000000:00:1C.0 Off | 0 |
| N/A 25C P0 24W / 70W | 2MiB / 15360MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
| 2 Tesla T4 Off | 00000000:00:1D.0 Off | 0 |
| N/A 25C P0 26W / 70W | 2MiB / 15360MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
| 3 Tesla T4 Off | 00000000:00:1E.0 Off | 0 |
| N/A 25C P0 25W / 70W | 2MiB / 15360MiB | 7% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
+-----------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=============================================================================|
| No running processes found |
+-----------------------------------------------------------------------------+Here is my command to start training:
#!/bin/bash
CUDA_LAUNCH_BLOCKING=1 torchrun --nproc_per_node=4 --master_port=9292 train.py \
--model_name_or_path ./models/llama-7b \
--data_path ./alpaca_data.json \
--fp16 True \
--output_dir alpaca_out \
--num_train_epochs 3 \
--per_device_train_batch_size 4 \
--per_device_eval_batch_size 4 \
--gradient_accumulation_steps 8 \
--evaluation_strategy "no" \
--save_strategy "steps" \
--save_steps 2000 \
--save_total_limit 1 \
--learning_rate 2e-5 \
--weight_decay 0. \
--warmup_ratio 0.03 \
--deepspeed "./configs/default_offload_opt_param.json" \
--tf32 FalseBut I got the following errors:
WARNING:torch.distributed.elastic.multiprocessing.api:Sending process 41388 closing signal SIGTERM
WARNING:torch.distributed.elastic.multiprocessing.api:Sending process 41389 closing signal SIGTERM
WARNING:torch.distributed.elastic.multiprocessing.api:Sending process 41390 closing signal SIGTERM
ERROR:torch.distributed.elastic.multiprocessing.api:failed (exitcode: -9) local_rank: 0 (pid: 41387) of binary: /home/ubuntu/ali/venv/bin/python3
Traceback (most recent call last):
File "/home/ubuntu/ali/venv/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/home/ubuntu/ali/venv/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 346, in wrapper
return f(*args, **kwargs)
File "/home/ubuntu/ali/venv/lib/python3.10/site-packages/torch/distributed/run.py", line 762, in main
run(args)
File "/home/ubuntu/ali/venv/lib/python3.10/site-packages/torch/distributed/run.py", line 753, in run
elastic_launch(
File "/home/ubuntu/ali/venv/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/home/ubuntu/ali/venv/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 246, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
========================================================
train.py FAILED
--------------------------------------------------------
Failures:
<NO_OTHER_FAILURES>
--------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2023-04-25_18:25:09
host : *********************
rank : 0 (local_rank: 0)
exitcode : -9 (pid: 41387)
error_file: <N/A>
traceback : Signal 9 (SIGKILL) received by PID 41387
========================================================Could you please help me?
Source: tatsu-lab/stanford_alpaca