Moore Threads MUSA Support
ms-swift supports Moore Threads GPUs through torch_musa and torchada. torchada redirects the torch.cuda.* APIs to torch.musa, and swift imports it automatically when both torch_musa and torchada are installed, so most training/inference scripts run on MUSA without modification.
This document uses Qwen3.5-4B as an example to demonstrate LoRA and full-parameter fine-tuning, inference, and LoRA merging with ms-swift on MTT S5000.
1. Environment Setup
1.1 Base Environment
Pull the Moore Threads training image (which ships torch_musa, DeepSpeed, flash-attention, MT-TransformerEngine, etc.) and start a container with the following command:
IMAGE_NAME=sh-harbor.mthreads.com/mcctest/training-suite:v2.1.7-rc4
docker pull ${IMAGE_NAME}
CONTAINER_NAME=swift_test
docker run -itd --privileged --network=host \
--env MTHREADS_VISIBLE_DEVICES=all \
--shm-size=80g \
-v /data:/data \
--name ${CONTAINER_NAME} \
${IMAGE_NAME} \
/bin/bash
docker exec -it ${CONTAINER_NAME} /bin/bash
Software versions used for verification in this document:
| Component | Version |
|---|---|
| GPU / Driver | MTT S5000 (80GB) × 8 / Driver 3.3.8-server |
| MUSA Toolkits | 4.3.8 |
| python | 3.10.12 |
| torch / torch_musa | 2.7.1 / 2.7.1.post1 |
| torchada | 0.1.90 |
| transformers | 5.16.1 |
| peft | 0.18.1 |
| deepspeed | 0.19.3 (MUSA build shipped in the image) |
1.2 Install ms-swift
Qwen3.5 requires transformers>=5.2.0, while the image ships an older transformers (4.57). To avoid breaking the environment that other components in the image (vLLM, sglang, etc.) depend on, we recommend creating a virtual environment that inherits the system packages and installing ms-swift there:
# Reuse torch/torch_musa/deepspeed from the image; only override the packages that need upgrading in the venv
python -m venv --without-pip --system-site-packages /home/venv_swift
source /home/venv_swift/bin/activate
git clone https://github.com/modelscope/ms-swift.git
cd ms-swift
python -m pip install -e . -r requirements/musa.txt qwen_vl_utils decord
# The image ships torchao 0.9.0 (required by sglang); peft>=0.19 requires torchao>=0.16, otherwise LoRA fails
python -m pip install peft==0.18.1
1.3 Environment Check
Confirm that torch_musa detects the GPUs:
python -c "import torch, torch_musa; print(torch.musa.is_available(), torch.musa.device_count())"
# output: True 8
Confirm that swift imports correctly (torchada is enabled automatically):
python -c "import swift, torch; print(swift.__version__, torch.cuda.device_count())"
# output: 4.6.0.dev0 8
Check GPU utilization and memory usage:
mthreads-gmi
---------------------------------------------------------------------
mthreads-gmi:2.3.3 Driver Version:3.3.8-server
---------------------------------------------------------------------
ID Name |PCIe |%GPU Mem
Device Type Perf |Pcie Lane Width |Temp MPC Capable
| ECC Mode
+-------------------------------------------------------------------+
0 MTT S5000 |00000000:18:00.0 |0% 0MiB(81920MiB)
Physical P0 |16x(16x) |47C YES
| On-die, EDC
+-------------------------------------------------------------------+
...
2. Examples
Notes:
Use
MUSA_VISIBLE_DEVICESto select the visible GPUs; for multi-GPU training setNPROC_PER_NODE, and swift launches the processes via torchrun automatically.--model Qwen/Qwen3.5-4Bbelow downloads the model from ModelScope automatically; it can also be replaced by a local model directory.The datasets follow the ms-swift examples (alpaca Chinese/English data plus self-cognition data);
--model_author/--model_nameare used for self-cognition training.
2.1 Single-GPU LoRA Fine-tuning of Qwen3.5-4B
PYTORCH_MUSA_ALLOC_CONF='expandable_segments:True' \
MUSA_VISIBLE_DEVICES=0 \
IMAGE_MAX_TOKEN_NUM=1024 \
VIDEO_MAX_TOKEN_NUM=128 \
FPS_MAX_FRAMES=12 \
swift sft \
--model Qwen/Qwen3.5-4B \
--tuner_type lora \
--dataset 'AI-ModelScope/alpaca-gpt4-data-zh#500' \
'AI-ModelScope/alpaca-gpt4-data-en#500' \
'swift/self-cognition#500' \
--model_author swift \
--model_name swift-robot \
--load_from_cache_file true \
--add_non_thinking_prefix true \
--loss_scale ignore_empty_think \
--split_dataset_ratio 0.01 \
--torch_dtype bfloat16 \
--num_train_epochs 1 \
--per_device_train_batch_size 4 \
--per_device_eval_batch_size 4 \
--learning_rate 1e-4 \
--lora_rank 8 \
--lora_alpha 32 \
--target_modules all-linear \
--gradient_accumulation_steps 4 \
--group_by_length true \
--output_dir output/Qwen3.5-4B-lora \
--eval_steps 50 \
--save_steps 50 \
--save_total_limit 2 \
--logging_steps 5 \
--max_length 2048 \
--warmup_ratio 0.05 \
--dataset_num_proc 4 \
--dataloader_num_workers 4
The single GPU uses about 15GiB of memory, and 1 epoch (94 steps) takes about 9 minutes.
2.2 Multi-GPU Full-parameter Fine-tuning of Qwen3.5-4B (DeepSpeed ZeRO-3)
The image ships a MUSA build of DeepSpeed, so --deepspeed zero3 can be used directly:
PYTORCH_MUSA_ALLOC_CONF='expandable_segments:True' \
NPROC_PER_NODE=4 \
MUSA_VISIBLE_DEVICES=0,1,2,3 \
IMAGE_MAX_TOKEN_NUM=1024 \
VIDEO_MAX_TOKEN_NUM=128 \
FPS_MAX_FRAMES=12 \
swift sft \
--model Qwen/Qwen3.5-4B \
--tuner_type full \
--dataset 'AI-ModelScope/alpaca-gpt4-data-zh#500' \
'AI-ModelScope/alpaca-gpt4-data-en#500' \
'swift/self-cognition#500' \
--model_author swift \
--model_name swift-robot \
--load_from_cache_file true \
--add_non_thinking_prefix true \
--loss_scale ignore_empty_think \
--split_dataset_ratio 0.01 \
--torch_dtype bfloat16 \
--freeze_vit true \
--freeze_aligner true \
--num_train_epochs 1 \
--per_device_train_batch_size 2 \
--per_device_eval_batch_size 2 \
--learning_rate 1e-5 \
--gradient_accumulation_steps 2 \
--gradient_checkpointing true \
--group_by_length true \
--output_dir output/Qwen3.5-4B-full \
--eval_steps 50 \
--save_steps 50 \
--save_total_limit 2 \
--save_only_model true \
--logging_steps 5 \
--max_length 2048 \
--warmup_ratio 0.05 \
--dataset_num_proc 4 \
--dataloader_num_workers 4 \
--deepspeed zero3
With 4-GPU ZeRO-3, each GPU uses about 26.5GiB of memory, and 1 epoch (93 steps) takes about 7 minutes. To use 8 GPUs, set NPROC_PER_NODE=8 and MUSA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7.
2.3 Inference and LoRA Merging
Inference with the LoRA weights (transformers backend):
MUSA_VISIBLE_DEVICES=0 \
swift infer \
--adapters output/Qwen3.5-4B-lora/vx-xxx/checkpoint-xxx \
--infer_backend transformers \
--enable_thinking false \
--stream true \
--max_new_tokens 2048
Inference with the full-parameter checkpoint:
MUSA_VISIBLE_DEVICES=0 \
swift infer \
--model output/Qwen3.5-4B-full/vx-xxx/checkpoint-xxx \
--infer_backend transformers \
--enable_thinking false \
--stream true \
--max_new_tokens 2048
Inference results after self-cognition training:
[QUERY] 你是谁?
[RESPONSE] 我是一个由swift开发的人工智能助手,名为swift-robot。我被设计用来回答您的问题,提供信息,并进行有趣的对话。如果您有任何问题或需要帮助,请随时向我提问。
[QUERY] who are you?
[RESPONSE] I am an artificial intelligence assistant named swift-robot, trained by swift. My purpose is to provide assistance, answer questions, and engage in conversation with users. ...
Merge the LoRA weights (written to checkpoint-xxx-merged):
MUSA_VISIBLE_DEVICES=0 \
swift export \
--adapters output/Qwen3.5-4B-lora/vx-xxx/checkpoint-xxx \
--merge_lora true