Moore Threads MUSA Support

ms-swift supports Moore Threads GPUs through torch_musa and torchada. torchada redirects the torch.cuda.* APIs to torch.musa, and swift imports it automatically when both torch_musa and torchada are installed, so most training/inference scripts run on MUSA without modification.

This document uses Qwen3.5-4B as an example to demonstrate LoRA and full-parameter fine-tuning, inference, and LoRA merging with ms-swift on MTT S5000.

1. Environment Setup

1.1 Base Environment

Pull the Moore Threads training image (which ships torch_musa, DeepSpeed, flash-attention, MT-TransformerEngine, etc.) and start a container with the following command:

IMAGE_NAME=sh-harbor.mthreads.com/mcctest/training-suite:v2.1.7-rc4
docker pull ${IMAGE_NAME}

CONTAINER_NAME=swift_test
docker run -itd --privileged --network=host \
    --env MTHREADS_VISIBLE_DEVICES=all \
    --shm-size=80g \
    -v /data:/data \
    --name ${CONTAINER_NAME} \
    ${IMAGE_NAME} \
    /bin/bash

docker exec -it ${CONTAINER_NAME} /bin/bash

Software versions used for verification in this document:

Component Version
GPU / Driver MTT S5000 (80GB) × 8 / Driver 3.3.8-server
MUSA Toolkits 4.3.8
python 3.10.12
torch / torch_musa 2.7.1 / 2.7.1.post1
torchada 0.1.90
transformers 5.16.1
peft 0.18.1
deepspeed 0.19.3 (MUSA build shipped in the image)

1.2 Install ms-swift

Qwen3.5 requires transformers>=5.2.0, while the image ships an older transformers (4.57). To avoid breaking the environment that other components in the image (vLLM, sglang, etc.) depend on, we recommend creating a virtual environment that inherits the system packages and installing ms-swift there:

# Reuse torch/torch_musa/deepspeed from the image; only override the packages that need upgrading in the venv
python -m venv --without-pip --system-site-packages /home/venv_swift
source /home/venv_swift/bin/activate

git clone https://github.com/modelscope/ms-swift.git
cd ms-swift
python -m pip install -e . -r requirements/musa.txt qwen_vl_utils decord
# The image ships torchao 0.9.0 (required by sglang); peft>=0.19 requires torchao>=0.16, otherwise LoRA fails
python -m pip install peft==0.18.1

1.3 Environment Check

  • Confirm that torch_musa detects the GPUs:

python -c "import torch, torch_musa; print(torch.musa.is_available(), torch.musa.device_count())"
# output: True 8
  • Confirm that swift imports correctly (torchada is enabled automatically):

python -c "import swift, torch; print(swift.__version__, torch.cuda.device_count())"
# output: 4.6.0.dev0 8
  • Check GPU utilization and memory usage: mthreads-gmi

---------------------------------------------------------------------
    mthreads-gmi:2.3.3           Driver Version:3.3.8-server
---------------------------------------------------------------------
ID   Name                 |PCIe                |%GPU  Mem
     Device Type   Perf   |Pcie Lane Width     |Temp  MPC Capable
                                               |      ECC Mode
+-------------------------------------------------------------------+
0    MTT S5000            |00000000:18:00.0    |0%    0MiB(81920MiB)
     Physical      P0     |16x(16x)            |47C   YES
                                               |      On-die, EDC
+-------------------------------------------------------------------+
...

2. Examples

Notes:

  • Use MUSA_VISIBLE_DEVICES to select the visible GPUs; for multi-GPU training set NPROC_PER_NODE, and swift launches the processes via torchrun automatically.

  • --model Qwen/Qwen3.5-4B below downloads the model from ModelScope automatically; it can also be replaced by a local model directory.

  • The datasets follow the ms-swift examples (alpaca Chinese/English data plus self-cognition data); --model_author/--model_name are used for self-cognition training.

2.1 Single-GPU LoRA Fine-tuning of Qwen3.5-4B

PYTORCH_MUSA_ALLOC_CONF='expandable_segments:True' \
MUSA_VISIBLE_DEVICES=0 \
IMAGE_MAX_TOKEN_NUM=1024 \
VIDEO_MAX_TOKEN_NUM=128 \
FPS_MAX_FRAMES=12 \
swift sft \
    --model Qwen/Qwen3.5-4B \
    --tuner_type lora \
    --dataset 'AI-ModelScope/alpaca-gpt4-data-zh#500' \
              'AI-ModelScope/alpaca-gpt4-data-en#500' \
              'swift/self-cognition#500' \
    --model_author swift \
    --model_name swift-robot \
    --load_from_cache_file true \
    --add_non_thinking_prefix true \
    --loss_scale ignore_empty_think \
    --split_dataset_ratio 0.01 \
    --torch_dtype bfloat16 \
    --num_train_epochs 1 \
    --per_device_train_batch_size 4 \
    --per_device_eval_batch_size 4 \
    --learning_rate 1e-4 \
    --lora_rank 8 \
    --lora_alpha 32 \
    --target_modules all-linear \
    --gradient_accumulation_steps 4 \
    --group_by_length true \
    --output_dir output/Qwen3.5-4B-lora \
    --eval_steps 50 \
    --save_steps 50 \
    --save_total_limit 2 \
    --logging_steps 5 \
    --max_length 2048 \
    --warmup_ratio 0.05 \
    --dataset_num_proc 4 \
    --dataloader_num_workers 4

The single GPU uses about 15GiB of memory, and 1 epoch (94 steps) takes about 9 minutes.

2.2 Multi-GPU Full-parameter Fine-tuning of Qwen3.5-4B (DeepSpeed ZeRO-3)

The image ships a MUSA build of DeepSpeed, so --deepspeed zero3 can be used directly:

PYTORCH_MUSA_ALLOC_CONF='expandable_segments:True' \
NPROC_PER_NODE=4 \
MUSA_VISIBLE_DEVICES=0,1,2,3 \
IMAGE_MAX_TOKEN_NUM=1024 \
VIDEO_MAX_TOKEN_NUM=128 \
FPS_MAX_FRAMES=12 \
swift sft \
    --model Qwen/Qwen3.5-4B \
    --tuner_type full \
    --dataset 'AI-ModelScope/alpaca-gpt4-data-zh#500' \
              'AI-ModelScope/alpaca-gpt4-data-en#500' \
              'swift/self-cognition#500' \
    --model_author swift \
    --model_name swift-robot \
    --load_from_cache_file true \
    --add_non_thinking_prefix true \
    --loss_scale ignore_empty_think \
    --split_dataset_ratio 0.01 \
    --torch_dtype bfloat16 \
    --freeze_vit true \
    --freeze_aligner true \
    --num_train_epochs 1 \
    --per_device_train_batch_size 2 \
    --per_device_eval_batch_size 2 \
    --learning_rate 1e-5 \
    --gradient_accumulation_steps 2 \
    --gradient_checkpointing true \
    --group_by_length true \
    --output_dir output/Qwen3.5-4B-full \
    --eval_steps 50 \
    --save_steps 50 \
    --save_total_limit 2 \
    --save_only_model true \
    --logging_steps 5 \
    --max_length 2048 \
    --warmup_ratio 0.05 \
    --dataset_num_proc 4 \
    --dataloader_num_workers 4 \
    --deepspeed zero3

With 4-GPU ZeRO-3, each GPU uses about 26.5GiB of memory, and 1 epoch (93 steps) takes about 7 minutes. To use 8 GPUs, set NPROC_PER_NODE=8 and MUSA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7.

2.3 Inference and LoRA Merging

Inference with the LoRA weights (transformers backend):

MUSA_VISIBLE_DEVICES=0 \
swift infer \
    --adapters output/Qwen3.5-4B-lora/vx-xxx/checkpoint-xxx \
    --infer_backend transformers \
    --enable_thinking false \
    --stream true \
    --max_new_tokens 2048

Inference with the full-parameter checkpoint:

MUSA_VISIBLE_DEVICES=0 \
swift infer \
    --model output/Qwen3.5-4B-full/vx-xxx/checkpoint-xxx \
    --infer_backend transformers \
    --enable_thinking false \
    --stream true \
    --max_new_tokens 2048

Inference results after self-cognition training:

[QUERY] 你是谁?
[RESPONSE] 我是一个由swift开发的人工智能助手,名为swift-robot。我被设计用来回答您的问题,提供信息,并进行有趣的对话。如果您有任何问题或需要帮助,请随时向我提问。

[QUERY] who are you?
[RESPONSE] I am an artificial intelligence assistant named swift-robot, trained by swift. My purpose is to provide assistance, answer questions, and engage in conversation with users. ...

Merge the LoRA weights (written to checkpoint-xxx-merged):

MUSA_VISIBLE_DEVICES=0 \
swift export \
    --adapters output/Qwen3.5-4B-lora/vx-xxx/checkpoint-xxx \
    --merge_lora true