ms-swift: The Chinese Fine-Tuning Standard Western Devs Don't Know (600+ Models, Apache-2.0)
Unsloth and axolotl dominate in the West. In China, the standard is ms-swift by ModelScope — 600+ models, CPT/SFT/DPO/GRPO, Apache-2.0.
TL;DR — You fine-tune with unsloth or axolotl. In China, the default is ms-swift (ModelScope, 15,500★, Apache-2.0, active daily): CPT/SFT/DPO/GRPO over 600+ models — including Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, plus Llama 4 and everything Western. It's the shortest path when your base model is Chinese.
The "because Z"
- One tool, 600+ models. ms-swift's model zoo covers the Chinese open models (Qwen, DeepSeek, GLM, MiniCPM, InternLM) and the Western ones. No adapter hunting per model family.
- It's what Chinese labs ship with. When a Qwen/GLM release lands, ms-swift has a fine-tuning recipe the same week. That speed matters for post-training on freshly released models.
- Apache-2.0 + daily activity. Verified via the GitHub API: license Apache-2.0, pushed 2026-09-03. It's not an abandoned research repo — it's production infrastructure.
The honest caveat
The documentation is China-first (though increasingly bilingual). If you only fine-tune Llama/Mistral models, unsloth's DX is smoother. ms-swift wins the moment your base model is Qwen, DeepSeek or GLM — which, given 45%+ of OpenRouter traffic now runs Chinese models, is more often than you'd think.
Verdict
- Base model chinois + besoin de post-training → ms-swift (CPT/SFT/DPO/GRPO, 600+ modèles).
- Workflow 100% Llama/Mistral occidentaux → unsloth reste excellent.
- Both are Apache/MIT — no license trap either way.
Data verified 2026-09-03 · Source: GitHub API (modelscope/ms-swift, license + pushed_at).
Was this article helpful?
Let us know to improve our AI generation.