Back to East→West AI Tools
east westeast-westfine-tuningmodelscope

ms-swift: The Chinese Fine-Tuning Standard Western Devs Don't Know (600+ Models, Apache-2.0)

Unsloth and axolotl dominate in the West. In China, the standard is ms-swift by ModelScope — 600+ models, CPT/SFT/DPO/GRPO, Apache-2.0.

Matthieu LebasSeptember 3, 20262 min read226 words

TL;DR — You fine-tune with unsloth or axolotl. In China, the default is ms-swift (ModelScope, 15,500★, Apache-2.0, active daily): CPT/SFT/DPO/GRPO over 600+ models — including Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, plus Llama 4 and everything Western. It's the shortest path when your base model is Chinese.

The "because Z"

  1. One tool, 600+ models. ms-swift's model zoo covers the Chinese open models (Qwen, DeepSeek, GLM, MiniCPM, InternLM) and the Western ones. No adapter hunting per model family.
  2. It's what Chinese labs ship with. When a Qwen/GLM release lands, ms-swift has a fine-tuning recipe the same week. That speed matters for post-training on freshly released models.
  3. Apache-2.0 + daily activity. Verified via the GitHub API: license Apache-2.0, pushed 2026-09-03. It's not an abandoned research repo — it's production infrastructure.

The honest caveat

The documentation is China-first (though increasingly bilingual). If you only fine-tune Llama/Mistral models, unsloth's DX is smoother. ms-swift wins the moment your base model is Qwen, DeepSeek or GLM — which, given 45%+ of OpenRouter traffic now runs Chinese models, is more often than you'd think.

Verdict

  • Base model chinois + besoin de post-training → ms-swift (CPT/SFT/DPO/GRPO, 600+ modèles).
  • Workflow 100% Llama/Mistral occidentaux → unsloth reste excellent.
  • Both are Apache/MIT — no license trap either way.

Data verified 2026-09-03 · Source: GitHub API (modelscope/ms-swift, license + pushed_at).

east-westfine-tuningmodelscopechina
Share this article:

Was this article helpful?

Let us know to improve our AI generation.