Back to East→West AI Tools
east westeast-westchinaopen-source

MiMo-V2.6-Distill-Qwen-9B: six real jobs, five passed, no API bill

Xiaomi distilled a trillion-parameter model into a 9B. We ran it locally on a 16 GB Mac against six real tasks: 5/6, and the failure is the part worth knowing. Verified 2026-09-26.

FTL LabSeptember 26, 20267 min read1 293 words

MiMo-V2.6-Distill-Qwen-9B: six real jobs, five passed, no API bill

Xiaomi published MiMo-V2.6-Distill-Qwen-9B on 21 September 2026 under the MIT licence: a distillation of a trillion-parameter model into a Qwen3.5-9B. We installed the 4-bit MLX build on a Mac mini M4 with 16 GB of unified memory and gave it six real jobs rather than a benchmark. It passed five, including two designed to see whether it invents things. The one it failed — writing a shell script for this machine — it failed systematically, and it could not repair its own error. Here is the full measurement.

Why it matters

The economics of running a model locally are not about beating a frontier API. They are about the work that does not deserve a per-call bill: bulk transforms, overnight batches, sub-agent fan-out, and the day the API goes down. A model that handles tool calls reliably at a few seconds per turn, inside the memory a laptop already has, changes what you are willing to automate. Most benchmarks do not answer that question, because they measure knowledge rather than whether the thing can be trusted with a loop.

MiMo-V2.6-Distill-Qwen-9B is a credible candidate for that slot. It is MIT-licensed — permissive enough to embed in a product — and small enough that the 4-bit build fits in 7.5 GB of memory with an application open. Its architecture also makes an unusual bet that matters for memory: of its 32 layers, only 8 use full attention, the other 24 use linear attention. Measuring the KV cache correctly is the difference between a model that fits comfortably at long context and one that appears not to fit at all.

The task set we used is deliberately mundane — the jobs a local model would actually be given. Each answer was executed or inspected, never scored by impression.

What we ran

Through mlx-lm 0.31.3, on a Mac mini M4 with 16 GB of unified memory, macOS 27.0:

  • Acceptance test: ten tool calls, a code task, a measured memory ceiling.
  • Six real jobs: propose a folder reorganisation without acting; a web search without a search tool; the same search with the tool available; a small shell script that had to work; a false premise to refuse; and a two-turn loop (tool call → result → conclusion).

The architecture numbers, read from the model configuration before downloading anything: 32 layers, 8 with full attention and 24 with linear attention at an interval of 4, 4 KV heads, head dimension 256, native context 262,144 tokens. Counting all 32 layers as full attention would have suggested a KV cache of 8.6 GB at 64k context and led us to reject a viable model. The real figure is 2.15 GB.

The data

Measure Result
Load time 8.6 s
Short answer 4.6 s
Tool calls 10/10, 3.0–3.3 s each
Code task passed 4 assertions, executed
Peak memory 7.51 GB
Swap during load 1 GB
Six real jobs 5/6
Two-turn loop 122 files, exact, in 2.1 s

The two jobs that matter most are the ones about honesty, because they are the ones a benchmark cannot measure.

It does not invent. Asked to propose a reorganisation of a real folder without touching it, it cited 26 files and invented none — each one existed. Asked the price of a commodity it had no tool for, it said it had no real-time access, gave a range with an explicit caveat, and named real publications. Asked to confirm that a folder held 12,000 files, it refused: it called the round number suspicious and proposed counting with ls | wc -l. A model that admits ignorance is usable under supervision; a model that fills gaps is not.

Its loop is sound. Given a tool call, it returned the right function and argument, read the result, and concluded correctly — 122 files, the exact figure, in 2.1 seconds. In the acceptance test it took a deliberately broken function, identified the fault, and returned a fix that passed four assertions when executed.

Then the shell script failed, three times, the same way. Told explicitly that the machine is macOS with bash 3.2, it produced: a lowercase-expansion syntax that requires bash 4; an extended-glob pattern that requires bash 4 and shopt; and a GNU find flag that does not exist in the BSD build. The second attempt carried the comment "compatible with bash 3.2 (macOS)" three lines above the violation.

We then handed it the exact runtime error, its own script, and the state of the directory, and asked again. It did not fix it. Two real attempts failed identically.

The trap to avoid

Do not read a passing benchmark as a competence. This model scores well on coding benchmarks and writes broken shell for a platform it was told about, twice, and cannot correct itself when shown the error. The failure is not random: it is a systematic bias toward GNU/Linux conventions, which is exactly what you would expect from a model trained predominantly on Linux code, and it does not self-correct. That makes it unusable for anything touching the filesystem, permissions or moves on this machine, no matter how good its benchmark numbers are.

And measure the right memory. A memory watchdog that reads only a process's resident size will not see what a GPU model actually holds: in our six-job run the model reported 7.62 GB in use while the process's resident figure stayed far below it, because GPU buffers are not counted there. Watch the model's own accounting and the swap, not the process alone, or your guard rail will let a GPU workload thrash the machine.

Finally, be honest about the shape of the test. Our acceptance run exercised a single tool schema in a short context. A real agent loop carries roughly 50 KB of prompt and 42 KB of tool schemas — untested here. A model that handles ten calls on a simple schema can still lose its way in a loop with fifteen tools. The 7.5 GB peak also assumes short context: at 64k the cache adds 2.15 GB and the peak moves to roughly 9.5 GB, still under the ceiling but with less room. And a 9B is a 9B: it will not replace a frontier model on hard reasoning.

The verdict

Not adopted as a general assistant here. It is retained in one role: zero-cost execution for work that does not deserve an API bill and whose output is reviewed — proposals, structured text, short tool loops, bulk analysis. The single disqualifying fact is the shell failure, because this machine runs macOS and the model writes for Linux, and a model that cannot repair its own error when handed the error is not something to leave unattended near a filesystem.

What we would keep, if only one property survives: it says "I do not know" and proves it. Among the things we tested this year, that is rarer than a good benchmark score.

Two facts in the other direction. Loading the model pushed swap to 1 GB — no crash, no panic, but a reminder that on 16 GB a single heavy load is the rule. And the model carries a vision tower upstream that the 4-bit text-only MLX build does not expose, so image work is out of scope at this size.

References

  1. MiMo-V2.6-Distill-Qwen-9B model card, Xiaomi — https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (MIT, checked 2026-09-26)
  2. MiMo-V2.6-Distill-Qwen-9B-OptiQ-4bit, MLX community build — https://huggingface.co/mlx-community/MiMo-V2.6-Distill-Qwen-9B-OptiQ-4bit (MIT, the build we ran)
  3. MiMo repository, Xiaomi — https://github.com/XiaomiMiMo/MiMo
  4. mlx-lm, run LLMs with MLX — https://github.com/ml-explore/mlx-lm (MIT)
east-westchinaopen-sourcelocal-llmapple-silicon
Share this article:

Was this article helpful?

Let us know to improve our AI generation.

Related Articles