What Xiaomi’s Cheapest AI Couldn’t Fix
MiMo-V2.6-Pro scored 46 on the Intelligence Index, beat Claude Opus 5 on terminal benchmarks, and cost 11 to 28 times less. But when a client’s terminal locked up at 2:14 a.m., the only benchmark that mattered was whether the API would answer.

The Queue at 2 A.M.
A story based on Xiaomi’s MiMo-V2.6 release.
I remember the exact time because the alert hit my phone at 2:14 a.m. A logistics client in Shah Alam couldn’t scan inbound pallets. Their terminal logs were a mess. Our old model, MiMo V2.5, was supposed to inspect those logs. It had scored 65.2 on Terminal Bench 2.1, which sounded fine in a slide. In practice, it missed the one line that mattered: a certificate expiry buried under three hundred retry errors.
So when Xiaomi open-sourced MiMo-V2.6 on September 22, I didn’t care about leaderboards. I cared about that line.
The release notes said Pro scored 46 on Artificial Analysis’s Intelligence Index v4.3.2. First among open-weight models. GLM-5.3 got 45. Kimi K3 got 44. Claude Fable 5.1 and GPT-6 Astra both got 53. Xiaomi acknowledged the gap. I acknowledged it too. Five points is five points.
But the jump from V2.5-Pro was twenty points. That was the number that made me download the weights.
My colleague Farah leaned over my desk. “Training cost?”
“$2.62 million for Pro,” I said. “Less than six days. Thirty steps. About 750,000 trajectories. Flash was $850,000.”
She whistled. “And the pass rates?”
“Up 25% and 12% over the old versions. On DeepSWE, Pro went from 48.8 to 65.7. Flash from 58.4 to 72.6.”
Farah ran the numbers. “Flash gained more than 19 points for $850,000. Pro spent three times that for 14 points. Flash is the better buy.”
“That’s what Goldman said.”
We deployed Flash for routine log triage that night. Pro stayed for the hard stuff. The API pricing was hard to ignore. Pro was $0.435 per million input tokens and $0.87 per million output tokens. Cache discount was 99%. Cached input cost $0.0036 per million. Flash was $0.14 in, $0.28 out. Claude Opus 5 listed at $5.00 input and $25.00 output. We were a small team. Cheap mattered.
The first test was a remote log inspection. I pointed Pro at a failing terminal session. It read the logs, found the certificate expiry, and suggested a renewal script. It took 40 seconds. V2.5 had taken four minutes and missed it.
I ran the benchmark anyway. Terminal Bench 2.1: 89.9. Above Claude Opus 5’s 89.1 and GPT-5.6 Sol’s 88.8. Terminal Bench 4.0 was harder. Pro scored 34.9. GLM-5.3 got 42.4. Grok 4.7 got 49.0. On AutomationBench v1.0.6, Pro scored 53.1. Claude Opus 5 got 50.3. GPT-5.6 Sol got 45.8.
So the model was near the top on classic terminal work. On newer agent tasks, it still had ground to cover. That matched what I saw. It could fix a log. It could not run a whole incident response.
At 2:14 a.m., I needed it to run a whole incident response.
The client’s terminal was locked. I opened our API console. I pasted the logs. Pro started generating. Then it stopped. The queue. Xiaomi had received more than 66,000 applications during the UltraSpeed beta. Ten queue slots per person per day. Thirty-minute session caps. Now the API was open. The traffic surge was real.
I waited. The cursor blinked. I thought about output efficiency. AA data said Pro generated 140 million output tokens while running the Intelligence Index. Median. “Somewhat verbose.” Output speed around 130 tokens per second. Xiaomi said its token efficiency was second only to Step 5, lower than the average of other frontier domestic models. If a model needs two or three times the tokens, a cheap unit price gets eaten by volume. I was watching it happen in real time.
The response came at 2:31 a.m. Seventeen minutes late. It found the issue. A bad pattern dataset from a mid-run intervention. Goldman’s report mentioned those. Infrastructure error reruns. Zero-gradient tasks. GPU memory shortages. The RL bottleneck had shifted to engineering optimization. I felt that shift in my chest.
I fixed the client’s terminal. Then I wrote a fallback rule. Flash for anything under 10,000 tokens. Pro for deep analysis. Local weights for anything that could not wait. The 1.02T-parameter MoE model with 42B active parameters was not cheap to serve. Xiaomi could open-source the weights. They could not open-source the GPUs.
The next morning, Farah asked if we should switch everything to Pro.
“No,” I said. “It’s the best open-weight model for some jobs. It’s not the best for a 2 a.m. queue.”
We kept the MIT weights. We kept the 7,000 RL task environments. We joined the agent frameworks: OpenClaw, OpenCode, KiloCode, Cline, BLACKBOXAI. We used the first-week free access. We read the technical report. We watched the Lean 4 formalization of “Period Three Implies Chaos” pass kernel verification. We saw the embodied simulation with the Franka Panda arm. We knew Xiaomi planned to invest more than 60 billion RMB over three years. We knew Goldman expected MiMo-V3 with HySparse in early 2027, sparsity ratio from 7:1 to 11:1.
But that was a prediction. My client’s pallets were a result.
I still use MiMo-V2.6-Pro. I just don’t use it when the queue is long. The price is good. The score is good. The hard part is serving it. And at 2 a.m., the hard part is the only part that matters.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.