01 logo

How a “Weaker” Chinese AI Model Stole DeepSeek’s Crown Overnight

Xiaomi’s MiMo-V2.5 just claimed the global #1 spot. It’s cheaper, multi-modal, and everywhere. But the real story is what happens when the price war ends.

By JinPublished 2 months ago 4 min read
https://mimo.xiaomi.com/mimo-v2-pro

On OpenRouter’s leaderboard, a quiet changing of the guard took place in the final week of July. Xiaomi’s MiMo‑V2.5 recorded 10.46 trillion tokens in weekly calls and 31.2 trillion in monthly calls, officially surpassing DeepSeek V4 Flash to claim the top spot for global model invocation volume.

Since June, its weekly call volume has soared from 1.46 trillion to 10.46 trillion tokens, a ~616% increase in just two months.

This was not a gradual climb. It was a near‑vertical curve.

Trailing behind it are DeepSeek V4 Flash, Hy3, Qwen‑Max 0728, and DeepSeek V4. The top five are all Chinese models. Chinese AI models now hold a combined 63.5% share on OpenRouter, surpassing US‑based models for the first time.

MiMo‑V2.5’s ascent is not an isolated case. It is a cross‑section of Chinese AI’s growing “combined‑arms advantage” in engineering execution and cost control.

But beneath that cross‑section lie far more complex layers.

I. The Core Contradiction: Weaker Coding Ability, Yet Higher Volume

Ask around in developer communities, and the answer is unanimous: in coding capability, MiMo‑V2.5 falls short of DeepSeek V4 Flash.

This is not subjective. Third‑party data from OpenCode Go shows that DeepSeek models together account for over 70% of all coding‑related calls, with V4 Flash alone taking more than half of the entire market. For complex chain‑of‑thought reasoning and advanced algorithmic implementation, developers still default to DeepSeek.

The explanation is native multi‑modality, a dimension most have overlooked.

Unlike DeepSeek V4 Flash, a pure text/code foundation model, MiMo‑V2.5 supports native video input. It requires no pre‑processing (e.g., frame slicing) and can directly perform long‑form video scene understanding, temporal reasoning, and event detection.

In use cases like AI‑powered video creation, automated content moderation, meeting summarization, and product demo analysis, MiMo‑V2.5 offers a solution that few alternatives can match: a sufficiently high intelligence floor, combined with the convenience of native video handling.

Your speculation, “large‑scale, batch video understanding,” is very likely the primary engine behind the volume spike. For AI video creators and automated processing pipelines, handling millions of hours of video streams translates to astronomical token consumption. And MiMo‑V2.5 sits precisely at the intersection of “capable enough and cheap enough.”

II. The Price Knife: Erasing Every Switching Cost

In late May, Xiaomi announced a permanent price cut for the MiMo‑V2.5 series API.

Input dropped to $0.40 per million tokens, output to $2.00 per million tokens, a reduction of up to 99%. This pricing directly mirrors DeepSeek V4 Flash, making the two models virtually identical in cost. For developers, the friction of switching models was eliminated.

But the truly anomalous data lies in the structure of call sources.

OpenRouter’s breakdown shows that MiMo‑V2.5’s Top 5 applications account for only 1.7% of its total calls. In contrast, DeepSeek V4 Flash or Hy3 typically see their Top 5 apps contributing around 20%.

This means MiMo‑V2.5’s volume is not driven by a handful of mega‑apps, but rather by tens of thousands of long‑tail developers and small automation scripts. This decentralised distribution proves that MiMo’s API has achieved a high level of versatility and ease‑of‑use: it can be embedded into countless niche digital workflows, from video subtitle generation and surveillance footage analysis to educational video summarisation and product presentation conversion.

But it also implies something else: no one can say with precision where all these calls are actually going. Even Xiaomi itself may not have a clear picture. This adds immense ambiguity to future market‑strategy formulation.

III. The “Understudy” Position: Volume Without Stickiness

The developer community’s positioning of MiMo‑V2.5 is consistent: a powerful “specialist understudy,” not the “absolute first‑choice.”

Its typical usage pattern looks like this. When a workflow suddenly requires processing a design mock‑up or a tutorial video, developers call MiMo for multi‑modal parsing, then pass the extracted text to DeepSeek for logical reasoning. “MiMo sees, DeepSeek thinks”—this model‑routing division is becoming a standard pattern in AI‑native applications.

This positioning brings a structural fragility: neither user stickiness nor irreplaceability is strong enough.

Outside of multi‑modality, MiMo has not built an absolute moat. If DeepSeek launches a version with equally native video support at a lower price, or if the next‑generation open‑source model catches up on multi‑modal capabilities, MiMo’s volume dominance could collapse within weeks.

The current high volume looks less like a vote for an irreplaceable model and more like a price‑sensitive stampede toward the best‑value multi‑modal option available today.

These are two fundamentally different business logics. The former means users can leave at any time; the latter means they cannot.

IV. The Profitability Question: What Comes After 31 Trillion Tokens?

Processing massive video streams at $0.40 per million tokens imposes obvious compute and bandwidth costs.

Top call volume does not equal top revenue. Ultra‑low unit prices demand extreme scale just to break even. If the current volume is predominantly driven by price‑sensitive, low‑loyalty batch scripts, then when subsidies taper off or prices are adjusted, the 31‑trillion‑monthly‑active bubble could burst quickly.

Xiaomi needs to answer a question not covered by any public data: how much of this volume represents paid, sustainable commercial activity, and how much is just testing or subsidy‑driven traffic?

No public information currently sheds light on this.

V. From Blitzkrieg to Trench Warfare: The Real Battle Is on “IQ” Ground

MiMo‑V2.5’s rise is a textbook case of precise market differentiation.

It did not confront DeepSeek head‑on in the coding arena. That battlefield has already been sealed by DeepSeek’s 70%+ stranglehold. Instead, it chose the under‑explored blue ocean of multi‑modal video understanding, using native capability and extreme pricing to force its way in.

It was a brilliant blitzkrieg.

But a blitzkrieg must be followed by positional warfare. The ultimate contest in foundation models remains the race for intelligence boundaries and reasoning depth. Multi‑modality gives a model “eyes” and “ears”; logical reasoning is the “brain.” MiMo has proven it has sharp eyes and keen ears. But to go from “first in volume” to “first in mind,” it still needs to prove itself on the evolutionary path of the brain.

The top spot on OpenRouter will change hands frequently. The truly meaningful question is: when the next model arrives with even lower prices and stronger multi‑modal abilities, will MiMo’s users stay, or migrate again?

The answer to that question is not in the call‑volume data. It lies in the IQ of the model’s next iteration.

social mediatech newsfact or fictionthought leadersproduct review

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin