01 logo

The 0.02 Yuan Price Cut Is Not the Interesting Part

I spent the morning after DeepSeek’s Flash price change doing the math developers actually care about. The input number is low. The output number is the wall.

By JinPublished 28 days ago • 3 min read

At 12:00 Beijing time on September 10, DeepSeek changed its Flash pricing. I opened the page and copied the numbers into a spreadsheet.

Idle hours: cache-hit input 0.02 yuan per million tokens, cache-miss input 1 yuan, output 4 yuan.

Peak hours double those. Cache-hit input 0.04 yuan, cache-miss input 2 yuan, output 8 yuan.

That is the headline. Headlines are not bills.

I run a small document search tool. Most of my input is repeated context: system prompts, retrieved passages, chat history. My cache hit rate is usually high. When I plugged in the new numbers, my input cost almost disappeared.

Before the change, 1 million tokens with a 99.23 percent hit rate cost about 0.0612 yuan. After the change, the same mix costs about 0.0275 yuan. That is a drop of more than 55 percent. The input column on my test invoice went from a small number to a smaller number. I almost missed it.

Then I looked at output. Output is still 4 yuan per million tokens. During peak hours, 8 yuan.

That is where the money goes. Input can be cached. Output cannot. Every new token is new. Every extra sentence has a price.

If you write long answers, translate documents, or generate reports, the input cut helps. It does not save you. This is the part the pricing page does not say out loud.

I learned this the hard way last month. I built a small tool that summarizes long reports for a client. The input side was fine. The output side was not. I had set the model to generate a detailed summary with bullet points and a conclusion. The client loved it. My bill did not. I ended up rewriting the prompt to force shorter, more structured answers. The quality dropped a little. The cost dropped a lot. That was the trade I had to make.

The new prices do not change that trade. They just make the input side cheaper. The output side is still the decision point.

The other part is the dates.

On September 8, DeepSeek’s V4.1 Flash internal testing announcement used eight words: stronger capability, faster speed, lower cost. The model name carried a suffix: deepseek-v4.1-flash-expires-on-0910.

On September 10 at noon, the new Flash prices went live.

Three dates. Two days apart. The old model’s price moved right as the new model’s test window closed. Developers see the price first. The strategy shows up later.

I am not saying this is a conspiracy. I am saying the timing is tight. Companies do this all the time. They lower the price of the old model to clear the way for the new one. It is a normal product move. But it still matters if you are the one paying the bill.

For developers, the split is simple. If your app is input-heavy, you breathe easier. Document search, code assistants, agents, and RAG pipelines all benefit. If your app is output-heavy, you still have to watch the bill. Writing tools, translation tools, and report generators will not feel the same relief.

You can move batch jobs to idle hours. You can shorten outputs. You can stream. But you cannot cache a sentence you have not written yet.

I talked to a friend who runs a translation service. He was excited about the price cut until he ran his own numbers. His input is short. His output is long. The input cut saved him almost nothing. He is now looking at other models. He does not want to switch. But he might have to.

Competitors have a harder problem. They can match the price. They cannot match the cost structure overnight. A 99.23 percent cache hit rate is not a marketing number. It is an engineering result. Without it, a 0.02 yuan input price is a loss. With it, the price is sustainable.

That is the real barrier. The price war is now a caching war.

I closed my spreadsheet around 2 a.m. The input column was almost empty. The output column was still there, staring back. The 0.02 yuan number is nice. It is not the number that decides whether I can scale.

That number is 4.

I do not know what DeepSeek will do next. Maybe they will cut output prices too. Maybe they will introduce a new tier. Maybe the next model will change the math again. For now, I am planning around 4 yuan per million tokens. That is the number I can control. The rest is noise.

If you are building something with these models, run your own numbers. Do not trust the headline. Trust your bill.

tech newsfact or fictionthought leadersproduct review

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin