01 logo

The Token Reckoning — Why Companies Are Quietly Rehiring the Programmers They Fired

AI was supposed to make human coders obsolete. Then the monthly token bill hit $221 million, and everything changed.

By JinPublished 2 months ago 4 min read

In June 2026, a screenshot circulated through developer chat groups. It showed a notice posted by the CTO of a mid-sized SaaS company in an internal tech channel: “We’ve calculated that our AI coding tool’s token expenses over the past three months have exceeded the total personnel cost of two mid-level engineers. Effective immediately, team Copilot seats are suspended, and we’re reopening Java positions.” The first reply in the group consisted of just four words: “Feng shui turns.”

The anecdote was soon compressed into a catchier one-liner: “Because you can’t stiff AI on its bills, companies are hiring programmers again.” Exaggerated, sure. But beneath it lies a stack of invoices.

Uber gave roughly 5,000 engineers access to Claude Code in November 2025. The idea wasn’t to replace anyone, just to let them “write faster.” By February 2026, adoption hit 32%; by March, 84%. In April, CTO Praveen Neppalli Naga said in an internal meeting: “This year’s AI budget—already spent.”

He offered his own case: a single two-hour demo burned through $1,200 in Claude Code fees. After that meeting, the company imposed a monthly token cap of $1,500 per person. It also scrapped an internal leaderboard that had been designed to encourage AI use, because “driving adoption” and “controlling spend” belonged to two different teams. The first team got bonuses; the second inherited the mess. To this day, one line in Uber’s financial model remains blank: “Unable to directly correlate token consumption with consumer feature output.”

Over at Meta, 85,000 employees burned through 73.7 trillion tokens in 30 days. At estimated market rates, the monthly cost came to roughly $221 million. Andrew Bosworth, the company’s technology chief, shut down an internal consumption leaderboard called “Claudeonomics.” A line from his statement was screenshotted and leaked: “Token usage alone does not measure impact.” In other words, the person with the biggest number on the board wasn’t necessarily producing the most. He might simply have been more willing than his colleagues to spend the company’s money.

Tencent’s moves were more granular. Previously, each employee received a token allowance worth roughly $2,000 a month, almost uniformly. Starting in June 2026, quotas became dynamically assigned by department, role, and task scenario. Most people’s allowances dropped to between 1,000 and 5,000 RMB. One employee posted a screenshot showing they were left with just 1,400 RMB. Another colleague, however, got an increased quota—because he was producing three times as much code. Internally, Tencent calls this an “efficiency bargain”: it’s not that you can’t use AI; it’s that you have to use it cost-effectively.

Why can layoffs be delayed and negotiated, while AI bills must be settled by the end of the month?

Human costs have given. Salaries can be paid a few days late, social insurance deferred for months, or two people’s work piled onto one. None of it stops the production line immediately. But AI services are utilities: miss a payment, and the service stops. The algorithm doesn’t do favors. It doesn’t take phone calls that say “we’ll cover it next week.”

That rigidity carves a different curve into a company’s cash flow than labor costs do. An engineer on a 300,000 RMB annual salary is a number you can predict at the start of the year. An engineer using AI coding tools might cost 200 RMB one day and 2,000 RMB the next, just because they ran a long-chain agent task. The budget becomes a blind box.

More important is the incentive structure. When using AI, the line between “using more” and “wasting” is film-thin. The internal leaderboards at Uber and Meta prove it: if “consumption” is silently accepted as a proxy for “effort,” employees will rationally choose to consume more. Human overtime at least has biological limits. AI has no limits, only prices.

Companies aren’t retreating to manual work. They’re doing something they once did during the cloud era: cost governance.

Uber’s $1,500 cap was crude rationing. Tencent’s dynamic allocation went a step further: it turned tokens from “free office supplies” into “budget assets that require justification.” Each department now has to explain the ratio of “effective tokens” (those that produce usable code, documents, decisions) to “waste tokens” (redundant calls, circular agent loops, consumption for leaderboard climbing). One project manager said that in weekly meetings now, the token consumption curve and the sales conversion curve are projected on the same slide.

The pattern is familiar. Between 2006 and 2012, enterprises lurched from “cloud carnival” to “cloud bill shock” and eventually to FinOps. They invented resource tagging, reserved instances, and budget alerts. Now the same script is playing out again, with AI in the lead role.

Programmers haven’t been “recalled” en masse. But how they work has been turned inside out.

One visible change: “knowing how to save tokens” is hardening into a concrete skill. The Tencent employee who got an extra quota because of high code output said something in an internal talk: “I spend more time writing prompts than writing code.” The implication: going forward, measuring an engineer won’t just mean asking “can they get it done?” It will also mean asking “how much did they cost the company to have the AI do this part?”

In McKinsey’s 2025 global AI survey, among nearly 2,000 companies, only 39% said AI had a clear positive contribution to profit. For the other 60-plus percent, AI spending was eating cash flow. A footnote next to that figure: at the companies where AI did contribute to profit, “output per token” was one of the internal evaluation metrics.

In a recent public appearance, Tencent’s Dowson Tang said most of the company’s code is now generated by AI, and that engineers spend more time on architecture design. He didn’t use the word “replace.” But one move stands out: they’ve begun splitting “defining the problem” and “writing the code” into two jobs. The first goes to humans; the second, to the models.

The CTO of that SaaS company, the one who suspended Copilot seats and restarted Java hiring, later did a retrospective in a community forum. His final line: “We’re not hiring people to write code. We’re hiring people to do the math—to figure out which lines are worth paying an AI to write, and which aren’t.”

That sentence contains no uplift, no tidy label. It is simply a crumpled token bill, with someone’s penciled notes on the back.

tech newsthought leaders

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin