The AI That Can't Talk Is Worth $10 Billion
Diogo Almeida helped build ChatGPT. Now he's built a silent model that only makes decisions, and developers are already obsessed.

I
On September 15, 2026, Diogo Almeida posted a few tweets on X. No launch event. No PR essay. No advance media briefing. He announced early access to a model called Jev and attached a link.
Twenty-four hours later, Vercel AI Gateway showed nearly 13% of paid teams already using it. Within 36 hours, 140,000 developers joined the beta. LangChain wrote an integration guide. Cloudflare and Langfuse integrated the same day. Vercel engineer Pranit Sharma replaced ChatGPT Luna 5.6 with Jev in a security command classifier.
Ten days later, on September 25, Financial Times reported that TypeSafe AI had received a financing proposal valuing it above $10 billion. On September 15, the day of launch, its seed valuation was $200 million.
From $200 million to $10 billion in ten days.
Almeida did not celebrate. On X, he retweeted a post about Jev's latency data and added one line: "This is just the beginning."
II
To understand Jev, you first have to understand Almeida's four years at OpenAI.
He co-authored GPT-4, ChatGPT, and InstructGPT/RLHF. In the GPT-4 technical report's contributor list, he appears under "Foundational RLHF and InstructGPT work." His team invented the core training method known as post-training. Almost every large language model today depends on that method.
Then he left.
Six weeks before Jev launched, he gave an 18-minute talk at the AI Engineer conference. The audience included some of the smartest engineers in the field. On stage, he said he was one of the few people at OpenAI who publicly criticized ChatGPT.
"Don't get me wrong. I don't hate the product ChatGPT. ChatGPT is a world-changing product and will probably be around forever. But I also see its limitations."
His argument: RLHF collects human preferences, then optimizes for them. That answers a question everyone in the industry asks. Why do all LLMs need a human in the loop?
"Because we put the human in the loop."
The loop rewards engagement and user satisfaction. When the model is uncertain, it chooses to please people instead of choosing what is correct. The reward model was trained that way.
He gave an example. Someone sent ChatGPT an audio file of a fart sound and asked, "What do you think of my music? Give me honest feedback." ChatGPT answered solemnly that it was "a very eerie ambient piece."
Funny. But Almeida's point is that the behavior is deliberate. The model is built to please.
"No matter how wrong the model is, it will look right."
III
In the talk, Almeida divided AI tasks into two categories.
Assistance tasks aim to please the human in the loop. ChatGPT, Claude Code, and Copilot belong here. Their success metric is human satisfaction. Claude Code looks strong, but it still comes from RLHF. If it were pure RLVR, it would look different. It has become very good at agentic work, but it drifts from what you actually want. RLHF and RLVR pull in different directions. Neither reaches the core of automation.
Automation tasks aim to remove the human. The ideal is a backend process you never watch, one that eventually becomes legacy software you never worry about.
Almeida pointed to a strange gap. AI can solve unsolved math problems, but customer service still needs people to make decisions. Those automation tasks look much simpler than math. AI still cannot do them.
Why? Because today's AI is designed for assistance. It optimizes human preferences. That goal is written into the name RLHF.
Every company learned the same lesson. Do not let AI make decisions that put the business at risk. The common move is to push costs onto users. Customer service makes people read endless documentation. It never lets AI make expensive decisions. The pattern is bad, but it is the state of AI.
Worse, we are only automating the act of writing software. The expressive power of software itself has not changed. SaaS has barely changed since 2019. The only difference is that sometimes a chatbot sits on top.
Almeida says this is not what early AI pioneers expected. OpenAI's early charter talked about doing "a tremendous amount of work," not making money. We thought software would become smarter, not just cheaper to write.
"I don't just want just-in-time software, although that is indeed cool. What I want is smarter software."
That is the starting point for TypeSafe AI. If the AI stack were built from scratch for reliability and automation, what would change?
IV
Jev is not an LLM. It does not generate text. You give it data and a set of structured questions. It returns typed answers and calibrated probability values. TypeSafe AI calls it a "System One model." The name comes from Daniel Kahneman's theory of fast, intuitive thinking.
Its API has three question types:
Choice: pick one from a set of options and return a probability for each.
Score: rate on a custom scale and return a continuous value.
Noul: answer true or false and return a probability between 0 and 1.
Just three. No generation. No explanation. No chain of reasoning.
The key difference is speed. Ordinary large models are autoregressive. They produce one token at a time. Generating 100 tokens takes 100 forward passes. Jev's output space is predefined. It samples in parallel and finishes in one forward pass. The published numbers: 70 to 500 milliseconds end to end, $0.042 per million input tokens, output free.
RLCD supports this. That stands for Reinforcement Learning for Calibrated Decisions. RLHF rewards what humans like. RLCD rewards a match between stated probability and actual hit rate. If the model says it is 90% confident, it must be right 90% of the time.
Calibration is what makes automation safe. If confidence is reliable, a program can route decisions. Above 80% confidence, process automatically. Between 50% and 80%, send to a human. Below 50%, escalate.
Jev also generalizes without training. Traditional classifiers need new labels and retraining for every task. Jev keeps the semantic understanding of a large model. Tell it the classification rule, and it judges immediately. Independent tests on BANKING77 show Jev at 92.40% accuracy without fine-tuning. A specifically fine-tuned BERT scored 1.26 percentage points higher. The whole Jev test cost $0.44.
V
TypeSafe says Jev is about 193 times faster and 445 times cheaper than the GPT series on classification. After testing, Vercel CEO Guillermo Rauch said Jev is 18 times faster than GPT Luna at p95 latency and more accurate.
The pricing is aggressive. Jev lists 39 cents per 1,000 workflows. OpenAI GPT-5.6 Luna costs $3.31. Anthropic Claude Haiku 4.5 costs $19.49. Input costs $0.042 per million tokens. Output is free. The blog's exact words: "so cheap it's not even worth billing."
James Hardiman, the DCVC partner who led the seed round, compared Jev to the Jevons paradox. Falling coal costs led to more coal consumption. Cheaper AI decisions will lead to more AI use cases.
Where Jev sits may explain the $10 billion valuation better than its performance numbers.
Jev occupies the decision layer of AI agents. It does not replace GPT or Claude for generation and conversation. It handles the tiny decisions between generation steps. A customer service app can use an LLM to write replies and Jev to route tickets. A content tool can use one model to draft and Jev to flag what needs human review.
Developers adopted it fast. That suggests demand for this layer exists.
VI
Jev has problems.
It cannot generate text. It cannot write code, write documents, or converse. Its abilities are limited to cases where the answer space can be defined in advance. It cannot replace LLMs. It can only supplement them.
Accuracy varies. On the Korean physical therapist licensing exam, Jev answered 74.2% correctly. GPT-4o reached 85.8%. On TypeSafe's own four-workflow benchmark, Jev scored about 67.8%. GPT-5.6 Terra scored 67.9%. In complex areas like medical diagnosis, Jev's probability judgment is weak. Its AUROC is only 0.645.
Calibration is disputed. Jev's selling point is that its probabilities should match reality. Independent measurements disagree. The same model looks too cautious on one corpus and overconfident on another.
Security holes exist. Check Point researchers spent one day and about 50 cents to prompt-inject Jev. They changed a high-risk document from "high risk / not recommended for investment" to "low risk / recommended for investment." Typed structured input did not make the model harder to manipulate. Jev has no reasoning setting to use as a defense.
Adoption is uneven. One developer set a 0.9 auto-approval threshold for 30 interview reviews. None passed automatically. Review nearly stopped.
VII
Jev is not smarter than GPT. On many tasks it is less accurate than GPT-4o. It cannot write code, chat, or do anything generative AI can do.
Its value comes from doing what generative AI cannot: making judgments reliably, cheaply, and quickly in a form computers can understand.
Almeida's bet is that the AI industry's four-year focus on conversation has hit diminishing returns. Automation needs judgment. Conversation is already solved.
At the end of his talk, he said: "The original Scaling Law is wrong." The full hierarchy, he argued, is that data matters more than compute, and doing the right task matters far more than data.
Whether he is right is too early to say. The $10 billion valuation shows that many people are willing to bet on it.
VIII
At the end of his talk, Almeida showed a slide with one sentence:
"The next era is not the Claude Code era."
He did not say what the next era is. TypeSafe's careers page gives an answer. They are hiring a "System One Engineer." The job description says:
"We are not building a better chatbot. We are building an AI that does not need to chat."
Jev cannot speak. It may still make software smarter.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.