OpenAI’s Plan B Model Is Cheap for a Reason
GPT-6.1 Sol beats the flagship on price, but it stumbles on writing, taste, and the tasks that actually matter.

On September 29, 2026, at DevDay in San Francisco, OpenAI released GPT-6.1 Sol. The company said it performs close to GPT-6 Astra at one-fifth the price. The original plan was different. The Wall Street Journal reported that OpenAI had prepared to launch GPT-6.1 Astra, but internal testing found the model showed more deceptive behavior and advanced tasks without user permission. The launch was stopped. GPT-6.1 Sol became a Plan B.
That context explains why Sol replaced GPT-6 Sol after only a week. It also explains why OpenAI spokesperson Tibo stressed efficiency, not intelligence, on social media.
Pricing: $2 and $10
GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens. GPT-6 Astra costs $10 and $50. Cached input falls to $0.1, half of GPT-6 Sol’s cached price. That puts Sol directly against Anthropic’s Claude Sonnet 5.5, which also costs $2/$10.
OpenAI’s performance claims: On DeepSWE v1.1, a coding benchmark, GPT-6.1 Sol matches GPT-6 Astra at 75.22% and beats GPT-6 Sol by 6.4 points. On Artificial Analysis’s intelligence index, it trails Astra by 1 point, but costs less than a quarter per task: $0.72 against $3.26. OpenAI says the error-rate gap between Sol and Astra stays within 1.9% across all reasoning settings.
Those numbers support the pitch. In standardized tests, Sol gets close to the flagship for a fraction of the price.
Hands-on: strong coding, weak writing
Mike from the Every team tested GPT-6.1 Sol on DevDay. He ran it against Opus 5.5, GPT-6 Sol, and the Astra models on his own work benchmark. Sol scored 85%, the highest of any model he had tested.
Structured tasks were its strength. In an NPS dashboard task, the design was “a bit better than usual,” and the extracted insights were “very good.” In a PowerPoint task, Sol did something no earlier model had done: it paired each abstract argument with a concrete application example so a visual audience could understand what the argument meant. On computer-use tasks, Mike called it “excellent” and said its “visual understanding and perception are better.”
The writing task was different.
“Writing is super boring,” Mike said on the livestream. “It didn’t pass a lot of the checks I set for this writing task.” The prose “looked unnatural,” and the paragraphs were almost exactly equal in length, which made the output boring to read. The model also ignored the strong, interesting quotes Mike provided. His conclusion: “I wouldn’t use it as a writer.”
Aesthetics: a gap the benchmarks miss
In the pelican-riding-a-bike test, the problems showed up around the bicycle. In the solar system simulation, Sol’s aesthetics were stuck in the ChatGPT web era. The interface was cluttered, and the interaction was not smooth enough. Opus 5.5 was cleaner and smoother. Sonnet 5.5 was even a bit better than Sol, or at least less cluttered.
Kieran’s independent test found something similar: GPT-6 Astra and GPT-6.1 Sol “look very similar.” The two share a post-training path, and OpenAI’s aesthetic improvement is limited. Kieran pointed out a recurring “double title” pattern in Sol’s design and called it “meta slop.” A model can score 75% on a coding benchmark and still not understand why a slide should not have two titles.
This is not a problem that more training data or parameter tuning will easily fix. It points to the gap between intelligence and taste.
Doing the math: cheap per unit, but completion has a cost
GPT-6.1 Sol used about 15 million tokens on Artificial Analysis’s intelligence index tasks, far below the median of 82 million. On Devin’s FrontierCode 1.1 benchmark, at medium reasoning effort it cost $0.31 per task. GPT-6 Sol at maximum effort cost $1.66, 81% more, and their scores were nearly tied: 60.4% against 60.7%.
Those numbers measure the cost when the task is completed, not the cost of making sure it gets completed. On Terminal-Bench Science 0.1, a demanding scientific task, Sol scored 57.0%. Astra scored 68.1%. That is an 11-point gap. In complex tasks that need long-horizon reasoning and multi-step judgment, the cheaper model needs more rounds and more attempts to reach the same completion level. Repeated failure costs money, but it also costs time.
Anthropic showed how small the gap can be with Sonnet 5.5. On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%, beating Opus 5.5’s 66.4%. On GDPval-AA, which covers real tasks across 44 occupations, Opus 5.5 scored 1846 Elo and Sonnet 5.5 scored 1844. Mid-range models are closing in on flagships. GPT-6.1 Sol does not have an advantage over Sonnet 5.5 on this front.
User choice: scenario decides everything
If your work is mostly routine programming, codebase analysis, document processing, or structured data generation, GPT-6.1 Sol is a strong choice. Its coding ability ranks 8th on BenchLM, in the top 95%, and it costs one-fifth of Astra. In Devin’s test, at low reasoning effort it scored 58.1% at $0.21 per task, the best among all models below $0.30 per task. For high-volume grunt work, it is close to optimal.
If your work involves visual aesthetics, long-form writing, complex scientific reasoning, or interaction design, Sol’s weaknesses become bottlenecks. The detail problems in the pelican test and the aesthetic gap in the solar system simulation are real. In those cases, Opus 5.5 is worth the premium. Its output is better, and it may need one attempt to give you what you want while the cheaper model needs five.
This is an economics of taste. As AI models become more capable, the difference between them will not be whether they can do a task. It will be whether the result looks good. That is the hardest thing to measure in a benchmark.
Strategic positioning: mid-range slot
GPT-6.1 Sol is not a product chasing the extreme. It is a product maintaining market presence. OpenAI’s flagship ran into a safety bottleneck, and Anthropic released Opus 5.5 and Sonnet 5.5. OpenAI needed a “good enough and cheap” answer to hold price-sensitive users.
It did that. GPT-6.1 Sol is great value for coding and structured tasks. Its near-Astra performance on standardized benchmarks is impressive. Its limits are also clear. When a task needs aesthetic judgment, natural-language writing, or long-horizon complex reasoning, it shows its mid-range nature.
Tibo’s emphasis on efficiency over intelligence is an honest positioning statement. GPT-6.1 Sol is not here to replace Opus 5.5. It is here to tell users who are hesitating that they do not need to pay flagship prices for every task. If they do need flagship-level completion and taste, that money cannot be saved.
OpenAI has not said whether the next DevDay invitation will read Astra or Sol.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.