01 logo

OpenAI Cut GPT-6 Prices in Half. The Benchmarks Show What You Give Up.

Sol and Luna are cheaper, faster, and stronger at coding. They also regressed on knowledge work, and some users say the models are quietly routed to cheaper versions.

By JinPublished 3 days ago • 5 min read

GPT-6 Sol and Luna: Cheaper, Faster, and Not Always Better

On September 22, 2026, OpenAI released GPT-6 Sol and GPT-6 Luna. Less than a month after GPT-6 Astra, the product line shrank from four tiers (Astra, Sol, Terra, Luna) to three: Astra, Sol, Luna. Terra is gone. Sol's price moved into the former Terra range. Luna's input price fell to $0.10 per million tokens.

Anthropic released Claude Opus 5.5 the same day, priced at $4 input and $20 output. Two leading companies moved 90 minutes apart. The competition shifted from who is smarter to how much the same work costs.

Price: Cost per Task Becomes the Main Measure

GPT-6 Sol's API pricing is $2 input and $10 output, a 50% drop from GPT-5.6 Sol's promotional price. Luna is $0.10 input and $0.50 output, also halved.

Artificial Analysis gives a clearer number. Sol costs $1.06 to run one Intelligence Index task at maximum reasoning effort. The previous generation cost $1.99. Luna costs $0.07 per task. The previous generation cost about $0.18. That is a drop of about 60%.

Luna's input price is lower than DeepSeek V4.1 Flash's off-peak price of $0.15. DeepSeek still wins on cache hits. Its off-peak cache-hit input is RMB 0.02 per million tokens. In cache-hit scenarios, Luna is not ahead.

OpenAI also changed its caching. Cache reads keep a 90% discount. Changing reasoning effort or switching tools no longer breaks the existing KV/prompt cache. Before, using low effort for simple steps and raising it for hard questions could invalidate the prefix. Now configuration_update lets the cache continue.

GitHub says Copilot now reprocesses more than 50% fewer prompt tokens. Manus reports cache hit rates rising from about 85% to a sustained 90% or higher. In agent work, one task may trigger a dozen tool calls. Each call reuses context. The actual bill can be lower than list price.

Cache lifetime rose to 30 minutes. OpenAI added explicit cache breakpoints, cache diagnostics, and cache prewarming. Developers can precompute system prompts, tool definitions, and reference material before a user request.

These infrastructure changes did not appear on the pricing page. They are what makes Luna at $0.10 per million tokens possible.

Performance: Flat Overall, Mixed in Specific Areas

Artificial Analysis found that Sol and Luna are roughly flat against GPT-5.6 on the overall Intelligence Index. Sol scores 48 at maximum reasoning effort. Luna scores 37, which matches GPT-5.6 Luna.

Price halved. Scores stayed still. Cost per unit of intelligence now matters more than the absolute score.

General benchmarks have hit diminishing returns. The changes are in specific areas.

Areas that declined:

Hallucination rate. Sol fell from 92% to 60%. Luna fell from 93% to 77%. Sol gets there by refusing to answer more often. It attempted only 83% of questions, down from 99%.

Knowledge work. On GDPval-AA v2.1, Sol's Elo rating fell by about 100 points. Luna fell by about 75. After reviewing hundreds of outputs, Artificial Analysis said the regression came from lower presentation quality and deliverables that missed parts of the scoring rubric.

Luna's coding tasks. Its coding agent index fell 2 points to 41. SWE-Atlas-QnA scored 44%, down from 49%. DeepSWE v1.1 scored 64%, down from 66%.

Health and medicine. Output length shrank. Sol's average answer fell from 1,764 characters to 977, a drop of nearly 45%. Luna shrank by 35%. Scores fell on HealthBench/Hard questions that require large amounts of medical detail.

Areas that improved:

Sol's coding agent index rose 2 points to 57. Terminal-Bench 4.0 rose from 37% to 43%. AutomationBench scored 62%, up from 60%. Luna scored 53% on AutomationBench, up from 50%.

Sol's coding agent index in the Codex environment also rose 2 points.

The split shows that GPT-6 Sol and Luna do not beat the previous generation everywhere. They are stronger at coding and automation. They give up depth and completeness in knowledge work. ZDNET's assessment: the selling point is price, not intelligence.

Inference Infrastructure: Async Tool Calls and Mid-Turn Steering

GPT-6 supports asynchronous tool calling. After calling a slow tool, the model does not wait. It can keep reasoning, call other tools, and handle other parts of the request.

Mid-turn steering lets a user insert an instruction while the model works. For example, "drop this direction, switch to B." The Responses API keeps the completed work and continues.

Dynamic reasoning effort lets a long session move among low, high, and max. Before, those changes could invalidate the prefix. Now they do not.

These features make GPT-6 feel more like a runtime than a request-response model. For agent developers, that means less waiting, less repeated computation, and smoother interaction.

Industry Impact: The Middle Tier Disappears, the Price War Escalates

The GPT-5.6 line had Sol (flagship), Terra (balanced), and Luna (lightweight). The GPT-6 line has Astra (top), Sol (mid-tier), and Luna (lightweight). Terra is gone. Sol's price moved into the former Terra band.

OpenAI decided that mid-tier and high-tier no longer need separate products. Sol covers both.

In the market, Sol's price sits at Claude Sonnet 5's level. Luna pushes into the low-cost market. On DeepSWE 1.1, Sol scores 68.8% in max mode. Claude Fable 5's highest score is 69.9%, a gap of 1.1 percentage points. Sol's cost per task is about 80% lower. Luna scores 66.6% on the same test, close to Opus 5 and Fable 5 at medium effort, at 93% to 96% lower cost.

When closed-source models cost less per task than open-source inference, open-source models lose their traditional cost advantage.

User Experience: Two Kinds of Feedback

A V2EX user wrote: "6 sol is pretty good. Thinking is shorter, it gets work done much more briskly, no longer in the unusable state of the 5.6 era, and the price is cheaper." Shorter output helps. Sol's average answer fell from 1,764 characters to 977, a drop of nearly 45%. For daily coding and quick questions, less filler is welcome.

Negative feedback is just as strong. Many users report dumbing down. One wrote: "I just refunded. In the past, dumbing down was hard to notice. Now the GPT-6 gap is too obvious." Another said, "sol is basically routed to luna too." That user offered a test: "Just try the candy question. A full-powered sol at any reasoning level can answer 21. A dumbed-down sol always says 2." Others believe "6-sol and 6-luna are probably distilled and quantized versions of the previous generation."

The low price has another side: service quality can vary. Some dumbing down may come from routing. Requests may go to cheaper models instead of Sol or Luna. For developers running production APIs, that uncertainty can matter more than price.

Conclusion: The Limits of Cost Efficiency

GPT-6 Sol and Luna show how hard OpenAI is pushing cost efficiency. Sol beats Opus 5 at about 9% of its cost per task. Luna reaches the previous flagship level at about 1/100 the cost.

At the same time, the overall Intelligence Index is flat. Knowledge work regressed. Some coding benchmarks fell. OpenAI chose to keep coding and automation strong while compressing depth and output length in knowledge work. That is a deliberate trade-off.

The developer decision is clearer now. For standardized workflows, cost sensitivity, and coding or automation, Sol and Luna are the most cost-effective choices on the market. For complex knowledge work that needs deep analysis and complete deliverables, Astra remains the better option. OpenAI's official statement: "If you want the best results and an uncompromising experience, we still recommend GPT-6 Astra."

The larger shift is the move toward price per task as the main competitive measure. The next question is how much room is left for cost reduction, and whether AI applications will expand as expected once cost stops being the bottleneck.


tech newsfact or fictionthought leaders

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin