OpenAI says it has made its newest GPT-5.6 models significantly more economical to use. In an update to its API pricing, two models—Luna and Terra—now cost less per input and output token, reflecting an ongoing push to improve efficiency across the company’s offerings.
For developers and teams that rely on automated workflows, the change matters because token-based pricing directly affects how much you can accomplish within a given budget or allowance. OpenAI also highlights that the new rates influence how usage is counted in tools like Codex and ChatGPT Work.
What changed in GPT-5.6 API pricing
OpenAI reduced the price of two GPT-5.6 models. The biggest drop is for Luna, while Terra also becomes cheaper, just not as steeply.
Luna: lower token costs
With the updated pricing, GPT-5.6 Luna now costs:
- $0.20 per million input tokens
- $1.20 per million output tokens
OpenAI notes these figures are down from $1 per million input tokens and $6 per million output tokens.
Terra: a smaller but meaningful cut
GPT-5.6 Terra also drops in price. Under the new rates, it costs:
- $2.00 per million input tokens (down from $2.50)
- $12.00 per million output tokens (down from $15)
Overall, the update positions both models as more budget-friendly options for production work where token volume can add up quickly.
How cost-efficient GPT-5.6 affects Codex and ChatGPT Work
OpenAI said the new prices impact how it counts usage in Codex and ChatGPT Work. That is important because many users don’t experience pricing purely as “cost per token”—they experience it as allowances or quotas that determine how much work they can run.
In simple terms, if your tasks use these GPT-5.6 models under the updated pricing, they can deduct less from your allowance. The result is that you may be able to complete more work within the same quota.
For teams managing multiple agents, recurring jobs, or iterative development loops, this kind of adjustment can translate into a noticeable increase in throughput—especially for workflows that generate large outputs, where output token pricing often becomes the dominant factor.
Upgrading Auto-review to GPT-5.6 Luna
OpenAI is also rolling forward model improvements in parts of its ecosystem. Specifically, it says it is upgrading Auto-review in the ChatGPT app and updating Codex CLI from GPT-5.4 to GPT-5.6 Luna.
OpenAI expects this shift to reduce costs by approximately ten times. While the exact savings depend on usage patterns, the direction is clear: more tasks can be run using the cost-efficient GPT-5.6 Luna model without the same token cost pressure.
GPT-5.6 Sol gets a Fast mode for the API
Beyond price cuts for Luna and Terra, OpenAI also describes a performance option for API customers: a Fast mode associated with GPT-5.6 Sol.
OpenAI says there are no changes to Sol’s standard pricing at this time. Instead, Fast mode gives customers a speed boost while charging a higher rate for that extra performance.
How much faster is Fast mode?
According to OpenAI, GPT-5.6 Sol Fast mode can be up to 2.5 times faster than standard processing. Importantly, OpenAI says this improvement does not reduce the model’s intelligence.
What does it cost?
The performance comes at a cost: Fast mode is priced at twice the standard API price. OpenAI frames this option as best suited for scenarios where time matters.
It’s particularly designed for time-sensitive coding, research, and “agentic” workloads—situations where waiting for results can slow down iteration and delivery. For many other use cases, OpenAI suggests Fast mode isn’t necessary, and standard processing may be the more cost-effective choice.
Why the company expects efficiency gains
OpenAI links these updates to improvements made to its models. The company says recent progress in GPT-5.6 Sol enabled the efficiency gains behind the price reductions for Luna and Terra.
In other words, OpenAI is positioning the pricing changes as more than a marketing adjustment. It’s presenting them as a result of real efficiency improvements in the underlying systems—changes that can reduce the compute cost of generating responses.
Efficiency and intelligence: Luna’s position in tests
OpenAI also references its own evaluation results. In those test outcomes, Luna is placed at the top of the company’s intelligence index among the compared models.
Notably, OpenAI states this occurs despite Luna having a substantially lower cost per task. For teams comparing options, that combination—strong performance alongside reduced token pricing—can be a compelling reason to shift workloads to Luna, especially when budgets are tight or when outputs are lengthy.
Practical takeaways for teams using GPT-5.6
If you currently build applications or workflows around GPT-5.6, the updates have several practical implications.
- Review token-heavy pipelines: Because output tokens often dominate spend, the Luna output token reduction can have outsized benefits for tasks that generate long responses.
- Re-check allowance behavior: OpenAI notes usage counting changes in Codex and ChatGPT Work. If you have quotas, you may get more runs per allowance now.
- Consider workflow upgrades: With Auto-review and Codex CLI moving to Luna, you may see improved cost dynamics without changing how you trigger review or command-line tasks.
- Use Fast mode selectively: Fast mode for GPT-5.6 Sol can help when deadlines are tight, but it charges double the standard API rate, so it may not be ideal for every task.
Overall, OpenAI’s messaging centers on one theme: cost-efficient GPT-5.6 should help customers do more with the same spend while still keeping model quality in focus.
Security note: test your detection layers
While this article focuses on pricing and performance, OpenAI’s update arrives alongside ongoing reminders from security teams about operational visibility. Some attack paths continue when defenders don’t alert, leaving gaps for threats to slip by.
One approach described in a security context is to validate detection coverage by running breach and attack simulation tests against your SIEM and EDR rules—before attackers find the weak spots. The goal is to ensure the monitoring stack catches the majority of successful attempts, rather than only a small fraction.
For organizations adopting new AI tooling, it’s a useful reminder to align new automation with strong monitoring and well-tested detection policies.
Conclusion
OpenAI’s latest changes make cost-efficient GPT-5.6 a clearer value proposition. With Luna priced down sharply and Terra reduced across both input and output tokens, developers can potentially complete more work within the same allowance—especially in Codex and ChatGPT Work where usage counting is affected.
At the same time, OpenAI is modernizing related features by upgrading Auto-review and Codex CLI to GPT-5.6 Luna, expecting major cost reductions. And for teams needing speed, GPT-5.6 Sol Fast mode offers up to 2.5x faster responses at twice the standard price.
If you run frequent AI tasks, it’s worth revisiting your model selection and budgeting assumptions to reflect the new token economics and performance options.
