OpenAI announced a price cut across the GPT-5.6 line on July 30: GPT-5.6 Luna down 80%, GPT-5.6 Terra down 20%, flagged by Simon Willison as a "huge price drop."
The stated mechanism is the interesting part. In a companion post, How GPT-5.6 fuses frontier intelligence with frontier efficiency, OpenAI credits GPT-5.6 Sol with enabling the reduction — using the model to optimize load balancing across its fleet and, more notably, to optimize inference itself.
Why it matters
If the story holds, it's a recursive-efficiency claim rather than a hardware or quantization one: the frontier model paying for its own serving costs by rewriting the serving stack. An 80% cut on Luna also resets the floor for anyone benchmarking cost-per-task against competitors.
In adjacent reading, a JuliaHub evaluation of GPT-5.6 versus Claude Fable 5 on physical-AI tasks picked up 98 points and 21 comments on Hacker News the same day.
Sources: Simon Willison, OpenAI, HN discussion