OpenAI closed September with three moves. On 3 September it announced Astra, the top of the GPT-6 family, and opened it to general use the next day. On 22 September came the family's middle and light models, Sol and Luna. And before the month was out, on 29 September, GPT-6.1 Sol.
At first glance it is a familiar story: a big model, a mid-size model, a small model. But this time the hero of the story isn't the top model; it's the one in the middle.
The family: Astra, Sol, Luna
- Astra: the most capable model in the family. Its context window exceeds a million tokens (1,050,000), so it can read long documents and large codebases in one go.
- Sol: the middle tier, for heavy everyday work.
- Luna: built for speed and cost; a light model that gives up a little accuracy.
Astra: power, with the safety brake on
Astra's release wasn't a straight road. OpenAI held the model back for a few weeks over cybersecurity risks and released it with additional safeguards. The version opened to general use refuses some cybersecurity-related requests. Early access went to organisations in OpenAI's application-based cybersecurity programme; it then rolled out to ChatGPT's higher plans, the API and AWS.
Two technical details stand out. According to OpenAI's vice president of research, Astra is the first model pre-trained on more than 100,000 GPUs at the Stargate site in Texas. It is also reported to use a technique referred to as "recurrent depth", which hides part of the model's reasoning steps from view. The second point is controversial, because being able to watch how a model thinks was valuable for safety.
Price: 10 dollars per million input tokens, 50 dollars per million output tokens. Top-tier pricing.
GPT-6.1 Sol: getting close to Astra at a fifth of the price
This is the real news of the month. According to OpenAI, GPT-6.1 Sol almost catches up with Astra in agentic coding, computer use and professional work. Its price is a fifth of Astra's: 2 dollars per million input tokens, 10 dollars per million output.
The published results:
- DeepSWE 1.1 (software engineering): the same score as Astra.
- OSWorld 2.0 (computer use): 2.1 points behind Astra.
- AutomationBench (business automation): 4.8 points ahead of the earlier GPT-6 Sol.
Independent measurement paints a similar picture. On the Artificial Analysis intelligence index, at the highest settings, GPT-6.1 Sol scores 52 and Astra 53; the cost per task is 0.72 dollars for Sol and 3.26 dollars for Astra.
One point of difference, four and a half times the price.
How to read these numbers
Most of the tests are ones OpenAI picked itself. Independent measurements seem to confirm the overall picture, but the details will differ from job to job. The ratio that matters to me: being able to do the same work at a fifth of the price, at very close quality.
What changes for me
I follow the models of both companies; I'm not sitting in the stands of a race. What matters is which job I give to which tool.
The picture in the GPT-6 family is almost identical to the one I saw at Anthropic the same month: top performance is moving down to a lower model at a much lower price. At Anthropic, Opus 5.5 reached Fable level; at OpenAI, GPT-6.1 Sol came close to Astra.
The practical consequence is simple:
- A mid-tier model for the default work. Long documents, multi-step jobs, code. GPT-6.1 Sol is the new reference in this class.
- The top model, when it's genuinely needed. Paying five times as much for a one-point difference makes no sense for most work.
- The real skill is routing. Knowing which job to give to which model is becoming more valuable than picking the best model.
Then there's the safety question. Astra being held back over cybersecurity, and part of its reasoning being invisible, is a reminder that these models are no longer ordinary software. I use these tools in my agency's technical infrastructure; the stronger they get, the more attention I pay to what I leave automated and what I check myself.
Conclusion
A year ago everyone was asking "which model is the best?" Today the right question is a different one: which is the most efficient model that is good enough for this job?
The ordinary looks at the summit. The rebellious looks at where the job ends.