DIY AI News

GPT-6 Sol and Claude Opus 5.5 rumours: what is confirmed

An X post points to imminent releases from OpenAI and Anthropic. The launch claims remain unconfirmed, but the proposed price-performance improvement raises a useful question for people paying for AI work.

GPT-6 Sol could arrive on Tuesday, 22 September, with Claude Opus 5.5 potentially appearing on 21 or 22 September, according to a post from @kimmonismus. The account explicitly cautions that the information is unconfirmed.

The post predicts substantially better price-performance for GPT-6 Sol than Astra and frames the possible Opus release as a response to OpenAI. Neither that competitive explanation nor the launch timetable has been independently verified by DIY AI.

For developers and businesses, the useful question is whether either model would make reliable work cheaper. A new name, a lower token price and a better result on a difficult task are three different things.

What the official model pages actually show

At the time of checks on 21 September 2026, OpenAI’s public model catalogue listed GPT-6 Astra alongside the GPT-5.6 family. It did not list GPT-6 Sol. Anthropic’s main model comparison listed Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5, rather than Opus 5.5.

Anthropic’s newsroom dates its Fable 5.1 and Mythos 5.1 announcement to 1 September. Those are documented releases; they should not be confused with the newer Opus rumour.

The current published API rates provide a baseline for assessing any future announcement:

Published modelInput per million tokensOutput per million tokens
GPT-6 AstraUS$10US$50
GPT-5.6 SolUS$4US$20
Claude Fable 5.1US$10US$50
Claude Opus 5US$5US$25

These are standard base rates for existing models, not leaked prices for the rumoured releases. Caching, processing options and separately charged tools can change the bill.

Why a cheaper token is not necessarily cheaper work

Using those base rates, a hypothetical workload consuming 100,000 uncached input tokens and 10,000 output tokens would cost US$1.50 on Astra, US$0.60 on GPT-5.6 Sol, US$1.50 on Fable 5.1 or US$0.75 on Opus 5. This is arithmetic for identical token consumption, not a measured comparison of the models completing the same task.

Identical requests need not produce identical amounts of work. A model might inspect more files, generate longer reasoning, repeat unsuccessful steps or require a second attempt before its answer is usable.

OpenAI’s published Astra evaluation already includes a relevant example. On Terminal-Bench 4.0, OpenAI reports 57.9% for Astra against 37.3% for GPT-5.6 Sol, with approximately 9% lower estimated API cost per task in the configurations shown.

That is a vendor-reported benchmark result, not a promise about every coding job. But it shows why Astra’s higher token rates don’t automatically make it the more expensive option for a completed task.

A future Sol release could be valuable without beating Astra at every difficult problem. Matching the quality needed for routine work while reducing total cost or waiting time would be useful in its own right. Conversely, a lower price would be less compelling if it came with more retries or manual correction.

The separate claim about “Bel” and AGI

The same X post also claims that a system called “Bel” is described internally as artificial general intelligence. DIY AI could not independently verify that assertion.

An alleged internal label is not an independently demonstrated capability. It does not establish that GPT-6 Sol is AGI, or that the two names refer to the same system. Evaluating such a claim would require a stated definition, documented tests and enough evidence to examine failures as well as successes.

It should therefore remain separate from practical questions about a product release, such as which users can access it and what they will pay.

What this means for ChatGPT and Claude users

Do not treat the rumour as a reason to upgrade a subscription or change production model settings. First establish whether a released model is available through the product you actually use: a chat application, a coding agent or an API integration.

For subscription users, the relevant test is how much useful work fits inside the allowance. For API customers, it is the total metered cost of accepted output. A change to one does not, by itself, establish an improvement in the other. DIY AI’s Claude Code pricing guide explains why those access routes need separate comparisons.

In a Reddit discussion about the release rumours, some commenters prioritised Codex usability and usable allowances over another model announcement, while others questioned the changing version numbers. These are individual reactions rather than representative customer research, but they identify questions that a launch announcement needs to answer.

What would make a release worth switching for?

For a coding team, a useful first comparison would run the new model against a contained bug fix, a multi-file change and a code review using the same starting repository and acceptance criteria. Record completed tasks, failed attempts, elapsed time, total usage and developer corrections. Keep approval requirements unchanged during the comparison.

The surrounding agent also deserves attention. Repository access, permission handling and recovery from mistakes can affect the outcome independently of the underlying model. Our Claude Code versus OpenAI Codex comparison covers those workflow differences.

The next evidence to look for is a provider announcement naming the model, documented access and pricing, and reproducible evaluations. Until then, the proposed releases are worth watching—not a confirmed change to what customers can buy or use.

Written by Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Back to AI News