GCC Exchange

GPT-6 Sol and Sol Pro: Where the Price Difference Actually Comes From

Two entries landed on the same day with the same family name and different prices, and the tempting explanation is the wrong one. Sol and Sol Pro were both added on 22 September 2026; both are frontier reasoning models from the same vendor; the second costs more. The natural reading — that the more expensive one is a newer, larger, better model — does not survive contact with the details. The weights are the same. The difference is how much thinking the model is allowed to do before it answers, which changes the token count, which changes how the rate is applied. In this month’s latest LLM models, this pair is the cleanest example of a price difference that is not a capability difference, and it is the reason a comparison table listing both rows as separate models will mislead you twice — once about quality, once about cost. On OrcaRouter the GPT-6 Sol entry for the base model exposes the rate it is billed at on our side, which is the number you would actually pay to run it here, and that is the only rate of ours quoted in this article.

The two rows on one list

Same vendor, same day, two names. One is the base model. The other carries a suffix that the vendor’s own documentation treats as an option on that model rather than as a separate release.

That is the crux: a suffix can mean a new model, a faster serving configuration, or a higher reasoning budget. Here it means the third. There is no new pretraining run, no parameter change, no separate benchmark row for a different set of weights — there is a model, and there is a way of running it that spends more compute per answer.

Once you know that, the two prices stop being a puzzle and start being informative. The gap between them is the vendor’s own estimate of what thorough reasoning is worth, expressed as a rate.

Why a higher reasoning budget costs more

The mechanism is worth spelling out because it drives everything downstream.

A reasoning model generates intermediate tokens before the final answer. Those tokens are produced by the same forward pass as any other token, and they consume the same compute. They are usually billed as output tokens whether or not the user ever sees them.

So a high-budget mode is not a markup added at the till. It is a genuinely more expensive computation, because it produces more tokens for the same request. Two things then follow.

First, the cost of a request depends on the mode even when the visible answer is identical in length. A short answer produced after four thousand hidden reasoning tokens is not a cheap request.

Second, the latency changes with it. Every hidden token takes time to generate, so the wait before the first visible character grows roughly with the reasoning budget. If you measure time-to-first-token on the base mode and then switch to the thorough mode, the number you measured stops describing what your users experience.

What the published score carries with it

The independent index we checked on 28 September 2026 records 47.53 for this family, in a configuration labelled `max`. That label is the whole caveat. A score at the highest reasoning setting is an upper bound on what the model can do, produced under the most expensive configuration available — and it is not the score you will get at the default setting, which is what most callers actually use.

This is the single most common error in comparison writing about reasoning models. Two models are placed side by side; one number was produced at maximum effort and the other at default; the reader concludes that the first is smarter. The configuration label is printed right there in the source and is skipped because the number is more interesting than its footnote.

The habit that fixes it is mechanical: never write a reasoning model’s score without the effort setting beside it, and never compare two scores whose settings differ. If you need a comparison, compare like with like — default against default, or max against max — and treat the max figures as a ceiling rather than a typical result.

The comparison that does answer the question

If you want to know whether the thorough mode is worth its rate, the per-token comparison will not tell you. The question is not which mode is cheaper per million tokens; it is which mode completes your task for less money and in less time.

That has an awkward but honest form. Take a set of real tasks from your own traffic — not benchmark questions, your traffic. Run each through both modes. For every task record three numbers: total tokens billed, wall-clock time to a usable answer, and whether the answer was accepted or had to be redone.

Then compute cost per accepted answer, including redo costs, for each mode.

The result is frequently counter-intuitive in both directions. On easy tasks the thorough mode loses badly: it burns hidden tokens to reach the same answer and adds latency for nothing. On tasks near the edge of the base model’s competence the thorough mode wins by a wide margin, because a failed answer costs you the whole request plus a retry, and the base mode’s failures cluster exactly where the tasks are hardest.

That cluster is the thing to find. Segment your traffic by difficulty — a proxy as crude as input length plus whether the task involves multiple constraints is good enough — and route the hard tail to the thorough mode and the easy bulk to the base. A single global choice wastes money at one end or accuracy at the other, and the split is where the actual saving lives.

What this pair teaches about every release calendar

The same fortnight contains three distinct naming patterns, and no calendar separates them.

A suffix can mean a serving tier of existing weights, sold faster at a premium. It can mean a reasoning mode of existing weights, sold as more thorough at a premium. Or it can mean an actual new generation. All three appear as a row with a name, a date and a price, and the price is the field that most reliably distinguishes them: roughly double the base rate with the same benchmark row is a serving tier; a higher rate with a higher effort setting on the same benchmark row is a reasoning mode; a new benchmark row plus a new context window is a genuine release.

Applying that test here takes a minute and changes what you do. Without it you would budget for two models, provision for two models, and report on two models, when you have one model and one option — and the day someone on the team enables the expensive option, your cost report becomes wrong in a way that looks like a traffic spike.