OpenAI’s GPT 6.1 Sol announcement places a new model between its highest-capability offering and its lighter everyday option. The September 29 release focuses on coding, computer use, and longer tasks that require a model to keep working across several steps. Its most immediately legible change is price: standard input and output tokens cost a fifth of those for GPT 6 Astra.

That gives developers a useful new candidate for work that repeats often. It also makes the distinction between a cheap request and a cheap completed task more consequential. A workflow can spend little on one model call and still become expensive through repeated attempts, unnecessary tool calls, or a result that needs substantial repair.

Availability comes before comparison

OpenAI says GPT 6.1 Sol is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Its model reference identifies the API model as gpt-6.1-sol. Work and Codex are the announced ChatGPT surfaces, which matters when someone looks for the model in a different part of the product.

The launch also describes an Ultrafast option planned for the following days. That future delivery should be kept separate from the model available at launch. A promised speed mode is not a measurement of the standard mode a team can use today.

A large language model generates a sequence of tokens in response to its context. In an agent workflow, some of that output can request a tool action. The tool returns information, the context grows, and another model call follows. The bill therefore belongs to the sequence of calls that produced an accepted result, rather than to the first answer alone.

Read the three prices separately

In the standard model pricing , GPT 6.1 Sol costs $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. Astra’s corresponding rates in OpenAI’s comparison are $10, $1, and $50. Input and output are one fifth of Astra’s rates. Cached input is one tenth, so a single fraction does not describe every kind of usage.

Consider an illustrative task with 100,000 uncached input tokens and 10,000 output tokens. Multiplying those volumes by the listed rates gives $0.30 for Sol and $1.50 for Astra. This is arithmetic using a hypothetical workload, not a benchmark or an estimate of what every job will consume. It excludes tools, storage, and other charges. Cache writes, larger-context requests, and different processing modes have separate pricing conditions in the model reference.

If the same 100,000 input tokens were all eligible cached input, the token subtotal would instead be $0.11 for Sol and $0.60 for Astra, with the same assumed output. The difference illustrates why the shape of a workload matters. A team with a large repeated instruction prefix may see a different price relationship from one sending short, varied prompts.

Caching rewards stable context

OpenAI’s prompt-caching guide explains that reuse depends on an exact matching prefix. Put stable instructions and examples before changing task details when that ordering also makes the task clear. Similar meaning is insufficient if the tokens at the beginning differ.

Caching also has eligibility and retention conditions. A repeated paragraph is not automatically a promise that every token will receive the cached rate. The operational check is the usage record, which reports how much input was actually cached. Expected reuse and observed reuse should be kept as separate quantities when estimating expenditure.

There is a second benefit to a stable prefix: it gives an evaluation a more consistent starting point. If the instructions change between runs, a difference in results can come from the prompt as well as the model. Keeping the common task definition fixed makes it easier to inspect the variable that a comparison is supposed to measure.

Benchmarks describe their own tasks

OpenAI reports gains over the previous Sol model on coding and computer-use evaluations. Those results are evidence about the vendor’s specified tasks and settings. They do not establish identical reliability across every repository, document format, or desktop workflow.

For a team considering migration, the practical evaluation starts with a small set of real tasks and explicit acceptance criteria. A code change might need to satisfy the existing tests and preserve an interface. A document task might need to reproduce figures correctly and place each conclusion next to its supporting evidence. A browser task might need a verified final state rather than a convincing account of an attempted action.

These definitions make failures visible. Without them, fluent output can look complete before anyone checks the work that matters. The model price is easy to read from a table. Whether the output meets the task’s standard requires a separate observation.

Measure cost at the point of acceptance

A useful comparison records all model calls for one task, then records whether the result was accepted, repaired, or abandoned. Keep tool expenditure and human review beside token usage. That allows a team to distinguish a model that spends less on equivalent work from one that simply produces a cheaper first attempt.

The arithmetic does not demand that every task use one model. A lower-cost model may handle work with clear, readily checked outcomes, while a more capable option remains useful for tasks where failed attempts or lengthy repair dominate the total. The choice can be made at the task boundary instead of becoming a blanket statement about one model’s intelligence.

GPT 6.1 Sol supplies a lower standard token price and another model to evaluate against actual work. Its value becomes clearest when the measurement follows the whole task: stable context where useful, visible tool actions, a defined acceptance test, and a cost record that ends only when the work is finished.