OpenAI puts the completed task at the center
In The Work Now Within Reach , published September 8, OpenAI connects its model development, product distribution and computing infrastructure to a business argument: more capable systems and more efficient delivery can make additional work economically practical. The essay is by Sarah Friar , whom OpenAI appointed chief financial officer in 2024.
It is a statement of the company's strategy, not independent proof that a customer's costs have fallen. Its most useful analytical opening is the completed task. A model request can be cheap while the overall job remains expensive, particularly when someone must repeat the request, inspect the result or repair what it produced.
Newsroom's contribution here is a practical way to evaluate that distinction. The examples below are hypothetical accounting exercises. They do not describe measured results from OpenAI products or claim that one provider is more cost-effective than another.
Define completion before comparing the bill
Suppose a team asks an AI system to prepare a short document from an approved set of sources. One possible definition of completion is simply that a file exists. A more useful definition might require accurate citations, the requested structure and acceptance by the person responsible for publication.
Those definitions produce different success counts. If ten files are generated but only six meet the agreed criteria, the denominator should not quietly become ten completed jobs. The rejected work consumed resources too. Leaving it outside the calculation would make the system appear cheaper without improving its actual usefulness.
For this hypothetical exercise, start with a fixed acceptance rule and a representative set of tasks. Record the attempts, the outputs accepted, the human review time and the repairs required. The method needs no elaborate dashboard to begin. It needs a clear definition that stays stable while alternatives are compared.
The broader idea of total cost of ownership is useful background: the initial or most visible charge is only one part of an overall cost. Applied here, the specific categories should match the work being assessed. A casual draft and a document that supports a consequential decision should not be evaluated as if they have identical review requirements.
Three cost boundaries should stay separate
There are at least three distinct questions in the discussion. What does it cost the provider to serve an interaction? What does the customer pay for access and usage? What does it cost the organization to obtain an acceptable result?
A change at one boundary does not automatically establish a change at the others. A provider's internal efficiency improvement could support lower prices, additional capacity, different service levels or other business choices. A customer's actual outcome depends on the product terms and the work performed. The September essay does not turn those possible paths into a universal customer saving.
The organizational boundary includes effort that may sit outside the AI invoice. In our document example, this could include preparing source material, explaining requirements, checking citations and correcting the final output. Some of that effort may already exist in the previous workflow. A fair comparison should account for both paths consistently.
Time also needs a clear definition. A task that runs unattended for an hour may consume less staff attention than a ten-minute interaction requiring continuous supervision. Conversely, an unattended run that blocks a deadline may be costly despite requiring little direct attention. Record elapsed time and human effort separately before deciding which matters most.
Hardware measurements answer a narrower question
OpenAI's essay points to its earlier Jalapeño inference-chip results . That August 25 report describes tests using InferenceX across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. Its comparisons normalize performance using published accelerator power ratings. OpenAI presents the results as first-party testing and says deployment is planned to begin by year-end.
That evidence concerns the tested systems and workloads. It should not be recast as a measurement of every customer's completed task, a guaranteed change to subscription prices or a demonstrated reduction in an organization's electricity consumption. Those claims would require different evidence and boundaries.
For the reader evaluating an agent workflow, the relevant bridge is straightforward to state but must still be measured: does an infrastructure change improve the experience of reaching an accepted result? Faster responses may matter, but the evaluation also needs to keep task quality, retries and review effort visible.
This is why a single throughput figure cannot finish the business comparison. It can inform one part of the explanation while leaving the practical acceptance test unresolved. The same discipline applies when the measurement comes from a competitor or an independent benchmark rather than the vendor itself.
A small comparison can reveal the trade-off
Return to the hypothetical document task. Imagine two approaches with the same acceptance criteria. One has a lower usage charge but requires repeated corrections. The other costs more per attempt but produces more acceptable drafts. Neither is automatically the better choice without knowing the size and frequency of those differences.
The comparison should also retain failures rather than silently removing them. A workflow that performs well on straightforward examples but regularly stalls on a common exception may need a separate fallback. The time spent recognizing and routing those exceptions belongs in the operational picture.
Our analysis of Nvidia's Personal AI Router discusses a related distinction between routing requests and the infrastructure that actually serves a model. Here, the corresponding lesson is to avoid letting one improved component stand in for the performance of the complete workflow.
OpenAI's new essay provides a clear corporate thesis about making more work viable. For a team deciding what to use, the next step is to test a narrower proposition: whether its own defined tasks reach an acceptable result with less total effort and an understood cost. That is a question a repeatable comparison can answer. The strategy statement alone cannot.
