CFOtech Ireland - Technology news for CFOs & financial decision-makers
Ireland
OpenAI sets AI value scorecard for corporate spend

OpenAI sets AI value scorecard for corporate spend

Mon, 20th Jul 2026 (Yesterday)
Mark Tarre
MARK TARRE News Chief

OpenAI has outlined a framework for measuring the return on corporate spending on artificial intelligence, centred on what it calls "Useful Intelligence per Dollar".

The framework is aimed at finance leaders and other executives trying to judge whether AI systems are creating enough value to justify their cost. It argues that common software metrics such as licences, active users and adoption rates do not clearly show how much useful work AI is completing inside an organisation.

Instead, companies should assess AI using four measures: whether the system completes work that matters, what each successful task costs, how dependable the result is, and whether the economics improve as usage grows.

Sarah Friar, Chief Financial Officer at OpenAI, framed the issue around the concerns she hears from finance chiefs weighing AI budgets against business outcomes. The central question, she said, is whether the value of work completed by AI rises faster than the cost of producing it.

The argument shifts attention away from simple technical pricing benchmarks such as cost per token. A cheaper model may require more attempts, more review or more employee time before it produces an acceptable answer, while a more expensive model may complete the same task in one pass.

That distinction matters most in business workflows where AI is used for more than drafting text. OpenAI pointed to tasks such as resolving customer support issues, shipping code changes, reviewing contracts and helping finance teams prepare for forecast reviews.

Under the model, organisations are encouraged to define what "done" means for a specific workflow, then measure outcomes in the system where the work happens. For a support team, that could mean a customer issue resolved; for an engineering team, a code change that passes tests.

Cost of work

A major part of the framework is measuring the full cost of completing a successful task. Businesses should add the total cost of doing the work, count the tasks that met the required quality threshold, and divide the cost by the number of successful tasks.

That calculation includes not only model pricing and compute use, but also employee time, human review, retries and rework. The cheapest model by token price, OpenAI argues, will not always deliver the lowest cost per useful outcome.

OpenAI also used the framework to describe how it sees product positioning across its latest model range. GPT-5.6, released last week, includes three tiers - Sol, Terra and Luna - intended for different balances of reasoning depth, speed and cost.

A customer might choose Luna for a high-volume workflow, Terra for tasks that need more depth, and Sol when stronger reasoning cuts the number of attempts required. In OpenAI's view, the economics of the whole task, rather than a model's sticker price, should determine which version a customer uses.

To support its case, OpenAI cited results from the Artificial Analysis Coding Agent Index, saying GPT-5.6 Sol with maximum reasoning set a new benchmark while using 54% fewer output tokens than another leading model. It presented the claim as evidence that token efficiency and useful work completed can improve at the same time.

Dependability focus

The third measure in the scorecard is dependability, which OpenAI described as central to whether companies expand their use of AI into more important workflows. It said adoption often starts with drafting tasks before moving into activities that involve finding context, reasoning across tools and data, and eventually taking action.

OpenAI argued that dependability has direct economic value because accurate and consistent results reduce the time people spend reviewing, correcting and repeating work. That, in turn, lowers the cost per successful task.

Teams should track whether an AI output is ready to use, needs correction or needs escalation to a person. OpenAI said those categories are more useful than model accuracy alone because they show whether AI is reducing the human effort needed to complete a project.

It also said companies need clear limits around data access, system permissions and the points at which a person must review or approve an action. Those controls, OpenAI argued, are necessary before AI moves from drafting to acting inside business systems.

In that context, the company linked ChatGPT Work to the security, privacy, compliance and workspace management tools already associated with ChatGPT Enterprise. It said those features allow organisations to give AI access to more context and workflows while keeping oversight in place.

Scale economics

The final part of the framework examines whether AI economics improve at scale. Companies should track the same workflow over time, compare total cost with the number of successful tasks, and assess whether completed work rises faster than spending while quality is maintained or improved.

Compute sits at the centre of that equation, OpenAI said, because it affects research, model quality, speed, dependability, availability and cost. Training compute shapes future models, while inference compute determines the useful work customers receive today.

OpenAI argued that improvements in models, inference efficiency, hardware, routing and product design should all raise the return on compute over time. It said the benefits should appear in practical terms such as better answers, faster results, fewer corrections and lower costs for completed work.

Friar said: "The question I hear from CFOs everywhere is simple: how do we get more value from our AI spend?"

She added: "Our job is to make that equation better with every generation: more capable models, faster and more dependable results, and lower costs for the work customers need done."