OpenAI proposes useful intelligence per dollar as a new AI metric
The story in five sentences
- OpenAI proposes Useful Intelligence per Dollar as an economic benchmark for AI.
- The key measure should be work actually completed rather than active users or consumed tokens.
- The scorecard examines value, cost per successful task, reliability, and scaling effects.
- Companies should measure outcomes inside each workflow and record corrections and escalations.
- Performance figures for OpenAI models cited in the proposal come from the vendor and are not independently verified.
Full explanation
Sarah Friar argues that conventional measures such as active users or consumed tokens do not show whether AI creates useful results. The proposed scorecard therefore focuses on relevant work completed above a defined quality threshold. It combines the value of that work with the full cost of a successful task, the reliability of the result, and the effects achieved when the system is scaled.
The cost calculation includes more than model usage. Employee time, reviews, retries, corrections, and follow-up work are counted before the total is divided by tasks that met the required quality level. Reliability is separated into outcomes that can be used immediately, outcomes that need correction, and cases that must be escalated to a person or another process.
OpenAI says companies should collect these measures inside the workflows where AI is used. Greater automation also depends on clear data-access boundaries, systems the AI may change, and explicit approvals. GPT-5.6 and ChatGPT Work are examples, together with performance figures from OpenAI. Those figures are vendor benchmarks rather than independent confirmation.