Skip to content

The Four-Denominator Ledger for Pricing AI Work Correctly

Token price is not total cost. The four-denominator ledger prices AI work by accepted, verified, resolved and sustained outcomes.

5 min readOperations and economics

Research area 43: Cost, latency, energy and capacity-economics engineering

In brief

  • Price a vendor invoices is not the cost an AI system actually incurs.
  • The four-denominator ledger prices work by accepted, verified, resolved and sustained outcomes.
  • Full-system cost has no single reproducible figure, only a range with stated assumptions.

Cost, latency, energy and capacity-economics engineering is the discipline of pricing AI work by the outcome it actually delivers, not by the token bill a vendor invoices. A lower price per call can sit on top of a higher cost per resolved case once retrieval, tool use, verification, human review, retries and idle capacity are counted in. The instrument for keeping that distinction honest is what I call the four-denominator ledger, and every board conversation about AI spend should pass through it before a number reaches a slide.

What it is

The four-denominator ledger tracks the same case through four widening definitions of done: an accepted answer, a verified task, a resolved case and a sustained outcome. Price is what a vendor charges per unit. Incurred cost is what the organization actually pays once energy, hardware time, storage and human effort are added. Business value is the outcome measured against a counterfactual. Treating any one as a stand-in for another is where AI budgets go wrong.

Four principles hold it together. Choose the denominator on purpose: define the population, time horizon, acceptance rule and unit of independence before picking one. Keep failed and abandoned work inside the cohort, since a case that never resolves still consumed a model call and often a reviewer's time, and dropping it manufactures a better number without changing the economics. Separate measured cost from modeled cost from financial assumption, each with its own stated confidence. Evaluate every optimization, caching, batching, a smaller model, learned routing, against that same denominator: a saving that arrives by weakening verification or excluding failed cases is not a saving.

Why it matters now

The 2026 evidence base finally has real measurements to check comfortable defaults against. A January 2026 study, Where Do the Joules Go?, measured accelerator energy across 46 models, seven tasks and 1,858 configurations on H100 and B200 hardware: newer hardware won 63 of 72 matched comparisons, with the rest ranging from 53 percent worse to 82 percent better depending on workload. A per-token energy average cannot stand in for energy measured against a declared boundary and a verified case.

The architecture

The four-denominator ledger moves a case through four stages, each stricter than the one before it, with one rule governing every transition: nothing is allowed to vanish between stages.

The animation opens on the area label and an accepted answer node appearing with a traveling dot. The dot then advances rightward through three more nodes, verified task, resolved case and sustained outcome, drawing each connecting flow as it arrives. A second, failed case starts alongside the first but diverts downward into a retained box rather than disappearing, showing that failed work stays counted. Four small markers light up one at a time beneath the nodes as each stage is reached, showing resources accumulating even as the case count narrows. The diagram fades to an empty stage and repeats.
The four-denominator ledger for pricing AI work by outcome

The accepted answer is the cheapest stage to measure and the easiest to mistake for the whole picture, since it is usually the only number a vendor's dashboard shows by default. The verified task is the output that passed whatever checks the organization requires, a validation rule, a second model, a human reviewer, and that check's cost belongs in the numerator from the moment it runs, not only when it fails. The resolved case is the verified task that held up, with no reopen or escalation after the fact: a capped-retry budget that assumes every attempt fails independently understates true cost, since the cases that keep failing tend to be the ones a system genuinely cannot handle. The sustained outcome is the resolution that held across the horizon that matters to the business, with no downstream cost surfacing after the ledger closed.

The failure mode that recurs at every stage is quietly changing the denominator, reporting cost per accepted answer when the question is cost per resolved case, or comparing a token price against a self-hosted variable cost without adding the fixed cost and utilization risk a hardware change reopens.

How to lead it

Ownership belongs to whoever controls the full resource contract, model spend, retrieval, tools, verification and the capacity plan, not to whichever team holds the invoice for model calls. Fund the ledger itself first: instrumentation that ties model, retrieval, tool, verification, human and incident cost per case to an invoice or a metered counter, not a catalog estimate. Measure cost per verified task and cost per resolved case as the primary pair, plus acceptance, abandonment and tail latency. Stop funding optimizations measured only against successful cases, and self-hosting proposals that omit utilization risk and the next model's validation cost.

Two decisions do not belong to an engineering team alone: where the acceptance standard sits, since cost pressure otherwise quietly lowers the bar, and the make-or-buy call between an API and self-hosted capacity, since that commits fixed cost and staffing regardless of price.

What it is worth

The money shows up in three places: a numerator that keeps growing through retries and rework while the organization measures against a narrower denominator, a capacity decision made on a nominal unit rate instead of a break-even volume, and an energy or carbon figure treated as measured when it was only modeled. Measure before and after the same way, cost per verified and per resolved case on the same replayed workload, with failed cases preserved from the baseline, one optimization at a time.

The 2026 evidence supports measuring energy against a declared boundary rather than a per-token average, reprofiling any predictor on new hardware before relying on it, and counting every retry attempt and its verification cost rather than a first-attempt failure rate. It does not support a single reproducible figure for full-system cost, including human work and embodied carbon, held together with sustained business value. Report a range with stated assumptions, and treat any point estimate offered to a board as a claim to test, not a fact to fund against.

Questions leaders ask

Is cost per token a reliable way to compare AI options?
No. Token price is an input rate, not a cost. It leaves out retrieval, tool use, verification, human review, retries and idle capacity. A lower price per call can sit on top of a higher cost per resolved case once those are added. Compare options on cost per verified task or cost per resolved case, on the same workload and the same acceptance standard, not on the invoice line alone.
Does newer accelerator hardware always use less energy?
No. A January 2026 study measured 46 models across seven tasks and 1,858 configurations on H100 and B200 hardware and found the newer chip won 63 of 72 matched comparisons, with the rest ranging from 53 percent worse to 82 percent better depending on the workload. A per-token energy average cannot substitute for energy measured against a declared boundary and a verified outcome.
Can a latency and energy predictor replace measuring a real deployment?
Not directly. An August 2026 predictor built from about 1,200 profiled configurations reported average errors of 10.98 percent for latency and 11.78 percent for energy on the hardware it was trained on, and needed fresh profiling before it transferred to new hardware. Modeled savings are useful for planning, but they are not measured savings until a deployment confirms them.
When does self-hosting a model beat paying per call?
Only past a break-even volume set by fixed monthly cost divided by the gap between the API rate and the self-hosted variable rate, and only once utilization, staffing and the next model's validation cost are added to the fixed side. The dossier behind this framework works illustrative scenarios rather than vendor quotes, and every one of them shifts once demand, price or model lifetime change.
Should reasoning models that generate many candidate answers be budgeted differently?
Yes. A 2026 study of test-time scaling evaluated 27 models against 18,543 questions and distinguished long single trajectories, terminal aggregation across candidates and feedback-driven search as separate inference regimes, each with its own budget and uncertainty accounting. Treat a system searching live in production as a different cost problem than the same system scored against a fixed bank of pre-generated candidates.

Want this thinking applied to your organization?