List prices are gross.
We compute net.
The true delivered cost of every inference request, computed from your fleet’s own telemetry. A number nobody can see today. Not even the operator.
Priced on fiction.
A GPU serving modern inference runs many requests through the same silicon at the same moment. Batches recompose every few milliseconds, and the published per-token price is an average that is wrong for every individual request. The real number lives in the operator’s own telemetry, and today nobody computes it.
Metering

How it works
Nothing touches your serving path.
Shadow mode on one cluster. 90 days. Read-only.
Observe.
A read-only collector ingests telemetry your fleet already emits, DCGM and vLLM metrics. No scheduler hooks, no proxies, nothing in the request path.
Attribute.
The engine allocates real GPU time share to every request across co-scheduled batches, so each request carries its true share of consumed silicon.
Meter.
A dollar figure per request, per customer, per model, reconciled against your actual fleet cost.
Measure.
True delivered cost per request, on the operator’s fleet.
Only the fleet can know.
Application-side tools estimate cost from published prices, so the estimate is the fiction restated. The true number requires the operator’s own telemetry, which never leaves the building.
Measurement becomes the market.
Every commodity market started with a measurement standard. Oil got Brent. Power got the kilowatt-hour. Compute gets its meter.
One fleet sets the standard.
We are selecting a single design partner for a 90 day, read-only shadow deployment. You see the true cost of every request on your own cluster. We calibrate the meter on production workloads.