The true cost of every request you serve.
Metron computes it from the metrics your fleet already produces. We only read. Your traffic never touches us.
Priced on fiction.
A GPU serving modern inference runs many requests through the same silicon at the same moment. Batches recompose every few milliseconds, and the published per-token price is an average that is wrong for every individual request. The real number lives in the operator’s own telemetry.
Metering

How it works
Nothing touches your serving path.
Shadow mode on one cluster. 90 days. Read-only.
Observe.
A read-only collector ingests telemetry your fleet already emits, DCGM and vLLM, SGLang metrics. No scheduler hooks, no proxies, nothing in the request path.
Attribute.
The engine allocates real GPU time share to every request across co-scheduled batches, so each request carries its true share of consumed silicon.
Meter.
A dollar figure per request, per customer, per model, reconciled against your actual fleet cost.
Measure.
True delivered cost per request, on the operator’s fleet.
What leaves your building.
Aggregate telemetry only: GPU utilization, memory, power draw, batch composition, request timing. No prompts, no completions, no customer content. The collector runs inside your perimeter, you control what it reads, and you can revoke it at any time.
Only the fleet can know.
Application-side tools estimate cost from published prices, so the estimate is the fiction restated. The true number requires the operator’s own telemetry, which never leaves the building.
Measurement becomes the market.
Every commodity market started with a measurement standard. Oil got Brent. Power got the kilowatt-hour. Compute gets its meter.
One fleet sets the standard.
We are selecting a single design partner for a 90 day, read-only shadow deployment. You see the true cost of every request on your own cluster. We calibrate the meter on production workloads.