The true cost of every request you serve.

Metron computes it from the metrics your fleet already produces. We only read. Your traffic never touches us.

Priced on fiction.

A GPU serving modern inference runs many requests through the same silicon at the same moment. Batches recompose every few milliseconds, and the published per-token price is an average that is wrong for every individual request. The real number lives in the operator’s own telemetry.

Metering

Human Hand Reaching For Right Side

How it works

Nothing touches your serving path.

Shadow mode on one cluster. 90 days. Read-only.

Observe.

A read-only collector ingests telemetry your fleet already emits, DCGM and vLLM, SGLang metrics. No scheduler hooks, no proxies, nothing in the request path.

Attribute.

The engine allocates real GPU time share to every request across co-scheduled batches, so each request carries its true share of consumed silicon.

Meter.

A dollar figure per request, per customer, per model, reconciled against your actual fleet cost.

Measure.

True delivered cost per request, on the operator’s fleet.

What leaves your building.

Aggregate telemetry only: GPU utilization, memory, power draw, batch composition, request timing. No prompts, no completions, no customer content. The collector runs inside your perimeter, you control what it reads, and you can revoke it at any time.

Metron

Customer meter

Request cost

GPU time share

Operator telemetry

GPU share

Telemetry

Only the fleet can know.

Application-side tools estimate cost from published prices, so the estimate is the fiction restated. The true number requires the operator’s own telemetry, which never leaves the building.

Measurement becomes the market.

Every commodity market started with a measurement standard. Oil got Brent. Power got the kilowatt-hour. Compute gets its meter.

Measure

Measure

True delivered cost

True delivered cost per request, on the operator’s fleet. The product today.

Operator fleet telemetry

Per-request delivered cost

Production workloads

Product today

Measure

True delivered cost

True delivered cost per request, on the operator’s fleet. The product today.

Operator fleet telemetry

Per-request delivered cost

Production workloads

Product today

Optimize

Waste recovery

The meter exposes waste. Recovering it is the second product.

Expose utilization waste

Recover margin

Improve routing decisions

Second product

Optimize

Waste recovery

The meter exposes waste. Recovering it is the second product.

Expose utilization waste

Recover margin

Improve routing decisions

Second product

Optimize

Index

Index

Reference price

Calibrated across fleets, the meter becomes a reference price for delivered inference. Horizon.

Calibrated across fleets

Reference price

Delivered inference

Horizon

Index

Reference price

Calibrated across fleets, the meter becomes a reference price for delivered inference. Horizon.

Calibrated across fleets

Reference price

Delivered inference

Horizon

One fleet sets the standard.

We are selecting a single design partner for a 90 day, read-only shadow deployment. You see the true cost of every request on your own cluster. We calibrate the meter on production workloads.

Metron

The metering layer for AI compute.

Metron

The metering layer for AI compute.

Metron

The metering layer for AI compute.