Monitor the Fleet
Modelplane runs an OpenTelemetry collector on every inference cluster. It
collects from every component Modelplane installs. This includes the inference
server engine, inference gateway and Envoy proxy, router, and the GPU exporter
your cloud provides. It renames each component’s series to a single
modelplane_* vocabulary and exports them to wherever you say - any backend the
collector has an exporter for, not only OTLP.
Modelplane allows you to write one destination for your metrics. You don’t need to manage per-deployment configurations or update your configuration when a deployment changes. The collector finds pods itself, so a leader/worker split or a prefill/decode pair is collected the same as a single pod.
Telemetry workflow
Every series carries cluster, job, and instance labels of the target
resource. A series about a deployment also carries deployment, replica,
namespace, engine, and role labels.
Each replica publishes its own series, so combine them in your query. Which combination is right follows from what the metric measures:
sum by (deployment), for anything counted, such as requests, tokens, or queue depth.avg by (deployment), for a ratio.max by (deployment), for a saturation figure an alert fires on.
To combine:
sum by (deployment) (rate(modelplane_frontend_request_duration_seconds_count[5m]))The replica is an index rather than a pod, so it’s bounded by the replica count
and survives a restart and a rolling update. Group by replica, not by instance:
# One line per replica, stable across rolling updates
max by (deployment, replica) (modelplane_kv_cache_utilization_ratio)instance is the pod’s address, which keeps two pods of the same replica apart. It
turns over on every rolling update, so group by replica rather than by instance.
Some examples of the available metrics:
| Metric | Means |
|---|---|
modelplane_frontend_ttft_seconds | Time to the first token, measured at the gateway |
modelplane_frontend_tpot_seconds | Time per output token, measured at the gateway |
modelplane_frontend_request_duration_seconds | What the caller waited, end to end |
modelplane_request_queue_seconds | How long a request waited before the engine started |
modelplane_requests_waiting | Queue depth per engine |
modelplane_kv_cache_utilization_ratio | KV-cache occupancy, averaged over replicas |
modelplane_request_input_tokens | Prompt size, as a histogram |
modelplane_request_output_tokens | Generated length, as a histogram |
modelplane_gpu_memory_used_bytes | Framebuffer memory in use, per GPU |
modelplane_energy_joules_total | Energy drawn since the driver last reloaded |
Latency appears twice on purpose. The frontend_ series are what your caller experienced,
measured at the gateway. The engine’s own series are what the engine spent. For
example, if the frontend metric is slow and the engine isn’t, you can
troubleshoot routing, queueing, or networking issues instead of the model.
Saturation gauges are per replica, so how you combine them decides what you see. Average
across a deployment to plan capacity, and take the maximum to alert: three replicas at 0.3
and one at 0.99 average to something comfortable while the fourth evicts and recomputes. A
high maximum beside modelplane_requests_preempted_total climbing is one replica thrashing. To alert on it:
max by (deployment) (modelplane_kv_cache_utilization_ratio) > 0.95Send telemetry to a destination
Create a TelemetryDestination for your OpenTelemetry-compatible endpoint:
apiVersion: modelplane.ai/v1alpha1
kind: TelemetryDestination
metadata:
name: default
spec:
sinks:
- name: primary
type: otlphttp
endpoint: https://otel.example.internaltype names a collector exporter, by the name OpenTelemetry gives it.
To authenticate with a bearer token, store the token in a Secret in
modelplane-system on your control plane. Create it once: Modelplane copies it to
every cluster running a collector, so you don’t put the credential on each GPU
cluster yourself.
kubectl create secret generic telemetry-credentials \
--namespace modelplane-system \
--from-literal=token=<your-token>Reference the Secret from the sink with secretRef, and set auth.bearerTokenKey to
the key that holds the token:
spec:
sinks:
- name: primary
type: otlphttp
endpoint: https://otel.example.internal
secretRef:
name: telemetry-credentials
auth:
bearerTokenKey: tokenModelplane configures the collector to send the token with every export. Rotating the token needs no restart.
If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there:
spec:
sinks:
- name: prometheus
type: prometheus_remote_write
endpoint: https://prom.example.internal/api/v1/writeIf you create more than one sink, all get the entire stream. Each sink carries its own credential, so a vendor and your own Prometheus don’t have to share a Secret.
spec:
sinks:
- name: vendor
type: otlphttp
endpoint: https://otel.vendor.example
secretRef:
name: vendor-token
auth:
bearerTokenKey: token
- name: prometheus
type: prometheus_remote_write
endpoint: https://prom.example.internal/api/v1/writeThat is two copies of the fleet’s metrics, so a vendor charging per sample charges for both.
Sinks can also come from more than one TelemetryDestination. Modelplane concatenates
them, so a team adding an export creates its own object rather than editing one somebody
else owns. Sink names are what the collector calls its exporters, so they have to be
unique across destinations; where two collide, the destination whose name sorts first
keeps it and Modelplane says so on the ServingStack.
Anything else the exporter takes goes under config, passed through as you wrote it:
- name: vendor
type: otlphttp
endpoint: https://otel.vendor.example
config:
compression: gzip
sending_queue:
queue_size: 10000
tls:
ca_file: /etc/ssl/certs/internal.pemModelplane doesn’t define a schema for an exporter’s settings so anything under
the config is passed to the collector exactly as written. TLS, retries,
queueing, compression and headers all work and any new settings in the collector
are respected and the sink keeps working.
To use an authentication scheme Modelplane doesn’t compose, define the extension
yourself under spec.extensions. Then reference it by its key from the sink’s
config.auth.authenticator. The auth block does the same wiring for you when
you use a bearer token.
spec:
sinks:
- name: vendor
type: otlphttp
endpoint: https://otel.vendor.example
secretRef:
name: vendor-oauth
config:
auth:
authenticator: oauth2client/vendor
extensions:
oauth2client/vendor:
client_id: modelplane
client_secret: ${env:CLIENT_SECRET}
token_url: https://issuer.example/oauth2/tokenModelplane doesn’t run any collectors until you create a
TelemetryDestination.
Creating a destination turns collection on everywhere at once, and there’s no per-deployment opt-out.
Your clusters reach the control plane, and only the control plane reaches your backend. A cluster with no route to your observability stack still reports, and the backend’s credential lives in one place instead of on every GPU cluster.
Computing rates, quantiles, and ratios
A collector transforms each measurement as it passes it on. It holds no history, so it produces no rates and no quantiles. Your backend does that. A fleet-wide p99:
histogram_quantile(0.99, sum by (le) (
rate(modelplane_frontend_ttft_seconds_bucket{deployment="qwen3-8b"}[5m])))Modelplane has no dashboards of its own. What it exports is counters and histogram buckets, and your backend derives the rates and quantiles at query time. To precompute them instead, export to Prometheus and write recording rules there.
Engines
Modelplane already knows vLLM’s and SGLang’s metric names and renames them for you, so
neither needs anything from you here. SGLang needs two flags: --enable-metrics to publish /metrics at all, and
--collect-tokens-histogram for the prompt and generation histograms behind
modelplane_request_input_tokens and modelplane_request_output_tokens. Without the
second it publishes those as plain counters and both series stay empty. vLLM needs
nothing.
apiVersion: modelplane.ai/v1alpha1
kind: ModelDeployment
metadata:
name: my-sglang-model
namespace: ml-team
spec:
template:
spec:
engines:
- name: engine
members:
- role: Standalone
template:
spec:
containers:
- name: engine
image: lmsysorg/sglang:v0.5.10.post1-runtime
command:
- /bin/sh
- -c
- >-
exec python3 -m sglang.launch_server
--model-path <model>
--host 0.0.0.0
--port 8000
--enable-metrics
--collect-tokens-histogramSGLang publishes no queue time per request and no preemption counters, so
modelplane_request_queue_seconds and modelplane_requests_preempted_total carry vLLM
only.
Any other OpenAI-compatible engine reports its frontend numbers with no configuration.
The gateway measures those, not the engine, so modelplane_frontend_* works for an
engine Modelplane has never seen.
To normalize that engine’s own metrics as well, create a MetricMapping:
apiVersion: modelplane.ai/v1alpha1
kind: MetricMapping
metadata:
name: my-engine
spec:
metrics:
- from: my_engine_queued_requests
to: modelplane_requests_waiting
- from: my_engine_kv_transfer_ms
fromUnit: Milliseconds
to: modelplane_request_kv_transfer_secondsModelplane renders every mapping into every cluster’s collector, so you write one once.
from is the name your engine emits and to is what Modelplane calls it.
Modelplane leaves the combining to your backend. A scrape of one replica is one batch, so a collector that added them up would be summing readings taken at different moments, and two readings of one cumulative counter come to twice the traffic that happened. Your backend holds every replica’s series and combines them at query time.
Say fromUnit whenever the engine measures in something other than the unit the name
claims, and Modelplane converts to the base one. Skipping this is the expensive mistake here:
a series named _seconds that holds milliseconds reads a thousand times fast, and nothing
downstream can tell.
Rename only where the measurements agree. Two engines’ histograms under one name are worth less than nothing if their buckets disagree, because a quantile over them is wrong rather than approximate.
Examples
apiVersion: modelplane.ai/v1alpha1
kind: MetricMapping
metadata:
name: my-engine
spec:
metrics:
# A plain rename.
- from: my_engine_queued_requests
to: modelplane_requests_waiting
# A unit conversion. The engine reports milliseconds; the name says seconds.
- from: my_engine_kv_transfer_ms
to: modelplane_request_kv_transfer_seconds
fromUnit: Milliseconds
# A request count taken out of a duration histogram. The histogram
# keeps its own name; this adds a counter beside it.
- from: my_engine_request_duration_seconds
part: Count
to: modelplane_requests_total
# Two counters folded into one name, told apart by a fixed label.
- from: my_engine_prompt_tokens_total
to: modelplane_tokens_total
labels:
- name: direction
value: input
- from: my_engine_generated_tokens_total
to: modelplane_tokens_total
labels:
- name: direction
value: output
# A label the engine already emits, renamed and its values translated.
- from: my_engine_finished_requests_total
to: modelplane_responses_total
labels:
- name: reason
from: finish_reason
values:
eos: stop
max_tokens: lengthWhy engine latency and gateway latency differ
modelplane_request_ttft_seconds comes from the engine, and engines bucket their
histograms differently. vLLM resolves down to a millisecond. SGLang resolves to a hundred
of them. A quantile across both is wrong, not approximate. Use the engine series to compare
one engine against itself, and the frontend_ series for anything fleet-wide.
Some measurements don’t translate at all. SGLang’s inter-token latency isn’t vLLM’s time per output token, so neither is renamed onto a shared name. The gateway measures time per output token for both.