Modelplane Modelplane docs

Monitor the Fleet

Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from every component Modelplane installs. This includes the inference server engine, inference gateway and Envoy proxy, router, and the GPU exporter your cloud provides. It renames each component’s series to a single modelplane_* vocabulary and exports them to wherever you say - any backend the collector has an exporter for, not only OTLP.

Modelplane allows you to write one destination for your metrics. You don’t need to manage per-deployment configurations or update your configuration when a deployment changes. The collector finds pods itself, so a leader/worker split or a prefill/decode pair is collected the same as a single pod.

Telemetry workflow

Every series carries cluster, job, and instance labels of the target resource. A series about a deployment also carries deployment, replica, namespace, engine, and role labels.

Each replica publishes its own series, so combine them in your query. Which combination is right follows from what the metric measures:

  • sum by (deployment), for anything counted, such as requests, tokens, or queue depth.
  • avg by (deployment), for a ratio.
  • max by (deployment), for a saturation figure an alert fires on.

To combine:

promql
sum by (deployment) (rate(modelplane_frontend_request_duration_seconds_count[5m]))

The replica is an index rather than a pod, so it’s bounded by the replica count and survives a restart and a rolling update. Group by replica, not by instance:

promql
# One line per replica, stable across rolling updates
max by (deployment, replica) (modelplane_kv_cache_utilization_ratio)

instance is the pod’s address, which keeps two pods of the same replica apart. It turns over on every rolling update, so group by replica rather than by instance.

Some examples of the available metrics:

MetricMeans
modelplane_frontend_ttft_secondsTime to the first token, measured at the gateway
modelplane_frontend_tpot_secondsTime per output token, measured at the gateway
modelplane_frontend_request_duration_secondsWhat the caller waited, end to end
modelplane_request_queue_secondsHow long a request waited before the engine started
modelplane_requests_waitingQueue depth per engine
modelplane_kv_cache_utilization_ratioKV-cache occupancy, averaged over replicas
modelplane_request_input_tokensPrompt size, as a histogram
modelplane_request_output_tokensGenerated length, as a histogram
modelplane_gpu_memory_used_bytesFramebuffer memory in use, per GPU
modelplane_energy_joules_totalEnergy drawn since the driver last reloaded

Latency appears twice on purpose. The frontend_ series are what your caller experienced, measured at the gateway. The engine’s own series are what the engine spent. For example, if the frontend metric is slow and the engine isn’t, you can troubleshoot routing, queueing, or networking issues instead of the model.

Saturation gauges are per replica, so how you combine them decides what you see. Average across a deployment to plan capacity, and take the maximum to alert: three replicas at 0.3 and one at 0.99 average to something comfortable while the fourth evicts and recomputes. A high maximum beside modelplane_requests_preempted_total climbing is one replica thrashing. To alert on it:

promql
max by (deployment) (modelplane_kv_cache_utilization_ratio) > 0.95

Send telemetry to a destination

Create a TelemetryDestination for your OpenTelemetry-compatible endpoint:

yaml
apiVersion: modelplane.ai/v1alpha1
kind: TelemetryDestination
metadata:
  name: default
spec:
  sinks:
  - name: primary
    type: otlphttp
    endpoint: https://otel.example.internal

type names a collector exporter, by the name OpenTelemetry gives it.

To authenticate with a bearer token, store the token in a Secret in modelplane-system on your control plane. Create it once: Modelplane copies it to every cluster running a collector, so you don’t put the credential on each GPU cluster yourself.

shell
kubectl create secret generic telemetry-credentials \
  --namespace modelplane-system \
  --from-literal=token=<your-token>

Reference the Secret from the sink with secretRef, and set auth.bearerTokenKey to the key that holds the token:

yaml
spec:
  sinks:
  - name: primary
    type: otlphttp
    endpoint: https://otel.example.internal
    secretRef:
      name: telemetry-credentials
    auth:
      bearerTokenKey: token

Modelplane configures the collector to send the token with every export. Rotating the token needs no restart.

If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there:

yaml
spec:
  sinks:
  - name: prometheus
    type: prometheus_remote_write
    endpoint: https://prom.example.internal/api/v1/write

If you create more than one sink, all get the entire stream. Each sink carries its own credential, so a vendor and your own Prometheus don’t have to share a Secret.

yaml
spec:
  sinks:
  - name: vendor
    type: otlphttp
    endpoint: https://otel.vendor.example
    secretRef:
      name: vendor-token
    auth:
      bearerTokenKey: token
  - name: prometheus
    type: prometheus_remote_write
    endpoint: https://prom.example.internal/api/v1/write

That is two copies of the fleet’s metrics, so a vendor charging per sample charges for both.

Sinks can also come from more than one TelemetryDestination. Modelplane concatenates them, so a team adding an export creates its own object rather than editing one somebody else owns. Sink names are what the collector calls its exporters, so they have to be unique across destinations; where two collide, the destination whose name sorts first keeps it and Modelplane says so on the ServingStack.

Anything else the exporter takes goes under config, passed through as you wrote it:

yaml
  - name: vendor
    type: otlphttp
    endpoint: https://otel.vendor.example
    config:
      compression: gzip
      sending_queue:
        queue_size: 10000
      tls:
        ca_file: /etc/ssl/certs/internal.pem

Modelplane doesn’t define a schema for an exporter’s settings so anything under the config is passed to the collector exactly as written. TLS, retries, queueing, compression and headers all work and any new settings in the collector are respected and the sink keeps working.

To use an authentication scheme Modelplane doesn’t compose, define the extension yourself under spec.extensions. Then reference it by its key from the sink’s config.auth.authenticator. The auth block does the same wiring for you when you use a bearer token.

yaml
spec:
  sinks:
  - name: vendor
    type: otlphttp
    endpoint: https://otel.vendor.example
    secretRef:
      name: vendor-oauth
    config:
      auth:
        authenticator: oauth2client/vendor
  extensions:
    oauth2client/vendor:
      client_id: modelplane
      client_secret: ${env:CLIENT_SECRET}
      token_url: https://issuer.example/oauth2/token

Modelplane doesn’t run any collectors until you create a TelemetryDestination.

Creating a destination turns collection on everywhere at once, and there’s no per-deployment opt-out.

Your clusters reach the control plane, and only the control plane reaches your backend. A cluster with no route to your observability stack still reports, and the backend’s credential lives in one place instead of on every GPU cluster.

Computing rates, quantiles, and ratios

A collector transforms each measurement as it passes it on. It holds no history, so it produces no rates and no quantiles. Your backend does that. A fleet-wide p99:

promql
histogram_quantile(0.99, sum by (le) (
  rate(modelplane_frontend_ttft_seconds_bucket{deployment="qwen3-8b"}[5m])))

Modelplane has no dashboards of its own. What it exports is counters and histogram buckets, and your backend derives the rates and quantiles at query time. To precompute them instead, export to Prometheus and write recording rules there.

Engines

Modelplane already knows vLLM’s and SGLang’s metric names and renames them for you, so neither needs anything from you here. SGLang needs two flags: --enable-metrics to publish /metrics at all, and --collect-tokens-histogram for the prompt and generation histograms behind modelplane_request_input_tokens and modelplane_request_output_tokens. Without the second it publishes those as plain counters and both series stay empty. vLLM needs nothing.

yaml
apiVersion: modelplane.ai/v1alpha1
kind: ModelDeployment
metadata:
  name: my-sglang-model
  namespace: ml-team
spec:
  template:
    spec:
      engines:
      - name: engine
        members:
        - role: Standalone
          template:
            spec:
              containers:
              - name: engine
                image: lmsysorg/sglang:v0.5.10.post1-runtime
                command:
                - /bin/sh
                - -c
                - >-
                  exec python3 -m sglang.launch_server
                  --model-path <model>
                  --host 0.0.0.0
                  --port 8000
                  --enable-metrics
                  --collect-tokens-histogram

SGLang publishes no queue time per request and no preemption counters, so modelplane_request_queue_seconds and modelplane_requests_preempted_total carry vLLM only.

Any other OpenAI-compatible engine reports its frontend numbers with no configuration. The gateway measures those, not the engine, so modelplane_frontend_* works for an engine Modelplane has never seen.

To normalize that engine’s own metrics as well, create a MetricMapping:

yaml
apiVersion: modelplane.ai/v1alpha1
kind: MetricMapping
metadata:
  name: my-engine
spec:
  metrics:
  - from: my_engine_queued_requests
    to: modelplane_requests_waiting
  - from: my_engine_kv_transfer_ms
    fromUnit: Milliseconds
    to: modelplane_request_kv_transfer_seconds

Modelplane renders every mapping into every cluster’s collector, so you write one once. from is the name your engine emits and to is what Modelplane calls it.

Modelplane leaves the combining to your backend. A scrape of one replica is one batch, so a collector that added them up would be summing readings taken at different moments, and two readings of one cumulative counter come to twice the traffic that happened. Your backend holds every replica’s series and combines them at query time.

Say fromUnit whenever the engine measures in something other than the unit the name claims, and Modelplane converts to the base one. Skipping this is the expensive mistake here: a series named _seconds that holds milliseconds reads a thousand times fast, and nothing downstream can tell.

Rename only where the measurements agree. Two engines’ histograms under one name are worth less than nothing if their buckets disagree, because a quantile over them is wrong rather than approximate.

Examples

yaml
apiVersion: modelplane.ai/v1alpha1
kind: MetricMapping
metadata:
  name: my-engine
spec:
  metrics:
  # A plain rename.
  - from: my_engine_queued_requests
    to: modelplane_requests_waiting

  # A unit conversion. The engine reports milliseconds; the name says seconds.
  - from: my_engine_kv_transfer_ms
    to: modelplane_request_kv_transfer_seconds
    fromUnit: Milliseconds

  # A request count taken out of a duration histogram. The histogram
  # keeps its own name; this adds a counter beside it.
  - from: my_engine_request_duration_seconds
    part: Count
    to: modelplane_requests_total

  # Two counters folded into one name, told apart by a fixed label.
  - from: my_engine_prompt_tokens_total
    to: modelplane_tokens_total
    labels:
    - name: direction
      value: input
  - from: my_engine_generated_tokens_total
    to: modelplane_tokens_total
    labels:
    - name: direction
      value: output

  # A label the engine already emits, renamed and its values translated.
  - from: my_engine_finished_requests_total
    to: modelplane_responses_total
    labels:
    - name: reason
      from: finish_reason
      values:
        eos: stop
        max_tokens: length

Why engine latency and gateway latency differ

modelplane_request_ttft_seconds comes from the engine, and engines bucket their histograms differently. vLLM resolves down to a millisecond. SGLang resolves to a hundred of them. A quantile across both is wrong, not approximate. Use the engine series to compare one engine against itself, and the frontend_ series for anything fleet-wide.

Some measurements don’t translate at all. SGLang’s inter-token latency isn’t vLLM’s time per output token, so neither is renamed onto a shared name. The gateway measures time per output token for both.