The llm-d-router exposes the following Prometheus metrics to monitor its behavior and performance, particularly concerning Encode/Prefill/Decode disaggregation.
All metrics are in the llm_d_inference_scheduler subsystem.
Metrics defined by llm-d Router are in addition to Inference Gateway metrics. For more details of seeing metrics, see the metrics and observability section.
- Type: Counter
- Labels:
model_name: string (the target model name, or "unknown" if empty)decision_type: string - one of:decode-only- the request used the decode-only path (no disaggregation)prefill-decode- the request was split into prefill and decode stages (P/D or EP/D)encode-decode- the request used encode disaggregation with local prefill+decode (E/PD)encode-prefill-decode- the request used the full three-stage pipeline (E/P/D)
- Release Stage: ALPHA
- Description: Counts the number of requests processed, broken down by the disaggregation routing decision.
- Usage: Provides a high-level view of how many requests are utilizing each disaggregation topology.
- Actionability:
- Monitor the distribution across decision types to understand engagement rates for each disaggregation mode.
- Sudden changes in ratios might indicate configuration issues, changes in workload patterns, or problems with the decision logic.
Deprecated: Use
disagg_decision_totalinstead.
- Type: Counter
- Labels:
model_name: string (the target model name, or "unknown" if empty)decision_type: string ("decode-only" or "prefill-decode")
- Release Stage: ALPHA
- Description: Counts the number of requests processed, broken down by the Prefill/Decode disaggregation decision. This metric only covers P/D disaggregation and does not account for encode disaggregation.
Note
This metric is maintained for backward compatibility with the deprecated
pd-profile-handler. New deployments should use disagg_decision_total.
Exposed when the flowControl feature gate is enabled. All carry the llm_d_epp_ prefix.
- Type: Histogram
- Labels:
fairness_id: string (the tenant or flow identifier for fairness rotation)priority: string (the priority band, e.g., "0", "10")outcome: string (Dispatched,RejectedCapacity,RejectedOther,EvictedTTL,EvictedContextCancelled,EvictedOther)inference_pool: stringmodel_name: stringtarget_model_name: string
- Release Stage: ALPHA
- Description: Total time a request spends in the Flow Control layer, from enqueue to final outcome.
- Usage: Primary latency signal for flow control. Rising p99 indicates backends are saturated or capacity limits are too tight.
- Type: Histogram
- Release Stage: ALPHA
- Description: Time taken for each internal dispatch cycle.
- Usage: Measures the overhead of the dispatch loop itself. Rising values indicate increasing cost per cycle from saturation detection, priority band iteration, or fairness evaluation.
- Type: Histogram
- Labels:
fairness_id: string (the tenant or flow identifier)priority: string (the priority band)outcome: string
- Release Stage: ALPHA
- Description: Time taken to enqueue a request into the Flow Control layer.
- Usage: Measures the time spent in capacity checks and queue insertion within the processor.
- Type: Gauge
- Labels:
fairness_id: string (the tenant or flow identifier)priority: string (the priority band)inference_pool: stringmodel_name: stringtarget_model_name: string
- Release Stage: ALPHA
- Description: Current number of requests actively held in the Flow Control queue.
- Usage: Tracks queue depth per priority band and tenant. A steadily growing value indicates the dispatch rate is lower than the arrival rate.
- Type: Gauge
- Labels:
fairness_id: string (the tenant or flow identifier)priority: string (the priority band)inference_pool: stringmodel_name: stringtarget_model_name: string
- Release Stage: ALPHA
- Description: Current total size in bytes of requests actively held in the Flow Control queue.
- Usage: Tracks memory pressure from queued requests. Compare against the configured
maxBytescapacity to gauge how close a band is to rejecting new requests.
- Type: Gauge
- Labels:
inference_pool: string
- Release Stage: ALPHA
- Description: Current saturation level of the inference pool (0.0 = empty, 1.0 = fully saturated).
- Usage: When saturation reaches the usage limit threshold, the dispatch cycle skips dispatching and requests remain queued. Sustained 1.0 indicates all backends are at capacity.
- Type: Counter
- Labels:
outcome: string — the terminal outcome of the request. One of:Dispatched— request was forwarded to a backendRejectedCapacity— request was rejected because the queue was at capacityRejectedNoEndpoints— request was rejected at the capacity boundary while the candidate pool had no endpoints (surfaces as HTTP 503 rather than 429)RejectedOther— request was rejected for another reason (e.g., controller shutdown)EvictedTTL— request exceeded its time-to-live while waiting in the queueEvictedContextCancelled— client disconnected before the request was dispatchedEvictedOther— request was evicted for another reason
priority: string (the priority band, e.g.,"0","10")inference_pool: string
- Release Stage: ALPHA
- Description: Total number of requests processed by the Flow Control layer, incremented once per request after its terminal outcome is determined.
- Usage: Provides a direct signal for rejection and eviction rates without log parsing. Unlike
flow_control_request_queue_duration_seconds_count, this counter also captures controller-level early rejections where no queue item is created (e.g., rejection during controller shutdown), covering cases the histogram misses. - Actionability:
- A rising rate of
outcome="RejectedCapacity"indicates the queue capacity limits are too tight or backends are persistently saturated — consider tuningmaxBytes/maxRequestsor scaling backends. - A rising rate of
outcome="RejectedNoEndpoints"indicates the inference pool has scaled to zero or all endpoints are unregistered — investigate pool health and scaling configuration. - A rising rate of
outcome="EvictedTTL"indicates requests are waiting longer than their TTL allows — investigate backend throughput or tighten admission. outcome="Dispatched"is the healthy baseline; compare it against total request rate to derive the acceptance ratio.
- A rising rate of
Three metrics covering ext_proc gRPC stream lifecycle. Disabled by default; enable with --enable-grpc-stream-metrics. These metrics are emitted under the llm_d_epp_ prefix (separate from llm_d_inference_scheduler_*).
- Type: Gauge
- Release Stage: ALPHA
- Description: Number of ext_proc gRPC streams currently open.
- Usage: Sized at one stream per Envoy worker per EPP backend. A persistent increase under steady load indicates streams are being opened faster than they close.
- Type: Histogram
- Release Stage: ALPHA
- Description: Duration an ext_proc gRPC stream stays open, in seconds.
- Usage: Long-lived streams are normal; the histogram surfaces the distribution. A sudden shift toward short durations can indicate Envoy reconnecting due to handler errors.
- Type: Counter
- Labels:
code: string — the gRPC status code at stream close (OK,Canceled,DeadlineExceeded,Internal, ...). Barecontext.Canceledandcontext.DeadlineExceededare classified to their canonical codes rather than collapsing intoUnknown.
- Release Stage: ALPHA
- Description: Total ext_proc gRPC streams completed, by gRPC status code.
- Usage: Rate of
code="OK"is the healthy stream-completion rate. A rising rate ofcode="Internal"orcode="Unknown"indicates handler errors.code="Canceled"is expected on Envoy restarts and rolling EPP updates.
In-flight load, emitted under the llm_d_epp_ prefix. Present only when an InFlightLoadProducer is
configured: the producer owns these metrics and registers them through the plugin metrics recorder. The
per-endpoint gauges are updated from the producer's live per-endpoint counters as requests are admitted
and released (the same source as the /debug/plugins/state dump and the token-load scorer); the
per-model request_inflight gauge is moved by the producer as requests are admitted and completed.
- Type: Gauge
- Labels:
endpoint_name: string — the target endpoint (pod) name.namespace: string — the endpoint's namespace.producer_name: string — the configuredInFlightLoadProducerinstance name, so multiple producers emit distinct series.
- Release Stage: ALPHA
- Description: Requests currently in flight on each endpoint (scheduled, not yet completed), as tracked by the in-flight load producer.
- Usage: Per-replica queue depth for load-aware routing and capacity analysis. Unlike the per-model
request_inflightgauge (admitted-but-not-completed, aggregated by model), this is broken down by endpoint so it shows which replica is loaded.
- Type: Gauge
- Labels:
endpoint_name: string — the target endpoint (pod) name.namespace: string — the endpoint's namespace.producer_name: string — the configuredInFlightLoadProducerinstance name.
- Release Stage: ALPHA
- Description: Tokens currently in flight on each endpoint — uncached prompt tokens, optionally plus estimated output tokens when the producer's
addEstimatedOutputTokensis set. - Usage: Per-replica token pressure, a finer load signal than request count when request sizes vary widely.
- Type: Gauge
- Labels:
model_name: string — the model named in the request body.target_model_name: string — the target model after traffic split.fairness_id: string — the flow-control fairness queue identity.priority: string — the request priority.
- Release Stage: ALPHA
- Description: Requests admitted to the endpoint picker but not yet completed, aggregated by model.
- Usage: Picker-wide concurrency by model. Unlike the per-endpoint
inflight_requestsgauge, this is aggregated across endpoints, so it answers "how much is in flight for this model" rather than "which replica is loaded".