The opensearch-go transport exposes a pull-based metrics API that returns a point-in-time snapshot of request counters, connection pool state, per-connection health, policy-level breakdowns, and router cache state. All fields are JSON-tagged for easy serialization.
Metrics are collected using atomic, per-request counters (e.g. requests, failures, and responses-by-status) recorded on the request hot path. Additional detailed metrics (connection-pool state, per-connection health, per-policy breakdowns, and router cache state) are lazily accumulated and read lock-free when you call Metrics(), so they cost nothing until you ask for them.
Metrics require no configuration. Construct a client and call Metrics(). When constructing through opensearchapi.NewClient, reach the method via apiClient.Client.Metrics().
client, err := opensearch.NewClient(opensearch.Config{
Addresses: []string{"https://localhost:9200"},
})
if err != nil {
log.Fatal(err)
}
m, err := client.Metrics()
if err != nil {
log.Fatal(err)
}
data, _ := json.MarshalIndent(m, "", " ")
fmt.Println(string(data))The Metrics() method lives on opensearch.Client. It returns an opensearchtransport.Metrics struct and an error -- non-nil only when a detailed-snapshot callback fails. It never errors merely because metrics are "disabled"; they are always available.
The top-level Metrics struct contains aggregate counters, per-connection details, per-policy pool snapshots, and router cache state.
| Field | JSON | Type | Description |
|---|---|---|---|
Requests |
requests |
int |
Total requests performed by the transport |
Failures |
failures |
int |
Total request failures |
Responses |
responses |
map[int]int |
Response count by HTTP status code |
| Field | JSON | Type | Description |
|---|---|---|---|
LiveConnections |
live_connections |
int |
Non-dead connections (active + standby) |
DeadConnections |
dead_connections |
int |
Connections in the dead list |
StandbyConnections |
standby_connections |
int |
Connections in the standby partition |
OverloadedServers |
overloaded_servers |
int |
Connections the client considers overloaded |
| Field | JSON | Type | Description |
|---|---|---|---|
ConnectionsPromoted |
connections_promoted |
int |
Dead to ready transitions (successful resurrections) |
ConnectionsDemoted |
connections_demoted |
int |
Ready to dead transitions |
ZombieConnections |
zombie_connections |
int |
Dead connections forcibly retried |
StandbyPromotions |
standby_promotions |
int |
Standby to active transitions |
StandbyDemotions |
standby_demotions |
int |
Active to standby transitions |
| Field | JSON | Type | Description |
|---|---|---|---|
HealthChecks |
health_checks |
int |
Baseline GET / health checks performed |
ClusterHealthChecks |
cluster_health_checks |
int |
GET /_cluster/health?local=true checks performed |
HealthChecksSuccess |
health_checks_success |
int |
Successful health check outcomes |
HealthChecksFailed |
health_checks_failed |
int |
Failed health check outcomes |
Recorded by the client-side DNS cache installed on the built-in transport (see DNSCacheRefresh). All zero when a custom Transport is supplied or caching is disabled.
| Field | JSON | Type | Description |
|---|---|---|---|
DNSLookups |
dns_lookups |
int |
Dials that consulted the DNS cache |
DNSCacheMisses |
dns_cache_misses |
int |
Lookups not served from cache (cold, or re-resolved after eviction) |
DNSLookupErrors |
dns_lookup_errors |
int |
Lookups that returned a resolution error |
Each connection produces a ConnectionMetric in Metrics.Connections. Connections are deduplicated: the same *Connection appearing in multiple policy pools is reported once.
| Field | JSON | Type | Description |
|---|---|---|---|
URL |
url |
string |
Node URL |
Failures |
failures |
int |
Failure count (omitted when zero) |
IsDead |
dead |
bool |
In the dead list |
IsStandby |
standby |
bool |
In the standby partition |
IsOverloaded |
overloaded |
bool |
Whether the connection is currently flagged overloaded |
IsWarmingUp |
warming_up |
bool |
In warmup phase after promotion |
IsHealthChecking |
health_checking |
bool |
Currently being health-checked |
NeedsCatUpdate |
needs_cat_update |
bool |
Shard placement data is stale |
Weight |
weight |
int |
Effective weight for selection |
DeadSince |
dead_since |
*time.Time |
When the connection was marked dead |
OverloadedSince |
overloaded_since |
*time.Time |
When overload was detected |
State |
state |
ConnState |
Packed connection state word |
Populated when request routing is active and the connection has observed traffic.
| Field | JSON | Type | Description |
|---|---|---|---|
RTTBucket |
rtt_bucket |
*int64 |
Quantized RTT tier (smaller values indicate lower observed RTT) |
RTTMedian |
rtt_median |
*string |
Median observed RTT as a human-readable duration |
EstLoad |
est_load |
*float64 |
Estimated per-connection load |
MCSR |
mcsr |
*int |
Adaptive max_concurrent_shard_requests value (nil when disabled) |
| Field | JSON | Type | Description |
|---|---|---|---|
Meta.ID |
meta.id |
string |
OpenSearch node ID |
Meta.Name |
meta.name |
string |
OpenSearch node name |
Meta.Roles |
meta.roles |
[]string |
Node roles (data, ingest, cluster_manager, etc.) |
Each policy with a connection pool produces a PolicySnapshot in Metrics.Policies. Policies register their snapshot callback at construction time; the Metrics() method invokes all callbacks on each poll.
| Field | JSON | Type | Description |
|---|---|---|---|
Name |
name |
string |
Policy name (e.g., role:data, roundrobin, coordinator, client) |
Enabled |
enabled |
bool |
Whether this policy is currently routing traffic |
ActiveCount |
active_count |
int |
Connections in the active partition |
StandbyCount |
standby_count |
int |
Connections in the standby partition |
DeadCount |
dead_count |
int |
Connections in the dead list |
ActiveListCap |
active_list_cap |
int |
Capacity of the active partition |
WarmingCount |
warming_count |
int |
Connections in warmup |
HealthCheckingCount |
health_checking_count |
int |
Connections being health-checked |
Requests |
requests |
int64 |
Connections returned by Next() |
Successes |
successes |
int64 |
Resurrections via OnSuccess() |
Failures |
failures |
int64 |
Demotions via OnFailure() |
WarmupSkips |
warmup_skips |
int64 |
Requests skipped during warmup |
WarmupAccepts |
warmup_accepts |
int64 |
Requests accepted during warmup |
The client pool snapshot is always present and represents the flat connection pool shared by all policies.
When request routing is active, Metrics.Router contains a RouterSnapshot with per-index routing state and the effective cache configuration.
| Field | JSON | Type | Description |
|---|---|---|---|
MinFanOut |
min_fan_out |
int |
Minimum candidates per routing decision |
MaxFanOut |
max_fan_out |
int |
Maximum candidates per routing decision |
DecayFactor |
decay_factor |
float64 |
Exponential decay for request rate smoothing |
FanOutPerReq |
fan_out_per_request |
float64 |
Fan-out growth per request |
IdleEvictionTTL |
idle_eviction_ttl |
string |
Time before idle index slots are evicted |
Each index with an active routing slot produces an entry in Router.Indexes.
| Field | JSON | Type | Description |
|---|---|---|---|
Name |
name |
string |
Index name |
FanOut |
fan_out |
int |
Current fan-out (candidate count) for this index |
ShardNodes |
shard_nodes |
int |
Nodes with known shard placement for this index |
RequestRate |
request_rate |
float64 |
Smoothed request rate |
IdleSince |
idle_since |
*time.Time |
When the index became idle (nil if active) |
Poll metrics on a timer for logging or export to an external monitoring system.
func safeFloat(f *float64) float64 {
if f == nil {
return 0
}
return *f
}
ticker := time.NewTicker(30 * time.Second)
defer ticker.Stop()
for range ticker.C {
m, err := client.Metrics()
if err != nil {
log.Printf("metrics error: %v", err)
continue
}
log.Printf("requests=%d failures=%d live=%d dead=%d standby=%d overloaded=%d",
m.Requests, m.Failures,
m.LiveConnections, m.DeadConnections,
m.StandbyConnections, m.OverloadedServers)
// Per-connection detail
for _, c := range m.Connections {
cm := c.(opensearchtransport.ConnectionMetric)
if cm.RTTMedian != nil {
log.Printf(" %s rtt=%s load=%.2f",
cm.URL, *cm.RTTMedian, safeFloat(cm.EstLoad))
}
}
// Per-policy breakdown
for _, p := range m.Policies {
log.Printf(" policy %s: active=%d standby=%d dead=%d req=%d fail=%d",
p.Name, p.ActiveCount, p.StandbyCount, p.DeadCount, p.Requests, p.Failures)
}
// Router cache (when routing is active)
if m.Router != nil {
for _, idx := range m.Router.Indexes {
log.Printf(" index %s: fan_out=%d shard_nodes=%d rate=%.1f",
idx.Name, idx.FanOut, idx.ShardNodes, idx.RequestRate)
}
}
}The Metrics struct is fully JSON-tagged. Serialize with encoding/json for structured logging or HTTP endpoints:
m, _ := client.Metrics()
data, _ := json.Marshal(m)
w.Header().Set("Content-Type", "application/json")
w.Write(data)The metrics API is pull-based: call client.Metrics() inside your collector's Collect() or OTEL callback. Map fields to gauges and counters as appropriate:
| Metrics field | Metric type | Suggested name |
|---|---|---|
Requests |
Counter | opensearch_client_requests_total |
Failures |
Counter | opensearch_client_failures_total |
Responses[code] |
Counter (labeled) | opensearch_client_responses_total{code="200"} |
LiveConnections |
Gauge | opensearch_client_connections_live |
DeadConnections |
Gauge | opensearch_client_connections_dead |
StandbyConnections |
Gauge | opensearch_client_connections_standby |
OverloadedServers |
Gauge | opensearch_client_connections_overloaded |
ConnectionMetric.EstLoad |
Gauge (per-node) | opensearch_client_node_est_load{node="..."} |
ConnectionMetric.MCSR |
Gauge (per-node) | opensearch_client_node_mcsr{node="..."} |
PolicySnapshot.ActiveCount |
Gauge (per-policy) | opensearch_client_policy_active{policy="..."} |
IndexRouterState.FanOut |
Gauge (per-index) | opensearch_client_index_fan_out{index="..."} |
For event-driven observability (as opposed to polling), implement the opensearchtransport.ConnectionObserver interface and set it on the Observer field of opensearch.Config. The observer receives callbacks for connection lifecycle events (promote, demote, overload), routing decisions, health checks, and shard map invalidations. See the routing guide for details on observer events.
The metrics API and observer API are complementary: metrics give you aggregate snapshots for dashboards, while the observer gives you per-event detail for tracing and debugging.