- Author(s): @gtcooke94
- Approver: @dfawley, @easwars, @ejona86, @matthewstevenson88, @markdroth
- Status: Approved
- Implemented in:
- Last updated: 2026-06-03
- Discussion at: https://groups.google.com/g/grpc-io/c/hwK5GWDvL20?e=48417069
gRPC's authentication stack has no telemetry. This document details adding non-per-call metrics to gRPC's SSL/TLS authentication stack.
gRPC's authentication stack currently lacks telemetry. This document outlines the addition of authentication-related non-per-call metrics, leveraging the existing telemetry infrastructure defined in A66 and the specific architecture for non-per-call metrics established in A79. Because authentication and handshakers are connection-level abstractions, they inherently require non-per-call instrumentation. Currently, application owners face challenges in diagnosing handshake failures, relying on verbose logging without a structured mechanism for aggregation and analysis.
- A66: OpenTelemetry Metrics/Stats
- A78: gRPC OTel Metrics for WRR, Pick First, and XdsClient
- A79: OpenTelemetry Non-Per-Call Metrics Architecture
- A89: Backend Service Metric Label
- A94: OTel metrics for Subchannels
- A107: TLS Private Key Offloading
All metrics will be scoped to TLS exclusively.
The TLS handshake result will be represented by an enum that indicates success
or provides information on why the handshake failed. This value must manage a
balance of low-cardinality while being fine-grained enough to be useful;
therefore, a set consisting of subdomains of authentication errors will be
created. These are strings that will be a label value in the handshake metrics.
These strings will be identical in all languages. In cases where we cannot
categorize an error or cannot get enough granularity in a given implementation
and/or language, UNKNOWN_FAILURE will be the catch-all error code.
"UNKNOWN_FAILURE",
"SUCCESS",
// Peer certificate verification failures.
"CERTIFICATE_VERIFICATION_FAILED",
"CERTIFICATE_REVOKED",
"CERTIFICATE_EXPIRED",
"CERTIFICATE_NOT_YET_VALID",
"CERTIFICATE_AUTHORITY_INVALID",
// TLS negotiation mismatch failures
"CERTIFICATE_HOSTNAME_MISMATCH",
"CERTIFICATE_MALFORMED",
"CIPHER_SUITE_MISMATCH",
"PROTOCOL_VERSION_UNSUPPORTED",
"INAPPROPRIATE_FALLBACK",
"NO_APPLICATION_PROTOCOL",
// Cryptographic failures
"SIGNATURE_VERIFICATION_FAILED",
"DECRYPTION_FAILED",
"KEY_EXCHANGE_FAILURE",
// Other failures
"UNEXPECTED_MESSAGE",
"HANDSHAKE_TIMEOUT",
"PEER_CONNECTION_CLOSED"
Further, whether the handshake is resumed or not is also critical for understanding authentication behavior. An enum describing the type of resumption used (or none) will be created - it can be extended in the future to make finer-grained distinctions between the type of resumption that is used (e.g., ticket-based resumption vs. session-based resumption). This is presented as a C++ enum below, but will be identical in all languages.
enum class TlsResumptionType {
FULL_HANDSHAKE,
RESUMED_HANDSHAKE,
};The following metrics are count metrics, with the primary information coming from the labels.
grpc.client.tls.handshakes(unit: {handshake} type: int64 counter. description: EXPERIMENTAL: The number of handshakes performed on the client side. The primary useful information comes from the the labels, in particular the result label contains information on why a handshake failed if it did not succeed.)
| Label Name | Required/Optional | Description |
|---|---|---|
grpc.tls.handshake.result |
Required | The TlsTelemetryHandshakeResult enum string indicating success or the reason for handshake failure. |
grpc.target |
Required | The target string (as defined in A66) passed to the channel. |
grpc.tls.handshake.resumed |
Optional | The TlsResumptionType enum string indicating if and how the handshake was resumed. |
grpc.lb.locality |
Optional | The locality to which the traffic is being sent (as defined in A78, A94). |
grpc.lb.backend_service |
Optional | The backend service to which the traffic is being sent (as defined in A89, A94). |
grpc.server.tls.handshakes(unit: {handshake} type: int64 counter. description: EXPERIMENTAL: The number of handshakes performed on the server side. The primary useful information comes from the the labels, in particular the result label contains information on why a handshake failed if it did not succeed.))
| Label Name | Required/Optional | Description |
|---|---|---|
grpc.tls.handshake.result |
Required | The TlsTelemetryHandshakeResult enum string indicating success or the reason for handshake failure. |
grpc.tls.handshake.resumed |
Optional | The TlsResumptionType enum string indicating if and how the handshake was resumed. |
The following metrics are non-per-call bucketed latency metrics that report the duration of offloaded cryptographic operations.
grpc.server.tls.offload_certificate_selection_duration(unit: s, type: float64 histogram - latency buckets defined in A66. description: EXPERIMENTAL: Measures the duration of the offloaded certificate selection operation.)
| Label Name | Required/Optional | Description |
|---|---|---|
grpc.status |
Required | Result of the certificate selection offloading, in the format of a gRPC status code (as defined in A66). |
Note - there is no associated client certificate selection metric. This is a server specific feature.
For the offloaded private key metrics, we specify signing because that is the only offloaded private key operation supported by gRPC (for example, signature offload using an EC or RSA key). Older TLS versions have the concept of other private key operations that gRPC does not support (for example, decryption offload using an RSA key). See A107 for more detail on private key signers.
grpc.client.tls.offload_private_key_signing_duration(unit: s, type: float64 histogram - latency buckets defined in A66. description: EXPERIMENTAL: Measures the duration of the offloaded private key signing operation.)
| Label Name | Required/Optional | Description |
|---|---|---|
grpc.status |
Required | Result of the private key signing offloading, in the format of a gRPC status code (as defined in A66). |
grpc.target |
Required | The target string (as defined in A66) passed to the channel. |
grpc.tls.private_key.offloader_name |
Required | A string identifying the private key signer implementation, e.g. "HSM" or "private_key_signer_service". This must be low-cardinality |
grpc.tls.private_key_algorithm |
Optional | An algorithm enum indicating how the offloaded private key signing was done, e.g. “RsaPkcs1Sha256”. |
grpc.lb.locality |
Optional | The locality to which the traffic is being sent (as defined in A78, A94). |
grpc.lb.backend_service |
Optional | The backend service to which the traffic is being sent (as defined in A89, A94). |
grpc.server.tls.offload_private_key_signing_duration(unit: s, type: float64 histogram - latency buckets defined in A66. description: EXPERIMENTAL: Measures the duration of the offloaded private key signing operation.)
| Label Name | Required/Optional | Description |
|---|---|---|
grpc.status |
Required | Result of the private key signing offloading, in the format of a gRPC status code (as defined in A66). |
grpc.tls.private_key.offloader_name |
Required | A string identifying the private key signer implementation e.g. "HSM" or "private_key_signer_service". This must be low-cardinality. |
grpc.tls.private_key_algorithm |
Optional | An algorithm enum indicating how the offloaded private key signing was done, e.g. “RsaPkcs1Sha256”. |
This feature will be explicitly configured by users, thus no environment variable protection is needed. If a user does not configure TLS metrics or offloading, telemetry won't be collected under this mechanism unless general telemetry/stats plugins are active.
The alternative of splitting generic handshaker metrics and TLS-specific metrics was considered, but opened many rabbit holes as to what would qualify as a handshaker (e.g. are HTTP Connect, TCP, etc. all in scope, or just alternative protocols to TLS such as ALTS). Thus, the decision was made to scope the metrics to TLS.
In the C/C++ implementation, the transport security interface (TSI) has historically been decoupled from gRPC. However, this design choice has been broken over time, and the two have been coupled for years now. The reasons behind decoupling TSI and gRPC are no longer relevant, therefore we will fully accept this coupling. Thus, the TSI code can contain gRPC monitoring specifics. In the few use-cases where TSI is not called via gRPC, we will ensure that metric incrementation is not performed.
We will add the Channel's CollectionScope as an optional argument to the SSL
TSI handshaker creation functions. This ensures that we don't break any existing
users of TSI and that we never increment metrics when TSI is used outside of
gRPC. When TSI is called from gRPC, we will pass this argument. This will be
stored on the handshaker, and in ssl_transport_security.cc we will access this
from the handshaker to increment metrics.
tsi_result tsi_ssl_client_handshaker_factory_create_handshaker(
tsi_ssl_client_handshaker_factory* factory,
const char* server_name_indication, size_t network_bio_buf_size,
size_t ssl_bio_buf_size,
std::optional<std::string> alpn_preferred_protocol_list,
+ grpc_core::RefCountedPtr<grpc_core::CollectionScope> collection_scope,
tsi_handshaker** handshaker);
tsi_result tsi_ssl_server_handshaker_factory_create_handshaker(
tsi_ssl_server_handshaker_factory* factory, size_t network_bio_buf_size,
- size_t ssl_bio_buf_size, tsi_handshaker** handshaker);
+ size_t ssl_bio_buf_size,
+ grpc_core::RefCountedPtr<grpc_core::CollectionScope> collection_scope,
+ tsi_handshaker** handshaker);For the implementation of the general handshake metric, we leverage gRPC-Go's existing transport-level architecture.
The abstraction layer for general handshake information is the transport
connection layer (internal/transport). Client handshakes are performed in a
single place: internal/transport/http2_client.go inside NewHTTP2Client via
transportCreds.ClientHandshake. Server handshakes are also performed in a
single place: internal/transport/http2_server.go inside NewServerTransport
via config.Credentials.ServerHandshake. We can increment the metrics here
only in the case where TLS is the protocol being used.
--- a/internal/transport/http2_client.go
+++ b/internal/transport/http2_client.go
@@ -294,7 +294,23 @@
if transportCreds != nil {
isTLS := transportCreds.Info().SecurityProtocol == "tls"
conn, authInfo, err = transportCreds.ClientHandshake(connectCtx, addr.ServerName, conn)
+ if isTLS {
+ <increment client/server handshakes metric>
+ }
if err != nil {
return nil, connectionErrorf(isTemporary(err), err, "transport: authentication handshake failed: %v", err)
}
To extract resumption information, we will need to augment the TLSInfo to include ConnectionState.DidResume.
The TLS specific offload metrics will go with their implementation. This feature is not yet written in Go, so we cannot discuss specific metric implementation details.
In gRPC-Java, to support shading where core classes are relocated per-consumer, all metric instruments used in core/ or transport modules (like netty/) must be defined within the api/ module.
Following the pattern established by MIN_RTT_INSTRUMENT in InternalTcpMetrics.java, we will:
- Define a new
InternalSecurityMetrics.javaclass in the api/ module under theio.grpcpackage. - Expose
MetricRecorderfromGrpcHttp2ConnectionHandlerin the netty/ module. - Instrument the Netty
ClientTlsHandlerandServerTlsHandlerinProtocolNegotiators.javato track handshake duration and record it.
The TLS specific offload metrics will go with their implementation. This feature is not yet written in Java, so we cannot discuss specific metric implementation details.
Wrapped languages that support non-per-call metrics will get the authentication telemetry features "for free" from the Core implementation. However, Python for example, does not currently support non-per-call metrics.