Skip to content

fix: override idle timeout with drain time in draining phase - #47704

Open
rudrakhp wants to merge 9 commits into
envoyproxy:mainfrom
rudrakhp:http_drain_idle_timeout
Open

rudrakhp wants to merge 9 commits into
envoyproxy:mainfrom
rudrakhp:http_drain_idle_timeout

Conversation

@rudrakhp

@rudrakhp rudrakhp commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Commit Message: fix: override idle timeout with drain time in draining phase
Additional Description: Override the idle timeout with drain time when proxy is in draining phase.
Risk Level: Low (new API)
Testing: Unit and Integ testing
Docs Changes: Yes
Release Notes: Yes
Platform Specific Features: N/A
Fixes #42305

@repokitteh-read-only

Copy link
Copy Markdown

CC @envoyproxy/api-shepherds: Your approval is needed for changes made to (api/envoy/|docs/root/api-docs/).
envoyproxy/api-shepherds assignee is @htuch
CC @envoyproxy/api-watchers: FYI only for changes made to (api/envoy/|docs/root/api-docs/).

🐱

Caused by: #47704 was opened by rudrakhp.

see: more, trace.

Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
@rudrakhp
rudrakhp force-pushed the http_drain_idle_timeout branch from d328b63 to 1ea113f Compare September 24, 2026 17:43

@htuch htuch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not entirely clear on why we need a new field in the API here - why can't we start draining immediately for idle connections? If there's a bug, we can fix that without an API change. I'd like to keep things operationally simple here, because this makes life more complicated for operators AFAICT.

@rudrakhp

Copy link
Copy Markdown
Member Author

Thanks @htuch, I agree that avoiding another knob would be preferable if we can settle on suitable default behavior.

The reason I took this approach was the HTTP/1 concern raised by @yanavlasov in the comment #42305 (comment): a connection that Envoy considers idle can race with the client sending its next request, so immediately closing all idle connections on drain can cause request failures.

So I was treating this as a way to decouple normal connection reuse from shutdown latency. If we consider immediately closing idle HTTP/1 connections during graceful drain an acceptable default despite that race (which I agree is what most require), I am open to pursuing the simpler behaviour instead.

@htuch

htuch commented Sep 30, 2026

Copy link
Copy Markdown
Member

Doesn't the race still exist in this PR, since you could race with the client after the drain idle duration? It seems an inherent property of H1. CC @yanavlasov

@rudrakhp

rudrakhp commented Sep 30, 2026 •

Copy link
Copy Markdown
Member Author

Yes, this doesn't guarantee that there won't be a race. But a smaller (nonzero) drain_idle_timeout can be useful to configure how long to wait before safely assuming no callers. For example, time needed in a service mesh setup during shutdown for updates about the terminating endpoint to propagate to all its callers.

About the default behavior (when unset) I can see why we should do it immediately rather than keep the existing behavior of waiting for idle_timeout.

@htuch

htuch commented Oct 1, 2026

Copy link
Copy Markdown
Member

I'm not suggesting to cut the connection at. the start of the drain. The point is more that at the end of the drain period, that's when it makes sense to cut; i.e. drain duration overrides idle timeout.

@rudrakhp

rudrakhp commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

drain duration overrides idle timeout

Ah I see using the existing drain time as the idle timeout makes sense. Making that change.

Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
@rudrakhp rudrakhp changed the title http: support different idle timeout when draining fix: override idle timeout with drain time in draining phase Oct 1, 2026
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
@rudrakhp
rudrakhp requested a review from htuch October 1, 2026 05:09
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
…dec connections

Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>
Signed-off-by: Rudrakh Panigrahi <rudrakh97@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

allow close downstream idle connection while draining

2 participants