Repository navigation
[Feature]: Support in-place gateway recreation to preserve ALB/IP and avoid DNS changes #3924
Description
Activity
Most items from #3959 are going to be a prerequisite. Working on them now
This issue is stale because it has been open for 30 days with no activity.
dstack0.21.1 allows you to migrate the gateway to a new instance through a sequence of scaling out and scaling in.- For gateways with ACM certificates, this can be done without downtime and without DNS changes. The same ALB is preserved.
- For gateways without a certificate, DNS changes are still required. Supporting redeployment without DNS changes for such gateways is a work in progress.
Instructions for updating a single-replica gateway with an ACM certificate
-
Update the server and the CLI to 0.21.1 or newer.
-
Locate the YAML configuration that was used to create the gateway, or restore it from
dstack gateway get --json <gateway name>. -
Set
replicas: 2in the configuration. -
Apply the configuration.
$ dstack apply -f my-gateway.dstack.yml Project main User admin Configuration my-gateway.dstack.yml Type gateway Backend aws Region eu-west-1 Domain example.com Replicas 2 Found gateway my-gateway. Detected changes that can be updated in-place: - replicas Update the gateway? [y/n]: y NAME BACKEND HOSTNAME DOMAIN DEFAULT STATUS my-gateway aws (eu-west-1) dstack-xaoxm1dz-lb-1834091824.eu-west-1.elb.amazonaws.com example.com ✓ runningNOTE: Make sure
dstack applysays the gateway can be updated. If it says the gateway cannot be updated, do not confirm, as that would result in gateway termination. -
Wait until the second replica is
running. This may take about two minutes.$ dstack gateway -w NAME BACKEND HOSTNAME DOMAIN DEFAULT STATUS my-gateway dstack-xaoxm1dz-lb-1834091824.eu-west-1.elb.amazonaws.com example.com ✓ running replica=0 aws (eu-west-1) 52.48.189.229 running replica=1 aws (eu-west-1) 3.248.191.93 running -
Verify that existing services have been registered on the second replica. This may take a few seconds.
$ # UPD: on 0.20.3+, use --within-gateway instead of --target-gateway $ dstack event --target-gateway my-gateway [2026-08-14 19:57:06] [👤admin] [gateway my-gateway] Gateway updated. Changed fields: replicas [2026-08-14 19:59:08] [run first-service, gateway my-gateway] Service registered on gateway replica 1 [2026-08-14 19:59:08] [run second-service, gateway my-gateway] Service registered on gateway replica 1 [2026-08-14 19:59:08] [job second-service-0-0, gateway my-gateway] Service replica registered on gateway replica 1 [2026-08-14 19:59:08] [job first-service-0-0, gateway my-gateway] Service replica registered on gateway replica 1 -
Verify that the services are responding as expected.
-
Set
replicas: 1in the configuration. -
Apply the configuration.
-
Wait until the first replica is
terminated.$ dstack gateway -w NAME BACKEND HOSTNAME DOMAIN DEFAULT STATUS my-gateway dstack-xaoxm1dz-lb-1834091824.eu-west-1.elb.amazonaws.com example.com ✓ running replica=0 aws (eu-west-1) 52.48.189.229 terminated replica=1 aws (eu-west-1) 3.248.191.93 running
Today's 0.21.3 release addresses the part about gateways without certificates — it is now possible to create gateways without a certificate while still using an ALB (#4219). Such gateways can then be redeployed without downtime using the sequence described in my comment above.
Existing
certificate: nullgateways can be migrated to the new model using the following sequence:- Create a new gateway with the same domain, but with load_balancer: { type: alb }.
- Move the services from the old gateway to the new one. This can be done without stopping the services (Support in-place update for service
gateway#4221). - Update the DNS records to point to the new gateway's ALB.
- Delete the old gateway.
Service downtime is expected during steps 2–3 and while the DNS changes propagate. However, this is a one-time disruptive migration; future redeployments of the new gateway can be performed without downtime thanks to the ALB
I believe the feature is ready, so I’ll close the ticket. If anything is missing, don’t hesitate to reopen it, open a new ticket, or contact us directly
Problem
Summary
When a dstack gateway needs to be recreated (e.g., to refresh the underlying EC2 instance), the current behavior forces downstream DNS or CNAME changes because the AWS Load Balancer and/or EC2 public IP change.
We'd like to request support for in-place gateway recreation that preserves the externally-visible endpoint.
Current Behavior
Gateway with cert (HTTPS)
Gateway without cert
Solution
Requested Behavior
Gateway with cert
Support in-place recreation that:
Result: no CNAME change required, minimal/zero downtime.
Gateway without cert
Support in-place restart/recreation that:
Result: no DNS change required.
Motivation / Background
We have several dstack gateways that were created some time ago and are now running Linux kernel versions that no longer meet our internal security compliance policy. The policy requires that the kernel release
date be within the last 180 days, with a 60-day remediation window once an asset is flagged.
To remediate, we need to refresh these gateways onto a current AMI/kernel. Today, doing so forces:
Because gateway recreation will be a recurring operational task (driven by quarterly patching cycles, not just one-time kernel upgrades), the DNS-change toil compounds significantly. In-place recreation would let
us meet our patching SLA without coordinating downstream DNS changes each cycle.
Workaround
No response
Would you like to help us implement this feature by sending a PR?
No