Steps to reproduce
- Create a targeted AWS On-Demand Capacity Reservation:
aws ec2 create-capacity-reservation --region us-east-1 --availability-zone us-east-1a \
--instance-type m5.large --instance-platform Linux/UNIX --instance-count 1 \
--instance-match-criteria targeted
- Apply an elastic fleet that sets the reservation:
type: fleet
name: odcr-test
backends: [aws]
regions: [us-east-1]
instance_types: [m5.large]
nodes: 0..1
reservation: cr-<id>
- Apply a task in that fleet without
reservation:
type: task
fleets: [odcr-test]
commands: [sleep 600]
Actual behaviour
The run provisions an m5.large, but outside the reservation: the EC2 instance has no CapacityReservationId, and the reservation's AvailableInstanceCount stays at 1. The instance can also land in a different AZ than the reservation's. Setting the same reservation on the run as well makes the instance land in the reservation.
Offers are selected using the effective requirements, where the fleet's reservation is used if the run doesn't set one:
|
reservation=get_single_value_optional( |
|
fleet_requirements.reservation, run_requirements.reservation |
|
), |
These requirements are also passed to run_job():
|
job_provisioning_data = await run_async( |
|
compute.run_job, |
|
run, |
|
job, |
|
offer, |
|
project_ssh_public_key, |
|
project_ssh_private_key, |
|
offer_volumes, |
|
placement_group_model_to_placement_group_optional(placement_group_model), |
|
requirements, |
|
extra_authorized_keys, |
But the default run_job() builds InstanceConfiguration.reservation from the run's job spec only, so it is None:
|
instance_config = InstanceConfiguration( |
|
project_name=run.project_name, |
|
instance_name=get_job_instance_name(run, job), |
|
user=run.user, |
|
ssh_keys=[SSHKey(public=project_ssh_public_key.strip())], |
|
volumes=volumes, |
|
reservation=job.job_spec.requirements.reservation, |
|
tags=run.run_spec.merged_profile.tags, |
|
) |
With no reservation, AWS create_instance() doesn't restrict subnets to the reservation's AZ and omits CapacityReservationSpecification. AWS then uses the default open preference, which never consumes a targeted reservation:
|
if reservation_id is not None: |
|
struct["CapacityReservationSpecification"] = { |
|
"CapacityReservationTarget": {"CapacityReservationId": reservation_id} |
|
} |
Expected behaviour
A run provisioned in a fleet with reservation launches into that reservation, whether or not the run sets reservation itself.
dstack version
master (0d578c8)
Additional information
Fleet-triggered provisioning (nodes.min > 0) isn't affected, since it builds the instance configuration from the fleet spec.
Open reservations mask the bug, because AWS places a matching instance into an open reservation without targeting it.
For AWS Capacity Blocks, the launch can't succeed without the reservation, so runs in an elastic Capacity Block fleet fail. #4344 fixes cluster placement groups for Capacity Blocks, but the run-triggered path still depends on this bug being fixed.
GCP also uses the default run_job() and reads InstanceConfiguration.reservation in create_instance(), so it is likely affected as well (not verified).
A possible fix is to take the reservation from the effective requirements in run_job(), and in the placement group check in jobs_submitted.py.
Steps to reproduce
reservation:Actual behaviour
The run provisions an
m5.large, but outside the reservation: the EC2 instance has noCapacityReservationId, and the reservation'sAvailableInstanceCountstays at 1. The instance can also land in a different AZ than the reservation's. Setting the samereservationon the run as well makes the instance land in the reservation.Offers are selected using the effective requirements, where the fleet's
reservationis used if the run doesn't set one:dstack/src/dstack/_internal/server/services/requirements/combine.py
Lines 71 to 73 in 0d578c8
These requirements are also passed to
run_job():dstack/src/dstack/_internal/server/background/pipeline_tasks/jobs_submitted.py
Lines 2512 to 2522 in 0d578c8
But the default
run_job()buildsInstanceConfiguration.reservationfrom the run's job spec only, so it isNone:dstack/src/dstack/_internal/core/backends/base/compute.py
Lines 412 to 420 in 0d578c8
With no reservation, AWS
create_instance()doesn't restrict subnets to the reservation's AZ and omitsCapacityReservationSpecification. AWS then uses the defaultopenpreference, which never consumes a targeted reservation:dstack/src/dstack/_internal/core/backends/aws/resources.py
Lines 212 to 215 in 0d578c8
Expected behaviour
A run provisioned in a fleet with
reservationlaunches into that reservation, whether or not the run setsreservationitself.dstack version
master (0d578c8)
Additional information
Fleet-triggered provisioning (
nodes.min > 0) isn't affected, since it builds the instance configuration from the fleet spec.Open reservations mask the bug, because AWS places a matching instance into an open reservation without targeting it.
For AWS Capacity Blocks, the launch can't succeed without the reservation, so runs in an elastic Capacity Block fleet fail. #4344 fixes cluster placement groups for Capacity Blocks, but the run-triggered path still depends on this bug being fixed.
GCP also uses the default
run_job()and readsInstanceConfiguration.reservationincreate_instance(), so it is likely affected as well (not verified).A possible fix is to take the reservation from the effective
requirementsinrun_job(), and in the placement group check injobs_submitted.py.