Repository navigation
Rebase the fork on upstream 0.22.2+ for Nebius Dynamic Spot pricing (keep fork features) - #1
Merged
Merged
Conversation
Introduce gateway in-place update mechanism. For
now, only `domain` can be updated.
```yaml
$ dstack apply -f test.dstack.yml
Found gateway test-gateway. Detected changes that can be updated in-place:
- domain
Update the gateway? [y/n]: y
NAME BACKEND HOSTNAME DOMAIN DEFAULT STATUS
test-gateway gcp (us-west4) 34.125.56.225 new.example.com running
```
Add new API methods:
- `/api/project/{project_name}/gateways/get_plan`
- `/api/project/{project_name}/gateways/apply`
Deprecate API methods:
- `/api/project/{project_name}/gateways/create`
- `/api/project/{project_name}/gateways/set_wildcard_domain`
Deprecate CLI arguments:
- `dstack gateway update --domain`
Fixes `AssertionError: assert '4 years ago' == '3 years ago'`
* shim(aws): retry EBS device resolution to tolerate attach latency * add debug log when retrying GetRealDeviceName
…stackai#4007) * Drop P100 in GCP and OCI * Expand image reference documentation * Update to cuda:13.0.3 * Set NCCL_VERSION=2.28.3-1 * Bump Verda OS image version * Fix runner/README.md * Set docker_base_image = "0.15"
Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com>
Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com>
* Refresh landing diagram and product popup Diagram: drop "Your data", give "Any model" real logos (GLM/DeepSeek/Qwen/ Kimi), relabel the middle band to "dstack orchestration" with a Docker mark, and switch to flat black fills + hairline borders (theme-aware, so it inverts in dark mode). Product popup / Get-started switcher: flat black tiles, tighter framing, Enterprise before Sky, and hover to switch. Mirror the same changes into the docs (mkdocs); the docs "Get started" popup is now click-to-open with matching external-link icons and button weight. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Invert docs "Get started" button with the theme The docs header primary button used a hardcoded rgba(0,0,0,.87) fill keyed on the primary color, so it stayed dark in dark mode and read as a void. Make it theme-aware (dark fill in light mode, light fill in dark mode) like the landing's primary button. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
We don't use `FullLoader` features
`memory: ..4GB` should be translated to `--mem=4096M`, not `--mem=8192M` (where 8GB is taken from `DEFAULT_MEMORY_SIZE.min`)
`cachetools<7.1` contstraint introduced in [1] doesn't help since pyright 1.1.410 (the last version it works is 1.1.409). This patch removes the version constraint, which no longer helps, and adds `# pyright: ignore` directive instead. [1]: dstackai#3846
) Most backend don't need it -- they provision machines/containers based on offers, which are already filtered against the effective Requirements, but both Slurm and Kubernetes backends completely ignore provided offers and rely on Requirements instead, which they previously took from the JobSpec, which is generated from the Run configuration only. Fixes: dstackai#4019
* Switch the backend from ComputeWithFilteredOffersCached to ComputeWithAllOffersCached. Kubernetes node offers don't depend on run requirements, so all node offers are now cached once and adjusted/filtered per requirements via get_offers_modifiers(), instead of re-querying the cluster for every distinct requirements set. * Emit a Kubernetes request of 0 when a resource range has no lower bound (e.g. `cpu: ..4`). Previously no request was emitted, and since Kubernetes defaults the request to the limit, `cpu: ..4` silently behaved like `cpu: 4`. An explicit 0 request is less surprising and matches how ranges behave elsewhere in dstack. * Introduce ResourceRequests/ResourceLimits dataclasses that centralize the translation between dstack ResourcesSpec and Kubernetes resource maps, and back (from_kubernetes_map). Previously this logic was duplicated and built ad hoc as dicts in _create_job_pod() and get_instance_offers(). * Add adjust_resources_by_resource_requests() to cap an offer's advertised resources to what was requested (used as the offer modifier), and to reflect the actual pod requests on the provisioned instance in run_job().
Handle termination errors to make use of `dstack`'s retry mechanism and avoid leaving orphan instances.
- For testing the auth token, use
`/api/v1/instances`.
- For finding an instance by ID, use
`/api/v0/instances/{id}`. On the upside, this
allows to issue only one request for a specific
instance, as opposed to traversing paginated
`/api/vX/instances` responses. On the downside,
if multiple instances are being provisioned at
once, each instance will require separate
`/api/v0/instances/{id}` requests, without an
option to reuse cache between instances. This
can result in hitting rate limits in
`VastAICompute.update_provisioning_data`, which
`dstack` easily recovers from by retrying on the
next pipeline tick.
* Add server access to tasks and dev environments * Update run response expectations * Preserve compatibility with older servers * Fix server access retry timeout for starting jobs Removing the job server connection on PROVISIONING/PULLING iterations reset the failure time tracked by the pool, so retry_timed_out() could never fire and jobs with a failing tunnel were never terminated by the server. Remove the connection only when the job leaves RUNNING; jobs_terminating covers the final cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Allow DSTACK_TOKEN to override the configured token DSTACK_TOKEN alone now overrides the token of the project configured in config.yml. DSTACK_SERVER_URL still takes effect only together with DSTACK_TOKEN, and a missing token or project now produces an explicit error instead of a silent fallback when no config is available. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Rename `server` to `dstack` in run configurations Enabling dstack access inside a run now reads like Docker-in-Docker (`docker: true`): `dstack: true`. No migration needed: the field is stored in run spec JSON only; pre-release DBs with the old field can be reset. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Rename docs section to `dstack` inside `dstack` Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Follow the server bind address for job server access The reverse forward targeted 127.0.0.1 while the server may be bound to a specific address via DSTACK_SERVER_HOST, making the forward target unreachable. Derive the target from the bind address and include it in the error when the server is not reachable through the tunnel. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Drop redundant pool lock in JobServerConnectionsPool _get_lock and remove_all only run await-free expressions, so the extra lock adds nothing under the single-threaded event loop. Make _get_lock synchronous and drop the lock. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Reject dstack access combined with inactivity_duration The persistent server connection counts as SSH activity, so a dev environment with both dstack access and inactivity_duration would never be detected as inactive. Reject the combination at configuration time instead of silently breaking inactivity_duration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Adopt an existing job server socket instead of overwriting it With multiple server replicas, each replica used to overwrite the job's /run/dstack/server.sock on every open, so ownership churned between replicas (they repeatedly stole the socket from each other) and every overwrite briefly broke access. Probe the socket first and (re)create the reverse forward only when it is missing or unreachable; otherwise adopt the existing one. This keeps a single stable owner. If the owner replica dies, another takes over on its next health check. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Support dstack access for services Allow `dstack: true` on service configurations, mirroring tasks and dev environments. Each replica is a separate job that gets its own `/run/dstack/server.sock`, so multiple replicas behave like multi-node job replicas. Exclude the field from run specs sent to older servers, same as for tasks and dev environments. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add DSTACK_FORBID_DSTACK_IN_RUNS server setting Let operators forbid `dstack: true` in run configurations, mirroring DSTACK_FORBID_SERVICES_WITHOUT_GATEWAY. When set, submitting a run with dstack access enabled is rejected server-side. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document dstack in dstack in Protips and note apply/attach Add a 'dstack in dstack' Protips section, and note in each dstack-access section that runs can also submit runs with dstack apply and attach to them with dstack attach, not just inspect them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Move dstack field to BaseRunConfiguration, simplify compat check The dstack field applied to exactly the run configuration types that extend BaseRunConfiguration, so define it there directly and drop the ConfigurationWithDstackParams mixin. The run spec exclude check then no longer needs the isinstance guard, since every run configuration has the field. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Simplify server access check --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Vast.ai offers interruptible (spot) instances, but the backend filtered
out every spot offer with an `extra_filter` and never bid on one, so only
on-demand instances could be provisioned.
Vast's `PUT /asks/{id}/` endpoint creates an interruptible instance when
the payload includes a `price` (per-machine bid in $/hour); omitting it
creates an on-demand instance. gpuhunt already emits Vast spot offers
(price set to the offer's `min_bid`), so no catalog changes are needed.
* Drop the `extra_filter` that removed spot offers in
`get_offers_by_requirements`.
* In `run_job`, bid the spot offer's price and pass it to
`create_instance`; on-demand offers pass no bid.
* Add the `price` field to the `create_instance` payload.
Interruption detection and retry are backend-agnostic (the server treats
a lost spot instance as `INTERRUPTED_BY_NO_CAPACITY`), so no further
changes are required.
- Detect and clean up stale spot instances that failed to start due to insufficient bid. - Update to gpuhunt 0.1.27, which includes better ordering and more precise pricing of spot offers. - Use backend-specific `min_bid` offer parameter introduced in gpuhunt 0.1.27.
Use `warning` instead of `error` to avoid reporting connection errors to Sentry. These errors can occur if the user deleted or stopped the gateway directly in the backend, in which case the error does not indicate a problem with `dstack`.
Detect spot interruptions and terminate the instance without waiting for provisioning timeout to elapse.
* Support OTel tracing * Fix otel.configure_tracing call site * Rename DB spans * Support OTel logs * Support OTel metrics * Track background task runs metric * Document Observability * Tests cleanup * Filter opentelemetry. logs
- Use a timeout in all requests to avoid requests hanging indefinitely. - Switch from URL-based to header-based auth token, which is the authentication method currently advertised in Vast.ai API docs.
* Add endpoint preset creation and reuse * Fix endpoint agent Windows launchers * Fix endpoint preset tests on Windows * Improve endpoint preset lifecycle * Fix formatting in endpoint preset list output * Improve endpoint preset apply output * Isolate endpoint interrupt test from SSH check * Refine endpoint preset terminology and output * Persist endpoint agent progress logs * Make endpoint output test width-independent * Document endpoint preset progress output * Preserve endpoint benchmark output in narrow terminals * Improve endpoint preset output and docs * Add endpoint preset creation time * Stabilize endpoint preset output test * Keep endpoint output test fix isolated * Add endpoint preset get command * Isolate endpoint CLI tests from SSH * Use existing Claude authentication by default * Document endpoint preset roadmap * Move endpoint preset code to CLI packages * Fix Windows type check for endpoint agent * Reuse planned endpoint preset configuration * Fix endpoint configuration reference --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com>
* Drop redundant `profile` argument * Flatten nested ifs * Eliminate dead code in helpers
* Set DSTACK_OTEL_METRICS_EXPORTERS=otlp by default * Unregister prometheus default collectors * Document OTEL_METRIC_EXPORT_INTERVAL
* CLI: add `--no-profile` option Disables default profile loading. Mutually exclusive with `--profile`. Supported by the following commands: * `dstack apply` (run configurations only) * `dstack preset create|apply` * `dstack offer` * Add `DSTACK_NO_PROFILE`
…i#4327) * Apply DSTACK_SERVICE_CLIENT_TIMEOUT to gateway nginx and app The server now passes the timeout to gateways on service and entrypoint registration. The default is raised from 60 to 300 seconds to match the previous nginx timeout. * Fix docstrings * Clean up tests
* Add Daytona backend * Shorten Daytona volume documentation * Validate Daytona API key permissions Document required permissions and remove redundant Daytona tests. * Pin gpuhunt 0.1.31 * Clarify Daytona volume sharing * Clarify Daytona volume documentation * Clean up Daytona docs, timeouts, and tests --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com>
Starting October 8, Nebius is introducing Dynamic Spot pricing for preemptible VMs. The price can change for existing instances during their lifetime. To prevent the price from increasing above the one `dstack` shows, `dstack` now caps the price at its current value at provisioning time using a pricing policy. If the price grows above the cap, the instance is interrupted by Nebius. The actual price may still drop below the one `dstack` shows. Preemptible L40S is still sold at a flat rate and is exempt from pricing policies. Full support for Nebius Dynamic Spot pricing (tracking the current price and showing it in `dstack`, user-specified caps, etc.) may be added later upon request.
Add an OCI backend setting that accepts a list of shape names to allow provisioning in addition to the standard supported instance families. Only works for shapes included in `dstack`'s pricing catalog (`gpuhunt`). This setting can be used to test various shapes that were not yet tested internally at `dstack`. ```yaml experimental_instance_types: [BM.GPU.RTXPRO.8] ```
* Add Hot Aisle bare metal support Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add a flag to release Hot Aisle bare metal without force Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Start Hot Aisle shim with nohup instead of disown The shim start is now retried until the SSH command succeeds, and `disown` fails in shells other than bash even though the shim starts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Connect to Hot Aisle bare metal via ssh_access The bare metal server's ip_address is private (100.64.0.0/10); SSH is exposed on ssh_access. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Don't fail Hot Aisle docs commands when nothing is in stock Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Tighten Hot Aisle minimum reservation note Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Silence curl progress in Hot Aisle docs commands Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Recommend fixed fleet nodes for prepaid Hot Aisle instances Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Mark Hot Aisle bare metal as experimental Keep the no-force flag enabled; drop it once offers carry the minimum reservation period and instances stay idle until it ends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make Hot Aisle bare metal opt-in via `bare_metal: true` Bare metal offers are suggested only if the backend sets `bare_metal: true`, like RunPod's `community_cloud`. Exclude the field from client requests when unset to keep older servers working. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Report `B300 PC` as `B300`. - Make the `gpu: B300` filter work reliably for both `B300 AC` and `B300 PC` variants. Before: ```shell $ dstack offer -b vastai --gpu B300 No matching instance offers available. Possible reasons: https://dstack.ai/docs/guides/troubleshooting/#no-offers ``` After: ```shell $ dstack offer -b vastai --gpu B300 # BACKEND RESOURCES INSTANCE TYPE PRICE 1 vastai (us-california) cpu=x86:256 mem=3095.7GB disk=25751.7GB gpu=B300:270GB:8 (spot) 53587615 $73.5377 2 vastai (us-california) cpu=x86:256 mem=3095.7GB disk=25751.7GB gpu=B300:270GB:8 53587615 $94.871 3 vastai (us-california) cpu=x86:128 mem=773.9GB disk=12875.8GB gpu=B300:270GB:4 (spot) 53587617 $36.7688 4 vastai (us-california) cpu=x86:128 mem=773.9GB disk=12875.8GB gpu=B300:270GB:4 53587617 $47.4355 5 vastai (hr-croatia) cpu=x86:43 mem=32.2GB disk=810.2GB gpu=B300:270GB:1 (spot) 53579010 $5.411 6 vastai (hr-croatia) cpu=x86:43 mem=32.2GB disk=810.2GB gpu=B300:270GB:1 53579010 $10.1485 7 vastai (hr-croatia) cpu=x86:86 mem=129GB disk=1620.5GB gpu=B300:270GB:2 (spot) 53579004 $10.822 8 vastai (hr-croatia) cpu=x86:86 mem=129GB disk=1620.5GB gpu=B300:270GB:2 53579004 $20.297 ``` The fix is in the `gpuhunt` package. This commit bumps the `gpuhunt` version in `dstack`.
* Update Azure GRID driver to 580.178.04 NVIDIA's 2026-09-30 security bulletin covers A10 vGPU. Azure has patched the NVadsA10_v5 hosts and now supports only the patched guest drivers: vGPU 19.6 (580.178.04) and vGPU 20.2. The 580 driver also adds CUDA 13.0 support on A10. With the previous 570 driver, CUDA 13 images failed with "driver too old" unless they bundled CUDA forward compatibility libraries, as dstack's base image does. This matches the 580 driver already used in dstack-cuda images. Also allow republishing an existing Azure image as a new gallery version, so dstack servers using dstack-grid-0.14 pick up the new driver without a release: - azure_variant input to build only some Azure image variants - azure_new_gallery_version input to replace the build image and publish the next gallery version; without it, publishing stops if a version exists Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Stop Azure image publish if gallery versions can't be listed Without this, a failed version listing looked like an empty gallery, so the script tried to publish 0.0.1 again instead of the next version. Also describe azure_new_gallery_version as a rebuild, since packer -force deletes and rebuilds the existing managed image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Andrey Cheptsov <andrey.cheptsov@github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Skip placement groups for AWS Capacity Blocks AWS rejects launches into Capacity Blocks that specify a placement group. Fixes dstackai#4344. * Drop AWS Capacity Block lookup cache and handle missing ReservationType AWS omits ReservationType for regular capacity reservations. * Use fleet reservation when provisioning runs
…i#4354) * Bound SSHTunnel control commands with timeouts and BatchMode * Drop ConnectTimeout from SSHTunnel control commands * Drop -n from SSHTunnel exec command and document control batch mode
(cherry picked from commit 49e8715)
(cherry picked from commit c140804)
(cherry picked from commit 3a5fd6e)
Adapted from fork commit f3cb27c (fix(plugin): remove example plugin). The pyproject.toml change (comment out the example plugin) is dropped: upstream now keeps dstack-plugin-server in its docs group, and docker/server/stgn/Dockerfile copies examples/ (fork commit 49e8715), so the editable plugin path resolves. Keeping pyproject.toml identical to upstream keeps uv.lock consistent for a reproducible image build. (cherry picked from commit f3cb27c, adapted) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…upstream Ports the net change of these fork commits to verda/compute.py onto the upstream file (per-instance startup scripts and SSH keys dstackai#3718, instance volumes dstackai#3758, Pydantic v2 dstackai#4077) with one 3-way merge, because a per-commit cherry-pick hit the same conflicts eight times: 792baf1 Add SFS (Shared File System) volume support for Verda/DataCrunch 530297b Fix shim to handle NFS volumes without block device (verda part) f714237 Fix _get_volume_by_id to handle SDK objects, not just dicts 9661390 Use direct HTTP request for volume API instead of SDK 5e27141 Add attach/detach volume methods for SFS volumes d5e7a61 Update Verda SFS volume support with actual API attach/detach 00987a4 Include user SSH key in addition to project SSH key e6212a7 Restore verbose logging and debug SSH key support Adaptations to upstream: - VerdaCompute keeps upstream ComputeWithInstanceVolumesSupport and adds ComputeWithVolumeSupport (register, attach, detach of SFS volumes). - run_job takes the upstream signature (requirements, extra_authorized_keys). The host keys are the project key plus extra_authorized_keys (upstream's source of the user key, was run_spec.ssh_key_pub). reservation comes from requirements. instance_offer.model_copy() (Pydantic v2). - create_instance keeps the upstream per-instance key and script lifecycle (cleanup on failure). The debug key from DSTACK_DEBUG_SSH_PUBLIC_KEY is a per-instance key too, so the cleanup removes it. The SFS NFS mount commands run before the shim commands, as in the fork. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Ports the net shim and server change of these fork commits: 530297b Fix shim to handle NFS volumes without block device (shim part) 526ee15 Add NFS mount support in shim for reused fleet instances Adaptations to upstream: - The volume code moved from runner/internal/shim/docker.go to runner/internal/shim/volumes.go upstream. The NFS branch goes at the top of formatAndMountVolume there. docker.go stays as upstream has it. - VolumeInfo (shim) and ShimVolumeInfo (server) gain nfs_host and nfs_pseudo. The server fills them from the Verda SFS volume backend_data (pseudo_path, location_code) in _volume_to_shim_volume_info. Checked: GOOS=linux go vet ./internal/shim/... clean, the shim builds for linux/amd64 and linux/arm64 with CGO_ENABLED=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Ports fork commit db1d670 (Fix UNIQUE constraint error when reattaching volume to same instance) to the upstream pipeline. The old process_submitted_jobs.py task is gone upstream. The attach moved to pipeline_tasks/jobs_submitted.py, which still inserts a second VolumeAttachmentModel row when a reused fleet instance already holds the volume, and the job fails with: UNIQUE constraint failed: volumes_attachments.volume_id, volumes_attachments.instance_id Now _process_volume_attachments takes the instance id on the existing-instance path. If the volume (loaded with its attachments) is already attached to that instance, it skips the backend attach call and returns a payload marked already_attached. The apply step inserts no row for it. The volume still counts in job_runtime_data.volume_names. Test first: test_reuses_existing_volume_attachment_on_existing_instance failed with the IntegrityError above (RED), passes now (GREEN). The other volume tests in test_submitted_jobs.py stay green. (cherry picked from commit db1d670, adapted) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Upstream supports the Nebius RTX PRO 6000 natively now (gpuhunt 0.1.32 maps gpu-rtx6000 to RTXPRO6000, and SUPPORTED_PLATFORMS has gpu-rtx6000). That supersedes the fork's live-offer injection and gpuhunt patch (99092f1, 66fc66f), which are not carried over. The deployed fork image also served the uk-south2 platform name gpu-rtx6000-a, which upstream does not list. gpuhunt maps it to RTXPRO6000 too, so one list entry is enough. Test first: TestSupportedInstances failed for gpu-rtx6000-a (RED), passes now (GREEN). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…g policy) Dry check for Nebius Dynamic Spot pricing (2026-10-08): builds the CreateInstanceRequest through resources.create_instance with the InstanceServiceClient mocked (no network). A spot create carries the recovery policy and the pricing_model oneof set to spot_pricing_policy with the given policy id. An on-demand create carries no pricing model. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…atures Tests for the fork features carried by the rebase: - _get_sfs_mount_commands: mounts a mapped SFS volume over NFS, skips an unmapped volume and a volume without backend data. - register_volume: rejects a non-shared volume, registers a shared one with its NFS backend data. - create_instance: the DSTACK_DEBUG_SSH_PUBLIC_KEY key is a per-instance key stored in the backend data (so the upstream cleanup removes it), no extra key when the env is unset, and the SFS mounts run before the shim. - run_job: the host gets the project key and the extra authorized keys. - _volume_to_shim_volume_info: nfs_host and nfs_pseudo from the SFS backend data, none for a block volume. Also reads instance_config.volumes with getattr, because upstream tests (and some callers) pass a config without volumes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ShimVolumeInfo.nfs_host and nfs_pseudo (fork, Verda SFS) serialized as null for every volume, which changed the shim submit body for all backends and broke TestShimClientV1::test_submit and TestShimClientV2::test_submit_task. They are excluded when None now, so a non-SFS volume body is identical to upstream. The Go side (VolumeInfo) already uses omitempty. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review BLOCKER: upstream turned the volume configuration into a union tagged by backend (Pydantic v2, d10edfd) with members for aws, gcp, runpod, kubernetes and daytona only. A Verda SFS volume (backend verda or datacrunch) then fails to parse in the CLI and the server, and the volume rows that the fork stored in the server DB fail to load. VerdaVolumeConfiguration takes region, volume_id and size like the Runpod member, accepts backend verda and datacrunch, and accepts (and drops) the availability_zone field that dstack 0.20.3 stored in each row. register_volume asserts the config type. Test first: test_verda_volume_configuration.py failed at import (RED), passes now (GREEN), including a row in the 0.20.3 shape. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review BLOCKER: upstream pyproject force-includes skills/*/SKILL.md in the wheel, so 'uv sync --extra all' fails in the stgn image build without them. Checked by building the wheel from the Dockerfile's copy set: without skills/ the build fails (rc 2), with it the build passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Operator decision 2026-10-08: the Nebius pricing policy caps a spot VM at the on-demand price of the same instance type in the same region (for example 1.79 USD/h for gpu-rtx6000 in us-central1), not at the current spot price. Nebius then preempts the VM only when the spot price grows above what the same VM costs on demand. The offer shows the cap, as upstream does, so a run's max_price still filters the offer before any create. Spot offers carry the matching on-demand price in their backend data. Without an on-demand row for that type and region, the cap stays the spot price (upstream behavior). Flat-rate spot offers keep no pricing policy. The cap and its source are logged at create time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…of-range Review fixes for aa99109 (review workflow wf_88a5d0da-f11, 3 lanes, a skeptic per finding): - BLOCKER: Nebius accepts a pricing policy bid only up to 0.01 USD per GPU-hour below the on-demand price (OUT_OF_RANGE above it). The cap is now the on-demand price per GPU minus 0.010, rounded down. It is never below the spot price cap. - MAJOR: the shown spot price is the cap, so the catalog order no longer held. get_all_offers_with_availability now returns the offers sorted by price (stable, so spot stays ahead of on-demand on a tie). The server merges the per-backend lists and expects each one sorted. - MINOR: a cap that Nebius refuses as out of range (for example after an on-demand price change that the offline catalog does not show yet) is retried once with the spot price cap, the upstream behavior. Spot offers keep their catalog spot price in backend data for that. - Operator fallback: a spot offer with no on-demand price takes the run's max_price as its cap (offer modifier). Its shown price does not change. Known effect, accepted: dstack shows and records the cap as the price of a Nebius spot VM, so run costs overstate the billed spot price. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fork moves to the upstream rebase (PR #1). The old master history stays reachable as the second parent. The tree is the rebased tree, unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Since 2026-10-08 Nebius sells preemptible VMs with Dynamic Spot pricing. A preemptible create needs a pricing model. The fork (dstack 0.20.3 + 19 fork commits) sends none, so Nebius spot runs no longer provision (for example
m002-labant-nebiusstays inno offers). Upstream fixed this in dstackai#4328 (release 0.22.2): each spot create gets its own Nebius pricing policy, capped at the price dstack shows, and the nebius SDK moves to>=0.6.13,<0.7.Operator decision: rebase the fork on upstream master (
9ebdd3f36, 0.22.2 + 8) and keep every fork feature.Fork commits: kept, adapted, skipped
extra_authorized_keys. The debug key is a per-instance key, so the upstream cleanup removes itrunner/internal/shim/volumes.go(upstream moved the code) and the server schema/client (a00c1e6). The NFS fields are sent only when set (933b852)pipeline_tasks/jobs_submitted.py, upstream still had the UNIQUE bug (a31d102, test first)gpu-rtx6000natively (gpuhunt 0.1.32). Addedgpu-rtx6000-a(uk-south2), which the deployed fork served (8910240)Review fixes on top:
VerdaVolumeConfigurationin the backend-tagged volume union, so Verda SFS volumes parse and the stored rows load (cdaa2ec).COPY skills skillsin the stgn Dockerfile, which upstream's pyproject needs (f9c4961).Tests
test_running_jobs.py::...instance_access_is_valid[pulling-sqlite]) also fails on clean upstream9ebdd3f36.spot_pricing_policy),gpu-rtx6000-a, Verda SFS (mount script, register, debug key, mount order, host keys), shim NFS fields, volume reattach, Verda volume config. Each written RED first.GOOS=linux go vet ./...clean, shim builds for linux/amd64 and linux/arm64.skills/(fails without it).--runpostgres), a real image build.Review
Review swarm, 5 lanes (fork features, Nebius spot path, DB migration and rollback, controlboard and shim compatibility, regression on other backends), a skeptic per BLOCKER/MAJOR: 3 BLOCKER (2 distinct, both fixed here), 13 MAJOR (operational, they go to the deploy plan: DB snapshot before the restart, a kept rollback image, never pass
DSTACK_VERSIONinto the container, the per-job log quota default, the CLI pin, the shim version plan, the uncapped spot VMs from the old line), 6 MINOR. 2 findings refuted.Deploy
Not part of this PR. The server restart runs 53 upstream DB migrations on Neon (one-way). The deploy plan is in the mission 019 comm folder and runs only on the operator's go.
🤖 Generated with Claude Code