Repository navigation
[air] Add rank-partitioned container support - #6841
Merged
Merged
Conversation
Collaborator
Integration test reportCommit: 61f4615
Top 6 slowest tests (at least 2 minutes):
|
# Conflicts: # cmd/air/runsubmit.go # cmd/air/runsubmit_test.go
vinchenzo-db
marked this pull request as ready for review
September 30, 2026 21:08
vinchenzo-db
requested review from
ben-hansen-db,
maggiewang-db and
renaudhartert-db
September 30, 2026 21:09
vinchenzo-db
commented
Sep 30, 2026
ben-hansen-db
approved these changes
Oct 1, 2026
ben-hansen-db
left a comment
Contributor
There was a problem hiding this comment.
A couple of comments to address. Nothing blocking
| items = append(items, uploadItem{secretEnvVarsName, data}) | ||
| } | ||
|
|
||
| for _, container := range cfg.Containers { |
Contributor
There was a problem hiding this comment.
The backend looks for hyperparameters.yaml beside each container’s command, but this code uploads it only to the parent launch directory.
We'll need to upload the shared hyperparameters beside every container command.
| if err := validateCommand(*container.Command); err != nil { | ||
| return fmt.Errorf("%s.command: %w", prefix, err) | ||
| } | ||
| if len(container.Ranks) == 0 { |
Contributor
There was a problem hiding this comment.
Do we have validation for "containers requires at least two entries"?
Contributor
Author
There was a problem hiding this comment.
no need, with one container it will fan out across all nodes. The reason is because in the future, it makes sense to keep the container format and remove the "singleton" format.
vinchenzo-db
enabled auto-merge
October 1, 2026 21:28
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 1, 2026
deco-sdk-tagging Bot
added a commit
that referenced
this pull request
Oct 7, 2026
## Release v1.20.0 ### Notable Changes * Remove the Terraform deployment engine. `bundle.engine: terraform` and `DATABRICKS_BUNDLE_ENGINE=terraform` now error, and a failed migration of existing Terraform state is reported as an error instead of falling back to Terraform. To keep deploying with Terraform, use Databricks CLI v1.19.x. ([#6888](#6888), [#6889](#6889)) ### CLI * `databricks aitools install` now supports Kiro, installing Databricks agent skills into its skills directory. ([#6908](#6908)) * Fixed `databricks api` corrupting integers larger than 2^53 (such as job and pipeline ids) — request bodies and responses now preserve them exactly. ([#6884](#6884)) * Added `--auth-mode` and `--set <plugin>.<resourceKey>.authMode=obo|sp|both` to `databricks apps init` so AppKit resources can be accessed on behalf of the user, by the service principal, or both. The default stays service principal. ([#6886](#6886)) * `databricks apps init` now requires a value for every field a service principal resource binding references, prompting for missing values in an interactive terminal and otherwise failing with the `--set` key to use, instead of creating a project with unset variables. ([#6903](#6903)) * Add `databricks apps init --package-manager <npm|pnpm>` to select the package manager for Node.js templates. Infer the default quietly from template lockfiles and AppKit version, check prerequisites before creating files, and preserve template formatting and pnpm version pins. ([#6902](#6902)) * Select npm or pnpm from `packageManager` declarations and lockfiles for `apps validate` and project validation during `apps deploy`. ([#6892](#6892)) * Fix `auth docker host` reporting the credential helper as configured when its executable is missing from `PATH`. ([#6880](#6880)) * Warn when the CLI binary was built more than 6 months ago and recommend updating. ([#6898](#6898)) ### AI Runtime * Add an experimental rank-partitioned container images to AI Runtime jobs. ([#6841](#6841)) * Support snapshot fields directly under `code_source` without requiring `type` or a nested `snapshot` block. ([#6927](#6927)) * Map AIR priority and Unity Catalog image fields when converting run configurations to bundles. ([#6905](#6905)) * Add workspace backend validation to `air run --dry-run`. ([#6934](#6934)) ### Bundles * Warn that `bundle.terraform` is deprecated and has no effect since the Terraform deployment engine was removed. ([#6940](#6940)) * Direct engine now detects and applies an explicitly configured zero-value boolean or float (e.g. `gcp_attributes.use_preemptible_executors: false`, `azure_attributes.spot_bid_max_price: 0`) added to a resource first deployed without the field, matching the existing handling of an explicit integer zero. ([#6882](#6882)) * Fix `bundle deployment migrate` failing with "no such file or directory" when the Terraform state has no resources or the configuration no longer declares any of them. ([#6958](#6958)) * `bundle run` and `pipelines run` now send the per-update `development` parameter for pipelines in development mode targets. Setting `development` on a pipeline is deprecated and now emits a warning; use `mode: development` instead. ([#6863](#6863)) * Remove the hidden `bundle debug terraform` command. ([#6933](#6933)) * Add support for `run_as.group_name` at the bundle and target levels for jobs and pipelines. ([#6676](#6676)) * Fix recreating a secret scope that was deleted outside of the bundle with the direct deployment engine. ([#6970](#6970)) * Accept title-case booleans (`True`/`False`, as rendered by Azure Pipelines) for boolean variables, and accept the same boolean strings (`yes`/`no`, `on`/`off`, ...) in Python bundles as in YAML. ([#6942](#6942)) ### Dependency Updates * Bump `github.com/databricks/databricks-sdk-go` from v0.182.0 to v0.185.0. ([#6928](#6928))
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes
Add multi-container support in CLI yaml. Now CLI will support both singleton tasks, as well as multi-container/image multi-node tasks. Included an example yaml.
Top level env vars are merged into each container. Each container now has it's own definition of:
I've also added example docker builds for the RL usecase I'm covering, in case another user wants to reproduce it.
KNOWN GAP:
Jobs page will not show container level workloads because Jobs UI reads from task level command_path to find the training workload, and has no concept of containers yet.
Why
For primarily RL usecase, users may want to combine inference workload with training workload. We want to have a way to support this multi-node multi-image usecase
Tests
Unit tests
E2E test:
Replace placeholder images in example with my own DAR images.
Dry run:
Launch run:
Job run: https://dbc-04ac0685-8857.staging.cloud.databricks.com/jobs/runs/286153620017076
Sneak peak at exciting logs:
