Repository navigation
Add libraries field to the clusters resource - #6831
Merged
Merged
Conversation
Sankalp-Mittal
added this pull request to stack #6832
September 24, 2026 09:43
Sankalp-Mittal
marked this pull request as draft
September 24, 2026 09:45
Sankalp-Mittal
marked this pull request as ready for review
September 24, 2026 09:48
Collaborator
Integration test reportCommit: 30c009e
Top 50 slowest tests (at least 2 minutes):
|
Sankalp-Mittal
force-pushed
the
sankalp-mittal/cluster-libraries
branch
from
September 24, 2026 10:29
e3ba12e to
1a0467a
Compare
Restarting an all-purpose cluster kills attached sessions and running work, which is too disruptive to trigger implicitly on deploy. Installs apply live on a running cluster; uninstalls are marked UNINSTALL_ON_RESTART and take effect at the cluster's next restart. Co-authored-by: Isaac <no-reply@databricks.com>
The Jobs service installs a task's libraries on its existing all-purpose cluster via the Libraries API when a run starts, and they persist there. Mirror that so cluster-status reports them locally, like on cloud. This makes the integration_whl/interactive_* regression from #6365 reproduce without a real workspace. The direct engine now reads the job's wheel as cluster drift on redeploy, so interactive_cluster records its second deploy per engine. Co-authored-by: Isaac <no-reply@databricks.com>
Covers declaring, drifting and removing cluster libraries alongside an out-of-band library, and asserts that no change restarts the cluster. Co-authored-by: Isaac <no-reply@databricks.com>
The cluster declares libraries and a wheel job targets it via existing_cluster_id. Removing a declared library uninstalls it together with the job's libraries and does not restart the cluster. Co-authored-by: Isaac <no-reply@databricks.com>
A cluster without a libraries section now behaves as before the field existed: the plan does not read the cluster's libraries, so libraries installed by job runs or out of band are not drift, and no library calls are made. Removing the section stops managing the libraries and leaves them installed; `libraries: []` still uninstalls them. To make the read depend on the config, add an optional DoReadWithState resource method that receives the node's local state. Planning uses it when implemented; reads without a local state (deletes, the post-apply refresh for remote references) keep using DoRead. Co-authored-by: Isaac <no-reply@databricks.com>
interactive_cluster declares no cluster libraries, so it is back to its original output on both engines. libraries-scenarios checks that the plan reads no libraries without a section and that removing the section makes no library calls. libraries-restart empties the list instead of dropping it, so it still exercises an uninstall. Co-authored-by: Isaac <no-reply@databricks.com>
Replace the DoReadWithState framework method from 5d0fe69 with a cluster-only override, as suggested in review. The cluster reads its libraries unconditionally again, and when the config has no libraries section the whole-field change is skipped with the new reason `unmanaged`: no drift, no "Updated clusters.*" line, and no install or uninstall calls. Removing the section also leaves the libraries installed and unmanaged, while `libraries: []` stays managed. config-remote-sync ignores skips other than remote_addition, so unmanaged libraries are not synced into the config. Co-authored-by: Isaac <no-reply@databricks.com>
The plan now reads the cluster's libraries without a section and skips them with reason `unmanaged`, so the step records the skipped change instead of asserting that no library read happens. Co-authored-by: Isaac <no-reply@databricks.com>
andrewnester
previously approved these changes
Sep 28, 2026
andrewnester
self-requested a review
September 28, 2026 09:45
| // and running work. Installs apply live on a running cluster; uninstalls are marked | ||
| // UNINSTALL_ON_RESTART and take effect at the cluster's next restart. | ||
| // A cluster edit restarts the cluster on its own, so only wait when there was no edit. | ||
| if !edited { |
Contributor
There was a problem hiding this comment.
Why do we need to do it?
Contributor
Author
There was a problem hiding this comment.
this is a remanant from the old restart design I'll refactor the code to remove all of these
andrewnester
approved these changes
Sep 28, 2026
andrewnester
left a comment
Contributor
There was a problem hiding this comment.
LGTM, please clean up the code from old restart logic before merging
The bundle no longer restarts clusters for library changes, so drop what only existed for that design: the `edited` gate that skipped waiting for installs when a cluster edit restarted the cluster, the restart reasoning in the DoUpdate and WaitAfterCreate comments, and the testserver's clusters/restart handler. waitForInstall now always runs after reconciling; it already returns early when the cluster is not running. Co-authored-by: Isaac <no-reply@databricks.com>
Rename libraries-restart to libraries-update and check the install and uninstall requests instead of restart requests. Remove the //clusters/restart checks, the "!Restarting cluster" checks for a log line that no longer exists, and the "no restart" wording elsewhere. With the testserver restart route gone, an accidental restart call now fails the tests locally. Co-authored-by: Isaac <no-reply@databricks.com>
deco-sdk-tagging Bot
added a commit
that referenced
this pull request
Sep 30, 2026
## Release v1.19.0 ### CLI * Honor `CLAUDE_CONFIG_DIR` in `aitools` commands. ([#6838](#6838)) * Added `--ttl` and `--no-expiry` flags to `databricks postgres create-branch` so a branch's expiration can be set without hand-writing a `--json` spec. `--ttl` accepts the REST API duration form (`604800s`), a Go duration (`168h`), or day/week units (`7d`, `3w`); `--no-expiry` creates a branch that never expires. One of `--ttl`, `--no-expiry`, or a spec expiration in `--json` is required. ([#6313](#6313)) * `databricks ssh connect` serverless sessions now provide Claude Code and Codex configured with Unity Gateway out of the box. ([#6885](#6885)) ### AI Runtime * Add `databricks air images push` (Preview) to configure Docker authentication and push container images to Databricks Artifact Registry. ([#6869](#6869)) * `air run` now grants the configured `permissions` on the MLflow experiment as well as the job. ([#6870](#6870)) ### Bundles * Add libraries field to clusters. ([#6831](#6831)) * Error out when a configured `workspace_id` does not match the connected workspace, instead of silently using it in resource URLs emitted by `bundle summary`. ([#6754](#6754)) * Fix spurious recreation of Lakebase (Postgres) branches, roles, and catalogs when the referenced project is updated in place: an in-place project change (e.g. `display_name`) no longer forces a delete + create of resources that reference the project's or branch's `name`. ([#6865](#6865)) * Direct engine now detects and applies an explicitly configured integer zero (e.g. `gcp_attributes.local_ssd_count: 0`) added to a resource first deployed without the field. ([#6867](#6867)) * Migrate existing Terraform deployment state to the direct engine before deploying (previously done after a Terraform deploy), so the deploy runs on the direct engine. ([#6749](#6749)) * The `postgres_snapshot_schedules` resource (introduced in [v1.16.0](https://github.com/databricks/cli/releases/tag/v1.16.0)) is now marked Beta and is no longer available in PyDABs, matching the other `postgres_*` resources; configure it in YAML instead. ([#6887](#6887))
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes
Re-lands #6365 (reverted in #6830) with the regression fix folded in. Adds
resources.clusters.<name>.librariesin the direct engine: libraries are installed on all-purpose clusters through the Libraries API (install,uninstall,cluster-status) as part of the cluster lifecycle.Compared to #6365:
librariessection. Without one, the cluster behaves as before the field existed: libraries installed by job runs or out of band are not drift, there is noUpdated clusters.*, and no install or uninstall calls are made. The plan still readscluster-statusand records the libraries as skipped with the new reasonunmanaged(done in the cluster'sOverrideChangeDesc).UNINSTALL_ON_RESTARTby the backend and take effect at the cluster's next restart.existing_cluster_idcluster, so this class of regression now fails locally instead of only in the CloudSlow nightly.Library behaviour by config:
librariessectionunmanagedlibrariessection removedlibraries: []Why
#6365 regressed the nightly
integration_whl/interactive_*tests. A job task on an existing cluster installs its libraries cluster-wide, and they persist after the run. The direct engine read them as drift even on clusters that declared no libraries, restarted the cluster on every deploy, and uninstalled them.Decisions, from the team discussion and review with @andrewnester:
librariessection must behave as before: no change detection, no uninstalls, no restarts. Change detection and uninstall calls happen only when the section is used.This replaces #6791, which is now closed; its testserver change and tests moved here.
How to review
git range-diff 76872b50a^..76872b50a 55f0c1626..1a0467ad3. Apart from the rewritten commit message and one surrounding line inRemapStatethat changed on main, the only difference is the PR number in the changelog fragment.libraries-scenariosacceptance test (local only)interactive_cluster_declared_libscloud testDoReadWithState, superseded by 7)OverrideChangeDesc(replaces theDoReadWithStateframework method from 5, per review)libraries-scenariosupdate for theunmanagedskipeditedgate aroundwaitForInstall(it now always runs after reconciling), restart comments, and the testserverclusters/restarthandlerlibraries-restarttolibraries-updateintegration_whl/interactive_clusteris back to its original script and output: it declares no cluster libraries, so the redeploy after a job run is2 unchangedon both engines, as before #6365.Tests
TestClusterOverrideChangeDescLibraries: no section and removed section areunmanaged;libraries: [], declared libraries and per-library removals stay managed.acceptance/bundle/resources/clusters/libraries-scenarios(local): no section (out-of-band library skipped asunmanaged), then declare, drift and remove the section alongside an out-of-band library. It covers theREADPLANpath.acceptance/bundle/integration_whl/interactive_cluster_declared_libs(Cloud + CloudSlow, both DMS variants): cluster-declared libraries plus a wheel job on the same cluster.resources/clusters/libraries-update(renamed fromlibraries-restart, Cloud + CloudSlow): checks the install request on create and when adding a library, and the uninstall request forlibraries: [].clusters/restart, so an accidental restart call fails the tests locally instead of needing explicit "no restart" checks.integration_whltree andresources/clusters/libraries*pass on AWS cloud (run before commits 9–10).This pull request and its description were written by Isaac.