Skip to content

[air] Grant MLflow experiment permissions on submit - #6870

Merged
ben-hansen-db merged 6 commits into
mainfrom
air/mlflow-experiment-permissions
Sep 29, 2026
Merged

ben-hansen-db merged 6 commits into
mainfrom
air/mlflow-experiment-permissions

Conversation

@ben-hansen-db

@ben-hansen-db ben-hansen-db commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

air run previously granted the configured permissions only on the underlying job. The MLflow experiment created by the AI Runtime backend was left with default permissions, so naming a collaborator in permissions gave them access to the job but not to the experiment.

This change keeps job permissions set remotely and additionally grants the same ACLs on the experiment from the client:

  • After submit, the run is resolved (already done for job permissions) and its full experiment name is read from GenAiComputeTask.MlflowExperimentName.
  • The experiment is resolved by name (get-by-name) and granted the configured ACLs via the experiment permissions API.
  • Job permission levels are mapped onto their nearest experiment equivalents, since experiments only support CAN_READ/CAN_EDIT/CAN_MANAGE:
    • CAN_VIEW → CAN_READ
    • CAN_MANAGE_RUN → CAN_EDIT
    • CAN_MANAGE / IS_OWNER → CAN_MANAGE
    • levels with no experiment equivalent are skipped.
  • The grant is best-effort and time-bounded (short timeout, poll get-by-name): a not-yet-resolvable experiment never delays or fails an otherwise-successful submit — it warns and moves on, mirroring the existing job-permission fallback.

Testing

`air run` only granted the configured `permissions` on the underlying job.
The MLflow experiment created by the AI Runtime backend was left with default
permissions, so a grant that named a collaborator gave them the job but not the
experiment.

Job permissions are still set remotely; the experiment is now resolved by name
(get-by-name) and granted the same ACLs client-side. Job permission levels are
mapped onto their nearest experiment equivalents (CAN_VIEW->CAN_READ,
CAN_MANAGE_RUN->CAN_EDIT, CAN_MANAGE/IS_OWNER->CAN_MANAGE). The grant is
best-effort and time-bounded so a not-yet-created experiment never delays or
fails an otherwise-successful submit.

Co-authored-by: Isaac <no-reply@databricks.com>
@github-actions github-actions Bot added the AIR Databricks AI Runtime CLI label Sep 28, 2026
ben-hansen-db and others added 4 commits September 28, 2026 22:07
Co-authored-by: Isaac <no-reply@databricks.com>
Avoids allocating a fresh timer per poll iteration that outlives a
ctx-cancelled wait.

Co-authored-by: Isaac <no-reply@databricks.com>
Move the permission grant out of submitWorkload to the caller, after the
"Submitted workload" result is printed, and wrap the bounded best-effort grant
in a "Granting permissions…" spinner (text mode). The success line is shown
immediately; the experiment resolve/grant runs behind the spinner instead of
delaying it.

Co-authored-by: Isaac <no-reply@databricks.com>
The AI Runtime backend creates the experiment at submit time, so get-by-name
resolves almost immediately. Tighten the poll to a 100ms interval capped at 2s,
just enough to cover a brief read-after-write consistency race.

Co-authored-by: Isaac <no-reply@databricks.com>
@eng-dev-ecosystem-bot

eng-dev-ecosystem-bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Integration test report

Commit: 9cfb317

Run: 36637244013

Env ✅​pass 🙈​skip Time
✅​ aws linux 276 16 5:31
✅​ aws windows 278 14 5:53
✅​ azure linux 275 16 5:35
✅​ azure windows 277 14 3:24
✅​ gcp linux 276 16 5:26
✅​ gcp windows 278 14 3:12
Top 6 slowest tests (at least 2 minutes):
duration env testname
5:51 aws windows TestAccept
3:58 gcp linux TestAccept
3:56 azure linux TestAccept
3:55 aws linux TestAccept
3:22 azure windows TestAccept
3:11 gcp windows TestAccept

@@ -0,0 +1 @@
* `air run` now grants the configured `permissions` on the MLflow experiment as well as the job. ([#6870](https://github.com/databricks/cli/pull/6870))

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want this changelog entry?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How about changing it to: Improve AIR workload sharing through the permissions field in YAML

@ben-hansen-db
ben-hansen-db marked this pull request as ready for review September 29, 2026 16:19
@ben-hansen-db
ben-hansen-db requested review from a team as code owners September 29, 2026 16:19
Comment thread cmd/air/runpermissions.go Outdated
if len(run.Tasks) == 0 {
return ""
}
task := run.Tasks[0].GenAiComputeTask

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you add a unit test for applySubmittedPermissions where GetRun returns an AiRuntimeTask instead of a GenAiComputeTask, and assert that experiment permissions are updated?

@ben-hansen-db ben-hansen-db Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated! Looks like this caught a bug

A submitted AIR run is an ai_runtime_task, so GetRun returns
ai_runtime_task, not gen_ai_compute_task. The previous code read the
experiment name off GenAiComputeTask, so on the real submit path it found
nothing and silently skipped the experiment grant; the acceptance test only
passed because it stubbed a gen_ai_compute_task.

Resolve the run's own MLflow experiment id from runs/get-output
(ai_runtime_task_output.mlflow_experiment_id) instead. This is the run's
actual experiment, so the grant can never touch a different user's
experiment — and it avoids reconstructing a /Users/<user>/<name> path (which
has no default for a service principal). Drops the get-by-name lookup.

Adds a test covering applySubmittedPermissions when GetRun returns an
ai_runtime_task, asserting both job and experiment permissions are granted,
and fixes the acceptance fixture to the real ai_runtime_task + get-output
shape.

Co-authored-by: Isaac <no-reply@databricks.com>
@ben-hansen-db
ben-hansen-db added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit 32efeaf Sep 29, 2026
31 checks passed
@ben-hansen-db
ben-hansen-db deleted the air/mlflow-experiment-permissions branch September 29, 2026 22:54
deco-sdk-tagging Bot added a commit that referenced this pull request Sep 30, 2026
## Release v1.19.0

### CLI

 * Honor `CLAUDE_CONFIG_DIR` in `aitools` commands. ([#6838](#6838))
 * Added `--ttl` and `--no-expiry` flags to `databricks postgres create-branch` so a branch's expiration can be set without hand-writing a `--json` spec. `--ttl` accepts the REST API duration form (`604800s`), a Go duration (`168h`), or day/week units (`7d`, `3w`); `--no-expiry` creates a branch that never expires. One of `--ttl`, `--no-expiry`, or a spec expiration in `--json` is required. ([#6313](#6313))
 * `databricks ssh connect` serverless sessions now provide Claude Code and Codex configured with Unity Gateway out of the box. ([#6885](#6885))

### AI Runtime

 * Add `databricks air images push` (Preview) to configure Docker authentication and push container images to Databricks Artifact Registry. ([#6869](#6869))
 * `air run` now grants the configured `permissions` on the MLflow experiment as well as the job. ([#6870](#6870))

### Bundles

 * Add libraries field to clusters. ([#6831](#6831))
 * Error out when a configured `workspace_id` does not match the connected workspace, instead of silently using it in resource URLs emitted by `bundle summary`. ([#6754](#6754))
 * Fix spurious recreation of Lakebase (Postgres) branches, roles, and catalogs when the referenced project is updated in place: an in-place project change (e.g. `display_name`) no longer forces a delete + create of resources that reference the project's or branch's `name`. ([#6865](#6865))
 * Direct engine now detects and applies an explicitly configured integer zero (e.g. `gcp_attributes.local_ssd_count: 0`) added to a resource first deployed without the field. ([#6867](#6867))
 * Migrate existing Terraform deployment state to the direct engine before deploying (previously done after a Terraform deploy), so the deploy runs on the direct engine. ([#6749](#6749))
 * The `postgres_snapshot_schedules` resource (introduced in [v1.16.0](https://github.com/databricks/cli/releases/tag/v1.16.0)) is now marked Beta and is no longer available in PyDABs, matching the other `postgres_*` resources; configure it in YAML instead. ([#6887](#6887))
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AIR Databricks AI Runtime CLI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants