> For the complete documentation index, see [llms.txt](https://documentation.immuta.com/saas/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.immuta.com/saas/~/changes/l3NnvynMHxi6VvqRtJhK/integrations/databricks-unity-catalog/how-to-guides/configure.md).

# Configure a Databricks Unity Catalog Integration

[Databricks Unity Catalog](https://docs.databricks.com/data-governance/unity-catalog/index.html) allows you to manage and access data in your Databricks account across all of your workspaces. With Immuta’s Databricks Unity Catalog integration, you can write your policies in Immuta and have them enforced automatically by Databricks across data in your Unity Catalog metastore.

This page details how to configure the integration. To configure the Databricks Unity Catalog integration and register data sources using the [simplified workflow](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/data-and-integrations/registering-a-host/simplified-registration-reference.md), see this [how-to guide](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/data-and-integrations/registering-a-host/how-to-guides/connect-unity-catalog.md).

## Permissions

Several different accounts are used to set up and maintain the Databricks Unity Catalog integration. The permissions required for each are outlined below.

* **Immuta account (required)**: This user configures the integration on the app settings page in Immuta. To access the app settings page, this user needs the following permission:
  * `APPLICATION_ADMIN` Immuta permission
* **Databricks service principal (required)**: This service principal is used continuously by Immuta to orchestrate Unity Catalog policies and maintain state between Immuta and Databricks. In the [automatic setup option](#automatic-setup), Immuta also uses this service principal to create the Immuta-managed catalog. This service principal needs the following Databricks privileges:
  * `CREATE CATALOG` privilege **on the Unity Catalog metastore**. *This is only required if you have Immuta* [*automatically configure*](#configure-the-databricks-unity-catalog-integration) *the integration in Databricks for you. If a separate user will run the Immuta script in Databricks to manually configure the integration, that Databricks user account needs this privilege instead.*
  * `OWNER` permission on the **Immuta catalog** you configure.
  * `OWNER` privilege **on one of the securables below** so that Immuta can administer Unity Catalog row-level and column-level security controls.
    * **on catalogs** with schemas and tables registered as Immuta data sources. This permission could also be applied by granting `OWNER` on a catalog to a Databricks group that includes the Immuta service principal to allow for multiple owners.
    * **on schemas** with tables registered as Immuta data sources.
    * **on all tables** registered as Immuta data sources - if the `OWNER` permission cannot be applied at the catalog- or schema-level. In this case, each table registered as an Immuta data source must individually have the `OWNER` permission granted to the Immuta service principal.
  * `USE CATALOG` and `USE SCHEMA` **on parent catalogs** and **schemas** of tables registered as Immuta data sources so that the Immuta service principal can `SELECT` and `MODIFY` securables within the parent catalog and schema.
  * `SELECT` and `MODIFY` **on all tables** registered as Immuta data sources so that the Immuta service principal can grant and revoke access to tables and apply Unity Catalog row- and column-level security controls.
  * For native query audit (optional)
    * `USE CATALOG` **on the `system` catalog**
    * `USE SCHEMA` **on the `system.access` schema**
    * `SELECT` **on the following system tables**:
      * `system.access.audit`
      * `system.access.table_lineage`
      * `system.access.column_lineage`
* **Databricks account (recommended)**: This user account can manually configure the integration in Databricks to create the Immuta-managed catalog. To do so, this account requires the following Databricks privileges:
  * `CREATE CATALOG` **on the Unity Catalog metastore**
  * `ACCOUNT ADMIN` **on the Unity Catalog metastore** for native query audit (optional)

## Requirements

Before you configure the Databricks Unity Catalog integration, ensure that you have fulfilled the following requirements:

* Unity Catalog [metastore created](https://docs.databricks.com/data-governance/unity-catalog/create-metastore.html) and attached to a Databricks workspace. Immuta supports configuring a single metastore for each configured integration, and that metastore may be attached to multiple Databricks workspaces.
* Unity Catalog enabled on your Databricks cluster or SQL warehouse. All SQL warehouses have Unity Catalog enabled if your workspace is attached to a Unity Catalog metastore. Immuta recommends linking a SQL warehouse to your Immuta tenant rather than a cluster for both performance and availability reasons.

{% hint style="info" %}
**Unity Catalog best practices**

Ensure your integration with Unity Catalog goes smoothly by following these guidelines:

* Use a Databricks SQL warehouse to configure the integration. Databricks SQL warehouses are faster to start than traditional clusters, require less management, and can run all the SQL that Immuta requires for policy administration. A serverless warehouse provides nearly instant startup time and is the preferred option for connecting to Immuta.
* Move all data into Unity Catalog before configuring Immuta with Unity Catalog. The default catalog used once Unity Catalog support is enabled in Immuta is the `hive_metastore`, which is not supported by the Unity Catalog native integration. Data sources in the Hive Metastore must be managed by the Databricks Spark integration. Existing data sources will need to be re-created after they are moved to Unity Catalog and the Unity Catalog integration is configured.
  {% endhint %}

## Migrate data to Unity Catalog

1. Ensure that all Databricks clusters that have Immuta installed are stopped and the Immuta configuration is removed from the cluster. Immuta-specific cluster configuration is no longer needed with the Databricks Unity Catalog integration.
2. Move all data into Unity Catalog before configuring Immuta with Unity Catalog. *Existing data sources will need to be re-created after they are moved to Unity Catalog and the Unity Catalog integration is configured.* If you don't move all data before configuring the integration, [metastore magic](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/integrations/databricks-spark/reference-guides/metastore-magic.md) will protect your existing data sources throughout the migration process.

## Create the Databricks service principal

In Databricks, [create a service principal](https://docs.databricks.com/en/admin/users-groups/service-principals.html#manage-service-principals-in-your-account) with the [privileges listed above](#permissions).

## Opt to enable native query audit for Unity Catalog

1. [Enable a system schema where the `<SCHEMA_NAME>` is `access`](https://docs.databricks.com/en/administration-guide/system-tables/index.html#enable-a-system-schema).
2. [Enable verbose audit logs in Unity Catalog](https://docs.databricks.com/en/administration-guide/account-settings/verbose-logs.html).
3. If you will configure the integration using the manual setup option, the Immuta script you will use includes the SQL statements for granting required privileges to the service principal, so you can skip this step and continue to the [manual setup section](#manual-setup). Otherwise, [manually grant the Immuta service principal access to the Databricks Unity Catalog system tables](https://docs.databricks.com/en/administration-guide/system-tables/index.html#grant-access-to-system-tables). For Databricks Unity Catalog audit to work, the service principal must have the following access at minimum:
   * `USE CATALOG` on the `system` catalog
   * `USE SCHEMA` on the `system.access` schema
   * `SELECT` on the following system tables:
     * `system.access.audit`
     * `system.access.table_lineage`
     * `system.access.column_lineage`

## Configure the Databricks Unity Catalog integration

{% hint style="info" %}
**Existing data source migration**: If you have existing Databricks data sources, complete these [migration](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/integrations/databricks-unity-catalog/how-to-guides/migrating.md) steps before proceeding.
{% endhint %}

You have two options for configuring your Databricks Unity Catalog integration:

* [Automatic setup](#automatic-setup): Immuta creates the catalogs, schemas, tables, and functions using the integration's configured service principal.
* [Manual setup](#manual-setup): Run the Immuta script in Databricks yourself to create the catalog. You can also modify the script to customize your storage location for tables, schemas, or catalogs. The user running the script must have the [Databricks privileges listed above](#permissions).

### Automatic setup

**Required permissions:** When performing an automatic setup, the Immuta service principal must have the [permissions listed above](#permissions).

1. Click the **App Settings** icon in the left sidebar.
2. Scroll to the **Global Integrations Settings** section and check the **Enable Databricks Unity Catalog support in Immuta** checkbox.
3. Click the **Integrations** tab.
4. Click **+ Add Native Integration** and select **Databricks Unity Catalog** from the dropdown menu.
5. Complete the following fields:
   * **Server Hostname** is the hostname of your Databricks workspace.
   * **HTTP Path** is the HTTP path of your Databricks cluster or SQL warehouse.
   * **Immuta Catalog** is the name of the catalog Immuta will create to store internal entitlements and other user data specific to Immuta. This catalog will only be readable for the Immuta service principal and should not be granted to other users. The catalog name may only contain letters, numbers, and underscores and cannot start with a number.
6. If using a proxy server with Databricks Unity Catalog, click the **Enable Proxy Support** checkbox and complete the **Proxy Host** and **Proxy Port** fields. The username and password fields are optional.
7. Opt to fill out the **Exemption Group** field with the name of a group in Databricks that will be excluded from having data policies applied and must not be changed from the default value. Create this account-level group for privileged users and service accounts that require an unmasked view of data before configuring the integration in Immuta.
8. Opt to scope the query audit ingestion by entering in Unity Catalog **Workspace IDs**. Enter a comma-separated list of the workspace IDs that you want Immuta to ingest audit records for. If left empty, Immuta will audit all tables and users in Unity Catalog.
9. [Unity Catalog query audit](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/detect-your-data/audit/reference-guides/query-audit-logs/databricks-uc.md) is enabled by default; you can disable it by clicking the **Enable Native Query Audit** checkbox. [Ensure you have enabled system tables in Unity Catalog and provided the required access to the Immuta service principal](#opt-to-enable-native-query-audit-for-unity-catalog).
   1. Configure the [audit frequency](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/detect-your-data/audit/reference-guides/query-audit-logs/databricks-uc.md#audit-frequency) by scrolling to **Integrations Settings** and find the **Unity Catalog Audit Sync Schedule** section.
   2. Enter how often, in hours, you want Immuta to ingest audit events from Unity Catalog as an integer between 1 and 24.
   3. Continue with your integration configuration.
10. Select your authentication method from the dropdown:
    * **Access Token**: Enter a **Databricks Personal Access Token**. This is the access token for the Immuta service principal. This service principal must have the [metastore privileges listed above](#permissions) for the metastore associated with the Databricks workspace. If this token is configured to expire, update this field regularly for the integration to continue to function.
    * **OAuth machine-to-machine (M2M)**:
      * **AWS Databricks**:
        * Follow [Databricks documentation to create a client secret](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html) for the Immuta service principal and assign this service principal the [privileges listed above](#permissions) for the metastore associated with the Databricks workspace.
        * Fill out the **Token Endpoint** with the full URL of the identity provider. This is where the generated token is sent. The default value is `https://<your workspace name>.cloud.databricks.com/oidc/v1/token`.
        * Fill out the **Client ID**. This is a combination of letters, numbers, or symbols, used as a public identifier and is the [client ID displayed in Databricks when creating the client secret for the service principal](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html#step-3-create-an-oauth-secret-for-a-service-principal).
        * Enter the **Scope** (string). The scope limits the operations and roles allowed in Databricks by the access token. See the [OAuth 2.0 documentation](https://oauth.net/2/scope/) for details about scopes.
        * Enter the **Client Secret** you created above. Immuta uses this secret to authenticate with the authorization server when it requests a token.
      * **Azure Databricks:**
        * Follow [Databricks documentation](https://learn.microsoft.com/en-us/azure/databricks/dev-tools/auth/oauth-m2m) to create a service principal within Azure and then populate to your Databricks account and workspace.
        * Assign this service principal the [privileges listed above](#permissions) for the metastore associated with the Databricks workspace.
        * Within Databricks, [create an OAuth client secret for the service principal](https://learn.microsoft.com/en-us/azure/databricks/dev-tools/auth/oauth-m2m#step-5-create-an-azure-databricks-oauth-secret-for-the-service-principal). This completes your Databricks-based service principal setup.
        * Within Immuta, fill out the **Token Endpoint** with the full URL of the identity provider. This is where the generated token is sent. The default value is `https://<your workspace name>.azuredatabricks.net/oidc/v1/token`.
        * Fill out the **Client ID**. This is a combination of letters, numbers, or symbols, used as a public identifier and is the [client ID displayed in Databricks when creating the client secret for the service principal](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html#step-3-create-an-oauth-secret-for-a-service-principal) (note that Azure Databricks uses the Azure SP Client ID; it will be identical).
        * Enter the **Scope** (string). The scope limits the operations and roles allowed in Databricks by the access token. See the [OAuth 2.0 documentation](https://oauth.net/2/scope/) for details about scopes.
        * Enter the **Client Secret** you created above. Immuta uses this secret to authenticate with the authorization server when it requests a token.
11. Click **Save**.

### Manual setup

**Required permissions:** When performing a manual setup, a service principal and a Databricks account must have the [permissions listed above](#permissions).

1. Click the **App Settings** icon in the left sidebar.
2. Scroll to the **Global Integrations Settings** section and check the **Enable Databricks Unity Catalog support in Immuta** checkbox.
3. Click the **Integrations** tab.
4. Click **+ Add Native Integration** and select **Databricks Unity Catalog** from the dropdown menu.
5. Complete the following fields:
   * **Server Hostname** is the hostname of your Databricks workspace.
   * **HTTP Path** is the HTTP path of your Databricks cluster or SQL warehouse.
   * **Immuta Catalog** is the name of the catalog Immuta will create to store internal entitlements and other user data specific to Immuta. This catalog will only be readable for the Immuta service principal and should not be granted to other users. The catalog name may only contain letters, numbers, and underscores and cannot start with a number.
6. If using a proxy server with Databricks Unity Catalog, click the **Enable Proxy Support** checkbox and complete the **Proxy Host** and **Proxy Port** fields. The username and password fields are optional.
7. Opt to fill out the **Exemption Group** field with the name of a group in Databricks that will be excluded from having data policies applied and must not be changed from the default value. Create this account-level group for privileged users and service accounts that require an unmasked view of data before configuring the integration in Immuta.
8. Opt to scope the query audit ingestion by entering in Unity Catalog **Workspace IDs**. Enter a comma-separated list of the workspace IDs that you want Immuta to ingest audit records for. If left empty, Immuta will audit all tables and users in Unity Catalog.
9. [Unity Catalog query audit](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/detect-your-data/audit/reference-guides/query-audit-logs/databricks-uc.md) is enabled by default; you can disable it by clicking the **Enable Native Query Audit** checkbox. [Ensure you have enabled system tables in Unity Catalog and provided the required access to the Immuta service principal](#opt-to-enable-native-query-audit-for-unity-catalog).
   1. Configure the [audit frequency](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/detect-your-data/audit/reference-guides/query-audit-logs/databricks-uc.md#audit-frequency) by scrolling to **Integrations Settings** and find the **Unity Catalog Audit Sync Schedule** section.
   2. Enter how often, in hours, you want Immuta to ingest audit events from Unity Catalog as an integer between 1 and 24.
   3. Continue with your integration configuration.
10. Select your authentication method from the dropdown:
    * **Access Token**: Enter a **Databricks Personal Access Token**. This is the access token for the Immuta service principal. This service principal must have the [metastore privileges listed above](#permissions) for the metastore associated with the Databricks workspace. If this token is configured to expire, update this field regularly for the integration to continue to function.
    * **OAuth machine-to-machine (M2M)**:
      * **AWS Databricks**:
        * Follow [Databricks documentation to create a client secret](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html) for the Immuta service principal and assign this service principal the [privileges listed above](#permissions) for the metastore associated with the Databricks workspace.
        * Fill out the **Token Endpoint** with the full URL of the identity provider. This is where the generated token is sent. The default value is `https://<your workspace name>.cloud.databricks.com/oidc/v1/token`.
        * Fill out the **Client ID**. This is a combination of letters, numbers, or symbols, used as a public identifier and is the [client ID displayed in Databricks when creating the client secret for the service principal](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html#step-3-create-an-oauth-secret-for-a-service-principal).
        * Enter the **Scope** (string). The scope limits the operations and roles allowed in Databricks by the access token. See the [OAuth 2.0 documentation](https://oauth.net/2/scope/) for details about scopes.
        * Enter the **Client Secret** you created above. Immuta uses this secret to authenticate with the authorization server when it requests a token.
      * **Azure Databricks:**
        * Follow [Databricks documentation](https://learn.microsoft.com/en-us/azure/databricks/dev-tools/auth/oauth-m2m) to create a service principal within Azure and then populate to your Databricks account and workspace.
        * Assign this service principal the [privileges listed above](#permissions) for the metastore associated with the Databricks workspace.
        * Within Databricks, [create an OAuth client secret for the service principal](https://learn.microsoft.com/en-us/azure/databricks/dev-tools/auth/oauth-m2m#step-5-create-an-azure-databricks-oauth-secret-for-the-service-principal). This completes your Databricks-based service principal setup.
        * Within Immuta, fill out the **Token Endpoint** with the full URL of the identity provider. This is where the generated token is sent. The default value is `https://<your workspace name>.azuredatabricks.net/oidc/v1/token`.
        * Fill out the **Client ID**. This is a combination of letters, numbers, or symbols, used as a public identifier and is the [client ID displayed in Databricks when creating the client secret for the service principal](https://docs.databricks.com/en/dev-tools/auth/oauth-m2m.html#step-3-create-an-oauth-secret-for-a-service-principal) (note that Azure Databricks uses the Azure SP Client ID; it will be identical).
        * Enter the **Scope** (string). The scope limits the operations and roles allowed in Databricks by the access token. See the [OAuth 2.0 documentation](https://oauth.net/2/scope/) for details about scopes.
        * Enter the **Client Secret** you created above. Immuta uses this secret to authenticate with the authorization server when it requests a token.
11. Select the **Manual** toggle and copy or download the script. You can modify the script to customize your storage location for tables, schemas, or catalogs.
12. Run the script in Databricks.
13. Click **Save**.

## Map Databricks users to Immuta

If the usernames in Immuta do not match usernames in Databricks, map each Databricks username to each Immuta user account to ensure Immuta properly enforces policies using one of the methods linked below:

* [Map the external IDs from an external identity manager](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/people/users-index/how-to-guides/external-user-mapping.md#configure-external-user-id-mapping-on-app-settings-page)
* [Manually map the external IDs on the user's profile page](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/people/users-index/how-to-guides/external-user-mapping.md#manually-configure-external-user-id-mapping-on-a-users-page)

If the Databricks user doesn't exist in Databricks when you configure the integration, [manually link their Immuta username to Databricks](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/people/users-index/how-to-guides/external-user-mapping.md#manually-configure-external-user-id-mapping-on-a-users-page) after they are created in Databricks. Otherwise, policies will not be enforced correctly for them in Databricks. Databricks user identities for Immuta users are automatically marked as invalid when the user is not found during policy application, preventing them from being affected by Databricks policy until their Immuta user identity is manually mapped to their Databricks identity.

## Opt to enable Databricks Unity Catalog tag ingestion

{% hint style="info" %}
**Design partner preview**

This feature is only available to select accounts. Reach out to your Immuta representative to enable this feature.
{% endhint %}

**Requirement**: A Databricks Unity Catalog integration must be configured for tags to be ingested.

To allow Immuta to automatically import table and column tags from Databricks Unity Catalog, enable Databricks Unity Catalog tag ingestion in the external catalog section of the Immuta app settings page.

1. Navigate to the **App Settings** page.
2. Scroll to **2 External Catalogs**, and click **Add Catalog**.
3. Enter a **Display Name** and select **Databricks Unity Catalog** from the dropdown menu.
4. Click **Save** and confirm your changes.

## Register data

[Register Databricks securables in Immuta](/saas/~/changes/l3NnvynMHxi6VvqRtJhK/data-and-integrations/registering-metadata/register-data-sources/databricks-tutorial.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://documentation.immuta.com/saas/~/changes/l3NnvynMHxi6VvqRtJhK/integrations/databricks-unity-catalog/how-to-guides/configure.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
