1 of 3

Data Sources

Payload Attribute Details

Audience: Data Engineers
Content Summary: This page contains details and examples of payload attributes for creating data sources.

`connectionKey`

The connectionKey is a unique identifier for the collection of data sources being created. If an existing connectionKey is used with new connection information, it will delete the old data sources and create new ones from the new information in the payload.

`connection`

Attribute

Description

Special Cases

Athena: Also requires region and queryResultLocationBucket. queryResultLocationDirectory is optional. authenticationMethod can be none, accessKey (default: username = access key, password = secret key), or instanceRole.
BigQuery: Does not require hostname and password. Requires sid, which is the GCP project ID, and userFiles with the keyName of KeyFilePath and the base64-encoded keyfile.json.
Databricks: Also requires httpPath. No username is required.
Hadoop: authenticationMethod can be none, userPassword, hdInsight, kerberos, or kerberosHdInsight.
Trino: authenticationMethod can be No Authentication, LDAP Authentication, or Kerberos Authentication.
Snowflake: Also requires warehouse. authenticationMethod can be userPassword or PRIV_KEY_FILE. If using PRIV_KEY_FILE, do not specify a password; userFiles is required with the keyName of PRIV_KEY_FILE and the base64-encoded Snowflake key.

`nameTemplate`

Available templates include

<tablename>
<schema>
<database>

All cases of the name in Immuta should be lowercase.

For example, consider a table TPC.CUSTOMER that is given the following nameTemplate:

dataSourceFormat: <schema> <tablename>
tableFormat: <tablename>
schemaFormat: <schema>
schemaProjectNameFormat: <schema>

This nameTemplate will produce a data source named tpc.customer in a schema project named tpc.

`options`

`owners`

`sources`

Best practice: Use Subscription Policies to Control Access

If you are not tagging individual columns, omit sources to create data sources for all tables in the schema or database, and then use Subscription Policies to control access to the tables instead of excluding them from Immuta.

This attribute configures which sources are created. If sources is not provided, all sources from the given connection will be created.

There are 3 types of sources than can be specified:

all tables
query
table

Recommended: Specify All Tables

If you specify any sources (either tables or queries), but you still want to create data sources for the rest of the tables in the schema or database, you can specify all as a source:

sources:
  - all: true

Best practice: Use schema monitoring

Excluding sources or specifying all: true will turn on automatic schema monitoring in Immuta. As tables are added or removed, Immuta will look for those changes on a schedule (by default, once a day) and either disable or delete data sources for removed tables or create data sources for new tables. New tables will be tagged New so that you can build a policy to restrict access to new tables until they are evaluated by data owners. Data owners will be notified of new tables, and all subscribers will be notified if data sources are disabled or deleted.

Specify a Query

Immuta recommends creating a view in your native database instead of using this option, but if that is not possible, you can create data sources based on SQL statements:

sources:
  - query: “select * from table”
    naming:
      datasource: “My Source”,
      table: “my_source”,
      schema: “queries”

Specify a Table

If you want to select specific tables to be created as data sources, or if you want to tag individual data sources or columns within a data source, you need to leverage this parameter:

sources:
  - table: name_of_table
    schema: name_of_schema

Additional Options

When specifying a table or query there are other options that can be specified:

Columns

If any columns are specified, those are the only columns that will be available in the data source.
If no columns are specified, Immuta will look for new or removed columns on a schedule (by default, once a day) and add or remove columns from the data sources automatically as needed.
New columns will be tagged New, so you can build a policy to automatically mask new columns until they are approved.
Data Owners will be notified when columns are added or removed.

columns is an array of objects for each column:

Column Descriptions

You can add descriptions to columns without having to specify all the columns in the data source. columnDescriptions is an array of objects with the following schema:

columnDescriptions:
  - columnName: acct_num
    description: The account number

Request Payload Examples

Audience: Data Engineers
Content Summary: This page contains example request payloads for creating data sources.

Basic Data Source

connectionKey: my-databricks
connection:
    hostname: your.databricks.hostname.com
    port: 443
    ssl: true
    database: tpc
    username: token
    password: "${DATABRICKS_PASSWORD}"
    httpPath: sql/protocolv1/o/0/11101101
    handler: Databricks

Data Source (with More Options)

connectionKey: my-databricks
nameTemplate:
  dataSourceFormat: Databricks <Tablename>
  tableFormat: <tablename>
  schemaFormat: databricks
connection:
  hostname: your.databricks.hostname.com
  port: 443
  ssl: true
  database: data
  username: token
  password: "${DATABRICKS_PASSWORD}"
  httpPath: sql/protocolv1/o/0/1110-11123
  handler: Databricks
sources:
  - table: credit_card_transactions
    schema: data
    tags:
      table:
        - PCI
        - SENSITIVE
      columns:
      - columnName: transaction_date
        tags:
          - PCI
          - DATE
  - table: crime_data
    schema: data
    naming:
        datasource: Crime Data
        table: crime_data
        schema: databricks

Databricks Data Source (Override Naming Convention)

connectionKey: ebock-databricks
nameTemplate:
  dataSourceFormat: Databricks <Tablename>
  tableFormat: <tablename>
  schemaFormat: databricks
connection:
  hostname: your.databricks.hostname.com
  port: 443
  ssl: true
  database: ebock
  username: token
  password: "${DATABRICKS_PASSWORD}"
  httpPath: sql/protocolv1/o/0/1110-185737-wove
  handler: Databricks
sources:
  - table: credit_card_transactions
    schema: ebock
  - table: crime_data_delta
    schema: ebock
    naming:
        datasource: Crime Data
        table: crime_data
        schema: databricks
  - table: hipaa_data
    schema: ebock

Impala Data Source (with `userFile`)

connectionKey: cdh-impala
nameTemplate:
  dataSourceFormat: Impala <Tablename>
  tableFormat: <tablename>
  schemaFormat: impala
connection:
  hostname: your.hadoop.hostname.example.com
  port: 21050
  ssl: true
  database: default
  handler: Apache Impala
  authenticationMethod: kerberos
  username: usera
  userFiles:
    - keyName: TrustedCerts
      content: <Base64 encoded contents of file go here>
      userFilename: tls-ca-bundle.pem
sources:
  - table: medical_records_parquet
    schema: default
  - table: nyc_taxi_fare_parquet
    schema: default
  - table: nyc_taxi_trip_parquet
    schema: default

Redshift Spectrum Data Sources

Your nativeSchemaFormat must contain _immuta to avoid schema name conflicts.

connectionKey: redshift
connection:
  hostname: your-redshift-cluster.djie25k.us-east-1.redshift.amazonaws.com
  port: 5439
  ssl: true
  database: your_database_with_external_schema
  username: awsuser
  password: your_password
  handler: Redshift
  schema: external_schema
nameTemplate:
  dataSourceFormat: <Tablename>
  schemaFormat: <schema>
  tableFormat: <tablename>
  schemaProjectNameFormat: <Schema>
  nativeSchemaFormat: <schema>_immuta
  nativeViewFormat: <tablename>
sources:
  - all: true

Snowflake Data Source (Specify Sources)

connectionKey: tpc-snowflake
nameTemplate:
  dataSourceFormat: Snowflake <Tablename>
  tableFormat: <tablename>
  schemaFormat: snowflake
connection:
  hostname: example.hostname.snowflakecomputing.com
  port: 443
  ssl: true
  database: TPC
  username: USERA
  password: "${SNOWFLAKE_PASSWORD}"
  schema: PUBLIC
  warehouse: IT_WH
  handler: Snowflake
sources:
  - table: CASE
    schema: PUBLIC
  - table: CASE2
    schema: PUBLIC
  - table: CUSTOMER
    schema: PUBLIC
  - table: WEB_SALES
    schema: PUBLIC

Payload Attribute Details

Audience: Data Engineers
Content Summary: This page contains details and examples of payload attributes for creating data sources.

`connectionKey`

`connection`

Attribute

Description

Special Cases

Athena: Also requires region and queryResultLocationBucket. queryResultLocationDirectory is optional. authenticationMethod can be none, accessKey (default: username = access key, password = secret key), or instanceRole.
BigQuery: Does not require hostname and password. Requires sid, which is the GCP project ID, and userFiles with the keyName of KeyFilePath and the base64-encoded keyfile.json.
Databricks: Also requires httpPath. No username is required.
Hadoop: authenticationMethod can be none, userPassword, hdInsight, kerberos, or kerberosHdInsight.
Trino: authenticationMethod can be No Authentication, LDAP Authentication, or Kerberos Authentication.
Snowflake: Also requires warehouse. authenticationMethod can be userPassword or PRIV_KEY_FILE. If using PRIV_KEY_FILE, do not specify a password; userFiles is required with the keyName of PRIV_KEY_FILE and the base64-encoded Snowflake key.

`nameTemplate`

Attribute

Description

Available templates include

<tablename>
<schema>
<database>

All cases of the name in Immuta should be lowercase.

For example, consider a table TPC.CUSTOMER that is given the following nameTemplate:

dataSourceFormat: <schema> <tablename>
tableFormat: <tablename>
schemaFormat: <schema>
schemaProjectNameFormat: <schema>

This nameTemplate will produce a data source named tpc.customer in a schema project named tpc.

`options`

Attribute

Description

`owners`

Attribute

Description

`sources`

Best practice: Use Subscription Policies to Control Access

This attribute configures which sources are created. If sources is not provided, all sources from the given connection will be created.

There are 3 types of sources than can be specified:

all tables
query
table

Recommended: Specify All Tables

If you specify any sources (either tables or queries), but you still want to create data sources for the rest of the tables in the schema or database, you can specify all as a source:

sources:
  - all: true

Best practice: Use schema monitoring

Specify a Query

Immuta recommends creating a view in your native database instead of using this option, but if that is not possible, you can create data sources based on SQL statements:

sources:
  - query: “select * from table”
    naming:
      datasource: “My Source”,
      table: “my_source”,
      schema: “queries”

Specify a Table

If you want to select specific tables to be created as data sources, or if you want to tag individual data sources or columns within a data source, you need to leverage this parameter:

sources:
  - table: name_of_table
    schema: name_of_schema

Additional Options

When specifying a table or query there are other options that can be specified:

Option

Description

Columns

If any columns are specified, those are the only columns that will be available in the data source.
If no columns are specified, Immuta will look for new or removed columns on a schedule (by default, once a day) and add or remove columns from the data sources automatically as needed.
New columns will be tagged New, so you can build a policy to automatically mask new columns until they are approved.
Data Owners will be notified when columns are added or removed.

columns is an array of objects for each column:

Attribute

Description

Column Descriptions

You can add descriptions to columns without having to specify all the columns in the data source. columnDescriptions is an array of objects with the following schema:

Attribute

Description

columnDescriptions:
  - columnName: acct_num
    description: The account number

Request Payload Examples

Audience: Data Engineers
Content Summary: This page contains example request payloads for creating data sources.

Basic Data Source

connectionKey: my-databricks
connection:
    hostname: your.databricks.hostname.com
    port: 443
    ssl: true
    database: tpc
    username: token
    password: "${DATABRICKS_PASSWORD}"
    httpPath: sql/protocolv1/o/0/11101101
    handler: Databricks

Data Source (with More Options)

connectionKey: my-databricks
nameTemplate:
  dataSourceFormat: Databricks <Tablename>
  tableFormat: <tablename>
  schemaFormat: databricks
connection:
  hostname: your.databricks.hostname.com
  port: 443
  ssl: true
  database: data
  username: token
  password: "${DATABRICKS_PASSWORD}"
  httpPath: sql/protocolv1/o/0/1110-11123
  handler: Databricks
sources:
  - table: credit_card_transactions
    schema: data
    tags:
      table:
        - PCI
        - SENSITIVE
      columns:
      - columnName: transaction_date
        tags:
          - PCI
          - DATE
  - table: crime_data
    schema: data
    naming:
        datasource: Crime Data
        table: crime_data
        schema: databricks

Databricks Data Source (Override Naming Convention)

connectionKey: ebock-databricks
nameTemplate:
  dataSourceFormat: Databricks <Tablename>
  tableFormat: <tablename>
  schemaFormat: databricks
connection:
  hostname: your.databricks.hostname.com
  port: 443
  ssl: true
  database: ebock
  username: token
  password: "${DATABRICKS_PASSWORD}"
  httpPath: sql/protocolv1/o/0/1110-185737-wove
  handler: Databricks
sources:
  - table: credit_card_transactions
    schema: ebock
  - table: crime_data_delta
    schema: ebock
    naming:
        datasource: Crime Data
        table: crime_data
        schema: databricks
  - table: hipaa_data
    schema: ebock

Impala Data Source (with `userFile`)

connectionKey: cdh-impala
nameTemplate:
  dataSourceFormat: Impala <Tablename>
  tableFormat: <tablename>
  schemaFormat: impala
connection:
  hostname: your.hadoop.hostname.example.com
  port: 21050
  ssl: true
  database: default
  handler: Apache Impala
  authenticationMethod: kerberos
  username: usera
  userFiles:
    - keyName: TrustedCerts
      content: <Base64 encoded contents of file go here>
      userFilename: tls-ca-bundle.pem
sources:
  - table: medical_records_parquet
    schema: default
  - table: nyc_taxi_fare_parquet
    schema: default
  - table: nyc_taxi_trip_parquet
    schema: default

Redshift Spectrum Data Sources

Your nativeSchemaFormat must contain _immuta to avoid schema name conflicts.

connectionKey: redshift
connection:
  hostname: your-redshift-cluster.djie25k.us-east-1.redshift.amazonaws.com
  port: 5439
  ssl: true
  database: your_database_with_external_schema
  username: awsuser
  password: your_password
  handler: Redshift
  schema: external_schema
nameTemplate:
  dataSourceFormat: <Tablename>
  schemaFormat: <schema>
  tableFormat: <tablename>
  schemaProjectNameFormat: <Schema>
  nativeSchemaFormat: <schema>_immuta
  nativeViewFormat: <tablename>
sources:
  - all: true

Snowflake Data Source (Specify Sources)

connectionKey: tpc-snowflake
nameTemplate:
  dataSourceFormat: Snowflake <Tablename>
  tableFormat: <tablename>
  schemaFormat: snowflake
connection:
  hostname: example.hostname.snowflakecomputing.com
  port: 443
  ssl: true
  database: TPC
  username: USERA
  password: "${SNOWFLAKE_PASSWORD}"
  schema: PUBLIC
  warehouse: IT_WH
  handler: Snowflake
sources:
  - table: CASE
    schema: PUBLIC
  - table: CASE2
    schema: PUBLIC
  - table: CUSTOMER
    schema: PUBLIC
  - table: WEB_SALES
    schema: PUBLIC

Data Sources

Payload Attribute Details

connectionKey

connection

Special Cases

nameTemplate

options

owners

sources

Recommended: Specify All Tables

Specify a Query

Specify a Table

Additional Options

Columns

Column Descriptions

Tags

Request Payload Examples

Basic Data Source

Data Source (with More Options)

Databricks Data Source (Override Naming Convention)

Impala Data Source (with userFile)

Redshift Spectrum Data Sources

Snowflake Data Source (Specify Sources)

Payload Attribute Details

connectionKey

connection

Special Cases

nameTemplate

options

owners

sources

Recommended: Specify All Tables

Specify a Query

Specify a Table

Additional Options

Columns

Column Descriptions

Tags

Request Payload Examples

Basic Data Source

Data Source (with More Options)

Databricks Data Source (Override Naming Convention)

Impala Data Source (with userFile)

Redshift Spectrum Data Sources

Snowflake Data Source (Specify Sources)

`connectionKey`

`connection`

`nameTemplate`

`options`

`owners`

`sources`

Impala Data Source (with `userFile`)

`connectionKey`

`connection`

`nameTemplate`

`options`

`owners`

`sources`

Impala Data Source (with `userFile`)