> For the complete documentation index, see [llms.txt](https://documentation.immuta.com/2024.2/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.immuta.com/2024.2/application-settings/reference-guides/metadata-ingestion.md).

# Metadata Ingestion

For Immuta to enforce policies, it needs to catalog the resources policies are being applied to by performing metadata ingestion. Metadata ingestion is the process that occurs when you [create an Immuta data source](/2024.2/data-and-integrations/registering-metadata/register-data-sources/query-backed-tutorial.md) where Immuta gathers details about your tables. However, Immuta does not need access to the data within the tables in order to protect it, with the exception of a few specific and advanced masking policies detailed below.

## Immuta Metadata Database

Immuta collects and stores the following kinds of information in Immuta's Metadata Database for policy enforcement. Further, policy information may be transmitted to data source host systems for enforcement purposes as part of a query or to enable the host system to perform policy enforcement.

* **Identity Management Information:** Usernames, group information, and other kinds of personal identifiers may be stored and referenced for the purposes of performing authentication and access control and may be retained in audit logs. When such information is relevant for access determination under policy, it may be retained as part of the policy definition.
* **Schema Information:** Data source metadata such as schema, column data types, and information about the host.

Immuta's Metadata Database can also contain the following forms of metadata for policy enforcement. These forms contain sample data from your tables and [can be disabled](/2024.2/data-and-integrations/registering-metadata/data-source-settings/how-to-guides/disable-data-sampling.md) if you do not want Immuta to have access to the data being protected.

* **Fingerprints:** When enabled, additional statistical queries made during the health check are distilled into summary statistics, called fingerprints. During this process, statistical query results and data samples (which may contain PII) are temporarily held in memory by the Fingerprint Service.
* **k-Anonymization Policies:** When a k-anonymization policy is applied, the columns under the k-anonymization policy are queried within a separate fingerprinting process which generates rules enforcing k-anonymity. The results of this query, which may contain PII, are temporarily held in memory by the Fingerprint Service. The final rules are stored for enforcement. Immuta requires that you opt in to use this masking policy type. To enable k-anonymization for your account, see the [k-anonymization section on the app settings how-to guide](/2024.2/application-settings/how-to-guides/config-builder-guide.md#k-anonymization).
* **Randomized Response Policies:** If the list of substitution values for a categorical column is not part of the policy specification (e.g., when specified via the API), a list is obtained via query and merged into the policy definition.

### Metadata Ingestion Workflow

If no metadata collection types have been disabled, data is processed in the following workflow to support data source creation, health checks, policy enforcement, and dictionary features.

1. A System Administrator configures the integration in Immuta.
2. A Data Owner registers data sources from their remote data platform with Immuta.
3. When a data source is created or updated, the Metadata Database pulls in and stores statistics about the data source, including row count and high cardinality calculations.
4. The data source health check runs daily to ensure existing tables are still valid.
5. If an external catalog is enabled, the daily health check will pull in data source attributes (e.g., tags and definitions) and store them in the Metadata Database.

### Immuta Privileges

Immuta requires certain privileges to perform metadata ingestion. The user connecting a table to Immuta as a data source must have privileges specific to their data platform to perform metadata ingestion.

For example, a user registering a Snowflake table as an Immuta data source must have the `REFERENCES` privilege to view the structure of the table and allow Immuta access to that information as well. This does not require the user (or Immuta) to have access to view the data itself.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://documentation.immuta.com/2024.2/application-settings/reference-guides/metadata-ingestion.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.