Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Important
This feature is in Beta. To use it, a workspace admin must turn on Feature Views Advanced Functionality from the Previews page. See Manage Azure Databricks previews.
Use a feature entity purge to remove the computed feature values of one or more Feature Views for selected entities. A purge removes matching values from applicable offline and online materialization tables in Unity Catalog and Lakebase. Run the purge in the workspace where the materializations you want to purge were created. It does not delete the Feature View definitions, materializations, or source data. Materializations created in other workspaces are not affected.
For example, you can purge a customer's materialized feature values after receiving a data deletion request.
Warning
Before you submit a purge, you must remove the selected entities from the feature source data and prevent the source from receiving new data for those entities. A purge does not change the source or stop scheduled and continuous materialization. If matching source data remains or arrives later, a materialization recreates the values.
Requirements
- The features must be registered Feature Views. Each Feature View must have at least one materialization created in the workspace where you submit the purge. To purge materializations created in different workspaces, submit a separate request in each workspace.
RequestSource,FeatureViewSource, and other nonmaterializable features are not supported. - Specify from 1 through 500 unique features by their three-part Unity Catalog names.
- No feature can have an active purge or backfill operation in progress when you submit the request.
- The entities table must be a Unity Catalog Delta table with at least one column. See Prepare the entities table.
- Use Databricks Feature Engineering client version 0.18.0 and above.
Permissions
For production, Databricks recommends running purges as the same service principal that creates and materializes the Feature Views. This principal normally already has the required privileges. If privileges have changed, verify the following requirements before submitting a purge.
To submit and complete a purge, you must have the following privileges:
MANAGEon every Feature View in the request.SELECTon the entities table.SELECTandMODIFYon each offline, partial-aggregate, and online table that the purge updates.- Access to the Lakebase instance or project used by the online store, including permission to read and update its online tables. See Grant project permissions.
- For a schema-registry Stream without an embedded schema, permission to read its ingestion-table metadata, such as
SELECTorREAD_METADATAwhere available. Purge validation reads the schema to resolve entity-key types, not the ingestion-table rows. USE CATALOGandUSE SCHEMAon the parent catalog and schema of each resource listed above.
For full Unity Catalog privilege descriptions, see Unity Catalog privileges reference. For Feature View permissions, see Materialize Feature Views.
Prepare the entities table
Create a Delta table that identifies the entities to purge. Each column must correspond to an entity key of every Feature View in the request. The column names and data types must match the Feature View entity keys. Don't include non-key columns.
The table can contain a subset of the entity keys that is a leading prefix of the Feature View's entities list. The purge matches all entities with the supplied key values, regardless of the remaining keys. Null key values are supported.
Important
The entities-table columns must contain all entity keys or a leading prefix that matches the Feature View's entities list order. The order of columns in the entities table does not matter.
For existing online tables, the primary-key order can differ from the Feature View's entities list order. In that case, validation uses the online table's primary-key prefix instead and reports the required order if a partial-key request does not match.
Requested Feature Views may have different entity-key sets, as long as the entities-table columns are a subset of every Feature View's entity-key set.
The purge job resolves one Delta table version and uses that snapshot for every target. Don't modify the entities table after you submit the request. To purge entities added after the job resolves the version, submit another request.
Submit and monitor a purge
Call FeatureEngineeringClient.purge_feature_entities. The method returns a long-running operation without waiting for the purge to finish.
from databricks.feature_engineering import FeatureEngineeringClient
fe = FeatureEngineeringClient()
operation = fe.purge_feature_entities(
features=[
"main.feature_store.customer_purchase_count_30d",
"main.feature_store.customer_purchase_total_30d",
],
entities_table="main.feature_store.entities_to_purge",
)
print(f"Purge operation: {operation.name()}")
# wait() polls until the operation finishes.
response = operation.wait()
print(f"Purge state: {response.state}")
for result in response.results or []:
print(
result.feature,
result.offline_state,
result.online_state,
result.error,
)
Use operation.done() to refresh the operation and make one status check without waiting for completion. Then use operation.metadata() to inspect its PENDING or RUNNING state.
The completed response contains one result per Feature View:
offline_statereports the outcome for offline materializations.online_statereports the combined outcome for online materializations.- An outcome is
SUCCEEDED,FAILED, orNOT_APPLICABLE. - An outcome is
NOT_APPLICABLEwhen the Feature View has no materialization target for that store type. For example, an online-only Feature View has an offline outcome ofNOT_APPLICABLE, and a Feature View without an online materialization has an online outcome ofNOT_APPLICABLE. errorcontains the Feature View-specific error when an applicable target fails.
Check every per-feature result. An operation can finish with different outcomes for different Feature Views or store types.
Also check response.state and response.error for an operation-level failure.
If a purge partially fails, inspect the per-feature errors and correct any access or configuration problems. Then submit a new purge request with a new request_id for only the Feature Views with failed outcomes. Keep the source data clean before retrying. If failures persist, contact Databricks support.
Retry a request safely
Set request_id to a UUID4 value and retain it until the request is accepted. If a network or client failure makes the submission result ambiguous, retry the same request with the same ID. Databricks returns the existing operation instead of starting a second purge.
The feature set and entities table must match the original request. Reusing an ID with different request parameters is rejected. The client generates an ID when you omit one, but an explicit ID is required to make application-level retries idempotent. This is also useful when an outer orchestration layer might retry the operation.
Attribute purge compute costs
Purge runs use serverless compute. With Databricks Feature Engineering client version 0.18.1 and above, use tags or budget_policy_id to attribute the compute cost:
operation = fe.purge_feature_entities(
features=["main.feature_store.customer_purchase_count_30d"],
entities_table="main.feature_store.entities_to_purge",
tags={"team": "recommendations", "workflow": "privacy-deletion"},
budget_policy_id="00000000-0000-0000-0000-000000000000",
)
The tags apply to purge compute, not to the Unity Catalog Feature resources. A request supports at most 25 tags.