Notă
Accesul la această pagină necesită autorizare. Puteți încerca să vă conectați sau să modificați directoarele.
Accesul la această pagină necesită autorizare. Puteți încerca să modificați directoarele.
Important
This feature is in Public Preview. Workspace admins can control access to this feature from the Previews page. See Manage Azure Databricks previews.
After you have created your Feature View definitions, which are stored in Unity Catalog, you can produce feature data from your source table using the feature definitions. This process is called materializing your features. Azure Databricks creates and manages Lakeflow pipelines to populate tables in Unity Catalog for model training and batch scoring or online serving.
For information about serving Feature Views, see Serve Feature Views.
Requirements
- Features must be created as Feature Views and stored in Unity Catalog.
- For version requirements, see Requirements.
Materialization support by feature type
Whether and where a feature can be materialized depends on its type:
ColumnSelectionfeatures can be materialized to online stores. See ColumnSelection materialization.- Batch
Lastfeatures with aRollingWindowcan be materialized to online stores. See Materialize freshness-bounded latest values. RequestSource,FeatureViewSource, andCustomUDFfeatures cannot be materialized. See On-demand features and materialization.
Features that materialize only to online stores are still usable offline: create_training_set and compute_features compute their point-in-time values directly from the source, so no offline materialization is needed.
ColumnSelection materialization
ColumnSelection features select the latest value of a single column per entity key without aggregation. They can only be materialized to online stores. For offline use cases (training and batch inference), ColumnSelection features are fetched directly from the source data at query time, so offline materialization is not needed.
Materialization behavior
- The pipeline writes the most recent row per entity key to the online table, with no aggregation window.
- Online materialization populates the online table with the current latest value per entity key.
Example
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
DeltaTableSource, Feature, ColumnSelection, TableTrigger, OnlineStoreConfig,
)
fe = FeatureEngineeringClient()
delta_source = DeltaTableSource(
catalog_name="catalog",
schema_name="schema",
table_name="transactions",
)
amount_feature = Feature(
source=delta_source,
function=ColumnSelection("amount"),
entity=["user_id"],
timeseries_column="transaction_time",
name="latest_transaction_amount",
)
# Register before materializing
amount_feature = fe.register_feature(
feature=amount_feature,
catalog_name="catalog",
schema_name="schema",
)
mfs = fe.materialize_features(
features=[amount_feature],
online_config=OnlineStoreConfig(
catalog_name="catalog",
schema_name="feats_online",
table_name_prefix="txn_",
online_store_name="lb_usw2"
),
trigger=TableTrigger(),
)
ColumnSelection features use TableTrigger, which runs the pipeline whenever the source Delta table receives a new commit. No offline_config is needed because ColumnSelection features are read directly from the source for offline use cases (training and batch inference).
Note
RequestSource features cannot be materialized because they represent data provided by the caller at inference time (or extracted from the labeled DataFrame at training time). There is no source table to read from. The values exist only in the request payload or training DataFrame.
Materialize freshness-bounded latest values
A batch Last aggregation with a RollingWindow gives an online feature a time-to-live (TTL): its value expires out of the online store when no source value falls within the window. This is useful when the age of a value determines whether it is safe to serve. For example, a device state from the last hour can be served, while an older state resolves to null instead of remaining available indefinitely.
A RollingWindow defines an explicit duration that gives the feature TTL-like behavior. For example, if the source publishes values daily and you want to keep only values from the last seven days, use a RollingWindow with a seven-day duration. On each trigger, materialization does a full replace of the online table, so values that fall outside the window are removed on a predictable schedule. This produces the same result as a streaming Last with a RollingWindow, but costs less because it does not run compute continuously.
This publish pattern is similar to a feature table with a TTL that is published in snapshot mode. If you maintain a set of feature tables and want to use them alongside other Feature Views, you can adapt them to the Feature Views authoring framework. Use a batch Last with a RollingWindow and a TableTrigger.
This combination has a special online-only materialization mode with the following requirements:
- The source must be a
DeltaTableSource. - The aggregation function must be
Last, and its window must be aRollingWindow. - The materialization must provide an
OnlineStoreConfig, omitOfflineStoreConfig, and useTableTrigger.
In the following example, latest_device_state_1h is a registered Feature that meets these requirements:
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import OnlineStoreConfig, TableTrigger
fe = FeatureEngineeringClient()
materialized = fe.materialize_features(
features=[latest_device_state_1h],
online_config=OnlineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="latest_device_state_serving",
online_store_name="device_state_store",
),
trigger=TableTrigger(),
)
For offline training and batch scoring, the feature engineering client computes the point-in-time value directly from the source. It does not read an offline materialization for this feature.
Considerations
- Expiry advances only on a trigger. The online value and its expiry advance when a source table commit triggers a materialization refresh. Time passing alone does not trigger a refresh. If the source stops publishing, the last materialized value remains in the online store until a later commit triggers a refresh.
- Align the timeseries column with the source publish time. The timeseries column must reflect when the source published the data. Otherwise, the online and offline values diverge, because the online store joins at the trigger time while offline reads join at the timeseries time.
- Set the window duration to a multiple of the publish cadence. If the
RollingWindowduration is not a multiple of the source publish cadence, some values are treated as expired during offline training while they are still visible online.
On-demand features and materialization
RequestSource values come from the training DataFrame or inference request, so there is no source table to materialize. FeatureViewSource features apply a CustomUDF to upstream feature values during training or serving. They do not store a precomputed result. Materialization is also unsupported for CustomUDF features backed by a Delta table.
For a dependency graph such as revenue_sum_7d and cost_sum_7d feeding a margin feature:
- For offline training, call
create_training_setwithmargin. It resolves the upstream features and computes point-in-time values, using compatible offline materializations when available. - For online serving, materialize the supported revenue and cost features to an online store. The endpoint looks them up and computes
marginfor each request. - Pass only the supported upstream features to
materialize_features, notmarginor any request-backed features. Materialization does not recursively materialize a derived feature's dependencies.
A graph that uses only request-backed features does not need an online store. See Train with FeatureViewSource features and Serve derived features.
Permissions
Materialization requires privileges on the feature and on the source and destination resources. For full Unity Catalog privilege descriptions, see Unity Catalog privileges reference.
Materializing a feature requires
MANAGE. Callingmaterialize_featurescreates and manages the backing Lakeflow pipelines and Unity Catalog tables, so it is a management operation. You must haveMANAGEon the feature, along withREAD FEATUREto read the feature definition being materialized.Deleting a materialized feature is restricted to its creator. Only the user who created a materialized feature can delete it with
delete_materialized_feature. This restriction is independent of Unity Catalog privileges:MANAGEon the feature or its parent schema does not let another user delete it.Reading source data requires
SELECT. For a feature that uses a Delta table source, you must haveSELECTon the source table. For a feature that uses a Stream source, you must haveSELECTon the Stream's ingestion table.For other permissions required by a Stream's authentication configuration, see Kafka authentication.
Creating destination tables requires
CREATE TABLE. You must haveCREATE TABLEon every schema specified byOfflineStoreConfigorOnlineStoreConfig. Offline and online destinations can be in different schemas or catalogs, and materialization requires privileges on each destination.Materializing to an online store requires
CAN USE. You must haveCAN USEon the Lakebase instance or project used by the online store. For information about Lakebase permissions, see Grant project permissions.Listing materialized features requires
READ FEATUREon the parent feature. To uselist_materialized_features, you must haveREAD FEATUREon the feature that was materialized.Reading materialized data requires
SELECT. You must haveSELECTon each offline or online output table that you access.READ FEATUREon the parent feature does not grant access to these tables.
For every Unity Catalog resource involved in materialization, you also need USE CATALOG on its parent catalog and USE SCHEMA on its parent schema. This requirement applies to the feature, each source or Stream ingestion table, and each configured destination. READ FEATURE and MANAGE granted on a schema or catalog apply to all current and future features it contains.
API data structures
OfflineStoreConfig
Configuration for the offline store where materialized features will be written. When materialize_features is called, the feature store backend creates tables using this prefix. Each pipeline run materializes the latest feature values to the table according to materialization schedule.
OfflineStoreConfig(
catalog_name: str, # Catalog name for the offline table where materialized features will be stored
schema_name: str, # Schema name for the offline table
table_name_prefix: str # Table name prefix for the offline table. The pipeline may create multiple tables with this prefix, each updated at different cadences
)
from databricks.feature_engineering.entities import OfflineStoreConfig
offline_store = OfflineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="customer_features"
)
OnlineStoreConfig
Configuration for the online store, which stores features used by model serving. Materialization creates Delta tables with the catalog.schema.table_name_prefix, and streams the tables to the Online Feature Store with the same name.
from databricks.feature_engineering.entities import OnlineStoreConfig
online_store = OnlineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="customer_features_serving",
online_store_name="customer_features_store"
)
MaterializedFeature
Represents a Feature View that has been materialized, that is, that has a precomputed representation available in Unity Catalog. There are separate materialized features for the offline table and online table. Typically, users will not instantiate a MaterializedFeature directly.
API function calls
materialize_features()
Materializes a list of Feature Views into either an offline Delta table or to an Online Feature Store. Features must be registered in Unity Catalog before calling this function (for example, using create_feature or register_feature). Locally constructed features that have not been registered will not work.
FeatureEngineeringClient.materialize_features(
*, # Arguments are keyword-only
features: List[Feature], # List of Feature Views to materialize
offline_config: Optional[OfflineStoreConfig] = None, # Offline store config (aggregation features only)
online_config: Optional[OnlineStoreConfig] = None, # Online store config
trigger: Union[CronSchedule, TableTrigger, StreamingMode], # Materialization trigger
tags: Optional[Dict[str, str]] = None, # Custom tags for cost attribution
budget_policy_id: Optional[str] = None, # Serverless usage policy for cost attribution
) -> List[MaterializedFeature]:
The method returns a list of materialized features, which contain metadata about when feature values are updated and the Unity Catalog tables where features are materialized.
If both an OnlineStoreConfig and an OfflineStoreConfig are provided, then two materialized features are returned per feature provided, one for each type of store.
The trigger parameter controls when the materialization pipeline runs:
CronSchedule: Runs on a schedule derived from the feature timing or on a caller-provided Quartz cron schedule. Supported for batch aggregation features (AggregationFunctionfromDeltaTableSource).TableTrigger: Runs when the upstream Delta table receives a commit. Supported forColumnSelectionfeatures and aggregation features (AggregationFunction) backed by aDeltaTableSource. For aggregation features, the pipeline is throttled to run at most once per half the feature's granularity (the slide duration for a sliding window, or the window length for a tumbling window), capped at 1 hour and never more often than every 5 minutes. For example, at most once every 30 minutes for a 1-hour granularity, or at most once an hour for a granularity of 2 hours or more. The interval is measured from the previous run, so a commit that arrives after it elapses still triggers a run promptly.StreamingMode: Runs as a continuous streaming pipeline. Required for features backed by aStreamSource.
You cannot mix features that require different trigger types in a single materialize_features call. Issue separate calls instead.
To attribute the cost of a materialization, pass tags, budget_policy_id, or both. Azure Databricks applies them to the job or pipeline it creates, so its spend carries your attribution in the billable usage system table. Both are applied at creation, so attributing an existing materialization differently means creating a new one. budget_policy_id takes the ID of a serverless usage policy. To create one and get its ID, see Create a serverless usage policy. For the tag limits, the resources each value reaches, and how to query the attributed spend, see Feature Store cost management.
Materialize to offline store
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
CronSchedule, OfflineStoreConfig,
)
fe = FeatureEngineeringClient()
materialized = fe.materialize_features(
features=features,
offline_config=OfflineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="customer_features"
),
trigger=CronSchedule(
quartz_cron_expression="0 0 * * * ?", # Hourly
timezone_id="UTC",
),
)
Materialize to online store
Note
To materialize most aggregation features to an online store, you must also materialize to an offline store. Both offline_config and online_config are required. The online_store_name must reference an existing Online Feature Store. For instructions on creating one, see Databricks Online Feature Stores.
ColumnSelection features do not require an OfflineStoreConfig. See ColumnSelection materialization.
The batch Last with RollingWindow special case is also online-only. See Materialize freshness-bounded latest values.
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
CronSchedule, OfflineStoreConfig, OnlineStoreConfig,
)
fe = FeatureEngineeringClient()
materialized = fe.materialize_features(
features=features,
offline_config=OfflineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="customer_features"
),
online_config=OnlineStoreConfig(
catalog_name="main",
schema_name="feature_store",
table_name_prefix="customer_features_serving",
online_store_name="customer_features_store"
),
trigger=CronSchedule(
quartz_cron_expression="0 0 * * * ?", # Hourly
timezone_id="UTC",
),
)
Materialize streaming features
Streaming features can only be materialized to online stores; the offline_config parameter is not supported. Offline materialization is not supported because streaming features require a real-time pipeline to ensure sub-second freshness. For offline training or evaluation, the feature engineering client will recompute the feature values based on each data point evaluated.
Streaming features cannot be mixed with batch features in the same materialize_features call.
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
OnlineStoreConfig, StreamingMode,
)
fe = FeatureEngineeringClient()
materialized = fe.materialize_features(
features=[streaming_feature],
online_config=OnlineStoreConfig(
catalog_name="my_catalog",
schema_name="my_schema",
table_name_prefix="streaming_features_serving",
online_store_name="feature_store_online"
),
trigger=StreamingMode(),
)
list_materialized_features()
Returns the materializations of a single feature, identified by its full name. feature_name is required and keyword-only. To review the materializations across many features, first list the features in a catalog or schema, then call list_materialized_features on each feature returned.
By default, a maximum of 100 materializations are returned. You can change this limit using the max_results parameter.
FeatureEngineeringClient.list_materialized_features(
*, # Arguments are keyword-only
feature_name: str, # Required: full name of the feature whose materializations to list
max_results: int = 100, # Maximum number of materializations to return
) -> List[MaterializedFeature]:
delete_materialized_feature()
Before deleting a materialized feature, remove or update any models or feature specs that reference the feature.
Deletes a materialized feature. The feature to pass depends on the feature type:
- Aggregation features: Pass the offline materialized feature. If there is an online materialized feature for the same feature, both are deleted. For an online-only batch
Lastfeature with aRollingWindow, pass the online materialized feature. ColumnSelectionfeatures: Pass the online materialized feature.ColumnSelectionfeatures are materialized only to the online store (see ColumnSelection materialization), so there is no paired offline feature.
As part of materialization, features are grouped together by data source and aggregation window for efficiency. ColumnSelection features have no aggregation window, so they are grouped only by data source. The materialization pipeline, offline table, and online table are not deleted until all grouped features have been deleted. When the last materialized feature in a group is deleted, the feature store schedules the associated resources for automatic cleanup by a background process. See Background resource cleanup.
To clean up materialized features, look at the table associated with a materialized feature. Each feature in the table (one per column) must be deleted before compute and Delta table resources are cleaned up.
Use list_materialized_features() to get the materialized_feature argument.
FeatureEngineeringClient.delete_materialized_feature(
materialized_feature: MaterializedFeature, # Required: The materialized feature to delete
) -> None
from databricks.feature_engineering import FeatureEngineeringClient
fe = FeatureEngineeringClient()
feature_names = [
"main.feature_store.amount_sum_sliding_7d_1d",
"main.feature_store.amount_sum_sliding_30d_1d",
"main.feature_store.transaction_count_sliding_7d_1d",
"main.feature_store.latest_transaction_amount",
"main.feature_store.latest_user_tier",
]
for name in feature_names:
mfs = fe.list_materialized_features(feature_name=name) # required, keyword-only
offline = [mf for mf in mfs if not mf.is_online]
for mf in (offline or mfs):
fe.delete_materialized_feature(materialized_feature=mf)
fe.delete_feature(full_name=name)
Background resource cleanup
When you delete a materialized feature, Databricks removes the feature metadata immediately. The associated infrastructure (tables, pipelines, and jobs) is cleaned up asynchronously by a background process.
Because multiple materialized features can share the same tables and pipelines, these shared resources are not removed until every materialized feature that references them has been deleted. When the last materialized feature sharing a set of tables is deleted, the background process automatically deletes the following resources:
- The offline Delta tables containing the materialized feature data
- The online tables, if the features were materialized to an online store
- The materialization pipeline
- The orchestration job
This background process uses a Databricks-managed system service principal to perform these cleanup actions on your behalf, including deleting tables, pipelines, and jobs in your workspace. No action is required from you. The cleanup is fully managed by the feature store.
Note
There might be a short delay between deleting the last materialized feature in a group and the removal of the associated tables and other resources.
View materialization status
To view the materialization status of your Feature Views in the Databricks UI and debug materialization errors, see Explore Feature Views in Unity Catalog.
Limitations
Batch features
- Batch materialization pipelines run as serverless Lakeflow pipelines.
- Batch rolling window features cannot be materialized, except for the online-only
Lastspecial case described in Materialize freshness-bounded latest values. For offline training or batch inference, rolling window features are computed from the source data for each point-in-time lookup. ColumnSelectionfeatures can only be materialized to online stores.RequestSource,FeatureViewSource, andCustomUDFfeatures cannot be materialized. See On-demand features and materialization.- Materialized features can only be deleted in the workspace in which they were created.
- Only the user who created a materialized feature can delete it, regardless of Unity Catalog privileges on the feature or its parent schema.
- For materialized aggregation features, the online materialized feature cannot be deleted directly. Delete the paired offline materialized feature, and the change propagates to both.
- For materialized aggregation features created before April 20, 2026, the materialization pipeline continues producing new feature values until all materialized features in the pipeline have been deleted, which triggers resource cleanup. To create an updated pipeline that supports per-feature delete, delete and re-materialize the feature.
- For materialized
ColumnSelectionfeatures, the materialization pipeline continues producing new feature values until all materialized features in the pipeline have been deleted, which triggers resource cleanup.
Streaming features
- Streaming features can only be materialized to online stores. Offline materialization is not needed because streaming features at training time are designed to be recomputed from historic events per data point to provide millisecond-level accuracy.
- Streaming features cannot be mixed with batch features in a single
materialize_featurescall. compute_featuresdoes not support streaming features.- The workspace must be in a region that supports Lakebase instances.
- Only JSON-serialized Kafka messages are supported. Message schemas must be provided directly in JSON Schema format. Schema registries (Confluent, Glue) are not formally supported during the preview, but if you provide the schema directly, pipelines can read from topics governed by a schema registry.
- Only
RollingWindowis supported for streaming aggregation features.TumblingWindowandSlidingWindowshould be used with batch features. - Only
Count,Avg,Sum,StddevPop,Max,Min,First,Last,FirstN,LastN,FirstDistinct, andLastDistinctaggregation functions are supported for streaming features. - Column selection features from streaming sources do not handle out-of-order messages. The latest event on the Kafka stream is shown, even if the timeseries column value is earlier than a previously received event.
- Streaming pipelines are restarted twice a week. Each restart can cause processing delays and startup times of up to 1 minute. Excluding restarts, the p99 freshness is 200ms.
- Feature backfill for materialization is not supported. When a feature is materialized, it calculates from that point forward. Newly created aggregations in the online store are inaccurate until their time window has passed.
- Only Databricks Online Feature Store is supported.
- Streaming materialization pipelines run as serverless Lakeflow pipelines.
- Enterprise tier workspaces only.