Megjegyzés
Az oldalhoz való hozzáféréshez engedély szükséges. Megpróbálhat bejelentkezni vagy módosítani a címtárat.
Az oldalhoz való hozzáféréshez engedély szükséges. Megpróbálhatja módosítani a címtárat.
Ez a cikk a csomagkonfigurációs fájlok szintaxisát ismerteti, amelyek deklaratív Automation-kötegeket (korábbi nevén Databricks-eszközcsomagokat) határoznak meg. Lásd : Mik azok a deklaratív automatizálási csomagok?.
A csomagok létrehozásához és használatához lásd: Deklaratív automation-csomagok fejlesztése.
Csomagkonfigurációs referencia: Konfigurációs referencia.
databricks.yml
A csomagnak tartalmaznia kell egy (és csak egy) konfigurációs fájlt, amely a köteg projektmappájának gyökerénél van elnevezve databricks.yml .
databricks.yml a csomagokat meghatározó fő konfigurációs fájl, de hivatkozhat más konfigurációs fájlokra, például erőforrás-konfigurációs fájlokra a include leképezésben. A csomagkonfiguráció YAML-ben van kifejezve. A YAML-ről további információt a hivatalos YAML-specifikációban talál.
A legegyszerűbben databricks.yml a csomag nevét határozza meg, amely egy szükséges legfelső szintű leképezés és egy céltelepítés.
bundle:
name: my_bundle
targets:
dev:
default: true
Az összes legfelső szintű megfeleltetés részleteiért lásd a konfigurációs referenciát.
Jótanács
A Deklaratív Automation-csomagok Python-támogatása lehetővé teszi erőforrások definiálását a Pythonban. Lásd : Csomagkonfiguráció a Pythonban.
Előírás
Az alábbi YAML-specifikáció legfelső szintű konfigurációs kulcsokat biztosít a Deklaratív Automation-csomagokhoz. A teljes konfigurációs referenciaért tekintse meg a konfigurációs referencia és a Deklaratív automation-csomagok erőforrásait.
# This is the default bundle configuration if not otherwise overridden in
# the "targets" top-level mapping.
bundle: # Required.
name: string # Required.
databricks_cli_version: string
cluster_id: string
deployment: Map
git:
origin_url: string
branch: string
# This is the identity to use to run the bundle
run_as:
- user_name: <user-name>
- service_principal_name: <service-principal-name>
# These are any additional configuration files to include.
include:
- '<some-file-or-path-glob-to-include>'
- '<another-file-or-path-glob-to-include>'
# These are any scripts that can be run.
scripts:
<some-unique-script-name>:
content: string
# These are any additional files or paths to include or exclude.
sync:
include:
- '<some-file-or-path-glob-to-include>'
- '<another-file-or-path-glob-to-include>'
exclude:
- '<some-file-or-path-glob-to-exclude>'
- '<another-file-or-path-glob-to-exclude>'
paths:
- '<some-file-or-path-to-synchronize>'
# These are the default artifact settings if not otherwise overridden in
# the targets top-level mapping.
artifacts:
<some-unique-artifact-identifier>:
build: string
dynamic_version: boolean
executable: string
files:
- source: string
path: string
type: string
# These are for any custom variables for use throughout the bundle.
variables:
<some-unique-variable-name>:
description: string
default: string or complex
lookup: Map
type: string # The only valid value is "complex" if the variable is a complex variable, otherwise do not define this key.
# These are the workspace settings if not otherwise overridden in
# the targets top-level mapping.
workspace:
artifact_path: string
host: string
profile: string
resource_path: string
root_path: string
state_path: string
# These are the permissions to apply to resources defined
# in the resources mapping.
permissions:
- level: <permission-level>
group_name: <unique-group-name>
- level: <permission-level>
user_name: <unique-user-name>
- level: <permission-level>
service_principal_name: <unique-principal-name>
# These are the resource settings if not otherwise overridden in
# the targets top-level mapping.
resources:
alerts:
<unique-alert-name>:
# alert settings
apps:
<unique-app-name>:
# app settings
catalogs:
<unique-catalog-name>:
# catalog settings
clusters:
<unique-cluster-name>:
# cluster settings
dashboards:
<unique-dashboard-name>:
# dashboard settings
database_catalogs:
<unique-database-catalog-name>:
# database catalog settings
database_instances:
<unique-database-instance-name>:
# database instance settings
experiments:
<unique-experiment-name>:
# experiment settings
jobs:
<unique-job-name>:
# job settings
model_serving_endpoints:
<unique-model-serving-endpoint-name>:
# model_serving_endpoint settings
pipelines:
<unique-pipeline-name>:
# pipeline settings
postgres_branches:
<unique-postgres-branch-name>:
# postgres branch settings
postgres_endpoints:
<unique-postgres-endpoint-name>:
# postgres endpoint settings
postgres_projects:
<unique-postgres-project-name>:
# postgres project settings
quality_monitors:
<unique-quality-monitor-name>:
# quality monitor settings
registered_models:
<unique-registered-model-name>:
# registered model settings
schemas:
<unique-schema-name>:
# schema settings
secret_scopes:
<unique-secret-scope-name>:
# secret scopes settings
sql_warehouses:
<unique-sql-warehouse-name>:
# sql warehouse settings
synced_database_tables:
<unique-synced-database-table-name>:
# synced database table settings
volumes:
<unique-volume-name>:
# volumes settings
# These are the targets to use for deployments and workflow runs. One and only one of these
# targets can be set to "default: true".
targets:
<some-unique-programmatic-identifier-for-this-target>:
artifacts:
# artifact build settings for this target
bundle:
# bundle settings for this target
default: boolean
git: Map
mode: string
permissions:
# permissions for this target
presets:
<preset>: <value>
resources:
# resource settings for this target
sync:
# sync settings for this target
variables:
<defined-variable-name>: <non-default-value> # value for this target
workspace:
# workspace settings for this target
run_as:
# run_as settings for this target
Példák
Ez a szakasz néhány alapvető példát tartalmaz a csomagok működésének és a konfiguráció strukturálásának megértéséhez.
Megjegyzés:
A csomagfunkciókat és a csomaghasználati eseteket bemutató konfigurációs példákért tekintse meg a Csomagkonfigurációs példákat és a Csomagpéldák adattárat a GitHubon.
Az alábbi példakonfiguráció egy olyan helyi fájlt hello.py határoz meg, amely ugyanabban a könyvtárban található, mint a csomagkonfigurációs fájl databricks.yml. Ezt a jegyzetfüzetet feladatként futtatja a megadott fürtazonosítóval rendelkező távoli fürtön. A távoli munkaterület URL-címe és a munkaterület hitelesítési hitelesítő adatai a hívó helyi konfigurációs profiljából lesznek beolvasva DEFAULT.
bundle:
name: hello-bundle
resources:
jobs:
hello-job:
name: hello-job
tasks:
- task_key: hello-task
existing_cluster_id: 1234-567890-abcde123
notebook_task:
notebook_path: ./hello.py
targets:
dev:
default: true
Az alábbi példa egy célhelyet prod ad hozzá, amely egy eltérő távoli munkaterület URL-címét és munkaterület-hitelesítési hitelesítő adatait használja, amelyeket a hívó fájl .databrickscfg fájljának megadott munkaterület URL-címéhez illeszkedő host bejegyzéséből olvas be. Ez a feladat ugyanazt a jegyzetfüzetet futtatja, de egy másik távoli fürtöt használ a megadott fürtazonosítóval.
Megjegyzés:
A Databricks azt javasolja, hogy a host leképezés helyett default a leképezést használja, ahol csak lehetséges, mivel ez hordozhatóbbá teszi a csomagkonfigurációs fájlokat. A host leképezés beállítása arra utasítja a Databricks parancssori felületét, hogy keressen egy megfelelő profilt a .databrickscfg fájlban, majd a profil mezőivel állapítsa meg, hogy melyik Databricks-hitelesítési típust használja. Ha több profil is létezik egyező host mezővel, akkor a --profile csomagparancsok beállításával meg kell adnia a használni kívánt profilt.
Figyelje meg, hogy nem kell deklarálnia a notebook_task leképezést a prod leképezésen belül, mivel az alapértelmezés szerint a notebook_task leképezést használja a legfelső szintű resources leképezésen belül, ha a notebook_task leképezést nincs explicit módon felülírva a prod leképezésen belül.
bundle:
name: hello-bundle
resources:
jobs:
hello-job:
name: hello-job
tasks:
- task_key: hello-task
existing_cluster_id: 1234-567890-abcde123
notebook_task:
notebook_path: ./hello.py
targets:
dev:
default: true
prod:
workspace:
host: https://<production-workspace-url>
resources:
jobs:
hello-job:
name: hello-job
tasks:
- task_key: hello-task
existing_cluster_id: 2345-678901-fabcd456
A következő csomagparancsokkal érvényesítheti, üzembe helyezheti és futtathatja ezt a feladatot a dev célon belül. A csomagok életciklusával kapcsolatos részletekért lásd: Deklaratív automation-csomagok fejlesztése.
# Because the "dev" target is set to "default: true",
# you do not need to specify "-t dev":
databricks bundle validate
databricks bundle deploy
databricks bundle run hello_job
# But you can still explicitly specify it, if you want or need to:
databricks bundle validate
databricks bundle deploy -t dev
databricks bundle run -t dev hello_job
A feladat érvényesítése, üzembe helyezése és futtatása a prod célon belül:
# You must specify "-t prod", because the "dev" target
# is already set to "default: true":
databricks bundle validate
databricks bundle deploy -t prod
databricks bundle run -t prod hello_job
A több csomagra kiterjedő definíciók és beállítások modularizálása és jobb újrafelhasználása érdekében ossza fel a csomagkonfigurációt külön fájlokra:
# databricks.yml
bundle:
name: hello-bundle
include:
- '*.yml'
# hello-job.yml
resources:
jobs:
hello-job:
name: hello-job
tasks:
- task_key: hello-task
existing_cluster_id: 1234-567890-abcde123
notebook_task:
notebook_path: ./hello.py
# targets.yml
targets:
dev:
default: true
prod:
workspace:
host: https://<production-workspace-url>
resources:
jobs:
hello-job:
name: hello-job
tasks:
- task_key: hello-task
existing_cluster_id: 2345-678901-fabcd456