GitLab connector reference

Important

This feature is in Beta. To use it, a workspace admin must turn on Lakeflow Connect for GitLab from the Previews page. See Manage Azure Databricks previews.

This page has reference information for the managed GitLab connector, including supported source tables and destination table schemas.

Supported source tables

The GitLab connector supports the following source tables, all under the default source schema:

Source table Primary key Description Sync mode Cursor field
groups id GitLab groups that the authenticating user is a member of. Full refresh only
projects id GitLab projects that the authenticating user is a member of. Full refresh only
issues id Issues across the member projects. Incremental updated_at
merge_requests id Merge requests across the member projects. Incremental updated_at
commits id Repository commits for each member project. Incremental created_at
pipelines id CI/CD pipelines for each project. Incremental updated_at
project_members project_id, id Members of each project. Full refresh only
branches project_id, name Repository branches for each project. Full refresh only
tags project_id, name Repository tags for each project. Full refresh only
releases project_id, tag_name Releases for each project. Full refresh only
deployments id Deployments for each project. Incremental updated_at
jobs id CI/CD jobs for each pipeline. Full refresh only
project_labels project_id, id Labels defined in each project. Full refresh only
project_milestones id Milestones for each project. Incremental updated_at
users project_id, id Users with access to each project. Full refresh only
group_labels group_id, id Labels defined in each group. Full refresh only
group_milestones id Milestones for each group. Incremental updated_at
group_members group_id, id Members of each group. Full refresh only
group_issue_boards group_id, id Issue boards for each group. Full refresh only
epics id Epics for each group. Requires a GitLab Premium or Ultimate group. Incremental updated_at
merge_request_commits project_id, merge_request_iid, id Commits on each merge request. Full refresh only
epic_issues group_id, epic_iid, epic_issue_id Issues linked to each epic. Requires a GitLab Premium or Ultimate group. Full refresh only
pipelines_extended project_id, id Extended detail for each pipeline. Full refresh only
group_iterations group_id, id Iterations for each group. Requires a GitLab Premium or Ultimate group. Incremental updated_at
epic_boards group_id, id Epic boards for each group. Requires a GitLab Premium or Ultimate group. Full refresh only
merge_request_discussions project_id, merge_request_iid, id Discussion threads on each merge request. Full refresh only
vulnerabilities project_id, id Security vulnerabilities for each project. Requires a GitLab Ultimate project. Full refresh only
compliance_frameworks group_id, id Compliance frameworks for each group. Requires a GitLab Ultimate group. Full refresh only

Connector options

For usage, see Examples.

Option Type Scope Required Applies to Description
start_datetime string Pipeline No All incremental tables The earliest record timestamp to ingest on the first sync, in yyyy-MM-dd'T'HH:mm:ss.SSS'Z' format. Defaults to 2018-01-01T00:00:00.000Z.

Destination table schemas

groups

Primary key: id Sync mode: Full refresh only

Field Data type
id BIGINT
name STRING
path STRING
full_path STRING
full_name STRING
description STRING
visibility STRING
web_url STRING
parent_id BIGINT
created_at TIMESTAMP

projects

Primary key: id Sync mode: Full refresh only

Field Data type
id BIGINT
name STRING
name_with_namespace STRING
path_with_namespace STRING
description STRING
visibility STRING
web_url STRING
default_branch STRING
archived BOOLEAN
topics ARRAY<STRING>
namespace STRUCT<id: BIGINT, name: STRING, path: STRING, kind: STRING, full_path: STRING, parent_id: BIGINT, avatar_url: STRING, web_url: STRING>
created_at TIMESTAMP
updated_at TIMESTAMP
last_activity_at TIMESTAMP

issues

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
project_id BIGINT
title STRING
description STRING
state STRING
author STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
assignees ARRAY<STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>>
labels ARRAY<STRING>
milestone STRUCT<id: BIGINT, iid: BIGINT, project_id: BIGINT, group_id: BIGINT, title: STRING, description: STRING, state: STRING, created_at: TIMESTAMP, updated_at: TIMESTAMP, due_date: DATE, start_date: DATE, expired: BOOLEAN, web_url: STRING>
web_url STRING
created_at TIMESTAMP
updated_at TIMESTAMP
closed_at TIMESTAMP

merge_requests

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
project_id BIGINT
title STRING
description STRING
state STRING
merge_status STRING
detailed_merge_status STRING
author STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
assignees ARRAY<STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>>
reviewers ARRAY<STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>>
source_branch STRING
target_branch STRING
sha STRING
merge_commit_sha STRING
labels ARRAY<STRING>
web_url STRING
created_at TIMESTAMP
updated_at TIMESTAMP
merged_at TIMESTAMP
closed_at TIMESTAMP

commits

Primary key: id Cursor field: created_at

Field Data type
id STRING
short_id STRING
project_id BIGINT
title STRING
message STRING
author_name STRING
author_email STRING
authored_date TIMESTAMP
committer_name STRING
committer_email STRING
committed_date TIMESTAMP
created_at TIMESTAMP
web_url STRING
parent_ids ARRAY<STRING>
stats STRUCT<additions: BIGINT, deletions: BIGINT, total: BIGINT>
trailers VARIANT
extended_trailers VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

pipelines

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
project_id BIGINT
sha STRING
ref STRING
status STRING
source STRING
name STRING
web_url STRING
created_at TIMESTAMP
updated_at TIMESTAMP

project_members

Primary key: project_id, id Sync mode: Full refresh only

Field Data type
id BIGINT
project_id BIGINT
username STRING
name STRING
state STRING
access_level BIGINT
avatar_url STRING
web_url STRING
created_by STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
expires_at STRING
created_at TIMESTAMP

branches

Primary key: project_id, name Sync mode: Full refresh only

Field Data type
project_id BIGINT
name STRING
merged BOOLEAN
protected BOOLEAN
developers_can_push BOOLEAN
developers_can_merge BOOLEAN
can_push BOOLEAN
default BOOLEAN
commit STRUCT<id: STRING, short_id: STRING, created_at: TIMESTAMP, parent_ids: ARRAY<STRING>, title: STRING, message: STRING, author_name: STRING, author_email: STRING, authored_date: TIMESTAMP, committer_name: STRING, committer_email: STRING, committed_date: TIMESTAMP, web_url: STRING>

tags

Primary key: project_id, name Sync mode: Full refresh only

Field Data type
project_id BIGINT
name STRING
message STRING
target STRING
protected BOOLEAN
release STRUCT<tag_name: STRING, description: STRING>
commit STRUCT<id: STRING, short_id: STRING, created_at: TIMESTAMP, parent_ids: ARRAY<STRING>, title: STRING, message: STRING, author_name: STRING, author_email: STRING, authored_date: TIMESTAMP, committer_name: STRING, committer_email: STRING, committed_date: TIMESTAMP, web_url: STRING>
created_at TIMESTAMP

releases

Primary key: project_id, tag_name Sync mode: Full refresh only

Field Data type
project_id BIGINT
name STRING
tag_name STRING
description STRING
created_at TIMESTAMP
released_at TIMESTAMP
author STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
commit STRUCT<id: STRING, short_id: STRING, created_at: TIMESTAMP, parent_ids: ARRAY<STRING>, title: STRING, message: STRING, author_name: STRING, author_email: STRING, authored_date: TIMESTAMP, committer_name: STRING, committer_email: STRING, committed_date: TIMESTAMP, web_url: STRING>
assets STRUCT<count: BIGINT, sources: ARRAY<STRUCT<format: STRING, url: STRING>>, links: ARRAY<STRUCT<id: BIGINT, name: STRING, url: STRING, direct_asset_url: STRING, link_type: STRING>>>
_links VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

deployments

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
project_id BIGINT
status STRING
ref STRING
sha STRING
created_at TIMESTAMP
updated_at TIMESTAMP
user STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
environment STRUCT<id: BIGINT, name: STRING, slug: STRING, external_url: STRING, created_at: TIMESTAMP, updated_at: TIMESTAMP>
deployable VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

jobs

Primary key: id Sync mode: Full refresh only

Field Data type
id BIGINT
project_id BIGINT
pipeline_id BIGINT
status STRING
stage STRING
name STRING
ref STRING
failure_reason STRING
created_at TIMESTAMP
started_at TIMESTAMP
finished_at TIMESTAMP
duration DOUBLE
queued_duration DOUBLE
user STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
commit STRUCT<id: STRING, short_id: STRING, created_at: TIMESTAMP, parent_ids: ARRAY<STRING>, title: STRING, message: STRING, author_name: STRING, author_email: STRING, authored_date: TIMESTAMP, committer_name: STRING, committer_email: STRING, committed_date: TIMESTAMP, web_url: STRING>

project_labels

Primary key: project_id, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
id BIGINT
name STRING
color STRING
description STRING
description_html STRING
text_color STRING
subscribed BOOLEAN
archived BOOLEAN
priority BIGINT
is_project_label BOOLEAN
open_issues_count BIGINT
closed_issues_count BIGINT
open_merge_requests_count BIGINT

project_milestones

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
project_id BIGINT
title STRING
description STRING
state STRING
created_at TIMESTAMP
updated_at TIMESTAMP
due_date DATE
start_date DATE
expired BOOLEAN

users

Primary key: project_id, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
id BIGINT
name STRING
username STRING
state STRING
avatar_url STRING
web_url STRING
locked BOOLEAN

group_labels

Primary key: group_id, id Sync mode: Full refresh only

Field Data type
group_id BIGINT
id BIGINT
name STRING
color STRING
description STRING
description_html STRING
text_color STRING
subscribed BOOLEAN
archived BOOLEAN
open_issues_count BIGINT
closed_issues_count BIGINT
open_merge_requests_count BIGINT

group_milestones

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
iid BIGINT
group_id BIGINT
title STRING
description STRING
state STRING
created_at TIMESTAMP
updated_at TIMESTAMP
due_date DATE
start_date DATE
expired BOOLEAN

group_members

Primary key: group_id, id Sync mode: Full refresh only

Field Data type
group_id BIGINT
id BIGINT
name STRING
username STRING
state STRING
avatar_url STRING
web_url STRING
access_level BIGINT
created_by STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
created_at TIMESTAMP
expires_at STRING
locked BOOLEAN

group_issue_boards

Primary key: group_id, id Sync mode: Full refresh only

Field Data type
id BIGINT
group_id BIGINT
name STRING
hide_backlog_list BOOLEAN
hide_closed_list BOOLEAN
project BIGINT
group STRUCT<id: BIGINT, name: STRING, web_url: STRING>
lists ARRAY<STRUCT<id: BIGINT, label: STRUCT<id: BIGINT, name: STRING, color: STRING, description: STRING, text_color: STRING>, position: BIGINT, max_issue_count: BIGINT, max_issue_weight: BIGINT, limit_metric: STRING>>

epics

Primary key: id Cursor field: updated_at

Field Data type
id BIGINT
work_item_id BIGINT
iid BIGINT
group_id BIGINT
parent_id BIGINT
parent_iid BIGINT
title STRING
description STRING
state STRING
confidential BOOLEAN
web_url STRING
reference STRING
references STRUCT<short: STRING, relative: STRING, full: STRING>
author STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
start_date DATE
start_date_is_fixed BOOLEAN
start_date_fixed DATE
start_date_from_milestones DATE
start_date_from_inherited_source DATE
end_date DATE
due_date DATE
due_date_is_fixed BOOLEAN
due_date_fixed DATE
due_date_from_milestones DATE
due_date_from_inherited_source DATE
created_at TIMESTAMP
updated_at TIMESTAMP
closed_at TIMESTAMP
labels ARRAY<STRING>
upvotes BIGINT
downvotes BIGINT
color STRING
_links VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

merge_request_commits

Primary key: project_id, merge_request_iid, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
merge_request_iid BIGINT
id STRING
short_id STRING
title STRING
message STRING
author_name STRING
author_email STRING
authored_date TIMESTAMP
committer_name STRING
committer_email STRING
committed_date TIMESTAMP
created_at TIMESTAMP
parent_ids ARRAY<STRING>
trailers VARIANT
extended_trailers VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

epic_issues

Primary key: group_id, epic_iid, epic_issue_id Sync mode: Full refresh only

Field Data type
group_id BIGINT
epic_iid BIGINT
epic_issue_id BIGINT
id BIGINT
iid BIGINT
project_id BIGINT
title STRING
description STRING
state STRING
author STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
assignees ARRAY<STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>>
milestone STRUCT<id: BIGINT, iid: BIGINT, project_id: BIGINT, group_id: BIGINT, title: STRING, description: STRING, state: STRING, created_at: TIMESTAMP, updated_at: TIMESTAMP, due_date: DATE, start_date: DATE, expired: BOOLEAN, web_url: STRING>
labels ARRAY<STRING>
confidential BOOLEAN
weight BIGINT
discussion_locked BOOLEAN
user_notes_count BIGINT
upvotes BIGINT
downvotes BIGINT
due_date DATE
web_url STRING
time_stats STRUCT<time_estimate: BIGINT, total_time_spent: BIGINT, human_time_estimate: STRING, human_total_time_spent: STRING>
_links VARIANT
created_at TIMESTAMP
updated_at TIMESTAMP
closed_at TIMESTAMP

Note

This table does not support SCD type 2 because it has VARIANT columns.

pipelines_extended

Primary key: project_id, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
id BIGINT
iid BIGINT
name STRING
sha STRING
before_sha STRING
ref STRING
tag BOOLEAN
status STRING
source STRING
created_at TIMESTAMP
updated_at TIMESTAMP
started_at TIMESTAMP
finished_at TIMESTAMP
committed_at TIMESTAMP
duration BIGINT
queued_duration DOUBLE
coverage STRING
yaml_errors STRING
web_url STRING
user STRUCT<id: BIGINT, username: STRING, public_email: STRING, name: STRING, state: STRING, locked: BOOLEAN, avatar_url: STRING, web_url: STRING>
detailed_status STRUCT<icon: STRING, text: STRING, label: STRING, group: STRING, tooltip: STRING, has_details: BOOLEAN, details_path: STRING, favicon: STRING>
archived BOOLEAN

group_iterations

Primary key: group_id, id Cursor field: updated_at

Field Data type
group_id BIGINT
id BIGINT
iid BIGINT
sequence BIGINT
title STRING
description STRING
state BIGINT
created_at TIMESTAMP
updated_at TIMESTAMP
start_date DATE
due_date DATE
web_url STRING

epic_boards

Primary key: group_id, id Sync mode: Full refresh only

Field Data type
group_id BIGINT
id BIGINT
name STRING
hide_backlog_list BOOLEAN
hide_closed_list BOOLEAN
group STRUCT<id: BIGINT, name: STRING, web_url: STRING>
labels ARRAY<STRUCT<id: BIGINT, name: STRING, color: STRING, description: STRING, text_color: STRING>>
lists ARRAY<STRUCT<id: BIGINT, label: STRUCT<id: BIGINT, name: STRING, color: STRING, description: STRING, text_color: STRING>, position: BIGINT, list_type: STRING>>

merge_request_discussions

Primary key: project_id, merge_request_iid, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
merge_request_iid BIGINT
id STRING
individual_note BOOLEAN
notes VARIANT

Note

This table does not support SCD type 2 because it has VARIANT columns.

vulnerabilities

Primary key: project_id, id Sync mode: Full refresh only

Field Data type
project_id BIGINT
id STRING
title STRING
description STRING
state STRING
severity STRING
reportType STRING
detectedAt TIMESTAMP
confirmedAt TIMESTAMP
resolvedAt TIMESTAMP
dismissedAt TIMESTAMP
dismissalReason STRING

compliance_frameworks

Primary key: group_id, id Sync mode: Full refresh only

Field Data type
group_id BIGINT
id STRING
name STRING
description STRING
color STRING
default BOOLEAN
pipelineConfigurationFullPath STRING