ai_transcribe function

Applies to: check marked yes Databricks SQL check marked yes Databricks Runtime

The ai_transcribe() function transcribes an audio file into text. It returns the transcript as a set of time-stamped segments and labels each segment with a speaker using automatic speaker diarization. The output is a VARIANT value, so you can chain it into other AI Functions.

Important

This feature is in Beta. To use it, a workspace admin must turn on AI Transcribe from the Previews page. See Manage Azure Databricks previews.

Data security

Your document data is processed within the Databricks security perimeter. Databricks does not store the parameters that are passed into the AI function calls, but does retain metadata run details, such as the Databricks Runtime version used.

Requirements

  • This function is only available in some regions, see AI function availability.
  • If you are using serverless compute, the following is also required:
    • The serverless environment version must be set to 3 or above, as this enables features like VARIANT.
    • Must use either Python or SQL. For additional serverless features and limitations, see Serverless compute limitations.
  • The ai_transcribe function is available using Databricks notebooks, SQL editor, Databricks workflows, jobs, or Lakeflow pipelines.
  • ai_transcribe costs are recorded as part of the AI_FUNCTIONS product.

Supported input file formats

For input data, use either a FILE type expression that references an audio file or a BINARY expression that contains the audio file's bytes. If the source files are stored in a Unity Catalog volume, generate the binary column using the Spark binaryFile format reader, as shown in the examples.

ai_transcribe transcribes audio only. Video files are not supported. During the Beta, the following audio formats are supported:

  • MP3
  • WAV
  • FLAC
  • Opus
  • M4A

Supported languages

During the Beta, ai_transcribe supports English and Spanish audio. Audio in other languages might return inaccurate transcripts.

Syntax

ai_transcribe(content)

Arguments

content is the only required argument.

  • content: A FILE type expression that references an audio file, or a BINARY expression that holds the audio file's bytes, such as the content column produced by read_files(..., format => 'binaryFile').

Returns

ai_transcribe returns a VARIANT value with the transcript broken into time-stamped segments. Each segment also carries a speaker_id label from automatic speaker diarization. See Speaker diarization.

The output has the following schema:

{
  "error_message": STRING,          // null on success; an error message string on failure
  "response": {
    "duration_seconds": DOUBLE,     // total audio duration, in seconds
    "segments": [
      {
        "start": DOUBLE,            // segment start time, in seconds
        "end": DOUBLE,              // segment end time, in seconds
        "speaker_id": STRING,       // speaker label ("1", "2", ...)
        "text": STRING              // transcript text for the segment
      }
    ]
  }
}

The response fields are as follows:

  • error_message: null when the call succeeds, or a string describing the failure.
  • response.duration_seconds: The total duration of the audio, in seconds.
  • response.segments: An array of transcript segments, in time order. To get the full transcript, concatenate the segments' text values in order.
  • segments[].start and segments[].end: The segment's start and end time, in seconds.
  • segments[].speaker_id: The speaker label for the segment. See Speaker diarization.
  • segments[].text: The transcript text for the segment.

Speaker diarization

ai_transcribe runs speaker diarization automatically: it detects distinct speakers in the audio and labels each segment with a speaker_id. Diarization is always on during the Beta and can't be turned on or off by users, so no configuration is required. Speakers are labeled "1", "2", "3", and so on in the order they first speak, and the labels are consistent across the whole recording.

Limitations

The following limitations apply during the Beta:

  • Each call accepts at most 1 hour of audio. Longer audio returns an error.
  • Each call accepts an input file of at most 512 MB. A larger file returns an error.
  • ai_transcribe transcribes audio only; video files are not supported.

Examples

The following example passes a FILE type column named audio from a table directly to ai_transcribe.

SELECT
  audio.uri AS path,
  ai_transcribe(audio) AS transcription
FROM my_catalog.my_schema.audio_files;

Note

The FILE type is in Beta.

The following example transcribes every audio file in a Unity Catalog volume. It reads the files as binary data with read_files(..., format => 'binaryFile') and passes the content column to ai_transcribe.

SELECT
  path,
  ai_transcribe(content) AS transcription
FROM READ_FILES('/Volumes/catalog/schema/volume/audio/', format => 'binaryFile');

The next example expands the transcript into one row per segment, with the speaker label, timestamps, and text as separate columns. It uses variant_explode table-valued function to un-nest the segments array.

WITH transcribed AS (
  SELECT
    path,
    ai_transcribe(content) AS result
  FROM READ_FILES('/Volumes/catalog/schema/volume/audio/', format => 'binaryFile')
)
SELECT
  transcribed.path,
  segment.value:start::DOUBLE AS start_seconds,
  segment.value:end::DOUBLE AS end_seconds,
  segment.value:speaker_id::STRING AS speaker_id,
  segment.value:text::STRING AS text
FROM
  transcribed,
  LATERAL variant_explode(transcribed.result:response.segments) AS segment;