VoiceAgentInputTranscription Class

Asynchronous input-audio transcription configuration. Extends the OpenAI Realtime transcription options with the Azure and MAI transcription models, custom speech models, and phrase hints.

Constructor

VoiceAgentInputTranscription(*args: Any, **kwargs: Any)

Variables

Name Description
language
str

The language of the input audio. Supplying the input language in ISO-639-1 (e.g. en) format will improve accuracy and latency.

languages

Possible languages of the input audio, in ISO-639-1 format. Supported by gpt-transcribe and gpt-live-transcribe.

keywords

Words or phrases to guide transcription of the input audio. Supported by gpt-transcribe and gpt-live-transcribe.

prompt
str

An optional text to guide the model's style or continue a previous audio segment. For whisper-1, the prompt is a list of keywords. For gpt-4o-transcribe models (excluding gpt-4o-transcribe-diarize), the prompt is a free text string, for example "expect words related to technology". Prompt is not supported with gpt-realtime-whisper in GA Realtime sessions.

delay
str or str or str or str or str

Controls how long the model waits before emitting transcription text. Higher values can improve transcription accuracy at the cost of latency. Only supported with gpt-realtime-whisper in GA Realtime sessions. Is one of the following types: Literal["minimal"], Literal["low"], Literal["medium"], Literal["high"], Literal["xhigh"]

model

The transcription model identifier. Configure customer custom speech deployments in custom_speech. Required. Known values are: "whisper-1", "gpt-realtime-whisper", "gpt-4o-transcribe", "gpt-4o-mini-transcribe", "gpt-4o-transcribe-diarize", "gpt-transcribe", "gpt-live-transcribe", "mai-transcribe", and "azure-speech".

custom_speech

Optional customer custom speech deployment configuration, keyed by locale.

phrase_list

Optional phrase hints that bias recognition toward domain terms.

Methods

as_dict

Return a dict that can be turned into json using json.dump.

clear

Remove all items from the dictionary.

copy
get

Get the value for key if key is in the dictionary, else default. :param str key: The key to look up. :param any default: The value to return if key is not in the dictionary. Defaults to None :returns: The value for key if key is in the dictionary, else default. :rtype: any

items
keys
pop

Removes specified key and return the corresponding value. :param str key: The key to pop. :param any default: The value to return if key is not in the dictionary :returns: The value corresponding to the key. :rtype: any :raises KeyError: If key is not found and default is not given.

popitem

Removes and returns some (key, value) pair :returns: The (key, value) pair. :rtype: tuple :raises KeyError: if the dictionary is empty.

setdefault

Return the value for key if key is in the dictionary; otherwise set the key to default and return default. :param str key: The key to look up. :param any default: The value to set if key is not in the dictionary :returns: The value for key if key is in the dictionary, else default. :rtype: any

update

Update the dictionary from a mapping or an iterable of key-value pairs. :param any args: Either a mapping object or an iterable of key-value pairs.

values

as_dict

Return a dict that can be turned into json using json.dump.

as_dict(*, exclude_readonly: bool = False) -> dict[str, Any]

Keyword-Only Parameters

Name Description
exclude_readonly

Whether to remove the readonly properties.

Default value: False

Returns

Type Description

A dict JSON compatible object

clear

Remove all items from the dictionary.

clear() -> None

copy

copy() -> Model

get

Get the value for key if key is in the dictionary, else default. :param str key: The key to look up. :param any default: The value to return if key is not in the dictionary. Defaults to None :returns: The value for key if key is in the dictionary, else default. :rtype: any

get(key: str, default: Any = None) -> Any

Parameters

Name Description
key
Required
default
Default value: None

items

items() -> ItemsView[str, Any]

Returns

Type Description

a set-like object providing a view on the mapping's items

keys

keys() -> KeysView[str]

Returns

Type Description

a set-like object providing a view on the mapping's keys

pop

Removes specified key and return the corresponding value. :param str key: The key to pop. :param any default: The value to return if key is not in the dictionary :returns: The value corresponding to the key. :rtype: any :raises KeyError: If key is not found and default is not given.

pop(key: str, default: ~typing.Any = <object object>) -> Any

Parameters

Name Description
key
Required
default

popitem

Removes and returns some (key, value) pair :returns: The (key, value) pair. :rtype: tuple :raises KeyError: if the dictionary is empty.

popitem() -> tuple[str, Any]

setdefault

Return the value for key if key is in the dictionary; otherwise set the key to default and return default. :param str key: The key to look up. :param any default: The value to set if key is not in the dictionary :returns: The value for key if key is in the dictionary, else default. :rtype: any

setdefault(key: str, default: ~typing.Any = <object object>) -> Any

Parameters

Name Description
key
Required
default

update

Update the dictionary from a mapping or an iterable of key-value pairs. :param any args: Either a mapping object or an iterable of key-value pairs.

update(*args: Any, **kwargs: Any) -> None

values

values() -> ValuesView[Any]

Returns

Type Description

an object providing a view on the mapping's values

Attributes

custom_speech

Optional customer custom speech deployment configuration, keyed by locale.

custom_speech: dict[str, str] | None

delay

Controls how long the model waits before emitting transcription text. Higher values can improve transcription accuracy at the cost of latency. Only supported with gpt-realtime-whisper in GA Realtime sessions. Is one of the following types: Literal["minimal"], Literal["low"], Literal["medium"], Literal["high"], Literal["xhigh"]

delay: Literal['minimal', 'low', 'medium', 'high', 'xhigh'] | None

keywords

Words or phrases to guide transcription of the input audio. Supported by gpt-transcribe and gpt-live-transcribe.

keywords: list[str] | None

language

The language of the input audio. Supplying the input language in ISO-639-1 (e.g. en) format will improve accuracy and latency.

language: str | None

languages

Possible languages of the input audio, in ISO-639-1 format. Supported by gpt-transcribe and gpt-live-transcribe.

languages: list[str] | None

model

The transcription model identifier. Configure customer custom speech deployments in custom_speech. Required. Known values are: "whisper-1", "gpt-realtime-whisper", "gpt-4o-transcribe", "gpt-4o-mini-transcribe", "gpt-4o-transcribe-diarize", "gpt-transcribe", "gpt-live-transcribe", "mai-transcribe", and "azure-speech".

model: str | _models.VoiceAgentInputTranscriptionModel

phrase_list

Optional phrase hints that bias recognition toward domain terms.

phrase_list: list[str] | None

prompt

An optional text to guide the model's style or continue a previous audio segment. For whisper-1, the prompt is a list of keywords. For gpt-4o-transcribe models (excluding gpt-4o-transcribe-diarize), the prompt is a free text string, for example "expect words related to technology". Prompt is not supported with gpt-realtime-whisper in GA Realtime sessions.

prompt: str | None