ContentAnalyzerConfig Class

Configuration settings for an analyzer.

Constructor

ContentAnalyzerConfig(*args: Any, **kwargs: Any)

Variables

Name Description
return_details

Return all content details.

locales

List of locale hints for speech transcription.

enable_ocr

Enable optical character recognition (OCR).

enable_layout

Enable layout analysis.

enable_figure_description

Enable generation of figure description.

enable_figure_analysis

Enable analysis of figures, such as charts and diagrams.

enable_formula

Enable mathematical formula detection.

table_format

Representation format of tables in analyze result markdown. Known values are: "html" and "markdown".

chart_format

Representation format of charts in analyze result markdown. Known values are: "chartJs" and "markdown".

annotation_format

Representation format of annotations in analyze result markdown. Known values are: "none" and "markdown".

disable_face_blurring

Disable the default blurring of faces for privacy while processing the content.

estimate_field_source_and_confidence

Return field grounding source and confidence.

content_categories

Map of categories to classify the input content(s) against.

enable_segment

Enable segmentation of the input by contentCategories.

segment_per_page

Force segmentation of document content by page.

omit_content

Omit the content for this analyzer from analyze result. Only return content(s) from additional analyzers specified in contentCategories, if any.

workflow

Workflow used for content analysis. Known values are: "default" and "agentic".

allow_input_truncation

When true, input that exceeds the service's processable-unit limit is truncated to the limit and returned as a partial result with a warning, instead of failing. Field extraction and segmentation, where configured, run over the processed content and may be inaccurate. Defaults to false. Overridable per request by the allowInputTruncation query parameter.

allow_in_page_segments

Enable sub-page segmentation. When true, segments may cover a portion of a page instead of full pages.

chunking_strategy

Strategy for chunking document content into smaller units for RAG scenarios. When omitted, chunking is disabled.

Methods

as_dict

Return a dict that can be turned into json using json.dump.

clear

Remove all items from the dictionary.

copy
get

Get the value for key if key is in the dictionary, else default. :param str key: The key to look up. :param any default: The value to return if key is not in the dictionary. Defaults to None :returns: The value for key if key is in the dictionary, else default. :rtype: any

items
keys
pop

Removes specified key and return the corresponding value. :param str key: The key to pop. :param any default: The value to return if key is not in the dictionary :returns: The value corresponding to the key. :rtype: any :raises KeyError: If key is not found and default is not given.

popitem

Removes and returns some (key, value) pair :returns: The (key, value) pair. :rtype: tuple :raises KeyError: if the dictionary is empty.

setdefault

Return the value for key if key is in the dictionary; otherwise set the key to default and return default. :param str key: The key to look up. :param any default: The value to set if key is not in the dictionary :returns: The value for key if key is in the dictionary, else default. :rtype: any

update

Update the dictionary from a mapping or an iterable of key-value pairs. :param any args: Either a mapping object or an iterable of key-value pairs.

values

as_dict

Return a dict that can be turned into json using json.dump.

as_dict(*, exclude_readonly: bool = False) -> dict[str, Any]

Keyword-Only Parameters

Name Description
exclude_readonly

Whether to remove the readonly properties.

Default value: False

Returns

Type Description

A dict JSON compatible object

clear

Remove all items from the dictionary.

clear() -> None

copy

copy() -> Model

get

Get the value for key if key is in the dictionary, else default. :param str key: The key to look up. :param any default: The value to return if key is not in the dictionary. Defaults to None :returns: The value for key if key is in the dictionary, else default. :rtype: any

get(key: str, default: Any = None) -> Any

Parameters

Name Description
key
Required
default
Default value: None

items

items() -> ItemsView[str, Any]

Returns

Type Description

a set-like object providing a view on the mapping's items

keys

keys() -> KeysView[str]

Returns

Type Description

a set-like object providing a view on the mapping's keys

pop

Removes specified key and return the corresponding value. :param str key: The key to pop. :param any default: The value to return if key is not in the dictionary :returns: The value corresponding to the key. :rtype: any :raises KeyError: If key is not found and default is not given.

pop(key: str, default: ~typing.Any = <object object>) -> Any

Parameters

Name Description
key
Required
default

popitem

Removes and returns some (key, value) pair :returns: The (key, value) pair. :rtype: tuple :raises KeyError: if the dictionary is empty.

popitem() -> tuple[str, Any]

setdefault

Return the value for key if key is in the dictionary; otherwise set the key to default and return default. :param str key: The key to look up. :param any default: The value to set if key is not in the dictionary :returns: The value for key if key is in the dictionary, else default. :rtype: any

setdefault(key: str, default: ~typing.Any = <object object>) -> Any

Parameters

Name Description
key
Required
default

update

Update the dictionary from a mapping or an iterable of key-value pairs. :param any args: Either a mapping object or an iterable of key-value pairs.

update(*args: Any, **kwargs: Any) -> None

values

values() -> ValuesView[Any]

Returns

Type Description

an object providing a view on the mapping's values

Attributes

allow_in_page_segments

Enable sub-page segmentation. When true, segments may cover a portion of a page instead of full pages.

allow_in_page_segments: bool | None

allow_input_truncation

When true, input that exceeds the service's processable-unit limit is truncated to the limit and returned as a partial result with a warning, instead of failing. Field extraction and segmentation, where configured, run over the processed content and may be inaccurate. Defaults to false. Overridable per request by the allowInputTruncation query parameter.

allow_input_truncation: bool | None

annotation_format

"none" and "markdown".

annotation_format: str | _models.AnnotationFormat | None

chart_format

"chartJs" and "markdown".

chart_format: str | _models.ChartFormat | None

chunking_strategy

Strategy for chunking document content into smaller units for RAG scenarios. When omitted, chunking is disabled.

chunking_strategy: _models.ChunkingStrategy | None

content_categories

Map of categories to classify the input content(s) against.

content_categories: dict[str, '_models.ContentCategoryDefinition'] | None

disable_face_blurring

Disable the default blurring of faces for privacy while processing the content.

disable_face_blurring: bool | None

enable_figure_analysis

Enable analysis of figures, such as charts and diagrams.

enable_figure_analysis: bool | None

enable_figure_description

Enable generation of figure description.

enable_figure_description: bool | None

enable_formula

Enable mathematical formula detection.

enable_formula: bool | None

enable_layout

Enable layout analysis.

enable_layout: bool | None

enable_ocr

Enable optical character recognition (OCR).

enable_ocr: bool | None

enable_segment

Enable segmentation of the input by contentCategories.

enable_segment: bool | None

estimate_field_source_and_confidence

Return field grounding source and confidence.

estimate_field_source_and_confidence: bool | None

locales

List of locale hints for speech transcription.

locales: list[str] | None

omit_content

Omit the content for this analyzer from analyze result. Only return content(s) from additional analyzers specified in contentCategories, if any.

omit_content: bool | None

return_details

Return all content details.

return_details: bool | None

segment_per_page

Force segmentation of document content by page.

segment_per_page: bool | None

table_format

"html" and "markdown".

table_format: str | _models.TableFormat | None

workflow

"default" and "agentic".

workflow: str | _models.ContentAnalyzerWorkflow | None