ai_extract 함수

적용 대상:체크 마크가 표시된 예 Databricks SQL 체크 마크가 표시된 예 Databricks Runtime

이 함수는 ai_extract() 사용자가 제공하는 스키마에 따라 텍스트 및 문서에서 구조화된 데이터를 추출합니다. 기본 추출에 간단한 필드 이름을 사용하거나, 중첩된 개체, 배열, 형식 유효성 검사 및 청구서, 계약 및 재무 서류와 같은 비즈니스 문서에 대한 필드 설명을 사용하여 복잡한 스키마를 정의할 수 있습니다.

이 함수는 종단 간 문서 처리를 위해 작성 가능한 워크플로를 사용하도록 설정하는 등 VARIANT다른 AI 함수의 텍스트 또는 ai_parse_document 출력을 허용합니다.

결과의 ai_extract유효성을 검사하고 반복하는 시각적 UI는 정보 추출을 참조하세요.

요구 사항

Apache 2.0 라이선스

현재 사용할 수 있는 기본 모델은 Apache 2.0 라이선스인 Copyright © The Apache Software Foundation에 따라 라이선스가 부여됩니다. 고객은 해당 모델 라이선스를 준수할 책임이 있습니다.

Databricks는 해당 조건을 준수하도록 이러한 라이선스를 검토할 것을 권장합니다. Databricks의 내부 벤치마크에 따라 더 나은 성능을 제공하는 모델이 향후에 나타날 경우 Databricks는 모델(및 이 페이지에 제공된 해당 라이선스 목록)을 변경할 수 있습니다.

이 함수를 구동하는 모델은 Model Serving Foundation 모델 API를 사용하여 사용할 수 있습니다. Databricks에서 사용할 수 있는 모델과 해당 모델의 사용을 제어하는 라이선스 및 정책에 대한 자세한 내용은 해당 모델 용어를 참조하세요.

Databricks의 내부 벤치마크에 따라 더 나은 성능을 제공하는 모델이 미래에 등장할 경우 Databricks는 모델을 변경하고 설명서를 업데이트할 수 있습니다.

  • 이 함수는 일부 지역에서만 사용할 수 있습니다. AI 함수 가용성을 참조하세요.
  • 향상된 보안 및 규정 준수 추가 기능이 있는 작업 영역의 경우
  • 이 함수는 Azure Databricks SQL 클래식에서 사용할 수 없습니다.
  • Databricks 런타임 15.4 LTS 이상이 필요합니다. 최고의 성능과 최신 기능 접근을 위해 Databricks Runtime 18.2 이상이 권장됩니다.
  • Databricks SQL 가격 페이지를 확인하세요.

Tip

Databricks는 버전 2.0 이상을ai_extract사용하는 것이 좋습니다. 버전 1.0은 이러한 기능을 지원하지 않으며 새 워크로드 또는 프로덕션 워크로드에는 권장되지 않는 레거시 인터페이스입니다.

버전 2.0 이상에서 지원합니다.

  • 중첩된 개체, 배열, 형식 유효성 검사 및 필드 설명이 있는 리치 JSON 스키마
  • 다음과 같은 업스트림 AI 함수에서 입력을 모두 STRINGVARIANT 수락합니다. ai_parse_document

버전 2.1은 인용 및 신뢰도 점수도 추가합니다.

버전을 명시적으로 고정하려면 .를 전달합니다 options => map('version', '2.1').

데이터 보안

문서 데이터는 Databricks 보안 경계 내에서 처리됩니다. Databricks는 AI 함수 호출에 전달되는 매개변수를 저장하지 않지만, Databricks 런타임 버전과 같은 메타데이터 실행 세부 사항은 유지합니다.

구문

ai_extract(content, schema [, options])

버전 2

ai_extract(content, schema [, options])

버전 1(레거시)

ai_extract(content, labels [, options])

논쟁

  • content: A VARIANT 또는 STRING 식입니다. 다음 중 하나를 수락합니다:

    • 원시 텍스트를 로
    • VARIANT 다른 AI 함수(예: ai_parse_document)에 의해 생성된 A
  • schema STRING: 추출을 위한 JSON 스키마를 정의하는 리터럴입니다. 스키마는 다음과 같습니다.

    • 단순 스키마: 필드 이름의 JSON 배열(문자열로 가정)
      "[\"vendor_name\", \"invoice_id\", \"total_amount\"]"
      
    • 고급 스키마: 형식 정보, 설명 및 중첩 구조체가 있는 JSON 개체
      • , string, integernumber및 boolean 형식을 enum지원합니다. 형식 유효성 검사를 수행합니다. 유효하지 않은 값은 오류가 발생합니다. 최대 500개 열거형 값입니다.
      • 다음을 사용하여 중첩된 개체를 "type": "object" 지원합니다. "properties"
      • 를 사용하여 기본 형식 또는 개체의 배열을 "type": "array" 지원합니다. "items"
      • 추출 품질을 안내하는 각 속성에 대한 선택적 "description" 필드
  • options: 구성 옵션을 포함하는 선택 사항 MAP<STRING, STRING> :

    • version: 마이그레이션을 지원하도록 버전 전환("2.1", "2.0", "1.0"). 기본값은 입력 형식을 기반으로합니다.
    • instructions: 추출 품질을 개선하기 위해 작업 및 도메인에 대한 전역 설명입니다. 20,000자 미만이어야 합니다.
    • mode: 로 설정 "precision" 되어 있으며, 긴 문서, 대용량 출력, 그리고 크고 깊이 중첩된 복잡한 스키마에 최적화되어 있습니다. 예를 들어, 정밀 모드를 사용해 송장에서 수백 개의 항목 항목을 추출하고, 여러 페이지에 걸쳐 명시된 계약 조건을 추출하며, 지표나 평가 계산과 같은 추론이 필요한 스키마를 채우는 식입니다.
    • enableCitations true: 추출 스키마의 각 필드에 대한 출력에 문서에서 추출된 출력을 나타내는 0개 이상의 인용 목록이 포함됩니다.
    • enableConfidenceScores true: 추출 스키마의 각 필드에 대한 출력에 0에서 1 사이의 신뢰도 점수가 포함되어 모델이 해당 값에 대해 얼마나 확실한지 나타냅니다. 적절한 신뢰도 임계값은 특정 사용 사례에 따라 달라지며 위험 및 오류에 대한 허용 오차와 일치하는 차단을 선택해야 합니다.

버전 2

  • content: A VARIANT 또는 STRING 식입니다. 다음 중 하나를 수락합니다:

    • 원시 텍스트를 로
    • VARIANT 다른 AI 함수(예: ai_parse_document)에 의해 생성된 A
  • schema STRING: 추출을 위한 JSON 스키마를 정의하는 리터럴입니다. 스키마는 다음과 같습니다.

    • 단순 스키마: 필드 이름의 JSON 배열(문자열로 가정)
      "[\"vendor_name\", \"invoice_id\", \"total_amount\"]"
      
    • 고급 스키마: 형식 정보, 설명 및 중첩 구조체가 있는 JSON 개체
      • , string, integernumber및 boolean 형식을 enum지원합니다. 형식 유효성 검사를 수행합니다. 유효하지 않은 값은 오류가 발생합니다. 최대 500개 열거형 값입니다.
      • 다음을 사용하여 중첩된 개체를 "type": "object" 지원합니다. "properties"
      • 를 사용하여 기본 형식 또는 개체의 배열을 "type": "array" 지원합니다. "items"
      • 추출 품질을 안내하는 각 속성에 대한 선택적 "description" 필드
  • options: 구성 옵션을 포함하는 선택 사항 MAP<STRING, STRING> :

    • version: 마이그레이션을 지원하도록 버전 스위치("1.0" v1 동작의 경우, "2.0" v2 동작의 경우). 기본값은 입력 형식을 기반으로 하지만 "1.0".
    • instructions: 추출 품질을 개선하기 위해 작업 및 도메인에 대한 전역 설명입니다. 20,000자 미만이어야 합니다.
    • mode: 로 "precision" 설정되어 있으며, 긴 문서, 대용량 출력, 복잡한 스키마 — 크고 깊이 중첩되었거나 추론이 필요한 상황에 최적화되어 있습니다. 예를 들어, 정밀 모드를 사용해 송장에서 수백 개의 항목 항목을 추출하고, 여러 페이지에 걸쳐 명시된 계약 조건을 추출하며, 지표나 평가 계산과 같은 추론이 필요한 스키마를 채우는 식입니다.

버전 1(레거시)

  • content STRING: 원시 텍스트를 포함하는 식입니다.

  • labels: ARRAY<STRING> 리터럴 상수. 각 요소는 추출할 엔터티의 형식입니다.

  • options: 구성 옵션을 포함하는 선택 사항 MAP<STRING, STRING> :

    • version: 마이그레이션을 지원하도록 버전 스위치("1.0" v1 동작의 경우, "2.0" v2 동작의 경우). 기본값은 입력 형식을 기반으로 하지만 ."1.0"

반품

포함하는 항목을 VARIANT 반환합니다.

{
  "response": {...},       // Extracted data matching the provided schema. Each leaf is returned as a Field object (see below).
  "error_message": null,   // null on success, or error message on failure
  "metadata": { ... }      // Metadata about the response, including version and citation details.
}

필드에는 response 스키마에 따라 추출된 구조화된 데이터가 포함됩니다.

  • 필드 이름 및 형식이 스키마 정의와 일치합니다.
  • 스키마의 구조는 응답에서 유지됩니다. 중첩된 개체와 배열은 원래 모양을 유지합니다. 추출 스키마의 각 "스칼라" 필드에는 다음 필드가 있는 출력 개체가 있습니다.
    • value: 스키마에 따라 입력된 추출된 값입니다. Null 필드를 추출할 수 없는 경우
    • citation_ids: 있는 경우에만 enableCitations 존재합니다 true. 인덱싱된 ID의 배열입니다 metadata.citations.
    • confidence_score: 있는 경우에만 enableConfidenceScores 존재합니다 true. 0에서 1 사이의 부동 소수입니다 .
  • 정수, 숫자, 부울 및 열거형 형식에 대해 형식 유효성 검사가 적용됩니다.
  • content가 NULL이면 결과는 NULL입니다.

필드에는 metadata 응답에 대한 메타데이터가 포함됩니다. 요청에서 설정된 경우 mode , metadata 사용된 것을 mode 포함합니다. 이 경우 enableCitationstrue추출된 값을 다시 입력의 response 위치로 추적하는 필드의 각 인용 ID에 대한 세부 정보도 포함됩니다.

입력 유형에 따라 인용은 다음 두 가지 유형 중 하나일 수 있습니다.

  • 원시 텍스트(STRING) 입력의 경우 인용은 원래 입력의 텍스트 범위입니다. metadata.citations의 각 개체에는 다음이 포함됩니다.
    • id: 필드의 citation_ids 항목과 일치하는 정수입니다.
    • start: 입력 문자열에 포함되는 0부터 시작하는 문자 오프셋입니다.
    • stop: 입력 문자열에 대한 배타적인 0부터 시작하는 문자 오프셋입니다.
  • PDF 문서 및 이미지(다운스트림을 사용하는 ai_extract 경우)의 ai_parse_document경우 인용은 원래 입력의 경계 상자입니다. 각 개체에는 metadata.citations 다음이 포함됩니다.
    • id: 필드의 항목과 citation_ids 일치하는 정수입니다.
    • bbox: 출력의 {coord, page_id} element.bbox와 도형이 동일한 개체의 ai_parse_document 배열입니다. coord 는 0부터 시작하는 페이지 인덱스처럼 [x0, y0, x1, y1]; page_id 페이지 이미지의 픽셀 좌표입니다.

버전 2

포함하는 항목을 VARIANT 반환합니다.

{
  "response": { ... },   // Extracted data matching the provided schema
  "metadata": {
    "version": "2.0"     // Function version used
  },
  "error_message": null  // null on success, or error message on failure
}

필드에는 response 스키마에 따라 추출된 구조화된 데이터가 포함됩니다.

  • 필드 이름 및 형식이 스키마 정의와 일치합니다.
  • 중첩된 개체 및 배열은 구조에 유지됩니다.
  • 필드를 찾을 수 null 없는 경우
  • , integernumber및 boolean 형식에 대해 enum형식 유효성 검사가 적용됩니다.

필드에는 metadata 응답에 대한 메타데이터가 포함됩니다.

content가 NULL이면 결과는 NULL입니다.

버전 1(레거시)

각 필드가 STRUCT 에 지정된 labels엔터티 형식에 해당하는 위치를 반환합니다. 각 필드에는 추출된 엔터티를 나타내는 문자열이 포함됩니다. 함수가 엔터티 형식에 대해 둘 이상의 후보를 찾으면 하나만 반환합니다.

예제

단순 스키마 - 필드 이름만

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '["invoice_id", "vendor_name", "total_amount", "invoice_date"]',
    options => map('version', '2.1')
  );
 {
   "response": {
     "invoice_id":   {"value": "12345"},
     "vendor_name":  {"value": "Acme Corp"},
     "total_amount": {"value": "1250.00"},
     "invoice_date": {"value": "2024-01-15"}
   },
   "error_message": null
 }

고급 스키마 - 형식 및 설명 포함

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '{
      "invoice_id": {"type": "string", "description": "Unique invoice identifier"},
      "vendor_name": {"type": "string", "description": "Legal business name"},
      "total_amount": {"type": "number", "description": "Total invoice amount"},
      "invoice_date": {"type": "string", "description": "Date in YYYY-MM-DD format"}
    }',
    options => map('version', '2.1')
  );
 {
   "response": {
     "invoice_id":   {"value": "12345"},
     "vendor_name":  {"value": "Acme Corp"},
     "total_amount": {"value": 1250.00},
     "invoice_date": {"value": "2024-01-15"}
   },
   "error_message": null
 }

중첩된 개체 및 배열

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp
     Line 1: Widget A, qty 10, $50.00 each
     Line 2: Widget B, qty 5, $100.00 each
     Subtotal: $1,000.00, Tax: $80.00, Total: $1,080.00',
    '{
      "invoice_header": {
        "type": "object",
        "properties": {
          "invoice_id": {"type": "string"},
          "vendor_name": {"type": "string"}
        }
      },
      "line_items": {
        "type": "array",
        "description": "List of invoiced products",
        "items": {
          "type": "object",
          "properties": {
            "description": {"type": "string"},
            "quantity": {"type": "integer"},
            "unit_price": {"type": "number"}
          }
        }
      },
      "totals": {
        "type": "object",
        "properties": {
          "subtotal": {"type": "number"},
          "tax_amount": {"type": "number"},
          "total_amount": {"type": "number"}
        }
      }
    }',
    options => map('version', '2.1')
  );
 {
   "response": {
     "invoice_header": {
       "invoice_id":  {"value": "12345"},
       "vendor_name": {"value": "Acme Corp"}
     },
     "line_items": [
       {"description": {"value": "Widget A"}, "quantity": {"value": 10}, "unit_price": {"value": 50.00}},
       {"description": {"value": "Widget B"}, "quantity": {"value": 5},  "unit_price": {"value": 100.00}}
     ],
     "totals": {
       "subtotal":     {"value": 1000.00},
       "tax_amount":   {"value": 80.00},
       "total_amount": {"value": 1080.00}
     }
   },
   "error_message": null
 }

을 사용하여 작성 ai_parse_document

> WITH parsed_docs AS (
    SELECT
      path,
      ai_parse_document(
        content,
        MAP('version', '2.0')
      ) AS parsed_content
    FROM READ_FILES('/Volumes/finance/invoices/', format => 'binaryFile')
  )
  SELECT
    path,
    ai_extract(
      parsed_content,
      '["invoice_id", "vendor_name", "total_amount"]',
      MAP('version', '2.1', 'instructions', 'These are vendor invoices.')
    ) AS invoice_data
  FROM parsed_docs;

열거형 사용

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp, amount: $1,250.00 USD',
    '{
      "invoice_id": {"type": "string"},
      "vendor_name": {"type": "string"},
      "total_amount": {"type": "number"},
      "currency": {
        "type": "enum",
        "labels": ["USD", "EUR", "GBP", "CAD", "AUD"],
        "description": "Currency code"
      },
      "payment_terms": {"type": "string"}
    }',
    options => map('version', '2.1')
  );
 {
   "response": {
     "invoice_id":     {"value": "12345"},
     "vendor_name":    {"value": "Acme Corp"},
     "total_amount":   {"value": 1250.00},
     "currency":       {"value": "USD"},
     "payment_terms":  {"value": null}
   },
   "error_message": null
 }

인용(STRING 입력, SPAN 인용)

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '{
      "invoice_id": {"type": "string", "description": "Unique invoice identifier"},
      "vendor_name": {"type": "string", "description": "Legal business name"},
      "total_amount": {"type": "number", "description": "Total invoice amount"},
      "invoice_date": {"type": "string", "description": "Date in YYYY-MM-DD format"}
    }',
   options => map(
     'version', '2.1',
     'enableCitations', 'true'
   )
  );
 {
   "response": {
     "invoice_id": {"citation_ids": [0], "value": "12345"},
     "vendor_name": {"citation_ids": [0], "value": "Acme Corp"},
     "total_amount": {"citation_ids": [1], "value": 1250.00},
     "invoice_date": {"citation_ids": [1], "value": "2024-01-15"}
   },
   "metadata": {
     "version": "2.1",
     "chunk_type": "span",
     "citations": [
       {"id": 0, "start": 0, "stop": 29},
       {"id": 1, "start": 29, "stop": 60}
     ]
   },
   "error_message": null
 }

인용(ai_parse_document VARIANT, BBOX 인용)

> WITH parsed AS (
    SELECT ai_parse_document(
             content,
             map('imageOutputPath', '/Volumes/main/default/parsed_images/')  // necessary for rendering bboxes
           ) AS doc
    FROM READ_FILES('/Volumes/main/default/invoices/invoice.pdf', format => 'binaryFile')
  )
  SELECT ai_extract(
    doc,
    '{"invoice_id":{"type":"string"}, "total_amount":{"type":"number"}}',
    options => map('version','2.1','enableCitations','true')
  ) AS extracted
  FROM parsed;
{
  "response": {
    "invoice_id":   {"citation_ids": [0], "value": "12345"},
    "total_amount": {"citation_ids": [1], "value": 1250.00}
  },
  "metadata": {
    "version": "2.1",
    "chunk_type": "bbox",
    "citations": [
      {"id": 0, "bbox": [{"coord": [120, 80,  240, 110], "page_id": 0}]},
      {"id": 1, "bbox": [{"coord": [400, 500, 560, 530], "page_id": 0}]}
    ],
    "pages": [{"id": 0, "image_uri": "/Volumes/main/default/parsed_images/6077ca79...f8efdb2ed05.jpg"}]
  },
  "error_message": null
}

신뢰도 점수


> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '{
      "invoice_id": {"type": "string", "description": "Unique invoice identifier"},
      "vendor_name": {"type": "string", "description": "Legal business name"},
      "total_amount": {"type": "number", "description": "Total invoice amount"},
      "invoice_date": {"type": "string", "description": "Date in YYYY-MM-DD format"}
    }',
   options => map(
    'version', '2.1',
    'enableConfidenceScores', 'true'
   )
  );
{
  "response": {
    "invoice_id": {"confidence_score": 0.95, "value": "12345"},
    "vendor_name": {"confidence_score": 0.62, "value": "Acme Corp"},
    "total_amount": {"confidence_score": 1.0, "value": 1250.00},
    "invoice_date": {"confidence_score": 0.99, "value": "2024-01-15"}
  },
  "error_message": null
}

정밀 모드

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '["invoice_id", "vendor_name", "total_amount"]',
   options => map(
    'version', '2.1',
    'mode', 'precision'
   )
  );
{
  "response": {
    "invoice_id":   {"value": "12345"},
    "vendor_name":  {"value": "Acme Corp"},
    "total_amount": {"value": 1250.00}
  },
  "metadata": {
    "version": "2.1",
    "mode": "precision"
  },
  "error_message": null
}

노트북 예제

다음 Notebook은 함수의 인용 출력을 분석하기 위한 시각적 디버깅 인터페이스를 ai_extract 제공합니다. 인용 메타데이터를 하위 문자열 코드 조각(STRING 입력) 또는 경계 상자 오버레이(VARIANT 입력)로 렌더링하고 SQL의 요소에 다시 조 ai_extract 인하여 수동 검토를 위해 ai_parse_document 저신뢰 추출에 플래그를 지정할 수 있는 방법을 보여 줍니다.

인용 렌더링 Notebook

노트북 받기

버전 2

단순 스키마 - 필드 이름만

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '["invoice_id", "vendor_name", "total_amount", "invoice_date"]'
  );
 {
   "response": {
     "invoice_id": "12345",
     "vendor_name": "Acme Corp",
     "total_amount": "1250.00",
     "invoice_date": "2024-01-15"
   },
   "metadata": {
     "version": "2.0"
   },
   "error_message": null
 }

고급 스키마 - 형식 및 설명 포함

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '{
      "invoice_id": {"type": "string", "description": "Unique invoice identifier"},
      "vendor_name": {"type": "string", "description": "Legal business name"},
      "total_amount": {"type": "number", "description": "Total invoice amount"},
      "invoice_date": {"type": "string", "description": "Date in YYYY-MM-DD format"}
    }'
  );
 {
   "response": {
     "invoice_id": "12345",
     "vendor_name": "Acme Corp",
     "total_amount": 1250.00,
     "invoice_date": "2024-01-15"
   },
   "metadata": {
     "version": "2.0"
   },
   "error_message": null
 }

중첩된 개체 및 배열

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp
     Line 1: Widget A, qty 10, $50.00 each
     Line 2: Widget B, qty 5, $100.00 each
     Subtotal: $1,000.00, Tax: $80.00, Total: $1,080.00',
    '{
      "invoice_header": {
        "type": "object",
        "properties": {
          "invoice_id": {"type": "string"},
          "vendor_name": {"type": "string"}
        }
      },
      "line_items": {
        "type": "array",
        "description": "List of invoiced products",
        "items": {
          "type": "object",
          "properties": {
            "description": {"type": "string"},
            "quantity": {"type": "integer"},
            "unit_price": {"type": "number"}
          }
        }
      },
      "totals": {
        "type": "object",
        "properties": {
          "subtotal": {"type": "number"},
          "tax_amount": {"type": "number"},
          "total_amount": {"type": "number"}
        }
      }
    }'
  );
 {
   "response": {
     "invoice_header": {
       "invoice_id": "12345",
       "vendor_name": "Acme Corp"
     },
     "line_items": [
       {"description": "Widget A", "quantity": 10, "unit_price": 50.00},
       {"description": "Widget B", "quantity": 5, "unit_price": 100.00}
     ],
     "totals": {
       "subtotal": 1000.00,
       "tax_amount": 80.00,
       "total_amount": 1080.00
     }
   },
   "metadata": {
     "version": "2.0"
   },
   "error": null
 }

을 사용하여 작성 ai_parse_document

> WITH parsed_docs AS (
    SELECT
      path,
      ai_parse_document(
        content,
        MAP('version', '2.0')
      ) AS parsed_content
    FROM READ_FILES('/Volumes/finance/invoices/', format => 'binaryFile')
  )
  SELECT
    path,
    ai_extract(
      parsed_content,
      '["invoice_id", "vendor_name", "total_amount"]',
      MAP('instructions', 'These are vendor invoices.')
    ) AS invoice_data
  FROM parsed_docs;

열거형 사용

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp, amount: $1,250.00 USD',
    '{
      "invoice_id": {"type": "string"},
      "vendor_name": {"type": "string"},
      "total_amount": {"type": "number"},
      "currency": {
        "type": "enum",
        "labels": ["USD", "EUR", "GBP", "CAD", "AUD"],
        "description": "Currency code"
      },
      "payment_terms": {"type": "string"}
    }'
  );
 {
   "response": {
     "invoice_id": "12345",
     "vendor_name": "Acme Corp",
     "total_amount": 1250.00,
     "currency": "USD",
     "payment_terms": null
   },
   "metadata": {
     "version": "2.0"
   },
   "error": null
 }

정밀 모드

> SELECT ai_extract(
    'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
    '["invoice_id", "vendor_name", "total_amount"]',
    options => map('version', '2.0', 'mode', 'precision')
  );
 {
   "response": {
     "invoice_id": "12345",
     "vendor_name": "Acme Corp",
     "total_amount": 1250.00
   },
   "metadata": {
     "version": "2.0",
     "mode": "precision"
   },
   "error_message": null
 }

버전 1(레거시)

> SELECT ai_extract(
    'John Doe lives in New York and works for Acme Corp.',
    array('person', 'location', 'organization')
  );
 {"person": "John Doe", "location": "New York", "organization": "Acme Corp."}

> SELECT ai_extract(
    'Send an email to jane.doe@example.com about the meeting at 10am.',
    array('email', 'time')
  );
 {"email": "jane.doe@example.com", "time": "10am"}

제한점

  • 이 함수는 Azure Databricks SQL 클래식에서 사용할 수 없습니다.
  • 이 함수는 뷰와 함께 사용할 수 없습니다.
  • 이 스키마는 최대 256개의 필드를 지원합니다.
  • 필드 이름은 최대 150자를 포함할 수 있습니다.
  • 스키마는 중첩된 필드에 대해 최대 12단계의 중첩을 지원합니다.
  • 열거형 필드는 최대 500개 값을 지원합니다.
  • 형식 유효성 검사는 , integer및 numberboolean 형식에 적용enum됩니다. 값이 지정된 형식과 일치하지 않으면 함수는 오류를 반환합니다.
  • 최대 전체 컨텍스트 크기는 100만 토큰입니다.

버전 2

  • 이 함수는 Azure Databricks SQL 클래식에서 사용할 수 없습니다.
  • 이 함수는 뷰와 함께 사용할 수 없습니다.
  • 이 스키마는 최대 256개의 필드를 지원합니다.
  • 필드 이름은 최대 150자를 포함할 수 있습니다.
  • 스키마는 중첩된 필드에 대해 최대 12단계의 중첩을 지원합니다.
  • 열거형 필드는 최대 500개 값을 지원합니다.
  • 형식 유효성 검사는 , integer및 numberboolean 형식에 적용enum됩니다. 값이 지정된 형식과 일치하지 않으면 함수는 오류를 반환합니다.
  • 최대 전체 컨텍스트 크기는 100만 토큰입니다.

버전 1(레거시)

  • 이 함수는 Azure Databricks SQL 클래식에서 사용할 수 없습니다.
  • 이 함수는 뷰와 함께 사용할 수 없습니다.
  • 콘텐츠에서 엔터티 형식에 대한 후보가 두 개 이상 있으면 하나의 값만 반환됩니다.