你当前正在访问 Microsoft Azure Global Edition 技术文档网站。 如果需要访问由世纪互联运营的 Microsoft Azure 中国技术文档网站,请访问 https://docs.azure.cn。
Azure OpenAI Batch API 有效处理大规模和大容量处理任务。 它处理具有单独配额的异步请求组,并提供 24 小时的目标周转时间,成本比全球标准低 50%。 通过批处理,可以在单个文件中发送大量请求,而不是一次发送一个请求。 全局批量请求具有单独的排队令牌配额,因此您的在线工作负载不会受到影响。
关键用例包括:
大规模数据处理: 快速并行分析大规模数据集。
内容生成: 创建大量文本,例如产品说明或文章。
文档审阅和摘要: 自动查看和汇总冗长的文档。
客户支持自动化: 同时处理大量查询以加快响应速度。
数据提取和分析: 从大量的非结构化数据中提取和分析信息。
自然语言处理(NLP)任务: 对大型数据集执行情绪分析或翻译等任务。
营销和个性化: 大规模生成个性化内容和建议。
提示
如果你的批处理作业规模太大,以至于即使已将部署配额用满,仍会触及排队令牌限制,那么某些区域现已支持一项新功能,允许你以指数退避方式将多个批处理作业加入队列。
在已排队令牌配额可用后,可以自动创建并启动下一个批处理作业。 若要了解详细信息,请参阅 使用指数退避自动重试大型批处理作业。
重要
该服务旨在在 24 小时内处理批处理请求,但不会使需要更长时间的作业过期。 可以随时 取消 任务。 取消作业时,服务将取消任何剩余工时,并返回任何已完成的工作。 你需要为已完成的工作付费。
静态数据仍保留在指定的 Azure 地理区域内,而数据可能会在任何 Azure OpenAI 位置进行推理处理。 详细了解数据存储地。
批处理支持
全局批处理模型可用性
| 地区 | gpt-5.4-mini, 2026-03-17 | gpt-5.4, 2026-03-05 | gpt-5.1, 2025-11-13 | gpt-5, 2025-08-07 | o3, 2025-04-16 | o4-mini, 2025-04-16 | gpt-4.1, 2025-04-14 | gpt-4.1-nano, 2025-04-14 | gpt-4.1-mini, 2025-04-14 | o3-mini, 2025-01-31 | gpt-4o, 2024-05-13 | gpt-4o, 2024-08-06 | gpt-4o, 2024-11-20 | gpt-4o-mini, 2024-07-18 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| australiaeast | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| brazilsouth | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| canadaeast | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| centralus | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| eastus | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| eastus2 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| francecentral | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| germanywestcentral | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| japaneast | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| koreacentral | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| northcentralus | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| norwayeast | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| polandcentral | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| southafricanorth | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| southcentralus | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| 南印度 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| swedencentral | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| switzerlandnorth | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| uksouth | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| westeurope | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| westus | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| westus3 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
注意
虽然 Global Batch 支持较旧的 API 版本,但某些模型需要更新的 API 版本。 例如,o3-mini由于在此日期之后发布,所以与2024-10-21不兼容。 若要使用全局 Batch 访问较新的模型,请使用 v1 API。
功能支持
目前不支持以下功能:
- 与 Assistants API 集成。
- 与 Azure OpenAI on Your Data 功能集成。
批量部署
注意
在 Microsoft Foundry 门户中,批处理部署类型显示为 Global-Batch 和 Data Zone Batch。 若要详细了解 Azure OpenAI 部署类型,请参阅 部署类型指南。
提示
建议为所有全局批处理模型部署启用 动态配额 ,以帮助避免由于排队令牌配额不足而导致的作业失败。 使用动态配额,您的部署可以在有额外容量时战略性地利用更多的配额。 将动态配额设置为关闭时,您的部署将只能处理不超过您在创建部署时定义的排队令牌上限的请求。
先决条件
- Azure订阅 - 免费创建一个订阅。
- 已部署类型为
GlobalBatch或DataZoneBatch的模型资源。
准备批处理文件
与 微调一样,批处理使用 JSON 行 (.jsonl) 格式的文件。 下面是一些包含不同类型的受支持内容的示例文件:
输入格式
响应 API
{"custom_id": "task-0", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was Microsoft founded, and by whom?"}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was XBOX merged into Microsoft?"}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "What is Visual Basic?"}}
聊天完成接口
{"custom_id": "task-0", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was Microsoft founded?"}]}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was the first XBOX released?"}]}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "What is Altair Basic?"}]}}
custom_id 是必需的,以便您可以识别哪个单独的批处理请求与给定的响应相对应。 响应不会按照批处理文件中.jsonl定义的顺序返回。
model 应将属性设置为与要针对推理响应的全局 Batch 部署的名称匹配。
重要
必须设置该 model 属性以匹配您希望针对推断响应的全局 Batch 部署的名称。
批处理文件的每一行上必须存在相同的全局批处理模型部署名称。如果要以其他部署为目标,则必须在单独的批处理文件/作业中执行此操作。
为了获得最佳性能,我们建议提交大型文件进行批处理,而不是大量小型文件,每个文件中只有几行。
创建输入文件
在本文中,我们将创建一 test.jsonl 个名为的文件,并将上述标准输入代码块中的内容复制到该文件。 需要修改全局批处理部署名称并将其添加到文件的每一行。
创建批处理作业
准备好输入文件后,首先需要上传该文件,然后才能启动批处理作业。 文件上传可以通过编程方式或通过 Microsoft Foundry 门户完成。 此示例演示如何将文件直接上传到 Azure OpenAI 资源。 或者,您可以为 Azure OpenAI Batch 配置 Azure Blob 存储。
-
登录到 Microsoft Foundry。 确保New Foundry开关处于打开状态。 这些步骤适用于 Foundry(新版)。
在右上角选择“生成”
从左侧窗格中选择“模型”
选择 “批处理作业>”创建批处理作业
跟踪批处理作业进度
创建作业后,可以通过选择最近创建的作业的作业 ID 来监视作业的进度。 默认情况下,系统会将你带到最近创建的批处理作业的状态页。
可以在右侧窗格中跟踪作业状态。
检索批处理作业输出文件
一旦作业完成或达到终端状态,它将生成一个错误文件和一个输出文件。您可以通过选择带有向下箭头图标的相应按钮下载并审阅这些文件。
取消批处理
取消正在进行的批处理任务。 批次的状态将在 cancelling 状态下维持最多 10 分钟,然后转换为 cancelled,届时输出文件中可能会有部分(如果有)结果可用。
先决条件
- Azure订阅 - 免费创建一个订阅。
- Python 3.8 或更高版本
- 以下Python库:
openai - Jupyter Notebook
- Azure OpenAI 资源,其中已部署一个部署类型为
Global-Batch的模型。 有关此过程的帮助,请参阅 资源创建和模型部署指南 。
本文中的步骤旨在按顺序在 Jupyter Notebook 中运行。 因此,我们将仅在示例开头实例化 Azure OpenAI 客户端一次。 如果您想要不按顺序运行某个步骤,通常需要在调用中设置一个 Azure OpenAI 客户端。
即使已安装 OpenAI Python库,也可能需要将安装升级到最新版本:
!pip install openai --upgrade
准备批处理文件
与 微调一样,全局批处理使用 JSON 行 (.jsonl) 格式的文件。 下面是一些包含不同类型的受支持内容的示例文件:
输入格式
响应 API
{"custom_id": "task-0", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was Microsoft founded, and by whom?"}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was XBOX merged into Microsoft?"}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "What is Visual Basic?"}}
聊天完成接口
{"custom_id": "task-0", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was Microsoft founded?"}]}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was the first XBOX released?"}]}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "What is Altair Basic?"}]}}
custom_id 是必需的,以便您可以识别哪个单独的批处理请求与给定的响应相对应。 响应不会按照批处理文件中.jsonl定义的顺序返回。
model 应将属性设置为与要针对推理响应的全局 Batch 部署的名称匹配。
重要
必须设置该 model 属性以匹配您希望针对推断响应的全局 Batch 部署的名称。
批处理文件的每一行上必须存在相同的全局批处理模型部署名称。如果要以其他部署为目标,则必须在单独的批处理文件/作业中执行此操作。
为了获得最佳性能,我们建议提交大型文件进行批处理,而不是大量小型文件,每个文件中只有几行。
创建输入文件
在本文中,我们将创建一 test.jsonl 个名为的文件,并将上述标准输入代码块中的内容复制到该文件。 需要修改全局批处理部署名称并将其添加到文件的每一行。 将此文件保存在执行Jupyter Notebook的同一目录中。
上传批处理文件
准备好输入文件后,首先需要上传该文件,然后才能启动批处理作业。 文件上传可以通过编程方式或通过 Microsoft Foundry 门户完成。 此示例演示如何将文件直接上传到 Azure OpenAI 资源。 或者,您可以为 Azure OpenAI Batch 配置 Azure Blob 存储。
import os
from datetime import datetime
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
token_provider = get_bearer_token_provider(
DefaultAzureCredential(), "https://ai.azure.com/.default"
)
client = OpenAI(
base_url = "https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/",
api_key=token_provider,
)
# Upload a file with a purpose of "batch"
file = client.files.create(
file=open("test.jsonl", "rb"),
purpose="batch",
extra_body={"expires_after":{"seconds": 1209600, "anchor": "created_at"}} # Optional you can set to a number between 1209600-2592000. This is equivalent to 14-30 days
)
print(file.model_dump_json(indent=2))
print(f"File expiration: {datetime.fromtimestamp(file.expires_at) if file.expires_at is not None else 'Not set'}")
file_id = file.id
通过取消注释和添加 extra_body={"expires_after":{"seconds": 1209600, "anchor": "created_at"}} ,可将上传文件设置为在 14 天内过期。 如果未设置过期,则每个资源的最大输入批处理文件数限制为 500 个。 通过设置过期值,每个资源的输入批处理文件数将增加到每个资源的 10,000 个文件。 若要删除批处理输入文件限制,请使用 Batch 和 Azure Blob 存储。
输出:
{
"id": "file-655111ec9cfc44489d9af078f08116ef",
"bytes": 176064,
"created_at": 1743391067,
"filename": "test.jsonl",
"object": "file",
"purpose": "batch",
"status": "processed",
"expires_at": 1744600667,
"status_details": null
}
File expiration: 2025-04-13 23:17:47
创建批处理作业
成功上传文件后,可以提交文件进行批处理。
# Submit a batch job with the file
batch_response = client.batches.create(
input_file_id=file_id,
endpoint="/chat/completions", # While passing this parameter is required, the system will read your input file to determine if the chat completions or responses API is needed.
completion_window="24h",
# extra_body={"output_expires_after":{"seconds": 1209600, "anchor": "created_at"}} # Optional you can set to a number between 1209600-2592000. This is equivalent to 14-30 days
)
# Save batch ID for later use
batch_id = batch_response.id
print(batch_response.model_dump_json(indent=2))
注意
当前,完成窗口必须设置为 24h。 如果设置除了 24h 以外的任何其他值,您的作业将失败。 超过 24 小时的作业将继续执行,直到取消。
输出:
{
"id": "batch_6caaf24d-54a5-46be-b1b7-518884fcbdde",
"completion_window": "24h",
"created_at": 1722476583,
"endpoint": null,
"input_file_id": "file-655111ec9cfc44489d9af078f08116ef",
"object": "batch",
"status": "validating",
"cancelled_at": null,
"cancelling_at": null,
"completed_at": null,
"error_file_id": null,
"errors": null,
"expired_at": null,
"expires_at": 1722562983,
"failed_at": null,
"finalizing_at": null,
"in_progress_at": null,
"metadata": null,
"output_file_id": null,
"request_counts": {
"completed": 0,
"failed": 0,
"total": 0
}
}
如果你的批处理作业太大,导致即使在用尽部署配额后仍然达到排队令牌限制,那么某些地区现在支持一种新的快速失败功能,该功能允许你以指数退避的方式,对多个批处理作业进行排队。因此一旦一个大型批处理作业完成,就可以自动启动下一个批处理作业。 若要详细了解哪些区域支持此功能以及如何调整代码以利用此功能,请参阅 排队批处理作业。
跟踪批处理作业进度
成功创建批处理作业后,可以在 Studio 中或以编程方式监视其进度。 检查批处理作业进度时,建议在每个状态调用之间等待至少 60 秒。
import time
import datetime
status = "validating"
while status not in ("completed", "failed", "canceled"):
time.sleep(60)
batch_response = client.batches.retrieve(batch_id)
status = batch_response.status
print(f"{datetime.datetime.now()} Batch Id: {batch_id}, Status: {status}")
if batch_response.status == "failed":
for error in batch_response.errors.data:
print(f"Error code {error.code} Message {error.message}")
输出:
2024-07-31 21:48:32.556488 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: validating
2024-07-31 21:49:39.221560 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: in_progress
2024-07-31 21:50:53.383138 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: in_progress
2024-07-31 21:52:07.274570 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: in_progress
2024-07-31 21:53:21.149501 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: finalizing
2024-07-31 21:54:34.572508 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: finalizing
2024-07-31 21:55:35.304713 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: finalizing
2024-07-31 21:56:36.531816 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: finalizing
2024-07-31 21:57:37.414105 Batch Id: batch_6caaf24d-54a5-46be-b1b7-518884fcbdde, Status: completed
可以使用以下状态值:
| 地位 | 描述 |
|---|---|
validating |
在批处理开始之前,正在验证输入文件。 |
failed |
输入文件未能通过验证过程。 |
in_progress |
输入文件已成功验证,批处理当前正在运行。 |
finalizing |
批处理已完成,正在准备结果。 |
completed |
批处理已完成,结果已准备就绪。 |
expired |
批次未能在 24 小时的时间窗口内完成。 |
cancelling |
批处理正在 cancelled(可能需要最多 10 分钟才能生效)。 |
cancelled |
批次为 cancelled. |
若要检查作业状态详细信息,可以运行:
print(batch_response.model_dump_json(indent=2))
输出:
{
"id": "batch_6caaf24d-54a5-46be-b1b7-518884fcbdde",
"completion_window": "24h",
"created_at": 1722476583,
"endpoint": null,
"input_file_id": "file-9f3a81d899b4442f98b640e4bc3535dd",
"object": "batch",
"status": "completed",
"cancelled_at": null,
"cancelling_at": null,
"completed_at": 1722477429,
"error_file_id": "file-c795ae52-3ba7-417d-86ec-07eebca57d0b",
"errors": null,
"expired_at": null,
"expires_at": 1722562983,
"failed_at": null,
"finalizing_at": 1722477177,
"in_progress_at": null,
"metadata": null,
"output_file_id": "file-3304e310-3b39-4e34-9f1c-e1c1504b2b2a",
"request_counts": {
"completed": 3,
"failed": 0,
"total": 3
}
}
观察到 error_file_id 和 output_file_id 都存在,output_file_id 还是一个单独的项。 使用 error_file_id 来帮助调试批处理作业发生的任何问题。
检索批处理作业输出文件
import json
output_file_id = batch_response.output_file_id
if not output_file_id:
output_file_id = batch_response.error_file_id
if output_file_id:
file_response = client.files.content(output_file_id)
raw_responses = file_response.text.strip().split('\n')
for raw_response in raw_responses:
json_response = json.loads(raw_response)
formatted_json = json.dumps(json_response, indent=2)
print(formatted_json)
输出:
为简洁起见,我们将仅包括输出中的单个聊天完成响应。 如果按照本文中的步骤操作,应有三个响应,如下所示:
聊天补全
{
"custom_id": "task-0",
"response": {
"body": {
"choices": [
{
"content_filter_results": {
"hate": {
"filtered": false,
"severity": "safe"
},
"self_harm": {
"filtered": false,
"severity": "safe"
},
"sexual": {
"filtered": false,
"severity": "safe"
},
"violence": {
"filtered": false,
"severity": "safe"
}
},
"finish_reason": "stop",
"index": 0,
"logprobs": null,
"message": {
"content": "Microsoft was founded on April 4, 1975, by Bill Gates and Paul Allen in Albuquerque, New Mexico.",
"role": "assistant"
}
}
],
"created": 1722477079,
"id": "chatcmpl-9rFGJ9dh08Tw9WRKqaEHwrkqRa4DJ",
"model": "gpt-4o-2024-05-13",
"object": "chat.completion",
"prompt_filter_results": [
{
"prompt_index": 0,
"content_filter_results": {
"hate": {
"filtered": false,
"severity": "safe"
},
"jailbreak": {
"filtered": false,
"detected": false
},
"self_harm": {
"filtered": false,
"severity": "safe"
},
"sexual": {
"filtered": false,
"severity": "safe"
},
"violence": {
"filtered": false,
"severity": "safe"
}
}
}
],
"system_fingerprint": "fp_a9bfe9d51d",
"usage": {
"completion_tokens": 24,
"prompt_tokens": 27,
"total_tokens": 51
}
},
"request_id": "660b7424-b648-4b67-addc-862ba067d442",
"status_code": 200
},
"error": null
}
响应 API
{
"custom_id": "task-0",
"response": {
"body": {
"id": "resp_0e5c78eb05ee70cf00690cc6d988e4819587556df17436a206",
"created_at": 1762445017.0,
"error": null,
"incomplete_details": null,
"instructions": null,
"metadata": {},
"model": "gpt-4.1-batch",
"object": "response",
"output": [
{
"id": "msg_0e5c78eb05ee70cf00690cc6da3c548195aae483031113df16",
"content": [
{
"annotations": [],
"text": "Microsoft was founded on **April 4, 1975** by **Bill Gates** and **Paul Allen**.",
"type": "output_text",
"logprobs": []
}
],
"role": "assistant",
"status": "completed",
"type": "message"
}
],
"parallel_tool_calls": true,
"temperature": 1.0,
"tool_choice": "auto",
"tools": [],
"top_p": 1.0,
"background": false,
"max_output_tokens": null,
"max_tool_calls": null,
"previous_response_id": null,
"prompt_cache_key": null,
"reasoning": {
"effort": null,
"summary": null
},
"safety_identifier": null,
"service_tier": "default",
"status": "completed",
"text": {
"format": {
"type": "text"
},
"verbosity": "medium"
},
"top_logprobs": 0,
"truncation": "disabled",
"usage": {
"input_tokens": 16,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens": 25,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 41
},
"user": null,
"content_filters": null,
"store": true
},
"request_id": "809b30c2-fa0b-4613-b5cc-c30f6b780c9a",
"status_code": 200
},
"error": null
}
其他批处理命令
取消批处理
取消正在进行的批处理任务。 批次的状态将在 cancelling 状态下维持最多 10 分钟,然后转换为 cancelled,届时输出文件中可能会有部分(如果有)结果可用。
client.batches.cancel("batch_abc123") # set to your batch_id for the job you want to cancel
列出批处理
列出特定Azure OpenAI 资源的批处理作业。
client.batches.list()
在 Python 库中,列表方法是分页的。
列出所有作业:
all_jobs = []
# Automatically fetches more pages as needed.
for job in client.batches.list(
limit=20,
):
# Do something with job here
all_jobs.append(job)
print(all_jobs)
列出批处理(预览版)
使用 REST API 列出具有其他排序/筛选选项的所有批处理作业。
在下面的示例中,我们提供了 generate_time_filter 函数来简化筛选器的构造。 如果不想使用此函数,筛选器字符串的格式将如下所示 created_at gt 1728860560 and status eq 'Completed'。
import requests
import json
from datetime import datetime, timedelta
from azure.identity import DefaultAzureCredential
token_credential = DefaultAzureCredential()
token = token_credential.get_token('https://ai.azure.com/.default')
endpoint = "https://{YOUR_RESOURCE_NAME}.openai.azure.com/"
api_version = "2025-03-01-preview"
url = f"{endpoint}openai/batches"
order = "created_at asc"
time_filter = lambda: generate_time_filter("past 8 hours")
# Additional filter examples:
#time_filter = lambda: generate_time_filter("past 1 day")
#time_filter = lambda: generate_time_filter("past 3 days", status="Completed")
def generate_time_filter(time_range, status=None):
now = datetime.now()
if 'day' in time_range:
days = int(time_range.split()[1])
start_time = now - timedelta(days=days)
elif 'hour' in time_range:
hours = int(time_range.split()[1])
start_time = now - timedelta(hours=hours)
else:
raise ValueError("Invalid time range format. Use 'past X day(s)' or 'past X hour(s)'")
start_timestamp = int(start_time.timestamp())
filter_string = f"created_at gt {start_timestamp}"
if status:
filter_string += f" and status eq '{status}'"
return filter_string
filter = time_filter()
headers = {'Authorization': 'Bearer ' + token.token}
params = {
"api-version": api_version,
"$filter": filter,
"$orderby": order
}
response = requests.get(url, headers=headers, params=params)
json_data = response.json()
if response.status_code == 200:
print(json.dumps(json_data, indent=2))
else:
print(f"Request failed with status code: {response.status_code}")
print(response.text)
输出:
{
"data": [
{
"cancelled_at": null,
"cancelling_at": null,
"completed_at": 1729011896,
"completion_window": "24h",
"created_at": 1729011128,
"error_file_id": "file-472c0626-4561-4327-9e4e-f41afbfb30e6",
"expired_at": null,
"expires_at": 1729097528,
"failed_at": null,
"finalizing_at": 1729011805,
"id": "batch_4ddc7b60-19a9-419b-8b93-b9a3274b33b5",
"in_progress_at": 1729011493,
"input_file_id": "file-f89384af0082485da43cb26b49dc25ce",
"errors": null,
"metadata": null,
"object": "batch",
"output_file_id": "file-62bebde8-e767-4cd3-a0a1-28b214dc8974",
"request_counts": {
"total": 3,
"completed": 2,
"failed": 1
},
"status": "completed",
"endpoint": "/chat/completions"
},
{
"cancelled_at": null,
"cancelling_at": null,
"completed_at": 1729016366,
"completion_window": "24h",
"created_at": 1729015829,
"error_file_id": "file-85ae1971-9957-4511-9eb4-4cc9f708b904",
"expired_at": null,
"expires_at": 1729102229,
"failed_at": null,
"finalizing_at": 1729016272,
"id": "batch_6287485f-50fc-4efa-bcc5-b86690037f43",
"in_progress_at": 1729016126,
"input_file_id": "file-686746fcb6bc47f495250191ffa8a28e",
"errors": null,
"metadata": null,
"object": "batch",
"output_file_id": "file-04399828-ae0b-4825-9b49-8976778918cb",
"request_counts": {
"total": 3,
"completed": 2,
"failed": 1
},
"status": "completed",
"endpoint": "/chat/completions"
}
],
"first_id": "batch_4ddc7b60-19a9-419b-8b93-b9a3274b33b5",
"has_more": false,
"last_id": "batch_6287485f-50fc-4efa-bcc5-b86690037f43"
}
对批处理作业进行排队
如果你的批处理作业太大,导致即使在用尽部署配额后仍然达到排队令牌限制,那么某些地区现在支持一种新的“快速失败”功能,该功能允许你以指数退避的方式,对多个批处理作业进行排队。 一个大型批处理作业完成后,排队令牌配额再次可用,可以自动创建和启动下一个批处理作业。
旧行为:
- 已运行并使用所有可用于部署的令牌的大型批处理作业。
- 提交的新批处理作业。
- 新的批处理作业进入验证阶段,最长可能需要几分钟时间。
- 根据当前可用的配额检查新作业的令牌计数。
- 新批处理作业失败,错误报告为超出令牌限制。
新行为:
- 已运行并使用所有可用于部署的令牌的大型批处理作业
- 新的批处理作业已提交
- 与当前可用的批处理配额作业相比,新作业的大致令牌计数会快速失败,因此你可以更轻松地以编程方式处理重试。
区域支持
以下区域支持新的快速失败行为:
- australiaeast
- eastus
- germanywestcentral
- italynorth
- northcentralus
- polandcentral
- swedencentral
- switzerlandnorth
- eastus2
- westus
下面的代码演示了处理快速失败行为的基本机制,以使用指数退避策略自动重试和对批处理作业进行排队。
根据批处理作业的大小,可能需要进一步增加 max_retries 或更改此示例。
import time
from openai import BadRequestError
max_retries = 10
retries = 0
initial_delay = 5
delay = initial_delay
while True:
try:
batch_response = client.batches.create(
input_file_id=file_id,
endpoint="/chat/completions",
completion_window="24h",
)
# Save batch ID for later use
batch_id = batch_response.id
print(f"✅ Batch created successfully after {retries} retries")
print(batch_response.model_dump_json(indent=2))
break
except BadRequestError as e:
error_message = str(e)
# Check if it's a token limit error
if 'token_limit_exceeded' in error_message:
retries += 1
if retries >= max_retries:
print(f"❌ Maximum retries ({max_retries}) reached. Giving up.")
raise
print(f"⏳ Token limit exceeded. Waiting {delay} seconds before retry {retries}/{max_retries}...")
time.sleep(delay)
# Exponential backoff - increase delay for next attempt
delay *= 2
else:
# If it's a different error, raise it immediately
print(f"❌ Encountered non-token limit error: {error_message}")
raise
输出:
⏳ Token limit exceeded. Waiting 5 seconds before retry 1/10...
⏳ Token limit exceeded. Waiting 10 seconds before retry 2/10...
⏳ Token limit exceeded. Waiting 20 seconds before retry 3/10...
⏳ Token limit exceeded. Waiting 40 seconds before retry 4/10...
⏳ Token limit exceeded. Waiting 80 seconds before retry 5/10...
⏳ Token limit exceeded. Waiting 160 seconds before retry 6/10...
⏳ Token limit exceeded. Waiting 320 seconds before retry 7/10...
✅ Batch created successfully after 7 retries
{
"id": "batch_1e1e7b9f-d4b4-41fa-bd2e-8d2ec50fb8cc",
"completion_window": "24h",
"created_at": 1744402048,
"endpoint": "/chat/completions",
"input_file_id": "file-e2ba4ccaa4a348e0976c6fe3c018ea92",
"object": "batch",
"status": "validating",
"cancelled_at": null,
"cancelling_at": null,
"completed_at": null,
"error_file_id": "",
"errors": null,
"expired_at": null,
"expires_at": 1744488444,
"failed_at": null,
"finalizing_at": null,
"in_progress_at": null,
"metadata": null,
"output_file_id": "",
"request_counts": {
"completed": 0,
"failed": 0,
"total": 0
}
}
先决条件
- Azure订阅 - 免费创建一个订阅。
- Azure OpenAI 资源,其中已部署一个部署类型为
Global-Batch的模型。 有关此过程的帮助,请参阅 资源创建和模型部署指南 。
准备批处理文件
与 微调一样,全局批处理使用 JSON 行 (.jsonl) 格式的文件。 下面是一些包含不同类型的受支持内容的示例文件:
输入格式
响应 API
{"custom_id": "task-0", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was Microsoft founded, and by whom?"}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "When was XBOX merged into Microsoft?"}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/responses", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "input": "What is Visual Basic?"}}
聊天完成接口
{"custom_id": "task-0", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was Microsoft founded?"}]}}
{"custom_id": "task-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "When was the first XBOX released?"}]}}
{"custom_id": "task-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "REPLACE-WITH-MODEL-DEPLOYMENT-NAME", "messages": [{"role": "system", "content": "You are an AI assistant that helps people find information."}, {"role": "user", "content": "What is Altair Basic?"}]}}
custom_id 是必需的,以便您可以识别哪个单独的批处理请求与给定的响应相对应。 响应不会按照批处理文件中.jsonl定义的顺序返回。
model 应将属性设置为与要针对推理响应的全局 Batch 部署的名称匹配。
重要
必须设置该 model 属性以匹配您希望针对推断响应的全局 Batch 部署的名称。
批处理文件的每一行上必须存在相同的全局批处理模型部署名称。如果要以其他部署为目标,则必须在单独的批处理文件/作业中执行此操作。
为了获得最佳性能,我们建议提交大型文件进行批处理,而不是大量小型文件,每个文件中只有几行。
创建输入文件
在本文中,我们将创建一 test.jsonl 个名为的文件,并将上述标准输入代码块中的内容复制到该文件。 需要修改全局批处理部署名称并将其添加到文件的每一行。
上传批处理文件
准备好输入文件后,首先需要上传该文件,然后才能启动批处理作业。 文件上传可以通过编程方式或通过 Microsoft Foundry 门户完成。 此示例演示如何将文件直接上传到 Azure OpenAI 资源。 或者,您可以为 Azure OpenAI Batch 配置 Azure Blob 存储。
重要
请谨慎使用 API 密钥。 不要直接在代码中包括 API 密钥,并且从不公开发布。 如果使用 API 密钥,请安全地将其存储在Azure 密钥保管库中。 有关在应用中安全地使用 API 密钥的详细信息,请参阅 API 密钥和 Azure 密钥保管库。
有关 AI 服务安全性的详细信息,请参阅 请求 Azure AI 服务的身份验证。
curl -X POST https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/files \
-H "Content-Type: multipart/form-data" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-F "purpose=batch" \
-F "file=@C:\\batch\\test.jsonl;type=application/json" \
-F "expires_after.seconds=1209600" \
-F "expires_after.anchor=created_at"
上述代码假定 test.jsonl 文件的特定文件路径。 根据需要调整本地系统的文件路径。
通过添加可选的 "expires_after.seconds=1209600" 和 "expires_after.anchor=created_at" 参数,您可以将上传文件设置为在 14 天后过期。 如果未设置过期,则每个资源的最大批输入文件数限制为 500 个。 通过设置过期值,每个资源的批处理文件数将增加到每个资源 10,000 个文件。 可以设置为介于 1209600-2592000 之间的数字。 这相当于 14-30 天。 若要删除批处理输入文件限制,请使用 Batch 和 Azure Blob 存储。
输出:
{
"status": "processed",
"bytes": 817,
"purpose": "batch",
"filename": "test.jsonl",
"expires_at": 1744607747,
"id": "file-7733bc35e32841e297a62a9ee50b3461",
"created_at": 1743398147,
"object": "file"
}
跟踪文件上传状态
根据上传文件的大小,可能需要一些时间才能完全上传和处理。 若要检查文件上传状态,请运行:
curl https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/files/{file-id} \
-H "api-key: $AZURE_OPENAI_API_KEY"
输出:
{
"status": "processed",
"bytes": 686,
"purpose": "batch",
"filename": "test.jsonl",
"expires_at": 1744607747,
"id": "file-7733bc35e32841e297a62a9ee50b3461",
"created_at": 1721408291,
"object": "file"
}
创建批处理作业
成功上传文件后,可以提交文件进行批处理。
curl -X POST https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/batches \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input_file_id": "file-abc123",
"endpoint": "/chat/completions",
"completion_window": "24h",
"output_expires_after": {
"seconds": 1209600
},
"anchor": "created_at"
}'
您可以选择在此处添加 `"output_expires_after":{"seconds": 1209600},` 和 `"anchor": "created_at"`,以使输出文件在 14 天后过期。
注意
当前,完成窗口必须设置为 24h。 如果设置除了 24h 以外的任何其他值,您的作业将失败。 超过 24 小时的作业将继续执行,直到取消。
输出:
{
"cancelled_at": null,
"cancelling_at": null,
"completed_at": null,
"completion_window": "24h",
"created_at": "2024-07-19T17:13:57.2491382+00:00",
"error_file_id": null,
"expired_at": null,
"expires_at": "2024-07-20T17:13:57.1918498+00:00",
"failed_at": null,
"finalizing_at": null,
"id": "batch_fe3f047a-de39-4068-9008-346795bfc1db",
"in_progress_at": null,
"input_file_id": "file-21006e70789246658b86a1fc205899a4",
"errors": null,
"metadata": null,
"object": "batch",
"output_file_id": null,
"request_counts": {
"total": null,
"completed": null,
"failed": null
},
"status": "Validating"
}
跟踪批处理作业进度
成功创建批处理作业后,可以在 Studio 中或以编程方式监视其进度。 检查批处理作业进度时,建议在每个状态调用之间等待至少 60 秒。
curl https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/batches/{batch_id} \
-H "api-key: $AZURE_OPENAI_API_KEY"
输出:
{
"cancelled_at": null,
"cancelling_at": null,
"completed_at": null,
"completion_window": "24h",
"created_at": "2024-07-19T17:33:29.1619286+00:00",
"error_file_id": null,
"expired_at": null,
"expires_at": "2024-07-20T17:33:29.1578141+00:00",
"failed_at": null,
"finalizing_at": null,
"id": "batch_e0a7ee28-82c4-46a2-a3a0-c13b3c4e390b",
"in_progress_at": null,
"input_file_id": "file-c55ec4e859d54738a313d767718a2ac5",
"errors": null,
"metadata": null,
"object": "batch",
"output_file_id": null,
"request_counts": {
"total": null,
"completed": null,
"failed": null
},
"status": "Validating"
}
可以使用以下状态值:
| 地位 | 描述 |
|---|---|
validating |
在批处理开始之前,正在验证输入文件。 |
failed |
输入文件未能通过验证过程。 |
in_progress |
输入文件已成功验证,批处理当前正在运行。 |
finalizing |
批处理已完成,正在准备结果。 |
completed |
批处理已完成,结果已准备就绪。 |
expired |
批次未能在 24 小时的时间窗口内完成。 |
cancelling |
批处理正在 cancelled(可能需要最多 10 分钟才能生效)。 |
cancelled |
批次为 cancelled. |
检索批处理作业输出文件
curl https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/files/{output_file_id}/content \
-H "api-key: $AZURE_OPENAI_API_KEY" > batch_output.jsonl
其他批处理命令
取消批处理
取消正在进行的批处理任务。 批次的状态将在 cancelling 状态下维持最多 10 分钟,然后转换为 cancelled,届时输出文件中可能会有部分(如果有)结果可用。
curl -X POST https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/batches/{batch_id}/cancel \
-H "api-key: $AZURE_OPENAI_API_KEY"
列出批处理
列出给定Azure OpenAI 资源的现有批处理作业。
curl https://YOUR_RESOURCE_NAME.openai.azure.com/openai/v1/batches \
-H "api-key: $AZURE_OPENAI_API_KEY"
列表 API 调用会分页。 响应包含一个布尔值 has_more ,用于指示是否有更多结果可以迭代。
列出批处理(预览版)
使用 REST API 列出具有其他排序/筛选选项的所有批处理作业。
curl "YOUR_RESOURCE_NAME.openai.azure.com/batches?api-version=2025-04-01-preview&$filter=created_at%20gt%201728773533%20and%20created_at%20lt%201729032733%20and%20status%20eq%20'Completed'&$orderby=created_at%20asc" \
-H "api-key: $AZURE_OPENAI_API_KEY"
为了避免错误 URL rejected: Malformed input to a URL function 空间被 %20替换为 。
批处理限制
| 限制名称 | 限制值 |
|---|---|
| 最大的批处理输入文件 - (无过期) | 500 |
| 最大批处理输入文件数 -(已设置过期时间) | 一万 |
| 最大输入文件大小 | 200 MB |
| 最大输入文件大小 - 自带存储(BYOS) | 1GB |
| 每个文件的最大请求数 | 100,000 |
注意
批处理文件限制不适用于输出文件(例如 result.jsonl,和 error.jsonl)。 若要删除批处理输入文件限制,请使用 Batch 和 Azure Blob 存储。
批处理配额
该表显示批次配额限制。 全局批处理的配额值以入队令牌表示。 提交文件以进行批处理时,将计算文件中的令牌数。 在批处理任务达到终端状态之前,这些令牌会计入你的总排队令牌限制。
全局批处理
| 模型 | 企业和 MCA-E | 默认 | 通过信用卡进行的月度订阅 | MSDN 订阅 | 面向学生的Azure免费试用版 |
|---|---|---|---|---|---|
gpt-4.1 |
5B | 200M | 50M | 90K | N/A |
gpt-4.1 mini |
15B | 1B | 50M | 90K | N/A |
gpt-4.1-nano |
15B | 1B | 50M | 90K | N/A |
gpt-4o |
5B | 200M | 50M | 90K | N/A |
gpt-4o-mini |
15B | 1B | 50M | 90K | N/A |
gpt-4-turbo |
300兆字节 | 80M | 40M | 90K | N/A |
gpt-4 |
150M | 30M | 5米 | 100K | N/A |
o3-mini |
15B | 1B | 50M | 90K | N/A |
o4-mini |
15B | 1B | 50M | 90K | N/A |
gpt-5 |
5B | 200M | 50M | 90K | N/A |
gpt-5.1 |
5B | 200M | 50M | 90K | N/A |
gpt-5.2 |
5B | 200M | 50M | N/A | N/A |
gpt-5.4 |
5B | 200M | 50M | N/A | N/A |
gpt-5.4-mini |
5B | 200M | 50M | N/A | N/A |
gpt-5.4-nano |
5B | 200M | 50M | N/A | N/A |
gpt-5.5 |
5B | 200M | 50M | 90K | N/A |
B = 十亿 |M = 百万 |K = 千
数据区批处理
| 模型 | 企业和 MCA-E | 默认 | 通过信用卡进行的月度订阅 | MSDN 订阅 | 面向学生的Azure免费试用版 |
|---|---|---|---|---|---|
gpt-4.1 |
500兆字节 | 30M | 30M | 90K | N/A |
gpt-4.1-mini |
15亿 | 100兆字节 | 50M | 90K | N/A |
gpt-4o |
500兆字节 | 30M | 30M | 90K | N/A |
gpt-4o-mini |
15亿 | 100兆字节 | 50M | 90K | N/A |
o3-mini |
15亿 | 100兆字节 | 50M | 90K | N/A |
gpt-5 |
5B | 200M | 50M | 90K | N/A |
gpt-5.1 |
5B | 200M | 50M | 90K | N/A |
gpt-5.4 |
5B | 200M | 50M | N/A | N/A |
gpt-5.4-mini |
5B | 200M | 50M | N/A | N/A |
gpt-5.5 |
5B | 200M | 50M | 90K | N/A |
批处理对象
| 财产 | 类型 | 定义 |
|---|---|---|
id |
字符串 | 批处理的标识符。 |
object |
字符串 | batch |
endpoint |
字符串 | 批处理使用的API端点。 |
errors |
对象 | 批处理的错误信息(如果有)。 |
input_file_id |
字符串 | 批次的输入文件ID。 |
completion_window |
字符串 | 应处理该批次的时间范围。 |
status |
字符串 | 批处理的当前状态。 可能的值:validating、、failed、in_progress、finalizingcompleted、expiredcancelling。 cancelled |
output_file_id |
字符串 | 包含成功执行请求的输出的文件的 ID。 |
error_file_id |
字符串 | 包含出现错误的请求输出的文件的 ID。 |
created_at |
整数 | 创建此批时的时间戳(以 Unix 纪元秒为单位)。 |
in_progress_at |
整数 | 这批开始进展时的时间戳(以 Unix 纪元秒为单位)。 |
expires_at |
整数 | 此批的过期时间戳(以 Unix 纪元秒表示)。 |
finalizing_at |
整数 | 此批开始完成时的时间戳(以 Unix 纪元秒为单位)。 |
completed_at |
整数 | 此批完成时的时间戳(以 Unix 纪元秒为单位)。 |
failed_at |
整数 | 此批失败时的时间戳(以 Unix 纪元秒为单位)。 |
expired_at |
整数 | 此批过期时的时间戳(以 Unix 纪元秒为单位)。 |
cancelling_at |
整数 | 此批启动时 cancelling 的时间戳(以 Unix 纪元秒为单位)。 |
cancelled_at |
整数 | 此批处理 cancelled 时的时间戳(以 Unix epoch 秒表示)。 |
request_counts |
对象 | 对象结构:total
整数 批处理中的请求总数。 completed
整数 批处理中成功完成的请求数。 failed
整数 批处理中失败的请求数。 |
metadata |
映射 | 一组可以附加到批处理的键-值对。 此属性可用于以结构化格式存储有关批的其他信息。 |
常见问题(常见问题解答)
图像是否可以与批处理 API 一起使用?
此功能仅限于某些多模式模型。 可以通过 图像 URL 或图像的 base64 编码表示形式提供图像作为输入。
是否可以将批处理 API 与微调的模型配合使用?
批处理 API 目前不支持微调的模型。
是否可以使用批处理 API 来嵌入模型?
批处理 API 目前不支持微调的模型。
内容筛选是否适用于全局 Batch 部署?
是的。 与其他部署类型类似,可以创建内容筛选器并将其与全局 Batch 部署类型相关联。
是否可以请求其他配额?
是的,可以从 Foundry 门户中的配额页面获取。 可以在 配额和限制文章中找到默认配额分配。
如果 API 未在 24 小时内完成请求,会发生什么情况?
我们的目标是在 24 小时内处理这些请求;我们不会使需要更长时间的作业过期。 您可以随时取消任务。 取消作业时,任何剩余的工作将被取消,任何已完成的工作将被返回。 你将需要为任何已完成的工作付费。
使用批处理,可以队列多少个请求?
可以批处理的请求数没有固定限制,但这取决于已排队令牌配额。 您的入列令牌限额包括一次可以入列的最大输入令牌数。
批处理请求完成后,由于输入令牌已被清除,批处理速率限制将被重置。 限制取决于队列中的全局请求数。 如果批处理 API 队列快速处理你的批处理,则批处理速率限制会更快地重置。
故障 排除
在 status 是 completed 时,作业是成功的。 成功的作业仍将生成一个 error_file_id,但它将与一个包含零字节的空文件相关联。
发生作业失败时,你将在属性中找到 errors 有关失败的详细信息:
{
"value": [
{
"id": "batch_80f5ad38-e05b-49bf-b2d6-a799db8466da",
"completion_window": "24h",
"created_at": 1725419394,
"endpoint": "/chat/completions",
"input_file_id": "file-c2d9a7881c8a466285e6f76f6321a681",
"object": "batch",
"status": "failed",
"cancelled_at": null,
"cancelling_at": null,
"completed_at": 1725419955,
"error_file_id": "file-3b0f9beb-11ce-4796-bc31-d54e675f28fb",
"errors": {
"object": "list",
"data": [
{
"code": "empty_file",
"message": "The input file is empty. Please ensure that the batch contains at least one request."
}
]
},
"expired_at": null,
"expires_at": 1725505794,
"failed_at": null,
"finalizing_at": 1725419710,
"in_progress_at": 1725419572,
"metadata": null,
"output_file_id": "file-ef12af98-dbbc-4d27-8309-2df57feed572",
"request_counts": {
"total": 10,
"completed": null,
"failed": null
}
}
]
}
错误代码
| 错误代码 | 定义 |
|---|---|
invalid_json_line |
输入文件中的一行(或多行)无法分析为有效的 json。 请确保没有拼写错误、正确的左括号和右括号以及根据 JSON 标准引用,并重新提交请求。 |
too_many_tasks |
输入文件中的请求数超过允许的最大值为 100,000。 请确保请求总数低于 100,000,并重新提交作业。 |
url_mismatch |
输入文件中的某一行具有与其余行不匹配的 URL,或者输入文件中指定的 URL 与预期的终结点 URL 不匹配。 请确保所有请求 URL 都相同,并且它们与与 Azure OpenAI 部署关联的终结点 URL 匹配。 |
model_not_found |
找不到在输入文件的 model 属性中指定的 Azure OpenAI 模型部署名称。请确保此名称指向有效的 Azure OpenAI 模型部署。 |
duplicate_custom_id |
此请求的自定义 ID 是另一个请求中的自定义 ID 的副本。 |
empty_file |
输入文件为空。 请确保批处理至少包含一个请求。 |
model_mismatch |
在输入文件中此请求的 model 属性中指定的 Azure OpenAI 模型部署名称与该文件的其余部分不匹配。请确保批处理中的所有请求都指向请求的 model 属性中的同一“Foundry 模型中的 Azure OpenAI”模型部署。 |
invalid_request |
输入行的架构无效或部署 SKU 无效。 请确保输入文件中的请求属性与预期的输入属性匹配,并且用于批处理 API 请求的 Azure OpenAI 部署 SKU 为 globalbatch。 |
input_modified |
提交批处理作业后,Blob 输入已被修改。 |
input_no_permissions |
无法访问输入 blob。 请检查 权限 以及 Azure OpenAI 帐户与 Azure 存储 帐户之间的网络访问。 |
已知问题
使用 Azure CLI 部署的资源无法直接在 Azure OpenAI 全局批处理中使用。 这是因为使用此方法部署的资源具有不遵循
https://your-resource-name.openai.azure.com格式的端点子域。 此问题的解决方法是使用其他常见部署方法之一部署新的 Azure OpenAI 资源,该方法将在部署过程中正确处理子域设置。不支持 UTF-8-BOM 编码
jsonl文件。 JSON 行文件应使用 UTF-8 进行编码。 JSON RFC 规范并不正式支持使用字节顺序标记(BOM)编码文件,而 Azure OpenAI 当前将 BOM 编码的文件视为无效。 UTF-8-BOM 编码文件当前将返回通用错误消息:“验证失败:无法从输入文件中提取有效的模型部署名称。 请确保输入文件中的每一行都具有“模型”字段中指定的有效部署名称,并且部署名称在所有行中都是一致的。当使用自己的存储批处理输入数据时,一旦提交了批处理作业,如果修改了输入 blob,则评分作业会被服务终止。