你如何向代理程式提出查詢,取決於提供其服務的代理伺服器。 請在下表中找到您的經紀人,然後依照對應的章節操作。 欲了解代理伺服器,請參閱 Agent Server。
| 你的代理程式 | 託管於 | 如何查詢 |
|---|---|---|
使用 DurableAgentServer |
Databricks Apps 上的 Agent Runtime | 位於 /api/invocations 的 Invocation API |
使用 MLflow AgentServer 或 LongRunningAgentServer(舊版) |
Databricks 應用程式 | Databricks OpenAI 用戶端或 OpenAI 回應 API,請參閱 /responses |
| 部署至模型服務 (舊版) | 模型服務端點 | Databricks OpenAI 用戶端、REST API ai_query或 AI Playground |
託管在 Databricks Apps 上的代理程式需要 Databricks OAuth 令牌。 個人存取權杖不適用於 Databricks Apps。 要從指令碼、服務主體、其他應用程式或筆記本產生 OAuth 權杖,請參閱 使用權杖驗證連線至 Databricks 應用程式 API。
查詢使用 DurableAgentServer 的代理程式
使用 DurableAgentServer 的代理程式會為呼叫 API 提供服務。 你用 Agent Bricks CLI 建立的代理程式使用 DurableAgentServer,並agentbricks deploy將其部署到 agent-bricks-<name>,作為名為 的應用程式。 每次對 API 的請求都會啟動一次代理執行,稱為叫用。
| 終點 | Description |
|---|---|
POST /api/invocations |
開始叫用。 預設情況下,此要求會等候處理完成後傳回結果。 將 stream 設為在事件發生時接收事件,或將 background 設為立即返回。 |
GET /api/invocations/<id> |
回傳調用狀態,完成後回傳其輸出。 |
GET /api/invocations/<id>/events?after=<event-id> |
串流傳輸發生在 <event-id> 之後的已儲存事件。 利用這個端點重新連接串流。 |
要求本文
POST /api/invocations 的請求本文接受以下欄位。 伺服器會拒絕包含其他欄位的請求。
| Field | Description |
|---|---|
id |
Required. 一個你為每次呼叫產生的 UUID。 伺服器將 ID 視為冪等性金鑰:使用相同 ID 重新發送相同請求時,會回傳現有的呼叫,而非重新執行代理程式。 重複使用一個 ID 來執行不同請求時,會傳回 409 錯誤。 |
session_id |
此叫用所屬的對話。 共用 session ID 的呼叫會依序一次執行一個。 由 CLI 範本產生的代理程式需要此欄位。 |
input |
你的代理程式的輸入。 由 CLI 範本產生的代理程式接受訊息清單,或具有 messages 清單的物件。 |
stream |
將其設為 true,即可接收 Server-Sent Events(SSE)事件。 |
background |
設定為 true,以立即傳回帶有狀態 URL 的 202 回應,然後輪詢結果。 |
你的代理程式的處理常式定義了 input 的結構。 從 CLI 範本產生的代理程式當 input 為物件時會讀取以下欄位:
| Field | Description |
|---|---|
messages |
對話會轉交給代理人。 |
actor |
代理人讀取與寫入其長期記憶的身分。 如果你沒有傳入 actor,代理會使用會話 ID,因此記憶不會帶到新的會話。 將 actor 設定為來自你應用程式中已登入的使用者,而非使用者輸入的文字。 |
model |
用於此次調用的模型,而非代理程式碼中設定的模型。 |
resume |
對代理暫停以供人工輸入(例如核准工具呼叫)的回應。 當代理人暫停時,status 叫用的狀態為 interrupted。 使用相同的resume在新的叫用中傳送session_id以繼續。 |
要讓代理在多個會話中回憶其對使用者的了解,請將使用者的 ID 傳遞為 actor:
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"session_id": "support-case-123",
"input": {
"messages": [{ "role": "user", "content": "What does Databricks do?" }],
"actor": "user-42"
}
}
特工磚塊
若要從您的終端機測試已部署的代理程式,請使用 agentbricks endpoint invoke。 這個指令會尋找應用程式,並用你的 CLI 設定檔進行驗證。
agentbricks --profile <profile> endpoint invoke agent-bricks-<name> \
--path /api/invocations \
--json "{\"id\":\"$(uuidgen)\",\"session_id\":\"$(uuidgen)\",\"input\":[{\"role\":\"user\",\"content\":\"Hello\"}]}"
要串流回應,將 "stream":true 加入 JSON 主體中,並傳遞 --sse。 若要測試代理程式使用 agentbricks dev 在本機執行時,請將應用程式名稱替換成 --url http://localhost:8000。
REST API
取得應用程式的網址。 輸出中的 URL 欄位是呼叫 API 的基礎 URL。
agentbricks --profile <profile> deployments get agent-bricks-<name>為你的設定檔取得一個 OAuth 權杖。 輸出在
access_token欄位中包含該標記。databricks auth token --profile <profile>發送請求:
curl --request POST \ --url <app-url>/api/invocations \ --header 'Authorization: Bearer <OAuth token>' \ --header 'content-type: application/json' \ --data '{ "id": "550e8400-e29b-41d4-a716-446655440000", "session_id": "support-case-123", "input": [{ "role": "user", "content": "What does Databricks do?" }] }'
回應包含呼叫 id、其 status,以及一個 output 欄位,欄位中值為你的代理人回傳的值。 伺服器不需要輸出結構:你的處理器可以回傳任何可序列化的 JSON 值。 由 CLI 範本產生的代理會回傳一個包含以下欄位的物件:
-
output:代理人在此次呼叫中產生的訊息。 -
status:completed,或如果代理人暫停以等待人工輸入,則為interrupted。
Python
以下範例使用 Databricks SDK 查詢應用程式 URL、產生 OAuth 權杖,然後呼叫 invocations API。
WorkspaceClient 必須使用 OAuth 認證。
import uuid
import requests
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
app_url = w.apps.get("agent-bricks-<name>").url
session_id = str(uuid.uuid4())
response = requests.post(
f"{app_url}/api/invocations",
headers=w.config.authenticate(),
json={
"id": str(uuid.uuid4()),
"session_id": session_id,
"input": [{"role": "user", "content": "What does Databricks do?"}],
},
)
response.raise_for_status()
print(response.json()["output"])
要繼續對話,請發送下一則訊息,並附上相同的 session_id 和新的 id。
串流、在背景執行,並重新連線
-
流:設定
"stream": true。 回應是包含run.started和run.failed(或delta)事件的 SSE 串流,加上你的代理所發出的事件,例如帶有串流文字的run.completed事件。 每個事件都有一個識別碼。 - 背景執行:設定
"background": true。 伺服器回傳帶有status_url的202回應。 持續輪詢GET /api/invocations/<id>直到狀態為completed。 如果你也設定了"stream": true,回應中包含一個可供你讀取事件的events_url。 -
重新連線:如果串流斷線,請呼叫
GET /api/invocations/<id>/events?after=<event-id>,並傳入你最後收到的事件 ID。
以下 Python 範例串流回應:
with requests.post(
f"{app_url}/api/invocations",
headers=w.config.authenticate(),
json={
"id": str(uuid.uuid4()),
"session_id": session_id,
"input": [{"role": "user", "content": "Summarize our last conversation."}],
"stream": True,
},
stream=True,
) as response:
response.raise_for_status()
for line in response.iter_lines(decode_unicode=True):
if line.startswith("data: "):
print(line[len("data: "):])
如果你部署代理時有多個實例,請在 X-Routing-Key 標頭中傳送會話 ID,將會話中的所有請求導向同一個實例。
查詢使用舊版 MLflow 的代理人 AgentServer
本節適用於您使用舊版代理伺服器部署在 Databricks Apps 上的代理:MLflow AgentServer 或 LongRunningAgentServer,搭配 ResponsesAgent 介面。 這些代理在 /responses 為 OpenAI Responses API 提供服務。
LongRunningAgentServer 服務相同的 API,因此以下範例同樣適用於此。 它也支援背景執行:在請求中將 background 設為 true,然後以 GET /responses/<response-id>?stream=true&starting_after=<sequence-number> 取回回應,串流該序列號之後的事件。
Databricks OpenAI 客戶端
Databricks 推薦這些代理程式使用 Databricks OpenAI 用戶端。 在型號名稱中加入 apps/ 前綴。
from databricks.sdk import WorkspaceClient
from databricks_openai import DatabricksOpenAI
input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
app_name = "<agent-app-name>"
# The WorkspaceClient must use OAuth authentication.
w = WorkspaceClient()
client = DatabricksOpenAI(workspace_client=w)
# Non-streaming request
response = client.responses.create(model=f"apps/{app_name}", input=input_msgs)
print(response)
# Streaming request
streaming_response = client.responses.create(
model=f"apps/{app_name}", input=input_msgs, stream=True
)
for chunk in streaming_response:
print(chunk)
若要傳遞 custom_inputs,請使用 extra_body 參數:
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_body={"custom_inputs": {"id": 5}},
)
要取得請求的追蹤 ID,請包含 x-mlflow-return-trace-id 標頭。 然後用 MLflow get_trace 取得完整的追蹤。
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_headers={"x-mlflow-return-trace-id": "true"},
)
trace_id = response.metadata["trace_id"]
trace = client.get_trace(trace_id)
REST API
將請求發送到應用程式 URL 的 /responses 路徑。 請求主體遵循 OpenAI 回應 API,因此你可以使用任何支援它的 HTTP 用戶端或工具。
curl --request POST \
--url <app-url>/responses \
--header 'Authorization: Bearer <OAuth token>' \
--header 'content-type: application/json' \
--data '{
"input": [{ "role": "user", "content": "hi" }],
"stream": true
}'
若要傳遞 custom_inputs,將它們加入請求主體:
curl --request POST \
--url <app-url>/responses \
--header 'Authorization: Bearer <OAuth token>' \
--header 'content-type: application/json' \
--data '{
"input": [{ "role": "user", "content": "hi" }],
"custom_inputs": { "id": 5 }
}'
要取得追蹤 ID,請包含 x-mlflow-return-trace-id: true 標頭。 回應體包含位於 metadata.trace_id 欄位中的 trace ID。 對於串流請求,追蹤 ID 會以獨立的 SSE 事件 data: {"trace_id": "tr-..."}(接近串流結尾處)形式出現在串流的末端。
查詢 Model Serving 上的舊版代理程式
使用此區塊用於 部署於模型服務端點的舊有代理。 你可以使用 Databricks OAuth 權杖或個人存取權杖來驗證身分。 若要將這些代理移至 Databricks 應用程式,請參見 「將代理從模型服務遷移至 Databricks 應用程式」。
Databricks OpenAI 客戶端
對於使用 ResponsesAgent 介面的代理,請呼叫 responses.create,並以端點名稱作為模型:
from databricks_openai import DatabricksOpenAI
input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"
client = DatabricksOpenAI()
# Non-streaming request. Calls predict.
response = client.responses.create(model=endpoint, input=input_msgs)
print(response)
# Streaming request. Calls predict_stream.
streaming_response = client.responses.create(model=endpoint, input=input_msgs, stream=True)
for chunk in streaming_response:
print(chunk)
對於使用傳統 ChatAgent 介面或 ChatModel 介面的代理程式,請使用聊天完成客戶端:
from databricks.sdk import WorkspaceClient
messages = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"
client = WorkspaceClient().serving_endpoints.get_open_ai_client()
response = client.chat.completions.create(model=endpoint, messages=messages)
print(response)
無論使用哪個用戶端,都可透過 extra_body 參數傳遞 custom_inputs 或 databricks_options。 例如,extra_body={"databricks_options": {"return_trace": True}} 會傳回包含回應的追蹤。
REST API
對於使用該ResponsesAgent介面的代理,請向 /serving-endpoints/responses 傳送請求,並以端點名稱作為模型:
curl --request POST \
--url https://<workspace-url>/serving-endpoints/responses \
--header 'Authorization: Bearer <token>' \
--header 'content-type: application/json' \
--data '{
"model": "<agent-endpoint-name>",
"input": [{ "role": "user", "content": "hi" }],
"stream": true
}'
對於使用 ChatAgent 或 input 介面的代理,請使用 messages 清單而非 ChatModel,將請求傳送至 /serving-endpoints/chat/completions。 若要傳遞 custom_inputs 或 databricks_options,將它們加入請求本文。 你也可以將請求傳送到端點的 /serving-endpoints/<agent-endpoint-name>/invocations URL。 請參見 「查詢端點後方的個別模型」。
AI 遊樂場
若想在 Model Serving 上與代理聊天而不寫程式碼,請開啟 AI Playground 並選擇代理的服務端點。 要將custom_inputs從 AI Playground 傳遞給代理人,請參閱在 AI Playground 和審核應用程式中提供custom_inputs。
使用 SQL 搭配
使用 ai_query 從 SQL 查詢 Model Serving 上的代理程式。 請參見 ai_query 函數 以了解語法與參數。
SELECT ai_query(
"<agent-endpoint-name>", question
) FROM (VALUES ('what is MLflow?'), ('how does MLflow work?')) AS t(question);