原生電腦應用

原生電腦使用允許回應模型請求有序的桌面或瀏覽器動作。 你的應用程式擁有執行環境,提供安全檢查以供核准,執行動作,擷取截圖,並將結果回傳給模型。

以下 Python 用戶端提供原生工具:

Client 工廠 Notes
OpenAIChatClient get_computer_tool() 使用 OpenAI 回應 API。
FoundryChatClient get_computer_tool() 需要 azure-ai-projects 2.3.0 或更新版本。

FoundryChatClient.get_computer_use_tool(...) 是一個獨立的預覽 API。 關於其設定選項,請參閱 Microsoft Foundry 模型提供者。

將工具新增至代理程式

建立提供者工具並交給客服人員:

from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

client = OpenAIChatClient()
agent = Agent(
    client=client,
    tools=[client.get_computer_tool()],
)

模型會回傳一個含有 type="computer_tool_call" 的 Content 項目。 該物品包括:

  • 一個項目 id 以及一個獨特的 call_id。
  • 一份有 actions 序的清單。
  • 可選的 pending_safety_checks。

未接聽的來電會出現在 AgentResponse.user_input_requests。 在應用程式執行任何動作前,先向使用者顯示動作和警告。 系統從不自動承認安全檢查。

回傳結果

當你的應用程式執行核准的動作後,請在工具訊息中回傳截圖。 以下片段假設 request 是電腦請求, screenshot_bytes 是由執行環境擷取的 PNG,且 confirmed_checks 僅包含使用者明確核准的檢查:

from agent_framework import Content, Message

result = Content.from_computer_tool_result(
    call_id=request.call_id,
    screenshot=Content.from_data(screenshot_bytes, "image/png"),
    acknowledged_safety_checks=confirmed_checks,
)

tool_message = Message(role="tool", contents=[result])
response = await agent.run(tool_message, session=session)

傳送結果時,請重複使用相同的 AgentSession。 你也可以提供帶有 Content.from_uri(...) 或 Content.from_hosted_file(...)的截圖。 OpenAI 和 Foundry 回應需要截圖,儘管共享結果類型允許提供者省略截圖。

ComputerSafetyCheck 類型以及 Content.from_computer_tool_call(...) 和 Content.from_computer_tool_result(...) 建構函式都是實驗性的代理程式框架 API。 呼叫和結果可透過 Content.to_dict() 持久化,並透過 Content.from_dict() 還原。

工作流程行為

當工作流程暫停進行原生電腦輸入時,代理框架會保留呼叫 ID、動作順序及跨檢查點的待執行安全檢查。 若本地函式呼叫在同一批次中完成,其結果會以原始呼叫順序回傳電腦結果。

如果代理待處理批次中的一個電腦請求被取消,該批次中剩餘的請求也會被取消。 已完成的結果仍留在終端輸出中,下一回合以全新代理人會話開始。