Microsoft Foundry Agent Testing and UI Experience

Umesh Dinkarrao Baradkar 0 Reputation points
2026-08-26T13:13:01.0966667+00:00

I'm currently working with ๐— ๐—ถ๐—ฐ๐—ฟ๐—ผ๐˜€๐—ผ๐—ณ๐˜ ๐—™๐—ผ๐˜‚๐—ป๐—ฑ๐—ฟ๐˜† ๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ and would love to learn from the community's real-world experience.

๐Ÿญ. ๐—Ÿ๐—ผ๐—ฎ๐—ฑ, ๐—ฆ๐˜๐—ฟ๐—ฒ๐˜€๐˜€ & ๐—ฃ๐—ฒ๐—ฟ๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐—ป๐—ฐ๐—ฒ ๐—ง๐—ฒ๐˜€๐˜๐—ถ๐—ป๐—ด

โ€ข How are you performing load, stress, and performance testing for Foundry Agents?

โ€ข Traces help with token usage, latency, and diagnostics, but how do you simulate ๐Ÿฎ๐Ÿฌ+ ๐—ฐ๐—ผ๐—ป๐—ฐ๐˜‚๐—ฟ๐—ฟ๐—ฒ๐—ป๐˜ ๐˜‚๐˜€๐—ฒ๐—ฟ๐˜€ and measure:

โ€ข Scalability

โ€ข Reliability

โ€ข Throughput

โ€ข Response times under load

โ€ข Are you using tools such as ๐—๐— ๐—ฒ๐˜๐—ฒ๐—ฟ, ๐—Ÿ๐—ผ๐—ฐ๐˜‚๐˜€๐˜, ๐—ธ๐Ÿฒ, ๐—”๐˜‡๐˜‚๐—ฟ๐—ฒ ๐—Ÿ๐—ผ๐—ฎ๐—ฑ ๐—ง๐—ฒ๐˜€๐˜๐—ถ๐—ป๐—ด, ๐—ผ๐—ฟ ๐—ฐ๐˜‚๐˜€๐˜๐—ผ๐—บ ๐—ฃ๐˜†๐˜๐—ต๐—ผ๐—ป ๐˜€๐—ฐ๐—ฟ๐—ถ๐—ฝ๐˜๐˜€?

๐Ÿฎ. ๐—”๐—ฑ๐—ฎ๐—ฝ๐˜๐—ถ๐˜ƒ๐—ฒ ๐—–๐—ฎ๐—ฟ๐—ฑ๐˜€ & ๐—•๐—ฒ๐˜๐˜๐—ฒ๐—ฟ ๐—จ๐—œ ๐—˜๐˜…๐—ฝ๐—ฒ๐—ฟ๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ

โ€ข Is there a recommended approach to use ๐—”๐—ฑ๐—ฎ๐—ฝ๐˜๐—ถ๐˜ƒ๐—ฒ ๐—–๐—ฎ๐—ฟ๐—ฑ๐˜€ or create a richer chatbot experience with Foundry Agents, similar to what we can achieve in Copilot Studio?

โ€ข Are you building custom web UIs (React, Teams apps, Bot Framework, etc.) on top of Foundry Agents?

Would appreciate any guidance, lessons learned, documentation, or examples from production implementations.

Foundry Agent Service
Foundry Agent Service

A fully managed platform in Microsoft Foundry for hosting, scaling, and securing AI agents built with any supported framework or model

0 comments No comments

2 answers

Sort by: Most helpful
  1. Manish Deshpande 8,055 Reputation points Microsoft External Staff Moderator
    2026-08-28T19:47:17.7+00:00

    Hello @Umesh Dinkarrao Baradkar

    You've framed it right โ€” Foundry traces give you token usage, latency and diagnostics, but they don't generate load. You need two separate planes: one to create load, and Foundry + App Insights to decompose it.

    1. Load and performance testing

    Check your subnet first. For network-isolated Agent Service, concurrent sessions map roughly 1:1 to usable subnet IPs: /27 โ†’ 32 IPs, ~27 usable, ~17 concurrent sessions; /26 โ†’ 64 IPs, ~59 usable, ~50 sessions (documented max). Your 20+ target won't fit on a /27. Microsoft recommends /24 for production. Verify with:

    az network vnet subnet show --ids <subnetId> --query "{prefix:addressPrefix, delegations:delegations[].serviceName}"
    

    Prefix should be /26+, in RFC 1918 space (public and CGNAT 100.64.0.0/10 aren't supported), delegated to Microsoft.App/environments. Azure reserves 5 IPs per subnet โ€” stay at or below 80% utilization.

    Agent type matters. Hosted agents consume subnet IPs (dedicated Micro VM + NIC each); prompt agents use a shared pool of ~10 IPs per project. Hosted limits: 100 active revisions, 1,000 total per agent, ~200 hosted agents per instance.

    Your real ceiling is usually model quota. Rate limiting is applied at the model deployment level, and one user turn fans out into several model calls, retrievals and tool calls โ€” so TPM/RPM typically binds before the agent runtime does. Sanity check: TPM รท (avg tokens/turn ร— turns per minute per user) should comfortably exceed 20. If not: backoff with jitter, or provisioned throughput.

    Generate load against your application endpoint, not the raw agent. Azure Load Testing supports JMeter and Locust only โ€” k6 must be self-hosted. Locust runs in LocalRunner mode on every engine; scale out by adding engines.

    Two signals that won't look like app errors: Azure Load Testing auto-stops when endpoints start throttling (that's a quota signal). And the portal doesn't expose IP utilization for delegated subnets โ€” treat data-proxy 5xx and session-creation 4xx as IP exhaustion, correlated against your VU ramp.

    Measure: p50/p95/p99 latency, conversations/sec, success vs. failure, 429/5xx + retries, token usage, per-tool-call latency, timeouts.

    2. Adaptive Cards / UI

    Keep the agent out of the presentation layer. Adaptive Cards aren't a native Foundry output format โ€” documented channels are Microsoft 365 Copilot and Teams; anything else needs custom integration. Two supported paths:

    • Publish from Foundry to Microsoft 365 โ€” auto-provisions Azure Bot Service + Entra ID. Fastest, minimal code.
    • Microsoft 365 Agents Toolkit proxy app โ€” when you need custom logic, SSO, or managed infra.

    Card interactions are handled by AgentApplication in the Microsoft 365 Agents SDK (Channel โ†’ Hosting layer โ†’ AgentApplication โ†’ your handlers), with microsoft-agents-hosting-teams for Teams-specific handlers. For a custom web UI, render your own components and stream โ€” samples/dotnet/azure-ai-streaming in github.com/microsoft/Agents is a good start.

    Worth weighing honestly: if rich UI is the priority and your orchestration is uncomplicated, Copilot Studio may be the better fit โ€” built-in orchestrator, reaches web/mobile/partner apps out of the box. Foundry wins when you need pro-code control and your own orchestrator.

    One caveat if you scale later: Foundry Agent Service has no native LLM load balancing โ€” you front it with APIM or a load balancer, and tool/grounding connections stay pinned to their original region even when traffic is routed elsewhere.

    Docs

    Thanks,
    Manish.

    Was this answer helpful?


  2. Allan Solomon Mejia 7,045 Reputation points
    2026-08-26T20:16:10.27+00:00

    Hello @Umesh Dinkarrao Baradkar

    For production testing, I would separate load generation from agent observability.

    1. Load, stress, and performance testing

    Microsoft Foundry tracing is useful for understanding individual agent executions: latency, exceptions, token consumption, tool calls, retrieval operations, etc. but, it isn't intended to simulate concurrent users. Microsoft documents Foundry tracing as an OpenTelemetry-based observability capability backed by Application Insights.

    For 20+ concurrent users, I would call the agent endpoint from an external load-testing tool. Azure Load Testing, k6, JMeter, Locust, or a custom async Python client can all work, depending on how your application exposes the agent.

    The useful pattern is:

    User's image

    During the test, measure at least:

    • End-to-end p50/p95/p99 response latency
    • Requests/conversations per second
    • Successful vs. failed requests
    • HTTP 429/5xx responses and retries
    • Model/token usage
    • Agent/tool-call latency
    • Dependency failures
    • Timeouts
    • Application-side CPU/memory if you host part of the solution yourself

    Foundry can then provide the per-request traces needed to determine where latency occurs. Microsoft recommends connecting Application Insights to the Foundry project; server-side tracing is then available for agents hosted in Foundry.

    Also remember that load testing an agent isn't exactly the same as testing a normal REST API. A single user request may result in several model calls, retrievals, and tool executions, so model quotas/rate limits and downstream dependencies can become the bottleneck before your application endpoint does.

    For some Foundry Agent Service configurations, concurrency is also a platform capacity consideration. For example, Microsoft documents regional concurrent-session capacity and subnet/IP requirements for network-isolated agent deployments.

    2. Adaptive Cards / richer UI

    I wouldn't make the Foundry Agent itself responsible for the presentation layer.

    A cleaner architecture is:

    User's image

    Foundry Agent Service provides the agent runtime, Responses API, models, tools, conversations, tracing, and related capabilities. Your application can provide the UX appropriate for the channel. (learn.microsoft.com)

    For a web application, React/Angular/etc. can render your own cards, tables, buttons, citations, progress indicators, streaming responses, and other components.

    For Teams/Microsoft 365 scenarios, Adaptive Cards are a good presentation option, but I would generate/render them in the application/channel integration layer, rather than treating Adaptive Cards as a native Foundry Agent UI capability.

    So for something comparable to the richer Copilot Studio experience, I'd typically use:

    Foundry Agent Service for intelligence/orchestration + your application or Teams layer for the UX + Application Insights/OpenTelemetry for observability.

    That also gives you much more control over production testing because you can load-test the same API/application path your real users will actually use.

    Sharing these references with you:

    Foundry Agent Service overview

    Agent tracing overview

    Set up tracing for Foundry agents

    Foundry Agent Service networking and concurrency

    Please "Accept the Answer" if this information helped you. This will help us and others in the community.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.